Thyroid ablation navigation method, model training method, electronic device, and storage medium
By automatically identifying standard sections in thyroid ultrasound video streams through a physical perception coding network and cross-attention mechanism, the problem of poor section consistency in existing technologies is solved, enabling accurate segmentation and safe distance calculation of thyroid nodules and surrounding structures, and supporting high-fidelity ablation path planning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-21
AI Technical Summary
In current thyroid nodule thermal ablation surgery, the selection of standard cut surfaces relies on the operator's experience, lacks spatiotemporal context stability analysis, and is difficult to handle ultrasound-specific artifacts, resulting in poor cut surface consistency and difficulty in accurately segmenting nodules and surrounding high-risk structures.
The image embedding features of the ultrasound video stream are extracted using a physical perception coding network. Combined with a standard anatomical memory bank and a cross-attention mechanism, standard sections are automatically identified through cross-attention maps and structural stability indices to generate multi-structure segmentation results, calculate safe distances, and plan ablation paths.
It achieves objective cross-section recognition without human intervention, improves the robustness of artifact region discrimination, ensures high-fidelity consistency of multi-structure segmentation, and supports more accurate pre-ablation planning.
Smart Images

Figure CN122423911A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to a thyroid ablation navigation method, a model training method, an electronic device, and a storage medium. Background Technology
[0002] Preoperative planning for thermal ablation of thyroid nodules relies on accurate identification of standard sections in ultrasound images and measurement of the spatial relationship between the nodule and surrounding important anatomical structures (such as the common carotid artery).
[0003] However, existing automated methods have the following shortcomings: the selection of standard sections largely depends on operator experience or simple single-frame classification, lacks spatiotemporal context stability analysis, resulting in poor section consistency; general segmentation models are difficult to handle ultrasound-specific artifacts such as acoustic shadowing and posterior enhancement, and are prone to boundary misjudgment; therefore, how to automatically and objectively screen out standard sections from thyroid ultrasound video streams and accurately segment nodules and surrounding high-risk structures is a technical problem that urgently needs to be solved in this field. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a thyroid ablation navigation method, a model training method, an electronic device, and a storage medium to at least partially improve the above-mentioned problems.
[0005] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, embodiments of the present invention provide a thyroid ablation navigation method, comprising: The acquired thyroid ultrasound video stream is input into a physical sensing coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream. The cross-attention between each of the image embedding features and all standard prototype key vectors in the preset standard anatomical memory is calculated to obtain the cross-attention map of each of the image embedding features; Based on the cross attention maps, a standard slice frame is determined from each of the image frames; The standard cross-section frame and the corresponding image embedding feature are input into the decoding network to obtain multi-structure segmentation results including the thyroid gland, nodules, trachea and common carotid artery; Based on the multi-structure segmentation results, the spatial distance between the nodule and the areas where the common carotid artery and recurrent laryngeal nerve run is located is calculated, and ablation path planning data is generated.
[0006] Optionally, before the step of inputting the acquired thyroid ultrasound video stream into a physical sensing coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream, the method further includes: Anisotropic diffusion filtering is applied to each frame of the acquired thyroid ultrasound video stream; the partial differential evolution equation of the anisotropic diffusion filtering is:
[0007]
[0008] in, Represents the image gradient. The diffusion coefficient function, The instantaneous gradient magnitude, This is the edge sensitivity threshold constant; Resample the filtered image to a preset resolution; The resampled image is subjected to channel duplication and Z-Score normalization to generate a normalized image frame.
[0009] Optionally, the physical sensing coding network includes a backbone coding network and an echo physical low-rank adapter. The step of inputting the acquired thyroid ultrasound video stream into the physical sensing coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream includes: For each image frame in the thyroid ultrasound video stream, the image frame is input into the backbone coding network to obtain the first coding feature; The first encoded feature is input into the echo physical low-rank adapter to obtain the second encoded feature; The first encoded feature and the second encoded feature are fused to obtain the image embedding feature.
[0010] Optionally, the echo physical low-rank adapter includes a dimensionality reduction layer, an asymmetric physical convolutional layer, and a dimensionality increase layer. The asymmetric physical convolutional layer includes a vertical convolutional kernel and a horizontal convolutional kernel. The step of inputting the first encoded feature into the echo physical low-rank adapter to obtain the second encoded feature includes: The first encoded feature is compressed into a low-dimensional space and rearranged into a spatial feature map through the dimensionality reduction layer; The spatial feature maps are respectively input into the vertical and horizontal convolution kernels of the asymmetric physical convolutional layer to obtain vertical and horizontal feature maps respectively; After fusing and activating the vertical and horizontal feature maps, the results are input into the up-dimensional layer to obtain the second encoded feature.
[0011] Optionally, determining a standard slice frame from each of the image frames based on each of the cross-attention maps includes: Calculate the Shannon entropy of each of the aforementioned cross-attention maps; Based on the image embedding features, calculate the model prediction confidence of each image frame as a standard cross-section; Based on the preset structural stability index calculation formula, the corresponding structural stability index is calculated according to the Shannon entropy and the prediction confidence of each model. The structural stability indices are arranged into a structural stability index curve according to the temporal order of the thyroid ultrasound video stream. Image frames with the largest structural stability index that are greater than a preset threshold, located at a local maximum, and have the largest structural stability index are selected from the structural stability index curves as standard cross-sectional frames.
[0012] Optionally, the step of embedding the standard cross-section frame and the corresponding image into the feature input decoding network to obtain multi-structure segmentation results including the thyroid gland, nodules, trachea, and common carotid artery includes: The standard cross-section frame is input into the target detection network to obtain the overall bounding box of the thyroid region. A hypoechoic circular region is located at the outer edge of the overall thyroid region boundary box, and a common carotid artery indicator box is generated. A strongly echogenic arc-shaped region is located below and inside the overall boundary box of the thyroid region to generate a tracheal prompt box; Locate the echogenic heterogeneous region within the overall thyroid region bounding box and generate a nodule prompt box; The common carotid artery prompt box, the trachea prompt box, and the nodule prompt box are encoded as sparse prompt embeddings; The sparse cue is embedded into the image embedding feature input decoding network corresponding to the standard section frame to obtain multi-structure segmentation results including thyroid, nodules, trachea and common carotid artery.
[0013] Optionally, based on the multi-structure segmentation results, calculating the spatial distance between the nodule and the areas along the common carotid artery and recurrent laryngeal nerve to generate ablation path planning data includes: Extract the nodule contour, common carotid artery contour, tracheal contour, and posterior thyroid capsule from the multi-structure segmentation results; Calculate the shortest Euclidean distance from all pixels on the nodule contour to the common carotid artery contour; When any of the shortest Euclidean distances is less than a preset safety threshold, a recommended injection path for the liquid isolation band is generated based on the arc length angle of the nodule and the common carotid artery. Based on the positional relationship between the tracheal contour and the posterior capsule of the thyroid gland, the course of the recurrent laryngeal nerve is determined, and the projected distance from the center of the nodule contour to the course of the recurrent laryngeal nerve is calculated. The shortest Euclidean distance, the projected distance, and the recommended injection path are output as ablation path planning data.
[0014] Secondly, embodiments of the present invention provide a method for training a thyroid ablation navigation model, wherein the thyroid ablation navigation model includes a physical sensing coding network and a decoding network, the physical sensing coding network includes a trunk coding network and an echo physical low-rank adapter, and the method includes: Acquire labeled thyroid ultrasound video data; the labeling includes segmentation masks of the thyroid gland, nodules, trachea, and common carotid artery in each frame of the image; The thyroid ultrasound video data is input into the physical sensing coding network to obtain the image embedding features of each image frame; Based on the segmentation mask in the annotation, an anatomical cue box is generated and encoded as a sparse cue embedding; The image embedding features and the sparse cue embedding are input together into the decoding network to obtain multi-structure segmentation results; Based on the total loss function, the total loss information is calculated according to the multi-structure segmentation results and the segmentation mask in the annotation; the total loss function is:
[0015]
[0016] in, This is a topological mutual exclusion term used to penalize pixel overlap in anatomically mutually exclusive structures. , , These represent the predicted probability diagrams for the common carotid artery, thyroid gland, and trachea, respectively. Marginal probability plot representing the thyroid gland, For pixel index; Freeze the pre-trained parameters of the backbone coding network, and update the parameters of the echo physical low-rank adapter and the decoding network through backpropagation based on the total loss information.
[0017] Thirdly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the method described in any of the above-mentioned embodiments.
[0018] Fourthly, embodiments of the present invention provide a storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any of the preceding claims.
[0019] This invention provides a thyroid ablation navigation method, model training method, electronic device, and storage medium. By explicitly modeling the inherent physical characteristics of ultrasound imaging through a physical perception coding network, the robustness of the model in identifying artifact regions is improved. Simultaneously, by combining a standard anatomical memory bank with a cross-attention mechanism and using attention entropy as a quantitative representation of structural stability, the system can automatically identify the most complete and clearest dynamic optimal frame of the anatomical structure in the video stream. This provides a high-fidelity and highly consistent image benchmark for subsequent multi-structure segmentation and safe distance calculation, ultimately supporting the generation of more accurate, interpretable, and clinically feasible pre-ablation planning.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A schematic structural block diagram of an electronic device provided in an embodiment of the present invention; Figure 2 This is a schematic flowchart of a thyroid ablation navigation method provided in an embodiment of the present invention; Figure 3 A flowchart illustrating step S210 provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of data processing in a physical sensing coding network provided in an embodiment of the present invention; Figure 5 A flowchart illustrating step S230 provided in an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the selection of a standard cross-sectional frame from a structural stability index curve, provided as an embodiment of the present invention. Figure 7 A flowchart illustrating step S240 provided in an embodiment of the present invention; Figure 8 A flowchart illustrating step S250 provided in an embodiment of the present invention; Figure 9 A schematic diagram illustrating a quantitative analysis of safety margin provided in an embodiment of the present invention; Figure 10This is a flowchart illustrating a thyroid ablation navigation model training method provided in an embodiment of the present invention. Figure 11 This is a schematic diagram of a topologically exclusive partitioning method provided in an embodiment of the present invention.
[0023] Icons: 100 - Electronic device; 101 - Memory; 102 - Communication interface; 103 - Processor; 104 - Communication bus. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0026] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this invention, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0027] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0028] To implement the process steps and functions of this invention, please refer to [link / reference]. Figure 1 , Figure 1This is a schematic structural block diagram of an electronic device 100 provided in an embodiment of the present invention. The electronic device 100 includes a memory 101 and a processor 103, which are electrically connected directly or indirectly to each other to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 can be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101, thereby performing various functional applications and data processing.
[0029] Electronic device 100 can be, but is not limited to, a personal computer (PC), a server, a distributed computer, etc. It is understood that electronic device 100 is not limited to a physical server, but can also be a virtual machine on a physical server, a virtual machine built on a cloud platform, or any other computer that can provide the same functionality as the server or virtual machine. The operating system of electronic device 100 can be, but is not limited to, Windows, Linux, etc.
[0030] The memory 101 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0031] The communication connection between the electronic device 100 and external devices is achieved through at least one communication interface 102 (which can be wired or wireless).
[0032] Processor 103 may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of this embodiment can be completed by integrated logic circuits in the hardware of processor 103 or by instructions in software form. Processor 103 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0033] Understandable. Figure 1 The structure shown is for illustrative purposes only; the electronic device 100 may also include components that are more advanced than those shown. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.
[0034] The thyroid ablation navigation method provided in the embodiments of the present invention will be described by way of example below. See also Figure 2 The subject executing this method can be one of the above. Figure 1 The electronic device 100 shown, the method includes as follows Figure 2 The following steps are described: S210: Input the acquired thyroid ultrasound video stream into the physical perception coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream.
[0035] S220: Calculate the cross-attention between each image embedding feature and all standard prototype key vectors in the preset standard anatomical memory to obtain the cross-attention map of each image embedding feature.
[0036] S230: Based on each cross-attention map, determine a standard cross-section frame from each image frame.
[0037] S240: The standard cross-section frame and the corresponding image are embedded into the feature input decoding network to obtain multi-structure segmentation results including the thyroid gland, nodules, trachea and common carotid artery.
[0038] S250: Based on the multi-structure segmentation results, calculate the spatial distance between the nodule and the areas where the common carotid artery and recurrent laryngeal nerve run, and generate ablation path planning data.
[0039] This method extracts image embedding features that integrate ultrasound physical properties through a physical perception coding network, enabling the model to have inherent robustness to artifacts such as tracheal acoustic shadowing and vascular enhancement. By calculating the cross-attention between the image embeddings of each frame and the standard anatomical memory, and determining the standard section frame accordingly, it achieves objective section recognition without manual intervention or reliance on a single-frame classifier. Based on the standard section frame and the corresponding image embedding, multi-structure segmentation results are generated to ensure that the spatial relationships of the thyroid gland, nodules, trachea, and common carotid artery conform to anatomical logic. Finally, based on the segmentation results, the spatial distance between the nodules and the common carotid artery and the recurrent laryngeal nerve course is calculated, and quantitative ablation path data that can be used to guide puncture path planning and safety boundary setting is directly output.
[0040] Before performing step S210, embodiments of the present invention may further perform standardized preprocessing on the acquired thyroid ultrasound video stream to improve image quality and enhance the stability of subsequent physical perception coding. This method further includes: Anisotropic diffusion filtering is applied to each frame of the acquired thyroid ultrasound video stream; the filtered image is resampled to a preset resolution; channel duplication and Z-score normalization are performed on the resampled image to generate a normalized image frame.
[0041] The thyroid ultrasound video stream is a continuous B-mode ultrasound video stream of the thyroid region acquired in real time through a standard ultrasound equipment interface. The video stream is organized in time series format and is denoted as... ,in The length of the time window (e.g., 90 frames). This is the original image resolution.
[0042] This invention does not employ Gaussian filtering, which blurs tissue edges, but instead constructs an anisotropic diffusion model for each frame of the image. Preprocessing is performed; this filtering process is driven by partial differential equations, and the evolution equation is as follows:
[0043]
[0044] in, Represents the image gradient. The diffusion coefficient function, The instantaneous gradient magnitude, This is the edge sensitivity threshold constant. It can be set to 25. The number of iterations for diffusion can be preset for each frame of the image.
[0045] The physical meaning of this formula is: in flat areas of the image ( diffusion coefficient Approaching 1, strong smoothing is performed to remove noise; while in peripheral areas such as the thyroid capsule or blood vessel walls ( ), The noise level rapidly approaches zero, halting the diffusion process. This step mathematically guarantees that high-frequency anatomical boundaries are preserved while suppressing noise.
[0046] The image is then resampled to a preset resolution (e.g., ...). Then, channel copying and Z-score normalization are performed to generate a normalized input tensor. This is to adapt to the input distribution requirements of subsequent deep networks.
[0047] Next, step S210 is executed, inputting each frame of image into the physical perception coding network. This network can use a pre-trained Visual Transformer (ViT) as its backbone architecture, embedding an echo-physical low-rank adapter within its self-attention module. This adapter is trained with ultrasound physics priors and has the ability to recognize ultrasound-specific acoustic phenomena. The network output is a high-dimensional image embedding feature corresponding to each frame of image. This image embedding feature explicitly encodes the acoustic physical properties of the ultrasound image, which is different from the output of a general visual encoder that treats ultrasound images as ordinary grayscale images.
[0048] In one alternative implementation, the physical sensing coding network includes a backbone coding network and an echo physical low-rank adapter, with parameters... Figure 3 S210 may include the following sub-steps: S211: For each image frame in the thyroid ultrasound video stream, the image frame is input into the backbone coding network to obtain the first coding feature.
[0049] S212: Input the first coding feature into the echo physical low-rank adapter to obtain the second coding feature.
[0050] The echo physical low-rank adapter includes a dimension reduction layer, an asymmetric physical convolutional layer, and a dimension increase layer. The asymmetric physical convolutional layer includes a vertical convolutional kernel and a horizontal convolutional kernel. S212 may include the following sub-steps: S2121: The first encoded features are compressed into a low-dimensional space and rearranged into a spatial feature map through a dimensionality reduction layer.
[0051] S2122: Input the spatial feature maps into the vertical and horizontal convolution kernels of the asymmetric physical convolutional layer respectively to obtain the vertical and horizontal feature maps respectively.
[0052] S2123: After fusing and activating the vertical and horizontal feature maps, the results are input into the up-dimensional layer to obtain the second encoded feature.
[0053] S213: Fuse the first coding feature with the second coding feature to obtain the image embedding feature.
[0054] To enable a general visual model to understand the physical propagation characteristics of ultrasound, this embodiment of the invention embeds a unique echo-physical low-rank adapter in parallel within each Transformer Block of the frozen-parameter backbone coding network. Unlike conventional LoRA, which only performs a simple linear mapping, this echo-physical low-rank adapter contains two parallel asymmetric physical convolutional paths.
[0055] For example, see Figure 4 Each standardized image frame from the thyroid ultrasound video stream is input into the backbone coding network, and the output is the first coding feature. ,in For the number of tokens, The feature dimension is then used for dimensionality reduction. The features are compressed into a low-dimensional space and rearranged into a spatial feature map, followed by vertical sound and shadow perception and horizontal layering perception: Using a core size of A high aspect ratio convolution kernel performs a vertical sliding operation on the spatial feature map, which can be mathematically expressed as:
[0056] The convolution kernel design has a clear physical directionality: ultrasound waves propagate along the vertical scan line and will produce posterior signal attenuation (sound shadow) when they encounter strong reflectors (such as tracheal cartilage rings or calcifications). The convolutional kernel can maximize the response to this vertical strip-shaped feature, thereby identifying signal-deficient regions of "non-anatomical structures".
[0057] Using a core size of The flattened convolutional kernel processes the spatial feature map horizontally, that is:
[0058] This approach is designed to respond to the horizontal, layered continuity of the neck anatomy as presented in cross-section.
[0059] Finally, the vertical feature map is... and horizontal feature map Fusion, GeLU activation and upscaling matrix After restoring the dimensions, the residual is injected into the self-attention output of the backbone coding network. This process is represented by the following formula:
[0060] in, This is the first coding feature. To freeze the weight, For a dimension reduction matrix, For an upgraded matrix, This is the activation function.
[0061] This design employs a special asymmetric convolution kernel. Vertical convolution kernels are specifically designed to respond to vertical acoustic shadowing and reverberation artifacts in ultrasound images, while horizontal convolution kernels respond to the boundaries of layered tissues. This design enables the model to "understand" ultrasound physics phenomena without disrupting its pre-trained knowledge.
[0062] Next, step S220 is executed to calculate the cross-attention of each image embedding feature. The preset standard anatomical memory bank is the key vector and value vector extracted from high-quality images that cover the typical standard cross-sectional anatomical configuration of the thyroid gland, which have been confirmed by domain experts, after being encoded by the same physical perception coding network.
[0063] For each image embedding feature, the image embedding feature is linearly transformed into a query vector, and then a scaled dot product cross-attention calculation is performed with all standard prototype key vectors in the memory to obtain a normalized attention weight matrix, which is the cross-attention map. This represents the semantic matching distribution between the current frame image and the standard anatomical prototype.
[0064] After obtaining the cross-attention maps of the embedded features of each image, step S230 is executed to perform global statistical analysis on the cross-attention maps generated for each frame, calculating their Structural Stability Index (SSI). This index reflects the clarity and consistency of the anatomical structure representation in the current frame: a higher SSI value indicates that key structures such as the thyroid gland, trachea, and common carotid artery in the image have highly matched standard prototypes in the memory bank, and that the attention response is concentrated and has low dispersion, corresponding to the clinically significant standard section that is "most stable and displays the most complete image." The SSI values of all frames are arranged in chronological order to form an SSI time curve; the single frame with the highest SSI value that meets the preset stability threshold is selected as the final locked standard section frame. In one optional implementation, see [link to implementation details]. Figure 5 S230 may include the following sub-steps: S231: Calculate the Shannon entropy of each cross attention map.
[0065] S232: Based on the embedding features of each image, calculate the model prediction confidence of each image frame as a standard cross-section.
[0066] S233: Based on the preset structural stability index calculation formula, calculate the corresponding structural stability index according to each Shannon entropy and each model prediction confidence.
[0067] S234: Arrange the structural stability indices into a structural stability index curve according to the time sequence of the thyroid ultrasound video stream.
[0068] S235: Select the image frame with the largest structural stability index from the structural stability index curve that is greater than the preset threshold, is at a local maximum point, and has the largest structural stability index as the standard cross-section frame.
[0069] First calculate each cross-attention map The Shannon entropy is calculated using the following formula:
[0070] in, It is a very small positive number, so avoid taking the logarithm of zero.
[0071] Next, for each image frame, a global average pooling operation is performed on its corresponding image embedding features to compress the spatial dimension feature map into a one-dimensional global feature vector. Then, the global feature vector is input into a multilayer perceptron consisting of two fully connected layers for feature dimensionality reduction and nonlinear mapping. Finally, the mapped features are input into a sigmoid activation function, which outputs a continuous probability value in the range [0,1]. This probability value is the model prediction confidence. The closer this value is to 1, the higher the probability that the current image frame belongs to a high-quality, non-deformed standard thyroid anatomical section.
[0072] The structural stability index is then calculated using the following formula:
[0073] in, For the first The rich entropy of frames This represents the maximum value of Shannon entropy. , It is the total number of positions in the cross-attention map. To predict the confidence level of the model, These are the weighting coefficients.
[0074] Then see Figure 6 All frame SSI values are arranged in chronological order to form a structural stability index curve. A unique frame that meets the following three conditions is selected from this curve as the standard cross-section frame: Threshold condition: Greater than a preset threshold (e.g., 0.85); Local maximum condition: This frame is a local maximum point, that is... and ; Maximum value condition: Select from all candidate frames that satisfy the threshold condition and the local maximum condition. The frame with the highest value.
[0075] The frame that meets the above conditions is the dynamic optimal section with the most complete and stable anatomical structure, and is used as the standard section frame for subsequent segmentation and planning.
[0076] After obtaining the standard cross-sectional frame, step S240 is executed, simultaneously inputting the standard cross-sectional frame and its image embedding features into the mask decoding network. This decoding network can employ a cue-driven architecture, receiving structural cue information automatically generated from anatomical topology priors (including spatial location and morphological constraints for the thyroid, nodules, trachea, and common carotid artery), and performing cross-modal fusion of cue embedding and image embedding. This forces the predicted masks for different anatomical structures to maintain anatomically reasonable relative positions in space, avoiding segmentation errors that violate medical common sense, such as blood vessels penetrating the thyroid parenchyma or nodules breaking through the thyroid capsule. The network ultimately outputs a set of pixel-level segmentation masks, accurately identifying the overall outline of the thyroid, the internal nodule region, the projection area of the tracheal cartilage rings, and the cross-section of the common carotid artery. In one optional implementation, see [link to implementation details]. Figure 7 S240 may include the following steps: S241: Input the standard cross-section frame into the target detection network to obtain the overall bounding box of the thyroid region.
[0077] S242: Locate a hypoechoic circular area at the outer edge of the overall thyroid region bounding box and generate a carotid artery prompt box.
[0078] S243: Locate a strongly echogenic arc-shaped region inside and below the overall thyroid region boundary box, and generate a tracheal cue box.
[0079] S244: Locate the echogenic heterogeneous region within the overall thyroid region bounding box and generate a nodule prompt box.
[0080] S245: Encode the common carotid artery cue box, trachea cue box, and nodule cue box as sparse cue embedding.
[0081] S246: The sparse cue is embedded into the image embedding feature corresponding to the standard slice frame and input into the decoding network to obtain multi-structure segmentation results including the thyroid gland, nodules, trachea and common carotid artery.
[0082] Input the standard slice frames locked by S234 into a lightweight object detection network (e.g., YOLOv8-Nano) to output the bounding box of the entire thyroid region. The bounding box covers the complete outline of the thyroid gland and serves as a unified spatial reference for subsequent anatomical prompts.
[0083] exist A region of interest is defined on the left or right side of the outer edge region (depending on the probe scanning orientation). This region is then segmented using a dark pixel threshold (e.g., the threshold is set to 0.35 times the global grayscale mean). Combined with Hough circle transform to detect circular structures, the center of the circular anechoic region is located. With radius Generate a carotid artery prompt box. In this embodiment of the invention, only the common carotid artery is identified as a blood vessel. The common carotid artery is closer to the lateral lobe of the thyroid gland than the internal jugular vein. When ablating the nodule, it is only necessary to ensure a safe distance between the nodule and the common carotid artery.
[0084] Because the trachea is located below and inside the thyroid gland and presents as a strongly echogenic arc, in Extract from the inner lower region Find the maximum gradient response arc segment in the gradient amplitude graph below the inner side and generate a tracheal prompt box. .
[0085] exist Internally, the texture heterogeneity index is calculated to locate the region where the echo texture differs most from the surrounding material, generating a nodule prompt box. .
[0086] The location information (coordinates, size) of each cue box is input into the encoder and converted into sparse embedding vectors (e.g., through positional encoding + MLP mapping). These embeddings have the same dimensionality as image embedding features. This enables the decoding network to "focus" on the anatomical region represented by the cue box.
[0087] The sparse cue embeddings and the image embedding features corresponding to the standard cross-section frames are input into the decoding network. The network fuses the cue and image features based on the cross-attention mechanism and outputs a four-channel segmentation mask, which accurately identifies the whole thyroid gland, internal nodules, tracheal cartilage ring projection and carotid artery cross section respectively.
[0088] Based on the segmentation mask obtained from S240, the closed contours of the nodule mask, the common carotid artery mask, and the tracheal and posterior thyroid capsule masks are extracted. First, the shortest Euclidean distance between the contours of the common carotid artery mask and all points on the nodule mask is calculated to obtain the minimum safe distance. Second, based on the spatial geometry of the tracheal contour and the posterior thyroid capsule, the high-risk course region of the recurrent laryngeal nerve in the ultrasound image plane is estimated, and the vertical projection distance from the center point of the nodule contour to this region is calculated. Finally, combining the minimum safe distance and the projection distance with a preset clinical safety threshold, ablation path planning data containing safe distance values and visual guide lines is generated to guide the operator in formulating puncture approach, energy placement, and fluid isolation band injection strategies. In one optional implementation, see [link to implementation details]. Figure 8 S250 may include the following steps: S251: Extract nodule contours, common carotid artery contours, tracheal contours, and posterior thyroid capsule from the multi-structure segmentation results.
[0089] S252: Calculate the shortest Euclidean distance from all pixels on the nodule contour to the common carotid artery contour.
[0090] S253: When any shortest Euclidean distance is less than a preset safety threshold, a recommended injection path for the liquid isolation zone is generated based on the arc length angle of the nodule and the common carotid artery.
[0091] S254: Based on the positional relationship between the tracheal outline and the posterior capsule of the thyroid gland, determine the course of the recurrent laryngeal nerve and calculate the projection distance from the center of the nodule outline to the course of the recurrent laryngeal nerve.
[0092] S255: Outputs the shortest Euclidean distance, projection distance, and recommended injection path as ablation path planning data.
[0093] Binary contours are extracted from the multi-structure segmentation mask output by S240: Nodule outline: composed of edge pixels of the nodule mask; Carotid artery outline: composed of edge pixels of the carotid artery mask; Tracheal outline: composed of edge pixels of the tracheal mask; Retrothyroid capsule: defined as the curve formed by the lowest consecutive non-zero pixel row of the thyroid mask in the vertical direction (y-axis).
[0094] See Figure 9 For each pixel on the nodule contour, calculate its Euclidean distance to the carotid artery contour, and take the minimum value of all Euclidean distances as the shortest Euclidean distance between the nodule and the carotid artery.
[0095] If the shortest Euclidean distance is less than a preset safety threshold (e.g., 2 mm, calibrated according to the physical resolution of the ultrasound image), then the fitting arc segment between the nodule contour and the common carotid artery contour with a distance less than the preset safety threshold is identified, the central angle of the arc segment on the nodule contour is calculated, and a recommended injection path for the liquid isolation zone is generated along the normal direction of the midpoint of the arc segment, 2 mm away from the outer edge of the nodule contour.
[0096] Based on the spatial relationship between the tracheal outline and the posterior capsule of the thyroid gland, the course of the recurrent laryngeal nerve is determined: This region is defined as a band-shaped area extending a predetermined distance (e.g., 3 mm) towards the posterior capsule of the thyroid gland, with the lateral edge of the trachea as the baseline. Its upper boundary is the extension line of the lower edge of the tracheal cartilage ring, and its lower boundary is the line connecting the lower pole of the thyroid gland. The vertical projection distance from the geometric center point of the nodule outline to this region is calculated.
[0097] Finally, the shortest Euclidean distance, projection distance, and recommended injection path are output as ablation path planning data for clinical workstation visualization and surgical navigation system.
[0098] Furthermore, this embodiment of the invention also provides a method for training a thyroid ablation navigation model. The thyroid ablation navigation model includes a physical sensing encoding network and a decoding network. The physical sensing encoding network includes a trunk encoding network and an echo physical low-rank adapter. See [link to documentation]. Figure 10 The method includes: S310: Acquire labeled thyroid ultrasound video data; the labels include segmentation masks for the thyroid gland, nodules, trachea, and common carotid artery in each frame.
[0099] S320: Input thyroid ultrasound video data into a physical perception coding network to obtain image embedding features for each image frame.
[0100] S330: Generate anatomical cue boxes based on the segmentation mask in the annotation and encode them as sparse cue embeddings.
[0101] S340: Image embedding features and sparse cue embeddings are input into the decoding network to obtain multi-structure segmentation results.
[0102] S350: Based on the total loss function, the total loss information is calculated according to the multi-structure segmentation results and the segmentation mask in the annotation; the total loss function is:
[0103]
[0104] in, This is a topological mutual exclusion term used to penalize pixel overlap in anatomically mutually exclusive structures. , , These represent the predicted probability diagrams for the common carotid artery, thyroid gland, and trachea, respectively. Marginal probability plot representing the thyroid gland, For pixel index.
[0105] S360: Freeze the pre-trained parameters of the backbone coding network and update the parameters of the echo physical low-rank adapter and decoding network through backpropagation based on the total loss information.
[0106] During model training, the decoder's loss function includes a topological mutual exclusion term. .like Figure 11As shown, this loss term applies a "repulsive force" to anatomical structures such as "thyroid-blood vessels" and "thyroid-trachea," forcing their segmentation results to be non-overlapping or to satisfy specific positional relationships (e.g., blood vessels must not invade the interior of the thyroid gland, and the trachea must be located on the outer edge of the thyroid gland). Through backpropagation, the decoder learns to generate segmentation masks that conform to anatomical topological logic.
[0107] In summary, the thyroid ablation navigation method, model training method, electronic device, and storage medium provided by this invention, through the construction of a standard anatomical memory bank, cross-attention map, and Shannon entropy-guided structural stability index mechanism, model the selection of standard sections in ultrasound videos as a dynamic optimal frame search problem under the dual constraints of spatiotemporal consistency and anatomical integrity, significantly improving the robustness and repeatability of section recognition. By explicitly modeling the physical degradation process of ultrasound imaging through a physical perception coding network, it effectively suppresses boundary misjudgments caused by artifacts such as acoustic shadowing and posterior enhancement. Combined with sparse cue embedding based on anatomical priors and a topological mutual exclusion loss function, it ensures that the segmentation results strictly conform to the anatomical logic of the neck. In ultrasound navigation, it achieves dual-dimensional quantitative evaluation of the shortest Euclidean distance between the nodule and the common carotid artery and the projected distance of the recurrent laryngeal nerve course area, and automatically generates recommended injection paths for the fluid isolation band accordingly, directly outputting structured planning data that can be called by the surgical navigation system, shifting preoperative safety assessment from experience-based judgment to evidence-driven assessment.
[0108] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0109] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0110] If the functionality is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0111] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
[0112] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A thyroid ablation navigation method, characterized in that, include: The acquired thyroid ultrasound video stream is input into a physical sensing coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream. The cross-attention between each of the image embedding features and all standard prototype key vectors in the preset standard anatomical memory is calculated to obtain the cross-attention map of each of the image embedding features; Based on the cross attention maps, a standard slice frame is determined from each of the image frames; The standard cross-section frame and the corresponding image embedding feature are input into the decoding network to obtain multi-structure segmentation results including the thyroid gland, nodules, trachea and common carotid artery; Based on the multi-structure segmentation results, the spatial distance between the nodule and the areas where the common carotid artery and recurrent laryngeal nerve run is located is calculated, and ablation path planning data is generated.
2. The method according to claim 1, characterized in that, Before the step of inputting the acquired thyroid ultrasound video stream into a physical perception coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream, the method further includes: Anisotropic diffusion filtering is applied to each frame of the acquired thyroid ultrasound video stream; the partial differential evolution equation of the anisotropic diffusion filtering is: in, Represents the image gradient. The diffusion coefficient function, The instantaneous gradient magnitude, This is the edge sensitivity threshold constant; Resample the filtered image to a preset resolution; The resampled image is subjected to channel duplication and Z-Score normalization to generate a normalized image frame.
3. The method according to claim 1, characterized in that, The physical sensing coding network includes a backbone coding network and an echo physical low-rank adapter. The process of inputting the acquired thyroid ultrasound video stream into the physical sensing coding network to obtain the image embedding features of each image frame in the thyroid ultrasound video stream includes: For each image frame in the thyroid ultrasound video stream, the image frame is input into the backbone coding network to obtain the first coding feature; The first encoded feature is input into the echo physical low-rank adapter to obtain the second encoded feature; The first encoded feature and the second encoded feature are fused to obtain the image embedding feature.
4. The method according to claim 3, characterized in that, The echo physical low-rank adapter includes a dimensionality reduction layer, an asymmetric physical convolutional layer, and a dimensionality increase layer. The asymmetric physical convolutional layer includes a vertical convolutional kernel and a horizontal convolutional kernel. The step of inputting the first encoded feature into the echo physical low-rank adapter to obtain the second encoded feature includes: The first encoded feature is compressed into a low-dimensional space and rearranged into a spatial feature map through the dimensionality reduction layer; The spatial feature maps are respectively input into the vertical and horizontal convolution kernels of the asymmetric physical convolutional layer to obtain vertical and horizontal feature maps respectively; After fusing and activating the vertical and horizontal feature maps, the results are input into the up-dimensional layer to obtain the second encoded feature.
5. The method according to claim 1, characterized in that, The step of determining a standard slice frame from each of the image frames based on each of the cross-attention maps includes: Calculate the Shannon entropy of each of the aforementioned cross-attention maps; Based on the image embedding features, calculate the model prediction confidence of each image frame as a standard cross-section; Based on the preset structural stability index calculation formula, the corresponding structural stability index is calculated according to the Shannon entropy and the prediction confidence of each model. The structural stability indices are arranged into a structural stability index curve according to the temporal order of the thyroid ultrasound video stream. Image frames with the largest structural stability index that are greater than a preset threshold, located at a local maximum, and have the largest structural stability index are selected from the structural stability index curves as standard cross-sectional frames.
6. The method according to claim 1, characterized in that, The step of embedding the standard cross-section frame and the corresponding image into the feature input decoding network to obtain multi-structure segmentation results including the thyroid gland, nodules, trachea, and common carotid artery includes: The standard cross-section frame is input into the target detection network to obtain the overall bounding box of the thyroid region. A hypoechoic circular region is located at the outer edge of the overall thyroid region boundary box, and a common carotid artery indicator box is generated. A strongly echogenic arc-shaped region is located below and inside the overall boundary box of the thyroid region to generate a tracheal prompt box; Locate the echogenic heterogeneous region within the overall thyroid region bounding box and generate a nodule prompt box; The common carotid artery prompt box, the trachea prompt box, and the nodule prompt box are encoded as sparse prompt embeddings; The sparse cue is embedded into the image embedding feature input decoding network corresponding to the standard section frame to obtain multi-structure segmentation results including thyroid, nodules, trachea and common carotid artery.
7. The method according to claim 1, characterized in that, Based on the multi-structure segmentation results, the spatial distance between the nodule and the areas along the common carotid artery and recurrent laryngeal nerve is calculated, and ablation path planning data is generated, including: Extract the nodule contour, common carotid artery contour, tracheal contour, and posterior thyroid capsule from the multi-structure segmentation results; Calculate the shortest Euclidean distance from all pixels on the nodule contour to the common carotid artery contour; When any of the shortest Euclidean distances is less than a preset safety threshold, a recommended injection path for the liquid isolation band is generated based on the arc length angle of the nodule and the common carotid artery. Based on the positional relationship between the tracheal contour and the posterior capsule of the thyroid gland, the course of the recurrent laryngeal nerve is determined, and the projected distance from the center of the nodule contour to the course of the recurrent laryngeal nerve is calculated. The shortest Euclidean distance, the projected distance, and the recommended injection path are output as ablation path planning data.
8. A method for training a thyroid ablation navigation model, characterized in that, The thyroid ablation navigation model includes a physical sensing coding network and a decoding network. The physical sensing coding network includes a trunk coding network and an echo physical low-rank adapter. The method includes: Acquire labeled thyroid ultrasound video data; the labeling includes segmentation masks of the thyroid gland, nodules, trachea, and common carotid artery in each frame of the image; The thyroid ultrasound video data is input into the physical sensing coding network to obtain the image embedding features of each image frame; Based on the segmentation mask in the annotation, an anatomical cue box is generated and encoded as a sparse cue embedding; The image embedding features and the sparse cue embedding are input together into the decoding network to obtain multi-structure segmentation results; Based on the total loss function, the total loss information is calculated according to the multi-structure segmentation results and the segmentation mask in the annotation; the total loss function is: in, This is a topological mutual exclusion term used to penalize pixel overlap in anatomically mutually exclusive structures. , , These represent the predicted probability diagrams for the common carotid artery, thyroid gland, and trachea, respectively. Marginal probability plot representing the thyroid gland, For pixel index; Freeze the pre-trained parameters of the backbone coding network, and update the parameters of the echo physical low-rank adapter and the decoding network through backpropagation based on the total loss information.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method described in any one of claims 1 to 8.