A Medical Ultrasound Image Segmentation Method Based on Geometric Guidance and Concept Awareness Fusion

By combining geometry-guided text generation and concept-aware cross-modal fusion modules with medical knowledge graph modeling, the problems of difficulty in producing high-quality text descriptions and insufficient cross-modal fusion are solved, thereby improving the accuracy and efficiency of medical ultrasound image segmentation.

CN122023813BActive Publication Date: 2026-06-30JIANGXI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGXI NORMAL UNIV
Filing Date
2026-04-13
Publication Date
2026-06-30

Smart Images

  • Figure CN122023813B_ABST
    Figure CN122023813B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image processing and proposes a medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion. By designing a geometrically guided text generation module, high-quality medical text descriptions are automatically generated based on connected component analysis and multidimensional geometric analysis, achieving a transformation from text-free segmentation to text-guided segmentation. This solves the problems of limited quantity, difficult production, and high cost of high-quality text descriptions. Furthermore, by designing a concept-aware cross-modal fusion module, multi-scale spatial alignment and medical knowledge graph modeling are utilized to address the problem of insufficient cross-modal fusion. Based on explicit modeling of prior medical knowledge, professional medical expertise is fully utilized to guide the fusion process and segmentation decisions. An adaptive segmentation decoding module is also proposed, dynamically selecting the optimal strategy based on feature complexity, improving efficiency while maintaining accuracy. This invention improves both the accuracy and efficiency of medical ultrasound image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing, and in particular to a medical ultrasound image segmentation method based on the fusion of geometric guidance and concept perception. Background Technology

[0002] Medical ultrasound imaging technology, as a core technology, plays a crucial role in precision medicine and clinical decision-making. Radiology staff need to manually segment abnormal areas from ultrasound images. However, this is not only a time-consuming process, but the quality of the segmentation results also depends heavily on the experience of the staff. Therefore, an automated ultrasound image segmentation model has become an urgent need.

[0003] In existing technologies, based on clinical practice, it is usually necessary to integrate image features and text descriptions for comprehensive judgment. This reflects the important value of multimodal information fusion. Therefore, visual-text image segmentation methods have emerged, which aim to improve segmentation accuracy by introducing text descriptions to guide visual feature learning. However, existing visual-text models face a fundamental challenge in medical image segmentation tasks: medical datasets rarely provide high-quality text descriptions for each image. Unlike natural image datasets, the professionalism and complexity of medical images make the production of large-scale text descriptions extremely difficult and costly. Although medical-specific models have alleviated this problem to some extent, their performance in automatically generating high-quality text descriptions is still limited. More importantly, existing methods have significant shortcomings in cross-modal fusion: (1) simple feature splicing or attention mechanisms cannot effectively integrate the deep semantic information of images and texts; (2) the lack of explicit modeling of medical prior knowledge makes it impossible to fully utilize medical expertise to guide the fusion process and segmentation decisions.

[0004] Therefore, designing a medical ultrasound image segmentation method to avoid the influence of the lack of text annotation in medical datasets and improve segmentation accuracy has become an urgent problem to be solved. Summary of the Invention

[0005] Based on this, this invention proposes a medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion. By designing a geometrically guided text generation module, high-quality medical text descriptions are automatically generated based on connected component analysis and multidimensional geometric analysis, achieving a transformation from text-free segmentation to text-guided segmentation. This solves the problems of limited quantity, difficult production, and high cost of high-quality text descriptions. Furthermore, by designing a concept-aware cross-modal fusion module, multi-scale spatial alignment and medical knowledge graph modeling are utilized to address the problem of insufficient cross-modal fusion. Based on explicit modeling of prior medical knowledge, professional medical expertise is fully utilized to guide the fusion process and segmentation decisions. An adaptive segmentation decoding module is also proposed, dynamically selecting the optimal strategy based on feature complexity, improving efficiency while maintaining accuracy. It preserves better details in complex lesion boundary regions and significantly reduces computational overhead in uniform background regions. This invention improves the accuracy and efficiency of medical ultrasound image segmentation.

[0006] This invention proposes a medical ultrasound image segmentation method based on geometry-guided and concept-aware fusion, comprising:

[0007] Acquire target medical ultrasound images and perform preprocessing. Input the preprocessed target ultrasound images into a medical ultrasound image segmentation model. The medical ultrasound image segmentation model includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module.

[0008] The text generation is performed according to the geometry-guided text generation module to obtain medical ultrasound text descriptions. The text generation includes geometric feature extraction and adaptive classification.

[0009] The medical ultrasound text description and the target ultrasound image are input into a multimodal feature encoding module to obtain ultrasound text features and ultrasound image features. The multimodal feature encoding module includes a text encoder and an image encoder.

[0010] Ultrasound text features and ultrasound image features are input into the concept-aware cross-modal fusion module. Feature enhancement is performed according to a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features. Feature fusion is performed according to a concept-aware fusion mechanism to obtain concept-aware fused features. The multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation.

[0011] The final segmentation result is obtained by decoding the data using an adaptive segmentation and decoding module, which is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction.

[0012] In summary, based on the aforementioned medical ultrasound image segmentation method using geometric guidance and concept-aware fusion, this invention designs a geometrically guided text generation module that automatically generates high-quality medical text descriptions based on connected component analysis and multidimensional geometric analysis, achieving a transformation from text-free segmentation to text-guided segmentation. This solves the problems of limited quantity, difficulty in production, and high cost of high-quality text descriptions. Furthermore, by designing a concept-aware cross-modal fusion module, which utilizes multi-scale spatial alignment and medical knowledge graph modeling, the problem of insufficient cross-modal fusion is addressed. Based on explicit modeling of prior medical knowledge, professional medical expertise is fully utilized to guide the fusion process and segmentation decisions. An adaptive segmentation decoding module is also proposed, which dynamically selects the optimal strategy based on feature complexity, improving efficiency while maintaining accuracy. It preserves better details in complex lesion boundary regions and significantly reduces computational overhead in uniform background regions. This invention improves the accuracy and efficiency of medical ultrasound image segmentation. Specifically, the process involves acquiring and preprocessing a target medical ultrasound image. The preprocessed image is then input into a medical ultrasound image segmentation model, which includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module. The geometry-guided text generation module generates text to obtain a medical ultrasound text description. This text generation includes geometric feature extraction and adaptive classification, achieving a transition from text-free segmentation to text-guided segmentation. This addresses the problems of limited quantity, difficulty in production, and high cost of high-quality text descriptions. The medical ultrasound text description and the target ultrasound image are then input into the multimodal feature encoding module to obtain ultrasound text features and ultrasound image features. This module includes a text encoder and an image encoder. The ultrasound text features and ultrasound image features are then input into the concept-aware cross-modal fusion module, which performs segmentation based on a multi-scale cross-modal attention mechanism. This invention improves the accuracy and efficiency of medical ultrasound image segmentation by performing feature enhancement to obtain multi-scale enhanced image features, and then performing feature fusion according to a concept-aware fusion mechanism to obtain concept-aware fused features. The multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation. By utilizing multi-scale spatial alignment and medical knowledge graph modeling, the problem of insufficient cross-modal fusion is solved. By explicitly modeling medical prior knowledge, professional knowledge in the medical field is fully utilized to guide the fusion process and segmentation decision. Decoding is performed according to an adaptive segmentation decoding module to obtain the final segmentation result. The adaptive segmentation decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction. It dynamically selects the optimal strategy according to feature complexity, improving efficiency while ensuring accuracy. It can maintain better details in complex lesion boundary regions and significantly reduce computational overhead in uniform background regions.

[0013] Furthermore, the step of generating text based on the geometry-guided text generation module to obtain a medical ultrasound text description specifically includes:

[0014] Connected component analysis is performed on the segmentation mask of the target medical ultrasound image to identify the geometric features of the target region. The specific algorithm for identifying these geometric features is as follows:

[0015] ,

[0016] ,

[0017] ,

[0018] ,

[0019] ,

[0020] ,

[0021] in, Represents a single connected component. Indicates area characteristics, Indicates roundness characteristics, This represents the perimeter of the outline of a connected component. Indicates aspect ratio characteristics. and This represents the minor axis length and major axis length of the ellipse fitting of the connected components. Indicates smoothness feature, and Let represent the approximate number of contour points and the original number of contour points of the connected component, respectively. Indicates density characteristics, This represents the area of ​​the convex hull of a connected component. Indicates location features, Represents a mapping function. and These represent the x-coordinate and y-coordinate of the centroid, respectively.

[0022] The target area is calculated and adaptively classified based on the geometric features. The specific algorithms for target area calculation and adaptive classification are as follows:

[0023] ,

[0024]

[0025] in, Indicates the target area. Indicates the total area. This indicates the target area size ratio: very_small represents a very small size ratio, small represents a small size ratio, moderate represents a medium size ratio, large represents a large size ratio, and very_large represents a very large size ratio. , , , These represent the 20th percentile threshold, 40th percentile threshold, 60th percentile threshold, and 80th percentile threshold of the dataset, respectively.

[0026] Geometrically guided medical ultrasound text descriptions are generated based on geometric features and adaptive classification results, and the BLIP model is fine-tuned with medical data.

[0027] Furthermore, the step of generating a geometry-guided medical ultrasound text description based on geometric features and adaptive classification results, and fine-tuning the BLIP model using medical data, specifically includes:

[0028] Using geometry-guided medical ultrasound text descriptions as supervision signals, the BLIP model is fine-tuned according to the LoRA parameter fine-tuning strategy and jointly trained on multiple medical datasets to obtain a fine-tuned and optimized BLIP model. The specific algorithm for obtaining the fine-tuned and optimized BLIP model is as follows:

[0029] ,

[0030] ,

[0031] in, This indicates the input text. This represents the fine-tuning optimization function. This represents the total number of training samples. Indicates the sample index. This represents the conditional probability of generating the target output text. This represents the output text of the BLIP model. Indicates the first One input image.

[0032] Furthermore, the step of inputting the medical ultrasound text description and the target ultrasound image into the multimodal feature encoding module to obtain ultrasound text features and ultrasound image features specifically includes:

[0033] The multimodal feature encoding module includes an image encoder and a text encoder;

[0034] The image encoder is based on a pyramid structure. Ultrasonic image features are obtained using the image encoder. The specific algorithm for obtaining these ultrasonic image features is as follows:

[0035] ,

[0036] in, Indicates ultrasound image features, Indicates an image encoder. Indicates the number of floors. Represents the target ultrasound image;

[0037] The ultrasonic text features are obtained based on a text encoder. The specific algorithm for obtaining the ultrasonic text features is as follows:

[0038] ,

[0039] in, Indicates ultrasonic text features. Indicates pooling, Indicates a text encoder. Indicates tokenization, This refers to a medical ultrasound text description.

[0040] Furthermore, the step of performing feature enhancement based on a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features specifically includes:

[0041] The ultrasound image features and ultrasound text features are uniformly projected into the same hidden dimension space. The specific algorithm for this uniform projection into the same hidden dimension space is as follows:

[0042] ,

[0043] ,

[0044] in, Indicates ultrasound image features, Indicates ultrasonic text features. This represents the features of the projected ultrasound image. This represents the features of the projected ultrasonic text. express activation, This indicates BatchNorm normalization. express convolution, express Regularization processing, This indicates LayerNorm normalization. Represents a linear transformation;

[0045] Learnable 2D positional codes are added to the features of the projected ultrasound image. The specific algorithm for adding learnable 2D positional codes is as follows:

[0046] ,

[0047] in, This indicates the addition of learnable 2D location-coded ultrasound image features. This indicates the addition of learnable 2D positional encoding. and Indicates the first The layer's feature height and feature width;

[0048] The image-text cross-modal association is calculated based on an adaptive attention mechanism, the specific algorithm of which is as follows:

[0049] ,

[0050]

[0051] in, , , These represent query features, key features, and value features, respectively. Indicates flattening process;

[0052] Spatial perception enhancement is performed based on a spatial modulation mechanism, and the specific algorithm for spatial perception enhancement is as follows:

[0053] ,

[0054] ,

[0055] in, Indicates spatial weights, This represents the sigmoid activation function. express convolution, Indicates global average pooling. This represents broadcast multiplication. Represents spatially-aware enhanced image features;

[0056] Multi-scale feature fusion is performed to obtain the multi-scale enhanced image features. The specific algorithm for multi-scale feature fusion is as follows:

[0057] ,

[0058] in, This represents multi-scale enhanced image features. This indicates splicing / merging.

[0059] Furthermore, the step of performing feature fusion based on the concept-aware fusion mechanism to obtain concept-aware fusion features specifically includes:

[0060] The attention weights for medical concepts are obtained, and the specific algorithm for these attention weights is as follows:

[0061] ,

[0062] in, Indicates the attention weight of medical concepts. express function, Represents a 100-dimensional linear transformation. express activation, Represents a 256-dimensional linear transformation. This indicates LayerNorm normalization. Indicates ultrasonic text features;

[0063] Based on graph convolution operations, each concept relationship type is assigned a unique learnable adjacency matrix to facilitate feature propagation through matrix multiplication and linear transformation.

[0064] Relationship weights are obtained based on a dynamic relationship attention mechanism. The specific algorithm for obtaining relationship weights is as follows:

[0065] ,

[0066] in, Represents relation weights. Represents a 5-dimensional linear transformation. Represents a 64-dimensional linear transformation;

[0067] The adjacency matrix is ​​obtained by weighting and combining the relation weights. This adjacency matrix is ​​dynamically adjusted based on the content of the medical ultrasound text description, activating the corresponding concept relation types through the medical ultrasound text description. The specific algorithm for obtaining the adjacency matrix is ​​as follows:

[0068] ,

[0069] in, Represents the adjacency matrix. Indicates the type of conceptual relationship. Represents the type of conceptual relationship The corresponding learnable adjacency matrix;

[0070] Concept embedding enhancement is performed based on a medical knowledge graph, and concept relationship modeling is carried out through graph convolution operations. The specific algorithm for concept relationship modeling is as follows:

[0071] ,

[0072] in, Indicates conceptual relationship, Represents the ordinal number of a concept. Indicates the first The medical concept attention weights corresponding to each concept Indicates the first Enhanced embedding representations corresponding to each concept;

[0073] The ultrasound image features are flattened into a sequence, and concept-guided image features are calculated using a multi-head attention mechanism. Then, image-text-concept multimodal information fusion is performed using an adaptive gating mechanism to obtain concept-aware fusion features. The specific algorithm for obtaining concept-aware fusion features is as follows:

[0074] ,

[0075] ,

[0076] ,

[0077] ,

[0078] ,

[0079] ,

[0080] ,

[0081] ,

[0082] ,

[0083] in, This represents the features of the flattened image. This indicates flattening. Indicates image processing, Represents the features of the input image. This indicates the query characteristics guided by concepts. and Representing key features and value features, This indicates adding a dimension with a size of 1, where the first dimension represents the position with index 1. This indicates multi-head attention-enhanced image features. Indicates the multi-head attention weight. This indicates a multi-head attention mechanism. Represents global image features. Indicates global average pooling. Represents global text features. Represents a linear transformation. These represent image fusion features, text fusion features, and concept fusion features, respectively. express function, Indicates a fusion gate. Indicates global fusion features, This indicates the characteristics of concept perception fusion. This represents the features of the original input image after image processing. express The reshaped spatial attention map This indicates extended processing. and These represent the feature height and feature width, respectively.

[0084] Furthermore, the step of performing decoding processing according to the adaptive segmentation decoding module to obtain the final segmentation result specifically includes:

[0085] The adaptive segmentation and decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction;

[0086] The adaptive upsampling calculates the feature complexity distribution based on a lightweight network. The specific algorithm for calculating the feature complexity distribution is as follows:

[0087] ,

[0088] in, Represents the feature complexity distribution. This represents the sigmoid activation function. express convolution, express activation, Indicates global average pooling. Indicates the decoding input features;

[0089] Path selection is performed based on feature complexity. Complex features are input into complex paths, and simple features are input into simple paths. Then, the output features of the complex and simple paths are adaptively weighted, fused, and refined to obtain adaptive upsampling refined features. The specific algorithm for obtaining adaptive upsampling refined features is as follows:

[0090] ,

[0091] ,

[0092] ,

[0093] ,

[0094] in, This represents the output characteristics of complex paths. This represents the output characteristics of a simple path. This represents transposed convolution. This indicates an upsampling step size adjustment operation. This indicates a fill operation. This indicates a scaling operation. This represents the bilinear interpolation operation. This represents adaptive weighted fusion features. This indicates adaptive upsampling to refine features. This indicates BatchNorm normalization. express convolution;

[0095] The feature pyramid refinement is based on feature refinement processing at three scales, and the specific algorithms for these three scales are as follows:

[0096] ,

[0097] ,

[0098] ,

[0099] in, , , This represents the refinement features at the first scale, the second scale, and the third scale. express convolution, Represents the fusion features of the input. This indicates a double adaptive upsampling. This indicates an 8x adaptive upsampling.

[0100] The context-aware prediction is based on contextual information and local detail features. The specific algorithm for the context-aware prediction is as follows:

[0101] ,

[0102] ,

[0103] ,

[0104] ,

[0105] ,

[0106] ,

[0107] in, Representing contextual features, Indicates local features, This represents a fusion of contextual and local features. and Indicates auxiliary supervision output, This indicates the final segmentation result.

[0108] This invention proposes a medical ultrasound image segmentation system based on geometry-guided and concept-aware fusion, comprising:

[0109] The preprocessing module is used to acquire and preprocess the target medical ultrasound image, and input the preprocessed target ultrasound image into the medical ultrasound image segmentation model. The medical ultrasound image segmentation model includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module.

[0110] A geometry-guided text generation module is used to generate text to obtain medical ultrasound text descriptions, the text generation including geometric feature extraction and adaptive classification;

[0111] A multimodal feature encoding module is used to acquire ultrasound text features and ultrasound image features. The multimodal feature encoding module includes a text encoder and an image encoder.

[0112] The concept-aware cross-modal fusion module is used to perform feature enhancement based on a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features, and to perform feature fusion based on a concept-aware fusion mechanism to obtain concept-aware fused features. The multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation.

[0113] An adaptive segmentation and decoding module is used to perform decoding processing to obtain the final segmentation result. The adaptive segmentation and decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction.

[0114] The present invention also provides a storage medium that stores one or more programs, which, when executed by a processor, implement the medical ultrasound image segmentation method based on geometry-guided and concept-aware fusion as described above.

[0115] The present invention also provides a computer device, the computer device including a memory and a processor, wherein:

[0116] The memory is used to store computer programs;

[0117] When the processor executes the computer program stored in the memory, it implements the medical ultrasound image segmentation method based on geometry guidance and concept perception fusion as described above. Attached Figure Description

[0118] Figure 1 This is a flowchart of the medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion proposed in the first embodiment of the present invention;

[0119] Figure 2 This is a schematic diagram of the structure of the medical ultrasound image segmentation system based on the fusion of geometry guidance and concept perception proposed in the second embodiment of the present invention.

[0120] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0121] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0122] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0123] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0124] Please see Figure 1 The diagram shows a flowchart of the medical ultrasound image segmentation method based on geometry-guided and concept-aware fusion proposed in the first embodiment of the present invention. This medical ultrasound image segmentation method based on geometry-guided and concept-aware fusion includes steps S01 to S05, wherein:

[0125] Step S01: Acquire the target medical ultrasound image and perform preprocessing, then input the preprocessed target ultrasound image into the medical ultrasound image segmentation model;

[0126] It should be noted that, in this embodiment, the medical ultrasound image segmentation model includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module.

[0127] Step S02: Generate text according to the geometry-guided text generation module to obtain a medical ultrasound text description;

[0128] It should be noted that in this embodiment, the text generation includes geometric feature extraction and adaptive classification. Connected component analysis is performed on the segmentation mask of the target medical ultrasound image to identify the geometric features of the target region. The specific algorithm for the geometric features is as follows:

[0129] ,

[0130] ,

[0131] ,

[0132] ,

[0133] ,

[0134] ,

[0135] in, Represents a single connected component. Indicates area characteristics, Indicates roundness characteristics, This represents the perimeter of the outline of a connected component. Indicates aspect ratio characteristics. and This represents the minor axis length and major axis length of the ellipse fitting of the connected components. Indicates smoothness feature, and Let represent the approximate number of contour points and the original number of contour points of the connected component, respectively. Indicates density characteristics, This represents the area of ​​the convex hull of a connected component. Indicates location features, Represents a mapping function. and These represent the x-coordinate and y-coordinate of the centroid, respectively.

[0136] The target area is calculated and adaptively classified based on the geometric features. The specific algorithms for target area calculation and adaptive classification are as follows:

[0137] ,

[0138]

[0139] in, Indicates the target area. Indicates the total area. This indicates the target area size ratio: very_small represents a very small size ratio, small represents a small size ratio, moderate represents a medium size ratio, large represents a large size ratio, and very_large represents a very large size ratio. , , , These represent the 20th percentile threshold, 40th percentile threshold, 60th percentile threshold, and 80th percentile threshold of the dataset, respectively.

[0140] Geometrically guided medical ultrasound text descriptions are generated based on geometric features and adaptive classification results, and the BLIP model is fine-tuned with medical data.

[0141] Using geometry-guided medical ultrasound text descriptions as supervision signals, the BLIP model is fine-tuned according to the LoRA parameter fine-tuning strategy and jointly trained on multiple medical datasets to obtain a fine-tuned and optimized BLIP model. The specific algorithm for obtaining the fine-tuned and optimized BLIP model is as follows:

[0142] ,

[0143] ,

[0144] in, This indicates the input text. This represents the fine-tuning optimization function. This represents the total number of training samples. Indicates the sample index. This represents the conditional probability of generating the target output text. This represents the output text of the BLIP model. Indicates the first One input image.

[0145] Step S03: Input the medical ultrasound text description and the target ultrasound image into the multimodal feature encoding module to obtain ultrasound text features and ultrasound image features;

[0146] It should be noted that in this embodiment, the multimodal feature encoding module includes a text encoder and an image encoder.

[0147] The image encoder is based on a pyramid structure. Ultrasonic image features are obtained using the image encoder. The specific algorithm for obtaining these ultrasonic image features is as follows:

[0148] ,

[0149] in, Indicates ultrasound image features, Indicates an image encoder. Indicates the number of floors. Represents the target ultrasound image;

[0150] The ultrasonic text features are obtained based on a text encoder. The specific algorithm for obtaining the ultrasonic text features is as follows:

[0151] ,

[0152] in, Indicates ultrasonic text features. Indicates pooling, Indicates a text encoder. Indicates tokenization, This refers to a medical ultrasound text description.

[0153] Step S04: Input the ultrasound text features and ultrasound image features into the concept-aware cross-modal fusion module, perform feature enhancement according to the multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features, and perform feature fusion according to the concept-aware fusion mechanism to obtain concept-aware fused features;

[0154] It should be noted that in this embodiment, the multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation, which projects ultrasound image features and ultrasound text features to the same hidden dimension space. The specific algorithm for this unified projection to the same hidden dimension space is as follows:

[0155] ,

[0156] ,

[0157] in, Indicates ultrasound image features, Indicates ultrasonic text features. This represents the features of the projected ultrasound image. This represents the features of the projected ultrasonic text. express activation, This indicates BatchNorm normalization. express convolution, express Regularization processing, This indicates LayerNorm normalization. Represents a linear transformation;

[0158] Learnable 2D positional codes are added to the features of the projected ultrasound image. The specific algorithm for adding learnable 2D positional codes is as follows:

[0159] ,

[0160] in, This indicates the addition of learnable 2D location-coded ultrasound image features. This indicates the addition of learnable 2D positional encoding. and Indicates the first The layer's feature height and feature width;

[0161] The image-text cross-modal association is calculated based on an adaptive attention mechanism, the specific algorithm of which is as follows:

[0162] ,

[0163]

[0164] in, , , These represent query features, key features, and value features, respectively. Indicates flattening process;

[0165] Spatial perception enhancement is performed based on a spatial modulation mechanism, and the specific algorithm for spatial perception enhancement is as follows:

[0166] ,

[0167] ,

[0168] in, Indicates spatial weights, This represents the sigmoid activation function. express convolution, Indicates global average pooling. This represents broadcast multiplication. Represents spatially-aware enhanced image features;

[0169] Multi-scale feature fusion is performed to obtain the multi-scale enhanced image features. The specific algorithm for multi-scale feature fusion is as follows:

[0170] ,

[0171] in, This represents multi-scale enhanced image features. This indicates splicing / merging.

[0172] The attention weights for medical concepts are obtained, and the specific algorithm for these attention weights is as follows:

[0173] ,

[0174] in, Indicates the attention weight of medical concepts. express function, Represents a 100-dimensional linear transformation. express activation, Represents a 256-dimensional linear transformation. This indicates LayerNorm normalization. Indicates ultrasonic text features;

[0175] Based on graph convolution operations, each concept relationship type is assigned a unique learnable adjacency matrix to facilitate feature propagation through matrix multiplication and linear transformation.

[0176] Relationship weights are obtained based on a dynamic relationship attention mechanism. The specific algorithm for obtaining relationship weights is as follows:

[0177] ,

[0178] in, Represents relation weights. Represents a 5-dimensional linear transformation. Represents a 64-dimensional linear transformation;

[0179] The adjacency matrix is ​​obtained by weighting and combining the relation weights. This adjacency matrix is ​​dynamically adjusted based on the content of the medical ultrasound text description, activating the corresponding concept relation types through the medical ultrasound text description. The specific algorithm for obtaining the adjacency matrix is ​​as follows:

[0180] ,

[0181] in, Represents the adjacency matrix. Indicates the type of conceptual relationship. Represents the type of conceptual relationship The corresponding learnable adjacency matrix;

[0182] Concept embedding enhancement is performed based on a medical knowledge graph, and concept relationship modeling is carried out through graph convolution operations. The specific algorithm for concept relationship modeling is as follows:

[0183] ,

[0184] in, Indicates conceptual relationship, Represents the ordinal number of a concept. Indicates the first The medical concept attention weights corresponding to each concept Indicates the first Enhanced embedding representations corresponding to each concept;

[0185] The ultrasound image features are flattened into a sequence, and concept-guided image features are calculated using a multi-head attention mechanism. Then, image-text-concept multimodal information fusion is performed using an adaptive gating mechanism to obtain concept-aware fusion features. The specific algorithm for obtaining concept-aware fusion features is as follows:

[0186] ,

[0187] ,

[0188] ,

[0189] ,

[0190] ,

[0191] ,

[0192] ,

[0193] ,

[0194] ,

[0195] in, This represents the features of the flattened image. This indicates flattening. Indicates image processing, Represents the features of the input image. This indicates the query characteristics guided by concepts. and Representing key features and value features, This indicates adding a dimension with a size of 1, where the first dimension represents the position with index 1. This indicates multi-head attention-enhanced image features. Indicates the multi-head attention weight. This indicates a multi-head attention mechanism. Represents global image features. Indicates global average pooling. Represents global text features. Represents a linear transformation. These represent image fusion features, text fusion features, and concept fusion features, respectively. express function, Indicates a fusion gate. Indicates global fusion features, This indicates the characteristics of concept perception fusion. This represents the features of the original input image after image processing. express The reshaped spatial attention map This indicates extended processing. and These represent the feature height and feature width, respectively.

[0196] Step S05: Perform decoding processing according to the adaptive segmentation decoding module to obtain the final segmentation result;

[0197] It should be noted that in this embodiment, the adaptive segmentation and decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction.

[0198] The adaptive upsampling calculates the feature complexity distribution based on a lightweight network. The specific algorithm for calculating the feature complexity distribution is as follows:

[0199] ,

[0200] in, Represents the feature complexity distribution. This represents the sigmoid activation function. express convolution, express activation, Indicates global average pooling. Indicates the decoding input features;

[0201] Path selection is performed based on feature complexity. Complex features are input into complex paths, and simple features are input into simple paths. Then, the output features of the complex and simple paths are adaptively weighted, fused, and refined to obtain adaptive upsampling refined features. The specific algorithm for obtaining adaptive upsampling refined features is as follows:

[0202] ,

[0203] ,

[0204] ,

[0205] ,

[0206] in, This represents the output characteristics of complex paths. This represents the output characteristics of a simple path. This represents transposed convolution. This indicates an upsampling step size adjustment operation. This indicates a fill operation. This indicates a scaling operation. This represents the bilinear interpolation operation. This represents adaptive weighted fusion features. This indicates adaptive upsampling to refine features. This indicates BatchNorm normalization. express convolution;

[0207] The feature pyramid refinement is based on feature refinement processing at three scales, and the specific algorithms for these three scales are as follows:

[0208] ,

[0209] ,

[0210] ,

[0211] in, , , This represents the refinement features at the first scale, the second scale, and the third scale. express convolution, Represents the fusion features of the input. This indicates a double adaptive upsampling. This indicates an 8x adaptive upsampling.

[0212] The context-aware prediction is based on contextual information and local detail features. The specific algorithm for the context-aware prediction is as follows:

[0213] ,

[0214] ,

[0215] ,

[0216] ,

[0217] ,

[0218] ,

[0219] in, Representing contextual features, Indicates local features, This represents a fusion of contextual and local features. and Indicates auxiliary supervision output, This indicates the final segmentation result.

[0220] Please see Figure 2 The figure shown is a schematic diagram of the medical ultrasound image segmentation system based on geometric guidance and concept-aware fusion proposed in the second embodiment of the present invention. The system includes:

[0221] Preprocessing module 10 is used to acquire target medical ultrasound images and perform preprocessing, and input the preprocessed target ultrasound images into a medical ultrasound image segmentation model. The medical ultrasound image segmentation model includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module.

[0222] The geometry-guided text generation module 20 is used to generate text to obtain medical ultrasound text descriptions, the text generation including geometric feature extraction and adaptive classification;

[0223] The multimodal feature encoding module 30 is used to acquire ultrasound text features and ultrasound image features. The multimodal feature encoding module includes a text encoder and an image encoder.

[0224] The concept-aware cross-modal fusion module 40 is used to perform feature enhancement based on a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features, and to perform feature fusion based on a concept-aware fusion mechanism to obtain concept-aware fusion features. The multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation.

[0225] An adaptive segmentation decoding module 50 is used to perform decoding processing to obtain the final segmentation result. The adaptive segmentation decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction.

[0226] The present invention also proposes a computer storage medium storing one or more programs that, when executed by a processor, implement the above-described medical ultrasound image segmentation method based on geometry-guided and concept-aware fusion.

[0227] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory to realize the above-mentioned medical ultrasound image segmentation method based on geometric guidance and concept perception fusion.

[0228] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0229] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0230] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0231] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0232] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion, characterized in that, include: Acquire target medical ultrasound images and perform preprocessing. Input the preprocessed target ultrasound images into a medical ultrasound image segmentation model. The medical ultrasound image segmentation model includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module. The text generation is performed according to the geometry-guided text generation module to obtain medical ultrasound text descriptions. The text generation includes geometric feature extraction and adaptive classification. Connected component analysis is performed on the segmentation mask of the target medical ultrasound image to identify the geometric features of the target region. The target area is calculated and adaptively classified based on the geometric features. A geometrically guided medical ultrasound text description is generated based on the geometric features and the adaptive classification results. The BLIP model is then fine-tuned with medical data. The medical ultrasound text description and the target ultrasound image are input into a multimodal feature encoding module to obtain ultrasound text features and ultrasound image features. The multimodal feature encoding module includes a text encoder and an image encoder. Ultrasound text features and ultrasound image features are input into the concept-aware cross-modal fusion module. Feature enhancement is performed according to a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features. Feature fusion is performed according to a concept-aware fusion mechanism to obtain concept-aware fused features. The multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation. Medical concept attention weights are obtained. Based on graph convolution operations, a unique learnable adjacency matrix is ​​assigned to each concept relationship type. Feature propagation is performed through matrix multiplication and linear transformation. Relationship weights are obtained based on a dynamic relationship attention mechanism. The relationship weights are weighted and combined to obtain an adjacency matrix. The adjacency matrix is ​​dynamically adjusted according to the content of the medical ultrasound text description. The corresponding concept relationship type is activated through the medical ultrasound text description. Concept embedding enhancement is performed based on the medical knowledge graph. Concept relationship modeling is performed through graph convolution operations. Ultrasound image features are flattened into a sequence form. Concept-guided image features are calculated through a multi-head attention mechanism. Finally, image-text-concept multimodal information fusion is performed through an adaptive gating mechanism to obtain concept-aware fusion features. The concept-aware fusion features are decoded by the adaptive segmentation decoding module to obtain the final segmentation result. The adaptive segmentation decoding module is based on adaptive upsampling, feature pyramid refinement and context-aware prediction.

2. The medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion according to claim 1, characterized in that, The step of generating text based on the geometry-guided text generation module to obtain a medical ultrasound text description specifically includes: Connected component analysis is performed on the segmentation mask of the target medical ultrasound image to identify the geometric features of the target region. The specific algorithm for identifying these geometric features is as follows: , , , , , , in, Represents a single connected component. Indicates area characteristics, Indicates roundness characteristics, This represents the perimeter of the outline of a connected component. Indicates aspect ratio characteristics. and This represents the minor axis length and major axis length of the ellipse fitting of the connected components. Indicates smoothness feature, and Let represent the approximate number of contour points and the original number of contour points of the connected component, respectively. Indicates density characteristics, This represents the area of ​​the convex hull of a connected component. Indicates location features, Represents a mapping function. and These represent the x-coordinate and y-coordinate of the centroid, respectively. The target area is calculated and adaptively classified based on the geometric features. The specific algorithms for target area calculation and adaptive classification are as follows: , , in, Indicates the target area. Indicates the total area. This indicates the target area size ratio: very_small represents a very small size ratio, small represents a small size ratio, moderate represents a medium size ratio, large represents a large size ratio, and very_large represents a very large size ratio. , , , These represent the 20th percentile threshold, 40th percentile threshold, 60th percentile threshold, and 80th percentile threshold of the dataset, respectively. Geometrically guided medical ultrasound text descriptions are generated based on geometric features and adaptive classification results, and the BLIP model is fine-tuned with medical data.

3. The medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion according to claim 2, characterized in that, The steps of generating geometry-guided medical ultrasound text descriptions based on geometric features and adaptive classification results, and fine-tuning the BLIP model with medical data, specifically include: Using geometry-guided medical ultrasound text descriptions as supervision signals, the BLIP model is fine-tuned according to the LoRA parameter fine-tuning strategy and jointly trained on multiple medical datasets to obtain a fine-tuned and optimized BLIP model. The specific algorithm for obtaining the fine-tuned and optimized BLIP model is as follows: , , in, This indicates the input text. This represents the fine-tuning optimization function. This represents the total number of training samples. Indicates the sample index. This represents the conditional probability of generating the target output text. This represents the output text of the BLIP model. Indicates the first One input image.

4. The medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion according to claim 1, characterized in that, The step of inputting the medical ultrasound text description and the target ultrasound image into the multimodal feature encoding module to obtain ultrasound text features and ultrasound image features specifically includes: The multimodal feature encoding module includes an image encoder and a text encoder; The image encoder is based on a pyramid structure. Ultrasonic image features are obtained using the image encoder. The specific algorithm for obtaining these ultrasonic image features is as follows: , in, Indicates ultrasound image features, Indicates an image encoder. Indicates the number of floors. Represents the target ultrasound image; The ultrasonic text features are obtained based on a text encoder. The specific algorithm for obtaining the ultrasonic text features is as follows: , in, Indicates ultrasonic text features. Indicates pooling, Indicates a text encoder. Indicates tokenization, This refers to a medical ultrasound text description.

5. The medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion according to claim 1, characterized in that, The step of performing feature enhancement based on a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features specifically includes: The ultrasound image features and ultrasound text features are uniformly projected into the same hidden dimension space. The specific algorithm for this uniform projection into the same hidden dimension space is as follows: , , in, Indicates ultrasound image features, Indicates ultrasonic text features. This represents the features of the projected ultrasound image. This represents the features of the projected ultrasonic text. express activation, This indicates BatchNorm normalization. express convolution, express Regularization processing, This indicates LayerNorm normalization. Represents a linear transformation; Learnable 2D positional codes are added to the features of the projected ultrasound image. The specific algorithm for adding learnable 2D positional codes is as follows: , in, This indicates the addition of learnable 2D location-coded ultrasound image features. This indicates the addition of learnable 2D positional encoding. and Indicates the first The layer's feature height and feature width; The image-text cross-modal association is calculated based on an adaptive attention mechanism, the specific algorithm of which is as follows: , in, , , These represent query features, key features, and value features, respectively. Indicates flattening; Spatial perception enhancement is performed based on a spatial modulation mechanism, and the specific algorithm for spatial perception enhancement is as follows: , , in, Indicates spatial weights, This represents the sigmoid activation function. express convolution, Indicates global average pooling. This represents broadcast multiplication. Represents spatially-aware enhanced image features; Multi-scale feature fusion is performed to obtain the multi-scale enhanced image features. The specific algorithm for multi-scale feature fusion is as follows: , in, This represents multi-scale enhanced image features. This indicates splicing / merging.

6. The medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion according to claim 1, characterized in that, The step of performing feature fusion based on the concept-aware fusion mechanism to obtain concept-aware fusion features specifically includes: The attention weights for medical concepts are obtained, and the specific algorithm for these attention weights is as follows: , in, Indicates the attention weight of medical concepts. express function, Represents a 100-dimensional linear transformation. express activation, Represents a 256-dimensional linear transformation. This indicates LayerNorm normalization. Indicates ultrasonic text features; Based on graph convolution operations, each concept relationship type is assigned a unique learnable adjacency matrix to facilitate feature propagation through matrix multiplication and linear transformation. Relationship weights are obtained based on a dynamic relationship attention mechanism. The specific algorithm for obtaining relationship weights is as follows: , in, Represents relation weights. Represents a 5-dimensional linear transformation. Represents a 64-dimensional linear transformation; The adjacency matrix is ​​obtained by weighting and combining the relation weights. This adjacency matrix is ​​dynamically adjusted based on the content of the medical ultrasound text description, activating the corresponding concept relation types through the medical ultrasound text description. The specific algorithm for obtaining the adjacency matrix is ​​as follows: , in, Represents the adjacency matrix. Indicates the type of conceptual relationship. Represents the type of conceptual relationship The corresponding learnable adjacency matrix; Concept embedding enhancement is performed based on a medical knowledge graph, and concept relationship modeling is carried out through graph convolution operations. The specific algorithm for concept relationship modeling is as follows: , in, Indicates conceptual relationship, Represents the ordinal number of a concept. Indicates the first The medical concept attention weights corresponding to each concept Indicates the first Enhanced embedding representations corresponding to each concept; The ultrasound image features are flattened into a sequence, and concept-guided image features are calculated using a multi-head attention mechanism. Then, image-text-concept multimodal information fusion is performed using an adaptive gating mechanism to obtain concept-aware fusion features. The specific algorithm for obtaining concept-aware fusion features is as follows: , , , , , , , , , in, This represents the features of the flattened image. This indicates flattening. Indicates image processing, Represents the features of the input image. This indicates the query characteristics guided by concepts. and Representing key features and value features, This indicates adding a dimension with a size of 1, where the first dimension represents the position with index 1. This indicates multi-head attention-enhanced image features. Indicates the multi-head attention weight. This indicates a multi-head attention mechanism. Represents global image features. Indicates global average pooling. Represents global text features. Represents a linear transformation. These represent image fusion features, text fusion features, and concept fusion features, respectively. express function, Indicates a fusion gate. Indicates global fusion features, This indicates the characteristics of concept perception fusion. This represents the features of the original input image after image processing. express The reshaped spatial attention map This indicates extended processing. and These represent the feature height and feature width, respectively.

7. The medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion according to claim 1, characterized in that, The step of performing decoding processing according to the adaptive segmentation decoding module to obtain the final segmentation result specifically includes: The adaptive segmentation and decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction; The adaptive upsampling calculates the feature complexity distribution based on a lightweight network. The specific algorithm for calculating the feature complexity distribution is as follows: , in, Represents the feature complexity distribution. This represents the sigmoid activation function. express convolution, express activation, Indicates global average pooling. Indicates the decoding input features; Path selection is performed based on feature complexity. Complex features are input into complex paths, and simple features are input into simple paths. Then, the output features of the complex and simple paths are adaptively weighted, fused, and refined to obtain adaptive upsampling refined features. The specific algorithm for obtaining adaptive upsampling refined features is as follows: , , , , in, This represents the output characteristics of complex paths. This represents the output characteristics of a simple path. This represents transposed convolution. This indicates an upsampling step size adjustment operation. This indicates a fill operation. This indicates a scaling operation. This represents the bilinear interpolation operation. This represents adaptive weighted fusion features. This indicates adaptive upsampling to refine features. This indicates BatchNorm normalization. express convolution; The feature pyramid refinement is based on feature refinement processing at three scales, and the specific algorithms for these three scales are as follows: , , , in, , , This represents the refinement features at the first scale, the second scale, and the third scale. express convolution, Represents the fusion features of the input. This indicates a double adaptive upsampling. This indicates an 8x adaptive upsampling. The context-aware prediction is based on contextual information and local detail features. The specific algorithm for the context-aware prediction is as follows: , , , , , , in, Representing contextual features, Indicates local features, This represents a fusion of contextual and local features. and Indicates auxiliary supervision output, This indicates the final segmentation result.

8. A medical ultrasound image segmentation system based on geometric guidance and concept-aware fusion, characterized in that, include: The preprocessing module is used to acquire and preprocess the target medical ultrasound image, and input the preprocessed target ultrasound image into the medical ultrasound image segmentation model. The medical ultrasound image segmentation model includes a geometry-guided text generation module, a multimodal feature encoding module, a concept-aware cross-modal fusion module, and an adaptive segmentation decoding module. A geometry-guided text generation module is used to generate text to obtain medical ultrasound text descriptions, the text generation including geometric feature extraction and adaptive classification; Connected component analysis is performed on the segmentation mask of the target medical ultrasound image to identify the geometric features of the target region. The target area is calculated and adaptively classified based on the geometric features. A geometrically guided medical ultrasound text description is generated based on the geometric features and the adaptive classification results. The BLIP model is then fine-tuned with medical data. A multimodal feature encoding module is used to acquire ultrasound text features and ultrasound image features. The multimodal feature encoding module includes a text encoder and an image encoder. The concept-aware cross-modal fusion module is used to perform feature enhancement based on a multi-scale cross-modal attention mechanism to obtain multi-scale enhanced image features, and to perform feature fusion based on a concept-aware fusion mechanism to obtain concept-aware fused features. The multi-scale cross-modal attention mechanism is based on an adaptive attention mechanism, and the concept-aware fusion mechanism is based on medical concept representation. Medical concept attention weights are obtained. Based on graph convolution operations, a unique learnable adjacency matrix is ​​assigned to each concept relationship type. Feature propagation is performed through matrix multiplication and linear transformation. Relationship weights are obtained based on a dynamic relationship attention mechanism. The relationship weights are weighted and combined to obtain an adjacency matrix. The adjacency matrix is ​​dynamically adjusted according to the content of the medical ultrasound text description. The corresponding concept relationship type is activated through the medical ultrasound text description. Concept embedding enhancement is performed based on the medical knowledge graph. Concept relationship modeling is performed through graph convolution operations. Ultrasound image features are flattened into a sequence form. Concept-guided image features are calculated through a multi-head attention mechanism. Finally, image-text-concept multimodal information fusion is performed through an adaptive gating mechanism to obtain concept-aware fusion features. An adaptive segmentation decoding module is used to decode concept-aware fusion features to obtain the final segmentation result. The adaptive segmentation decoding module is based on adaptive upsampling, feature pyramid refinement, and context-aware prediction.

9. A storage medium, characterized in that, The storage medium stores one or more programs that, when executed by a processor, implement the medical ultrasound image segmentation method based on geometric guidance and concept-aware fusion as described in any one of claims 1-7.

10. A computer device, characterized in that, The computer device includes a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the medical ultrasound image segmentation method based on geometric guidance and concept perception fusion as described in any one of claims 1-7.