MedLP and HAFB collaboratively-driven fetal level-II ultrasound image classification method
Through the fetal level II ultrasound image classification method driven by MedLP and HAFB, the problems of complex anatomical structure, data isomerism and insufficient medical knowledge in fetal ultrasound image classification are solved, and high-precision automated classification and diagnostic support are achieved.
Patent Information
- Application Number
- CN202510317106.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-04
AI Technical Summary
When existing fetal ultrasound image classification methods face problems such as low ultrasound image quality, high anatomical similarity, data heterogeneity and inconsistent labeling, it is difficult to achieve high-precision automated classification, and lack medical knowledge guidance, resulting in insufficient generalization capabilities of models and low doctor trust.
Using a collaboratively driven method between MedLP and HAFB, a fine-grained anatomical description and multi-scale feature extraction in the medical field are introduced by designing a prompt word learner. Combined with the adaptive fusion module, the MedLP-HAFB-CLIP model is built to optimize model parameters to improve the recognition ability of anatomical structures.
It significantly improves the accuracy and efficiency of fetal ultrasound image classification, especially in identifying complex and similar anatomical structures, enhances the generalization ability and clinical applicability of the model, and improves the reliability and accuracy of the diagnosis.
Smart Images

Figure CN120259739A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a method for classifying fetal level-II ultrasound images driven by the cooperation of MedLP and HAFB. Background Art
[0002] In traditional prenatal ultrasound examinations during the second trimester (22-26 weeks of gestation), the accurate identification and classification of standard fetal anatomical sections are the core diagnostic basis for evaluating fetal development status, and their accuracy directly affects the screening efficiency of birth defects and the timeliness of clinical intervention. According to the "Guidelines for Prenatal Ultrasound Examination of the Chinese Medical Doctor Association Ultrasonography Branch", 10 standard section images of the fetal head and face, spine, abdomen, etc. need to be systematically collected during the second-trimester level-II screening, covering key anatomical parts such as the transverse abdominal circumference section, the transverse lateral ventricle section, and the transverse thalamus section. Taking the transverse lateral ventricle section as an example, its measured value is directly related to the judgment of the degree of ventricular dilation, while the transverse thalamus section is the gold standard for evaluating cranial symmetry and the integrity of the midline structure. The standardized acquisition and classification of these standard sections not only provide a reliable basis for biometric measurements (such as biparietal diameter, head circumference), but also are the core basis for screening nervous system malformations (such as spina bifida, hydrocephalus). The automated implementation of standardized section classification can provide a reliable basis for the early screening of high-risk fetuses by reducing operator dependence and subjective bias. Therefore, achieving automated high-precision classification of fetal ultrasound standard sections has important clinical value for improving the quality of prenatal diagnosis and reducing the birth defect rate.
[0003] In recent years, algorithms such as large models and deep learning have developed rapidly in the field of medical image processing. In the scenario of fetal ultrasound image classification, deep learning algorithms such as convolutional neural networks (CNNs) and vision transformers (ViTs) have shown great potential in fetal section classification. CNNs can automatically learn the high-level semantic features of ultrasound images. For example, classic models based on CNNs such as VGGNet and ResNet have achieved high accuracy in classification tasks.
[0004] Although the existing technology provides new possibilities for fetal ultrasound image classification, its practical application still faces multiple technical bottlenecks. First, the physical characteristics of ultrasound images and fetal dynamic interference result in low data quality: maternal tissue attenuation, fetal position changes, and equipment parameter differences (such as probe frequency) introduce significant noise (reverberation artifacts, acoustic shadow occlusion), causing key anatomical structures to be blurred or geometrically deformed (such as the distortion of the anterior horn of the lateral ventricle). For example, tissues and organs in the pregnant woman's body, as well as the activities of the fetus, etc., will all generate artifacts, masking the effective information in the image and making it difficult to clearly display the fetal anatomical structure. Moreover, the posture and position of the fetus in the mother's body vary greatly, and the presentation angles and morphological differences of the same anatomical structure in different images are significant. Second, highly similar anatomical cross-sections lead to classification confusion: the transverse sections of the lateral ventricle and thalamus share common features such as the strong echo ring of the skull and the cavum septum pellucidum, and their differentiation relies on subtle differences (the triangular shape of the anterior horn vs. the symmetry of the thalamus). However, traditional models lack medical knowledge guidance and it is difficult to effectively distinguish such fine-grained features only through data-driven methods. In addition, multi-center data heterogeneity and inconsistent annotations severely restrict the generalization ability of the model: different hospitals use different ultrasound equipment models (such as GE E8 and Samsung WS80A) with different parameter settings, and the collected images have obvious differences in resolution, contrast, image format, etc. These differences make it difficult to directly integrate and share data, increase the difficulty of model training, and reduce the generalization ability of the model. Moreover, the scarcity of rare samples (such as single ventricle) and the long-tailed distribution of common cross-sections (such as four-chamber heart) exacerbate the overfitting risk, and the lack of unified anatomical annotation standards further weakens the clinical applicability of the model. Finally, the "black box" characteristic of existing methods hinders clinical implementation: the model cannot quantitatively analyze the correlation between anatomical features and clinical indicators, resulting in insufficient doctor trust.
[0005] In view of this, the present invention proposes a method for classifying fetal grade II ultrasound images jointly driven by MedLP and HAFB. Summary of the Invention
[0006] The object of the present invention is to propose a method for classifying fetal grade II ultrasound images jointly driven by MedLP and HAFB, which can accurately identify ultrasound images of different anatomical structures, significantly improve the accuracy and efficiency of fetal ultrasound image classification, and provide stronger support for clinical prenatal diagnosis.
[0007] To achieve the above object, the technical solution of the present invention is: a method for classifying fetal grade II ultrasound images jointly driven by MedLP and HAFB, specifically including the following steps:
[0008] S1: Collect fetal grade II ultrasound sectional images of key parts, including transverse abdominal circumference section, transverse binocular section, sagittal section of both kidneys, transverse section of both kidneys, transverse cerebellum section, transverse lateral ventricle section, median sagittal facial section, coronal nasolabial section, longitudinal spinal section, and transverse thalamus section; perform sectional marking and preprocessing on the collected ultrasound images, and randomly shuffle the order of all images and their corresponding labels for model training;
[0009] S2: Based on the CLIP model, fuse the MedLP strategy and the HAFB module to construct the MedLP-HAFB-CLIP model;
[0010] The MedLP strategy generates prompts matching anatomical categories by designing a prompt learner using learnable context vectors to enable the model to understand anatomical features in ultrasound images; for difficult categories, introduce fine-grained anatomical descriptions in the medical field and construct an auxiliary loss function to enhance the model's ability to distinguish difficult samples;
[0011] The HAFB module extracts image features from different levels through a multi-scale feature extraction unit to capture detailed and semantic information in ultrasound images; the adaptive fusion module dynamically adjusts the fusion weights of features at different levels according to task requirements to generate a more discriminative feature representation;
[0012] S3: Train the constructed MedLP-HAFB-CLIP model and optimize the model parameters by minimizing the total loss function;
[0013] S4: Input the ultrasound image into the trained MedLP-HAFB-CLIP model. The input ultrasound image generates a unified feature representation after feature extraction, fusion, and processing, and is matched with the predefined anatomical category text prompt features to generate a similarity score matrix. According to the similarity score matrix, obtain the category index with the highest similarity assigned by the model for each image as the prediction result of the input ultrasound image.
[0014] Preferably, the preprocessing of the collected ultrasound images specifically includes:
[0015] S1.1: Uniformly process the image size and scale the shortest side of the image to a preset number of pixels to provide a relatively consistent basic size for subsequent processing;
[0016] S1.2: Convert the image to RGB mode to adapt to the general image processing process;
[0017] S1.3: Convert the image to a tensor and perform normalization processing to make the data have a reasonable distribution;
[0018] S1.4: Crop a fixed-size area from the center of the image to ensure that the image sizes input to the model are exactly the same.
[0019] Preferably, by designing a prompt word learner, learnable context vectors are used to generate prompt words that match the anatomical categories, specifically including the following steps:
[0020] S2.1.1: Randomly initialize a set of context vectors where n ctx represents the number of context words, and d is the feature dimension; the context vectors are used as learnable parameters to adaptively adjust the prompt words to match the image features of different anatomical categories;
[0021] S2.1.2: According to the set of anatomical category names L = {l0, l1, …, l9}, generate an initial set of prompt words The set of anatomical category names specifically includes: l0 is the transverse section of the abdominal circumference, l1 is the transverse section of both eyes, l2 is the sagittal section of both kidneys, l3 is the transverse section of both kidneys, l4 is the transverse section of the cerebellum, l5 is the transverse section of the lateral ventricle, l6 is the median sagittal section of the face, l7 is the coronal section of the nose and lip, l8 is the longitudinal section of the spine, l9 is the transverse section of the thalamus. These anatomical category names serve as the basis for generating the initial prompt words and provide key information for the model to learn the features of different ultrasound sections; each prompt word p i is composed of a context word prefix and a category name suffix, that is, p i = prefix + l i ;
[0022] S2.1.3: Encode the initial prompt words through the word embedding layer of the CLIP model to obtain an embedding representation where n tokens is the number of tokens of the prompt word; the embedding representation is divided into a prefix context and a suffix in three parts; during the forward propagation of the model, the prompt word learner dynamically generates the final prompt word embedding according to the context vector C and the fixed prefix and suffix embeddings
[0023] Preferably, for difficult categories, fine-grained anatomical descriptions in the medical field are introduced, and an auxiliary loss function is constructed to enhance the model's ability to distinguish difficult samples, specifically:
[0024] S2.2.1: Medical experts, based on clinical anatomical knowledge, conduct a fine-grained definition of the ultrasonic imaging features of difficult categories, and input the content of the fine-grained definition as initial prompt words into the model to enhance the model's learning ability for complex features; the difficult categories include Category 5 and Category 9, namely the transverse section of the lateral ventricle and the transverse section of the thalamus;
[0025] S2.2.2: Construct an auxiliary loss function for Category 5 and Category 9; during the model training process, when the label of the input image is Category 5 or Category 9, the features of the input image and the corresponding prompt word features are processed separately:
[0026] For the image features of Category 5 and Category 9 Extract the corresponding text prompt word features where n 5 / 9 is the number of images in Category 5 and Category 9; calculate the similarity score between each image belonging to Category 5 or Category 9 and the text prompt word features corresponding to Category 5 and Category 9, and obtain a similarity score matrix where the number of rows n 5 / 9 corresponds to the number of images, and the number of columns is 2, respectively representing the similarity scores between each image and the text prompt word features of Category 5 and Category 9;
[0027] Use the cross-entropy loss function to supervise the similarity scores:
[0028]
[0029] where is the true label of Category 5 or Category 9.
[0030] Preferably, the dot product similarity is used to calculate the similarity score between each image belonging to Category 5 or Category 9 and the text prompt word features corresponding to Category 5 and Category 9; first, perform and Normalize by row so that the norm of each vector is 1, and then calculate the similarity score through matrix multiplication to obtain a similarity score matrix
[0031] Preferably, step S3 adopts a selective parameter update strategy for model training, only allowing the parameters of the visual encoder and the prompt word learner to be updated, and freezing the remaining parameters of the model; the total loss function includes a contrastive loss, a classification loss for the overall model training, and an auxiliary loss for difficult categories; the total loss function L total is expressed as:
[0032] L total =L + βL 5 / 9 =αL contrastive +(1 - α)L classification+βL 5 / 9
[0033] L = αL contrastive +(1 - α)L classification
[0034] where β is a hyperparameter used to control the weight of the hard class loss; L contrastive and L classification are the contrast loss and the classification loss respectively, and α is a hyperparameter used to balance the weights of the contrast loss and the classification loss.
[0035] Preferably, the contrast loss uses the cosine similarity function, specifically as follows:
[0036] For the input image I ∈ R b×c×h×w and the corresponding label y ∈ Z b , first use the visual encoder of the CLIP model to extract the image feature F I ∈ R b×d , and perform normalization, where b is the batch size, c is the number of channels, and h and w are the height and width of the image; then select the corresponding prompt word embedding P final [y] according to the label y, use the text encoder to extract the text feature F T ∈ R b×d , and perform normalization; the contrast loss is used to maximize the cosine similarity between the image feature F I and the corresponding text feature F T , and the formula is:
[0037]
[0038] where sim(·, ·) is the cosine similarity function used to measure the similarity between two vectors; τ is the temperature parameter used to control the sensitivity of the model to similarity judgment during contrast learning; the index j is used to traverse the negative examples of all samples within the batch.
[0039] Preferably, the classification loss uses the cross - entropy loss function, and the formula is:
[0040]
[0041] where p(y i |F I , P final ) is the probability that the model predicts the label y I given the image feature F final and the prompt word embedding P i ; the index i is used to traverse all samples within the batch.
[0042] Preferably, the HAFB module extracts image features from different levels through a multi-scale feature extraction unit to capture the detailed information and semantic information in the ultrasonic image; through the adaptive fusion module, the fusion weights of features at different levels are dynamically adjusted according to the task requirements to generate a more discriminative feature representation, as follows:
[0043] The HAFB module is constructed based on the CLIP vision encoder, and the CLIP vision encoder adopts the ViT-B / 16 architecture and consists of 12 layers of Transformer blocks; the HAFB module extracts multi-scale features from the 1st, 3rd, 5th, 7th, 9th, and 11th layers of Transformer blocks and extracts the CLS Token feature F of the corresponding layer l As the hierarchical semantic representation, where F l = X l [0, :] ∈ R d and l ∈ {1, 3, 5, 7, 9, 11}, X l is the feature output of the l-th layer of Transformer block;
[0044] Let the current processing level be l and the current level feature be The accumulated fusion feature is where n is the batch size and d prev is the historical feature dimension, and the fusion process is as follows:
[0045] First, the historical feature and the current feature are concatenated along the channel dimension to retain the integrity of multi-scale information, obtaining F concat :
[0046]
[0047] Next, the concatenated feature is dimension-reduced through the learnable parameter and the bias to obtain the transformed feature F trans :
[0048]
[0049] where, d new is the unified feature dimension;
[0050] To balance the importance of the historical fusion feature F trans and the current original feature F l , a learnable parameter θ l ∈ R is introduced, and the sigmoid function is applied to obtain the weight coefficient:
[0051]
[0052] where, σ represents the sigmoid function;
[0053] Finally, according to the fusion weight α l for F trans and the original feature F of the current level l perform weighted summation to obtain the fused feature F new :
[0054] F new =(1 - α l (F trans +α l F l .
[0055] Preferably, in step S4, the generated unified feature representation is matched with the predefined anatomical category text prompt feature F T through the cosine similarity function sim(·,·) to generate a similarity score matrix S∈R m×n , where m is the number of input images and n is the total number of anatomical categories. The matrix element S p,q represents the semantic matching degree between the p-th image and the q-th category; by applying the argmax operation row by row, the model assigns the category index with the highest similarity as the prediction result for each image; a confidence threshold mechanism is introduced. If the highest similarity score is lower than the preset threshold, the model will mark the result as an uncertain classification or trigger an artificial review process.
[0056] Compared with the prior art, the present invention has the following beneficial effects:
[0057] The present invention innovatively proposes the MedLP (Medical Learnable Prompt) strategy; by designing a prompt word learner, using learnable context vectors to generate prompt words that match anatomical categories, and integrating medical domain knowledge into the model training process skillfully. This strategy enables the model to understand the anatomical features in ultrasound images more deeply and effectively improves the classification ability for complex images. At the same time, for difficult anatomical categories such as the transverse section of the lateral ventricle and the transverse section of the thalamus, fine-grained anatomical descriptions in the professional medical field are introduced, and an auxiliary loss function is constructed to further enhance the model's discrimination ability for difficult samples.
[0058] The present invention also introduces the HierarchicalAdaptive Feature Block (HAFB) module; this module extracts image features from different levels through a multi-scale feature extraction unit, which can fully capture rich detail information and semantic information in ultrasound images. The adaptive fusion module dynamically adjusts the fusion weights of features at different levels according to task requirements to generate a more discriminative feature representation, further improving the classification performance of the model.
[0059] In data processing, for the problems of data imbalance and multi - center data heterogeneity, pre - processing methods such as oversampling and data standardization are adopted to improve the availability of data and the generalization ability of the model. Brief Description of the Drawings
[0060] Figure 1 It is a schematic diagram of the overall architecture of the MedLP - HAFB - CLIP model of the present invention. Detailed Embodiments
[0061] Next, in combination with the drawings, the technical solutions of the present invention will be specifically described.
[0062] The present invention proposes a method for classifying fetal level - II ultrasound images jointly driven by MedLP and HAFB, which specifically includes the following steps:
[0063] S1: Collect fetal level - II ultrasound cross - sectional images of key parts, including abdominal circumference cross - section, binocular cross - section, double - kidney sagittal section, double - kidney cross - section, cerebellum cross - section, lateral ventricle cross - section, mid - facial median sagittal section, nasolabial coronal section, spinal cord longitudinal section, thalamus cross - section; perform section marking and pre - processing on the collected ultrasound images, and randomly shuffle the order of all images and their corresponding labels for model training, so as to improve the generalization ability of the model and avoid the model learning the order deviation of the data. To ensure data quality, all images are collected by doctors with professional training according to clinical quality standards, and at the same time, samples with blurred images caused by fetal heart or limb movement are excluded; to ensure the consistency and accuracy of annotation, the annotation work is participated by multiple doctors, strictly follows the standardized annotation protocol, and undergoes strict review to ensure compliance with medical standards.
[0064] S2: Based on the CLIP model, fuse the MedLP strategy and the HAFB module to construct the MedLP - HAFB - CLIP model;
[0065] The MedLP strategy generates prompts matching anatomical categories by designing a prompt word learner using learnable context vectors, so that the model can understand the anatomical features in ultrasound images; for difficult categories, fine - grained anatomical descriptions in the medical field are introduced, and an auxiliary loss function is constructed to enhance the model's ability to distinguish difficult samples;
[0066] The HAFB module extracts image features from different levels through a multi - scale feature extraction unit to capture the detailed information and semantic information in ultrasound images; through the adaptive fusion module, the fusion weights of features at different levels are dynamically adjusted according to task requirements to generate a more discriminative feature representation;
[0067] S3: Train the constructed MedLP-HAFB-CLIP model and optimize the model parameters by minimizing the total loss function;
[0068] S4: Input the ultrasound image into the trained MedLP-HAFB-CLIP model. The input ultrasound image generates a unified feature representation after feature extraction, fusion, and processing, and is matched with the predefined anatomical category text prompt features to generate a similarity score matrix. According to the similarity score matrix, obtain the category index with the highest similarity assigned by the model for each image as the prediction result of the input ultrasound image.
[0069] In this embodiment, the preprocessing of the collected ultrasound images specifically includes:
[0070] S1.1: Uniformly process the image size and scale the shortest side of the image to a preset pixel (e.g., 224 pixels); provide a relatively consistent basic size for subsequent processing;
[0071] S1.2: Convert the image to the RGB mode to adapt to the general image processing process;
[0072] S1.3: Convert the image to a tensor and perform normalization processing to make the data have a reasonable distribution;
[0073] S1.4: Crop a fixed-size area from the center of the image to ensure that the image sizes input into the model are exactly the same.
[0074] In this embodiment, the MedLP strategy includes prompt word optimization and loss function customization for medical ultrasound image classification:
[0075] 1. Dynamic prompt word generation mechanism based on context vectors
[0076] To enable the model to better understand the anatomical features in medical ultrasound images, the present invention designs a Prompt Learner, and the position of this component in the overall model architecture is as Figure 1 shown.
[0077] First, randomly initialize a set of context vectors where n ctx represents the number of context words, and d is the feature dimension. These context vectors will be used as learnable parameters to adaptively adjust the prompt words to match the image features of different anatomical categories. According to the anatomical category name set L = {l0, l1,..., l9}, generate the initial prompt word set The anatomical category name set of the present invention covers the ultrasonic sections of 10 key parts, including: l0 is the abdominal cross section; l1 is the eye cross section; l2 is the bilateral kidney sagittal section, l3 is the bilateral kidney cross section, l4 is the cerebellum cross section, l5 is the lateral ventricle cross section, l6 is the facial median sagittal section, l7 is the nasolabial coronal section, l8 is the spinal longitudinal section, and l9 is the thalamus cross section. These anatomical category names serve as the basis for generating initial prompt words, providing key information for the model to learn the characteristics of different ultrasonic sections. Each prompt word p i It is composed of the context word prefix and the category name suffix, that is, p i = prefix + l i The initial prompt word is encoded through the word embedding layer of the CLIP model to obtain the embedded representation where n tokens is the number of word segments of the prompt word. The embedded representation is divided into prefix Context and suffix Three parts. During the forward propagation of the model, the prompt word learner dynamically generates the final prompt word embedding based on the context vector C and the fixed prefix and suffix embeddings This dynamic generation mechanism enables the model to adaptively adjust the cue words to fit the image features of different anatomical categories.
[0078] 2. Double reinforcement strategy for difficult categories
[0079] In the task of medical ultrasound image classification, category 5 (lateral ventricle cross section) and category 9 (thalamus cross section) are identified as difficult categories due to their high anatomical structure similarity and large overlap of image features. Aiming at the high confusion characteristics of lateral ventricle and thalamus cross sections, a dual enhancement strategy is designed: fine-grained anatomical description and dual supervision signals.
[0080] Professional medical domain knowledge is introduced in the fine-grained anatomical description. With their rich clinical experience and in-depth understanding of anatomy, medical experts are able to point out those anatomical structures that are easily confused and the key fine-grained features that distinguish them. Through cooperation with medical experts, the ultrasound image features of category 5 (lateral ventricle cross section) and category 9 (thalamus cross section) are finely defined based on clinical anatomical knowledge:
[0081] Category 5 (cross-section of the lateral ventricle): (1) hyperechoic skull ring; (2) centrally located falx cerebri; (3) cavum septum pellucidum; (4) anterior horn (a triangular hypoechoic area with its tip forward) and posterior horn (a linear hypoechoic structure extending posteriorly) of the lateral ventricle; (5) choroid plexus within the lateral ventricle (appears as a uniform hyperechoic band within the posterior horn). The key distinguishing features are the triangular shape of the anterior horn, the linear extension of the posterior horn, and the fixed spatial distribution of the choroid plexus.
[0082] Category 9 (transverse section of the thalamus): (1) strong echo ring of the skull; (2) midline falx cerebri; (3) cavum septum pellucidum; (4) bilaterally symmetric thalamus (oval, homogeneous, medium echo structure); (5) anterior horns of the bilateral lateral ventricles (appearing as thin-slit-like hypoecho); (6) sylvian fissure (linear hyperecho lateral to the thalamus). The key differential features are the symmetry of the thalamus, its adjacent relationship with the anterior horns of the lateral ventricles (the thalamus is located posterior to the anterior horns), and the anatomical landmark position of the sylvian fissure.
[0083] Directly input the above fine-grained definition content as the initial prompt into the model to enhance the model's learning ability for these complex features.
[0084] In terms of the dual supervision signals, the present invention constructs a supervision system from two levels: overall model training and special treatment of difficult categories. In the overall model training, contrastive loss and classification loss functions are used as the first layer of supervision.
[0085] To make full use of the feature extraction ability of the pre-trained CLIP model while avoiding overfitting problems, the present invention adopts a selective parameter update strategy. Most of the parameters of the CLIP model are frozen, and only the parameters of the visual encoder and the prompt learner are allowed to be updated. This can efficiently optimize for the medical ultrasound image classification task on the basis of retaining the powerful feature representation ability of the pre-trained model. For the input image I ∈ R b×c×h×w and the corresponding label y ∈ Z b (where b is the batch size, c is the number of channels, h and w are the height and width of the image), first use the visual encoder of the CLIP model to extract the image feature F I ∈ R b×d , and perform normalization processing. Then, select the corresponding prompt embedding P final [y] according to the label y, use the text encoder to extract the text feature F T ∈ R b×d , and also perform normalization processing.
[0086] The contrastive loss aims to maximize the cosine similarity between the image feature F I and the corresponding text feature F T , and the formula is:
[0087]
[0088] Among them, sim(·,·) is the cosine similarity function, which is used to measure the similarity between two vectors; τ is the temperature parameter, which is used to control the sensitivity of the model to similarity judgment during contrastive learning. A smaller τ will make the model more strict in distinguishing similarity, while a larger τ will make the model more lenient in distinguishing similarity; the index j is used to traverse the negative examples of all samples within the batch.
[0089] The classification loss uses the cross-entropy loss function:
[0090]
[0091] Among them, p(y i |F I , P final ) is the probability that the model predicts the label y I given the image feature F final and the prompt embedding P i ; the index i is used to traverse all samples within the batch.
[0092] By combining the contrastive loss and the classification loss function, the first-level supervision algorithm is obtained:
[0093] L = αL contrastive + (1 - α)L classification
[0094] Among them, α is a hyperparameter used to balance the weights of the contrastive loss and the classification loss.
[0095] For categories 5 and 9, an auxiliary loss function is specially constructed as the second-level supervision. During the training process, when the label of the input image is category 5 or category 9, the features of these images and the corresponding prompt features are processed separately. Specifically, for the image features of categories 5 and 9 (where n 5 / 9 is the number of images in categories 5 and 9), the corresponding text prompt features are extracted Here, the dot product similarity is used to calculate the similarity scores between each image belonging to category 5 or category 9 (a total of n 5 / 9 images) and the text prompt features corresponding to categories 5 and 9. The specific approach is to first and perform normalization processing row by row so that the norm of each vector is 1, and then calculate the similarity scores through matrix multiplication to obtain a similarity score matrix where the number of rows n 5 / 9 of the matrix corresponds to the number of images, and the number of columns is 2, respectively representing the similarity scores between each image and the text prompt features of categories 5 and 9.
[0096] The cross-entropy loss function is used to supervise these scores:
[0097]
[0098] where is the true label of class 5 or class 9.
[0099] The final loss function is composed of contrastive loss, classification loss, and auxiliary loss for difficult classes. The auxiliary loss function L 5 / 9 is added to the final loss function:
[0100] L total = L + βL 5 / 9 = αL contrastive + (1 - α)L classification + βL 5 / 9
[0101] where β is a hyperparameter used to control the weight of the loss for difficult classes. Through this carefully constructed loss function, the model can closely align the anatomical descriptions embedded in the text prompts with the image features by means of the contrastive learning mechanism, and focus on enhancing the attention weights of discriminative features such as the anterior / posterior horn morphology, choroid plexus distribution (class 5), and thalamus symmetry, sylvian fissure position (class 9). These fine-grained semantic information provides a quantifiable anatomical prior basis for constructing class-specific auxiliary loss functions, effectively improving the model's ability to distinguish similar structures and significantly enhancing the classification performance for difficult classes.
[0102] In this embodiment, the hierarchical adaptive feature module HAFB significantly improves the model's perception ability of the subtle differences and global semantics of medical images through a multi-scale feature extraction and adaptive fusion mechanism.
[0103] 1. Multi-scale feature extraction
[0104] In the field of medical image analysis, multi-scale feature extraction is crucial for accurately understanding the complex information in images. Features at different scales can capture multi-level information from subtle tissue structures to macroscopic organ morphology and pathological changes. HAFB is constructed based on a pre-trained CLIP visual encoder, which adopts the ViT-B / 16 architecture and consists of 12 Transformer blocks. The core of HAFB is to extract multi-scale features from different levels of the Transformer blocks to comprehensively analyze medical images. In Figure 1 it can be seen the connection relationship between the feature extraction processes at different levels and other parts of the model.
[0105] Let the input image generate an initial feature sequence X0 ∈ R (n+1)×d, where n = 196 is the number of image patches; d = 768 is the feature dimension, which determines the richness of the feature information contained in each image patch; +1 corresponds to the CLS Token (used to aggregate global semantics) and plays a key role in subsequent feature extraction and classification tasks. The feature output of the l-th layer Transformer block can be expressed as:
[0106] X l = Transformer l (X l-1 ) ∈ R (n+1)×d , l ∈ {1, 3, 5, 7, 9, 11}
[0107] HAFB extracts the CLS Token features F of each layer l = X l [0, :] ∈ R d as the hierarchical semantic representation. During the extraction process, directly select the row vector with index 0 in X l (i.e., the feature vector corresponding to the CLS Token). After being calculated by each layer of the Transformer block, this vector fuses the local and global information of the corresponding layer.
[0108] This design is based on the hierarchical abstraction characteristics of ViT, and the multi-level design of HAFB realizes the complete modeling of anatomical features from micro to macro:
[0109] Detail enhancement (shallow features, l = 1, 3): Shallow features retain the high-frequency details of the original image and avoid information loss in the deep network. This enables it to encode local anatomical details, such as tissue boundaries and textures. Taking fetal cranial ultrasound images as an example, the features of the first layer are highly sensitive to the gradient changes of the strong echo ring of the skull and can accurately locate the boundary of the cavum septum pellucidum.
[0110] Semantic abstraction (middle layer features, l = 5, 7): Middle layer features establish topological associations between anatomical landmarks through multi-head self-attention (MSA). The formula is as follows:
[0111] MSA(Q, K, V) = Concat(head1, …, head k )W O
[0112] where, d k = d / h (h is the number of heads, h = 12 in ViT-B / 16, d k = 64). W O ∈ R d×d, which is the output projection matrix. This mechanism enables the model to autonomously focus on discriminative regions (such as choroid plexus vs thalamus). Middle-level features utilize this mechanism to model organ-level morphology, such as the geometry of the lateral ventricles. In fetal cranial ultrasound images, the features at layer 5 associate the anterior horn (triangular hypoechoic region) and posterior horn (linear extension structure) of the lateral ventricles through the self-attention mechanism, quantifying their morphological differences.
[0113] Pathology perception (deep features, l = 9, 11): Deep features capture long-range dependencies through self-attention, providing high-order semantic support for complex pathology classification. This enables deep features to capture global pathology patterns, such as the spatial distribution of ventricular dilation. In fetal cranial ultrasound images, the features at layer 11 model the topological relationship between the thalamus and the lateral ventricles, enhancing the discriminative ability for anatomical variations (such as thalamic asymmetry).
[0114] Taking fetal cranial ultrasound images as an example, traditional single-scale feature extraction methods are difficult to simultaneously consider both fine structures and global information. The different-level features extracted by HAFB cooperate with each other and can analyze images from multiple dimensions. In liver ultrasound analysis, the multi-scale features respectively correspond to the liver capsule edge (shallow layer), vascular texture (middle layer), and echo uniformity (deep layer), demonstrating its cross-organ generalization ability.
[0115] In summary, by extracting multi-scale features from specific levels of the CLIP visual encoder, HAFB makes full use of the hierarchical abstraction characteristics of the ViT architecture and can effectively capture multi-level information in medical images from local anatomical details to global pathology patterns. These multi-scale features provide a rich information basis for subsequent medical image analysis tasks. Next, how to fuse these multi-scale features through an adaptive fusion module to further improve the model performance will be introduced in detail.
[0116] 2. Adaptive Fusion Module
[0117] To effectively integrate multi-scale features and enhance the model's discriminative ability for the complex semantics of medical images, HAFB designs an Adaptive Fusion Module, whose core lies in dynamically weighing the contribution degrees of different-level features to achieve progressive optimization of anatomical semantics. The schematic diagram of the overall structure of HAFB is as Figure 1 shown.
[0118] Let the current processing level be l, and the accumulated fused features be (n is the batch size, d prev is the historical feature dimension), and the current-level features be The fusion process is as follows:
[0119] First, concatenate the historical features and the current features along the channel dimension to retain the integrity of multi-scale information, obtaining (Fconcat )
[0120]
[0121] This operation avoids information loss of shallow details in deep transmission, ensuring the collaborative utilization of boundary textures (such as the strong echo ring of the skull) and global semantics (such as the morphology of the cerebral ventricle).
[0122] Next, through learnable parameters and biases the concatenated features are dimensionally reduced to obtain the transformed features (F trans )
[0123]
[0124] where d new is the unified feature dimension, reducing the interference of redundant information and improving the calculation efficiency.
[0125] To balance the importance of the historical fusion feature F trans and the current original feature F l a learnable parameter θ l ∈R is introduced, and the sigmoid function is applied to obtain the weight coefficient
[0126]
[0127] where (σ) represents the sigmoid function The weight (α l ) is restricted to the interval ([0,1]) to ensure the rationality and effectiveness of the weight.
[0128] Finally, according to the fusion weight (α l ) the (F trans ) and the current hierarchical original feature (F l ) are weighted and summed to obtain the fused feature (F new )
[0129] F new =(1 - αl)F trans +α l F l
[0130] The F new obtained through the above adaptive fusion module effectively integrates the image feature information at different levels, has stronger discriminability and semantic expression ability, and lays a solid foundation for subsequent classification decisions.
[0131] In this embodiment, during the model training stage, parameter optimization is achieved by minimizing the total loss function L totalImplementation. Based on the backpropagation algorithm, the gradient signal is backpropagated to the trainable parameter nodes of the visual encoder and the prompt word learner, and the network weights are dynamically adjusted according to the gradient descent rule, driving the model to gradually converge to the optimal state, thereby improving the accuracy of subsequent prediction tasks.
[0132] In the inference stage, the input ultrasound image is processed through feature extraction, fusion, and generation of a unified feature representation F. new . This feature is matched with the predefined anatomical category text prompt feature F. T through the cosine similarity function sim(·,·) to generate a similarity score matrix S ∈ R. m×n , where m is the number of input images and n is the total number of anatomical categories. The matrix element S. p,q represents the semantic matching degree between the p-th image and the q-th category. By applying the argmax operation row by row, the model assigns the category index with the highest similarity as the prediction result for each image. For example, for a transverse section ultrasound image of the lateral ventricle, its score for the corresponding category in the matrix is significantly higher than other categories, and the correct classification result is output after the argmax operation. To ensure the reliability of classification, this study introduces a confidence threshold mechanism: if the highest similarity score is lower than the preset threshold (0.6), the model will mark the result as "uncertain classification" or trigger an artificial review process. This mechanism effectively balances the automation efficiency and diagnostic safety, avoiding the risk of misjudgment caused by low-confidence predictions.
[0133] The present invention also proposes a fetal grade II ultrasound image classification system jointly driven by MedLP and HAFB, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned method for jointly driving the fetal grade II ultrasound image classification by MedLP and HAFB.
[0134] To comprehensively evaluate the model performance, a series of experiments were carried out, including comparison with advanced architectures such as Swin Transformer, EfficientNet, and ResNet-50, and ablation studies to analyze the contributions of each component of the model. The experimental results show that the model of the present invention performs excellently in the fetal ultrasound image classification task. The overall recognition accuracy reaches 99.3%, and it has significant advantages over the comparison models in terms of precision, recall, and F1 score. Especially in distinguishing similar anatomical structures such as the lateral ventricle section and the thalamus section, and in dealing with low-quality images with noise interference and fuzzy features, it demonstrates powerful performance. The ablation experiments fully verify the effectiveness of each component of the model. In the simulated clinical application scenario, the diagnostic accuracy of the model itself is as high as 100%, and it can serve as a powerful auxiliary tool for doctors, significantly improving the diagnostic efficiency of doctors. In summary, the present invention has made important progress in the field of fetal ultrasound image classification and has specifically made the following contributions:
[0135] 1. Propose the MedLP strategy: The MedLP strategy is proposed to address the problems of complex anatomical structures and insufficient utilization of medical knowledge in fetal ultrasound images. It generates context-aware prompts, introduces professional medical domain knowledge for fine-grained anatomical descriptions, and enhances the model's understanding and discrimination ability of image features, especially showing outstanding performance in the fine-grained distinction of similar anatomical structures.
[0136] 2. Design the Hierarchical Adaptive Feature Block (HAFB): The HAFB is designed to solve the problems of feature extraction and adaptability of existing models. It accurately captures the semantic and detailed information of ultrasound images through multi-scale feature extraction and adaptive fusion, significantly improving the classification performance.
[0137] 3. Integrate multi-dimensional optimization techniques: Integrate techniques such as data augmentation, normalization, and customized loss functions. Data augmentation expands the dataset through rotation, scaling, cropping, and brightness adjustment. Normalization improves the data quality, and the customized loss function adapts to image features, effectively reducing overfitting and enhancing the model's accuracy and generalization ability.
[0138] The above are the preferred embodiments of the present invention. All changes made in accordance with the technical solutions of the present invention that do not exceed the scope of the technical solutions of the present invention in terms of the functions and effects produced belong to the protection scope of the present invention.
Claims
1. A fetal grade II ultrasound image classification method jointly driven by MedLP and HAFB, characterized in that Specifically, it includes the following steps: S1: Collect fetal grade II ultrasound cross-sectional images of key parts, including abdominal circumference cross-section, bilateral eye cross-section, bilateral kidney sagittal section, bilateral kidney cross-section, cerebellum cross-section, lateral ventricle cross-section, median sagittal section of the face, nasolabial coronal section, spinal cord longitudinal section, thalamus cross-section; perform section marking and preprocessing on the collected ultrasound images, and randomly shuffle the order of all images and their corresponding labels for model training; S2: Based on the CLIP model, fuse the MedLP strategy and the HAFB module to construct the MedLP-HAFB-CLIP model; The MedLP strategy generates prompts that match the anatomical categories by designing a prompt learner and using learnable context vectors, so that the model can understand the anatomical features in the ultrasound images; for difficult categories, fine-grained anatomical descriptions in the medical field are introduced, and an auxiliary loss function is constructed to enhance the model's ability to distinguish difficult samples; The HAFB module extracts image features from different levels through a multi-scale feature extraction unit to capture the detailed information and semantic information in the ultrasound images; the adaptive fusion module dynamically adjusts the fusion weights of features at different levels according to the task requirements to generate a more discriminative feature representation; S3: Train the constructed MedLP-HAFB-CLIP model, and optimize the model parameters by minimizing the total loss function; S4: Input the ultrasound image into the trained MedLP-HAFB-CLIP model. The input ultrasound image generates a unified feature representation after feature extraction, fusion and processing, and is matched with the predefined anatomical category text prompt features to generate a similarity score matrix. According to the similarity score matrix, the model obtains the category index with the highest similarity for each image as the prediction result of the input ultrasound image.
2. The method for classifying fetal grade II ultrasound images co-driven by MedLP and HAFB according to claim 1, wherein The preprocessing of the collected ultrasound images specifically includes: S1.1: Uniformly process the image size and scale the shortest side of the image to the preset number of pixels; S1.2: Convert the image to RGB mode; S1.3: Convert the image to a tensor and perform normalization processing; S1.4: Crop a fixed-size area from the center of the image to ensure that the image sizes input into the model are exactly the same.
3. The fetal level II ultrasound image classification method jointly driven by MedLP and HAFB according to claim 1, wherein, The step of generating prompts that match the anatomical categories by designing a prompt learner and using learnable context vectors specifically includes the following steps: S2.1.1: Randomly initialize a set of context vectors where n ctx represents the number of context words, and d is the feature dimension; the context vectors are used as learnable parameters to adaptively adjust the prompt words to match the image features of different anatomical categories; S2.1.2: Generate an initial set of prompt words based on the set of anatomical category names \(L = \{l_0, l_1, \ldots, l_9\}\) The set of anatomical category names specifically includes: \(l_0\) is the transverse cross-section of the abdominal circumference, \(l_1\) is the transverse cross-section of both eyes, \(l_2\) is the sagittal cross-section of both kidneys, \(l_3\) is the transverse cross-section of both kidneys, \(l_4\) is the transverse cross-section of the cerebellum, \(l_5\) is the transverse cross-section of the lateral ventricles, \(l_6\) is the median sagittal cross-section of the face, \(l_7\) is the coronal cross-section of the nose and lips, \(l_8\) is the longitudinal cross-section of the spine, \(l_9\) is the transverse cross-section of the thalamus; each prompt word \(p\) i is composed of a context word prefix and a category name suffix, that is, \(p\) i = prefix + \(l\) i ; S2.1.3: Encode the initial prompt through the word embedding layer of the CLIP model to obtain an embedding representation where n tokens is the number of word segments of the prompt; divide the embedding representation into a prefix context and a suffix in three parts; during the forward propagation of the model, the prompt learner dynamically generates the final prompt embedding according to the context vector C and the fixed prefix and suffix embeddings 4. The fetal grade II ultrasound image classification method co-driven by MedLP and HAFB according to claim 1, wherein The step of introducing fine-grained anatomical descriptions in the medical field and constructing an auxiliary loss function for difficult categories to enhance the model's ability to distinguish difficult samples is specifically: S2.2.1: Medical experts define the fine-grained ultrasound image features of difficult categories based on clinical anatomical knowledge, and input the fine-grained definition content as the initial prompt into the model to enhance the model's learning ability for complex features; the difficult categories include category 5 and category 9, namely the lateral ventricle cross-section and the thalamus cross-section; S2.2.2: Construct an auxiliary loss function for Class 5 and Class 9; during the model training process, when the label of the input image is Class 5 or Class 9, the features of the input image and the corresponding prompt word features are processed separately: For the image features of category 5 and category 9 Extract the corresponding text prompt features where n 5 / 9 is the number of images in category 5 and category 9; Calculate the similarity scores between each image belonging to category 5 or category 9 and the text prompt features corresponding to category 5 and category 9 to obtain a similarity score matrix where the number of rows n of the matrix 5 / 9 corresponds to the number of images, and the number of columns is 2, representing the similarity scores between each image and the text prompt features of category 5 and category 9 respectively; Use the cross-entropy loss function to supervise the similarity score: wherein is the true label of class 5 or class 9.
5. The method for classifying fetal grade II ultrasound images co-driven by MedLP and HAFB according to claim 4, wherein The dot product similarity is used to calculate the similarity scores between each image belonging to class 5 or class 9 and the text prompt features corresponding to class 5 and class 9; first, and are normalized row by row to make the norm of each vector equal to 1, and then the similarity scores are calculated through matrix multiplication to obtain a similarity score matrix 6. The method for classifying fetal grade II ultrasound images co-driven by MedLP and HAFB according to any one of claims 4 or 5, characterized in that, Step S3 uses a selective parameter update strategy for model training, only allowing the parameters of the visual encoder and the prompt learner to be updated, and freezing the remaining parameters of the model; the total loss function includes a contrastive loss, a classification loss for overall model training, and an auxiliary loss for difficult classes; the total loss function L total is expressed as: L total = L + βL 5 / 9 = αL contrastive + (1 - α)L classification + βL 5 / 9 L = αL contrastive +(1 - α)L classification Among them, β is a hyperparameter used to control the weight of the hard class loss; L contrastive and L classification are the contrastive loss and the classification loss respectively. α is a hyperparameter used to balance the weights of the contrastive loss and the classification loss.
7. The method for classifying fetal grade II ultrasound images co-driven by MedLP and HAFB according to claim 6, wherein The contrastive loss uses the cosine similarity function, specifically as follows: For the input image \(I\in\mathbb{R}\) b×c×h×w and the corresponding label \(y\in\mathbb{Z}\) b , first, use the visual encoder of the CLIP model to extract the image feature \(F\) I \(\in\mathbb{R}\) b×d , and perform normalization, where \(b\) is the batch size, \(c\) is the number of channels, and \(h\) and \(w\) are the height and width of the image; then select the corresponding prompt embedding \(P\) final [y], use the text encoder to extract the text feature \(F\) T \(\in\mathbb{R}\) b×d , and perform normalization; the contrastive loss is used to maximize the cosine similarity between the image feature \(F\) I and the corresponding text feature \(F\) T , and the formula is: Where sim(·,·) is the cosine similarity function, used to measure the similarity between two vectors; τ is the temperature parameter, used to control the sensitivity of the model's similarity judgment during contrastive learning; the index j is used to iterate over the negative examples of all samples within the batch.
8. The fetal level II ultrasound image classification method co-driven by MedLP and HAFB according to claim 6, characterized in that, The classification loss uses the cross-entropy loss function, and the formula is: where p(y i | F I , P final ) is the probability that the model predicts the label y I given the image feature F final and the prompt embedding P i ; the index i is used to iterate over all samples in the batch.
9. The fetal grade II ultrasound image classification method jointly driven by MedLP and HAFB according to claim 1, characterized in that The HAFB module extracts image features from different levels through a multi-scale feature extraction unit to capture the detailed information and semantic information in the ultrasound image; through the adaptive fusion module, the fusion weights of features at different levels are dynamically adjusted according to the task requirements to generate a more discriminative feature representation, specifically as follows: The HAFB module is constructed based on the CLIP visual encoder, which adopts the ViT-B / 16 architecture and consists of 12 layers of Transformer blocks. The HAFB module extracts multi-scale features from the 1st, 3rd, 5th, 7th, 9th, and 11th layers of Transformer blocks and extracts the CLS Token feature F of the corresponding layer l As the hierarchical semantic representation, where F l = X l [0, :] ∈ R d and l ∈ {1, 3, 5, 7, 9, 11}, X l is the feature output of the l-th layer of Transformer blocks; Let the current processing level be l, and the current level feature be The accumulated fusion feature is where n is the batch size, and d prev is the historical feature dimension. The fusion process is as follows: First, the historical features and the current features are concatenated along the channel dimension to preserve the integrity of multi-scale information, resulting in F concat : Next, the concatenated features are dimensionally reduced by learnable parameters and biases to obtain the transformed feature F trans : where d new is the unified feature dimension; To balance the importance of the historical fusion feature F trans and the current original feature F l we introduce a learnable parameter θ l ∈R, and apply the sigmoid function to obtain the weight coefficient: Where, σ represents the sigmoid function; Finally, according to the fusion weight α l weight the F trans and the original feature F at the current level l by weighted summation to obtain the fused feature F new : F new = (1 - α l )F trans + α l F l .
10. The fetal level II ultrasound image classification method jointly driven by MedLP and HAFB according to claim 1, wherein, In step S4, the generated unified feature representation is matched with the predefined anatomical category text prompt feature F T through the cosine similarity function sim(·,·) to generate a similarity score matrix S ∈ R m×n , where m is the number of input images and n is the total number of anatomical categories, and the matrix element S p,q represents the semantic matching degree between the p-th image and the q-th category; by applying the argmax operation row by row, the model assigns the category index with the highest similarity to each image as the prediction result; a confidence threshold mechanism is introduced. If the highest similarity score is lower than the preset threshold, the model will mark the result as an uncertain classification or trigger a manual review process.
Citation Information
Cited By
Ultrasonic image classification method, device and equipment and readable storage medium
CN121121315A
An ultrasonic image classification method, device, equipment and readable storage medium
CN121121315B