Ultrasonic image quality evaluation method based on semantic and topological consistency

Through the ultrasonic image quality evaluation method with semantic and topological consistency, the spatial transformation network and text description are used to generate embedded vectors, and the model is optimized to reduce image deformation, which solves the problem of pseudo-label in the traditional method and realizes automatic evaluation of high accuracy and robustness.

CN120278965APending Publication Date: 2025-07-08HUNAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510348520.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional ultrasound image quality evaluation methods rely on pseudo-labels, which have problems with low evaluation accuracy, especially when imaging angles and operation techniques change, anatomical structure morphology leads to inaccurate pseudo-labels.

Method used

Using a method based on semantic and topological consistency, the ultrasonic image features are aligned through spatial transformation networks, embedded vectors are generated in combination with text descriptions, and the network model is optimized using a comprehensive loss function to reduce image deformation and enhance semantic information.

Benefits of technology

Automatic evaluation without manual labeling of ultrasonic image quality tags is achieved, reducing pseudo-label bias, improving evaluation accuracy, and showing robustness under a variety of imaging conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278965A_ABST
    Figure CN120278965A_ABST
Patent Text Reader

Abstract

The invention relates to an ultrasonic image quality evaluation method and device based on semantic and topological consistency, computer equipment and a storage medium, and the method comprises the steps: obtaining ultrasonic image data which comprises a reference ultrasonic image set and a to-be-aligned ultrasonic image set; performing model training according to ultrasonic images in a reference ultrasonic image set and a to-be-aligned ultrasonic image set, aligning feature maps by using a spatial transformation network through a network model, generating an embedded vector in combination with text description, and combining the feature maps with the embedded vector through a cross-modal attention mechanism. And performing parameter optimization on the model by using a comprehensive loss function combining a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function and a reconstruction loss function to obtain a trained network model. The ultrasonic image quality can be automatically evaluated, manual labeling of ultrasonic image quality labels is not needed, and the evaluation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image processing, and particularly to an ultrasonic image quality evaluation method based on semantic and topological consistency. Background Art

[0002] Ultrasonic imaging is a non-invasive and real-time medical imaging technology, which is widely used in clinical diagnosis and medical research. However, the quality of ultrasonic images is affected by various factors, such as equipment performance, operator skills, and imaging environment. High-quality ultrasonic images are crucial for accurate clinical diagnosis and medical research.

[0003] The traditional unsupervised ultrasonic image quality assessment method based on pseudo-labels (CRL-UIQA) realizes the automatic assessment of ultrasonic image quality through pseudo-label generation and joint learning of consistency and correlation, without relying on the quality labels manually annotated for ultrasonic images. However, this method also has its drawbacks: the pseudo-labels are generated based on the feature distribution distance between standard ultrasonic images and the ultrasonic images to be evaluated. Due to the possible presence of noise and outliers in the feature space, as well as changes in probe angle or operation techniques, such as scaling, rotation, and slight displacement, significant morphological changes may occur in anatomical structures. The above reasons will generate inaccurate pseudo-labels, resulting in deviations in the quality evaluation of ultrasonic images and the disadvantage of low evaluation accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide an ultrasonic image quality evaluation method, device, computer device, and storage medium based on semantic and topological consistency that can improve the evaluation accuracy for the above problems.

[0005] In a first aspect, the present application provides an ultrasonic image quality evaluation method based on semantic and topological consistency, including:

[0006] Obtain ultrasonic image data, where the ultrasonic image data includes a reference ultrasonic image set and an ultrasonic image set to be aligned;

[0007] Perform model training based on the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, and optimize the parameters of the model in combination with a comprehensive loss function to obtain a trained network model; the comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function;

[0008] Input the ultrasonic image to be evaluated into the network model to obtain a quality score for evaluating the quality of the ultrasonic image;

[0009] Among them, the network model includes:

[0010] An image input layer that receives ultrasound images from a reference ultrasound image set and a to-be-aligned ultrasound image set;

[0011] A feature extraction layer that extracts feature representations of ultrasound images based on a deep convolutional neural network to obtain feature maps of the ultrasound images;

[0012] An alignment network layer that performs feature alignment on different ultrasound images of the same anatomical structure through a spatial transformation network to obtain aligned feature maps;

[0013] A text guidance layer that uses text descriptions to parse the current ultrasound image to generate embedding vectors, and combines the aligned feature maps with the embedding vectors through a cross-modal attention mechanism to obtain updated feature maps;

[0014] A quality evaluation layer that uses a quality scoring formula to perform quality scoring on the updated feature maps; the quality scoring formula is obtained according to the negative cosine similarity loss function, the normalized cross-correlation loss function, and the smoothness loss function.

[0015] In one embodiment, model training is performed based on the ultrasound images in the reference ultrasound image set and the to-be-aligned ultrasound image set, and the parameters of the model are optimized by combining a comprehensive loss function to obtain a trained network model, including:

[0016] Extract feature maps from the ultrasound images in the reference ultrasound image set and the to-be-aligned ultrasound image set, transform the feature maps from the source space to the target space, use the negative cosine similarity loss function and the normalized cross-correlation loss function to perform feature alignment on the extracted features in the target space, repeat feature extraction and feature alignment a preset number of times, and retain the feature maps in the source space after each alignment;

[0017] Use the text encoder pre-trained by the CLIP model to extract the embedding vectors of the text descriptions. After splicing the feature maps in the source space after multiple alignments, use the cross-modal attention mechanism to fuse the spliced feature maps with the embedding vectors to obtain updated feature maps;

[0018] Reconstruct the image from the updated feature map through a reconstructed image decoder, compare the reconstructed image with the original updated feature map, calculate the reconstruction loss using the mean square error loss function, and then optimize the parameters of the model using a comprehensive loss function composed of the negative cosine similarity loss function, the normalized cross-correlation loss function, the smoothness loss function, and the reconstruction loss function to obtain a trained network model.

[0019] In one embodiment, transforming the feature map from the source space to the target space includes:

[0020]

[0021] Among them, and are the coordinates of the feature map in the source space s, representing the position of the i-th feature point before transformation, and are the coordinates of the feature map in the target space t, representing the position of the i-th feature point after transformation; Θ represents the affine transformation matrix, and θ 11 -θ 23 are the respective parameters in the affine transformation matrix.

[0022] In one embodiment, the negative cosine similarity loss function is:

[0023]

[0024] Among them, f i A,t is the i-th layer feature of the aligned feature map A in the target space t, and f i B,t is the i-th layer feature of the aligned feature map B in the target space t. θ represents the spatial transformation network, and ||f i A,t || 2 and ||f i B,t || 2 respectively represent the L2 norms of the vectors f i A,t and f i B,t ;

[0025] The normalized cross-correlation loss function is:

[0026]

[0027] Among them, f i A,t (p) and f i B,t (p) are the feature values at the position p of the i-th layer of the feature maps A and B, is the predicted feature value at the position p of the i-th layer of the feature maps A and B, and Ω i represents the set of positions of the i-th layer.

[0028] In one embodiment, the smoothness loss function includes:

[0029]

[0030] Among them, is the gradient operator, and ψ(p i ) is the displacement field at the position p;

[0031] The comprehensive loss function is as follows:

[0032] L total = L sim + L NCC + L smooth + L rec

[0033] where L rec is the mean squared error loss function.

[0034] In one embodiment, a cross-modal attention mechanism is used to fuse the spliced feature map and the embedding vector, including:

[0035]

[0036] where f b,t , are respectively the feature map in the target space after splicing and the feature map after feature fusion, e text is the embedding vector of the text description, (W k e text ) · represents the transpose operation of W k e text , W q , W k , W v are linear projection matrices respectively used for query, key and value, and d k is the scaling factor.

[0037] In one embodiment, the quality scoring formula is:

[0038] Q score = 1 - (α·φ(sim) + β·φ(NCC) + γ·φ(smooth))

[0039] where α, β and γ are weight coefficients, and φ(m) represents the normalized loss value, and the calculation formula is: m belongs to {sim, NCC, smooth}.

[0040] In a second aspect, the present application provides an ultrasonic image quality evaluation device based on semantic and topological consistency, including:

[0041] A data acquisition module, configured to acquire ultrasonic image data, where the ultrasonic image data includes a reference ultrasonic image set and a to-be-aligned ultrasonic image set;

[0042] A model training module, which is used to perform model training based on the ultrasound images in the reference ultrasound image set and the ultrasound images to be aligned, and optimize the parameters of the model in combination with a comprehensive loss function to obtain a trained network model; the comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function;

[0043] An image evaluation module, which is used to input the ultrasound image to be evaluated into the network model to obtain a quality score for evaluating the quality of the ultrasound image;

[0044] Wherein, the network model includes:

[0045] An image input layer, which receives the ultrasound images in the reference ultrasound image set and the ultrasound images to be aligned;

[0046] A feature extraction layer, which extracts the feature representation of the ultrasound image based on a deep convolutional neural network to obtain a feature map of the ultrasound image;

[0047] An alignment network layer, which aligns the features of different ultrasound images of the same anatomical structure through a spatial transformation network to obtain an aligned feature map;

[0048] A text guidance layer, which uses text description to parse the current ultrasound image to generate an embedding vector, and combines the aligned feature map with the embedding vector through a cross-modal attention mechanism to obtain an updated feature map;

[0049] A quality evaluation layer, which uses a quality scoring formula to perform quality scoring on the updated feature map; the quality scoring formula is obtained according to the negative cosine similarity loss function, the normalized cross-correlation loss function, and the smoothness loss function.

[0050] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0051] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0052] The above ultrasonic image quality evaluation method, device, computer device and storage medium based on semantic and topological consistency train a model according to the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, and optimize the parameters of the model in combination with a comprehensive loss function to obtain a trained network model. The comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function. The spatial transformation network is used by the network model to align the feature maps, reduce the deformation of ultrasonic images caused by changes in imaging angles and operation techniques, generate embedding vectors in combination with text descriptions, and combine the feature maps with the embedding vectors through a cross-modal attention mechanism to enhance the semantic information of the feature maps. The parameters of the model are optimized using a comprehensive loss function that combines a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function to obtain a trained network model. The ultrasonic image to be evaluated is input into the trained network model, and a quality score for evaluating the quality of the ultrasonic image can be obtained, which can automatically evaluate the quality of ultrasonic images, without the need to manually label the quality labels of ultrasonic images, and can effectively reduce pseudo-label bias and improve the evaluation accuracy. Description of the Drawings

[0053] Figure 1 It is a flowchart of the ultrasonic image quality evaluation method based on semantic and topological consistency in one embodiment;

[0054] Figure 2 It is a flowchart of training a model according to the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, and optimizing the parameters of the model in combination with a comprehensive loss function to obtain a trained network model in one embodiment;

[0055] Figure 3 It is a structural block diagram of the ultrasonic image quality evaluation device based on semantic and topological consistency in one embodiment;

[0056] Figure 4 It is an internal structure diagram of a computer device in one embodiment. Detailed Embodiments

[0057] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0058] Currently, general ultrasound image quality assessment methods generally use deep neural network architectures to automate and improve the accuracy of ultrasound image quality assessment, such as using Faster R-CNN, ResNet18, and bilinear convolutional neural networks. These methods have the following problems: First, a large amount of manual annotation work is required, which is time-consuming and prone to introducing human errors; second, the global geometric relationships of the images cannot be comprehensively understood, resulting in local detection errors and affecting the accuracy of the evaluation results.

[0059] Traditional unsupervised ultrasound image quality assessment methods based on pseudo-labels generate pseudo-labels based on the feature distribution distance between standard ultrasound images and the ultrasound images to be evaluated. Due to the possible presence of noise and outliers in the feature space, as well as changes in probe angle or operation techniques, such as scaling, rotation, and slight displacement, significant morphological changes may occur in anatomical structures. The above reasons cause the system to generate inaccurate pseudo-labels, resulting in deviations in the quality assessment of ultrasound images. Given the limitations of the existing technology, the key issues that need to be addressed in ultrasound image quality assessment methods are to reduce the dependence on manual annotation, enhance the understanding of global structures and context geometric relationships, and enhance the robustness to spatial deformation.

[0060] Based on this, the present application provides a method for evaluating the quality of ultrasound images based on semantic and topological consistency. The spatial transformation network is used by the network model to align the feature maps, reducing the deformation of ultrasound images caused by imaging angle and operation technique changes, and generating embedding vectors in combination with text descriptions. The feature maps and the embedding vectors are combined through a cross-modal attention mechanism to enhance the semantic information of the feature maps. The model is optimized by using a comprehensive loss function that combines a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function to obtain a trained network model. The ultrasound image to be evaluated is input into the trained network model, and the quality score of the ultrasound image is indirectly obtained according to the obtained comprehensive loss function to evaluate the quality of the ultrasound image. By integrating topological consistency priors into ultrasound image similarity measurements, the quality of ultrasound images can be automatically evaluated. This method does not rely on the quality labels manually annotated for ultrasound images, and can effectively solve the problem of image spatial deformation caused by imaging condition changes, and enhance the robustness of the model through text-guided cross-modal reconstruction.

[0061] In one embodiment, as Figure 1 shown, a method for evaluating the quality of ultrasound images based on semantic and topological consistency is provided, including:

[0062] Step S110: Obtain ultrasonic image data, which includes a reference ultrasonic image set and an ultrasonic image set to be aligned. Specifically, ultrasonic videos of second- and third-trimester fetal ultrasounds from multiple ultrasonic devices can be collected, and various anatomical slices can be extracted from them, such as upper abdominal, four-chamber heart, double kidney, and facial slices of the fetus. A small number of high-quality standard ultrasonic slices (reference ultrasonic images) are selected to obtain the reference ultrasonic image set, and the remaining images are divided into a training set and a test set. The training set serves as the ultrasonic image set to be aligned for model training, and the test set contains ultrasonic images to be evaluated.

[0063] Step S120: Train the model based on the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, and optimize the parameters of the model by combining a comprehensive loss function to obtain a trained network model. The comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function. Among them, the network model includes:

[0064] An image input layer that receives the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned.

[0065] A feature extraction layer that extracts the feature representation of the ultrasonic image based on a deep convolutional neural network to obtain the feature map of the ultrasonic image.

[0066] An alignment network layer that aligns the features of different ultrasonic images of the same anatomical structure through a spatial transformation network to obtain an aligned feature map;

[0067] A text-guided layer that uses text descriptions to parse the current ultrasonic image to generate an embedding vector, and combines the aligned feature map with the embedding vector through a cross-modal attention mechanism to obtain an updated feature map.

[0068] A quality evaluation layer that uses a quality scoring formula to score the updated feature map; the quality scoring formula is obtained based on the negative cosine similarity loss function, the normalized cross-correlation loss function, and the smoothness loss function.

[0069] Specifically, the input images of the network model are divided into two groups: Fixed Images and Moving Images. The Fixed Images are high-quality standard cross-sectional images (i.e., reference ultrasound images) selected by clinicians, and the Moving Images are other images of the same anatomical structure (i.e., ultrasound images to be aligned). The size of the input images is uniformly scaled to 224×224 for image processing. The feature extraction layer extracts the feature representation of the images based on a deep convolutional neural network. The alignment network layer contains several (e.g., three) stages of spatial transformation network blocks for aligning the semantic and topological features of the same anatomical structure, alleviating the spatial deformation caused by changes in imaging angles and operation techniques. The text guidance layer uses text descriptions to parse the current slice, assisting the network model in semantic alignment and enhancing the semantic learning ability of the network model through cross-modal reconstruction. The quality evaluation layer scores the quality of the query images based on anatomical consistency, perceptual consistency, and topological consistency.

[0070] Specifically, as Figure 2 shown, step S120 includes steps S122 to S126.

[0071] Step S122: Extract feature maps from the ultrasound images in the reference ultrasound image set and the ultrasound images to be aligned. Transform the feature maps from the source space to the target space. Use the negative cosine similarity loss function and the normalized cross-correlation loss function to perform feature alignment on the features extracted in the target space. Repeat the feature extraction and feature alignment a preset number of times, and retain the feature maps in the source space after each alignment.

[0072] During the training process of the network model, extract the feature representation from the reference ultrasound image set and the ultrasound images to be aligned. Specifically, a convolutional neural network can be used to encode the input ultrasound images to generate feature maps. Use the spatial transformation network to predict the affine transformation matrix and map the features of the ultrasound images to be aligned to the feature space of the reference ultrasound images. Using the spatial transformation network to align the feature maps can reduce the deformation of ultrasound images caused by changes in imaging angles and operation techniques. In addition, through multi-scale feature alignment, gradually refine the alignment accuracy of the features, and ensure the consistency of the global and local structures. Among them, the global consistency between the source features and the target features is quantitatively evaluated through the negative cosine similarity loss function; the consistency of the image in the local structure is measured through the normalized cross-correlation loss function; the affine transformation matrix is constrained through the smoothness loss function to ensure the continuity and smoothness of the transformation and prevent excessive deformation.

[0073] Step S124: Use the text encoder pre-trained by the CLIP (Contrastive Language-Image Pre-training) model to extract the embedding vector of the text description. After splicing the feature maps in the source space after multiple alignments, use the cross-modal attention mechanism to fuse the spliced feature maps with the embedding vector to obtain an updated feature map. Among them, taking the spatial transformation network block with three stages in the alignment network layer as an example, the feature maps in the source space after three alignments are spliced, and the cross-modal attention mechanism is used to fuse the spliced feature maps with the embedding vector to obtain an updated feature map.

[0074] The CLIP model converts the image label into image text description information and is a language model for image-text matching and training based on contrastive learning. To further enhance the semantic learning ability of the network model, after the feature alignment between the source ultrasound image and the ultrasound image to be aligned, a text-guided cross-modal reconstruction mechanism is added to further enhance the features: use the pre-trained CLIP text encoder to extract the embedding vector of the text description, fuse the text features with the image features through the attention mechanism, calculate the attention weights between the feature map and the embedding vector, and update the feature map according to the attention weights.

[0075] Step S126: Reconstruct the image from the updated feature map through the reconstruction image decoder, compare the reconstructed image with the original updated feature map, calculate the reconstruction loss using the mean square error loss function, and then optimize the parameters of the model using the comprehensive loss function composed of the negative cosine similarity loss function, the normalized cross-correlation loss function, the smoothness loss function, and the reconstruction loss function to obtain a trained network model.

[0076] Generate a reconstructed image through the reconstruction image decoder and compare it with the original image to calculate the reconstruction error as the reconstruction loss function. By defining a comprehensive loss function, including the negative cosine similarity loss function, the normalized cross-correlation loss function, the smoothness loss function, and the reconstruction loss function, the model is optimized. In the comprehensive loss function, the negative cosine similarity loss generated by using the text description to cross-modally enhance the semantic information of the feature map is used to measure the semantic consistency of the images by calculating the cosine similarity between the features; the normalized cross-correlation loss generated during the alignment of the feature maps measures the local structural consistency of the images; the smooth loss generated by the spatial transformation network to transform the feature map of the input ultrasound image ensures the smoothness of the image in geometric transformation; the reconstruction loss generated by image reconstruction measures the visual perception quality of the image by calculating the mean square error between the reconstructed image and the original image.

[0077] Specifically, during the training process of the network model, each group of ultrasound images input into the network model first passes through the feature extraction layer. ResNet18 is used as the backbone network to extract rough global features, and then part of the layers of ResNet18 are used to extract features. The affine transformation matrix is predicted through the fully connected layer. The affine transformation matrix is a 2×3 matrix used to represent geometric transformations in the two-dimensional space, such as translation, rotation, and scaling. The affine matrix output at the initial stage of training is [[1,0,0],[0,1,0]], indicating no spatial transformation. After predicting the affine transformation matrix through the fully connected layer, the feature maps of the fixed image and the moving image are transformed from the source space to the target space by means of the affine transformation matrix plus the grid generator to generate a sampling grid. The differentiable sampler samples from the feature maps of the fixed image and the moving image according to the sampling grid generated by the grid generator to generate the transformed feature maps. In this embodiment, the principle of transforming the feature map from the source space to the target space through the affine transformation matrix is as follows:

[0078]

[0079] Among them, and are the coordinates of the feature map in the source space s, representing the position of the i-th feature point before transformation, and are the coordinates of the feature map in the target space t, representing the position of the i-th feature point after transformation; Θ represents the affine transformation matrix, and θ 11 -θ 23 are the respective parameters in the affine transformation matrix. The features of the fixed image and the moving image extracted in the target space are aligned using the negative cosine similarity loss function and the normalized cross-correlation loss function. The aligned feature maps are reused twice to extract the feature layers, extracting medium-scale features and local detail features respectively. After each feature extraction, the negative cosine similarity loss function and the normalized cross-correlation loss function are used for alignment, and the feature maps in the source space after each alignment are retained.

[0080] The negative cosine similarity loss function is used to evaluate the global consistency between the aligned features. The formula of the negative cosine similarity loss function is:

[0081]

[0082] Among them, the fixed image is called A, the moving image is called B, and f i A,t is the i-th layer feature of the aligned feature map A in the target space t, and f i B,tThe i-th layer feature of the aligned feature map B in the target space t. In this embodiment, the network performs 3 alignments, so i is 3, θ represents the spatial transformation network, ||f i A,t || 2 and ||f i B,t || 2 respectively represent the L2 norms of the vectors f i A,t and f i B,t .

[0083] Furthermore, the normalized cross-correlation loss function is used to evaluate the consistency of the local structure. The formula of the normalized cross-correlation loss function is:

[0084]

[0085] where f i A,t (p) and f i B,t (p) are the feature values at position p in the i-th layer of feature map A and feature map B, is the predicted feature value at position p in the i-th layer of feature map A and feature map B, and Ω i represents the set of positions in the i-th layer.

[0086] In addition, the smoothness loss function is used to regularize the affine transformation matrix to ensure the continuity and smoothness of the transformation. The formula of the smoothness loss function is:

[0087]

[0088] where is the gradient operator, and ψ(p i ) is the displacement field at position p, that is, the feature map obtained through model training.

[0089] Furthermore, in order to enhance the semantic understanding ability of the network model, the text encoder pre-trained by the CLIP model is used to extract the features of the text description. The feature maps in the source space extracted in the above three stages are concatenated together, and then the cross-modal attention mechanism is used to fuse the feature maps with the text features. For each feature layer, the attention weights are calculated and the features are updated, and the obtained new features are concatenated with the above feature maps. The cross-modal attention mechanism is used to fuse the concatenated feature maps with the embedding vector, including:

[0090]

[0091] where f b,t and They are respectively the feature map of the spliced moving image in the target space and the feature map after feature fusion, e text is the embedding vector of the text description, (W k e text ) · represents the transpose operation of W k e text ; W q , W k , W v are the linear projection matrices for query, key, and value respectively, and their initial values are all random. The final weights are obtained when the loss value of the comprehensive loss function is minimized during model training. d k is the scaling factor, which is used to improve the numerical stability in the attention mechanism.

[0092] Use the updated feature map to reconstruct the image through a reconstructed image decoder, compare the reconstructed image with the original image, and calculate the reconstruction loss using the mean square error loss function. The training of the entire model relies on the comprehensive loss function composed of the negative cosine similarity loss function, the normalized cross-correlation loss function, the smoothness loss function, and the reconstruction loss function for optimization. The comprehensive loss function is defined as:

[0093] L total = L sim + L NCC + L smooth + L rec

[0094] where L rec is the mean square error loss function.

[0095] Step S130: Input the ultrasonic image to be evaluated into the network model to obtain the quality score for quality evaluation of the ultrasonic image. After training the network model, input the ultrasonic image to be evaluated into the model, and indirectly obtain the quality score of the ultrasonic image according to the obtained comprehensive loss function to conduct quality evaluation on the ultrasonic image.

[0096] Specifically, after training, use the network model structure except the reconstructed image decoder for ultrasonic image quality evaluation. Input the ultrasonic image to be evaluated into the trained network, and also use a small number of optimal standard ultrasonic images as reference ultrasonic images. Align the features of different ultrasonic images of the same anatomical structure through the spatial transformation network to obtain the aligned feature maps. Then generate the embedding vector using the text description, combine the feature maps with the embedding vector through the cross-modal attention mechanism, and evaluate the image quality based on the aligned ultrasonic images. The quality scoring formula used is:

[0097] Q score = 1 - (α·φ(sim) + β·φ(NCC) + γ·φ(smooth))

[0098] Among them, α, β, and γ are weight coefficients used to balance the influence of different loss terms, and they can all be set to 1 / 3; φ(m) represents the normalized loss value, and the calculation formula is: m belongs to {sim, NCC, smooth}, that is, the maximum loss function value and the minimum loss function value in the model training history are used to calculate with the currently obtained loss value. Through the negative cosine similarity loss value, the normalized cross-correlation loss value, and the smoothness loss value generated by normalization, the semantic consistency, local structure correspondence, and smoothness of the ultrasound image are evaluated, so as to indirectly obtain the quality evaluation score of the ultrasound image.

[0099] The above ultrasound image quality evaluation method based on semantic and topological consistency uses a small number of optimal standard ultrasound images as reference ultrasound images, aligns the features of different ultrasound images on the same plane through a spatial transformation network, introduces a CLIP text encoder to assist the model in semantic understanding of ultrasound images, and evaluates the quality of ultrasound images based on anatomical consistency and perceptual consistency; using the above method to evaluate the quality of ultrasound images can automatically evaluate the quality of ultrasound images, without manual annotation of ultrasound images, can effectively reduce pseudo-label bias, and improve the evaluation accuracy; evaluate the quality of ultrasound images in a clinical environment and is highly consistent with clinical judgments; shows robustness under various imaging conditions and performs better in terms of accuracy and robustness compared with traditional methods.

[0100] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0101] Based on the same inventive concept, an embodiment of the present application further provides an ultrasonic image quality evaluation device based on semantic and topological consistency for implementing the above-mentioned ultrasonic image quality evaluation method based on semantic and topological consistency. The implementation solutions provided by this device to solve problems are similar to the implementation solutions recorded in the above method. Therefore, the specific limitations in one or more embodiments of the ultrasonic image quality evaluation device based on semantic and topological consistency provided below can refer to the limitations on the ultrasonic image quality evaluation method based on semantic and topological consistency in the above text, and will not be elaborated here.

[0102] In one embodiment, as Figure 3 shown, an ultrasonic image quality evaluation device based on semantic and topological consistency is further provided, including: a data acquisition module 110, a model training module 120, and an image evaluation module 130, where:

[0103] The data acquisition module 110 is configured to acquire ultrasonic image data, and the ultrasonic image data includes a reference ultrasonic image set and a to-be-aligned ultrasonic image set.

[0104] The model training module 120 is configured to perform model training based on the ultrasonic images in the reference ultrasonic image set and the to-be-aligned ultrasonic image set, and optimize the parameters of the model in combination with a comprehensive loss function to obtain a trained network model; the comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function.

[0105] The image evaluation module 130 is configured to input the ultrasonic image to be evaluated into the network model to obtain a quality score for evaluating the quality of the ultrasonic image.

[0106] Among them, the network model includes:

[0107] An image input layer that receives the ultrasonic images in the reference ultrasonic image set and the to-be-aligned ultrasonic image set.

[0108] A feature extraction layer that extracts feature representations of ultrasonic images based on a deep convolutional neural network to obtain a feature map of the ultrasonic images.

[0109] An alignment network layer that aligns features of different ultrasonic images of the same anatomical structure through a spatial transformation network to obtain an aligned feature map.

[0110] A text guidance layer that uses text descriptions to parse the current ultrasonic image to generate an embedding vector, and combines the aligned feature map with the embedding vector through a cross-modal attention mechanism to obtain an updated feature map.

[0111] The quality evaluation layer uses a quality scoring formula to perform quality scoring on the updated feature map; the quality scoring formula is obtained based on the negative cosine similarity loss function, the normalized cross-correlation loss function, and the smoothness loss function.

[0112] In one embodiment, the model training module 120 is configured to extract feature maps from the ultrasound images in the reference ultrasound image set and the ultrasound image set to be aligned, transform the feature maps from the source space to the target space, and use the negative cosine similarity loss function and the normalized cross-correlation loss function to perform feature alignment on the extracted features in the target space. Feature extraction and feature alignment are repeated a preset number of times, and the feature maps in the source space after each alignment are retained; the embedding vectors of the text descriptions are extracted using the text encoder pre-trained by the CLIP model. After splicing the feature maps in the source space after multiple alignments, the spliced feature maps and the embedding vectors are fused using a cross-modal attention mechanism to obtain updated feature maps; the updated feature maps are reconstructed into images through a reconstruction image decoder, and the reconstructed images are compared with the original updated feature maps. The mean squared error loss function is used to calculate the reconstruction loss, and then the model is optimized for parameters using a comprehensive loss function composed of the negative cosine similarity loss function, the normalized cross-correlation loss function, the smoothness loss function, and the reconstruction loss function to obtain a trained network model.

[0113] Each module in the above ultrasonic image quality evaluation device based on semantic and topological consistency can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0114] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 4As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an ultrasonic image quality evaluation method based on semantic and topological consistency. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, a touchpad, or a mouse, etc.

[0115] Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0116] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0117] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0118] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, the steps in the above method embodiments are implemented.

[0119] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0120] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0122] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. An ultrasonic image quality evaluation method based on semantic and topological consistency, characterized in that, Including: Obtain ultrasonic image data, where the ultrasonic image data includes a reference ultrasonic image set and an ultrasonic image set to be aligned; Perform model training based on the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, and optimize the model parameters in combination with a comprehensive loss function to obtain a trained network model; the comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function; Input the ultrasonic image to be evaluated into the network model to obtain a quality score for evaluating the quality of the ultrasonic image; Wherein, the network model includes: An image input layer that receives the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned; A feature extraction layer that extracts the feature representation of the ultrasonic image based on a deep convolutional neural network to obtain a feature map of the ultrasonic image; An alignment network layer that aligns the features of different ultrasonic images of the same anatomical structure through a spatial transformation network to obtain an aligned feature map; A text guidance layer that uses text descriptions to parse the current ultrasonic image to generate an embedding vector, and combines the aligned feature map with the embedding vector through a cross-modal attention mechanism to obtain an updated feature map; A quality evaluation layer that uses a quality scoring formula to perform quality scoring on the updated feature map; the quality scoring formula is obtained according to the negative cosine similarity loss function, the normalized cross-correlation loss function, and the smoothness loss function.

2. The method according to claim 1, wherein Performing model training based on the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, and optimizing the model parameters in combination with a comprehensive loss function to obtain a trained network model, including: Extract feature maps from the ultrasonic images in the reference ultrasonic image set and the ultrasonic image set to be aligned, transform the feature maps from the source space to the target space, and use the negative cosine similarity loss function and the normalized cross-correlation loss function to perform feature alignment on the features in the target space. Repeat the feature extraction and feature alignment a preset number of times, and retain the feature maps in the source space after each alignment; Use the text encoder pre-trained by the CLIP model to extract the embedding vector of the text description. After splicing the feature maps in the source space after multiple alignments, use the cross-modal attention mechanism to fuse the spliced feature maps with the embedding vector to obtain an updated feature map; Reconstruct the image by passing the updated feature map through a reconstructed image decoder, compare the reconstructed image with the original updated feature map, calculate the reconstruction loss using the mean square error loss function, and then optimize the model parameters using the comprehensive loss function composed of the negative cosine similarity loss function, the normalized cross-correlation loss function, the smoothness loss function, and the reconstruction loss function to obtain a trained network model.

3. The method according to claim 2, characterized in that, Transforming the feature map from the source space to the target space includes: Among them, and are the coordinates of the feature map in the source space s, representing the position of the i-th feature point before transformation, and are the coordinates of the feature map in the target space t, representing the position of the i-th feature point after transformation; Θ represents the affine transformation matrix, and θ 11 -θ 23 are the respective parameters in the affine transformation matrix.

4. The method according to claim 2, wherein The negative cosine similarity loss function is: Among them, f i A,t is the feature of the aligned feature map A at the i-th layer in the target space t, and f i B,t is the feature of the aligned feature map B at the i-th layer in the target space t. θ represents the spatial transformation network, and ||f i A,t || 2 and ||f i B,t || 2 represent the L2 norms of the vectors f i A,t , f i B,t respectively; The normalized cross-correlation loss function is: Among them, f i A,t (p), f i B,t (p) is the feature value at position p of feature map A and feature map B in the i-th layer, is the predicted feature value at position p of feature map A and feature map B in the i-th layer, Ω i represents the set of positions in the i-th layer.

5. The method according to claim 4, wherein The smoothness loss function is: Among them, is the gradient operator, and ψ(p i ) is the displacement field at position p; The comprehensive loss function is: L total = L sim + L NCC + L smooth + L rec Among them, L rec is the mean squared error loss function.

6. The method according to claim 2, characterized in that Using the cross-modal attention mechanism to fuse the spliced feature map with the embedding vector includes: Among them, f b,t and are the feature maps in the target space after splicing and the feature maps after feature fusion respectively, e text is the embedding vector of the text description, (W k e text ) · represents the transpose operation of W k e text , W q , W k , W v are the linear projection matrices for query, key and value respectively, d k is the scaling factor.

7. The method according to any one of claims 1-6, characterized in that The quality scoring formula is: Q score = 1 - (α·φ(sim) + β·φ(NCC) + γ·φ(smooth)) where α, β, and γ are weight coefficients, and φ(m) represents the normalized loss value, with the calculation formula being: m belongs to {sim, NCC, smooth}.

8. An ultrasonic image quality evaluation device based on semantic and topological consistency, characterized in that, Including: A data acquisition module for acquiring ultrasonic image data, where the ultrasonic image data includes a reference ultrasonic image set and a to-be-aligned ultrasonic image set; A model training module for training a model based on the ultrasonic images in the reference ultrasonic image set and the to-be-aligned ultrasonic image set, and optimizing the parameters of the model in combination with a comprehensive loss function to obtain a trained network model; the comprehensive loss function includes a negative cosine similarity loss function, a normalized cross-correlation loss function, a smoothness loss function, and a reconstruction loss function; An image evaluation module for inputting an ultrasonic image to be evaluated into the network model to obtain a quality score for quality evaluation of the ultrasonic image; Among them, the network model includes: An image input layer for receiving the ultrasonic images in the reference ultrasonic image set and the to-be-aligned ultrasonic image set; A feature extraction layer for extracting a feature representation of the ultrasonic image based on a deep convolutional neural network to obtain a feature map of the ultrasonic image; An alignment network layer for performing feature alignment on different ultrasonic images of the same anatomical structure through a spatial transformation network to obtain an aligned feature map; A text guidance layer for parsing the current ultrasonic image using a text description to generate an embedding vector, and combining the aligned feature map with the embedding vector through a cross-modal attention mechanism to obtain an updated feature map; A quality evaluation layer for performing quality scoring on the updated feature map using a quality scoring formula; the quality scoring formula is obtained according to the negative cosine similarity loss function, the normalized cross-correlation loss function, and the smoothness loss function.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Product evaluation method and system and storage medium

    CN120997620A