Ultrasonic contrast video processing method and device, electronic equipment and storage medium

By converting ultrasound contrast imaging videos into projection images and utilizing time-series attention extraction networks and feature extraction networks, the shortcomings of hardware cost and inference time in existing technologies are overcome, achieving efficient utilization of ultrasound contrast imaging video features and accurate classification diagnosis.

CN121033720APending Publication Date: 2025-11-28CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510954565.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing ultrasound contrast imaging video feature extraction methods are insufficient in balancing hardware cost and inference time, making it difficult to efficiently utilize necessary feature information.

Method used

The ultrasound contrast imaging video is converted into a projection image. Temporal attention information is extracted through a time series attention extraction network, and the temporal attention information guides the feature extraction network to extract ultrasound contrast imaging features. Finally, the images are input into a classification model for diagnosis.

Benefits of technology

With limited hardware costs and inference time, the accuracy and computational efficiency of feature extraction were improved, resulting in more accurate ultrasound contrast video classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121033720A_ABST
    Figure CN121033720A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an ultrasonic contrast video processing method and device, electronic equipment and a storage medium, and the method comprises the steps: converting an ultrasonic contrast video into a projection image according to the brightness of each frame of image in the ultrasonic contrast video; extracting time attention information of the projection image through a time sequence attention extraction network; guiding a feature extraction network through the time attention information, and extracting ultrasound contrast features of the projection image; and inputting the ultrasound contrast features into a classification model, and outputting a classification result of the ultrasound contrast video. Therefore, according to the embodiment of the invention, the necessary feature information of the ultrasonic contrast can be efficiently utilized under the condition of considering limited hardware cost and reasoning time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to an ultrasound contrast video processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] Contrast-enhanced Ultrasound (CEUS) is an extension of traditional ultrasound imaging, which can identify the perfusion of blood in the target tissue by injecting contrast agents into the vein. In clinical practice, it can significantly improve the diagnostic accuracy, sensitivity and specificity. At present, ultrasound contrast has been widely used in the examination of breast, thyroid, liver, gallbladder, pancreas, spleen, kidney and other parts.

[0003] At present, the diagnosis result of ultrasound contrast mainly depends on the artificial analysis of doctors, but due to the personal experience and professional level of doctors, there may be subjective and misjudgment risks. Deep learning is considered to be one of the important means to solve such problems, which can gradually extract effective features of ultrasound contrast by using its processing and learning ability, and finally realize automatic diagnosis and other common tasks.

[0004] Among them, the complete ultrasound contrast is mostly video data, and the existing feature extraction methods are mainly divided into the following three categories:

[0005] The first category is to extract a specific frame of ultrasound contrast image by a certain method such as using domain knowledge, and then use a two-dimensional convolutional neural network (2D-CNN) to extract features. This method may lose important time sequence information and dynamic change information, affecting the accuracy of the final diagnosis, and may need domain knowledge to determine which frame is the most representative.

[0006] The second category is to input each frame of ultrasound contrast image into a two-dimensional convolutional neural network in turn for feature extraction, and may combine time series analysis technology to analyze the dynamic changes between frames. Among them, the second category method has high computational complexity and slow processing speed, and the two-dimensional convolutional neural network itself is not good at processing time continuity information.

[0007] The third type is to directly use a three-dimensional convolutional neural network (3D-CNN) to directly extract features from several frames or all frames of the contrast ultrasound video. Although this type of method can better capture the spatiotemporal information in the video, it is difficult to realize rapid deployment due to the large number of parameters, difficult training, and high computational cost, especially in environments that need to protect patient privacy and implement offline work. After completing the feature extraction operation, a fully connected network or other machine learning techniques are usually used to realize automatic diagnosis.

[0008] Therefore, the current common method for extracting features from a contrast ultrasound video cannot simultaneously take into account efficient use of necessary feature information of the contrast ultrasound video under limited hardware costs and inference time. SUMMARY

[0009] Embodiments of the present application provide a method and device for processing a contrast ultrasound video, an electronic device, and a storage medium, so as to efficiently use necessary feature information of the contrast ultrasound video under limited hardware costs and inference time.

[0010] In a first aspect, embodiments of the present application provide a method for processing a contrast ultrasound video, the method comprising:

[0011] converting the contrast ultrasound video into a projection image according to brightness of each frame image in the contrast ultrasound video;

[0012] extracting time attention information of the projection image through a time series attention extraction network;

[0013] extracting contrast ultrasound features of the projection image through a feature extraction network guided by the time attention information;

[0014] inputting the contrast ultrasound features into a classification model to output a classification result of the contrast ultrasound video.

[0015] In a second aspect, embodiments of the present application provide a device for processing a contrast ultrasound video, the device comprising:

[0016] a conversion module configured to convert the contrast ultrasound video into a projection image according to brightness of each frame image in the contrast ultrasound video;

[0017] an extraction module configured to:

[0018] extract time attention information of the projection image through a time series attention extraction network;

[0019] extract contrast ultrasound features of the projection image through a feature extraction network guided by the time attention information;

[0020] a classification module configured to input the contrast ultrasound features into a classification model and output a classification result of the contrast ultrasound video.

[0021] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a transceiver, and a processor:

[0022] a memory configured to store a computer program; a transceiver configured to transceive data under control of the processor; and a processor configured to read the computer program in the memory and perform the processing method of the contrast ultrasound video according to the first aspect.

[0023] In a fourth aspect, an embodiment of the present application provides a readable storage medium, the readable storage medium storing a program or instructions, the program or instructions being executed by a processor to implement the processing method of the contrast ultrasound video according to the first aspect.

[0024] In the embodiments of the present application, in some embodiments of the present application, the contrast ultrasound video can be converted into a projection image according to brightness of each frame image in the contrast ultrasound video, so that the time attention information of the projection image is extracted through the time sequence attention extraction network, and the contrast ultrasound features of the projection image are extracted through the time attention information guided feature extraction network, and then the contrast ultrasound features are input into the classification model to output the classification result of the contrast ultrasound video.

[0025] In the embodiments of the present application, in some embodiments of the present application, the contrast ultrasound video can be converted into a projection image according to brightness of each frame image in the contrast ultrasound video, so that the time attention information of the projection image is extracted through the time sequence attention extraction network, and the contrast ultrasound features of the projection image are extracted through the time attention information guided feature extraction network, and then the contrast ultrasound features are input into the classification model to output the classification result of the contrast ultrasound video. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor under the premise of the drawings.

[0027] Figure 1 The flowchart of the processing method of the contrast ultrasound video provided by the embodiments of the present application;

[0028] Figure 2 A flowchart of a specific embodiment of the method for processing the ultrasound contrast video of the present application is shown.

[0029] Figure 3 A schematic diagram of converting the ultrasound contrast video into a projection image according to an embodiment of the present application is shown.

[0030] Figure 4 A structural block diagram of the processing device for ultrasound contrast video provided by an embodiment of the present application is shown.

[0031] Figure 5 A structural block diagram of the electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0032] The term "and / or" in the embodiments of the present application describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the three cases of A alone, A and B together, and B alone. The character " / " generally represents an "or" relationship between the associated objects before and after it.

[0033] The term "multiple" in the embodiments of the present application means two or more, and other quantifiers are similar.

[0034] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0035] The embodiments of the present application provide a method and device for processing ultrasound contrast video, electronic equipment and storage medium, so as to efficiently utilize the necessary feature information of ultrasound contrast under the condition of limited hardware cost and inference time.

[0036] The method and device are based on the same application concept. Since the principles of the method and device for solving problems are similar, the implementation of the device and the method can be mutually referred to, and the repeated parts will not be described again.

[0037] Figure 1 A flowchart of a method for processing ultrasound contrast video provided by an embodiment of the present application is shown, which can include the following steps 101 to 104:

[0038] Step 101: converting the ultrasound contrast video into a projection image according to the brightness of each frame image in the ultrasound contrast video.

[0039] Among them, contrast-enhanced ultrasound (CEUS) is an extension of traditional ultrasound imaging. By intravenous injection of contrast agent, it can identify the perfusion of blood in the target tissue. In clinical practice, it can significantly improve the diagnostic accuracy, sensitivity and specificity. For example, malignant tumors usually show abnormal blood flow patterns, such as irregular blood vessel distribution and high blood flow velocity, while benign tumors usually have more uniform and stable blood flow. This difference helps doctors assess the characteristics of the tumor and effectively assist doctors in improving diagnostic accuracy.

[0040] In addition, the ultrasound contrast video belongs to three-dimensional data, and the projection image belongs to two-dimensional data. Therefore, the three-dimensional ultrasound contrast video can be converted into a two-dimensional projection image through step 101. It can be understood that the projection image converted from the ultrasound contrast video can also be referred to as an ultrasound contrast projection image.

[0041] In addition, according to the brightness of each frame image in the ultrasound contrast video, the ultrasound contrast video is converted into a two-dimensional projection image. Not only can the three-dimensional information of the video be effectively simplified and converted into a two-dimensional visual and processable projection image, but also the key information (i.e. brightness information) in the video can be retained at the same time.

[0042] Step 102: extracting time attention information of the projection image through a time series attention extraction network.

[0043] Among them, time series attention is a deep learning technique for time series data, which focuses on key time points by dynamically allocating weights to improve the model's ability to capture key information in long sequences. Therefore, in some embodiments of the present application, the time series attention extraction network can capture the time series features (i.e. time attention information) of the projection image.

[0044] In addition, the time attention information extracted by the time series attention extraction network can also be understood as the time attention information of the ultrasound contrast video.

[0045] Optionally, the time series attention extraction network includes at least one three-dimensional convolution layer, at least one three-dimensional pooling layer, a fully connected layer and an activation function.

[0046] The above-mentioned time attention information is used to represent the importance of different regions of the projection image, or to represent the importance of different frames or different time periods of the projection video.

[0047] Step 103: extracting ultrasound contrast features of the projection image through the time attention information guided feature extraction network.

[0048] The time attention information belongs to time sequence characteristics of the projection image or the contrast video, and is used to represent importance of different regions of the projection image, or is used to represent importance of different frames or different time periods of the projection video. Therefore, when the time attention information is used to extract the contrast feature of the projection image by the feature extraction network, the important regions (i.e., important frames of the contrast video) of the projection image can be focused on, so that the extracted feature is more accurate.

[0049] In addition, the feature extraction network can include at least one two-dimensional convolution layer and at least one two-dimensional pooling layer, or the feature extraction network can include at least one two-dimensional deep learning model.

[0050] Step 104: inputting the contrast feature into a classification model to output a classification result of the contrast video.

[0051] After the contrast feature is obtained, the contrast feature is input into the classification model, and then the classification result of the contrast video can be output. For example, when the contrast video is used to diagnose malignant tumors or benign tumors, the classification model can output a probability of belonging to a malignant tumor or a probability of belonging to a benign tumor.

[0052] Optionally, the classification model includes at least one two-dimensional convolution layer, at least one two-dimensional pooling layer, and a fully connected layer, or the classification model includes at least one two-dimensional deep learning model.

[0053] As described above, in some embodiments of the present application, the contrast video can be converted into a projection image according to brightness of each frame of the contrast video, so that the time attention information of the projection image is extracted by the time sequence attention extraction network, and the contrast feature of the projection image is extracted by the feature extraction network guided by the time attention information. Then, the contrast feature is input into the classification model to output the classification result of the contrast video.

[0054] According to brightness of each frame of the contrast video, the three-dimensional contrast video is converted into a two-dimensional projection image, which improves the calculation efficiency and retains important feature information. Moreover, according to the time attention information, the contrast feature of the projection image can be extracted by the feature extraction network, which can focus on important regions (i.e., important frames of the contrast video) of the projection image, so that the extracted feature is more accurate, and a more accurate classification result can be obtained by the classification model. Therefore, the embodiments of the present application can efficiently utilize necessary feature information of the contrast video under the condition of limited hardware cost and inference time.

[0055] Optionally, in step 101, the converting the contrast ultrasound video into the projection image according to the brightness of each frame image in the contrast ultrasound video comprises steps A-1 to A-3 as follows:

[0056] Step A-1: sliding the first step length horizontally and the second step length vertically in the j1th frame image of the contrast ultrasound video through the pre-determined window to obtain X windows overlaid on the j1th frame image, wherein j1 is an integer from 1 to f, f represents the total number of frame images in the contrast ultrasound video and the height of the projection image.

[0057] Step A-2: determining the projection value of the pixel point in the i1th column of the j1th row in the projection image according to the brightness value of the pixel point in the i1th window overlaid on the j1th frame image, i1 is an integer from 1 to X, X represents the number of windows overlaid on a frame image of the contrast ultrasound video and the width of the projection image.

[0058] Step A-3: obtaining the projection value of each pixel point in the projection image until j1=f and i1=X.

[0059] Wherein, the first step length is the length of the window and the second step length is the width of the window.

[0060] Optionally, in step A-2, the average value of the brightness value of the pixel point in the i1th window in the j1th frame image can be used as the projection value of the pixel point in the i1th column of the j1th row of the contrast ultrasound image, or the sum of the brightness value of the pixel point in the i1th window in the j1th frame image can be used as the projection value of the pixel point in the i1th column of the j1th row of the contrast ultrasound image.

[0061] Optionally, the length of the window is The width of the window is w represents the width of a frame image of the contrast ultrasound video and h represents the height of a frame image of the contrast ultrasound video.

[0062] From the above steps A-1 to A-3, the process of converting the contrast ultrasound video into the projection image can be described as follows:

[0063] Firstly, define a blank projection image P∈R X×f X represents the width of the projection image, for convenience of calculation, X should be a perfect square, and f represents the height of the projection image and the number of frames of the contrast ultrasound video.

[0064] Then, define a window with size a×b, wherein w represents the width of the frame of the contrast ultrasound video; h represents the height of the ultrasound contrast video frame; that is, the width and height of the ultrasound contrast video frame are divided by the square root of the projection image width respectively, and the width a and height b of the window are obtained by rounding down, which can ensure that the window can effectively cover each ultrasound contrast video frame in the sliding calculation operation.

[0065] Then, a projection function f(W) is defined, and on this window, the function can be expressed as: wherein Brightness(W m,n ) represents the brightness value of the pixel of the ultrasound contrast video frame located at the position (m, n) of the window. That is, the projection function f(W) calculates the brightness average value located in the window area, so that the ultrasound contrast video frame located in the window area can be compressed to the projection value.

[0066] Finally, for the j1th frame of ultrasound contrast video frame, the above defined window is slid with a horizontal step of a and a vertical step of b, and the window slides in the order of horizontal first and then vertical; if a part of the window exceeds the boundary of the video frame, the coverage area of the window will be ignored; finally, the single frame ultrasound contrast video frame will be covered by X windows.

[0067] Therefore, for the i1th window of the j1th frame, the expression of the projection value on the projection map P(j1, i1) is:

[0068] Brightness(m, n) represents the brightness value of the pixel in the i1th window on the j1th frame image. That is, the average value of the brightness value of the pixel in the i1th window of the j1th frame is taken as the projection value at the position of the j1th row and the i1th column of the projection map.

[0069] wherein i1 ∈ [1, X], X is the width of the projection map, and is also the number of window coverage on the single frame ultrasound contrast video frame, j1 ∈ [1, f], f is the height of the projection map, and is also the number of effective ultrasound contrast video frames.

[0070] As can be seen, through the above process, the ultrasound contrast video is converted into a two-dimensional projection map, while the key information in the video is retained, that is, enough original video information is retained. Through the definition of the size of the projection map and the window sliding operation, the ultrasound contrast video frame can be reasonably divided and covered in space, the brightness average value of each region is calculated, and finally the corresponding projection value is formed. This way can ensure that the three-dimensional information of the video is effectively simplified and converted into a two-dimensional visual and processable projection map, so as to facilitate subsequent analysis and processing.

[0071] Optionally, in step 102, the time attention information of the projection image is extracted by the time sequence attention extraction network, including the following steps B-1 to B-2:

[0072] Step B-1: converting the projection image into a column-priority representation projection image;

[0073] Step B-2: inputting the column-priority representation projection image into the time sequence attention extraction network to obtain the time attention information.

[0074] Therefore, when the time attention information of the projection image is extracted by the time sequence attention extraction network, the projection image can be first converted into a column-priority representation projection image, and then the column-priority representation projection image is input into the time sequence attention extraction network for processing, which can reduce the calculation amount and shorten the processing time of extracting the time attention information of the projection image.

[0075] Optionally, in step B-1, the projection image is converted into a column-priority representation projection image, including the following steps B-1.1 to B-1.3:

[0076] Step B-1.1: splitting the projection image by column to obtain f X 1 first vectors;

[0077] Step B-1.2: expanding each first vector into a 1 X X 1 second vector;

[0078] Step B-1.3: concatenating each second vector according to the first dimension to obtain a f X X 1 matrix, and determining the image represented by the f X X 1 matrix as the column-priority representation projection image;

[0079] Wherein, f represents the height of the projection image, and X represents the width of the projection image.

[0080] From the above steps B-1.1 to B-1.3 and step B-2, the process of extracting the time attention information of the projection image by the time sequence attention extraction network can be as follows:

[0081] The projection image P ∈ R X×f is split by column to form f column vectors p i4 ∈ R X×1 , where i4 ∈ [1, f]; the column vectors are expanded to size 1 X X 1 to obtain the expanded vector p' i4 ∈ R 1×X×1 ; and the expanded vector p' i4 is concatenated according to the first dimension to form a new matrix P' ∈ R f×X×1At this time, P' is a projection image in column-major representation; then, the projection image P' in column-major representation is input into the time sequence attention extraction network, and a time attention vector a of size Hx1 can be obtained t Each element of a is greater than zero and less than 1. t

[0082] Optionally, the time attention information includes a time attention vector of Hx1; in step 103, the feature extraction network is guided by the time attention information to extract the ultrasound contrast feature of the projection image, including the following steps C-1 to C-3:

[0083] Step C-1: input the projection image into the feature extraction network to obtain a feature map of C1xHxW1, wherein C1 represents the number of channels of the feature map, H represents the height of the feature map, and W1 represents the width of the feature map;

[0084] Step C-2: extend the time attention vector to a third vector of 1xHx1;

[0085] Step C-3: perform outer product calculation on the third vector and the feature map to obtain the ultrasound contrast feature of the projection image.

[0086] As can be seen from steps C-1 to C-3, the process of guiding the feature extraction network by the time attention information to extract the ultrasound contrast feature of the projection image can be as follows:

[0087] Input the projection image P into the feature extraction network to obtain a feature map F of C1xHxW1, wherein C1 is the number of channels of the feature map F, H is the height of the feature map F, and W1 is the width of the feature map F; extend the time attention vector a from Hx1 to 1xHx1; then, perform outer product calculation on the vector of 1xHx1 and the feature map F of C1xHxW1 to obtain the ultrasound contrast feature F'. t

[0088] Optionally, the method further includes the following steps D-1 to D-10:

[0089] Step D-1: obtain a training data set, wherein the training data set includes N ultrasound contrast videos, and N is an integer greater than 1;

[0090] Step D-2: convert the i2th ultrasound contrast video in the training data set into the i2th projection image, and i2 is an integer from 1 to N; it should be noted that the specific implementation process of step D-2 is the same as that of steps A-1 to A-3, and will not be repeated here.

[0091] ​​Step D-3: extracting time attention information of the i2th projection image by the time sequence attention extraction network; it should be noted that the specific implementation process of step D-3 here is the same as steps B-1 to B-2 described above, and will not be repeated here.

[0092] Step D-4: extracting ultrasound contrast features of the i2th projection image by the feature extraction network according to the time attention information of the i2th projection image; it should be noted that the specific implementation process of step D-4 here is the same as steps C-1 to C-3 described above, and will not be repeated here.

[0093] Step D-5: inputting the ultrasound contrast features of the i2th projection image into the classification model to output a classification result of the i2th ultrasound contrast video;

[0094] Step D-6: obtaining quantized attention information of the i2th projection image based on the brightness of each frame image in the i2th ultrasound contrast video;

[0095] Step D-7: until i2=N, obtaining classification results of N ultrasound contrast videos and quantized attention information of N projection images;

[0096] Step D-8: calculating a first loss value according to the classification results of N ultrasound contrast videos and classification labels of N ultrasound contrast videos;

[0097] Step D-9: calculating a second loss value according to the time attention information and the quantized attention information of N projection images;

[0098] Step D-10: training the time sequence attention extraction network, the feature extraction network and the classification model according to the first loss value and the second loss value.

[0099] As can be seen from the above steps D-1 to D-10, in the process of training the time sequence attention extraction network, the feature extraction network and the classification model, the quantized attention analysis can be performed on the projection image converted from each ultrasound contrast video in the training data set, that is, the quantized attention information of each projection image is extracted, and then the loss value calculation is performed based on the quantized attention information and the time attention information.

[0100] It should be noted that the quantized attention information refers to the process of converting the abstract state of human or model attention intensity, distribution, etc. into a measurable numerical index or structured data, and the core is to establish the mapping relationship between attention and mathematical expression. Therefore, the quantized attention information of the i2th projection image obtained based on the brightness of each frame image in the i2th ultrasound contrast video can represent the relationship between the attention of the doctor observing the ultrasound contrast video and the brightness change of the video frame.

[0101] Moreover, for the ultrasound contrast video, the doctor mainly focuses on the contrast agent filling and washout process. The filling process refers to the change from dark to light in the video, and the washout process refers to the change from light to dark. Therefore, the quantitative attention information of the projection image can reflect the area that needs to be focused on in the projection image or the time period that needs to be focused on in the ultrasound contrast video; that is, the quantitative attention information can accurately identify and quantify the focus of the doctor in the key stage of the contrast agent filling and washout.

[0102] It can be seen that in some embodiments of the present application, the quantitative attention information of the projection image can be obtained by combining the prior knowledge of the brightness change of the video in the contrast agent filling and washout process, and the deviation between the quantitative attention information and the above-mentioned time attention information can represent the accuracy of the time sequence attention extraction network. Therefore, in the above-mentioned model training process, the second loss value calculated based on the quantitative attention information and the above-mentioned time attention information is introduced, so that the time sequence attention extraction network can be guided by the prior knowledge, so that when the extracted time attention information is applied to the ultrasound contrast feature, it can more accurately focus on the key frame in the contrast video.

[0103] Optionally, the time attention information includes a time attention vector of Hx1, and the quantitative attention information includes a quantitative attention vector of Hx1; in the step D-9, the second loss value is calculated according to the time attention information and the quantitative attention information of the N projection images, including:

[0104] According to the second formula, the second loss value Loss K is calculated.

[0105] Wherein, the second formula is:

[0106] Wherein, a i2 represents the quantitative attention vector of the i2th projection image, a t,i2 represents the time attention vector of the i2th projection image.

[0107] Optionally, in the step D-8, the first loss value is calculated according to the classification results of the N ultrasound contrast videos and the classification labels of the N ultrasound contrast videos, including:

[0108] According to the third formula, the first loss value Loos C is calculated.

[0109] Wherein, the third formula is:

[0110] G i2 represents the label of the i2th ultrasound contrast video, wherein the positive class ultrasound contrast video is G i2= 1, negative class ultrasound contrast video is G i2 = 0, P i represents the probability that the i2th ultrasound contrast video is predicted as a positive class.

[0111] Optionally, in step D-10, the time series attention extraction network, the feature extraction network, and the classification model are trained according to the first loss value and the second loss value, including:

[0112] The sum Loss1 of the first loss value and the second loss value is calculated, and the time series attention extraction network, the feature extraction network, and the classification model are trained according to Loss1;

[0113] Alternatively,

[0114] The first loss value and the second loss value are weighted and summed to obtain Loss2 according to the predetermined weight of the first loss value and the weight of the second loss value, and the time series attention extraction network, the feature extraction network, and the classification model are trained according to Loss2.

[0115] In addition, during the training process, until Loss1 or Loss2 no longer decreases, the model performance evaluation can be performed.

[0116] The model performance evaluation refers to the performance of the test model and the effectiveness of the algorithm, wherein the accuracy, true positive rate, true negative rate, and precision four calculation indicators can be used to compare and measure based on the validation data set. The calculation methods of the four indicators are as follows:

[0117] Accuracy (Accuracy, Acc): used to measure the overall prediction ability of the model, calculated as the proportion of correctly predicted positive and negative class samples to the total number of samples, and the formula is:

[0118] True Positive Rate (True Positive Rate, TPR) / Recall: reflects the identification sensitivity of the model to positive class samples, and the formula is:

[0119] True Negative Rate (True Negative Rate, TNR): reflects the discrimination specificity of the model to negative class samples, and the formula is:

[0120] Precision (Precision, Pre): reflects the reliability of the positive class prediction result, and the formula is:

[0121] Wherein, TP represents the number of positive class samples predicted as positive class, TN represents the number of negative class samples predicted as negative class, FP represents the number of negative class samples predicted as positive class, and FN represents the number of positive class samples predicted as negative class.

[0122] Optionally, in the step D-6, the obtaining of the quantified attention information of the i2th projection image based on the brightness of each frame image in the i2th ultrasound contrast video comprises the following steps D-6.1 to D-6.4:

[0123] Step D-6.1: dividing each frame image of the i2th ultrasound contrast video into H groups, wherein H is an integer greater than 1;

[0124] Step D-6.2: determining a first parameter of the i3th group, wherein the first parameter is the sum of the brightness values of each frame image in a group, and the brightness value of a frame image is the sum of the brightness values of each pixel point in the frame image;

[0125] Step D-6.3: determining a first number of the group corresponding to the maximum first parameter in the H groups;

[0126] Step D-6.4: determining the quantified attention information of the i2th projection image according to the first number and the first parameters of the H groups.

[0127] Optionally, in the step D-6.4, the determining of the quantified attention information of the i2th projection image according to the first number and the first parameters of the H groups comprises:

[0128] determining the quantified attention values corresponding to the H groups according to the first formula;

[0129] determining a vector composed of the quantified attention values corresponding to the H groups as the quantified attention information of the i2th projection image;

[0130] Wherein, the first formula is:

[0131]

[0132] Wherein, a i3 represents the quantified attention value corresponding to the i3th group, Gb i3 represents the first parameter of the i3th group, i max represents the first number.

[0133] As can be seen from steps D-6.1 to D-6.4, the process of obtaining the quantified attention information of the i2th projection image based on the brightness of each frame image in the i2th ultrasound contrast video is as follows:

[0134] For contrast-enhanced ultrasound video, the physician is mainly interested in the contrast agent filling and washout process, regardless of whether the lesion is benign or malignant. The filling process refers to the change from dark to bright in the video, while the washout process refers to the change from bright to dark. Therefore, there is always a frame of the contrast-enhanced ultrasound video frame in which the contrast agent filling is most sufficient, and this frame is also the brightest in brightness among all frames. Therefore, the contrast-enhanced ultrasound video c∈R w×h×f When evenly divided into H groups in the video frame dimension, the expression of the number of frames Gf in each group is:

[0135] where f is the number of frames of the contrast-enhanced ultrasound video c, which is obtained by dividing the total number of frames f by the downward integer of the number of groups H, and subtracting the sum of the number of frames of the first H-1 groups from the total number of frames f, so that the number of frames in each group is as evenly as possible. Each group of video frames can represent a time period of the video and reflect the overall brightness and filling information of the video frames in that time period.

[0136] Again, the brightness sum Gb of the video frames in each of the H groups is calculated; specifically, Gb is the sum of the brightness values of all video frames in each group, and the expression of Gb is:

[0137]

[0138] where Gb i3 represents the brightness sum of the i3th group of video frames, brightness(c i3,j3 ) represents the brightness value of the j3th frame of the i3th group of video frames, and the brightness value of a video frame is the sum of the brightness values of all pixel points in the video frame, Gfi i3 represents the number of video frames in the i3th group of video frames.

[0139] Then, for Gb1 to Gb H , the group with the largest brightness sum is marked as Gb max , and its subscript is i max ; thus, based on Gb1 to Gb H and i max , the quantized attention value corresponding to the i3th group can be obtained according to the following formula:

[0140]

[0141] In this way, a1 to a H quantify the relationship between the physician's attention and the brightness change of the video frames when observing the contrast-enhanced ultrasound video.

[0142] Optionally, the method further comprises:

[0143] cropping the contrast-enhanced ultrasound video to retain the contrast-enhanced ultrasound information.

[0144] Therefore, the ultrasound contrast video in the step 101 or the ultrasound contrast video in the training data set can be cropped to retain the ultrasound contrast information, such as cropping the device number, patient data, and other non-valid information.

[0145] Optionally, the ultrasound contrast video includes a video clip of contrast agent filling and disappearing.

[0146] Therefore, the ultrasound contrast video in the step 101 or the ultrasound contrast video in the training data set can include a video clip of contrast agent filling and disappearing to ensure that the ultrasound contrast video is valid.

[0147] In addition, a certain number of ultrasound contrast videos can be selected as a verification data set to verify the time sequence attention extraction network, the feature extraction network, and the classification model after training. It can be understood that the ultrasound contrast video in the verification data set can also be cropped to retain the ultrasound contrast information, and / or include a video clip of contrast agent filling and disappearing.

[0148] In order to facilitate the understanding of the above-mentioned ultrasound contrast video processing method, the specific implementation of the ultrasound contrast video processing method is as follows:

[0149] (1) The training process, as shown in FIG. 2, mainly includes the following steps 201 to 208: Figure 2

[0150] Step 201: Constructing a data set, the specific process is as follows:

[0151] Although ultrasound contrast has been widely used in the diagnosis of various diseases, there is no publicly available ultrasound contrast data set. In order to improve the present embodiment, 580 breast ultrasound contrast samples were collected from a cooperative institution, including 330 benign data and 250 malignant data. Each sample includes a breast ultrasound contrast video C2 containing the change process of contrast agent from filling to complete disappearance, and its corresponding pathological diagnosis result as a binary classification label G of ultrasound contrast video benign and malignant classification.

[0152] First, the breast ultrasound contrast video in the 580 samples is processed by a specific cropping method to retain only the ultrasound contrast information in the video frame and delete invalid information such as device number and patient information. It should be noted that the width w and the height h of different ultrasound contrast videos after cropping can be different, but the width w and the height h of the video frames in the same ultrasound contrast video need to be the same.

[0153] ​Secondly, the intercepted contrast video should cover the whole process from the beginning of contrast agent filling to the complete disappearance, and each video is intercepted 224 frames, so an effective contrast video is represented as c e R w×h×224 .

[0154] Then, according to the pathological diagnosis result, a binary classification label is marked for each sample of the breast contrast video, indicating that the sample is benign or malignant.

[0155] Finally, 520 cases of processed sample data are randomly selected as the training data set, and the remaining 60 cases are used as the validation data set.

[0156] Step 202: Convert the contrast video into a projection map, the specific process is as follows:

[0157] First, for each contrast video, define a blank projection map P e R 224×224 224 corresponding to the width and height of the video.

[0158] Then, for each contrast video, use the corresponding window to calculate the sliding. It should be noted that because each contrast video is processed by a specific cropping method, the width w and height h of different contrast videos are not necessarily the same, so the window size a x b of each contrast video can not be the same, where the window width a can be set to 224, and the window height b can be set to 224.

[0159] For each contrast video, starting from the first video frame, use the above-mentioned pre-defined window to slide the window in the contrast video frame with a horizontal step of a and a vertical step of b, and slide first horizontally and then vertically; and if a part of the window exceeds the boundary of the video frame, the coverage area of the window will be ignored; in this way, X windows are covered on a video frame.

[0160] Wherein, in each sliding, the average value of the brightness value of the pixel points in the window is calculated using the projection function f(W), and the value is filled into the corresponding projection map P in turn, so that after the window sliding on the first video frame of a contrast video is completed, the first row of the projection map P will be filled. By analogy, until the window sliding calculation on the 224th video frame is completed, all rows and columns of the projection map P will be filled. Exemplarily, an example of converting a contrast video into a projection map is shown in Figure 3 .

[0161] ​Finally, the total of 580 three-dimensional breast contrast-enhanced ultrasound videos in the training set and the validation set were converted into 580 two-dimensional breast contrast-enhanced ultrasound projection images. Compared with directly analyzing three-dimensional ultrasound projection images using a three-dimensional convolutional network, analyzing a single two-dimensional projection image by a two-dimensional convolutional network is expected to reduce about 90% of the computational overhead, hardware cost and inference time.

[0162] Step 203: Extracting the temporal attention information of the projection image, the specific process is as follows:

[0163] First, the projection image P e R 224×224 Split by column, 224 column vectors p of size 224 x 1 will be formed i4 e R 224×1 , where i e [1, 224]; extend these column vectors to size 1 x 224 x 1 to obtain the extended vector p' i3 e R 1×224×1 ; so as to obtain the extended vector p' i4 According to the first dimension, a new matrix P' e R 224×224×1 is formed, at this time, P' is a column-first representation of the projection image.

[0164] Secondly, input P' into the pre-constructed time series attention extraction network: that is, input P' into a three-dimensional convolutional layer with a kernel size of 3 x 3 x 1 and a step of 1 x 1 x 1, output a 32-channel feature map; then, the 32-channel feature map is sequentially passed through a three-dimensional pooling layer with a size of 2 x 2 x 1 and a three-dimensional convolutional layer with a kernel size of 3 x 3 x 1 and a step of 1 x 1 x 1, output a 64-channel feature map; the 64-channel feature map is sequentially passed through a 2 x 2 x 1 three-dimensional pooling layer and a three-dimensional convolutional layer with a kernel size of 3 x 3 x 1 and a step of 1 x 1 x 1, output a 128-channel feature map; the 128-channel feature map is passed through a 2 x 2 x 1 three-dimensional pooling layer, and a linear rectifier (ReLU) activation function is used after each pooling operation for processing; finally, a fully connected layer is connected, output 7 nodes, and a sigmoid activation function is used for processing; in this way, the 7 nodes activated by the sigmoid activation function are the temporal attention vector a t of the projection image, where each element of a t is greater than zero and less than 1.

[0165] It should be noted that the above-mentioned temporal attention information can also be referred to as the temporal attention information of the contrast-enhanced ultrasound video.

[0166] Step 204: Based on the above-mentioned temporal attention information, extract the contrast-enhanced ultrasound features of the projection image, and perform binary classification based on the contrast-enhanced ultrasound features, the specific process is as follows:

[0167] Firstly, a two-dimensional deep learning model (ResNet-50) is used to construct a feature extraction network by using the convolutional layer and the pooling layer, all residual blocks, and the completed projection image P is input into the feature extraction network to obtain a feature map F with a size of 512x7x7; wherein 512 is the channel number of the feature map F, 7 is the height of the feature map, and 7 is the width of the feature map;

[0168] Secondly, the time attention vector a t is obtained in step 203 is expanded from 7x1 to 1x7x1 and is calculated with the outer product of the 512x7x7 feature map F to obtain F';

[0169] Then, F' is input into a classification model, that is, F' is extracted by a 3x3 convolutional layer to output 64 feature maps and a ReLU activation function is applied; then, a 2x2 max pooling layer is used for down-sampling, followed by adding a convolutional layer with a 3x3 convolution kernel to output 128 feature maps and a ReLU activation function is used again, followed by another 2x2 max pooling layer; then, after the feature map is flattened, it is input into a fully connected layer to output 64 nodes and a ReLU activation function is used for processing, and finally, an output layer is connected to output 2 nodes and a Softmax activation function is used for binary classification to output the prediction result of the ultrasound contrast video, that is, the probability of belonging to a malignant tumor or the probability of belonging to a benign tumor.

[0170] It should be noted that the steps 203 to 204 above extract the features on the time axis of the ultrasound contrast video based on prior knowledge, use a specific method to calculate the time attention, and apply it to the ultrasound contrast projection image, which can enhance the explainability, efficiency and accuracy of feature extraction.

[0171] Step 205: According to the prediction result of each ultrasound contrast video in the training data set and the binary classification label, a first loss value is calculated, and the specific process is as follows:

[0172] The first loss value Loos is calculated using the following cross-entropy loss function: C :

[0173]

[0174] G i2 represents the binary classification label of the i2th ultrasound contrast video, wherein the positive class ultrasound contrast video is G i2 = 1, and the negative class ultrasound contrast video is G i2 = 0, P i represents the probability that the i2th ultrasound contrast video is predicted as a positive class.

[0175] Step 206: Extracting the quantized attention information of the ultrasound contrast video, the specific process is as follows:

[0176] Firstly, each valid video frame of the breast ultrasound contrast video is evenly divided into 7 groups; the number of frames Gf in the ith3 group of an ultrasound contrast video is i3 The expression is: Here H = 7, f = 224, then it can be calculated that there are 32 video frames in each group.

[0177] Secondly, according to the expression The brightness of the ith3 group of video frames is calculated Gb i3 , brightness(c i3,j3 ) represents the brightness value of the jth3 frame of the ith3 group of video frames, and the brightness value of a video frame is the sum of the brightness values of all pixel points in the video frame.

[0178] Then, for Gb1 to Gb7, find the group with the maximum brightness value and mark it as Gb max , and its subscript is i max ; thus, based on Gb1 to Gb7 and i max , the quantized attention value corresponding to the ith3 group can be obtained according to the following formula:

[0179]

[0180] Taking an ultrasound contrast video as an example, its quantized attention vector is (1, 0, 1, 1, 1, 1, 0).

[0181] Step 207: Calculating the second loss value according to the time attention information and the quantized attention information, the specific process is as follows:

[0182] As can be seen from the above steps 203 and 206, each ultrasound contrast video can obtain a time attention information and a quantized attention information, so that the 520 ultrasound contrast videos in the above training data set can obtain 520 time attention information and 520 quantized attention information. The time attention information and the quantized attention information of these ultrasound contrast videos can calculate the second loss value Loos K according to the following L2 loss function:

[0183] Where a i2 represents the quantized attention vector of the ith2 ultrasound contrast video, a t,i2 represents the time attention vector of the ith2 ultrasound contrast video, and here N = 520.

[0184] Step 208: calculate the sum of the first loss value and the second loss value Loss1, and then train the model according to Loss1, that is, train the time series attention extraction network, the feature extraction network and the classification model, until Loss1 no longer decreases, and then evaluate the performance of the model based on the above validation data set in terms of accuracy Acc, true positive rate TPR, true negative rate TNR, precision Pre and recall Rec.

[0185] It should be noted that the second loss value Loos K In this way, the time attention vector a t is closer to the attention change of the doctor in the actual diagnosis and treatment of the ultrasound contrast video frames, that is, the time series attention extraction network more accurately focuses on the key frames in the breast contrast video, and the key columns on the breast ultrasound contrast projection map.

[0186] (ii) Application process, that is, after the training of the above time series attention extraction network, the feature extraction network and the classification model is completed, these models can be used to identify whether the ultrasound contrast video belongs to the ultrasound contrast video of malignant tumor or the ultrasound contrast video of benign tumor, which is described in detail as follows:

[0187] Convert the ultrasound contrast video to be identified into a projection map, wherein the specific conversion process can be referred to in the foregoing description and will not be described here again.

[0188] Extract the time attention information of the projection image converted from the ultrasound contrast video to be identified by the time series attention extraction network, wherein the extraction process of the time attention information can be referred to in the foregoing description and will not be described here again.

[0189] Based on the time attention information, extract the ultrasound contrast features of the projection image by the feature extraction network, wherein the specific process of extracting the ultrasound contrast features can be referred to in the foregoing description and will not be described here again.

[0190] Input the ultrasound contrast features into the classification model, and output the prediction result of whether the ultrasound contrast video belongs to the ultrasound contrast video of malignant tumor or the ultrasound contrast video of benign tumor.

[0191] In addition, it should be noted that a two-dimensional convolutional neural network (2D-CNN) is a convolutional neural network mainly used to process two-dimensional data, such as images and single frames of videos. The basic principle is to use a specified two-dimensional convolution kernel to slide on the image to extract features of a local region of the image to capture spatial features of the image, so it is often used for image classification, image segmentation and other tasks. The two-dimensional convolutional neural network mainly processes static images and cannot directly capture changes in time series. Therefore, for the analysis of videos or other dynamic scenes, it is difficult to extract features in time series, and it is difficult to complete tasks in three-dimensional scenes well.

[0192] A three-dimensional convolutional neural network (3D-CNN) is a convolutional neural network mainly used to process three-dimensional data, such as video sequences, medical imaging and object volumes. The basic principle is to use a specified three-dimensional convolution operation to slide in three dimensions to capture dynamic changes and features in time or space. The three-dimensional convolution network is particularly suitable for processing time series information and is often used for video analysis, action recognition and medical image analysis. Because the convolution kernel slides in three dimensions, this results in a significant increase in parameters and computational complexity, which may require more powerful hardware support, thereby increasing the time and cost of training and inference. In addition, three-dimensional data usually occupies more memory, which may bring challenges to model training and deployment in the case of limited memory resources.

[0193] After testing, compared with the scheme of inputting three-dimensional breast ultrasound contrast video into a two-dimensional convolution network and a three-dimensional convolution network for processing, the embodiment of the present application uses a two-dimensional breast ultrasound contrast projection map, reduces the computing overhead by nearly 90%, and the analysis speed is nearly 10 times faster, but only reduces the TNR by about 0.2%, and the accuracy Acc hardly decreases; and after adding time series attention, the accuracy Acc is at least increased by 3%, the true positive rate TPR and the true negative rate TNR are each increased by about 2%, and the precision Pre and the recall are each significantly improved.

[0194] In summary, the embodiments of the present application have the following advantages:

[0195] 1、In the embodiments of the present application, a three-dimensional contrast-enhanced ultrasound video can be converted into a two-dimensional projection image, and in the kernel process, a proper window is constructed to simulate the selection of the region of interest (ROI) in actual diagnosis; at the same time, the average brightness value in each window is extracted by using the sliding window technology, so as to effectively capture the key information in the video frame; in this way, the three-dimensional contrast-enhanced ultrasound video is converted into a two-dimensional projection image, which not only improves the calculation efficiency, but also retains important feature information, and is suitable for subsequent analysis and processing.

[0196] 2、In the embodiments of the present application, a feature extraction network and a time series attention extraction network are constructed to extract key features in the contrast-enhanced ultrasound video; wherein the contrast-enhanced ultrasound projection image is split and expanded into a new matrix according to columns, so as to perform time series attention calculation; and by uniformly grouping the contrast-enhanced ultrasound video frames, the filling and regression of the contrast agent and the relationship between the brightness and the contrast agent are analyzed, the brightness sum of each group is calculated, and the maximum brightness group is found, so that the quantified attention can be generated based on the brightness comparison; in addition, in the model training process, the loss function is used to guide the generation of time attention, so that the feature extraction can pay more attention to the key frames, thereby improving the accuracy of feature extraction and effectively improving the analysis accuracy and interpretability of the contrast-enhanced ultrasound video.

[0197] 3、In the embodiments of the present application, the time attention information extracted is used for outer product calculation with the feature map to generate a feature containing attention, and then the feature is input into the constructed classification model to output a benign or malignant classification prediction result, and the loss between the calculated prediction result and the classification label is used to optimize the model. In addition, a plurality of parameters for evaluating the training effect of the model are proposed to ensure that the final output diagnosis model can effectively analyze and diagnose the contrast-enhanced ultrasound.

[0198] 4、In the embodiments of the present application, a sufficient number of contrast-enhanced ultrasound videos and their corresponding benign or malignant classification labels are collected to remove irrelevant information and ensure the integrity and effectiveness of the video content. In addition, the data is randomly and uniformly divided into a training set and a validation set to ensure the balance of the samples, thereby providing a reliable foundation for the accuracy and performance of the subsequent model.

[0199] In summary, in order to solve the problem that the commonly used method cannot simultaneously consider the efficient use of spatial features and time series features of contrast-enhanced ultrasound under the condition of limited hardware cost and inference time when using a deep learning model to analyze contrast-enhanced ultrasound video, the embodiments of the present application utilize the imaging characteristics of contrast-enhanced ultrasound and the prior knowledge of the quantification of clinical diagnosis theory or habits, and propose a more efficient feature extraction method for contrast-enhanced ultrasound. Through accurate model training and evaluation, the model is more portable, the recognition is faster, and it is easier to deploy, and it has important clinical application prospects.

[0200] The processing method of the ultrasound contrast video provided by the embodiments of the present application is introduced above, and the processing device of the ultrasound contrast video provided by the embodiments of the present application will be introduced below in combination with the drawings.

[0201] Referring to Figure 4 The embodiments of the present application further provide a processing device of an ultrasound contrast video, and the device comprises:

[0202] A conversion module 401 is configured to convert the ultrasound contrast video into a projection image according to the brightness of each frame image in the ultrasound contrast video.

[0203] An extraction module 402 is configured to:

[0204] extract time attention information of the projection image through a time sequence attention extraction network;

[0205] extract ultrasound contrast features of the projection image through the time attention information guided feature extraction network;

[0206] A classification module 403 is configured to input the ultrasound contrast features into a classification model and output a classification result of the ultrasound contrast video.

[0207] Optionally, the conversion module 401 is specifically configured to:

[0208] slide in a first step length horizontally and in a second step length vertically in the j1th frame image of the ultrasound contrast video through a predetermined window, to obtain X windows covering the j1th frame image, wherein j1 is an integer from 1 to f, f represents the total number of frame images in the ultrasound contrast video and the height of the projection image;

[0209] determine the projection value of the pixel point in the j1th row and the i1th column of the projection image according to the brightness value of the pixel point in the i1th window covering the j1th frame image, i1 is an integer from 1 to X, X represents the number of windows covering one frame image of the ultrasound contrast video and the width of the projection image;

[0210] until j1=f and i1=X, the projection value of each pixel point in the projection image is obtained;

[0211] wherein the first step length is the length of the window and the second step length is the width of the window.

[0212] Optionally, the length of the window is

[0213] the width of the window is w represents a width of a frame image of the ultrasound contrast video, and h represents a height of the frame image of the ultrasound contrast video.

[0214] Optionally, the extraction module 402 extracts time attention information of the projection image through a time sequence attention extraction network, including:

[0215] converting the projection image into a column-priority representation projection image;

[0216] inputting the column-priority representation projection image into the time sequence attention extraction network to obtain the time attention information.

[0217] Optionally, the extraction module 402 converts the projection image into a column-priority representation projection image, including:

[0218] splitting the projection image by column to obtain f X 1 first vectors;

[0219] extending each of the first vectors into a 1 X X 1 second vector;

[0220] splicing each of the second vectors according to the first dimension to obtain a f X X 1 matrix, and determining the image represented by the f X X 1 matrix as the column-priority representation projection image;

[0221] wherein f represents a height of the projection image, and X represents a width of the projection image.

[0222] Optionally, the time attention information includes a H X 1 time attention vector;

[0223] The extraction module 402 extracts ultrasound contrast features of the projection image through a feature extraction network guided by the time attention information, including:

[0224] inputting the projection image into the feature extraction network to obtain a C1 X H X W1 feature map, wherein C1 represents a number of channels of the feature map, H represents a height of the feature map, and W1 represents a width of the feature map;

[0225] extending the time attention vector into a 1 X H X 1 third vector;

[0226] performing outer product calculation on the third vector and the feature map to obtain the ultrasound contrast features of the projection image.

[0227] Optionally, the device further includes a training module for:

[0228] obtaining a training data set, wherein the training data set includes N ultrasound contrast videos, and N is an integer greater than 1;

[0229] convert the i2th ultrasound contrast video into an i2th projection image according to brightness of each frame image in the i2th ultrasound contrast video, i2 being an integer from 1 to N;

[0230] extract time attention information of the i2th projection image through the time sequence attention extraction network;

[0231] extract ultrasound contrast features of the i2th projection image through the feature extraction network according to the time attention information of the i2th projection image;

[0232] input the ultrasound contrast features of the i2th projection image into a classification model, and output a classification result of the i2th ultrasound contrast video;

[0233] obtain quantized attention information of the i2th projection image based on brightness of each frame image in the i2th ultrasound contrast video;

[0234] until i2=N, obtain classification results of N ultrasound contrast videos and quantized attention information of N projection images;

[0235] calculate a first loss value according to the classification results of the N ultrasound contrast videos and classification labels of the N ultrasound contrast videos;

[0236] calculate a second loss value according to the time attention information and the quantized attention information of the N projection images;

[0237] train the time sequence attention extraction network, the feature extraction network and the classification model according to the first loss value and the second loss value.

[0238] Optionally, the training module obtains the quantized attention information of the i2th projection image based on brightness of each frame image in the i2th ultrasound contrast video, and the obtaining includes:

[0239] divide each frame image of the i2th ultrasound contrast video into H groups, wherein H is an integer greater than 1;

[0240] determine a first parameter of an i3th group, wherein the first parameter is a sum of brightness values of each frame image in a group, the brightness value of a frame image is a sum of brightness values of each pixel in the frame image, and i3 is an integer from 1 to H;

[0241] determine a first number of a group corresponding to the maximum first parameter in the H groups;

[0242] determine the quantized attention information of the i2th projection image according to the first number and the first parameters of the H groups.

[0243] Optionally, the training module determines quantized attention information of the i2th projection image according to the first number and the first parameters of the H groups, including:

[0244] determining quantized attention values corresponding to the H groups according to the first formula;

[0245] determining a vector composed of the quantized attention values corresponding to the H groups as the quantized attention information of the i2th projection image;

[0246] wherein the first formula is:

[0247]

[0248] wherein a i3 represents the quantized attention value corresponding to the i3th group, Gb i3 represents the first parameter of the i3th group, i max represents the first number.

[0249] Optionally, the device further includes:

[0250] a clipping module configured to clip the contrast ultrasound video to retain the contrast ultrasound information.

[0251] Optionally, the contrast ultrasound video includes a video clip of contrast agent filling and washout.

[0252] It should be noted that the division of units in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, there can be another division manner. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0253] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a processor-readable storage medium. Based on such an understanding, the technical solutions of the present application, essentially or in other words, the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store program codes.

[0254] It should be noted that the above-mentioned device provided by the embodiments of the present application can realize all the method steps realized by the method embodiments and achieve the same technical effects. Therefore, the same parts and beneficial effects of the method embodiments will not be described in detail.

[0255] The embodiments of the present application also provide an electronic device, as shown in the figure, the electronic device includes a memory 520, a transceiver 510, a processor 500; Figure 5 The transceiver 510 is used for receiving and sending data under the control of the processor 500.

[0256] The memory 520 is used for storing a computer program.

[0257] The transceiver 510 is used for receiving and sending data under the control of the processor 500.

[0258] The processor 500 is used for reading the computer program in the memory 520 and executing the processing method of the ultrasonic contrast video of the first aspect.

[0259] Wherein, in Figure 5In this regard, bus architecture can include any number of interconnected buses and bridges, specifically, various circuitry linking the one or more processors represented by processor 500 and memory represented by memory 520. Bus architecture can also link various other circuitry such as peripheral devices, voltage regulators, and power management circuitry, which are well known in the art and thus, not further described herein. Bus interface provides an interface to the bus architecture. Transceiver 510 can be a plurality of elements, i.e., including a transmitter and a receiver, providing a means for communicating with various other apparatus over a transmission medium, including a wireless channel, a wired channel, optical cable, etc. Processor 500 is responsible for managing the bus architecture and general processing, and memory 520 can store data used by processor 500 in executing operations.

[0260] Processor 500 can be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD), and processor 500 can also take a multi-core architecture.

[0261] It should be noted that the above apparatus provided by the embodiments of the present application can realize all the method steps achieved by the above method embodiments, and achieve the same technical effects. Therefore, the same parts and beneficial effects of the method embodiments will not be described in detail.

[0262] The embodiments of the present application also provide a readable storage medium, the readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to realize the above-mentioned ultrasound contrast video processing method.

[0263] The computer readable storage medium can be any available medium or data storage that can be accessed by a processor including both volatile and nonvolatile media, removable and non-removable media, computer readable storage media, and computer readable transmission media. By way of example, and not limitation, computer readable media can comprise the following: magnetic storage media, such as custom (e.g., floppy disks), hard disks, magnetic tapes, and magnetic-based optical disks (e.g., Magneto-Optical (MO) disks); optical storage media, such as optical disks, including Compact Disk Read Only Memory (CD-ROM), Digital Versatile Disk (DVD), Blu-Ray Disk (BD), and holographic light disks; and semiconductor storage media, such as Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory, and solid state drives (SSDs). Other computer readable media that can store data include the whole of or parts of machine, human, or computer generated data signals, computer generated data signals embodied in carrier waves, transitory signals, electromagnetic signals, or inductions signals (e.g., a data signal in which data is carried by a carrier wave, a transitory signal, an electromagnetic signal, or an induction signal).

[0264] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, magnetic disks, optical storage media, and the like) embodying computer readable program code.

[0265] The computer readable program code can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the computer readable program code which execute on the computer, other programmable data processing apparatus, or other device implement the functions specified in the flowchart block or blocks. Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to the present application. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable Figure 1 The flowchart and / or block diagram in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to the present application. In this regard, each block in the flowchart and / or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable

[0266] These processor-executable instructions can also be stored in a processor-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the processor-readable memory produce an article of manufacture including instruction means which implement the function specified in a flowchart Figure 1 of flows or multiple flows and / or blocks Figure 1 of blocks or multiple blocks.

[0267] These processor-executable instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the function specified in a flowchart Figure 1 of flows or multiple flows and / or blocks Figure 1 of blocks or multiple blocks.

[0268] Obviously, numerous modifications and variations of the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims and their equivalents, the application can be practiced otherwise than as specifically described.

Claims

1. A method for processing ultrasound contrast imaging video, characterized in that, The method includes: Based on the brightness of each frame in the ultrasound contrast video, the ultrasound contrast video is converted into a projection image. Temporal attention information of the projected image is extracted using a time-series attention extraction network; The time attention information guides the feature extraction network to extract ultrasound contrast features from the projected image; The ultrasound contrast imaging features are input into the classification model, and the classification results of the ultrasound contrast imaging video are output.

2. The method according to claim 1, characterized in that, The step of converting the ultrasound contrast imaging video into a projection image based on the brightness of each frame in the video includes: By sliding a predetermined window horizontally for a first step length and then vertically for a second step length in the j1st frame of the ultrasound contrast imaging video, X windows are obtained covering the j1st frame, where j1 is an integer from 1 to f, and f represents the total number of frames in the ultrasound contrast imaging video and the height of the projected image. Based on the brightness value of the pixel in the i1th window covering the j1th frame image, determine the projection value of the pixel in the j1th row and i1th column of the projected image, where i1 is an integer from 1 to X, and X represents: the number of windows covered on a frame of the ultrasound contrast video, and the width of the projected image. Until j1 = f and i1 = X, the projection values ​​of each pixel in the projected image are obtained; Wherein, the first step length is the length of the window, and the second step length is the width of the window.

3. The method according to claim 2, characterized in that, The length of the window The width of the window w represents the width of a frame of the ultrasound contrast imaging video, and h represents the height of a frame of the ultrasound contrast imaging video.

4. The method according to claim 1, characterized in that, The step of extracting temporal attention information from the projected image using a time-series attention extraction network includes: Convert the projected image into a column-major representation; The column-first represented projected image is input into the time-series attention extraction network to obtain the time attention information.

5. The method according to claim 4, characterized in that, The step of converting the projected image into a column-major representation includes: The projected image is split into columns to obtain f X×1 first vectors; Each of the first vectors is expanded into a second vector of 1×X×1; Each of the second vectors is concatenated according to the first dimension to obtain an f×X×1 matrix, and the image represented by the f×X×1 matrix is ​​determined as the column-major projection image; Where f represents the height of the projected image and X represents the width of the projected image.

6. The method according to claim 1, characterized in that, The temporal attention information includes an H×1 temporal attention vector; The step of extracting ultrasound contrast features of the projected image through the time attention information-guided feature extraction network includes: The projected image is input into the feature extraction network to obtain a C1×H×W1 feature map, where C1 represents the number of channels in the feature map, H represents the height of the feature map, and W1 represents the width of the feature map. The time attention vector is expanded into a third vector of 1×H×1; The ultrasound contrast features of the projected image are obtained by performing an outer product calculation between the third vector and the feature map.

7. The method according to claim 1, characterized in that, The method further includes: Obtain a training dataset, wherein the training dataset includes N ultrasound contrast videos, where N is an integer greater than 1; Based on the brightness of each frame in the i2th ultrasound contrast imaging video, the i2th ultrasound contrast imaging video is converted into the i2th projection image, where i2 is an integer from 1 to N; The temporal attention information of the i2th projection image is extracted using the time-series attention extraction network. Based on the temporal attention information of the i2th projection image, the ultrasound contrast features of the i2th projection image are extracted through the feature extraction network; Input the ultrasound contrast features of the i2th projection image into the classification model and output the classification result of the i2th ultrasound contrast video; Based on the brightness of each frame in the i2th ultrasound contrast video, obtain the quantized attention information of the i2th projection image; Until i2 = N, we obtain the classification results of N ultrasound contrast videos and the quantized attention information of N projection images; Calculate the first loss value based on the classification results of N ultrasound contrast imaging videos and the classification labels of N ultrasound contrast imaging videos; Calculate the second loss value based on the temporal attention information and quantized attention information of N projected images; The time-series attention extraction network, the feature extraction network, and the classification model are trained based on the first loss value and the second loss value.

8. The method according to claim 7, characterized in that, The process of obtaining quantized attention information for the i2th projected image based on the brightness of each frame in the i2th ultrasound contrast-enhanced video includes: Divide each frame of the i2th ultrasound contrast video into H groups, where H is an integer greater than 1; Determine the first parameter of the i3th group, where the first parameter is the sum of the brightness values ​​of each frame image in a group, the brightness value of a frame image is the sum of the brightness values ​​of each pixel in that frame image, and i3 is an integer from 1 to H. Determine the first number of the group corresponding to the largest first parameter among the H groups; Based on the first number and the first parameters of the H groups, the quantization attention information of the i2th projected image is determined.

9. The method according to claim 8, characterized in that, The step of determining the quantized attention information of the i2th projected image based on the first number and the first parameters of the H groups includes: Based on the first formula, determine the quantized attention values ​​corresponding to the H groups; The vector formed by the quantized attention values ​​corresponding to the H groups is determined as the quantized attention information of the i2th projection image; The first formula is: Among them, a i3 This represents the quantized attention value corresponding to the i3rd group, in Gb. i3 This represents the first parameter of the i3rd group, i max This indicates the first number.

10. The method according to any one of claims 1 to 9, characterized in that, The method further includes: The ultrasound contrast video is cropped to retain the ultrasound contrast information.

11. The method according to any one of claims 1 to 9, characterized in that, The ultrasound contrast video includes video segments showing the contrast agent filling and dissipating.

12. A device for processing ultrasound contrast imaging video, characterized in that, The device includes: The conversion module is used to convert the ultrasound contrast imaging video into a projection image based on the brightness of each frame in the video. Extraction module, used for: Temporal attention information of the projected image is extracted using a time-series attention extraction network; The time attention information guides the feature extraction network to extract ultrasound contrast features from the projected image; The classification module is used to input the ultrasound contrast imaging features into the classification model and output the classification result of the ultrasound contrast imaging video.

13. An electronic device, characterized in that, Includes memory, transceiver, and processor: A memory for storing computer programs; a transceiver for sending and receiving data under the control of the processor; and a processor for reading the computer programs in the memory and executing the ultrasound contrast video processing method according to any one of claims 1 to 11.

14. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the method for processing ultrasound contrast imaging video as described in any one of claims 1 to 11.