Classification methods, devices, equipment, and storage media for ultrasound videos of kidney tumors.

By acquiring and processing multimodal ultrasound videos, performing feature extraction and modality fusion, and using the YOLOX network to predict the location and type of kidney tumors, the problem of low accuracy in kidney tumor ultrasound video classification was solved, achieving higher diagnostic accuracy and reliability.

CN116740609BActive Publication Date: 2026-01-30SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310695617.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-12
Publication Date
2026-01-30
Estimated Expiration
2043-06-12

AI Technical Summary

Technical Problem

In existing technologies, the classification accuracy of ultrasound videos of kidney tumors is low, and the information between multimodal images and multiple slices is not effectively utilized, which affects the accuracy and reliability of diagnosis.

Method used

Multimodal ultrasound videos of kidney tumors were acquired, labeled, and processed to generate unified ultrasound image data. Modal fusion was performed through feature extraction and attention mechanisms. The location and category of kidney tumors were predicted using the YOLOX target detection network, and target features were selected for classification based on confidence scores.

Benefits of technology

It improves the classification accuracy of renal tumor ultrasound videos, reduces the detection error rate, and enhances the reliability and accuracy of renal tumor diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116740609B_ABST
    Figure CN116740609B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, device, and storage medium for classifying ultrasound videos of renal tumors. The method includes acquiring multimodal ultrasound videos of renal tumors and annotating multiple multimodal ultrasound videos to obtain a dataset; processing the data in the dataset to generate unified ultrasound image data; extracting features from the ultrasound image data to obtain multi-scale features; performing modal fusion processing on the multi-scale features based on an attention mechanism to obtain fused features; predicting the location and category of the renal tumor in any frame of the multimodal ultrasound video based on the fused features to obtain renal tumor prediction results on multiple single frames; filtering target features from the renal tumor prediction results based on confidence levels and performing prediction processing based on the target features to obtain target classification results. This application is beneficial for improving the classification accuracy of ultrasound videos of renal tumors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for classifying ultrasound videos of kidney tumors. Background Technology

[0002] Kidney tumors are a common type of tumor, generally classified as benign or malignant. Malignant kidney tumors, also known as renal cell carcinoma, are the most common malignant tumor, ranking first among urinary system malignancies. Because early symptoms are often subtle, most patients are diagnosed at an advanced stage, making early diagnosis crucial. Statistics show that the 5-year survival rate for early-stage renal cell carcinoma is over 90%, while the 5-year survival rate for late-stage renal cell carcinoma is only around 10-20%. This means that early detection and diagnosis of kidney tumors significantly improves patient survival. Although early-stage kidney tumors often have no obvious symptoms, as the tumor grows, patients may experience symptoms such as lower back pain, hematuria, and abdominal masses. However, these symptoms are not always caused by kidney tumors, requiring further examination and diagnosis to confirm the diagnosis. In conclusion, kidney tumors are a common and dangerous type of tumor. Early diagnosis and treatment are key to improving patient survival rates. Therefore, strengthening the understanding and research of kidney tumors and exploring more effective, convenient, and reliable diagnostic methods are of great significance for improving the survival rate and quality of life of patients with kidney tumors.

[0003] Computer-aided diagnosis (CAD) refers to a technology that uses computer technology to analyze and process medical images, providing digital image classification results to assist doctors in diagnosing and treating diseases. With the continuous development of computer and artificial intelligence technologies, CAD technology will play an increasingly important role in medical imaging diagnosis, providing better support and assurance for doctors' diagnosis and treatment. Ultrasound diagnosis is a commonly used method for the early diagnosis and surgical treatment of kidney tumors, offering advantages such as being non-invasive, radiation-free, and inexpensive. However, due to limitations in ultrasound image quality and resolution, doctors often struggle to accurately determine the location, shape, and size of kidney tumors, affecting the accuracy and reliability of the diagnosis. CAD technology can help doctors more accurately determine the location, shape, and size of tumors, improving the accuracy and reliability of subsequent diagnoses. Current research utilizes imaging technologies such as CT, enhanced CT, and multiphase MRI combined with deep learning to classify kidney tumors as benign or malignant. This approach can achieve automatic detection and segmentation of tumor regions. However, it does not employ multimodal imaging to improve the model's detection level, nor does it utilize information from multiple slices to reduce the model's error rate, resulting in low classification accuracy of kidney tumor ultrasound videos. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, device, and storage medium for classifying ultrasound videos of kidney tumors, so as to improve the classification accuracy of ultrasound videos of kidney tumors, which is currently low.

[0005] To address the aforementioned technical problems, embodiments of this application provide a method for classifying renal tumor ultrasound videos, comprising:

[0006] Multimodal ultrasound videos of kidney tumors were acquired, and multiple multimodal ultrasound videos were annotated to obtain a dataset;

[0007] Data processing is performed on the data in the dataset to generate unified ultrasound image data;

[0008] Multi-scale features are obtained by extracting features from the ultrasound image data;

[0009] The multi-scale features are subjected to modal fusion processing based on an attention mechanism to obtain fused features;

[0010] Based on the fusion features, the location and type of kidney tumor in any frame of the multimodal ultrasound video are predicted to obtain the kidney tumor prediction result on a single frame image.

[0011] Target features in the kidney tumor prediction results are selected based on confidence level, and prediction processing is performed based on the target features to obtain target classification results.

[0012] To address the aforementioned technical problems, embodiments of this application provide a classification device for renal tumor ultrasound videos, comprising:

[0013] The dataset acquisition unit is used to acquire multimodal ultrasound videos of kidney tumors and to annotate multiple multimodal ultrasound videos to obtain a dataset.

[0014] A data preprocessing unit is used to process the data in the dataset to generate unified ultrasound image data.

[0015] The feature extraction unit is used to extract features from the ultrasound image data to obtain multi-scale features;

[0016] The modality fusion unit is used to perform modality fusion processing on the multi-scale features based on an attention mechanism to obtain fused features;

[0017] The prediction result generation unit is used to predict the location and category of the kidney tumor in any frame of the multimodal ultrasound video based on the fusion features, and obtain the kidney tumor prediction result on a single frame image.

[0018] The classification result generation unit is used to filter out target features in the kidney tumor prediction results based on confidence level, and perform prediction processing based on the target features to obtain target classification results.

[0019] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is to provide a computer device, including one or more processors; and a memory for storing one or more programs, such that the one or more processors implement the classification method of renal tumor ultrasound video as described in any one of the above-mentioned methods.

[0020] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is: a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the classification method of renal tumor ultrasound video as described in any one of the above-mentioned methods.

[0021] This invention provides a method, apparatus, device, and storage medium for classifying ultrasound videos of kidney tumors. The invention acquires multimodal ultrasound videos of kidney tumors and annotates multiple such videos to obtain a dataset. Data processing is performed on the dataset to generate unified ultrasound image data. Feature extraction is performed on the ultrasound image data to obtain multi-scale features. Modality fusion processing is then performed on the multi-scale features based on an attention mechanism to obtain fused features. Based on the fused features, the location and category of the kidney tumor in any frame of the multimodal ultrasound video are predicted to obtain a kidney tumor prediction result for a single frame. Target features are selected from the kidney tumor prediction results based on confidence levels, and prediction processing is performed based on these target features to obtain a target classification result. This invention generates unified ultrasound image data, performs feature extraction and modality fusion based on the ultrasound image data to obtain fused features, and then predicts the location and category of the kidney tumor based on these fused features. This enables the detection of kidney tumor regions and utilizes information analysis from multiple frames, reducing the error rate and thus improving the classification accuracy of kidney tumor ultrasound videos. Attached Figure Description

[0022] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1A flowchart illustrating the implementation of a method for classifying renal tumor ultrasound videos according to an embodiment of this application;

[0024] Figure 2 This is a flowchart illustrating an implementation of a sub-process in the kidney tumor ultrasound video classification method provided in this application embodiment;

[0025] Figure 3 This is a schematic diagram of a dataset example provided in an embodiment of this application;

[0026] Figure 4 This is another implementation flowchart of a sub-process in the classification method of renal tumor ultrasound video provided in the embodiments of this application;

[0027] Figure 5 This is another implementation flowchart of a sub-process in the classification method of renal tumor ultrasound video provided in the embodiments of this application;

[0028] Figure 6 This is another implementation flowchart of a sub-process in the classification method of renal tumor ultrasound video provided in the embodiments of this application;

[0029] Figure 7 This is another implementation flowchart of a sub-process in the classification method of renal tumor ultrasound video provided in the embodiments of this application;

[0030] Figure 8 This is a schematic diagram of a kidney tumor ultrasound video classification device provided in an embodiment of this application;

[0031] Figure 9 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0033] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0034] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0035] It should be noted that the kidney tumor ultrasound video classification method provided in this application embodiment is generally executed by a server, and correspondingly, the kidney tumor ultrasound video classification device is generally configured in the server.

[0036] Please see Figure 1 , Figure 1 This paper illustrates one specific implementation of a method for classifying ultrasound videos of kidney tumors.

[0037] It should be noted that if substantially the same result is obtained, the method of this invention is not based on... Figure 1 Limited to the order of the processes shown, this method includes the following steps:

[0038] S1. Acquire multimodal ultrasound videos of kidney tumors and annotate multiple multimodal ultrasound videos to obtain a dataset.

[0039] Specifically, this application embodiment uses B-mode and contrast-enhanced ultrasound videos to classify renal tumors. During video collection: firstly, dual-frame ultrasound diagnostic videos are recorded for each case, with B-mode and contrast-enhanced images captured simultaneously. The videos reflect the physician's actions and focus during the contrast-enhanced ultrasound examination, and the video length covers all three phases of the contrast-enhanced ultrasound. Then, the dual images are cropped, and the B-mode and contrast-enhanced videos are saved separately as a single video. It should be noted that the collected cases include both benign and malignant renal tumors.

[0040] Please see Figure 2 and Figure 3 , Figure 2 One specific implementation of step S1 is shown. Figure 3 The following is a detailed illustration of a dataset example provided in the embodiments of this application:

[0041] S11. Acquire B-mode video and contrast modality video of kidney tumors.

[0042] S12. The B-mode video and the contrast-enhanced video are cropped so that the B-mode video and the contrast-enhanced video are merged into one video to obtain the multimodal ultrasound video.

[0043] S13. Return multiple multimodal ultrasound videos to the user terminal so that the user terminal can select and annotate the multiple multimodal ultrasound videos to obtain the dataset.

[0044] Specifically, for each ultrasound diagnostic video, the renal tumor region needs to be labeled. This embodiment requires returning multiple multimodal ultrasound videos to the user terminal, allowing the user terminal to select and label the multiple multimodal ultrasound videos to obtain the dataset. In one specific embodiment, this application uses the labeling software Pair to label the bounding rectangles and categories of renal tumors in the ultrasound diagnostic videos. Specifically, the bounding rectangles of the renal tumor region are labeled at 3-5 frame intervals on each video, and they are classified into malignant and benign types. To ensure the accuracy and reliability of the labeling, multiple doctors can participate in the labeling work, and the labeling results are reviewed and corrected. Finally, this embodiment obtains a dataset as shown below. Figure 3 The dataset shown contains ultrasound diagnostic video data from multiple cases and time phases, with accurate bounding rectangles and category annotations for the renal tumor regions.

[0045] S2. Data processing is performed on the data in the dataset to generate unified ultrasound image data.

[0046] Please see Figure 4 , Figure 4 A specific implementation of step S2 is shown below:

[0047] S21. The image frames in the dataset are scaled and cropped based on the image center to obtain cropped image data.

[0048] S22. Perform data augmentation on the cropped image data using a preset data augmentation method to obtain augmented image data.

[0049] S23. Perform pixel normalization processing on the enhanced image data, and perform mean and variance normalization processing on each channel in the enhanced image data to generate the ultrasound image data with unified data.

[0050] Specifically, this application embodiment uses a deep learning model to classify renal tumors from videos. In actual data collection, differences in the examination equipment and operating standards used can lead to inconsistent image resolutions, aspect ratios, etc., while the input size of the deep learning model remains constant. Therefore, to reduce the impact of data inconsistency on model training and evaluation results, this application embodiment performs data preprocessing. First, this application embodiment divides each ultrasound diagnostic video into two videos: a B-mode and a contrast-enhanced mode, and standardizes the aspect ratio, size, etc., of all videos in both modes.

[0051] Specifically, in this embodiment, the shorter side of the image is scaled to 640 pixels while maintaining the aspect ratio, and then a 640*640 region is extracted using center cropping. Secondly, to increase the robustness and generalization ability of the network, this embodiment employs a preset data augmentation method to augment the cropped image data, resulting in enhanced image data. The preset data augmentation methods used in this embodiment include random rotation, random flipping, and MixUP. Finally, before inputting the data into the deep learning model, this embodiment requires standardization of the input data. Specifically, this embodiment normalizes the pixel values ​​of the data to between 0 and 1, and performs mean and variance normalization on each channel to avoid differences between different channels affecting the model's training effect. Through the above data preprocessing steps, this embodiment can obtain a set of ultrasound image data with consistent size, aspect ratio, and pixel count, providing a guarantee for the subsequent training and evaluation of the deep learning model.

[0052] S3. Multi-scale features are obtained by extracting features from the ultrasound image data.

[0053] Furthermore, this application provides a specific embodiment of step S3: the ultrasound image data is downsampled step by step using two Swin-Transformer models with shared weights, according to multiple preset multiples, so as to extract features of different modalities in the ultrasound image data respectively, and obtain the multi-scale features.

[0054] Specifically, this application embodiment uses two Swin-Transformer models with shared weights to extract image features from the B-mode and the contrast modality, respectively. In this embodiment, the Swin-Transformer model uses a local attention mechanism instead of a global attention mechanism, thereby reducing computational cost and memory usage. The Swin strategy allows the local attention mechanism to have a global receptive field. This Swin-Transformer model performs excellently in image feature extraction, accurately extracting the required features while considering both local and global information, which helps improve the accuracy and reliability of kidney tumor diagnosis. Specifically, the Swin-Transformer model uses a hierarchical construction method similar to that in convolutional neural networks, progressively downsampling the input ultrasound image to 4x, 8x, 16x, and 32x, thus obtaining multi-scale features. Obtaining four feature maps of different sizes through the above steps helps improve the model's adaptability to targets of varying sizes.

[0055] S4. The multi-scale features are subjected to modal fusion processing based on the attention mechanism to obtain fused features.

[0056] Please see Figure 5 , Figure 5 A specific implementation of step S4 is shown below:

[0057] S41. Based on the self-attention mechanism, extract the unique features of the multi-scale features corresponding to the modalities to obtain the unique features of the two modalities.

[0058] S42. Extract complementary features between the corresponding modalities of the multi-scale features based on the cross-attention mechanism.

[0059] S43. The unique features of the two modes are added to the complementary features respectively to obtain the summed features of the two modes.

[0060] S44. The summation features of the two modalities are concatenated to obtain the fused features.

[0061] Specifically, in this embodiment, multi-scale features are extracted from the second, third, and fourth levels of two Swin-Transformer models for inter-modal feature fusion, aiming to improve the robustness of the model under different tumor sizes. Specifically, the Swin-Transformer model downsamples a 640*640 image to 80*80, 40*40, and 20*20 features, and then performs feature fusion at different scales. This embodiment extracts unique features corresponding to the multi-scale features based on the self-attention mechanism, obtaining unique features for both modalities; then, based on the cross-attention mechanism, complementary features between the multi-scale features are extracted; the unique features of each modality are then added to the complementary features to obtain the summed features of the two modalities; finally, the summed features of the two modalities are concatenated to obtain the fused features. In a specific embodiment, the input multi-scale feature x is mapped to three vectors in the attention mechanism: query vector q, key vector k, and value vector v. Then, the attention weight w is calculated as the similarity between the query vector q and all key vectors k, and these weights are added to the corresponding value vector v to obtain the final output.

[0062] S5. Based on the fusion features, predict the location and type of the kidney tumor in any frame of the multimodal ultrasound video to obtain the kidney tumor prediction result on a single frame image.

[0063] Please see Figure 6 , Figure 6 A specific implementation of step S5 is shown below:

[0064] S51. For a single frame image in the multimodal ultrasound video, the fused features are input into the YOLOX target detection network, wherein the single frame image includes a real label.

[0065] S52. The classification head of the YOLOX object detection network outputs the prediction result at each pixel position in the single frame image based on the fusion features, and an anchor network is constructed based on the prediction result at each pixel position.

[0066] In this network, each pixel position serves as an anchor point, and each anchor point records the probability value corresponding to each category.

[0067] S53. The category corresponding to the highest probability value among the anchor points is taken as the kidney tumor category.

[0068] S54. Predict the location information of the kidney tumor at each anchor point using the regression head of the YOLOX target detection network.

[0069] S55. Calculate the intersection-union ratio (IUU) of the location information and the real label to calculate the four coordinates of the bounding rectangle of the kidney tumor, and obtain the location information of the kidney tumor.

[0070] Specifically, this embodiment uses the PA-FPN in the YOLOX object detection network proposed by Ge et al. for multi-scale feature fusion. PA-FPN is a commonly used network structure in the field of object detection, used to extract features at different scales, thus enabling the network to have good robustness when detecting objects of different sizes. Specifically, in this embodiment, the PA-FPN input is the features after modality fusion, and the output is the feature pyramid generated by PA-FPN. The head module in the YOLOX object detection network is used to predict the location and category of kidney tumors. This embodiment uses the decoupled object detection head of YOLOX, which adopts an anchor-free scheme, that is, it outputs the prediction result at each pixel position of the input feature map, forming an anchor point grid, and each pixel position in the grid is called an anchor point. The YOLOX object detection head is divided into a classification head and a regression head. The classification head outputs the probability value of each category at each anchor point. In this embodiment, the category corresponding to the highest probability value is taken as the classification result of the kidney tumor. The regression head predicts the location information of the kidney tumor (including center coordinates and side length) and its intersection-over-union ratio (IoU) with the true label at each anchor point, and calculates the four coordinates of the bounding rectangle of the kidney tumor. Through the combined action of the classification head and the regression head, this embodiment can obtain the location and category information of the kidney tumor in the image.

[0071] S6. Based on the confidence level, target features in the kidney tumor prediction results are selected, and prediction processing is performed based on the target features to obtain the target classification results.

[0072] Please see Figure 7 , Figure 7 A specific implementation of step S6 is shown below:

[0073] S61. Multiply the maximum probability value of each anchor point by the intersection-union ratio to obtain the product result, and use the product result as the confidence level.

[0074] S62. Construct a confidence grid based on the confidence scores corresponding to the preset frame images.

[0075] S63. Using non-maximum suppression, a preset number of confidence scores are selected from the confidence score grid, and the features at the corresponding positions of the preset number of confidence scores are obtained to obtain the target features.

[0076] S64. Perform prediction processing based on the target features to obtain the target classification result.

[0077] Specifically, after detecting a kidney tumor on a single frame image, this embodiment of the application needs to combine features from multiple frames to provide a more accurate prediction of the benign or malignant nature of the kidney tumor. This embodiment uses the product of the predicted maximum probability value of each anchor point in the single-frame target detection head and the intersection-union ratio (IUGR) as the confidence level, forming a confidence grid. To improve the network's operating efficiency, this embodiment selects the results of the top 750 points based on the confidence level of each point in the confidence grid of the single-frame target detection head, and then uses non-maximum suppression to further select the features of the top 30 high-confidence locations. In one specific embodiment, after performing the above operations on eight frames of images, the top 30 high-confidence features on each single frame image are obtained. Finally, this embodiment uses a classification head composed of fully connected layers to predict the classification of the kidney tumor based on the target features, obtaining the target classification result. The target classification result is the classification of the kidney tumor region as benign or malignant in the kidney tumor ultrasound video.

[0078] This application embodiment acquires multimodal ultrasound videos of kidney tumors and annotates multiple multimodal ultrasound videos to obtain a dataset. Data processing is performed on the data in the dataset to generate unified ultrasound image data. Feature extraction is performed on the ultrasound image data to obtain multi-scale features. Modality fusion processing is performed on the multi-scale features based on an attention mechanism to obtain fused features. Based on the fused features, the location and category of the kidney tumor in any frame of the multimodal ultrasound video are predicted to obtain a kidney tumor prediction result for a single frame image. Target features in the kidney tumor prediction results are selected based on confidence levels, and prediction processing is performed based on the target features to obtain a target classification result. This application embodiment generates unified ultrasound image data, performs feature extraction and modality fusion based on the ultrasound image data to obtain fused features, and then predicts the location and category of the kidney tumor based on the fused features. This enables the detection of kidney tumor regions, and by utilizing information analysis from multiple frames, the error rate of detection is reduced, thereby improving the classification accuracy of kidney tumor ultrasound videos.

[0079] Please refer to Figure 8 As a response to the above Figure 1 The implementation of the method shown in this application provides an embodiment of a kidney tumor ultrasound video classification device, which is similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various computer devices.

[0080] like Figure 8As shown, the kidney tumor ultrasound video classification device of this embodiment includes: a dataset acquisition unit 71, a data preprocessing unit 72, a feature extraction unit 73, a modality fusion unit 74, a prediction result generation unit 75, and a classification result generation unit 76, wherein:

[0081] Data set acquisition unit 71 is used to acquire multimodal ultrasound videos of kidney tumors and to annotate multiple multimodal ultrasound videos to obtain a dataset;

[0082] The data preprocessing unit 72 is used to process the data in the dataset to generate unified ultrasound image data.

[0083] The feature extraction unit 73 is used to extract features from the ultrasound image data to obtain multi-scale features;

[0084] Modality fusion unit 74 is used to perform modality fusion processing on the multi-scale features based on an attention mechanism to obtain fused features;

[0085] The prediction result generation unit 75 is used to predict the location and type of kidney tumor in any frame of the multimodal ultrasound video based on the fusion features, and obtain the kidney tumor prediction result on a single frame image.

[0086] The classification result generation unit 76 is used to filter out the target features in the kidney tumor prediction results based on confidence, and perform prediction processing based on the target features to obtain the target classification result.

[0087] Furthermore, the kidney tumor prediction result on the single-frame image includes the kidney tumor category and the location information of the kidney tumor; the prediction result generation unit 75 includes:

[0088] The feature input unit is used to input the fused features into the YOLOX target detection network for a single frame image in the multimodal ultrasound video, wherein the single frame image includes a real label;

[0089] An anchor network architecture unit is used to output the prediction result at each pixel position in the single frame image based on the fusion feature through the classification head of the YOLOX object detection network, and to construct an anchor network based on the prediction result at each pixel position, wherein each pixel position in the anchor network is used as an anchor point, and the anchor point records the probability value corresponding to each category.

[0090] A kidney tumor category prediction unit is used to determine the category corresponding to the highest probability value among the anchor points as the kidney tumor category.

[0091] A location information prediction unit is used to predict the location information of the kidney tumor at each anchor point using the regression head of the YOLOX target detection network.

[0092] The location information determination unit is used to calculate the intersection-union ratio (IUU) of the location information and the real label to calculate the four coordinates of the bounding rectangle of the kidney tumor and obtain the location information of the kidney tumor.

[0093] Furthermore, the classification result generation unit 76 includes:

[0094] The confidence calculation unit is used to multiply the maximum probability value of the category of each anchor point with the intersection-union ratio to obtain the product result, and use the product result as the confidence score.

[0095] A confidence grid construction unit is used to construct a confidence grid based on the confidence scores corresponding to preset frame images.

[0096] The confidence filtering unit is used to filter out a preset number of confidence scores from the confidence score grid using a non-maximum suppression method, and obtain the features at the corresponding positions of the preset number of confidence scores to obtain the target features;

[0097] The target feature prediction unit is used to perform prediction processing based on the target features to obtain the target classification result.

[0098] Furthermore, the dataset acquisition unit 71 includes:

[0099] The video acquisition unit is used to acquire B-mode video and contrast-enhanced video of kidney tumors;

[0100] The video merging unit is used to trim the B-mode video and the contrast modal video so that the B-mode video and the contrast modal video are merged into one video to obtain the multimodal ultrasound video;

[0101] The video annotation unit is used to return multiple multimodal ultrasound videos to the user terminal, so that the user terminal can select and annotate the multiple multimodal ultrasound videos to obtain the dataset.

[0102] Furthermore, the data preprocessing unit 72 includes:

[0103] The image cropping unit is used to scale the image frames in the dataset and crop them based on the image center to obtain cropped image data.

[0104] The data augmentation unit is used to augment the cropped image data using a preset data augmentation method to obtain augmented image data;

[0105] The normalization processing unit is used to perform pixel normalization processing on the enhanced image data, and to perform mean and variance normalization processing on each channel in the enhanced image data to generate the ultrasound image data with unified data.

[0106] Furthermore, the feature extraction unit 73 includes:

[0107] The downsampling processing unit is used to downsample the ultrasound image data step by step using two Swin-Transformer models with shared weights, according to multiple preset multiples, so as to extract features of different modes in the ultrasound image data and obtain the multi-scale features.

[0108] Furthermore, the modal fusion unit 74 includes:

[0109] A unique feature extraction unit is used to extract unique features of the multi-scale feature corresponding to the modality based on the self-attention mechanism, so as to obtain unique features of the two modalities.

[0110] A complementary feature extraction unit is used to extract complementary features between the corresponding modalities of the multi-scale features based on the cross-attention mechanism.

[0111] The feature addition unit is used to add the unique features of the two modes to the complementary features respectively to obtain the added features of the two modes;

[0112] The feature splicing unit is used to splice the additive features of two modalities to obtain the fused feature.

[0113] This application embodiment acquires multimodal ultrasound videos of kidney tumors and annotates multiple multimodal ultrasound videos to obtain a dataset. Data processing is performed on the data in the dataset to generate unified ultrasound image data. Feature extraction is performed on the ultrasound image data to obtain multi-scale features. Modality fusion processing is performed on the multi-scale features based on an attention mechanism to obtain fused features. Based on the fused features, the location and category of the kidney tumor in any frame of the multimodal ultrasound video are predicted to obtain a kidney tumor prediction result for a single frame image. Target features in the kidney tumor prediction results are selected based on confidence levels, and prediction processing is performed based on the target features to obtain a target classification result. This application embodiment generates unified ultrasound image data, performs feature extraction and modality fusion based on the ultrasound image data to obtain fused features, and then predicts the location and category of the kidney tumor based on the fused features. This enables the detection of kidney tumor regions, and by utilizing information analysis from multiple frames, the error rate of detection is reduced, thereby improving the classification accuracy of kidney tumor ultrasound videos.

[0114] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 9 , Figure 9 This is a basic structural block diagram of the computer device in this embodiment.

[0115] Computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that only a computer device 8 with three components—memory 81, processor 82, and network interface 83—is shown in the figure; however, it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0116] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through methods such as keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0117] The memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the memory 81 may also include both internal storage units and external storage devices of the computer device 8. In this embodiment, the memory 81 is typically used to store the operating system and various application software installed on the computer device 8, such as program code for a classification method of renal tumor ultrasound video. In addition, the memory 81 may also be used to temporarily store various types of data that have been output or will be output.

[0118] In some embodiments, processor 82 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, processor 82 is used to run program code stored in memory 81 or process data, for example, to run the program code for the above-described method for classifying renal tumor ultrasound videos, to implement various embodiments of the method for classifying renal tumor ultrasound videos.

[0119] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 8 and other electronic devices.

[0120] This application also provides another embodiment, namely, a computer-readable storage medium storing a computer program that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described method for classifying ultrasound videos of kidney tumors.

[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.

[0122] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A method of classifying a renal tumor ultrasound video, the method comprising: The method comprises the following steps: Collecting multi-modal ultrasound videos of kidney tumors, and data labeling multiple multi-modal ultrasound videos to obtain a data set; Data processing is performed on the data in the data set to generate ultrasound image data with uniform data; Feature extraction is performed on the ultrasound image data to obtain multi-scale features; Based on the attention mechanism, the multi-scale features are subjected to modal fusion processing to obtain fusion features; Based on the fusion features, the position and category of the kidney tumor in any frame image of the multi-modal ultrasound video are predicted to obtain a kidney tumor prediction result on a single frame image; Based on the confidence, target features are screened out from the kidney tumor prediction result, and prediction processing is performed based on the target features to obtain a target classification result; The kidney tumor prediction result on the single frame image includes the category of the kidney tumor and the position information of the kidney tumor; the prediction processing of the position and category of the kidney tumor in any frame image of the multi-modal ultrasound video based on the fusion features to obtain the kidney tumor prediction result on the single frame image comprises: For a single frame image in the multi-modal ultrasound video, the fusion features are input into a YOLOX target detection network, wherein the single frame image includes a real label; The classification head of the YOLOX target detection network outputs a prediction result on each pixel position in the single frame image based on the fusion features, and an anchor network is constructed based on the prediction result on each pixel position, wherein each pixel position in the anchor network is an anchor point, and the anchor point records a probability value corresponding to each category; The category corresponding to the maximum probability value in the anchor point is taken as the category of the kidney tumor; The regression head of the YOLOX target detection network predicts the position information of the kidney tumor on each anchor point; The intersection over union of the position information and the real label is calculated to calculate four coordinates of a surrounding rectangle of the kidney tumor, thereby obtaining the position information of the kidney tumor; The target features are screened out from the kidney tumor prediction result based on the confidence, and prediction processing is performed based on the target features to obtain a target classification result, which comprises: The category maximum probability value of each anchor point is multiplied by the intersection over union to obtain a product result, and the product result is taken as the confidence; A confidence grid is constructed based on the confidence corresponding to a preset frame image; In a non-maximum suppression manner, a preset number of confidences are screened out from the confidence grid, and features at positions corresponding to the preset number of confidences are obtained to obtain target features; Prediction processing is performed based on the target features to obtain the target classification result.

2. The method of classifying a renal tumor ultrasound video of claim 1, wherein, The method comprises the following steps: Collecting B-mode videos and contrast mode videos of kidney tumors; The B-mode video and the contrast mode video are cropped to combine the B-mode video and the contrast mode video into one video to obtain the multi-modal ultrasound video; Return multiple multi-modal ultrasound videos to a user terminal, so that the user terminal frames and labels multiple multi-modal ultrasound videos to obtain the data set.

3. The method of classifying a renal tumor ultrasound video of claim 1, wherein, The data in the data set is processed to generate data-unified ultrasound image data, including: scaling the image frames in the data set, and cropping based on the image center to obtain cropped image data; performing data enhancement on the cropped image data using a preset data enhancement method to obtain enhanced image data; performing pixel normalization on the enhanced image data, and performing mean and variance normalization on each channel of the enhanced image data to generate the data-unified ultrasound image data.

4. The method of classifying a renal tumor ultrasound video of claim 1, wherein, The multi-scale features are extracted from the ultrasound image data, including: The ultrasound image data is down-sampled by two Swin-Transformer models with shared weights in multiple preset multiples to extract features of different modalities in the ultrasound image data to obtain the multi-scale features.

5. The method of classifying a renal tumor ultrasound video according to any one of claims 1 to 4, characterized in that, The attention mechanism includes self-attention mechanism and cross-attention mechanism; the multi-scale features are modality fused based on the attention mechanism to obtain fusion features, including: extracting unique features of corresponding modalities of the multi-scale features based on the self-attention mechanism to obtain unique features of two modalities; extracting complementary features between corresponding modalities of the multi-scale features based on the cross-attention mechanism; adding the unique features of the two modalities to the complementary features to obtain added features of the two modalities; splicing the added features of the two modalities to obtain the fusion features.

6. An apparatus for classifying a renal tumor ultrasound video, the apparatus comprising: including: a data set acquisition unit for acquiring multi-modal ultrasound videos of kidney tumors, and data labeling multiple multi-modal ultrasound videos to obtain a data set; a data preprocessing unit for processing data in the data set to generate data-unified ultrasound image data; a feature extraction unit for extracting multi-scale features from the ultrasound image data; a modality fusion unit for modality fusion of the multi-scale features based on an attention mechanism to obtain fusion features; a prediction result generation unit for predicting the position and type of kidney tumors in any frame of image of the multi-modal ultrasound video based on the fusion features to obtain a single-frame image kidney tumor prediction result; a classification result generation unit for screening target features from the kidney tumor prediction result based on confidence, and predicting based on the target features to obtain a target classification result; The single-frame image kidney tumor prediction result includes kidney tumor type and kidney tumor position information; the prediction result generation unit includes: a feature input unit for inputting the fusion features into a YOLOX target detection network for a single frame of image in the multi-modal ultrasound video, wherein the single frame of image includes a real label; Anchoring network architecture unit, configured to output a prediction result of each pixel position in the single frame image based on the fusion feature through a classification head of the YOLOX target detection network, and construct an anchoring network based on the prediction result of each pixel position, wherein each pixel position in the anchoring network is an anchor point, and the anchor point records a probability value corresponding to each category; A kidney tumor category prediction unit configured to take a category corresponding to a maximum probability value in the anchor point as the kidney tumor category; A position information prediction unit configured to predict position information of a kidney tumor on each anchor point through a regression head of the YOLOX target detection network; A position information determination unit configured to calculate an intersection over union of the position information and a real label to calculate four coordinates of a bounding rectangle of the kidney tumor, and obtain the position information of the kidney tumor; The classification result generation unit comprises: A confidence calculation unit configured to multiply a maximum probability value of a category of each anchor point with the intersection over union to obtain a product result, and take the product result as a confidence value; A confidence grid construction unit configured to construct a confidence grid based on a confidence value corresponding to a preset frame image; A confidence screening unit configured to screen a preset number of confidence values from the confidence grid in a non-maximum suppression manner, and obtain a target feature corresponding to a position of the preset number of confidence values; A target feature prediction unit configured to perform a prediction process based on the target feature to obtain the target classification result.

7. A computer device, characterized by A memory and a processor, wherein the memory stores a computer program, and the processor implements the kidney tumor ultrasound video classification method according to any one of claims 1 to 5 when executing the computer program.

8. A computer-readable storage medium, characterized in that, A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the kidney tumor ultrasound video classification method according to any one of claims 1 to 5.