Method, system and terminal for classifying echocardiography videos
By constructing a classification model network, the problem of insufficient consideration of view information in echocardiogram video classification was solved, achieving higher classification accuracy and meeting real-time requirements.
Patent Information
- Application Number
- CN202511281412.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies fail to fully consider view information when classifying echocardiographic videos, resulting in inadequate processing of spatial and temporal features and low accuracy.
A classification model network is constructed, including a feature extraction module, a feature enhancement module, and a feature aggregation module. Through interpolation and feature extraction, spatial and temporal features are aggregated and enhanced. Frame-level feature weighted fusion and gated feature selection are then performed to obtain fused features for classification.
It improves the accuracy of echocardiogram video classification, meeting the clinical needs for real-time performance and information integrity.
Smart Images

Figure CN121033732A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an echocardiogram video classification method, system and terminal. BACKGROUND
[0002] Due to its non-invasiveness, affordability, convenience, real-time observation of cardiac structure movement changes, and real-time observation of cardiac blood flow through Doppler effect, echocardiogram is widely used in clinical diagnosis.
[0003] However, some groups (e.g. children) have fast heartbeats, and the quality of ultrasound images is low, resulting in problems such as blurred or missing structures in echocardiograms. Moreover, when classifying echocardiogram videos, the view information is not fully considered, the error of the labeled cross-section caused by the difference in operation methods of different doctors, and the processing of spatial features and temporal features of echocardiogram videos is not perfect, which leads to low accuracy of echocardiogram video classification in the prior art.
[0004] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0005] The main purpose of the present application is to provide an echocardiogram video classification method, system, terminal and computer readable storage medium, which aims to solve the problem of low accuracy of echocardiogram video classification in the prior art when classifying echocardiogram videos, without fully considering the view information, and the processing of spatial features and temporal features of echocardiogram videos is not perfect.
[0006] To achieve the above purpose, the present application provides an echocardiogram video classification method, which comprises the following steps:
[0007] Constructing a classification model network, the classification model network comprising a feature extraction module, a feature enhancement module and a feature aggregation module;
[0008] Obtaining an echocardiogram video, obtaining a plurality of standard cross-sectional views from the echocardiogram video, inputting the plurality of standard cross-sectional views into the feature extraction module of the classification model network, performing interpolation processing and feature extraction on the plurality of standard cross-sectional views through the feature extraction module, and outputting a plurality of video features;
[0009] Inputting the plurality of video features into the feature enhancement module of the classification model network, performing spatial feature and temporal feature aggregation enhancement on the plurality of video features through the feature enhancement module, and outputting a plurality of enhanced features;
[0010] Multiple enhanced features are input into the feature aggregation module of the classification model network. The feature aggregation module performs frame-level feature weighted fusion on the multiple enhanced features to obtain multiple key frame features. Relevant features are selected from the multiple key frame features. Based on the relevant features, fused features are obtained. The fused features are classified to obtain the classification result of the echocardiogram video.
[0011] Optionally, the echocardiogram video classification method, wherein acquiring the echocardiogram video and obtaining multiple standard cross-sectional views based on the echocardiogram video specifically includes:
[0012] Two-dimensional echocardiogram videos of the target user are acquired, and the acoustic windows of the two-dimensional echocardiogram videos are scanned according to different standards to obtain multiple standard cross-sectional views with different cross-sections.
[0013] The standard cross-sectional views include: parasternal short-axis cross-sectional view, parasternal long-axis cross-sectional view, apical four-chamber cross-sectional view, and apical five-chamber cross-sectional view.
[0014] Optionally, the echocardiogram video classification method, wherein the step of interpolating and extracting features from multiple standard cross-sectional views using the feature extraction module to output multiple video features specifically includes:
[0015] Multiple standard cross-sectional views are input into the feature extraction module, and the frame number of the multiple standard cross-sectional views is interpolated according to the cardiac cycle pattern to obtain a standard cross-sectional view with a unified frame number.
[0016] After unifying the number of frames, the standard cross-sectional view is input into the pre-trained ResNet18 model for feature extraction, resulting in feature maps of different scales. The feature maps of different scales are then normalized to obtain multiple video features.
[0017] Optionally, in the echocardiogram video classification method, the feature enhancement module includes multiple spatiotemporal feature aggregation enhancement sub-modules, and the aggregation enhancement sub-modules include temporal feature enhancement branches and spatial feature enhancement branches;
[0018] The feature enhancement module aggregates and enhances multiple video features through spatial and temporal features, outputting multiple enhanced features, specifically including:
[0019] Multiple video features are input into the spatiotemporal feature aggregation and enhancement submodule. The multiple video features are then averaged and pooled through the temporal feature enhancement branch to obtain multiple averaged pooled features.
[0020] The plurality of average-pooled features are fed back to an SE unit to obtain a plurality of time weighting factors, and the plurality of average-pooled features are fed back to a sliding window unit after a convolution layer to learn to obtain a plurality of multi-scale time sequence features; the plurality of time weighting factors and the plurality of multi-scale time sequence features are fused respectively to obtain a plurality of time sequence enhancement features;
[0021] The plurality of video features are convolved by the spatial feature enhancement branch to obtain a plurality of convolution features, multi-scale attention entropy of the plurality of convolution features is calculated, the plurality of multi-scale attention entropy is fed back to a gate selection unit to obtain a plurality of spatial weighting factors, and the plurality of spatial weighting factors and the plurality of video features are respectively multiplied by a matrix to obtain a plurality of spatial enhancement features.
[0022] The plurality of time sequence enhancement features and the plurality of spatial enhancement features are respectively added by a matrix and then down-sampled to obtain a plurality of intermediate features, and the plurality of intermediate features are respectively input again to a next spatio-temporal feature aggregation enhancement sub-module for aggregation and enhancement of spatial features and time sequence features, until all spatio-temporal feature aggregation enhancement sub-modules are passed through to obtain a plurality of enhancement features.
[0023] Optionally, the echocardiogram video classification method, wherein the feature aggregation module comprises a frame-level feature weighted fusion sub-module and a multi-view gated feature selection sub-module.
[0024] Optionally, the echocardiogram video classification method, wherein the frame-level feature weighted fusion of the plurality of enhancement features by the feature aggregation module to obtain a plurality of key frame features specifically comprises:
[0025] The plurality of enhancement features are respectively input to a frame-level feature weighted fusion sub-module, the plurality of enhancement features are processed by a linear layer, a plurality of attention entropy is calculated according to the processing result, a plurality of frame-level information quantity regression results are obtained by frame-level information quantity regression of the plurality of attention entropy, and a plurality of normalization results are obtained by normalization of the plurality of frame-level information quantity regression results.
[0026] A frame index with the highest probability value is obtained from each of the plurality of normalization results, and a key frame feature is obtained by slicing the normalization result corresponding to each of the frame indexes.
[0027] Optionally, the echocardiogram video classification method, wherein the multi-view gated feature selection sub-module comprises a cross-view cross-attention layer and a gated feature selection layer.
[0028] The relevant features are selected from the plurality of key frame features, and the fusion features are obtained according to the relevant features, specifically comprising:
[0029] The plurality of key frame features are input into a multi-view gated feature selection submodule, the plurality of key frame features are processed through a linear layer, the processing result is input into the cross-view cross-attention layer for learning and interaction, and an interaction feature is obtained;
[0030] The interaction feature is input into the gated feature selection layer for feature extraction, a related feature is obtained, the related feature is expanded through a linear layer, and the expanded feature is fused to obtain a fused feature.
[0031] In addition, to achieve the above-mentioned purpose, the application further provides an echocardiogram video classification system, wherein the echocardiogram video classification system comprises:
[0032] A model construction module is configured to construct a classification model network, wherein the classification model network comprises a feature extraction module, a feature enhancement module and a feature aggregation module.
[0033] The feature extraction module is configured to obtain an echocardiogram video, obtain a plurality of standard cross-sectional views from the echocardiogram video, input the plurality of standard cross-sectional views into the feature extraction module of the classification model network, perform interpolation processing and feature extraction on the plurality of standard cross-sectional views through the feature extraction module, and output a plurality of video features.
[0034] The feature enhancement module is configured to input the plurality of video features into the feature enhancement module of the classification model network, perform spatial feature and time sequence feature aggregation enhancement on the plurality of video features through the feature enhancement module, and output a plurality of enhanced features.
[0035] The feature aggregation module is configured to input the plurality of enhanced features into the feature aggregation module of the classification model network, perform frame-level feature weighted fusion on the plurality of enhanced features through the feature aggregation module, obtain a plurality of key frame features, select a related feature from the plurality of key frame features, obtain a fused feature according to the related feature, classify the fused feature, and obtain a classification result of the echocardiogram video.
[0036] In addition, to achieve the above-mentioned purpose, the application further provides a terminal, wherein the terminal comprises a memory, a processor, and an echocardiogram video classification program stored on the memory and executable on the processor, and the echocardiogram video classification program is used to implement the steps of the echocardiogram video classification method as described above when executed by the processor.
[0037] In addition, to achieve the above object, the application further provides a computer readable storage medium, wherein the computer readable storage medium stores an echocardiogram video classification program, and the echocardiogram video classification program realizes the steps of the echocardiogram video classification method when executed by a processor.
[0038] In the application, a classification model network is constructed, which includes a feature extraction module, a feature enhancement module and a feature aggregation module; an echocardiogram video is obtained, a plurality of standard cross-section views are obtained according to the echocardiogram video, interpolation processing and feature extraction are performed on the plurality of standard cross-section views by the feature extraction module, and a plurality of video features are output; the plurality of video features are input into the feature enhancement module for spatial feature and time sequence feature aggregation enhancement, and a plurality of enhanced features are output; the plurality of enhanced features are input into the feature aggregation module for frame-level feature weighted fusion, a plurality of key frame features are obtained, relevant features are selected from the plurality of key frame features, fusion features are obtained according to the relevant features, the fusion features are classified, and a classification result of the echocardiogram video is obtained. The application effectively improves the accuracy of echocardiogram video classification. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 is a flowchart of a preferred embodiment of the echocardiogram video classification method of the application;
[0040] Figure 2 is a general architecture diagram of the classification model network in the echocardiogram video classification method of the application;
[0041] Figure 3 is an architecture diagram of a spatio-temporal feature aggregation enhancement submodule in the echocardiogram video classification method of the application;
[0042] Figure 4 is an architecture diagram of a frame-level feature weighted fusion submodule in the echocardiogram video classification method of the application;
[0043] Figure 5 is an architecture diagram of a multi-view gating feature selection submodule in the echocardiogram video classification method of the application;
[0044] Figure 6 is a structure diagram of a preferred embodiment of the echocardiogram video classification system of the application;
[0045] Figure 7 is a structure diagram of a preferred embodiment of the terminal of the application. DETAILED DESCRIPTION
[0046] The application provides an echocardiogram video classification method and system and a terminal. To make the purpose, technical solutions and effects of the application clearer and more explicit, the application is further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and do not limit the application.
[0047] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood as having meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.
[0048] In addition, if the description of "first", "second" and the like is involved in the embodiments of the application, the description of "first", "second" and the like is only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include at least one of the features. In addition, the technical solutions of each embodiment can be combined with each other, but it must be based on the realization of ordinary skilled in the art, when the combination of technical solutions appears contradictory or unachievable, it should be considered that the combination of technical solutions does not exist, nor within the protection scope required by the application.
[0049] The echocardiogram video classification method described in the preferred embodiment of the application comprises the following steps: Figure 1 As shown in the echocardiogram video classification method, the echocardiogram video classification method comprises the following steps:
[0050] Step S10, constructing a classification model network, the classification model network comprising a feature extraction module, a feature enhancement module and a feature aggregation module.
[0051] As shown in the echocardiogram video classification method, the echocardiogram video classification method comprises the following steps: Figure 2 The classification model network constructed by the application is used to realize the echocardiogram video classification, and the classification model network comprises a feature extraction module, a feature enhancement module and a feature aggregation module.
[0052] The feature extraction module is configured to perform multi-scale feature extraction using a pre-trained ResNet18; the feature enhancement module is configured to take the single attempt extracted features as the input of the feature enhancement module, the time branch of the feature enhancement module uses a combination of a sliding window and an SE network structure to establish the correlation between frames, and the spatial branch uses attention entropy and a gating module to further represent and strengthen the original spatial features; the feature aggregation module is configured to take the enhanced features of the single attempt through a frame-level feature weighting fusion module to select key frame features; finally, the key frame features of the multi-view are fed into a gated feature selection module to obtain a final classification result, which can be used as a feature for diagnosing VSD (Ventricular Septal Defect).
[0053] In step S20, an echocardiogram video is obtained, a plurality of standard cross-section views are obtained from the echocardiogram video, the plurality of standard cross-section views are input into the feature extraction module of the classification model network, interpolation processing and feature extraction are performed on the plurality of standard cross-section views by the feature extraction module, and a plurality of video features are output.
[0054] The echocardiogram video is obtained, and a plurality of standard cross-section views are obtained from the echocardiogram video, specifically including:
[0055] A two-dimensional echocardiogram video of a target user is obtained, and the acoustic window of the two-dimensional echocardiogram video is scanned according to different standards to obtain a plurality of standard cross-section views of different cross sections.
[0056] The standard cross-section views include a parasternal short-axis view, a parasternal long-axis view, an apical four-chamber view, and an apical five-chamber view.
[0057] In this embodiment, for one two-dimensional echocardiogram video, the acoustic window of the echocardiogram video is scanned according to different standards, and different cross-section echocardiogram videos can be obtained after the acoustic window is scanned according to different standards for clinical ultrasonic examination, Figure 2 The parasternal short-axis view and the apical five-chamber view are taken as examples.
[0058] Further, the interpolation processing and feature extraction of the plurality of standard cross-section views by the feature extraction module to output a plurality of video features specifically include:
[0059] The plurality of standard cross-section views are input into the feature extraction module respectively, the frame numbers of the plurality of standard cross-section views are interpolated according to the heart cycle rule, and a plurality of standard cross-section views with uniform frame numbers are obtained.
[0060] The multiple frame number unified standard cross-section views are input into a pre-trained ResNet18 model for feature extraction to obtain feature maps of different scales, and the feature maps of different scales are normalized to obtain multiple video features.
[0061] In the embodiment, first, for a two-dimensional echocardiogram video, the frame number is interpolated according to the law of the cardiac cycle to realize the uniformity of the frame number, and then the frames are processed in time sequence, and the size of each frame is 224*224. Then, a ResNet18 model pre-trained on ImageNet is used to extract multi-scale features from each frame of the frame number unified standard cross-section view for subsequent processing.
[0062] In step S30, the multiple video features are input into the feature enhancement module of the classification model network, the multiple video features are aggregated and enhanced by the feature enhancement module in spatial features and time sequence features, and multiple enhanced features are output.
[0063] It can be understood that feature enhancement is a key link to improve the representation ability of video features. The multi-scale feature maps obtained from the feature extraction module only contain spatial information of each frame, and the time sequence of the video is not modeled. The five spatial-temporal feature aggregation and enhancement sub-modules used in the application capture the relationship between adjacent frames in the video. The network structure of the spatial-temporal feature aggregation and enhancement sub-module is as shown in Figure 3 The feature enhancement module includes multiple spatial-temporal feature aggregation and enhancement sub-modules, and the aggregation and enhancement sub-modules include time feature enhancement branches and spatial feature enhancement branches. The aggregation and enhancement sub-modules will enhance the feature maps multiple times to obtain multi-scale enhanced features. For example, the initial features will be enhanced by the spatial-temporal feature enhancement module to obtain the first enhanced feature map, which is down-sampled and used as the input of the similar next spatial-temporal enhancement module, that is, the intermediate feature. The intermediate feature is enhanced by the spatial-temporal feature enhancement module to obtain the second enhanced feature map. Compared with the first enhanced feature map, the spatial resolution is halved to obtain more abstract semantic feature representation. The spatial-temporal enhancement module is used for 5 times to enhance the features of 5 scales.
[0064] The multiple video features are input into the spatial-temporal feature aggregation and enhancement sub-module, the multiple video features are average-pooled by the time feature enhancement branch to obtain multiple average-pooled features.
[0065] The multiple video features are input into the spatial-temporal feature aggregation and enhancement sub-module, the multiple video features are average-pooled by the time feature enhancement branch to obtain multiple average-pooled features.
[0066] The plurality of average-pooled features are fed back to an SE unit to obtain a plurality of time weighting factors, and the plurality of average-pooled features are fed back to a sliding window unit after a convolution layer to learn a plurality of multi-scale time sequence features, the plurality of time weighting factors and the plurality of multi-scale time sequence features are fused respectively to obtain a plurality of time sequence enhancement features;
[0067] The plurality of video features are convolved by the spatial feature enhancement branch to obtain a plurality of convolution features, a multi-scale attention entropy of the plurality of convolution features is calculated, the plurality of multi-scale attention entropies are fed back to a gate selection unit to obtain a plurality of spatial weighting factors, and the plurality of spatial weighting factors and the plurality of video features are respectively multiplied in matrix to obtain a plurality of spatial enhancement features.
[0068] The plurality of time sequence enhancement features and the plurality of spatial enhancement features are respectively added in matrix and then down-sampled to obtain a plurality of intermediate features, and the plurality of intermediate features are respectively input again to a next spatio-temporal feature aggregation enhancement sub-module for spatial feature and time sequence feature aggregation enhancement, until all the spatio-temporal feature aggregation enhancement sub-modules are passed through to obtain a plurality of enhancement features.
[0069] In the embodiment, first, in the time feature enhancement branch, the plurality of video features F i (i=1, 2, 3, 4, F1, F2, F3, F4 respectively correspond to video features of four different standard section views) are average-pooled and fed to an SE module to obtain time weighting factors The average-pooled features are fed to a sliding window module after a convolution layer, the sliding window module selects 1, 2, and 4 as the window size, and different window sizes can enable the model to learn time sequence information of different time scales, and then are multiplied in matrix to obtain enhanced fusion time sequence information features
[0070]
[0071] Wherein, F i indicates the i-th level feature of a single view, AvgPool indicates average pooling, SE represents an SE module, the SE module is used for enhancing the time features after average pooling, enriching the representation of the time features, and SW indicates a sliding window module.
[0072] Further, in the spatial feature enhancement branch, a multi-scale attention entropy is calculated from the convolved multi-scale spatial features, and a spatial weighting factor is obtained from the gate selection module The weighting coefficient and the original spatial feature are multiplied in matrix to obtain enhanced spatial features
[0073]
[0074] In this context, MAE refers to the multi-scale attention entropy calculation module, Gate refers to the gating module, and * indicates matrix multiplication.
[0075] Furthermore, the enhanced temporal characteristics and Matrix addition yields enhanced features for feature aggregation in subsequent tasks.
[0076]
[0077] Here, + indicates matrix addition.
[0078] Step S40: Input multiple enhanced features into the feature aggregation module of the classification model network, perform frame-level feature weighted fusion on the multiple enhanced features through the feature aggregation module to obtain multiple key frame features, select relevant features from the multiple key frame features, obtain fused features based on the relevant features, classify the fused features, and obtain the classification result of the echocardiogram video.
[0079] In this embodiment, multiple enhanced features are respectively subjected to frame-level feature weighted fusion and gated feature selection to achieve feature aggregation. Specifically, the feature aggregation module includes a frame-level feature weighted fusion submodule and a multi-view gated feature selection submodule.
[0080] like Figure 4 As shown, the step of performing frame-level feature weighted fusion of multiple enhanced features through the feature aggregation module to obtain multiple key frame features specifically includes:
[0081] Multiple enhancement features are input into a frame-level feature weighted fusion submodule, and processed through a linear layer. Multiple attention entropies are calculated based on the processing results. Frame-level information content regression is then performed on these attention entropies to obtain multiple frame-level information content regression results. These regression results are then normalized to obtain multiple normalized results p. i ;
[0082] The frame index with the highest probability value is obtained from the multiple normalization results, and the normalization results corresponding to each of the multiple frame indices are sliced to obtain multiple key frame features.
[0083] Understandably, feature enhancement Temporal information needs to be aggregated. The frame-level feature weighted fusion module calculates the attention entropy, assuming that the amount of effective information contained in each frame of the video varies. For example, the end of systole and end of diastole in the heart contains more information. The frame-level feature weighted fusion submodule proposed in this application regresses the amount of information at the frame level, thereby selecting the key frame features F related to VSD. key Used for subsequent view-level blending.
[0084] The calculation process of performing frame-level feature weighted fusion of multiple enhanced features through the feature aggregation module to obtain multiple key frame features is expressed as follows:
[0085]
[0086] Where FC represents linear layer processing, AE refers to the attention entropy calculation module, Reg represents the information regression module, Norm refers to the normalization operation, Argmax refers to taking the frame index with the highest probability value after normalization, and F i [.] indicates that keyframe features are extracted by slicing based on the index.
[0087] Furthermore, such as Figure 5 As shown, the multi-view gating feature selection submodule includes: a cross-view cross-attention layer and a gating feature selection layer.
[0088] The step of selecting relevant features from multiple keyframe features and obtaining fused features based on the relevant features specifically includes:
[0089] Multiple keyframe features are input into the multi-view gating feature selection submodule, and the multiple keyframe features are processed through a linear layer. The processing result is input into the cross-view cross-attention layer for learning and interaction to obtain interactive features.
[0090] The interactive features are input into the gated feature selection layer for feature extraction to obtain relevant features. The relevant features are then expanded through a linear layer, and the expanded features are fused to obtain fused features.
[0091] In this embodiment, keyframe features from multiple cross-views are learned using a cross-view attention layer, enabling the features from multiple views to be fully fused and interacted. The most relevant features are then extracted by a gating selection module and fed into a linear layer to expand the feature representation, finally yielding the fused feature F. out It will be sent to a classifier to complete the VSD-based video image classification.
[0092] The calculation process of selecting relevant features from multiple keyframe features and obtaining fused features based on the relevant features is expressed as follows:
[0093]
[0094] in, The keyframe features of the i-th view are represented by a sequence of keyframe features from multiple views. CCA represents the cross-view attention layer, GFS represents the gated feature selection layer, and F represents the keyframe feature selection layer. out This represents the fused features ultimately used for classification.
[0095] Furthermore, classifying the target's echocardiographic video using the classification model network further includes: constructing a dataset and training the classification model network using the dataset.
[0096] Specifically, this application uses the private dataset Echo-VSD-P, containing 2338 samples. Each sample contains only four different standard views (parasternal short-axis view, parasternal long-axis view, apical five-chamber view, and apical four-chamber view), and each view contains two modalities (grayscale ultrasound and color Doppler ultrasound). Each sample was annotated with ventricular septal defects by a professional echocardiologist. In the Echo-VSD-P dataset, the data was randomly divided into training and test sets at a 4:1 ratio, and five-fold cross-validation was used on the training set. The model was trained using the PyTorch framework on a TITAN RTX 3090 GPU with 24GB of memory. During training, the initial learning rate was set to 1*10. -4 and utilizing 10 -3 The learning rate decay mechanism of the decay coefficient was described. Furthermore, to optimize model performance, the AdamW optimizer was used, and cross-entropy loss was used as the objective function during model training. The experiment consisted of 50 batches, with a batch size of 32 in each batch.
[0097] As can be seen, this application proposes a classification network for multimodal echocardiographic videos with multiple views, ensuring information integrity while meeting the real-time requirements of clinical ultrasound image classification. Considering the importance of temporal information, an aggregation and enhancement module for spatial and temporal features, as well as a frame-level feature weighted fusion module, are proposed to model the motion information of pediatric cardiac structures and avoid structural blurring caused by using a single image. Furthermore, considering the comprehensiveness of information from multiple views, a gated feature selection module is proposed to perform weighted fusion of features from four views. Validation of the proposed model on a private dataset demonstrates that the method exhibits better performance than existing models.
[0098] Furthermore, such as Figure 6 As shown, based on the above-mentioned echocardiogram video classification method, the present invention also provides an echocardiogram video classification system, wherein the echocardiogram video classification system includes:
[0099] The model construction module 51 is used to construct a classification model network, which includes a feature extraction module, a feature enhancement module, and a feature aggregation module.
[0100] Feature extraction module 52 is used to acquire echocardiogram video, obtain multiple standard cross-sectional views based on the echocardiogram video, input the multiple standard cross-sectional views into the feature extraction module of the classification model network, and perform interpolation processing and feature extraction on the multiple standard cross-sectional views through the feature extraction module to output multiple video features;
[0101] The feature enhancement module 53 is used to input multiple video features into the feature enhancement module of the classification model network, and to perform spatial and temporal feature aggregation enhancement on the multiple video features through the feature enhancement module, and output multiple enhanced features;
[0102] The feature aggregation module 54 is used to input multiple enhanced features into the feature aggregation module of the classification model network, perform frame-level feature weighted fusion on the multiple enhanced features through the feature aggregation module to obtain multiple key frame features, select relevant features from the multiple key frame features, obtain fused features based on the relevant features, classify the fused features, and obtain the classification result of the echocardiogram video.
[0103] Furthermore, such as Figure 7 As shown, based on the above-mentioned echocardiogram video classification method and system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 7 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0104] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores an echocardiogram video classification program 40, which can be executed by the processor 10 to implement the echocardiogram video classification method of this application.
[0105] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the echocardiogram video classification method.
[0106] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0107] In one embodiment, when the processor 10 executes the echocardiogram video classification program 40 in the memory 20, the following steps are performed:
[0108] A classification model network is constructed, which includes a feature extraction module, a feature enhancement module, and a feature aggregation module;
[0109] An echocardiogram video is acquired, and multiple standard cross-sectional views are obtained from the echocardiogram video. The multiple standard cross-sectional views are input into the feature extraction module of the classification model network. The feature extraction module performs interpolation processing and feature extraction on the multiple standard cross-sectional views and outputs multiple video features.
[0110] Multiple video features are input into the feature enhancement module of the classification model network. The feature enhancement module performs spatial and temporal feature aggregation enhancement on the multiple video features and outputs multiple enhanced features.
[0111] Multiple enhanced features are input into the feature aggregation module of the classification model network. The feature aggregation module performs frame-level feature weighted fusion on the multiple enhanced features to obtain multiple key frame features. Relevant features are selected from the multiple key frame features. Based on the relevant features, fused features are obtained. The fused features are classified to obtain the classification result of the echocardiogram video.
[0112] The acquisition of echocardiographic video, specifically including obtaining multiple standard cross-sectional views from the echocardiographic video, includes:
[0113] Two-dimensional echocardiogram videos of the target user are acquired, and the acoustic windows of the two-dimensional echocardiogram videos are scanned according to different standards to obtain multiple standard cross-sectional views with different cross-sections.
[0114] The standard cross-sectional views include: parasternal short-axis cross-sectional view, parasternal long-axis cross-sectional view, apical four-chamber cross-sectional view, and apical five-chamber cross-sectional view.
[0115] Specifically, the step of interpolating and extracting features from multiple standard cross-sectional views using the feature extraction module to output multiple video features includes:
[0116] Multiple standard cross-sectional views are input into the feature extraction module, and the frame number of the multiple standard cross-sectional views is interpolated according to the cardiac cycle pattern to obtain a standard cross-sectional view with a unified frame number.
[0117] After unifying the number of frames, the standard cross-sectional view is input into the pre-trained ResNet18 model for feature extraction, resulting in feature maps of different scales. The feature maps of different scales are then normalized to obtain multiple video features.
[0118] The feature enhancement module includes multiple spatiotemporal feature aggregation enhancement sub-modules, and the aggregation enhancement sub-modules include temporal feature enhancement branches and spatial feature enhancement branches;
[0119] The feature enhancement module aggregates and enhances multiple video features through spatial and temporal features, outputting multiple enhanced features, specifically including:
[0120] Multiple video features are input into the spatiotemporal feature aggregation and enhancement submodule. The multiple video features are then averaged and pooled through the temporal feature enhancement branch to obtain multiple averaged pooled features.
[0121] Multiple average pooled features are fed back to the SE unit to obtain multiple time weighting factors. Multiple average pooled features are then fed back to the sliding window unit after passing through a convolutional layer to obtain multiple multi-scale temporal features. Multiple time weighting factors and multiple multi-scale temporal features are fused to obtain multiple temporal enhancement features.
[0122] The spatial feature enhancement branch convolves multiple video features to obtain multiple convolutional features. The multi-scale attention entropy of the multiple convolutional features is calculated. The multiple multi-scale attention entropy is fed back to the gating selection unit to obtain multiple spatial weighting factors. The multiple spatial weighting factors are matrix-multiplied with the multiple video features respectively to obtain multiple spatial enhancement features.
[0123] Multiple temporal enhancement features and multiple spatial enhancement features are respectively matrix-added and then downsampled to obtain multiple intermediate features. These intermediate features are then input into the next spatiotemporal feature aggregation enhancement submodule for aggregation enhancement of spatial and temporal features, until multiple enhanced features are obtained after passing through all spatiotemporal feature aggregation enhancement submodules.
[0124] The feature aggregation module includes a frame-level feature weighted fusion submodule and a multi-view gated feature selection submodule.
[0125] Specifically, the step of performing frame-level feature weighted fusion of multiple enhanced features through the feature aggregation module to obtain multiple key frame features includes:
[0126] Multiple enhancement features are input into the frame-level feature weighted fusion submodule, and the multiple enhancement features are processed through a linear layer. Multiple attention entropies are calculated based on the processing results. Frame-level information regression is performed on the multiple attention entropies to obtain multiple frame-level information regression results. The multiple frame-level information regression results are normalized to obtain multiple normalized results.
[0127] The frame index with the highest probability value is obtained from the multiple normalization results, and the normalization results corresponding to each of the multiple frame indices are sliced to obtain multiple key frame features.
[0128] The multi-view gated feature selection submodule includes: a cross-view cross-attention layer and a gated feature selection layer;
[0129] The step of selecting relevant features from multiple keyframe features and obtaining fused features based on the relevant features specifically includes:
[0130] Multiple keyframe features are input into the multi-view gating feature selection submodule, and the multiple keyframe features are processed through a linear layer. The processing result is input into the cross-view cross-attention layer for learning and interaction to obtain interactive features.
[0131] The interactive features are input into the gated feature selection layer for feature extraction to obtain relevant features. The relevant features are then expanded through a linear layer, and the expanded features are fused to obtain fused features.
[0132] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores an echocardiogram video classification program, which, when executed by a processor, implements the steps of the echocardiogram video classification method as described above.
[0133] In summary, this invention provides a method, system, and terminal for classifying echocardiographic videos. The method includes: constructing a classification model network, which includes a feature extraction module, a feature enhancement module, and a feature aggregation module; acquiring echocardiographic videos; obtaining multiple standard cross-sectional views from the echocardiographic videos; performing interpolation processing and feature extraction on the multiple standard cross-sectional views through the feature extraction module to output multiple video features; inputting the multiple video features into the feature enhancement module for aggregation and enhancement of spatial and temporal features to output multiple enhanced features; inputting the multiple enhanced features into the feature aggregation module for frame-level feature weighted fusion to obtain multiple keyframe features; selecting relevant features from the multiple keyframe features; obtaining fused features based on the relevant features; classifying the fused features to obtain the classification result of the echocardiographic video. This invention effectively improves the accuracy of echocardiographic video classification.
[0134] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal that includes that element.
[0135] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0136] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.
Claims
1. A method for classifying echocardiographic videos, characterized in that, The aforementioned echocardiogram video classification method includes: A classification model network is constructed, which includes a feature extraction module, a feature enhancement module, and a feature aggregation module; An echocardiogram video is acquired, and multiple standard cross-sectional views are obtained from the echocardiogram video. The multiple standard cross-sectional views are input into the feature extraction module of the classification model network. The feature extraction module performs interpolation processing and feature extraction on the multiple standard cross-sectional views and outputs multiple video features. Multiple video features are input into the feature enhancement module of the classification model network. The feature enhancement module performs spatial and temporal feature aggregation enhancement on the multiple video features and outputs multiple enhanced features. Multiple enhanced features are input into the feature aggregation module of the classification model network. The feature aggregation module performs frame-level feature weighted fusion on the multiple enhanced features to obtain multiple key frame features. Relevant features are selected from the multiple key frame features. Based on the relevant features, fused features are obtained. The fused features are classified to obtain the classification result of the echocardiogram video.
2. The echocardiogram video classification method according to claim 1, characterized in that, The acquisition of echocardiographic video, and the generation of multiple standard cross-sectional views from the echocardiographic video, specifically includes: Two-dimensional echocardiogram videos of the target user are acquired, and the acoustic windows of the two-dimensional echocardiogram videos are scanned according to different standards to obtain multiple standard cross-sectional views with different cross-sections. The standard cross-sectional views include: parasternal short-axis cross-sectional view, parasternal long-axis cross-sectional view, apical four-chamber cross-sectional view, and apical five-chamber cross-sectional view.
3. The echocardiogram video classification method according to claim 1, characterized in that, The step of interpolating and extracting features from multiple standard cross-sectional views using the feature extraction module to output multiple video features specifically includes: Multiple standard cross-sectional views are input into the feature extraction module, and the frame number of the multiple standard cross-sectional views is interpolated according to the cardiac cycle pattern to obtain a standard cross-sectional view with a unified frame number. After unifying the number of frames, the standard cross-sectional view is input into the pre-trained ResNet18 model for feature extraction, resulting in feature maps of different scales. The feature maps of different scales are then normalized to obtain multiple video features.
4. The echocardiogram video classification method according to claim 1, characterized in that, The feature enhancement module includes multiple spatiotemporal feature aggregation enhancement sub-modules, and the aggregation enhancement sub-modules include temporal feature enhancement branches and spatial feature enhancement branches; The feature enhancement module aggregates and enhances multiple video features through spatial and temporal features, outputting multiple enhanced features, specifically including: Multiple video features are input into the spatiotemporal feature aggregation and enhancement submodule. The multiple video features are then averaged and pooled through the temporal feature enhancement branch to obtain multiple averaged pooled features. Multiple average pooled features are fed back to the SE unit to obtain multiple time weighting factors. Multiple average pooled features are then fed back to the sliding window unit after passing through a convolutional layer to obtain multiple multi-scale temporal features. Multiple time weighting factors and multiple multi-scale temporal features are fused to obtain multiple temporal enhancement features. The spatial feature enhancement branch convolves multiple video features to obtain multiple convolutional features. The multi-scale attention entropy of the multiple convolutional features is calculated. The multiple multi-scale attention entropy is fed back to the gating selection unit to obtain multiple spatial weighting factors. The multiple spatial weighting factors are matrix-multiplied with the multiple video features respectively to obtain multiple spatial enhancement features. Multiple temporal enhancement features and multiple spatial enhancement features are respectively matrix-added and then downsampled to obtain multiple intermediate features. These intermediate features are then input into the next spatiotemporal feature aggregation enhancement submodule for aggregation enhancement of spatial and temporal features, until multiple enhanced features are obtained after passing through all spatiotemporal feature aggregation enhancement submodules.
5. The echocardiogram video classification method according to claim 1, characterized in that, The feature aggregation module includes a frame-level feature weighted fusion submodule and a multi-view gated feature selection submodule.
6. The echocardiogram video classification method according to claim 5, characterized in that, The step of performing frame-level feature weighted fusion of multiple enhanced features through the feature aggregation module to obtain multiple key frame features specifically includes: Multiple enhancement features are input into the frame-level feature weighted fusion submodule, and the multiple enhancement features are processed through a linear layer. Multiple attention entropies are calculated based on the processing results. Frame-level information regression is performed on the multiple attention entropies to obtain multiple frame-level information regression results. The multiple frame-level information regression results are normalized to obtain multiple normalized results. The frame index with the highest probability value is obtained from the multiple normalization results, and the normalization results corresponding to each of the multiple frame indices are sliced to obtain multiple key frame features.
7. The echocardiogram video classification method according to claim 5, characterized in that, The multi-view gated feature selection submodule includes: a cross-view cross-attention layer and a gated feature selection layer; The step of selecting relevant features from multiple keyframe features and obtaining fused features based on the relevant features specifically includes: Multiple keyframe features are input into the multi-view gating feature selection submodule, and the multiple keyframe features are processed through a linear layer. The processing result is input into the cross-view cross-attention layer for learning and interaction to obtain interactive features. The interactive features are input into the gated feature selection layer for feature extraction to obtain relevant features. The relevant features are then expanded through a linear layer, and the expanded features are fused to obtain fused features.
8. An echocardiogram video classification system, characterized in that, The echocardiographic video classification system includes: The model construction module is used to build a classification model network, which includes a feature extraction module, a feature enhancement module, and a feature aggregation module. The feature extraction module is used to acquire echocardiogram videos, obtain multiple standard cross-sectional views from the echocardiogram videos, input the multiple standard cross-sectional views into the feature extraction module of the classification model network, and perform interpolation processing and feature extraction on the multiple standard cross-sectional views through the feature extraction module to output multiple video features; The feature enhancement module is used to input multiple video features into the feature enhancement module of the classification model network, and to perform spatial and temporal feature aggregation enhancement on the multiple video features through the feature enhancement module, and output multiple enhanced features; The feature aggregation module is used to input multiple enhanced features into the feature aggregation module of the classification model network, perform frame-level feature weighted fusion on the multiple enhanced features through the feature aggregation module to obtain multiple key frame features, select relevant features from the multiple key frame features, obtain fused features based on the relevant features, classify the fused features, and obtain the classification result of the echocardiogram video.
9. A terminal, characterized in that, The terminal includes: a memory, a processor, and an echocardiogram video classification program stored in the memory and executable on the processor, wherein the echocardiogram video classification program, when executed by the processor, implements the steps of the echocardiogram video classification method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a filter detection program based on a temperature sensor, which, when executed by a processor, implements the steps of the filter detection method based on a temperature sensor as described in any one of claims 1-7.
Citation Information
Cited By
Feature fusion method for intelligent identification of pulmonary valve stenosis echocardiogram
CN121661455A
A feature fusion method for intelligent identification of pulmonary valve stenosis echocardiogram
CN121661455B