A medical video classification method based on multi-objective segmentation prior drive

By constructing a multi-objective segmentation-driven medical video classification method, using deep residual convolution network and multi-head self-attention mechanism, the problem of difficulty in capturing spatiotemporal semantic interactions in medical video classification is solved, and high accuracy and robustness recognition in complex scenarios is achieved.

CN120411862BActive Publication Date: 2025-08-26FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510905487.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-26
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing medical video classification methods are difficult to effectively capture spatiotemporal semantic interactions, and cannot meet the dual requirements of precision and interpretability of medical scenarios, especially in dynamic multi-objective scenarios, which are difficult to identify key features.

Method used

Using a medical video classification method based on multi-objective segmentation prior drive, a multi-objective segmentation-classification data set is constructed, and a deep residual convolution network and a multi-head self-attention mechanism is used to perform feature extraction, timing encoding and fusion, thereby improving the capture ability of long-distance dependencies between video frames.

Benefits of technology

It significantly enhances the accuracy and robustness of medical video classification, can effectively identify multi-objective interaction patterns, improves the model's understanding of complex scenarios, and has good clinical application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411862B_ABST
    Figure CN120411862B_ABST
Patent Text Reader

Abstract

The present invention relates to a medical video classification method based on multi-target segmentation prior-driven classification, belonging to the field of video recognition technology. The method comprises the following steps: acquiring and preprocessing a medical video frame sequence to obtain a raw video frame sequence; constructing a multi-target segmentation-classification dataset and dividing it into a training set and a test set; constructing a medical video classification model based on multi-target segmentation prior-driven classification, the medical video classification model comprising a feature extraction-temporal coding-fusion network and a classifier; training the model using data in the training set; during the training process, optimizing the model using a loss function to obtain a trained model; and inputting data from the test set into the trained model to obtain a medical video classification result. The present invention can enhance the accuracy and generalization ability of medical video classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video recognition, and in particular relates to a medical video classification method based on multi-target segmentation prior drive. Background Art

[0002] Medical videos, as important carriers of dynamic visual information in clinical diagnosis, are characterized by complex scene changes, significant multi-target coupling and linkage, and a high degree of interweaving of spatiotemporal information. Compared with static images, medical videos can more comprehensively capture the dynamic characteristics of lesions, helping to identify the morphological evolution of organs at different stages, and providing a more explanatory basis for the accurate auxiliary diagnosis of systemic diseases such as the heart. Therefore, medical video classification methods must not only extract local structural features within the frame but also be able to model the coordinated movement and deformation patterns of organ tissues in the temporal dimension, thereby supporting the recognition and discrimination of diverse disease patterns and meeting clinical needs for interpretability and practicality.

[0003] Although video classification research has made progress in model architecture, feature expression, and context modeling in recent years, medical videos still face several key challenges in practical applications. For example, due to the presence of dynamic background interference, multi-target structural linkage, and long-term dependencies, existing methods find it difficult to effectively capture spatiotemporal semantic interactions, which in turn limits the model's diagnostic discrimination capabilities. In addition, medical videos often have problems such as blurred boundaries and insignificant target features, which are particularly evident in dynamic multi-target scenarios. For example, in cardiac ultrasound videos, key features such as subtle abnormalities in cardiac chamber motion, changes in local myocardial contraction, or mild coronary artery stenosis often contain important diagnostic signals, but are extremely difficult to accurately identify with traditional methods in low-contrast or high-noise environments.

[0004] While existing methods have attempted to introduce temporal modeling mechanisms, such as recurrent neural networks (RNNs) and temporal convolutional networks (TCNs), to model local temporal dependencies, these methods typically perform well in short time series and struggle to capture structural changes over long time series. Furthermore, when processing multi-target videos, existing models make limited use of structural priors, making it difficult to achieve decoupled representations and semantic fusion of different organ targets. Furthermore, current mainstream video classification models for natural scenes often perform poorly when transferred to the medical field due to a lack of medical semantic modeling capabilities and structural constraints, failing to meet the dual requirements of accuracy and interpretability in medical scenarios. Summary of the Invention

[0005] In order to solve the above problems, the present invention provides a medical video classification method based on multi-target segmentation prior drive.

[0006] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions:

[0007] The present invention provides a medical video classification method based on multi-target segmentation prior drive, comprising the following steps:

[0008] S1. Acquire and preprocess medical video frame sequences to obtain original video frame sequences;

[0009] S2. Construct a multi-object segmentation-classification dataset and divide it into a training set and a test set. The dataset includes a video original frame sequence, a single-object segmentation mask sequence, a multi-object segmentation mask sequence, and a category label;

[0010] S3. Construct a medical video classification model based on multi-target segmentation prior-driven modeling, comprising a feature extraction-temporal coding-fusion network and a classifier; the feature extraction-temporal coding-fusion network comprises a feature extraction module, a temporal coding module, and a feature fusion module; the temporal coding module comprises a temporal information embedding layer, a spatiotemporal fusion information encoding layer, and an MLP projection layer; and train the model using data from the training set;

[0011] S4. During the training process, the model is optimized using the loss function to obtain a trained model.

[0012] S5. Input the data in the test set into the trained model to obtain the medical video classification results.

[0013] Furthermore, in step S1, the sampling frequency and resolution of each sample frame sequence of the collected original medical video are unified, and each frame image is preprocessed by histogram equalization and pixel intensity standardization to obtain the original frame sequence of the video, and the category label is obtained according to the diagnosis result report.

[0014] Furthermore, step S2 specifically includes:

[0015] The existing video target segmentation algorithm is used to segment the original video frame sequence. Then, through the method of inspection and correction by domain experts, a single-object segmentation mask sequence for the target anatomical structure and a multi-object segmentation mask sequence for the visual objects of interest in the scene are obtained. The multiple single-object segmentation mask sequences and one multi-object segmentation mask sequence are combined into a matrix as the input of the model.

[0016] Furthermore, the feature extraction module in step S3 specifically includes:

[0017] The feature extraction module uses the first five layers of the deep residual convolutional network ResNet50 as the backbone network. In the first five layers of the deep residual convolutional network ResNet50, the first convolution block uses a 5×5 convolution kernel for feature extraction; the second to fifth convolution blocks use Inception convolution, and channel-space mixed attention is added after the fifth convolution block to assign different weights to the pixel values ​​of each part of the frame image; the step size of the second and third convolution blocks is set to 2, and the step size of the fourth and fifth convolution blocks is set to 1;

[0018] The parallel single-object segmentation mask sequence and multi-object segmentation mask sequence are extracted by the feature extraction module from each frame image of each segmentation mask sequence to generate a single-target segmentation feature map sequence. and multi-target segmentation feature map sequence ;

[0019] In the process of inputting the single object segmentation mask sequence into the feature extraction module to extract the frame image features, the calculation formulas of channel attention and spatial attention are expressed as follows:

[0020] ,

[0021] ,

[0022] in, represents the output of the fifth convolutional block, represents the channel attention weight, represents the spatial attention weight, represents a multilayer perceptron, represents average pooling along the channel dimension, represents the maximum pooling along the channel dimension, Indicates that the convolution kernel size is The convolution operation, Represents a splicing operation, Represents the Sigmoid activation function; the output of the fifth convolution block is distributed through attention weights to obtain a sequence of single-target segmentation feature maps. , the formula is as follows:

[0023] ,

[0024] in, Represents element-by-element multiplication; similarly, we get a sequence of multi-target segmentation feature maps .

[0025] Furthermore, the time information embedding layer in step S3 specifically includes:

[0026] Single object segmentation feature map sequence and multi-target segmentation feature map sequence The parallel input is processed into the time information embedding layer. By introducing the temporal position encoding, the time sequence information is given to the feature map sequence, and the spatiotemporal fusion information feature representation of the single target segmentation feature map sequence and the spatiotemporal fusion information feature representation of the multi-target segmentation feature map sequence are obtained. The formula is expressed as follows:

[0027] ,

[0028] ,

[0029] in, Represents a positional encoding containing timing information, Indicates the position number of the video frame in the video frame sequence. Indicates the length of the video sequence; Represents the spatiotemporal fusion information feature representation of a single target segmentation feature map sequence, Represents the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences.

[0030] Furthermore, the spatiotemporal fusion information coding layer in step S3 specifically includes:

[0031] The spatiotemporal fusion information coding layer includes multiple channels, each channel is composed of several VTN encoder layers stacked in series, and for each VTN encoder layer, multi-head self-attention is used to perform spatiotemporal information coding fusion calculation; the spatiotemporal fusion information feature representation of the single target segmentation feature map sequence The calculation process in the first VTN encoder layer of the spatiotemporal fusion information encoding layer is as follows:

[0032] ,

[0033] ,

[0034] ,

[0035] ,

[0036] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively; , , Respectively represent the first weight matrix, the second weight matrix, and the third weight matrix that can be learned; represents the scaling factor; for Attention heads, the self-attention calculation formula is expressed as: ; The outputs of all attention heads are concatenated and then input into the linear layer for further fusion transformation and dimension processing to obtain the output of the multi-head self-attention layer , the formula is as follows:

[0037] ,

[0038] in, Represents a splicing operation, Represents the trainable fourth weight matrix; the output of the multi-head self-attention layer is processed through residual connection and layer normalization to obtain the normalized feature representation , the formula is as follows:

[0039] ,

[0040] in, Normalization operation of the representation layer; input the normalized feature representation into the feedforward neural network to extract and transform the features to obtain the features output by the feedforward neural network. The formula is as follows:

[0041] ,

[0042] in, represents the characteristics of the feedforward neural network output, represents a feedforward neural network, represents the bias vector of the first layer of the feedforward neural network, represents the weight matrix of the second layer of the feedforward neural network, represents the bias vector of the second layer of the feedforward neural network, Represents the ReLU activation function; the characteristics of the feedforward neural network output And the normalized feature representation Through residual connection and layer normalization, we get The output of the VTN encoding module is expressed as follows:

[0043] ,

[0044] in, represents the output of the first layer VTN encoding module of the kth frame of the i-th single target segmentation feature map sequence;

[0045] Stack the above process Layer, get the global spatiotemporal feature matrix of the single target segmentation feature map sequence ,in, , n represents the number of single target segmentation feature map sequences; similarly, the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences After processing through the spatiotemporal fusion information coding layer, the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence is obtained .

[0046] Furthermore, the MLP projection layer in step S3 specifically includes:

[0047] The MLP projection layer includes two fully connected layers and a ReLU activation function;

[0048] Global spatiotemporal feature matrix of single target segmentation feature map sequence and the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence After processing by the MLP projection layer, the single target segmentation fusion feature is obtained and multi-target segmentation fusion features .

[0049] Furthermore, the feature fusion module in step S3 specifically includes:

[0050] Fusion features for single target segmentation And multi-target segmentation fusion feature output Perform splicing to obtain the splicing matrix , the formula is as follows:

[0051] ,

[0052] The stitching matrix Generate the overall spatiotemporal features of the video through a fully connected layer and activation function , the formula is as follows:

[0053] ,

[0054] in, represents the trainable weight matrix of the fully connected layer, represents the trainable bias matrix, Represents the ReLU activation function.

[0055] Furthermore, in step S3, the classifier is a fully connected layer including a Softmax activation function, and the overall spatiotemporal features of the video The classifier outputs the category prediction results of the medical video, and the formula is as follows:

[0056] ,

[0057] in, represents the weight matrix of the classifier, represents the bias vector, represents the Softmax activation function, Represents the output category prediction value.

[0058] Furthermore, in step S4, the loss function uses cross entropy loss as the objective function to train the model, and the formula is as follows:

[0059] ,

[0060] in, represents the category prediction value of the jth sample, represents the category label of the jth sample, represents the number of training set samples, represents the cross entropy loss.

[0061] The advantages of the present invention are:

[0062] Aiming at the problem of medical video classification in complex scenes, the present invention proposes a medical video classification method based on multi-target segmentation prior-driven. By structurally segmenting the target of interest in the video and constructing a multi-way parallel feature extraction-temporal coding-fusion network, the spatiotemporal feature extraction and comprehensive modeling of single-object and multi-object segmentation mask sequences are realized. This method can effectively capture the long-distance dependency between video frames, characterize the dynamic evolution process and behavior pattern of the target, and then deconstruct the overall spatiotemporal features of the video into the behavior characteristics and structural change characteristics of multiple main targets, thereby transforming the traditional video classification task into a semantically guided target behavior analysis and scene understanding process. The multi-channel parallel structure constructed by the present invention can fine-grainedly mine the spatiotemporal dynamic features of each main target in the video, and realize the effective recognition of multi-target interaction patterns in complex scenes. By introducing the cross-modal attention mechanism, the VTN encoder adopted not only has the ability to perceive spatial features within the frame, but also can understand its evolution law in the time dimension, thereby enhancing the model's global understanding and causal reasoning capabilities of the video content. At the mechanism level, this method can effectively improve the ability to represent multi-target coupling and complex spatiotemporal relationships, significantly enhance the accuracy, robustness and generalization ability in medical video classification tasks, and has good clinical application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.

[0064] Figure 1 is a flow chart of the steps of the method of the present invention;

[0065] Figure 2 is the ROC curve of the classification of each category by the method of the present invention;

[0066] Figure 3 The figure shows the ROC curve comparison between the method of the present invention and the existing method during the experiment. DETAILED DESCRIPTION

[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0068] Example 1

[0069] In this embodiment, Figure 1 As shown, the present invention provides a medical video classification method based on multi-target segmentation prior drive, and the specific steps include:

[0070] S1. Collect and preprocess medical video frame sequences to obtain original video frame sequences.

[0071] Specifically, medical videos collected in real clinical environments have characteristics such as high noise, low contrast, and the presence of artifacts, and different instruments have different sampling densities, resolutions, and pixel intensities. To maintain consistency in information scale during video analysis, the method of the present invention performs a unified sampling frequency and resolution on each sample frame sequence of the collected original medical video, performs histogram equalization and pixel intensity standardization on each frame image, obtains the original video frame sequence, and obtains the category label based on the diagnosis result report; the unified sampling frequency and resolution preprocessing takes into account the differences in physical signs of different subjects and sampling environments. For ease of processing, each video sample is resampled so that each video sample in the sample set has the same number of video frames in the sampling interval and consistent frame image resolution; the histogram equalization processing takes into account the grayscale value differences and noise effects of different samples, and performs histogram preprocessing on each frame image to equalize the grayscale value distribution of the image and enhance local contrast; the pixel intensity standardization preprocessing performs intensity scaling on the video frame image, adjusting the pixel intensity value of the original sample from the range of -1000 to 2000 to the standardized range of 0 to 255.

[0072] S2. Construct a multi-object segmentation-classification dataset and divide it into a training set and a test set. The dataset includes a video original frame sequence, a single-object segmentation mask sequence, a multi-object segmentation mask sequence, and a category label.

[0073] Specifically, an existing video target segmentation algorithm is used to perform target segmentation on the original video frame sequence, and then a single-object segmentation mask sequence for the target anatomical structure and a multi-object segmentation mask sequence for the visual objects of interest in the scene are obtained through a method of inspection and correction by domain experts; the existing video target segmentation algorithms such as the Vivim model, the MemSAM model and the GraphEcho model.

[0074] S3. Construct a medical video classification model based on multi-target segmentation prior-driven, the medical video classification model includes a feature extraction-temporal coding-fusion network and a classifier; the feature extraction-temporal coding-fusion network includes a feature extraction module, a temporal coding module and a feature fusion module; the temporal coding module includes a time information embedding layer, a spatiotemporal fusion information coding layer, and an MLP projection layer; the model is trained using data in the training set.

[0075] Specifically, the feature extraction module uses the first five layers of the deep residual convolutional network ResNet50 as the backbone network. In the first five layers of the deep residual convolutional network ResNet50, the first convolution block uses a 5×5 convolution kernel for feature extraction; the second to fifth convolution blocks use Inception convolution, and channel-space mixed attention is added after the fifth convolution block to assign different weights to the pixel values ​​of each part of the frame image; this enables the model to focus on important information and aggregate contextual information for the target of interest to enhance the significance of target features and scene structure features; the stride of the second and third convolution blocks is set to 2, and the stride of the fourth and fifth convolution blocks is set to 1;

[0076] The parallel single-object segmentation mask sequence and multi-object segmentation mask sequence are extracted by the feature extraction module from each frame image of each segmentation mask sequence to generate a single-target segmentation feature map sequence. and multi-target segmentation feature map sequence ;

[0077] In the process of inputting the single object segmentation mask sequence into the feature extraction module to extract the frame image features, the calculation formulas of channel attention and spatial attention are expressed as follows:

[0078] ,

[0079] ,

[0080] in, represents the output of the fifth convolutional block, represents the channel attention weight, represents the spatial attention weight, represents a multilayer perceptron, represents average pooling along the channel dimension, represents the maximum pooling along the channel dimension, Indicates that the convolution kernel size is The convolution operation, Represents a splicing operation, Represents the Sigmoid activation function; the output of the fifth convolution block is distributed through attention weights to obtain a sequence of single-target segmentation feature maps. , the formula is as follows:

[0081] ,

[0082] in, Represents element-by-element multiplication; similarly, we get a sequence of multi-target segmentation feature maps .

[0083] Specifically, the time information embedding layer is to effectively integrate the time information and time dependency in the feature sequence, and adopt a time-aware feature representation method to segment the feature map sequence of a single target. and multi-target segmentation feature map sequence The parallel input is processed into the time information embedding layer. By introducing the temporal position encoding, the time sequence information is given to the feature map sequence, and the spatiotemporal fusion information feature representation of the single target segmentation feature map sequence is obtained. and multi-target segmentation feature map sequence spatiotemporal fusion information feature representation , the formula is as follows:

[0084] ,

[0085] ,

[0086] in, Represents a positional encoding containing timing information, Indicates the position number of the video frame in the video frame sequence. Indicates the length of the video sequence.

[0087] Specifically, the spatiotemporal fusion information coding layer includes multiple channels, each channel is composed of several VTN encoder layers stacked in series, and for each VTN encoder layer, multi-head self-attention is used to perform spatiotemporal information coding fusion calculation; here, the multi-target feature map sequence is the same as the multi-head self-attention calculation method used for each single-target feature map sequence, sharing the learning mechanism and calculation parameters.

[0088] Feature representation of spatiotemporal information fusion based on single target segmentation feature map sequence The calculation process in the first VTN encoder layer of the spatiotemporal fusion information encoding layer is as follows:

[0089] ,

[0090] ,

[0091] ,

[0092] ,

[0093] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively; , , Respectively represent the first weight matrix, the second weight matrix, and the third weight matrix that can be learned; Indicates the scaling factor, which is usually taken as 64; for the Attention heads, the self-attention calculation formula is expressed as: ; The outputs of all attention heads are concatenated and then input into the linear layer for further fusion transformation and dimension processing to obtain the output of the multi-head self-attention layer , the formula is as follows:

[0094] ,

[0095] in, Represents a splicing operation, Represents the trainable fourth weight matrix; the output of the multi-head self-attention layer is processed through residual connection and layer normalization to obtain the normalized feature representation , the formula is as follows:

[0096] ,

[0097] in, Normalization operation of the representation layer; input the normalized feature representation into the feedforward neural network to extract and transform the features to obtain the features output by the feedforward neural network. The formula is as follows:

[0098] ,

[0099] in, represents the characteristics of the feedforward neural network output, represents a feedforward neural network, represents the bias vector of the first layer of the feedforward neural network, represents the weight matrix of the second layer of the feedforward neural network, represents the bias vector of the second layer of the feedforward neural network, Represents the ReLU activation function; the characteristics of the feedforward neural network output And the normalized feature representation Through residual connection and layer normalization, we get The output of the VTN encoding module is expressed as follows:

[0100] ,

[0101] in, represents the output of the first layer VTN encoding module of the kth frame of the i-th single target segmentation feature map sequence;

[0102] Stack the above process Layers are used to extract richer and more significant spatiotemporal feature representations, and gradually refine the understanding of these features at each level to obtain the global spatiotemporal feature matrix of the single target segmentation feature map sequence ,in, , n represents the number of single target segmentation feature map sequences; similarly, the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences After processing through the spatiotemporal fusion information coding layer, the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence is obtained .

[0103] Specifically, the MLP projection layer includes two fully connected layers and a ReLU activation function. The output of the L-th layer VTN encoding module undergoes a linear transformation through the MLP projection layer, aiming to project the global spatiotemporal feature matrix into a suitable dimension to ensure its compactness and applicability for subsequent fusion processing steps.

[0104] Global spatiotemporal feature matrix of single target segmentation feature map sequence and the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence After processing by the MLP projection layer, the single target segmentation fusion feature is obtained and multi-target segmentation fusion features , used to characterize the dynamic behavior of each target and the evolution of scene structure.

[0105] Specifically, in the feature fusion module, the single target segmentation fusion feature And multi-target segmentation fusion feature output Perform splicing to obtain the splicing matrix , the formula is as follows:

[0106] ,

[0107] The stitching matrix Generate the overall spatiotemporal features of the video through a fully connected layer and activation function , the formula is as follows:

[0108] ,

[0109] in, represents the trainable weight matrix of the fully connected layer, represents the trainable bias matrix, Represents the ReLU activation function.

[0110] Specifically, the classifier is a fully connected layer containing a Softmax activation function. The overall spatiotemporal features of the video are passed through the classifier to output the category prediction result of the medical video. The formula is as follows:

[0111] ,

[0112] in, represents the weight matrix of the classifier, represents the bias vector, represents the Softmax activation function, Represents the output category prediction value.

[0113] S4. During the training process, the model is optimized through the loss function to obtain a trained model.

[0114] Specifically, the cross entropy loss is used as the objective function to train the model, and the formula is as follows:

[0115] ,

[0116] in, represents the category prediction value of the jth sample, represents the category label of the jth sample, represents the number of training set samples, represents the cross entropy loss.

[0117] S5. Input the data in the test set into the trained model to obtain the medical video classification results.

[0118] Example 2

[0119] Echocardiographic videos, based on the physical properties of ultrasound and the differences in ultrasound reflection characteristics of different tissues, present the changing process of ultrasound echo signals in the form of a graphical sequence. This intuitively reflects the anatomical structure of the human cardiac system, the morphological changes of various organs and tissues during the rhythmic cycle, and the motion characteristics. Abnormal changes in the spatiotemporal dimensions of individual and systemic structures reflect the pathological characteristics of different diseases. This paper uses a medical video classification method driven by multi-target segmentation. Based on the differences in spatiotemporal characteristics and symptom patterns presented by different heart diseases in echocardiographic videos, it classifies and diagnoses four types of heart disease: type 2 diabetes with coronary heart disease (T2DM with CHD), isolated coronary heart disease (isolated CHD), non-coronary heart disease (non-CHD), and normal heart disease.

[0120] In this embodiment, a self-built medical echocardiography video multi-objective segmentation-classification dataset is used to conduct a verification experiment on the effectiveness and advancement of the medical video classification model established by the present invention, and the performance of the proposed video classification model in the medical field is evaluated based on the experimental results.

[0121] Experimental Dataset: This medical echocardiography video multi-target segmentation and classification dataset includes echocardiography video samples collected from 928 subjects. The method first preprocesses the original video. Then, MemSAM, a mainstream segmentation model suitable for echocardiography videos, is used to segment seven cardiac targets sensitive to heart disease diagnosis: the left ventricle, right atrium, mitral valve, tricuspid valve, ventricular wall, atrial septum, and ventricular septum. Manual inspection and correction are then performed to obtain a video multi-target segmentation dataset. Class labels for each sample are then obtained based on the subject's diagnostic report. All of these video segmentation maps have been meticulously annotated and verified by medical experts. The sample distribution of the experimental dataset is shown in Table 1.

[0122] Table 1 Experimental sample distribution

[0123]

[0124] Implementation details: Experimental environment and hyperparameter settings. The constructed model, also referred to as the proposed model, was implemented using PyTorch and MONAI and trained on a cluster of 4 NVIDIA A100 GPUs. The learning rate was set to 0.0001 and the decay was 0.01. All input video frame images were normalized to have zero mean and unit standard deviation with non-zero pixels. During training, the label intensity was adjusted to the same output range for consistency and standardization purposes. The batch size for each GPU was set to 1. All models were trained for a total of 100 epochs using the Mini-Batch SGD optimizer. An early stopping mechanism was used to prevent overfitting during model training; the decay rate was applied every 30 epochs.

[0125] Method validation experiments: In training a medical video classification model driven by multi-objective segmentation priors, the sample set was randomly divided into two groups based on the proportion of disease types: one group of 660 samples formed the training set, and the other group of 268 samples served as the test set. The model training error accuracy was set to 0.01, the batch size was 16, and the number of training rounds was 100. Accuracy, recall, precision, and F1 score were used as model performance evaluation metrics. The test set samples were classified and recognized. The classification results for the four disease categories and the various evaluation metrics are shown in Table 2.

[0126] Table 2 Experimental results of the method of the present invention

[0127]

[0128] As can be seen from Table 2, there are significant differences in the classification performance of each category. The normal category has an accuracy of 88.80%, a recall rate of 90.91%, a precision of 89.29%, and an F1-Score of 90.09%, which is the best performance. This is not only due to the obvious physiological characteristics and low pathological interference of the normal state itself, but also reflects that the model can accurately capture the stable pattern of the normal cardiac system. The classification performance of T2DM combined with CHD is relatively low, with an accuracy of only 78.73%, a recall rate of 73.61%, a precision of 76.12%, and an F1-Score of 74.45%, suggesting that the model has a problem of missing detection when identifying this category, which may be due to the lack of information. The accuracy of the CHD category was 78.31%, the recall was 74.68%, the precision was 77.33%, and the F1-Score was 80.60%, indicating that this category is prone to misdiagnosis in practice, possibly due to the overlap of its pathological features with other abnormal conditions (such as non-CHD). The non-CHD category performed relatively stably, with an accuracy of 78.90%, a recall of 85.94%, a precision of 7.14%, and an F1-Score of 80.60%, demonstrating strong sensitivity to this category, but also exhibiting some false positives. Overall, the performance differences between the proposed model categories primarily reflect the discrimination and internal variability of each category in the spatiotemporal features of echocardiographic videos. The normal category performed best due to its stable and easily identifiable features, while the pathological category was relatively difficult to identify due to the similarity and confusion of some features.

[0129] The overall recognition accuracy of the method of the present invention reached 80.38%. According to the evaluation of medical experts, this is a good result in the current automatic classification diagnosis of the four types of coronary heart disease and has clinical application value. This is because in the diagnosis of diseases based on medical images, the method of the present invention can effectively extract and represent the behavioral characteristics and correlation change characteristics of multiple targets in the motion process that are sensitive to disease classification diagnosis. The application of mixed attention and multi-head self-attention mechanisms enhances the ability to focus on the target of interest and the comprehensiveness of feature extraction, while reducing the influence of background, noise and other non-sensitive targets. Based on medical diagnostic knowledge, the method of the present invention uses the local and global spatiotemporal feature information of multiple targets sensitive to disease diagnosis to understand complex scenes, and more accurately realizes the classification diagnosis of coronary heart disease.

[0130] Classification Model Performance Analysis Based on the Receiver Operating Characteristic (ROC) Curve: The ROC (Receiver Operating Characteristic) curve is an important method for evaluating the performance and clinical applicability of medical classification diagnostic models. It plots the performance at different classification thresholds using sensitivity (1 Positive Rate, TPR) on the vertical axis and 1 minus specificity (1-specificity, 0 Positive Rate, FPR) on the horizontal axis. The curve provides a visual reflection of the model's ability to distinguish between positive and negative samples. The area under the curve (AUC) is a key metric for measuring the model's overall classification performance, ranging from 0.5 (random guessing) to 1 (perfect classification). In medical scenarios, a higher AUC value indicates a better overall model performance in distinguishing between different classes. Sensitivity (TPR) indicates the model's ability to correctly identify positive examples, i.e., the proportion of samples that are actually positive that are correctly classified. High sensitivity indicates a higher disease detection rate. Specificity (1-FPR) indicates the model's ability to correctly exclude negative examples, i.e., the proportion of samples that are actually negative that are correctly classified. High specificity reduces the risk of misdiagnosis and is crucial for clinical decision-making.

[0131] like Figure 2 As shown in the figure, receiver operating characteristic (ROC) curves plotted based on the experimental results demonstrate that the proposed method performs well in the four disease classification tasks. The areas under the curve (AUCs) for T2DM combined with CHD, CHD alone, non-CHD, and normal categories were 0.798, 0.786, 0.809, and 0.867, respectively. The normal category had the highest AUC, reaching 0.867, indicating that the model has high sensitivity and specificity in distinguishing normal from abnormal states. The AUC for T2DM combined with CHD was 0.798, slightly lower than that for the other categories, suggesting a high diversity of features within this category. The AUC for CHD alone was 0.786, suggesting that the model's ability to capture its pathological features was slightly weaker. The AUC for non-CHD was 0.809, showing stable performance, indicating that the model has strong discriminatory power.

[0132] Comparative Experiments: In this comparative experiment, the proposed method was compared with the mainstream video classification models VTN (Video Transformer Network), Small-Big Network (Integrating Core and Contextual Views for Video Classification), ViViT (Video Vision Transformer), and MViT (Multiscale Vision Transformer) on a medical echocardiography video segmentation and classification dataset. The inputs of the four mainstream video classification comparison models were preprocessed echocardiography video frame image sequences, while the proposed method used single-object segmentation map sequences and multi-object segmentation map sequences from the medical echocardiography video multi-target segmentation and classification dataset. The evaluation metrics, training set, and test set were the same as those in the "Method Validation Experiment." The comparative experimental results are shown in Table 3.

[0133] Table 3 Comparative experimental results

[0134]

[0135] As shown in Table 3, the experimental results of the method of the present invention performed best across all evaluation indicators. This is because the classification method established by the present invention, through multi-target segmentation, focuses on the behavioral characteristics and spatial correlation variation characteristics of organs and tissues closely related to the classification and discrimination of the four types of coronary heart disease. This is consistent with the discrimination rules relied upon by medical experts in diagnosis. It also reduces the complexity of the video scene and the algorithm's requirements for the completeness of the training dataset to a certain extent, thereby improving the algorithm's generalization ability and robustness. The four models, VTN, Small-Big, ViViT, and MViT, are all "end-to-end" deep learning models that directly classify entire sequences of video frames. They lack detailed exploration of the appearance and motion characteristics of sensitive organs and tissues, targeted processing of the unique characteristics of medical images, and inspiration and utilization of medical rules. This reduces the models' ability to detect the spatiotemporal details and diverse abnormal patterns of medical videos that are closely related to disease classification. Furthermore, they lack the utilization of medical rules, which in turn leads to poor generalization of the four compared models.

[0136] Comparative analysis of the performance of different methods based on the receiver operating characteristic curve (ROC): The ROC curves of the five comparison models drawn according to the experimental results are as follows Figure 3As shown in the figure. Comparative experimental results show that the AUC of the proposed method is 0.815, which is significantly better than VTN (0.716), Small-Big (0.709), ViViT (0.742), and MViT (0.763). This result shows that the proposed method effectively improves sensitivity and specificity through multi-target segmentation and spatiotemporal feature fusion, and can more accurately capture the dynamic characteristics of organs and tissues that are closely related to disease classification. In addition, the overall trend of the ROC curve shows that the proposed method can still maintain high sensitivity within a high specificity range, further verifying its superior performance in medical video classification tasks.

[0137] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A medical video classification method based on multi-target segmentation prior drive, characterized by: The following steps are involved: S1. Acquire and preprocess medical video frame sequences to obtain original video frame sequences; S2. Construct a multi-object segmentation-classification dataset and divide it into a training set and a test set. The dataset includes a video original frame sequence, a single-object segmentation mask sequence, a multi-object segmentation mask sequence, and a category label; S3. Constructing a medical video classification model based on multi-target segmentation prior drive, the medical video classification model includes a feature extraction-temporal coding-fusion network and a classifier; The feature extraction-temporal coding-fusion network includes a feature extraction module, a temporal coding module and a feature fusion module; The temporal coding module includes a time information embedding layer, a spatiotemporal fusion information coding layer, and an MLP projection layer; the model is trained using data from the training set; The feature extraction module uses the first five layers of the deep residual convolutional network ResNet50 as the backbone network. In the first five layers of the deep residual convolutional network ResNet50, the first convolution block uses a 5×5 convolution kernel for feature extraction; The second to fifth convolution blocks use Inception convolution, and channel-spatial mixed attention is added after the fifth convolution block to assign different weights to the pixel values ​​of each part of the frame image; The stride of the second and third convolution blocks is set to 2, and the stride of the fourth and fifth convolution blocks is set to 1; The parallel single-object segmentation mask sequence and multi-object segmentation mask sequence are extracted by the feature extraction module from each frame image of each segmentation mask sequence to generate a single-target segmentation feature map sequence. and multi-target segmentation feature map sequence ; In the process of inputting the single object segmentation mask sequence into the feature extraction module to extract the frame image features, the calculation formulas of channel attention and spatial attention are expressed as follows: , , in, represents the output of the fifth convolutional block, represents the channel attention weight, represents the spatial attention weight, represents a multilayer perceptron, represents average pooling along the channel dimension, represents the maximum pooling along the channel dimension, Indicates that the convolution kernel size is The convolution operation, Represents a splicing operation, Represents the Sigmoid activation function; the output of the fifth convolution block is distributed through attention weights to obtain a sequence of single-target segmentation feature maps. , the formula is as follows: , in, Represents element-by-element multiplication; similarly, we get a sequence of multi-target segmentation feature maps ; S4. During the training process, the model is optimized using the loss function to obtain a trained model. S5. Input the data in the test set into the trained model to obtain the medical video classification results.

2. The medical video classification method based on multi-target segmentation prior drive according to claim 1 is characterized in that: In step S1, the sampling frequency and resolution of each sample frame sequence of the collected original medical video are unified, and each frame image is preprocessed by histogram equalization and pixel intensity standardization to obtain the original video frame sequence, and the category label is obtained according to the diagnosis result report.

3. The medical video classification method based on multi-target segmentation prior drive according to claim 2 is characterized in that: Step S2 specifically includes: The existing video target segmentation algorithm is used to segment the original video frame sequence. Then, through the method of inspection and correction by domain experts, a single-object segmentation mask sequence for the target anatomical structure and a multi-object segmentation mask sequence for the visual objects of interest in the scene are obtained. The multiple single-object segmentation mask sequences and one multi-object segmentation mask sequence are combined into a matrix as the input of the model.

4. The medical video classification method based on multi-target segmentation prior drive according to claim 3 is characterized in that: The time information embedding layer in step S3 specifically includes: Single object segmentation feature map sequence and multi-target segmentation feature map sequence The parallel input is processed into the time information embedding layer. By introducing the temporal position encoding, the time sequence information is given to the feature map sequence, and the spatiotemporal fusion information feature representation of the single target segmentation feature map sequence and the spatiotemporal fusion information feature representation of the multi-target segmentation feature map sequence are obtained. The formula is expressed as follows: , , in, Represents a positional encoding containing timing information, Indicates the position number of the video frame in the video frame sequence. Indicates the length of the video sequence; Represents the spatiotemporal fusion information feature representation of a single target segmentation feature map sequence, Represents the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences.

5. The medical video classification method based on multi-target segmentation prior drive according to claim 4 is characterized in that: The spatiotemporal fusion information coding layer in step S3 specifically includes: The spatiotemporal fusion information coding layer includes multiple channels, each channel is composed of several VTN encoder layers stacked in series, and for each VTN encoder layer, multi-head self-attention is used to perform spatiotemporal information coding fusion calculation; the spatiotemporal fusion information feature representation of the single target segmentation feature map sequence The calculation process in the first VTN encoder layer of the spatiotemporal fusion information encoding layer is as follows: , , , , Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively; , , Respectively represent the first weight matrix, the second weight matrix, and the third weight matrix that can be learned; represents the scaling factor; for Attention heads, the self-attention calculation formula is expressed as: ; The outputs of all attention heads are concatenated and then input into the linear layer for further fusion transformation and dimension processing to obtain the output of the multi-head self-attention layer , the formula is as follows: , in, Represents a splicing operation, Represents the trainable fourth weight matrix; the output of the multi-head self-attention layer is processed through residual connection and layer normalization to obtain the normalized feature representation , the formula is as follows: , in, Normalization operation of the representation layer; input the normalized feature representation into the feedforward neural network to extract and transform the features to obtain the features output by the feedforward neural network. The formula is as follows: , in, represents the characteristics of the feedforward neural network output, represents a feedforward neural network, represents the bias vector of the first layer of the feedforward neural network, represents the weight matrix of the second layer of the feedforward neural network, represents the bias vector of the second layer of the feedforward neural network, Represents the ReLU activation function; the characteristics of the feedforward neural network output And the normalized feature representation Through residual connection and layer normalization, we get The output of the VTN encoding module is expressed as follows: , in, represents the output of the first layer VTN encoding module of the kth frame of the i-th single target segmentation feature map sequence; Stack the above process Layer, get the global spatiotemporal feature matrix of the single target segmentation feature map sequence ,in, , n represents the number of single target segmentation feature map sequences; similarly, the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences After processing through the spatiotemporal fusion information coding layer, the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence is obtained .

6. The medical video classification method based on multi-target segmentation prior drive according to claim 5, characterized in that: The MLP projection layer in step S3 specifically includes: The MLP projection layer includes two fully connected layers and a ReLU activation function; Global spatiotemporal feature matrix of single target segmentation feature map sequence and the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence After processing by the MLP projection layer, the single target segmentation fusion feature is obtained and multi-target segmentation fusion features .

7. The medical video classification method based on multi-target segmentation prior drive according to claim 6, characterized in that: The feature fusion module in step S3 specifically includes: Fusion features for single target segmentation And multi-target segmentation fusion feature output Perform splicing to obtain the splicing matrix , the formula is as follows: , The stitching matrix Generate the overall spatiotemporal features of the video through a fully connected layer and activation function , the formula is as follows: , in, represents the trainable weight matrix of the fully connected layer, represents the trainable bias matrix, Represents the ReLU activation function.

8. The medical video classification method based on multi-target segmentation prior drive according to claim 7 is characterized in that: In step S3, the classifier is a fully connected layer including a Softmax activation function, and the overall spatiotemporal features of the video The classifier outputs the category prediction results of the medical video, and the formula is as follows: , in, represents the weight matrix of the classifier, represents the bias vector, represents the Softmax activation function, Represents the output category prediction value.

9. The medical video classification method based on multi-target segmentation prior drive according to claim 8, characterized in that: In step S4, the loss function uses cross entropy loss as the objective function to train the model. The formula is as follows: , in, represents the category prediction value of the jth sample, represents the category label of the jth sample, represents the number of training set samples, represents the cross entropy loss.

Citation Information

Patent Citations

  • Pedestrian behavior category detection method based on image sequence

    CN113688761A

  • Air traffic controller unsafe behavior classification method based on monitoring video

    CN115410035A