Medical video classification method based on multi-target segmentation prior driving
By constructing a medical video classification method driven by multi-objective segmentation prior, using deep residual convolution network and multi-head self-attention mechanism, the problem of temporal and spatial semantic interaction relationships in medical video classification is solved, and multi-objective interaction pattern recognition is achieved in complex scenarios, which improves diagnostic accuracy and robustness.
Patent Information
- Application Number
- CN202510905487.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing medical video classification methods are difficult to effectively capture spatiotemporal semantic interactions, especially in dynamic multi-objective scenarios, and the lack of medical semantic modeling capabilities, resulting in insufficient diagnostic accuracy and interpretability.
Using a medical video classification method based on multi-objective segmentation prior drive, a multi-objective segmentation-classification data set is constructed, and a deep residual convolution network and a multi-head self-attention mechanism is used to perform feature extraction, timing encoding and fusion, combining time information embedding and feature fusion, to improve the model's understanding of video content.
Effectively capture the long-distance dependence between video frames, improves the accuracy, robustness and generalization ability of medical video classification, can mine the space-time and dynamic characteristics of each major target in a fine-grained manner, and enhances the ability to recognize multi-objective interactive modes in complex scenarios.
Smart Images

Figure CN120411862A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video recognition, and in particular relates to a medical video classification method based on multi-target segmentation prior drive. Background Art
[0002] Medical videos, as important carriers of dynamic visual information in clinical diagnosis, are characterized by complex scene changes, significant multi-target coupling and linkage, and a high degree of interweaving of spatiotemporal information. Compared with static images, medical videos can more comprehensively capture the dynamic characteristics of lesions, helping to identify the morphological evolution of organs at different stages, and providing a more explanatory basis for the accurate auxiliary diagnosis of systemic diseases such as the heart. Therefore, medical video classification methods must not only extract local structural features within the frame but also be able to model the coordinated movement and deformation patterns of organ tissues in the temporal dimension, thereby supporting the recognition and discrimination of diverse disease patterns and meeting clinical needs for interpretability and practicality.
[0003] Although video classification research has made progress in model architecture, feature expression, and context modeling in recent years, medical videos still face several key challenges in practical applications. For example, due to the presence of dynamic background interference, multi-target structural linkage, and long-term dependencies, existing methods find it difficult to effectively capture spatiotemporal semantic interactions, which in turn limits the model's diagnostic discrimination capabilities. In addition, medical videos often have problems such as blurred boundaries and insignificant target features, which are particularly evident in dynamic multi-target scenarios. For example, in cardiac ultrasound videos, key features such as subtle abnormalities in cardiac chamber motion, changes in local myocardial contraction, or mild coronary artery stenosis often contain important diagnostic signals, but are extremely difficult to accurately identify with traditional methods in low-contrast or high-noise environments.
[0004] While existing methods have attempted to introduce temporal modeling mechanisms, such as recurrent neural networks (RNNs) and temporal convolutional networks (TCNs), to model local temporal dependencies, these methods typically perform well in short time series and struggle to capture structural changes over long time series. Furthermore, when processing multi-target videos, existing models make limited use of structural priors, making it difficult to achieve decoupled representations and semantic fusion of different organ targets. Furthermore, current mainstream video classification models for natural scenes often perform poorly when transferred to the medical field due to a lack of medical semantic modeling capabilities and structural constraints, failing to meet the dual requirements of accuracy and interpretability in medical scenarios. Summary of the Invention
[0005] In order to solve the above problems, the present invention provides a medical video classification method based on multi-target segmentation prior drive.
[0006] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: The present invention provides a medical video classification method driven by multi-object segmentation priors, comprising the following steps: S1. Collect and preprocess the medical video frame sequence to obtain the original video frame sequence; S2. Construct a multi-object segmentation-classification dataset and divide it into a training set and a test set. The dataset includes the original video frame sequence, single-object segmentation mask sequences, multi-object segmentation mask sequences, and class labels; S3. Construct a medical video classification model driven by multi-object segmentation priors. The medical video classification model includes a feature extraction-temporal encoding-fusion network and a classifier; the feature extraction-temporal encoding-fusion network includes a feature extraction module, a temporal encoding module, and a feature fusion module; the temporal encoding module includes a time information embedding layer, a spatio-temporal fusion information encoding layer, and an MLP projection layer; use the data in the training set to train the model; S4. During the training process, optimize the model through a loss function to obtain a trained model; S5. Input the data in the test set into the trained model to obtain the medical video classification result.
[0007] Further, in step S1, perform preprocessing on each sample frame sequence of the collected original medical video, including unifying the sampling frequency and resolution, performing histogram equalization and pixel intensity normalization on each frame image, to obtain the original video frame sequence, and obtain the class label according to the diagnostic result report.
[0008] Further, step S2 specifically includes: Adopt an existing video object segmentation algorithm to perform object segmentation on the original video frame sequence, and then obtain the single-object segmentation mask sequence for the target anatomical structure and the multi-object segmentation mask sequence for the visible objects of interest in the scene through the method of checking and correcting by domain experts; form a matrix with multiple single-object segmentation mask sequences and 1 multi-object segmentation mask sequence as the input of the model.
[0009] Further, the feature extraction module in step S3 specifically includes: The feature extraction module uses the first five layers of the deep residual convolutional network ResNet50 as the backbone network. In the first five layers of the deep residual convolutional network ResNet50, the first convolutional block uses a convolutional kernel with a size of 5×5 for feature extraction; the second to fifth convolutional blocks use Inception convolution, and a channel-spatial hybrid attention is added after the fifth convolutional block to assign different weights to the pixel values of each part of the frame image; the stride of the second and third convolutional blocks is set to 2, and the stride of the fourth and fifth convolutional blocks is set to 1; The parallel single-object segmentation mask sequence and multi-object segmentation mask sequence are input into the feature extraction module to extract the spatial features of each frame image in each segmentation mask sequence, generating a single-object segmentation feature map sequence and a multi-object segmentation feature map sequence ; During the process of inputting the single-object segmentation mask sequence into the feature extraction module to extract the frame image features, the calculation formulas of channel attention and spatial attention are as follows: , , where, represents the output of the fifth convolutional block, represents the channel attention weight, represents the spatial attention weight, represents a multi-layer perceptron, represents average pooling along the channel dimension, represents max pooling along the channel dimension, represents a convolutional operation with a convolutional kernel size of ; represents a concatenation operation, represents a Sigmoid activation function; the output of the fifth convolutional block is subjected to attention weight allocation to obtain a single-object segmentation feature map sequence , and the formula is as follows: , where, represents element-wise multiplication; similarly, a multi-object segmentation feature map sequence is obtained.
[0010] Furthermore, the time information embedding layer in step S3 specifically includes: The single-object segmentation feature map sequence and the multi-object segmentation feature map sequence are input in parallel into the time information embedding layer for processing. By introducing temporal position encoding, temporal order information is given to the feature map sequence, obtaining the spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence and the spatio-temporal fusion information feature representation of the multi-object segmentation feature map sequence. The formula is as follows: , , where, represents the position encoding containing temporal information, represents the position serial number of the video frame in the video frame sequence, represents the length of the video sequence; Represents the spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence. Represents the spatio-temporal fusion information feature representation of the multi-object segmentation feature map sequence.
[0011] Furthermore, the spatio-temporal fusion information encoding layer in step S3 specifically includes: The spatio-temporal fusion information encoding layer includes multiple channels, and each channel is serially stacked by several VTN encoder layers. For each VTN encoder layer, multi-head self-attention is used for spatio-temporal information encoding and fusion calculation; the spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence The calculation process in the first VTN encoder layer of the spatio-temporal fusion information encoding layer is as follows: , , , , Among them, Q, K, and V respectively represent the query matrix, the key matrix, and the value matrix. , , respectively represent the learnable first weight matrix, the second weight matrix, and the third weight matrix. Represents the scaling factor; for the th attention head, the self-attention calculation formula is expressed as: ; The outputs of all attention heads are concatenated and then input into the linear layer for further fusion transformation and dimension processing to obtain the output of the multi-head self-attention layer , and the formula is as follows: , Among them, represents the concatenation operation, represents the trainable fourth weight matrix; the output of the multi-head self-attention layer is processed through residual connection and layer normalization to obtain the normalized feature representation , and the formula is as follows: , Among them, represents the layer normalization operation; the normalized feature representation is input into the feed-forward neural network for feature extraction and transformation to obtain the feature output by the feed-forward neural network, and the formula is as follows: , Among them, represents the feature output by the feed-forward neural network, represents the feed-forward neural network, represents the bias vector of the first layer of the feedforward neural network, represents the weight matrix of the second layer of the feedforward neural network, represents the bias vector of the second layer of the feedforward neural network, Represents the ReLU activation function; the characteristics of the feedforward neural network output And the normalized feature representation Through residual connection and layer normalization, we get The output of the VTN encoding module is expressed as follows: , in, represents the output of the first layer VTN encoding module of the kth frame of the i-th single target segmentation feature map sequence; Stack the above process Layer, get the global spatiotemporal feature matrix of the single target segmentation feature map sequence ,in, , n represents the number of single target segmentation feature map sequences; similarly, the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences After processing through the spatiotemporal fusion information coding layer, the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence is obtained .
[0012] Furthermore, the MLP projection layer in step S3 specifically includes: The MLP projection layer includes two fully connected layers and a ReLU activation function; Global spatiotemporal feature matrix of single target segmentation feature map sequence and the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence After processing by the MLP projection layer, the single target segmentation fusion feature is obtained and multi-target segmentation fusion features .
[0013] Furthermore, the feature fusion module in step S3 specifically includes: Fusion features for single target segmentation And multi-target segmentation fusion feature output Perform splicing to obtain the splicing matrix , the formula is as follows: , The stitching matrix Generate the overall spatiotemporal features of the video through a fully connected layer and activation function , the formula is as follows: , in, represents the trainable weight matrix of the fully connected layer, represents the trainable bias matrix, represents the ReLU activation function.
[0014] Furthermore, in step S3, the classifier is a fully connected layer including the Softmax activation function, and the overall spatio-temporal feature of the video passes through the classifier to output the class prediction result of the medical video, and the formula is as follows: , wherein, represents the weight matrix of the classifier, represents the bias vector, represents the Softmax activation function, represents the output class prediction value.
[0015] Furthermore, in step S4, the cross-entropy loss is used as the objective function of the loss function to train the model, and the formula is as follows: , wherein, represents the class prediction value of the j-th sample, represents the class label of the j-th sample, represents the number of samples in the training set, represents the cross-entropy loss.
[0016] The advantages of the present invention are as follows: Aiming at the problem of medical video classification in complex scenes, the present invention proposes a medical video classification method based on multi-target segmentation prior-driven. By structurally segmenting the target of interest in the video and constructing a multi-way parallel feature extraction-temporal coding-fusion network, the spatiotemporal feature extraction and comprehensive modeling of single-object and multi-object segmentation mask sequences are realized. This method can effectively capture the long-distance dependency between video frames, characterize the dynamic evolution process and behavior pattern of the target, and then deconstruct the overall spatiotemporal features of the video into the behavior characteristics and structural change characteristics of multiple main targets, thereby transforming the traditional video classification task into a semantically guided target behavior analysis and scene understanding process. The multi-channel parallel structure constructed by the present invention can fine-grainedly mine the spatiotemporal dynamic features of each main target in the video, and realize the effective recognition of multi-target interaction patterns in complex scenes. By introducing the cross-modal attention mechanism, the VTN encoder adopted not only has the ability to perceive spatial features within the frame, but also can understand its evolution law in the time dimension, thereby enhancing the model's global understanding and causal reasoning capabilities of the video content. At the mechanism level, this method can effectively improve the ability to represent multi-target coupling and complex spatiotemporal relationships, significantly enhance the accuracy, robustness and generalization ability in medical video classification tasks, and has good clinical application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0018] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 is the ROC curve of the classification of each category by the method of the present invention; Figure 3 The figure shows the ROC curve comparison between the method of the present invention and the existing method during the experiment. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] Example 1 In this embodiment, Figure 1 As shown, the present invention provides a medical video classification method based on multi-target segmentation prior drive, and the specific steps include: S1. Collect and preprocess the medical video frame sequence to obtain the original video frame sequence.
[0021] Specifically, the medical videos collected in the clinical real environment have characteristics such as high noise, low contrast, and artifacts, and there are differences in the sampling density, resolution, and pixel intensity of different instruments. To maintain the consistency of the information scale in video analysis, the method of the present invention preprocesses each sample frame sequence of the collected original medical videos by unifying the sampling frequency and resolution, performing histogram equalization and pixel intensity standardization on each frame image to obtain the original video frame sequence, and obtaining class labels according to the diagnostic result report; the preprocessing of unifying the sampling frequency and resolution is to consider the differences in the physical signs of different subjects and the sampling environment. For the convenience of processing, each video sample is resampled so that each video sample in the sample set has the same number of video frames and the same frame image resolution in the sampling interval; the histogram equalization process is to consider the differences in the gray values of different samples and the influence of noise, and respectively perform histogram preprocessing on each frame image to equalize the gray value distribution of the image and enhance the local contrast; the pixel intensity standardization preprocessing is to perform intensity scaling on the video frame image, and adjust the original sampled pixel intensity value from the range of -1000 to 2000 to the standardized range of 0 to 255.
[0022] S2. Construct a multi-object segmentation-classification dataset and divide it into a training set and a test set. The dataset includes the original video frame sequence, single-object segmentation mask sequence, multi-object segmentation mask sequence, and class labels.
[0023] Specifically, use the existing video object segmentation algorithm to perform object segmentation on the original video frame sequence, and then obtain the single-object segmentation mask sequence for the target anatomical structure and the multi-object segmentation mask sequence for the visible objects of interest in the scene through the method of checking and correcting by domain experts; the existing video object segmentation algorithms such as the Vivim model, MemSAM model, and GraphEcho model.
[0024] S3. Construct a medical video classification model based on the prior-driven multi-object segmentation. The medical video classification model includes a feature extraction-temporal encoding-fusion network and a classifier; the feature extraction-temporal encoding-fusion network includes a feature extraction module, a temporal encoding module, and a feature fusion module; the temporal encoding module includes a time information embedding layer, a spatio-temporal fusion information encoding layer, and an MLP projection layer; use the data in the training set to train the model.
[0025] Specifically, the feature extraction module uses the first five layers of the deep residual convolutional network ResNet50 as the backbone network. In the first five layers of the deep residual convolutional network ResNet50, the first convolutional block uses a convolutional kernel with a size of 5×5 for feature extraction; the second to fifth convolutional blocks use Inception convolution, and a channel-spatial hybrid attention is added after the fifth convolutional block to assign different weights to the pixel values of each part of the frame image; enabling the model to focus on important information and aggregate context information for the target of interest to enhance the saliency of the target features and scene structure features; the stride of the second and third convolutional blocks is set to 2, and the stride of the fourth and fifth convolutional blocks is set to 1; The parallel single-object segmentation mask sequence and multi-object segmentation mask sequence pass through the feature extraction module to extract the spatial features of each frame image in each segmentation mask sequence, generating a single-object segmentation feature map sequence and a multi-object segmentation feature map sequence ; During the process of inputting the single-object segmentation mask sequence into the feature extraction module to extract the frame image features, the calculation formulas for channel attention and spatial attention are as follows: , , where, represents the output of the fifth convolutional block, represents the channel attention weight, represents the spatial attention weight, represents the multi-layer perceptron, represents average pooling along the channel dimension, represents max pooling along the channel dimension, represents a convolutional operation with a convolutional kernel size of , represents the concatenation operation, represents the Sigmoid activation function; the output of the fifth convolutional block is subjected to attention weight assignment to obtain the single-object segmentation feature map sequence ,and the formula is as follows: , where, represents element-wise multiplication; similarly, the multi-object segmentation feature map sequence is obtained.
[0026] Specifically, the temporal information embedding layer is to effectively integrate the temporal information and temporal dependence relationships in the feature sequence, and adopts a temporal-aware feature representation method. The single-object segmentation feature map sequence and the multi-object segmentation feature map sequence It is processed by being input into the temporal information embedding layer in parallel. By introducing temporal position encoding, temporal order information is given to the sequence of feature maps, and the spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence is obtained. and the spatio-temporal fusion information feature representation of the multi-object segmentation feature map sequence , which is expressed by the following formula: , , where, represents the position encoding containing temporal information, represents the position serial number of this video frame in the video frame sequence, represents the length of the video sequence.
[0027] Specifically, the spatio-temporal fusion information encoding layer includes multiple channels, and each channel is serially stacked by several VTN encoder layers. For each VTN encoder layer, multi-head self-attention is used for spatio-temporal information encoding and fusion calculation; here, the multi-head self-attention calculation methods used for the multi-object feature map sequence and each single-object feature map sequence are the same, sharing the learning mechanism and calculation parameters.
[0028] The spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence The calculation process in the first VTN encoder layer of the spatio-temporal fusion information encoding layer is as follows: , , , , where Q, K, and V represent the query matrix, key matrix, and value matrix respectively; , , respectively represent the learnable first weight matrix, second weight matrix, and third weight matrix; represents the scaling factor, generally taken as 64; for the th attention head, the self-attention calculation formula is expressed as: ; The outputs of all attention heads are concatenated and then input into the linear layer for further fusion transformation and dimension processing to obtain the output of the multi-head self-attention layer , which is expressed by the following formula: , where, represents the concatenation operation, Represents the trainable fourth weight matrix; the output of the multi-head self-attention layer is processed through residual connection and layer normalization to obtain the normalized feature representation , the formula is as follows: , in, Normalization operation of the representation layer; input the normalized feature representation into the feedforward neural network to extract and transform the features to obtain the features output by the feedforward neural network. The formula is as follows: , in, represents the characteristics of the feedforward neural network output, represents a feedforward neural network, represents the bias vector of the first layer of the feedforward neural network, represents the weight matrix of the second layer of the feedforward neural network, represents the bias vector of the second layer of the feedforward neural network, Represents the ReLU activation function; the characteristics of the feedforward neural network output And the normalized feature representation Through residual connection and layer normalization, we get The output of the VTN encoding module is expressed as follows: , in, represents the output of the first layer VTN encoding module of the kth frame of the i-th single target segmentation feature map sequence; Stack the above process Layers are used to extract richer and more significant spatiotemporal feature representations, and the understanding of these features is gradually refined at each level to obtain the global spatiotemporal feature matrix of the single target segmentation feature map sequence ,in, , n represents the number of single target segmentation feature map sequences; similarly, the spatiotemporal fusion information feature representation of multi-target segmentation feature map sequences After processing through the spatiotemporal fusion information coding layer, the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence is obtained .
[0029] Specifically, the MLP projection layer includes two fully connected layers and a ReLU activation function. The output of the L-th layer VTN encoding module undergoes a linear transformation through the MLP projection layer, aiming to project the global spatiotemporal feature matrix into a suitable dimension to ensure its compactness and applicability for subsequent fusion processing steps. Global spatiotemporal feature matrix of single target segmentation feature map sequence and the global spatiotemporal feature matrix of the multi-target segmentation feature map sequence After being processed by the MLP projection layer, single-object segmentation fusion features are obtained and multi-object segmentation fusion features , which are used to characterize the dynamic behaviors of each object and the evolution of the scene structure.
[0030] Specifically, in the feature fusion module, the single-object segmentation fusion features and the outputs of the multi-object segmentation fusion features are concatenated to obtain a concatenation matrix , and the formula is as follows: , The concatenation matrix generates the overall spatio-temporal features of the video through a fully connected layer and an activation function , and the formula is as follows: , where represents the trainable weight matrix of the fully connected layer, represents the trainable bias matrix, represents the ReLU activation function.
[0031] Specifically, the classifier is a fully connected layer containing the Softmax activation function. The overall spatio-temporal features of the video pass through the classifier to output the class prediction result of the medical video, and the formula is as follows: , where represents the weight matrix of the classifier, represents the bias vector, represents the Softmax activation function, represents the output class prediction value.
[0032] S4. During the training process, the model is optimized through the loss function to obtain a trained model.
[0033] Specifically, the cross-entropy loss is used as the objective function to train the model, and the formula is as follows: , where represents the class prediction value of the j-th sample, represents the class label of the j-th sample, represents the number of samples in the training set, represents the cross-entropy loss.
[0034] S5. The data in the test set is input into the trained model to obtain the medical video classification result.
[0035] Embodiment 2 Echocardiographic videos, based on the physical properties of ultrasound and the differences in ultrasound reflection characteristics of different tissues, present the changing process of ultrasound echo signals in the form of a graphical sequence. This intuitively reflects the anatomical structure of the human cardiac system, the morphological changes of various organs and tissues during the rhythmic cycle, and the motion characteristics. Abnormal changes in the spatiotemporal dimensions of individual and systemic structures reflect the pathological characteristics of different diseases. This paper uses a medical video classification method driven by multi-target segmentation. Based on the differences in spatiotemporal characteristics and symptom patterns presented by different heart diseases in echocardiographic videos, it classifies and diagnoses four types of heart disease: type 2 diabetes with coronary heart disease (T2DM with CHD), isolated coronary heart disease (isolated CHD), non-coronary heart disease (non-CHD), and normal heart disease.
[0036] In this embodiment, a self-built medical echocardiography video multi-objective segmentation-classification dataset is used to conduct a verification experiment on the effectiveness and advancement of the medical video classification model established by the present invention, and the performance of the proposed video classification model in the medical field is evaluated based on the experimental results.
[0037] Experimental Dataset: This medical echocardiography video multi-target segmentation and classification dataset includes echocardiography video samples collected from 928 subjects. The method first preprocesses the original video. Then, MemSAM, a mainstream segmentation model suitable for echocardiography videos, is used to segment seven cardiac targets sensitive to heart disease diagnosis: the left ventricle, right atrium, mitral valve, tricuspid valve, ventricular wall, atrial septum, and ventricular septum. Manual inspection and correction are then performed to obtain a video multi-target segmentation dataset. Class labels for each sample are then obtained based on the subject's diagnostic report. All of these video segmentation maps have been meticulously annotated and verified by medical experts. The sample distribution of the experimental dataset is shown in Table 1.
[0038] Table 1 Experimental sample distribution Implementation details: Experimental environment and hyperparameter settings. The constructed model, also referred to as the proposed model, was implemented using PyTorch and MONAI and trained on a cluster of 4 NVIDIA A100 GPUs. The learning rate was set to 0.0001 and the decay was 0.01. All input video frame images were normalized to have zero mean and unit standard deviation with non-zero pixels. During training, the label intensity was adjusted to the same output range for consistency and standardization purposes. The batch size for each GPU was set to 1. All models were trained for a total of 100 epochs using the Mini-Batch SGD optimizer. An early stopping mechanism was used to prevent overfitting during model training; the decay rate was applied every 30 epochs.
[0039] Method effectiveness verification experiment: In the training of a medical video classification model driven by multi-object segmentation priors, the sample set was randomly divided into 2 groups according to the disease type ratio. Among them, one group with 660 samples constituted the training set, and the other group with 268 samples was used as the test set. The training error precision of the model was set to 0.01, the batch size was 16, and the number of training epochs was 100. Accuracy, recall, precision, and F1-score were used as the model performance evaluation indicators. The classification results of the 4 types of diseases and various evaluation indicators for the test set samples are shown in Table 2.
[0040] Table 2 Experimental results of the method of the present invention As can be seen from Table 2, there are obvious differences in the classification performance of each category. The accuracy of the normal category reaches 88.80%, the recall is 90.91%, the precision is 89.29%, and the F1-Score is 90.09%, showing the best performance. This is not only due to the obvious physiological characteristics and low pathological interference of the normal state itself, but also reflects that the model can capture the stable pattern of the normal heart system more accurately; while the classification performance of T2DM combined with CHD is relatively low, the accuracy is only 78.73%, the recall is 73.61%, the precision is 76.12%, and the F1-Score is 74.45%, indicating that there is a problem of missed detection when the model identifies this category, which may be due to the relatively diverse internal feature distribution; the accuracy of simple CHD is 78.31%, the recall is 74.68%, the precision is 77.33%, and the F1-Score is 80.60%, indicating that this category is prone to misjudgment in actual diagnosis, probably because its pathological features overlap with those of other abnormal states (such as non-CHD); the non-CHD category is relatively stable, with an accuracy of 78.90%, a recall of 85.94%, a precision of 7.14%, and an F1-Score of 80.60%, showing that the model has strong sensitivity to this category, but there are also certain false alarms. Generally speaking, the performance differences of the model proposed in the present invention among different categories mainly reflect the discrimination and internal variability of the spatio-temporal features of echocardiogram videos in each category. Among them, the normal state performs the best because of its stable and easy-to-identify features, while the identification difficulty is relatively large among pathological categories due to the similarity and confusion of some features.
[0041] The overall recognition accuracy of the method of the present invention reaches 80.38%. Evaluated by medical experts, this is a good result in the current automatic classification diagnosis of four types of coronary heart disease and has clinical application value. This is because in the disease diagnosis based on medical images, the method of the present invention can effectively extract and represent the behavioral characteristics and associated change characteristics in the movement processes of multiple targets sensitive to disease classification diagnosis in terms of mechanism. The application of the hybrid attention and multi-head self-attention mechanisms enhances the aggregation ability of the targets of interest and the comprehensiveness of feature extraction, while reducing the influence of background, noise and other non-sensitive targets. The method of the present invention understands complex scenarios based on medical diagnosis knowledge and uses the local and global spatio-temporal feature information of multiple targets sensitive to disease diagnosis, and more accurately realizes the classification diagnosis of coronary heart disease.
[0042] Performance analysis of the classification model based on the Receiver Operating Characteristic (ROC) curve: The ROC (Receiver Operating Characteristic) curve is an important method for evaluating the clinical usability of the performance of a medical classification diagnosis model. It takes the sensitivity (True Positive Rate, TPR) as the vertical axis and 1 minus the specificity value (1 - Specificity, False Positive Rate, FPR) as the horizontal axis, and draws a curve through the performance under different classification thresholds, intuitively reflecting the ability of the model to distinguish positive and negative samples. Among them, the Area Under Curve (AUC) is the core index for measuring the overall classification performance of the model, and its value range is from 0.5 (random guess) to 1 (perfect classification). In the medical scenario, the larger the AUC value, the better the comprehensive performance of the model in distinguishing different categories. Sensitivity (TPR) represents the ability of the model to correctly identify positive examples, that is, the proportion of samples that are actually positive and are correctly classified. High sensitivity means a higher detection rate of the disease. Specificity (1 - FPR) represents the ability of the model to correctly exclude negative examples, that is, the proportion of samples that are actually negative and are correctly classified. High specificity can reduce the risk of misdiagnosis and is crucial for clinical decision-making.
[0043] such as Figure 2As shown, the ROC curve drawn based on the experimental results indicates that the method of the present invention performs excellently in the four - disease classification tasks. For the categories of T2DM combined with CHD, simple CHD, non - CHD, and normal, the area under the curve (AUC) is 0.798, 0.786, 0.809, and 0.867 respectively. Among them, the AUC of the normal category is the highest, reaching 0.867, indicating that the model has high sensitivity and specificity in distinguishing normal and abnormal states; the AUC of T2DM combined with CHD is 0.798, slightly lower than other categories, indicating that it is related to the relatively high internal feature diversity of this category; the AUC of simple CHD is 0.786, suggesting that the model has a slightly weaker ability to capture its pathological features; the AUC of non - CHD is 0.809, showing stable performance, indicating that the model has a strong ability to distinguish its features.
[0044] Comparative experiment: In the comparative experiment, the method of the present invention was compared with the mainstream video classification VTN (Video Transformer Network) model, Small - Big network (Integrating Core and Contextual Views for Video Classification), ViViT (Video Vision Transformer) model, and MViT (Multiscale Vision Transformer) model on the medical echocardiogram video segmentation - classification dataset. The input of the 4 mainstream video classification comparative models is the pre - processed echocardiogram video frame image sequence, and the method of the present invention uses the single - object segmentation map sequence and multi - object segmentation map sequence in the medical echocardiogram video multi - target segmentation - classification dataset; the evaluation metrics, training set, and test set are the same as those set in the "Method Effectiveness Verification Experiment". The results of the comparative experiment are shown in Table 3: Table 3 Comparative experiment result data As can be seen from Table 3, the experimental results of the method of the present invention perform best among all evaluation indicators. This is because the classification method established in the present invention focuses on the behavioral characteristics and spatial correlation change characteristics of organ tissues closely related to the classification and discrimination of four types of coronary heart disease through multi-object segmentation, which is consistent with the discrimination rules relied on by medical experts in diagnosis. At the same time, to a certain extent, it reduces the complexity of the video scene and the requirements of the algorithm for the completeness of the training data set, and improves the generalization ability and robustness of the algorithm. The four models of VTN, Small-Big, ViViT, and MViT are all "end-to-end" deep learning models that directly classify the entire sequence of video frames, lacking the detailed exploration of the appearance and motion characteristics of each sensitive organ tissue, the targeted processing of the unique characteristics of medical images, and the inspiration and utilization of medical rules, reducing the model's ability to discover the spatio-temporal detail characteristics and diverse abnormal patterns of medical videos closely related to disease typing, and lacking the utilization of medical rules, which also leads to the poor generalization ability of the four comparison models.
[0045] Performance comparison and analysis of different methods based on the receiver operating characteristic curve (ROC): The ROC curves drawn by the five comparison models according to the experimental results are as Figure 3 shown. The comparison experimental results show that the AUC of the method of the present invention is 0.815, significantly superior to VTN (0.716), Small-Big (0.709), ViViT (0.742), and MViT (0.763). This result indicates that the method of the present invention effectively improves the sensitivity and specificity through multi-object segmentation and spatio-temporal feature fusion, and can more accurately capture the dynamic characteristics of organ tissues closely related to disease classification. In addition, the overall trend of the ROC curve shows that the method of the present invention can still maintain high sensitivity within a relatively high specificity range, further verifying its excellent performance in the medical video classification task.
[0046] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A medical video classification method driven by multi-object segmentation prior, characterized in that, It includes the following steps: S1. Collect and preprocess the medical video frame sequence to obtain the original video frame sequence; S2. Construct a multi-object segmentation-classification dataset and divide it into a training set and a test set. The dataset includes the original video frame sequence, single-object segmentation mask sequence, multi-object segmentation mask sequence, and class labels; S3. Construct a medical video classification model based on the prior-driven multi-object segmentation. The medical video classification model includes a feature extraction-temporal encoding-fusion network and a classifier; The feature extraction-temporal encoding-fusion network includes a feature extraction module, a temporal encoding module, and a feature fusion module; The temporal encoding module includes a time information embedding layer, a spatio-temporal fusion information encoding layer, and an MLP projection layer; use the data in the training set to train the model; S4. During the training process, optimize the model through the loss function to obtain a trained model; S5. Input the data in the test set into the trained model to obtain the medical video classification result.
2. The medical video classification method based on multi-object segmentation prior driving according to claim 1, wherein In step S1, preprocess each sample frame sequence of the collected original medical video with a unified sampling frequency and resolution, perform histogram equalization and pixel intensity normalization on each frame image to obtain the original video frame sequence, and obtain class labels according to the diagnostic result report.
3. A medical video classification method based on multi-object segmentation prior driving according to claim 2, characterized in that Step S2 specifically includes: Adopt an existing video object segmentation algorithm to perform object segmentation on the original video frame sequence, and then obtain a single-object segmentation mask sequence for the target anatomical structure and a multi-object segmentation mask sequence for the visible objects of interest in the scene through the method of domain expert inspection and correction; form a matrix with multiple single-object segmentation mask sequences and one multi-object segmentation mask sequence as the input of the model.
4. A medical video classification method based on multi-object segmentation prior drive according to claim 3, characterized in that The feature extraction module in step S3 specifically includes: The feature extraction module uses the first five layers of the deep residual convolutional network ResNet50 as the backbone network. In the first five layers of the deep residual convolutional network ResNet50, the first convolutional block uses a 5×5 convolutional kernel for feature extraction; the second to fifth convolutional blocks use Inception convolution, and a channel-spatial hybrid attention is added after the fifth convolutional block to assign different weights to the pixel values of each part of the frame image; the stride of the second and third convolutional blocks is set to 2, and the stride of the fourth and fifth convolutional blocks is set to 1; The parallel single-object segmentation mask sequence and multi-object segmentation mask sequence pass through the feature extraction module to extract the spatial features of each frame image of each segmentation mask sequence, generating a single-object segmentation feature map sequence and a multi-object segmentation feature map sequence ; During the process of inputting the single-object segmentation mask sequence into the feature extraction module to extract frame image features, the calculation formulas of channel attention and spatial attention are as follows: , , Among them, represents the output of the fifth convolutional block, represents the channel attention weight, represents the spatial attention weight, represents the multi-layer perceptron, represents average pooling along the channel dimension, represents max pooling along the channel dimension, represents that the convolutional kernel size is for the convolutional operation, represents the concatenation operation, represents the Sigmoid activation function; the output of the fifth convolutional block undergoes attention weight allocation to obtain the single-object segmentation feature map sequence , and the formula is as follows: , Among them, represents element-wise multiplication; similarly, a multi-object segmentation feature map sequence is obtained .
5. A medical video classification method based on multi-object segmentation prior driving according to claim 4, characterized in that, The time information embedding layer in step S3 specifically includes: Single-object segmentation feature map sequence and multi-object segmentation feature map sequence are input into the temporal information embedding layer in parallel for processing. By introducing temporal position encoding, temporal order information is given to the feature map sequences, and the spatio-temporal fusion information feature representations of the single-object segmentation feature map sequence and the multi-object segmentation feature map sequence are obtained. The formula is as follows: , , Among them, represents the position encoding containing temporal information, represents the position serial number of the video frame in the video frame sequence, represents the length of the video sequence; represents the spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence, represents the spatio-temporal fusion information feature representation of the multi-object segmentation feature map sequence.
6. A medical video classification method based on multi-object segmentation prior drive according to claim 5, characterized in that The spatio-temporal fusion information encoding layer in step S3 specifically includes: The spatio-temporal fusion information encoding layer includes multiple channels, each channel is serially stacked by a number of VTN encoder layers. For each VTN encoder layer, multi-head self-attention is used for spatio-temporal information encoding and fusion calculation; spatio-temporal fusion information feature representation of the single-object segmentation feature map sequence The calculation process in the first VTN encoder layer of the spatio-temporal fusion information encoding layer is as follows: , , , , Among them, Q, K, and V represent the query matrix, key matrix, and value matrix respectively; , , represent the learnable first weight matrix, second weight matrix, and third weight matrix respectively; represents the scaling factor; for the th attention head, the self-attention calculation formula is expressed as: ; Concatenate the outputs of all attention heads, then input them into the linear layer for further fusion transformation and dimensional processing to obtain the output of the multi-head self-attention layer , and the formula is as follows: , Among them, represents a splicing operation, represents a trainable fourth weight matrix; the output of the multi-head self-attention layer is processed through residual connection and layer normalization to obtain a normalized feature representation , and the formula is expressed as follows: , Among them, represents the layer normalization operation; the normalized feature representation is input into the feed-forward neural network for feature extraction and transformation to obtain the features output by the feed-forward neural network, and the formula is as follows: , Among them, represents the features output by the feedforward neural network, represents the feedforward neural network, represents the bias vector of the first layer of the feedforward neural network, represents the weight matrix of the second layer of the feedforward neural network, represents the bias vector of the second layer of the feedforward neural network, represents the ReLU activation function; the features output by the feedforward neural network and the normalized feature representation through residual connection and layer normalization processing, obtain the output of the layer VTN encoding module, and the formula is expressed as follows: , Among them, represents the output of the first-layer VTN encoding module of the k-th frame of the i-th single-object segmentation feature map sequence; Stack the above process for layers to obtain the global spatio-temporal feature matrix of the single-object segmentation feature map sequence , where , and n represents the number of single-object segmentation feature map sequences; similarly, the spatio-temporal fusion information feature representation of the multi-object segmentation feature map sequence is processed through the spatio-temporal fusion information encoding layer to obtain the global spatio-temporal feature matrix of the multi-object segmentation feature map sequence .
7. A medical video classification method based on multi-object segmentation prior driving according to claim 6, characterized in that, The MLP projection layer in step S3 specifically includes: The MLP projection layer includes two fully connected layers and a ReLU activation function; Global spatio-temporal feature matrix of single-object segmentation feature map sequence and global spatio-temporal feature matrix of multi-object segmentation feature map sequence are processed by the MLP projection layer to obtain single-object segmentation fusion features and multi-object segmentation fusion features .
8. A medical video classification method based on multi-object segmentation prior drive according to claim 7, characterized in that, The feature fusion module in step S3 specifically includes: Single-object segmentation fusion features and multi-object segmentation fusion feature output are concatenated to obtain a concatenated matrix , which is expressed by the formula as follows: , The splicing matrix generates the overall spatio-temporal features of the video through a fully connected layer and an activation function , and the formula is as follows: , Among them, represents the trainable weight matrix of the fully connected layer, represents the trainable bias matrix, represents the ReLU activation function.
9. A medical video classification method based on multi-object segmentation prior driving according to claim 8, characterized in that In step S3, the classifier is a fully connected layer containing a Softmax activation function, and the overall spatio-temporal features of the video pass through the classifier to output the class prediction result of the medical video, which is expressed by the following formula: , Among them, represents the weight matrix of the classifier, represents the bias vector, represents the Softmax activation function, represents the predicted class value of the output.
10. A medical video classification method based on multi-object segmentation prior driving according to claim 9, characterized in that In step S4, the cross-entropy loss is used as the objective function of the loss function to train the model, and the formula is as follows: , Among them, represents the predicted class value of the j-th sample, represents the class label of the j-th sample, represents the number of training set samples, represents the cross-entropy loss.
Citation Information
Patent Citations
Pedestrian behavior category detection method based on image sequence
CN113688761A
Air traffic controller unsafe behavior classification method based on monitoring video
CN115410035A
Alzheimer disease classification method and system based on multi-view spatio-temporal topology fusion
CN118899076A
Action segmentation network optimization method based on feature similarity
CN119672588A
Traffic performance index prediction method and device based on GIS map information
WO2022142418A1
Cited By
Medical image sequence positioning method and system
CN121544625A