Automated spontaneous imaging recognition method based on enhanced spatiotemporal feature network

Through an automated spontaneous development recognition method based on enhanced spatiotemporal feature network, the problems of high data cost and low recognition accuracy in the prior art are solved, efficient and accurate identification of spontaneous development is achieved, and clinical diagnosis is supported.

CN119399566BActive Publication Date: 2025-06-06SUZHOU UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510007159.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

The existing spontaneous development recognition technology has the problems of high data cost and low recognition accuracy. Especially when relying on multivariate variable analysis, it is difficult to effectively capture the dynamic changing characteristics of spontaneous development in echocardiography videos.

Method used

An automated spontaneous development and recognition method based on enhanced spatiotemporal feature network is adopted. By pre-processing transesophageal echocardiography videos, a collection of video frames is extracted, and a trained automated spontaneous development and recognition model is used, including attention map creation module, spatiotemporal feature extraction module and classification module, to output spontaneous development and recognition results.

Benefits of technology

It effectively reduces the difficulty and identification cost of clinical data, improves identification accuracy, and can quickly and accurately capture the spatial structure and dynamic changes of spontaneous development, supporting doctors' accurate assessment of the risk of thromboembolic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399566B_ABST
    Figure CN119399566B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of video recognition and automatic analysis of echocardiogram videos, and specifically provides an automatic spontaneous imaging recognition method based on an enhanced spatiotemporal feature network, including preprocessing a video to be recognized to obtain a video frame set; wherein the video to be recognized is a transesophageal echocardiogram video; the video frame set is input into a trained automatic spontaneous imaging recognition model, and the automatic spontaneous imaging recognition model is used to output a corresponding spontaneous imaging recognition result; wherein the automatic spontaneous imaging recognition model is obtained by training based on a plurality of training data, and the training data includes a transesophageal echocardiogram video; the automatic spontaneous imaging recognition model includes an attention map creation module, a spatiotemporal feature extraction module, and a classification module. The automatic spontaneous imaging recognition method based on an enhanced spatiotemporal feature network of the present invention can effectively reduce data costs and has high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video recognition and automatic analysis of ultrasound cardiogram videos, and in particular to an automatic spontaneous imaging recognition method based on an enhanced spatiotemporal feature network. Background Art

[0002] Spontaneous Echo Contrast (SEC) is an echo phenomenon resembling a smoke-like vortex observed in transesophageal echocardiography (TEE) videos. Spontaneous echo contrast is common in the left atrium or left atrial appendage of patients with atrial fibrillation and is usually associated with arrhythmia, slow blood flow, and hypercoagulable state. Patients with atrial fibrillation are often at risk of thromboembolism, resulting in high morbidity and mortality. A large number of studies have shown that spontaneous contrast is highly correlated with thromboembolic events and is an important independent indicator for assessing the risk of thromboembolism. Among them, dense spontaneous contrast is considered to be a precursor to thrombosis. Therefore, in order to prevent thrombosis and guide the preoperative and postoperative evaluation of patients with atrial fibrillation, it is crucial to accurately and efficiently identify the severity of spontaneous contrast.

[0003] In the prior art, the identification of spontaneous imaging mainly relies on the doctor's visual inspection of the transesophageal echocardiography video, which has the problems of low accuracy, time-consuming and strong subjectivity. There are also identification technologies that identify spontaneous imaging by using echocardiography-related parameters and other clinical data for multivariate analysis, but they have the problems of high data cost and low identification accuracy. Summary of the invention

[0004] The automatic spontaneous imaging recognition method based on enhanced spatiotemporal feature network provided by the embodiment of the present invention at least solves the problems of high data cost and low recognition accuracy of existing spontaneous imaging recognition, effectively reduces data cost and improves recognition accuracy.

[0005] The present invention provides an automated spontaneous imaging recognition method based on an enhanced spatiotemporal feature network, comprising preprocessing a video to be recognized to obtain a video frame set; wherein the video to be recognized is a transesophageal echocardiography video; the video frame set is input into a trained automated spontaneous imaging recognition model, and the automated spontaneous imaging recognition model is used to output a corresponding spontaneous imaging recognition result; wherein the automated spontaneous imaging recognition model is obtained by training based on a plurality of training data, and the training data includes a transesophageal echocardiography video; the automated spontaneous imaging recognition model includes an attention map creation module, a spatiotemporal feature extraction module and a classification module.

[0006] In one embodiment of the present invention, the video frame set is input into a trained automatic spontaneous development recognition model, and the automatic spontaneous development recognition model is used to output a corresponding spontaneous development recognition result, including: inputting the video frame set into the attention map creation module, creating an attention map of a spontaneous development key area for each video frame of the video frame set through the attention map creation module, and obtaining a dual-channel video frame set; transmitting the dual-channel video frame set to the spatiotemporal feature extraction module, processing the dual-channel video frame set through the spatiotemporal feature extraction module, capturing the spatial structural features and temporal dynamic features of spontaneous development, and obtaining a spatiotemporal representation of spontaneous development; transmitting the spatiotemporal representation of spontaneous development to the video frame set; and transmitting the spatiotemporal representation of spontaneous development to the video frame set. The classification module is provided, and the spatiotemporal representation of the spontaneous development is processed by the classification module to obtain the spontaneous development recognition result; wherein the spontaneous development recognition result includes the spontaneous development type and the spontaneous development degree; the classification module includes a main classification head and an auxiliary classification head, the main classification head is used to distinguish the spontaneous development type, the spontaneous development type includes a dense type and a non-dense type, the auxiliary classification head includes a first auxiliary classification head and a second auxiliary classification head, the first auxiliary classification head is used to distinguish the first degree, and the second auxiliary classification head is used to distinguish the second degree, the first degree is the spontaneous development degree corresponding to the non-dense type, and the second degree is the spontaneous development degree corresponding to the dense type.

[0007] In one embodiment of the present invention, an attention map of the spontaneous development key area is created for each video frame of the video frame set by the attention map creation module to obtain a dual-channel video frame set, including: processing the video frame set according to the pixel intensity of the spontaneous development key area, assigning a corresponding pixel intensity weight to each pixel, and obtaining a mean weight map; processing the video frame set according to the distance from the spontaneous development key area to a control point, assigning a corresponding distance weight to each pixel, and obtaining a distance weight map; multiplying each video frame of the video frame set with the mean weight map and the distance weight map element by element to obtain an attention map video frame set; splicing each video frame of the video frame set with the attention map video frame corresponding to the attention map video frame set along the channel dimension to obtain the dual-channel video frame set.

[0008] In one embodiment of the present invention, the video frame set is processed according to the pixel intensity of the spontaneous development key area, and a corresponding pixel intensity weight is assigned to each pixel to obtain a mean weight map, including: performing a local average pooling operation on the video frame set to obtain a local mean map; performing a minimum and maximum normalization operation on the local mean map to obtain a normalized mean map; processing the normalized mean map according to a mean weight function to obtain the mean weight map; the mean weight function It is expressed as:

[0009] ,

[0010] in, is the pixel value scaling factor, , is the threshold value, , exp function is the exponential function;

[0011] and / or,

[0012] The video frame set is processed according to the distance from the spontaneous development key area to the control point to obtain a distance weight map, including: calculating the distance from each pixel in the video frame set to the control point, assigning a corresponding distance weight to each pixel, and obtaining a distance map; performing a minimum and maximum normalization operation on the distance map to obtain a normalized distance map; processing the normalized distance map according to a distance weight function to obtain the distance weight map; the distance weight function It is expressed as:

[0013] ,

[0014] in, is the distance weight adjustment factor, ; is the distance scale scaling factor, .

[0015] In one embodiment of the present invention, the dual-channel video frame set is processed by the spatiotemporal feature extraction module to capture the spatial structural features and temporal dynamic features of spontaneous development, and obtain the spatiotemporal representation of spontaneous development, including: processing the dual-channel video frame set by a convolutional neural network to obtain spatial features; processing the spatial features by a multi-head self-attention mechanism to obtain spatiotemporal features; adding the spatial features to the spatiotemporal features through jump connections to obtain the spatiotemporal representation of spontaneous development.

[0016] In one embodiment of the present invention, the dual-channel video frame set is processed by a convolutional neural network to obtain spatial features, which are expressed as:

[0017] ,

[0018] in, is the spatial feature, , is the number of samples, is the channel dimension, is the number of frames, is the frame height, is the frame width; The function is a spatial convolution with a kernel size of 1×3×3. is the dual-channel video frame set;

[0019] The spatial features are processed by the multi-head self-attention mechanism to obtain the spatiotemporal features , expressed as;

[0020] ,

[0021] in, ; The function is the multi-head self-attention mechanism; the multi-head self-attention mechanism includes:

[0022] The spatial features are transformed and linearly mapped to obtain the query matrix , key matrix Sum Matrix ;in, , , , is the number of attention heads;

[0023] Calculate the attention score based on the query matrix and the key matrix , expressed as:

[0024] ,

[0025] The attention score is processed using a normalized exponential function to obtain the attention weight , expressed as:

[0026] ,

[0027] in, , Function is the normalized exponential function, is an element in one of the vectors in the last dimension of the attention score, is the number of elements in the corresponding vector;

[0028] According to the attention weight and the value matrix, the untransformed spatiotemporal features are calculated , expressed as:

[0029] ,

[0030] in, ;

[0031] The untransformed spatiotemporal features are transformed into the spatiotemporal features. .

[0032] In one embodiment of the present invention, the spontaneous development spatiotemporal representation is processed by the classification module to obtain the spontaneous development recognition result, including: performing a global average pooling operation on the spontaneous development spatiotemporal representation to obtain a classification input feature; transmitting the classification input feature to the main classification head, processing the classification input feature by the main classification head to obtain the spontaneous development type; according to the spontaneous development type, transmitting the corresponding classification input feature to the corresponding auxiliary classification head, processing the classification input feature by the auxiliary classification head to obtain the spontaneous development degree; according to the spontaneous development type and the spontaneous development degree, obtaining the spontaneous development recognition result.

[0033] In one embodiment of the present invention, before the video frame set is input into a trained automatic spontaneous imaging recognition model and the corresponding spontaneous imaging recognition result is outputted using the automatic spontaneous imaging recognition model, the method further includes: acquiring training data, wherein the training data includes historical transesophageal echocardiography videos and historical recognition results; wherein the historical transesophageal echocardiography videos are transesophageal echocardiography videos for which spontaneous imaging recognition results have been confirmed, and the historical recognition results are spontaneous imaging recognition results corresponding to the historical transesophageal echocardiography videos; preprocessing the historical transesophageal echocardiography videos to obtain a training video frame set; inputting the training video frame set into a pre-constructed automatic spontaneous imaging recognition model to be trained, and outputting the corresponding training recognition result using the automatic spontaneous imaging recognition model to be trained; obtaining a loss value of the training data using a loss function according to the historical recognition results and the training recognition results; and judging whether the loss value meets a preset requirement. If so, the training of the automatic spontaneous imaging recognition model is completed.

[0034] In one embodiment of the present invention, according to the historical recognition result and the training recognition result, using the loss function, the loss value of the training data is obtained, including: calculating the loss value by the following formula: :

[0035] ,

[0036] ,

[0037] in, is the main classification head loss, is the first auxiliary classification head loss, is the second auxiliary classification head loss, is the total number, is the number of non-dense types, is the number of dense types, , is the predicted score of the true category; is the loss weight adjustment factor.

[0038] In one embodiment of the present invention, a video to be identified is preprocessed to obtain a video frame set, including: performing frame extraction on the video to be identified to obtain an initial video frame set; performing color conversion processing on the initial video frame set to obtain a grayscale frame set; performing noise data removal processing on the grayscale frame set to obtain the video frame set.

[0039] The above technical solution of the present invention has the following beneficial effects compared with the prior art:

[0040] The automated spontaneous imaging recognition method based on the enhanced spatiotemporal feature network described in the present invention, first of all, only requires one type of data, the transesophageal echocardiography video, as input data, and no additional clinical data is required. This effectively reduces the difficulty of obtaining clinical data, and also eliminates the need to clean data with complex sources and uneven quality, effectively reducing the recognition cost. Secondly, by preprocessing the transesophageal echocardiography video, it is convenient for the subsequent processing of the automated spontaneous imaging recognition model to improve the recognition accuracy. Finally, through the automated spontaneous imaging recognition model, it can adapt to the characteristics of the transesophageal echocardiography video, and effectively capture the spatial structure and dynamic change characteristics of spontaneous imaging in the transesophageal echocardiography video, so as to quickly and accurately obtain the spontaneous imaging recognition results, help doctors clinically diagnose the risk of thromboembolism, and provide strong support for doctors' decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without creative work. In the drawings:

[0042] Figure 1 It is a flow chart of an automated spontaneous imaging recognition method based on an enhanced spatiotemporal feature network in a preferred embodiment of the present invention.

[0043] Figure 2 It is a partial flow chart of preprocessing in a preferred embodiment of the present invention.

[0044] Figure 3 It is a comparative schematic diagram before and after the noise data removal process in the preferred embodiment of the present invention.

[0045] Figure 4 It is a structural schematic diagram of the automated spontaneous development recognition model in a preferred embodiment of the present invention.

[0046] Figure 5 It is a flow chart of the attention map creation module in the preferred embodiment of the present invention.

[0047] Figure 6 It is a comparative schematic diagram before and after processing by the attention map creation module in the preferred embodiment of the present invention.

[0048] Figure 7 It is a data statistical diagram of the training set and the test set in the preferred embodiment of the present invention.

[0049] Figure 8 It is a schematic diagram of model performance verification in a preferred embodiment of the present invention.

[0050] Fig. 9 It is a schematic diagram of module performance verification in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0051] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not intended to limit the scope of protection of the present invention.

[0052] It should be noted that in the process of manual visual inspection of transesophageal echocardiography videos and identification of spontaneous imaging, due to the complex dynamic characteristics of spontaneous imaging, the accuracy of identification depends to a large extent on the doctor's professional knowledge and experience, resulting in significant subjectivity.

[0053] Some studies have proposed the use of other indicators to assist in the identification of spontaneous imaging. For example, spontaneous imaging can be quantitatively assessed by comparing the integrated backscatter (IBS) intensity of the left atrium and left ventricle. However, this method requires standardized views, gain calibration, and selection of regions of interest, which requires high expertise and experience from the operator. In addition, some studies have proposed predicting the occurrence of spontaneous imaging based on the left atrial diameter, and found that the incidence of spontaneous imaging increased significantly when the left atrial diameter exceeded a certain value. This requires additional measurement of the left atrial diameter, which is time-consuming and susceptible to human error.

[0054] In summary, the spontaneous imaging identification methods that require manual intervention have the problems of low accuracy, time-consuming and high subjectivity.

[0055] The existing spontaneous imaging recognition and automated analysis of echocardiography have the problems of high data cost and low recognition accuracy.

[0056] Specifically, the spontaneous imaging recognition method based on multivariate analysis relies on a large amount of clinical data as support, and clinical data is relatively difficult to obtain. In addition, due to the complex data sources and uneven quality, a large amount of data cleaning work is required, which greatly increases the implementation cost of the method. At the same time, spontaneous imaging is a complex cardiac abnormality detected in echocardiographic videos with significant dynamic characteristics. However, multivariate analysis relies on static clinical data for analysis, which often exhibits fixed numerical characteristics and lacks spatiotemporal continuity. Therefore, traditional static analysis methods cannot effectively capture the dynamic change characteristics of spontaneous imaging, limiting their accuracy and applicability in identifying and diagnosing spontaneous imaging.

[0057] There is also a way to directly transfer the model from the field of general video recognition to echocardiography video analysis, but it cannot effectively characterize the characteristics of echocardiography videos. This is because general video recognition methods are usually designed for natural scene videos and are suitable for clear and well-structured images, while echocardiography videos contain complex medical features and the imaging quality is easily affected by factors such as noise and artifacts. In addition, key features in echocardiography videos, such as hemodynamic changes, are crucial for identifying spontaneous imaging, and general video recognition methods often cannot fully capture these subtle differences.

[0058] In summary, directly migrating these methods usually leads to poor recognition results and is difficult to meet clinical requirements for accuracy and professionalism.

[0059] In order to solve the problems of high data cost and low recognition accuracy of existing spontaneous imaging recognition, Figure 1 As shown, the present invention provides an automated spontaneous imaging recognition method based on an enhanced spatiotemporal feature network, comprising:

[0060] The video to be identified is preprocessed to obtain a video frame set. The video to be identified is a transesophageal echocardiography video, which is converted into a video frame set by preprocessing the transesophageal echocardiography video. The video frame set includes multiple video frames, so that the subsequent processing of the automated spontaneous imaging recognition model can improve the recognition accuracy.

[0061] Since only one type of data, namely transesophageal echocardiography video, is required as input data, on the one hand, the difficulty of obtaining clinical data is effectively reduced; on the other hand, the cleaning work of data with complex sources and uneven quality is eliminated, effectively reducing the recognition cost.

[0062] After obtaining the video frame set, the video frame set is input into the trained automatic spontaneous imaging recognition model, and the automatic spontaneous imaging recognition model can be used to output the corresponding spontaneous imaging recognition results, so as to prevent thrombosis and guide the preoperative and postoperative evaluation of patients with atrial fibrillation based on the spontaneous imaging recognition results.

[0063] Among them, the automated spontaneous imaging recognition model is trained based on multiple training data, including transesophageal echocardiography videos.

[0064] The automated spontaneous imaging recognition model includes an attention map creation module, a spatiotemporal feature extraction module, and a classification module. The attention map creation module is used to create an attention map of the key area (i.e., the left atrial area) where spontaneous imaging occurs in the video frame of the transesophageal echocardiography video, thereby enhancing the information features of the key area and suppressing the information features of the invalid area. The spatiotemporal feature extraction module is used to accurately and efficiently extract the spatiotemporal features of spontaneous imaging based on the movement characteristics of the heart, i.e., spatial locality and long-term dependence, combined with the spatial convolutional neural network and the temporal multi-head self-attention mechanism, so that the classification module can use it for the final spontaneous imaging recognition classification and perform hierarchical classification and recognition of the severity of spontaneous imaging.

[0065] The automated spontaneous imaging recognition method based on the enhanced spatiotemporal feature network described in the present invention, first of all, only requires one type of data, the transesophageal echocardiography video, as input data, and no additional clinical data is required. This effectively reduces the difficulty of obtaining clinical data, and also eliminates the need to clean data with complex sources and uneven quality, effectively reducing the recognition cost. Secondly, by preprocessing the transesophageal echocardiography video, it is convenient for the subsequent processing of the automated spontaneous imaging recognition model to improve the recognition accuracy. Finally, through the automated spontaneous imaging recognition model, it can adapt to the characteristics of the transesophageal echocardiography video, and effectively capture the spatial structure and dynamic change characteristics of spontaneous imaging in the transesophageal echocardiography video, so as to quickly and accurately obtain the spontaneous imaging recognition results, help doctors clinically diagnose the risk of thromboembolism, and provide strong support for doctors' decision-making.

[0066] It should be noted that after preprocessing, each video to be identified can obtain at least one video frame set. That is, in some cases, after preprocessing, a video to be identified can obtain multiple video frame sets. By inputting these video frame sets corresponding to the same video to be identified into the automated spontaneous imaging recognition model, and then obtaining the corresponding spontaneous imaging recognition results, these spontaneous imaging recognition results are comprehensively measured to obtain a more accurate spontaneous imaging recognition result.

[0067] Reference Figure 2As shown, the automatic spontaneous imaging recognition method based on enhanced spatiotemporal feature network of the present invention, in some embodiments, preprocesses the video to be recognized to obtain a video frame set, including:

[0068] Frames of the video to be identified are extracted to obtain an initial video frame set.

[0069] The initial video frame set is subjected to color conversion processing to obtain a grayscale frame set.

[0070] The grayscale frame set is subjected to noise data removal processing to obtain a video frame set.

[0071] Specifically, by extracting frames from the input video to be identified, an initial video frame set can be obtained, and the initial video frame set includes multiple initial video frames, which facilitates frame-by-frame processing, thereby reducing subsequent data processing volume, improving efficiency and recognition accuracy.

[0072] For example, the video to be identified, i.e., the transesophageal echocardiography video to be identified, is subjected to frame extraction by the OpenCV visual processing library to obtain an initial video frame set. .in, For the Initial video frames, is the total number of frames of the input TEE video.

[0073] Considering that spontaneous development recognition is independent of color, and the initial video frame is usually a three-channel color video frame, in order to reduce the amount of data processing and improve recognition efficiency, the initial video frame set is color converted to obtain a grayscale frame set. The grayscale frame set includes multiple single-channel grayscale video frames.

[0074] Considering that the video frame includes other useless information, such as electrocardiogram data, date data, and angle data, in addition to the cardiac structure information required for spontaneous imaging recognition, in order to prevent these invalid information from affecting the extraction of spontaneous imaging features by the automated spontaneous imaging recognition model and affecting the final recognition accuracy, they need to be removed from the image and only the cardiac structure information is retained.

[0075] Therefore, after obtaining the grayscale frame set, it is necessary to perform noise data removal processing on it to obtain a video frame set after the noise data is removed. Specifically, in the grayscale video frame, the shape of the target area that does not need to be subjected to noise data removal is a sector. When performing noise data removal, the noise data removal can be divided into two parts, namely, removing the electrocardiogram data outside the arc of the sector and removing the electrocardiogram data outside the two radii of the sector.

[0076] Since the center position and radius length of the sector are known, the sector arc equation can be directly solved, and then the electrocardiogram data outside the arc can be removed, which will not be described here.

[0077] However, depending on the angle of the transesophageal echocardiography examination, the position of the sector in the grayscale video frame is also different. Therefore, when removing the data outside the two radii of the sector, it is necessary to calculate the equation of the straight line where the radius is located. Specifically, it includes:

[0078] The grayscale frame set is subjected to edge detection to obtain an edge binary image. The edge binary image is a binary image obtained by extracting edge information from the original image through an edge detection algorithm. In this image, edge points are marked as target colors, while non-edge areas are marked as background colors. Preferably, edge detection uses the Canny edge detection method.

[0079] The edge binary image is subjected to Hough line transformation to obtain the edge line and obtain the endpoint coordinate set of the edge line. Through Hough line transformation, the straight line can be detected accurately and effectively.

[0080] Traverse the endpoint coordinate set until two target lines are obtained. The target line is collinear with the target center of the target area. Specifically, determine whether each line is collinear with the target center of the target area. If so, the line is the target line, and determine whether two target lines have been obtained. If two target lines have been obtained, stop the current traversal and start calculating the equation; if two target lines have not been obtained, continue to determine the next edge line. If not, continue to determine the next edge line.

[0081] After finding the target straight line, the target straight line equation can be calculated based on two points on the target straight line. The two points used to calculate the equation are the center of the target circle and the intersection of the straight line and the fan-shaped arc.

[0082] After obtaining the target straight line equation, traverse each pixel of the grayscale frame set and set the pixel values ​​outside the target straight line equation to zero, thereby removing the useless data outside the two radii of the fan. Figure 3 shown.

[0083] After removing the electrocardiogram data outside the arc of the sector and the electrocardiogram data outside the two radii of the sector, a video frame set can be obtained. After the initial video frame set is processed by color conversion and noise data removal, the subsequent data processing volume can be effectively reduced, the efficiency can be improved, and invalid information can be prevented from affecting the automatic spontaneous imaging recognition model to extract spontaneous imaging features and improve recognition accuracy.

[0084] Reference Figure 4As shown, the automatic spontaneous imaging recognition method based on enhanced spatiotemporal feature network of the present invention, in some embodiments, inputs a video frame set into a trained automatic spontaneous imaging recognition model, and uses the automatic spontaneous imaging recognition model to output a corresponding spontaneous imaging recognition result, including:

[0085] The video frame set is input into the attention map creation module, and the attention map creation module creates an attention map of the spontaneous development key area for each video frame of the video frame set to obtain a dual-channel video frame set.

[0086] In this embodiment, the key area for spontaneous imaging is set to the left atrium (including the left atrial appendage), which is the main area where spontaneous imaging occurs and is crucial for distinguishing different degrees of spontaneous imaging.

[0087] However, due to the imaging principle of transesophageal echocardiography, the pixel intensity of the key area of ​​spontaneous imaging is low, while the pixel intensity of irrelevant areas (such as myocardial areas) is high, which hinders the automatic spontaneous imaging recognition model from extracting spontaneous imaging features. Therefore, before the model extracts features, an attention map is first created for each video frame through the attention map creation module to enhance the pixel intensity of the key area of ​​spontaneous imaging, so that subsequent extraction and recognition can be performed quickly and efficiently.

[0088] Since the movement of the heart in the transesophageal echocardiography video has the characteristics of spatial position locality and long-range dependence of the movement cycle, in order to make full use of this knowledge, so as to more efficiently and accurately extract the spatiotemporal features of spontaneous imaging, after creating the attention map, the obtained dual-channel video frame set is transmitted to the spatiotemporal feature extraction module, and the dual-channel video frame set is processed by the spatiotemporal feature extraction module to capture the spatial structural features and temporal dynamic features of spontaneous imaging, and obtain the spatiotemporal representation of spontaneous imaging. Preferably, multiple spatiotemporal feature extraction modules are stacked to capture the spatiotemporal representation of spontaneous imaging.

[0089] Finally, the spatiotemporal representation of spontaneous imaging is transmitted to the classification module. The spatiotemporal representation of spontaneous imaging is processed by the classification module to obtain the spontaneous imaging recognition result. Therefore, thrombosis can be prevented and preoperative and postoperative evaluation of patients with atrial fibrillation can be guided based on the spontaneous imaging recognition result.

[0090] The spontaneous imaging identification results include the spontaneous imaging type and the degree of spontaneous imaging. The spontaneous imaging type includes dense type and non-dense type.

[0091] The classification module includes a main classification head and an auxiliary classification head. The main classification head is used to distinguish whether the spontaneous development type is a dense type or a non-dense type. The auxiliary classification head includes a first auxiliary classification head and a second auxiliary classification head. The first auxiliary classification head is used to distinguish the first degree, that is, the degree of spontaneous development corresponding to the non-dense type; the second auxiliary classification head is used to distinguish the second degree, that is, the degree of spontaneous development corresponding to the dense type. Preferably, the first degree includes three severity levels, namely, level 0, level 1 and level 2; the second degree includes two severity levels, namely, level 0 and level 1. Figure 4 In the figure, non-dense level 0 means that the corresponding spontaneous development belongs to the non-dense type, and its spontaneous development degree is level 0; the rest are the same and will not be repeated.

[0092] Preferably, after the spontaneous imaging spatiotemporal representation is transmitted to the classification module, the spontaneous imaging spatiotemporal representation is first subjected to a global average pooling operation to obtain the classification input feature , Subsequently, the classification input features are transmitted to the main classification head, which processes the classification input features and outputs the spontaneous imaging type. The spontaneous imaging type includes the category scores of non-dense spontaneous imaging and dense spontaneous imaging. ,in, is the score of the non-dense spontaneous imaging category, It is the score of dense spontaneous imaging category.

[0093] Subsequently, according to the output result of the main classification head, the corresponding first auxiliary classification head or second auxiliary classification head is selected. The auxiliary classification head processes the classification input features to obtain the degree of spontaneous development. Specifically, the first auxiliary classification head outputs the first degree score corresponding to the non-dense type of spontaneous development. ,in, and The second auxiliary classification head outputs the second degree score corresponding to the dense type spontaneous development. ,in, and The two severity scores of dense spontaneous imaging are respectively indicated. This design can not only improve the classification accuracy but also be consistent with the clinical diagnosis process.

[0094] Preferably, after obtaining the category scores and degree scores corresponding to multiple video frame sets corresponding to the same video to be identified (i.e., the type of spontaneous development and the degree of spontaneous development), these scores are comprehensively weighed and their average values ​​are calculated, thereby outputting the spontaneous development identification results corresponding to the video to be identified, and the identification is more accurate.

[0095] Further, see Figure 5As shown, the automatic spontaneous imaging recognition method based on enhanced spatiotemporal feature network of the present invention, in some embodiments, creates an attention map of the spontaneous imaging key area for each video frame of the video frame set through an attention map creation module to obtain a dual-channel video frame set, including:

[0096] According to the characteristics of sparse pixel distribution and low pixel intensity in the key area of ​​spontaneous imaging in transesophageal echocardiography, the video frame set is processed, and the corresponding pixel intensity weight is assigned to each pixel to obtain the mean weight map. Specifically, it includes:

[0097] First, each video frame in the video frame set is subjected to local average pooling operation in turn to obtain the local mean map , , is the frame height, is the frame width.

[0098] Secondly, the local mean map is normalized to the minimum and maximum values ​​so that the local mean map The value range is scaled to , get the normalized mean graph , .

[0099] Finally, according to the mean weight function Processing Normalized Mean Plot , get the mean weight map , .

[0100] Mean weight function Monotonically decreasing, which can be expressed as:

[0101] ,

[0102] in, is the pixel value scaling factor, , is the threshold value, , The function is an exponential function. Preferably, Set to 3, Set to 0.3.

[0103] According to the characteristic that the distance between the pixels in the key area of ​​spontaneous imaging and the control point (i.e., the top of the heart view) is short in transesophageal echocardiography, the video frame set is processed, and the corresponding distance weight is assigned to each pixel to obtain a distance weight map. Specifically, it includes:

[0104] First, the distance from each pixel in the video frame set to the reference point is calculated to obtain the distance map , For example, the coordinates of the reference point, i.e., the top of the heart view, are assumed to be , and then the Euclidean distance from each pixel of each video frame to the coordinate is calculated.

[0105] Secondly, the distance map is normalized to the minimum and maximum values. The value range is scaled to , and obtain the normalized distance map , .

[0106] Finally, according to the distance weight function Processing normalized distance map , get the distance weight map , . Distance weight function It is expressed as:

[0107] ,

[0108] in, is the distance weight adjustment factor, ; is the distance scale scaling factor, Preferably, Set to 2, Set to 1.5.

[0109] In the mean weight graph and distance weight graph Then, each video frame in the video frame set , and the mean weight graph And the distance weight map Multiply element by element to get the attention map video frame set. The attention map video frame set includes multiple attention map video frames . It is expressed as:

[0110] ,

[0111] Among them, the symbol In the attention map video frame, pixels in the spontaneously developing key area are enhanced, while pixels in irrelevant areas are suppressed.

[0112] Reference Figure 6 As shown, each video frame in the video frame set is spliced ​​with the attention map video frame corresponding to the attention map video frame set along the channel dimension, that is, the enhanced attention map video frame is spliced ​​to the original video frame as key information supplement to obtain a dual-channel video frame set.

[0113] Furthermore, the method for automatically identifying spontaneous imaging based on an enhanced spatiotemporal feature network of the present invention, in some embodiments, processes a dual-channel video frame set through a spatiotemporal feature extraction module to capture the spatial structural features and temporal dynamic features of spontaneous imaging, and obtains a spatiotemporal representation of spontaneous imaging, including:

[0114] First, in the spatial dimension, a convolutional neural network is used to process a set of dual-channel video frames. , establish the local correlation of the heart spatial structure and obtain the spatial features .

[0115] Specifically, a convolutional neural network is used to perform spatial convolution on the dual-channel video frame set frame by frame, with a convolution kernel size of 1×3×3 to obtain the spatial features . It is expressed as:

[0116] ,

[0117] in, , is the number of samples, is the channel dimension, is the number of frames, is the frame height, is the frame width. The function is a spatial convolution with a kernel size of 1×3×3. is a collection of dual-channel video frames.

[0118] Secondly, in the time dimension, spatial features are processed through a multi-head self-attention mechanism , capturing long-term dependencies and obtaining spatiotemporal features . Expressed as;

[0119] ,

[0120] in, ; The function is a multi-head self-attention mechanism.

[0121] The multi-head self-attention mechanism includes:

[0122] The spatial features Perform morphological transformation and linear mapping to obtain the query matrix , key matrix Sum Matrix Among them, after morphological transformation .

[0123] in, , , , is the number of attention heads.

[0124] Calculate the attention score based on the query matrix and the key matrix , expressed as:

[0125] ,

[0126] It should be noted that here Represents the dot product of two matrices.

[0127] Use the normalized exponential function to process the attention score and get the attention weight , expressed as:

[0128] ,

[0129] in, is the attention weight, , The function is a normalized exponential function. is the attention score the elements of one of the vectors in the last dimension, is the number of elements in the vector.

[0130] According to the second attention weight and value matrix, the untransformed spatiotemporal features are calculated , expressed as:

[0131] ,

[0132] in, .

[0133] The untransformed spatiotemporal features are transformed into morphological form to obtain spatiotemporal features. .

[0134] Finally, the spatial features Adding spatiotemporal features via skip connections The spatiotemporal characterization of spontaneous imaging was obtained , expressed as:

[0135] .

[0136] By setting up the spatiotemporal feature extraction module, on the one hand, the characteristics of cardiac motion in transesophageal echocardiography videos can be fully utilized to effectively capture the spatiotemporal characteristics of spontaneous imaging. On the other hand, compared with the model that uses the global multi-head self-attention mechanism in both spatial and temporal dimensions, this design significantly reduces the number of model parameters and computational complexity, making it more efficient.

[0137] The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network of the present invention, in some embodiments, before inputting the video frame set into the trained automatic spontaneous imaging recognition model and outputting the corresponding spontaneous imaging recognition result using the automatic spontaneous imaging recognition model, further comprises:

[0138] Acquire training data. The training data includes historical transesophageal echocardiography videos and historical recognition results. The historical transesophageal echocardiography videos are transesophageal echocardiography videos with confirmed spontaneous imaging recognition results. The historical recognition results are spontaneous imaging recognition results corresponding to the historical transesophageal echocardiography videos. The historical recognition results include spontaneous imaging types and corresponding spontaneous imaging degrees. For example, a certain historical recognition result is a non-dense type of spontaneous imaging with a severity level of 1.

[0139] The historical transesophageal echocardiography video is preprocessed to obtain a training video frame set. In this embodiment, the preprocessing method is the same as the preprocessing method during actual recognition, that is, frame extraction, color conversion and noise data removal are performed, which will not be repeated here. Preferably, in the training stage, only one training video frame set is selected for each historical transesophageal echocardiography video, rather than selecting multiple sets and calculating the average value as in actual recognition.

[0140] The training video frame set is input into the pre-built automatic spontaneous imaging recognition model to be trained, and the corresponding training recognition result is output by the automatic spontaneous imaging recognition model to be trained. In this embodiment, the method in which the automatic spontaneous imaging recognition model to be trained outputs the training recognition result is the same as the pre-processing method during actual recognition, that is, the training video frame set is processed by the attention map creation module, the spatiotemporal feature extraction module and the classification module, which will not be repeated here.

[0141] According to the historical recognition results and the training recognition results, the loss value of the training data is obtained using the loss function. Specifically, it includes:

[0142] The loss value is calculated by the following formula :

[0143] ,

[0144] ,

[0145] in, is the main classification head loss, is the first auxiliary classification head loss, is the second auxiliary classification head loss, is the total number, is the number of non-dense types, is the number of dense types, . is the predicted score of the true category. Taking the main classification head as an example, the main classification head will output a non-dense spontaneous imaging category score and a dense spontaneous imaging category score according to the spatiotemporal representation of spontaneous imaging. Assuming that the true category corresponding to the spatiotemporal representation of spontaneous imaging is a non-dense type of spontaneous imaging, the non-dense spontaneous imaging category score is the predicted score of the true category. is the loss weight adjustment factor.

[0146] It is determined whether the loss value meets the preset requirements. If so, the training of the automatic spontaneous imaging recognition model is completed.

[0147] It is worth noting that during the training process, the main classification head calculates the loss of all training video frames; the first auxiliary classification head only calculates the loss of the corresponding non-dense type training video frame set; the second auxiliary classification head only calculates the loss of the corresponding dense type training video frame set. The total loss value is obtained by adding up the loss values Used for back propagation.

[0148] Preferably, if the preset requirements are not met, the training data of the model is readjusted, and training is continued by increasing the number of training rounds, expanding the data set, etc., until the loss value meets the preset requirements.

[0149] Reference Figure 7 As shown, in order to train the automatic spontaneous imaging recognition model to be trained, a total of 1106 transesophageal echocardiography videos used in the clinic of the First Affiliated Hospital of Soochow University from 2018 to 2023 were collected and divided into training set and test set in the ratio of 80% and 20%. Figure 7 In the figure, non-dense level 0 means that the corresponding spontaneous development belongs to the non-dense type, and its spontaneous development degree is level 0; the rest are the same and will not be repeated.

[0150] In order to verify the effectiveness of the automatic spontaneous imaging recognition method based on the enhanced spatiotemporal feature network described in the present invention, experiments were conducted based on the test set. Among them, all experiments were conducted on a single NVIDIA Tesla V100 GPU.

[0151] During the training phase, 48 frames were randomly selected from the complete video of the training set and edited with a step size of 2, and then a region with a spatial resolution of 216×309 was cropped (cropped from the spatial resolution of 240×344) to obtain the input historical transesophageal echocardiography video, which was a 48×216×309 clip.

[0152] In the testing phase, historical transesophageal echocardiography videos were used as input using all clips from the full video of the testing set, with only one center crop.

[0153] To evaluate the performance, several representative spatiotemporal representation methods were compared during the validation, including P3D63, SlowFastR50, TimeSformer, VideoSwinTransformer-T, and VideoSwinTransformer-S. Among them, P3D63 and SlowFastR50 are models based on pure convolutional neural networks; TimeSformer, VideoSwinTransformer-T, and VideoSwinTransformer-S are models based on pure multi-head self-attention mechanisms.

[0154] When measuring classification and recognition performance, Accuracy, Sensitivity (Recall) and Precision are used. Among them, Accuracy is used to evaluate the overall classification accuracy. Sensitivity is used to evaluate the recall rate of the positive class, that is, the successful detection rate of dense spontaneous imaging, which is of great significance for clinical diagnosis. Precision is used to evaluate the recognition accuracy of the positive class, that is, the precise recognition performance of dense spontaneous imaging. When calculating these indicators, dense type spontaneous imaging is regarded as a positive class, and non-dense type spontaneous imaging is regarded as a negative class. The calculation method of each evaluation indicator is expressed as:

[0155] ,

[0156] ,

[0157] ,

[0158] in, True Positive, that is, the number of test samples that are actually dense type spontaneous display images that are correctly identified; True Negative, that is, the number of test samples that are actually non-dense spontaneous images that are correctly identified; is a false positive example, that is, the number of test samples that are actually dense type spontaneous display images but are incorrectly identified; is the false negative example (FalseNegative), that is, the number of test samples that are actually non-dense type spontaneous display images that are incorrectly identified.

[0159] Experimental results refer to Figure 8As shown, it can be seen that the automatic spontaneous imaging recognition method based on enhanced spatiotemporal feature network of the present invention has high accuracy. At the same time, it also shows that for characteristic medical imaging videos such as transesophageal echocardiography videos, it does not have global correlation in space, but has long-range dependence in time (motion cycle).

[0160] In addition, refer to Fig. 9 As shown, ablation experiments were also conducted on different spatiotemporal feature extraction modules. Among them, module No. 1 is the spatiotemporal feature extraction module described in the above embodiment. The difference between module No. 2 and module No. 1 is that it first performs multi-head self-attention mechanism processing and then processes through convolutional neural network. The difference between module No. 3 and module No. 1 is that it independently performs multi-head self-attention mechanism processing and convolutional neural network processing, and then weights and sums the results.

[0161] According to the results of the ablation experiment, the first module has the best performance. The reason is that for the second module, the global temporal correlation is first performed, which may weaken the local spatial correlation, thus affecting the subsequent spatial feature aggregation; for the third module, the independent modeling of spatial and temporal features may destroy the integrity of spatiotemporal features.

[0162] It should be noted that the term "including" and its variations used in the embodiments of the present invention are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".

[0163] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0164] The various steps described in the method implementation methods provided by the embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method implementation methods may include additional steps and / or omit the steps shown. The scope of protection of the present invention is not limited in this respect.

[0165] The term "embodiment" in this specification refers to specific features, structures or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily mean the same embodiment, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. The various embodiments in this specification are described in a related manner, and the same or similar parts between the various embodiments refer to each other. In particular, for the device, equipment, and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts refer to the partial description of the method embodiment.

[0166] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of protection. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the attached claims.

Claims

1. An automated spontaneous imaging recognition method based on an enhanced spatiotemporal feature network, characterized in that: include: Preprocessing the video to be identified to obtain a video frame set; wherein the video to be identified is a transesophageal echocardiography video; The video frame set is input into a trained automatic spontaneous development recognition model, and the automatic spontaneous development recognition model is used to output the corresponding spontaneous development recognition result; wherein, the video frame set is input into an attention map creation module of the automatic spontaneous development recognition model, and an attention map of a key area of ​​spontaneous development is created for each video frame of the video frame set by the attention map creation module to obtain a dual-channel video frame set; the dual-channel video frame set is transmitted to a spatiotemporal feature extraction module of the automatic spontaneous development recognition model, and the dual-channel video frame set is processed by the spatiotemporal feature extraction module to capture the spatial structural features and temporal dynamic features of spontaneous development and obtain a spatiotemporal representation of spontaneous development; the spatiotemporal representation of spontaneous development is extracted from the spatiotemporal feature extraction module. The spontaneous development characteristics are transmitted to the classification module of the automatic spontaneous development recognition model, and the spontaneous development spatiotemporal representation is processed by the classification module to obtain the spontaneous development recognition result; the spontaneous development recognition result includes the spontaneous development type and the spontaneous development degree; the classification module includes a main classification head and an auxiliary classification head, the main classification head is used to distinguish the spontaneous development type, the spontaneous development type includes a dense type and a non-dense type, the auxiliary classification head includes a first auxiliary classification head and a second auxiliary classification head, the first auxiliary classification head is used to distinguish the first degree, and the second auxiliary classification head is used to distinguish the second degree, the first degree is the spontaneous development degree corresponding to the non-dense type, and the second degree is the spontaneous development degree corresponding to the dense type; The automated spontaneous echocardiography recognition model is trained based on a plurality of training data, wherein the training data includes transesophageal echocardiography videos.

2. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 1, characterized in that: The attention map creation module creates an attention map of the spontaneous development key area for each video frame of the video frame set to obtain a dual-channel video frame set, including: Processing the video frame set according to the pixel intensity of the spontaneous development key area, assigning a corresponding pixel intensity weight to each pixel, and obtaining a mean weight map; Processing the video frame set according to the distance from the spontaneous development key area to the control point, assigning a corresponding distance weight to each pixel, and obtaining a distance weight map; Multiply each video frame of the video frame set by the mean weight map and the distance weight map element by element to obtain an attention map video frame set; Each video frame of the video frame set is spliced ​​with the attention map video frame corresponding to the attention map video frame set along the channel dimension to obtain the dual-channel video frame set.

3. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 2, characterized in that: The video frame set is processed according to the pixel intensity of the spontaneous development key area, and a corresponding pixel intensity weight is assigned to each pixel to obtain a mean weight map, including: Performing a local average pooling operation on the video frame set to obtain a local mean map; Performing a minimum-maximum normalization operation on the local mean map to obtain a normalized mean map; The normalized mean map is processed according to the mean weight function to obtain the mean weight map; the mean weight function It is expressed as: , in, is the pixel value scaling factor, , is the threshold value, , exp function is the exponential function; and / or, Processing the video frame set according to the distance from the spontaneous development key area to the control point to obtain a distance weight map, including: Calculating the distance from each pixel in the video frame set to the reference point, assigning a corresponding distance weight to each pixel, and obtaining a distance map; Performing a minimum-maximum normalization operation on the distance map to obtain a normalized distance map; Processing the normalized distance graph according to the distance weight function to obtain the distance weight graph; the distance weight function It is expressed as: , in, is the distance weight adjustment factor, ; is the distance scale scaling factor, .

4. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 1, characterized in that: The dual-channel video frame set is processed by the spatiotemporal feature extraction module to capture the spatial structural features and temporal dynamic features of the spontaneous development, and obtain the spatiotemporal representation of the spontaneous development, including: Processing the dual-channel video frame set through a convolutional neural network to obtain spatial features; Processing the spatial features through a multi-head self-attention mechanism to obtain spatiotemporal features; The spatial features are added to the spatiotemporal features through jump connections to obtain the spontaneous development spatiotemporal representation.

5. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 4, characterized in that: The dual-channel video frame set is processed by a convolutional neural network to obtain spatial features, which are expressed as: , in, is the spatial feature, , is the number of samples, is the channel dimension, is the number of frames, is the frame height, is the frame width; The function is a spatial convolution with a kernel size of 1×3×3. is the dual-channel video frame set; The spatial features are processed by the multi-head self-attention mechanism to obtain the spatiotemporal features , expressed as: , in, ; The function is the multi-head self-attention mechanism; the multi-head self-attention mechanism includes: The spatial features are transformed and linearly mapped to obtain the query matrix , key matrix Sum Matrix ;in, , , , is the number of attention heads; Calculate the attention score based on the query matrix and the key matrix , expressed as: , The attention score is processed using a normalized exponential function to obtain the attention weight , expressed as: , in, , Function is the normalized exponential function, is an element in one of the vectors in the last dimension of the attention score, is the number of elements in the corresponding vector; According to the attention weight and the value matrix, the untransformed spatiotemporal features are calculated , expressed as: , in, ; The untransformed spatiotemporal features are transformed into the spatiotemporal features. .

6. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 1, characterized in that: Processing the spatiotemporal representation of the spontaneous imaging by the classification module to obtain the spontaneous imaging recognition result includes: Performing a global average pooling operation on the spontaneous imaging spatiotemporal representation to obtain a classification input feature; Transmitting the classification input feature to the main classification head, processing the classification input feature by the main classification head to obtain the spontaneous development type; According to the spontaneous development type, the corresponding classification input feature is transmitted to the corresponding auxiliary classification head, and the classification input feature is processed by the auxiliary classification head to obtain the spontaneous development degree; The spontaneous development recognition result is obtained according to the spontaneous development type and the spontaneous development degree.

7. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 1, characterized in that: Before inputting the video frame set into a trained automatic spontaneous imaging recognition model and using the automatic spontaneous imaging recognition model to output a corresponding spontaneous imaging recognition result, the method further includes: Acquire training data, wherein the training data includes historical transesophageal echocardiography videos and historical recognition results; wherein the historical transesophageal echocardiography videos are transesophageal echocardiography videos for which spontaneous imaging recognition results have been confirmed, and the historical recognition results are spontaneous imaging recognition results corresponding to the historical transesophageal echocardiography videos; Preprocessing the historical transesophageal echocardiography video to obtain a training video frame set; Inputting the training video frame set into a pre-built automatic spontaneous imaging recognition model to be trained, and using the automatic spontaneous imaging recognition model to be trained to output corresponding training recognition results; According to the historical recognition result and the training recognition result, using a loss function, obtaining a loss value of the training data; It is determined whether the loss value meets the preset requirement. If so, the training of the automatic spontaneous imaging recognition model is completed.

8. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 7, characterized in that: According to the historical recognition result and the training recognition result, using the loss function, the loss value of the training data is obtained, including: The loss value is calculated by the following formula: : , , in, is the main classification head loss, is the first auxiliary classification head loss, is the second auxiliary classification head loss, is the total number, is the number of non-dense types, is the number of dense types, , is the predicted score of the true category; is the loss weight adjustment factor.

9. The method for automatic spontaneous imaging recognition based on enhanced spatiotemporal feature network according to claim 1, characterized in that: Preprocess the video to be identified to obtain a set of video frames, including: Extract frames from the video to be identified to obtain an initial video frame set; Performing color conversion processing on the initial video frame set to obtain a grayscale frame set; The grayscale frame set is subjected to noise data removal processing to obtain the video frame set.

Citation Information

Patent Citations

  • Behavior recognition method, device and system and storage medium

    CN119091496A