Heart Echocardiogram Assistant Judgment Device Based on Semantic Segmentation and Video Understanding

Through semantic segmentation and video understanding algorithms based on BEIT, RAFT and R(2+1)D, a cardiac color ultrasound assisted judgment device is constructed, which solves the problem of time-consuming traditional heart disease judgment, achieves rapid and accurate heart disease judgment, and optimizes the use of medical resources.

CN115908287BActive Publication Date: 2025-07-11FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211382707.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-07-11
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

The traditional heart disease determination process is time-consuming and expensive, and medical resources are seriously wasted, so it is difficult for the existing technology to quickly and accurately extract cardiac dynamic and static indicators from cardiac color ultrasound videos.

Method used

Using visual technologies such as BEIT, RAFT and R(2+1)D, combined with semantic segmentation and video understanding algorithms, a cardiac color ultrasound assisted judgment device is constructed, and the cardiac dynamic and static indicators are obtained through cardiac color ultrasound videos, and evaluation is used for evaluation to achieve fast and accurate diagnosis of the disease.

Benefits of technology

It greatly reduces the time for doctors to judge, optimizes medical resources, and quickly and accurately identify heart diseases. It only takes 0.5 seconds to identify videos, taking into account both accuracy and speed to complete auxiliary diagnosis of cardiac color ultrasound.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908287B_ABST
    Figure CN115908287B_ABST
Patent Text Reader

Abstract

The present invention proposes a cardiac color Doppler ultrasound assisted judgment device based on semantic segmentation and video understanding, based on a computer system, including: a cardiac color Doppler ultrasound video semantic segmentation module for performing semantic segmentation on a cardiac color Doppler ultrasound video using an encoder-decoder based on BEIT to obtain a binary cardiac contour video; a cardiac static index acquisition module for analyzing the cardiac contour video obtained by the cardiac color Doppler ultrasound video semantic segmentation module using a neural network model based on RAFT to obtain cardiac static indexes from the maximum contour frame and the minimum contour frame; a cardiac dynamic index acquisition module for analyzing the cardiac color Doppler ultrasound video obtained by the cardiac color Doppler ultrasound video semantic segmentation module using a video understanding algorithm based on R(2+1)D to obtain cardiac dynamic indexes; and an evaluation module for evaluating the cardiac static indexes and the cardiac dynamic indexes using an evaluation function and calculating the relationship with the type of heart disease using a linear regression algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of machine learning, medical auxiliary equipment, etc., and particularly relates to a cardiac color Doppler ultrasound auxiliary judgment device based on semantic segmentation and video understanding. Background Art

[0002] With the continuous development of technology and social progress, the combination of AI and medicine has become one of the relatively popular AI applications at present. Among them, for the determination of the type of traditional heart disease, doctors need to perform a series of examinations including cardiac color Doppler ultrasound videos, routine blood analysis indicators, and risk factors to make a determination. The entire process takes a long time and costs a lot. Summary of the Invention

[0003] To solve the problems of defects and deficiencies existing in the prior art, the present invention proposes a cardiac color Doppler ultrasound auxiliary judgment device based on semantic segmentation and video understanding, and proposes the design of a cardiac color Doppler ultrasound auxiliary evaluation device based on visual technologies such as BEIT, RAFT, R(2+1)D, etc. By obtaining the cardiac dynamic indicators and cardiac static indicators of the patient from the cardiac color Doppler ultrasound video, it can significantly reduce the doctor's judgment time, provide a very good reference for the doctor to make a judgment, optimize medical resources, and reduce medical accidents.

[0004] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0005] A cardiac color Doppler ultrasound auxiliary judgment device based on semantic segmentation and video understanding, characterized in that, based on a computer system, it includes:

[0006] A cardiac color Doppler ultrasound video semantic segmentation module, which is used to perform semantic segmentation on the cardiac color Doppler ultrasound video using an encoder-decoder based on BEIT to obtain a binary cardiac contour video;

[0007] A cardiac static index acquisition module, which is used to analyze the cardiac contour video obtained by the cardiac color Doppler ultrasound video semantic segmentation module using a neural network model based on RAFT, and obtain cardiac static indicators from the maximum contour frame and the minimum contour frame;

[0008] A cardiac dynamic index acquisition module, which is used to analyze the cardiac color Doppler ultrasound video using a video understanding algorithm based on R(2+1)D to obtain cardiac dynamic indicators;

[0009] An evaluation module, which is used to evaluate the cardiac static indicators and cardiac dynamic indicators using an evaluation function to obtain corresponding scores, and calculate the relationship with the type of heart disease using a linear regression algorithm.

[0010] Furthermore, the construction and working process of the cardiac color Doppler ultrasound video semantic segmentation module specifically include the following steps:

[0011] Step S11: Conduct training for the heart contour. Use Labelimg to annotate echocardiogram videos to obtain relevant annotations of video contour training data;

[0012] Step S12: Employ a 12-layer BEIT as the backbone network. Add cross-attention and depth convolutional kernels to the Block of BEIT to form a BEIT adapter; Regularize the downsampling layer, increase the receptive field through spatial pyramid pooling, and replace the third local attention kernel with a global attention kernel to enable capturing global features. At the same time, perform a cross-attention mechanism between the fifth local attention kernel and the third global attention kernel. Replace the sixth and seventh local attention kernels with two convolutional kernels of size 5×5 and stride 2, and finally connect a 1×1 convolutional kernel for dimensionality reduction to prevent information loss while reducing the number of parameters;

[0013] Step S13: Use the BEIT adapter in Step S12 to replace the feature extraction backbone of the semantic segmentation model deeplabv3. After the result processed by the SPP structure, extract image features at multiple levels through the Upernet network, and output four feature layers of different sizes (7, 7), (14, 14), (28, 28), (56, 56) through the connection with the extraction result of the backbone feature extraction network;

[0014] Step S14: Obtain the regression loss function KIOU by considering the overlap between the label and the prediction and the distance from the boundary Loss It is defined as follows:

[0015]

[0016] where IOU is the ratio of the intersection of the predicted region and the target region to the union of the predicted region and the target region, Distance_2 2 is the Euclidean distance between the centers of the predicted region and the target region, α is the weight coefficient for Distance_2 2 and IOU; β is a weight coefficient for the influencing factor, and k is a parameter measuring the aspect ratio consistency, which is defined as follows:

[0017]

[0018] where π is the pi, tan represents the tangent function, w gt represents the width of the target region, h gt is the height of the target region, w p represents the width of the predicted region, h p represents the height of the predicted region, where γ is the weight coefficient for adjusting the measurement of aspect ratio consistency;

[0019] Step S15: Use the semantic segmentation network based on BEIT trained through Steps S11 - S14 to perform echocardiogram video segmentation on the echocardiogram video, obtaining a binary heart contour video, where the pixels in the heart region are white and the pixels in the background region are black.

[0020] Furthermore, the construction and working process of the heart static index acquisition module specifically includes the following steps:

[0021] Step S21: Use the RAFT network as the network for analyzing the heart contour video, use CSPDarknet53 as the feature extraction backbone network, replace the 8th convolutional kernel with a depthwise separable convolutional kernel, set the stride to 2, change the stride of the secondary downsampling feature map from (2, 2) to (1, 1), add an average pooling layer to facilitate feature input into the RNN, and connect a convolutional kernel to prevent information loss, finally outputting the image convolutional feature map of two frames to achieve feature extraction;

[0022] Step S22: Calculate the similarity of the feature map vectors of two frames. The similarity is calculated by dot product, and the similarity Similarity is defined as follows:

[0023]

[0024] where Similarity is used to calculate the similarity between two vectors. When the value of Similarity is closer to 1, it indicates that the two frame feature maps are closer; i represents the index position in the vector, n represents the length of the vector, x represents the feature map vector of the previous frame, and y represents the feature map vector of the next frame;

[0025] Step S23: Label the heart contour video obtained by the echocardiogram video semantic segmentation module through the given end - diastolic frame index and end - systolic frame index for training;

[0026] Step S24: Use the trained RAFT to predict the heart contour video obtained by the echocardiogram video semantic segmentation module to obtain the required end - diastolic frame and end - systolic frame, and count the number of white pixels in the diastolic frame image and the systolic frame image through opencv to obtain the static indexes of the echocardiogram, namely the end - diastolic volume EDV and the end - systolic volume ESV;

[0027] Step S25: Calculate the ejection fraction EF of the heart of this object through the static indexes. The ejection fraction EF is defined as follows:

[0028] EF = (EDV - ESV) / EDV * 100%

[0029] where EF is the ejection fraction, EDV is the end - diastolic volume of the heart, and ESV is the end - systolic volume of the heart.

[0030] Furthermore, the construction and working process of the cardiac dynamic index acquisition module include the following steps:

[0031] Step S31: Manually annotate the peak velocity during cardiac systole and the velocity during cardiac diastole for each cardiac ultrasound video;

[0032] Step S32: Use R(2+1)D as the network model for video feature extraction, and replace the 2D convolution therein with separable 3D convolution. The first 7×7 convolution kernel with a stride of 2 is replaced by three 3×3 convolution kernels with a stride of 1, so that the output size of the last feature map remains 7×7. In addition, connect a 1×1 convolution with a stride of 1 to prevent information loss;

[0033] Replace the average pooling layer with a convolution layer of size 5×5 to better obtain the probability map and binary map;

[0034] Step S33: After obtaining the probability map and binary map, calculate the losses of the probability map and binary map through binary cross-entropy. The loss function is defined as follows:

[0035] L = arctan(θ × L s + δ × L b )

[0036] where θ is the weight parameter of the probability map loss, L s is the loss value of the probability map; δ is the weight function of the binary map loss value, L b represents the binary map loss value, and arctan represents the arctangent function;

[0037] Step S35: Perform Monte Carlo partitioning on the dataset labeled in S31, use the training set to train the improved R(2+1)D model, save the model with the highest accuracy on the validation set, and perform prediction on the test set to obtain the dynamic indexes of the cardiac ultrasound video, namely the peak velocity V during diastole stretch and the peak velocity V during systole shrink .

[0038] Furthermore, the construction and working process of the evaluation module specifically include the following steps:

[0039] Step S41: Use the dynamic indexes and static indexes obtained by the cardiac dynamic index acquisition module and the cardiac static index acquisition module as training sample indexes, and label the type of heart disease corresponding to each heart as a label, and use a random forest classifier for training;

[0040] Step S42: Obtain the trained parameters, and its final score is Score; it is defined as follows:

[0041]

[0042] Where ρ is the weight parameter of the ejection fraction EF, and EF is the cardiac ejection fraction. is the weight parameter of the peak velocity in diastole of the heart, and V stretch is the peak velocity in diastole of the heart, and τ is the weight parameter of the peak velocity in systole of the heart, and V shrink is the peak velocity in systole of the heart;

[0043] Step S43: Linearize the obtained score to obtain the classification relationship between the score and the type of heart disease.

[0044] Compared with the prior art, the present invention and its preferred solutions have the following beneficial effects:

[0045] 1. Replace part of the local attention mechanism in the BEIT structure with a global attention mechanism, which not only retains the global features but also strengthens the extraction of local features, and improves the problem of blurred segmentation edges.

[0046] 2. Using the SPP structure can convert multi-scale feature maps into fixed-size feature vectors, mainly improving the problem of image distortion and the blur problem caused by size scaling.

[0047] 3. In the field of traditional cardiac color Doppler ultrasound diagnosis, a new objective algorithm model for rapid evaluation combining ejection fraction and cardiac flow velocity is proposed.

[0048] 4. A cardiac color Doppler ultrasound auxiliary judgment device based on semantic segmentation and video understanding is proposed, which only takes 0.5 s to recognize a video, achieving a balance between accuracy and speed.

[0049] 5. Combining multiple computer vision algorithms, it has good detection accuracy in videos, completes part of the abnormal interference processing, and can better complete the cardiac color Doppler ultrasound auxiliary diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0051] Figure 1 is the implementation flowchart of the construction method of the device in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0053] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0054] Note that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0055] As Figure 1 shown, based on the present invention, a cardiac color Doppler ultrasound assisted judgment device based on semantic segmentation and video understanding is provided. The construction and development process of this embodiment is summarized into the following steps:

[0056] Step S1: Use an encoder-decoder based on BEIT to perform semantic segmentation on the cardiac color Doppler ultrasound video to obtain a binary cardiac contour video;

[0057] Step S11: Perform training for the cardiac contour. Use Labelimg to annotate the cardiac color Doppler ultrasound video to obtain relevant annotations of the video contour training data;

[0058] Step S12: Adopt a 12-layer BEIT as the backbone network. Add cross-attention and depth convolutional kernels to the Block of BEIT to form a BEIT adapter. Regularize the downsampling layer, increase the receptive field through spatial pyramid pooling, and replace the third local attention kernel with a global attention kernel so as to be able to capture global features. At the same time, perform a cross-attention mechanism on the fifth local attention kernel and the third global attention kernel, replace the sixth and seventh local attention kernels with two convolutional kernels of size 5×5 and stride 2, and finally connect a 1×1 convolutional kernel for dimensionality reduction to prevent information loss while reducing the number of parameters.

[0059] Step S13: Use the BEIT adapter in S12 to replace the feature extraction backbone of the traditional semantic segmentation model deeplabv3. After the result processed by the SPP structure, use the Upernet network to extract image features at multiple levels, and output four feature layers of different sizes (7, 7), (14, 14), (28, 28), (56, 56) through the connection with the result extracted by the backbone feature extraction network.

[0060] Step S14: Obtain the regression loss function KIOU by considering the overlap between the label and the prediction and the distance from the boundary LossThe definitions are as follows:

[0061]

[0062] Among them, IOU is the ratio of the intersection of the predicted region and the target region to the union of the predicted region and the target region, and Distance_2 2 is the Euclidean distance between the center points of the predicted region and the target region, α is for Distance_2 2 and the weight coefficient of IOU. β is a weight coefficient for the influencing factor, and k is a parameter for measuring the aspect ratio consistency, which is defined as follows:

[0063]

[0064] Among them, π is the mathematical constant pi, tan represents the tangent function in mathematics, w gt represents the width of the target region, h gt the height of the target region, w p represents the width of the predicted region, h p represents the height of the predicted region, where γ is the weight coefficient for adjusting the measurement of aspect ratio consistency.

[0065] Step S15: Use the above-trained semantic segmentation network based on BEIT to perform echocardiogram video segmentation to obtain a binary heart contour video, where the pixels in the heart region are white and the pixels in the background region are black.

[0066] Step S2: Use the neural network model based on RAFT to analyze the heart contour video obtained in Step S1, and obtain the static heart indicators from the maximum contour frame and the minimum contour frame. The specific steps are as follows:

[0067] Step S21: Use the RAFT network as the network for analyzing the heart contour video, use CSPDarknet53 as the feature extraction backbone network, replace the 8th convolutional kernel with a depthwise separable convolutional kernel, set the stride to 2, and at the same time change the stride of the secondary downsampling feature map from (2, 2) to (1, 1), add an average pooling layer to facilitate the feature input into the RNN, and connect a convolutional kernel to prevent information loss, and finally output the image convolutional feature map of two frames to achieve feature extraction;

[0068] Step S22: Calculate the similarity of the feature map vectors of two frames. The similarity is calculated by dot product, and the similarity Similarity is defined as follows:

[0069]

[0070] Among them, Similarity is used to calculate the similarity between two vectors. When the value of Similarity is closer to 1, it indicates that the two frame feature maps are closer. i represents the index position in the vector, n represents the length of the vector, x represents the feature map vector of the previous frame, and y represents the feature map vector of the next frame;

[0071] Step S23: Annotate the cardiac contour video obtained in S1 through the given end-diastolic frame index and end-systolic frame index for training.

[0072] Step S24: Use the trained RAFT to predict the cardiac contour video obtained in S1 to obtain the required end-diastolic frame and end-systolic frame. Count the number of white pixels in the diastolic frame image and systolic frame image through opencv to obtain the static indicators of the cardiac ultrasound, namely end-diastolic volume EDV and end-systolic volume ESV.

[0073] Step S25: Calculate the ejection fraction EF of the heart of this subject through the static indicators. The ejection fraction EF is defined as follows:

[0074] EF = (EDV - ESV) / EDV * 100%

[0075] Among them, EF is the ejection fraction, EDV is the end-diastolic volume of the heart, and ESV is the end-systolic volume of the heart.

[0076] Step S3: Use the video understanding algorithm based on R(2 + 1)D to analyze the cardiac ultrasound video to obtain cardiac dynamic indicators. Specifically, it includes the following steps:

[0077] Step S31: Manually annotate the peak velocity in cardiac systole and the velocity in cardiac diastole of each cardiac ultrasound video.

[0078] Step S32: Adopt R(2 + 1)D as the network model for video feature extraction, and replace the 2D convolution in it with separable 3D convolution. The first convolution kernel with a size of 7×7 and a stride of 2 is replaced by three convolution kernels with a size of 3×3 and a stride of 1; so that the output size of the last feature map remains 7×7; in addition, connect a 1×1 convolution with a stride of 1 to prevent information loss.

[0079] Adopt a convolution layer with a size of 5×5 to replace the average pooling layer, which can better obtain the probability map and binary map.

[0080] Step S34: After obtaining the probability map and binary map, calculate the loss between the probability map and the binary map through binary cross-entropy. The loss function is defined as follows:

[0081] L = arctan(θ × L s +δ × Lb )

[0082] where θ is the weight parameter of the probability graph loss, and L s is the loss value of the probability graph. δ is the weight function of the binary graph loss value, and L b represents the binary graph loss value, and arctan represents the arctangent function.

[0083] Step S35: Perform Monte Carlo partitioning on the dataset labeled in S31, use the training set to train the improved R(2+1)D model, save the model with the highest validation set accuracy, and test the test set to obtain the dynamic indexes of the echocardiogram video, that is, the diastolic peak velocity V stretch and the systolic peak velocity V shrink .

[0084] Step S4: Use the evaluation function to evaluate the cardiac static indexes and cardiac dynamic indexes to obtain the corresponding scores, and use the linear regression algorithm to calculate the relationship with the types of heart diseases, so as to realize the auxiliary judgment of the types of heart diseases through the echocardiogram video. Specifically, it includes the following steps:

[0085] Step S41: Use the obtained dynamic indexes and static indexes as the training sample indexes, and label the types of heart diseases corresponding to each heart as labels, and use the random forest classifier of traditional machine learning for training.

[0086] Step S42: Obtain the trained parameters, and its final score is Score. It is defined as follows:

[0087]

[0088] where ρ is the weight parameter of the ejection fraction EF, EF is the cardiac ejection fraction, is the weight parameter of the diastolic peak velocity of the heart, V stretch is the diastolic peak velocity of the heart, τ is the weight parameter of the systolic peak velocity of the heart, V shrink is the systolic peak velocity of the heart.

[0089] Step S43: Linearize the obtained scores to obtain the classification relationship between the scores and the types of heart diseases, and through the use of computer vision algorithms, an end-to-end process of judging the types of heart diseases from the input echocardiogram video is realized.

[0090] Based on the above design, in the form of code, a cardiac echocardiogram auxiliary judgment device of the present invention based on semantic segmentation and video understanding is formed in the computer system, including:

[0091] The echocardiogram video semantic segmentation module is used to perform semantic segmentation on the echocardiogram video using an encoder-decoder based on BEIT to obtain a binary heart contour video;

[0092] The heart static index acquisition module is used to analyze the heart contour video obtained by the echocardiogram video semantic segmentation module using a neural network model based on RAFT, and obtain heart static indexes from the maximum contour frame and the minimum contour frame;

[0093] The heart dynamic index acquisition module is used to analyze the echocardiogram video using a video understanding algorithm based on R(2+1)D to obtain heart dynamic indexes;

[0094] The evaluation module is used to evaluate the heart static indexes and heart dynamic indexes using an evaluation function to obtain corresponding scores, and calculate the relationship with the types of heart diseases using a linear regression algorithm.

[0095] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0096] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or multiple processes and / or blocks. Figure 1 One process or multiple processes and / or blocks Figure 1 Steps for realizing the functions specified in one block or multiple blocks.

[0099] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation of the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

[0100] This patent is not limited to the above best implementation manner. Anyone can obtain various other forms of the cardiac color Doppler ultrasound auxiliary judgment device based on semantic segmentation and video understanding under the inspiration of this patent. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage of this patent.

Claims

1. A cardiac ultrasound assisted judgment device based on semantic segmentation and video understanding, characterized in that, A computer system, comprising: A cardiac ultrasound video semantic segmentation module for performing semantic segmentation on a cardiac ultrasound video using an encoder-decoder based on BEIT to obtain a binary cardiac contour video; A cardiac static index acquisition module for analyzing the cardiac contour video obtained by the cardiac ultrasound video semantic segmentation module using a neural network model based on RAFT to obtain cardiac static indices from the maximum contour frame and the minimum contour frame; A cardiac dynamic index acquisition module for analyzing a cardiac ultrasound video using a video understanding algorithm based on R(2+1)D to obtain cardiac dynamic indices; An evaluation module for evaluating the cardiac static indices and cardiac dynamic indices using an evaluation function to obtain corresponding scores and calculating the relationship with the type of heart disease using a linear regression algorithm; The construction and working process of the cardiac ultrasound video semantic segmentation module specifically includes the following steps: Step S11: Conduct training for the cardiac contour, annotate the cardiac ultrasound video using Labelimg to obtain relevant annotations of the video contour training data; Step S12: Use a 12-layer BEIT as the backbone network, add cross-attention and depth convolution kernels to the Block of BEIT to form a BEIT adapter; regularize the downsampling layer, increase the receptive field through spatial pyramid pooling, and replace the third local attention kernel with a global attention kernel so as to be able to capture global features. At the same time, perform a cross-attention mechanism between the fifth local attention kernel and the third global attention kernel, replace the sixth and seventh local attention kernels with two convolutional kernels of size 5×5 and stride 2, and finally connect a 1×1 convolutional kernel for dimensionality reduction to prevent information loss while reducing the number of parameters; Step S13: Use the BEIT adapter in Step S12 to replace the feature extraction backbone of the semantic segmentation model deeplabv3. After the result processed by the SPP structure, extract image features at multiple levels through the Upernet network, and output four feature layers of different sizes (7, 7), (14, 14), (28, 28), (56, 56) through connection with the result extracted by the backbone feature extraction network; Step S14: The regression loss function KIOU obtained by considering the overlap between the label and the prediction and the distance from the boundary is defined as follows: Loss is defined as follows: where IOU is the ratio of the intersection of the predicted region and the target region to the union of the predicted region and the target region, and Distance_2 2 is the Euclidean distance between the center points of the predicted region and the target region, α is the weight coefficient for Distance_2 2 and IOU; β is a weight coefficient for the influencing factor, and k is a parameter for measuring the aspect ratio consistency, which is defined as follows: where π is the pi, tan represents the tangent function, w gt represents the width of the target area, h gt represents the height of the target area, w p represents the width of the prediction area, h p represents the height of the prediction area, where γ is the weight coefficient for adjusting the consistency of aspect ratio; Step S15: Use the semantic segmentation network based on BEIT trained through Steps S11 - S14 to segment the cardiac ultrasound video to obtain a binary cardiac contour video, where the pixels in the cardiac region are white and the pixels in the background region are black; The construction and working process of the cardiac static index acquisition module specifically includes the following steps: Step S21: Use the RAFT network as the network for analyzing the cardiac contour video, use CSPDarknet53 as the feature extraction backbone network, replace the eighth convolutional kernel with a depthwise separable convolutional kernel, set the stride to 2, change the stride of the secondary downsampling feature map from (2, 2) to (1, 1), add an average pooling layer to facilitate the input of features into the RNN, and connect a convolutional kernel to prevent information loss, and finally output the image convolutional feature maps of two frames to achieve feature extraction; Step S22: Calculate the similarity between the feature map vectors of two frames. The similarity is calculated by dot product, and the similarity Similarity is defined as follows: where Similarity is used to calculate the similarity between two vectors. When the value of Similarity is closer to 1, it indicates that the two frame feature maps are closer. i represents the index position in the vector, n represents the length of the vector, x represents the feature map vector of the previous frame, and y represents the feature map vector of the next frame; Step S23: Label the cardiac contour video obtained by the cardiac color Doppler ultrasound video semantic segmentation module through the given end-diastolic frame index and end-systolic frame index, and conduct training; Step S24: Use the trained RAFT to predict the cardiac contour video obtained by the cardiac color Doppler ultrasound video semantic segmentation module to obtain the required end-diastolic frame and end-systolic frame. Use opencv to count the number of white pixels in the diastolic frame image and systolic frame image to obtain the static indicators of the cardiac color Doppler ultrasound, namely end-diastolic volume EDV and end-systolic volume ESV; Step S25: Calculate the ejection fraction EF of the object's heart through the static indicators. The ejection fraction EF is defined as follows: EF = (EDV - ESV) / EDV * 100% where EF is the ejection fraction, EDV is the end-diastolic volume of the heart, and ESV is the end-systolic volume of the heart; The construction and working process of the cardiac dynamic index acquisition module include the following steps: Step S31: Manually label the peak velocity in cardiac systole and the velocity in cardiac diastole for each cardiac color Doppler ultrasound video; Step S32: Use R(2+1)D as the network model for video feature extraction, and replace the 2D convolution therein with separable 3D convolution. The first 7×7 convolution kernel with a stride of 2 is replaced by three 3×3 convolution kernels with a stride of 1, so that the output size of the last feature map remains 7×7. In addition, connect a 1×1 convolution with a stride of 1 to prevent information loss; Use a convolutional layer with a size of 5×5 to replace the average pooling layer to better obtain the probability map and binary map; Step S33: After obtaining the probability map and binary map, calculate the loss between the probability map and the binary map through binary cross-entropy. The loss function is defined as follows: L = arctan(θ × L s + δ × L b ) where θ is the weight parameter of the probability graph loss, and L s is the loss value of the probability graph; δ is the weight function of the binary graph loss value, and L b represents the binary graph loss value, and arctan represents the arctangent function; Step S35: Perform Monte Carlo partitioning on the dataset labeled in S31, use the training set to train the improved R(2+1)D model, save the model with the highest validation set accuracy, make predictions on the test set, and obtain the dynamic indicators of the echocardiogram video, namely the diastolic peak velocity V stretch and the systolic peak velocity V shrink ; The construction and working process of the evaluation module specifically include the following steps: Step S41: Use the dynamic indicators and static indicators obtained by the cardiac dynamic index acquisition module and the cardiac static index acquisition module as training sample indicators, and label the heart disease type corresponding to each heart as a label, and use a random forest classifier for training; Step S42: Obtain the trained parameters, and its final score is Score. It is defined as follows: where ρ is the weight parameter of the ejection fraction EF, and EF is the cardiac ejection fraction, is the weight parameter of the peak velocity in diastole of the heart, V stretch is the peak velocity in diastole of the heart, and τ is the weight parameter of the peak velocity in systole of the heart, V shrink is the peak velocity in systole of the heart; Step S43: Linearly process the obtained score to obtain the division relationship between the score and the heart disease type.