Medical video classification methods, electronic devices and storage media
By combining object detection, curve fitting, segmentation, and classification models, the problem of reliance on human experience in aortic valve leaflet diagnosis has been solved, improving diagnostic efficiency and accuracy, and achieving efficient medical video classification.
Patent Information
- Application Number
- CN202210531192.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-16
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-05-16
AI Technical Summary
In existing technologies, the diagnosis of the number and morphology of aortic valve leaflets mainly relies on the experience analysis of medical personnel or medical experts, which is greatly affected by random factors, resulting in insufficient diagnostic efficiency and accuracy.
A target detection model is used to extract the region of interest (ROI) of the target organ or tissue in the medical video. Curve fitting and correction are performed, and the ROI image of the target organ or tissue is cropped. The segmentation model is used for segmentation, and the classification model is combined with the classification model for identification and classification to obtain the classification result of the medical video.
It improves the accuracy of target organ tissue masking and medical video classification, reduces the variability caused by human factors, assists doctors in improving diagnostic efficiency, reduces the risk of errors, and realizes an efficient end-to-end algorithm process.
Smart Images

Figure CN117115697B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a medical video classification method, electronic device, and storage medium. Background Technology
[0002] Aortic valve disease is a common and dangerous cardiovascular disease that seriously endangers human health. Echocardiography is the preferred method for diagnosing and assessing the severity of valvular heart disease, and the number and morphology of the aortic valve leaflets are important tools for analyzing cardiac function. Currently, the number and morphology of the aortic valve leaflets are mainly determined using obtained echocardiographic images. These images are analyzed by specially trained medical personnel or specialists to determine the number of aortic valve leaflets, thereby identifying whether bicuspid or quadruple valve malformations are present. However, determining the number of aortic valve leaflets through the analysis of cardiac images by specially trained medical personnel or specialists requires a certain level of experience from the physician and is subject to significant random factors.
[0003] It should be noted that the information disclosed in the background section of this invention is intended only to enhance the understanding of the general background of this invention, and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention
[0004] The purpose of this invention is to provide a medical video classification method, electronic device, and storage medium that can effectively identify whether target organ tissues in medical videos are normal, thereby better assisting doctors in improving diagnostic efficiency.
[0005] To achieve the above objectives, the present invention provides a medical video classification method, comprising:
[0006] An object detection model is used to extract the region of interest (ROI) of the target organ or tissue in each frame of the acquired medical video, so as to obtain the location information of the ROI of the target organ or tissue corresponding to each frame of the medical video.
[0007] Curve fitting is performed based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical images, and the location information of the region of interest of the target organ tissue corresponding to each frame of medical images is corrected based on the fitting results.
[0008] Based on the location information of the corrected target organ tissue region of interest corresponding to each frame of medical image, the corresponding target organ tissue region of interest is cropped from each frame of medical image to obtain the corresponding target organ tissue region of interest image;
[0009] A segmentation model is used to segment the region of interest (ROI) image of the target organ tissue corresponding to each frame of medical image in order to obtain the corresponding target organ tissue mask;
[0010] A classification model is used to identify and classify the target organ and tissue masks corresponding to each frame of medical images in order to obtain the classification results of the target organ and tissue masks corresponding to each frame of medical images.
[0011] The classification result of the medical video is obtained based on the classification result of the target organ tissue mask corresponding to each frame of medical image.
[0012] Optionally, the step of performing curve fitting based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical images, and correcting the location information of the region of interest of the target organ tissue corresponding to each frame of medical images based on the fitting result, includes:
[0013] Based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical image extracted by the target detection model, curve fitting is performed to obtain the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue.
[0014] Based on the correspondence between the fitted image frames and the location information of the target organ tissue region of interest, the location information of the target organ tissue region of interest corresponding to each frame of medical image is corrected to obtain the corrected location information of the target organ tissue region of interest corresponding to each frame of medical image.
[0015] Optionally, the step of correcting the location information of the target organ's region of interest corresponding to each frame of the medical image based on the correspondence between the fitted image frames and the location information of the target organ's region of interest, to obtain the corrected location information of the target organ's region of interest corresponding to each frame of the medical image, includes:
[0016] For each frame of the medical image:
[0017] Based on the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue, the location information of the fitted region of interest of the target organ tissue corresponding to the medical image frame is obtained.
[0018] The first positional deviation information corresponding to the medical image is obtained based on the absolute value of the difference between the positional information of the target organ tissue region of interest corresponding to the frame of medical image extracted by the target detection model and the positional information of the fitted target organ tissue region of interest corresponding to the frame of medical image.
[0019] Based on the first positional deviation information corresponding to the medical image frame and the confidence probability value of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model, the second positional deviation information corresponding to the medical image frame is obtained.
[0020] Based on the second positional deviation information corresponding to the medical image frame, it is determined whether the positional information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is accurate;
[0021] If so, the location information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is used as the corrected location information of the target organ tissue region of interest corresponding to the medical image frame.
[0022] If not, then based on the location information of the target organ tissue region of interest corresponding to the medical image in the previous frame with accurate location information and the location information of the target organ tissue region of interest corresponding to the medical image in the next frame with accurate location information, the corrected location information of the target organ tissue region of interest corresponding to the medical image in the current frame is obtained.
[0023] Optionally, the step of obtaining the second positional deviation information corresponding to the medical image frame based on the first positional deviation information corresponding to the frame and the confidence probability value of the target organ tissue region of interest extracted by the target detection model, includes:
[0024] The second positional deviation information corresponding to this frame of medical image is obtained according to the following formula:
[0025] e i =E i *(1-p i )
[0026] In the formula, e i E represents the second positional deviation corresponding to the i-th frame of the medical image. i p represents the first positional deviation corresponding to the i-th frame of the medical image. i This represents the confidence probability value of the region of interest corresponding to the target organ tissue in the i-th frame of the medical image extracted by the target detection model.
[0027] Optionally, determining whether the location information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is accurate based on the second positional deviation information corresponding to the medical image frame includes:
[0028] Based on the first position deviation information corresponding to each frame of medical image, the mean first position deviation information corresponding to the medical video is obtained;
[0029] The average confidence probability value of the target organ tissue region of interest corresponding to each frame of medical image is extracted based on the target detection model to obtain the average confidence probability value corresponding to the medical video.
[0030] Based on the mean first positional deviation information corresponding to the medical video and the mean confidence probability corresponding to the medical video, the mean second positional deviation information corresponding to the medical video is obtained.
[0031] The position judgment threshold is obtained based on the preset multiple threshold and the average value of the second position deviation corresponding to the medical video;
[0032] For each frame of the medical video, based on the second positional deviation information corresponding to that frame of the medical image and the positional judgment threshold, it is determined whether the positional information of the target organ tissue region of interest corresponding to that frame of the medical image extracted by the target detection model is accurate.
[0033] Optionally, the step of performing curve fitting based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical image extracted by the target detection model, to obtain the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue, includes:
[0034] Based on the x-coordinate and y-coordinate information of the first corner point and the second corner point of the target organ tissue region of interest corresponding to each frame of medical image extracted by the target detection model, curve fitting is performed on the x-coordinate, y-coordinate, and y-coordinate of the first corner point of the target organ tissue region of interest, respectively. This is to obtain the correspondence between the fitted image frame and the x-coordinate of the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the second corner point of the target organ tissue region of interest, and the correspondence between the fitted image frame and the second corner point of the target organ tissue region of interest.
[0035] Optionally, obtaining the classification result of the medical video based on the classification result of the target organ tissue mask corresponding to each frame of medical image includes:
[0036] The classification results of the medical video are obtained by using the classification results of the target organ tissue mask corresponding to each frame of medical image and the confidence probability values of the region of interest of the target organ tissue in each frame of medical image extracted by the target detection model.
[0037] Optionally, the step of obtaining the classification result of the medical video based on the classification result of the target organ / tissue mask corresponding to each frame of medical images and the confidence probability value of the region of interest of the target organ / tissue in each frame of medical images extracted by the target detection model includes:
[0038] Based on the classification results of the target organ tissue mask corresponding to each frame of medical image, the probability value of the target organ tissue mask corresponding to each frame of medical image being judged as normal target organ tissue is obtained;
[0039] For each frame of a medical image, based on the probability value that the target organ tissue mask corresponding to that frame of the medical image is determined to be normal, and the confidence probability value of the region of interest of the target organ tissue extracted by the target detection model for that frame of the medical image, the first probability value corresponding to that frame of the medical image is calculated according to the following formula:
[0040] P segancls1i =P 1i *p i
[0041] In the formula, P segandcls1i Let P represent the first probability value corresponding to the i-th frame of the medical image. 1i p represents the probability value that the target organ tissue mask corresponding to the i-th frame of the medical image is judged as normal. i This represents the confidence probability value of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model;
[0042] Based on the first probability value corresponding to each frame of medical image, obtain the first probability mean corresponding to the medical video;
[0043] If the mean of the first probability is greater than a first preset threshold, the classification result of the medical video is determined to be that the target organ tissue is normal; otherwise, the classification result of the medical video is determined to be that the target organ tissue is abnormal.
[0044] Optionally, the step of obtaining the classification result of the medical video based on the classification result of the target organ / tissue mask corresponding to each frame of medical images and the confidence probability value of the region of interest of the target organ / tissue in each frame of medical images extracted by the target detection model includes:
[0045] Based on the classification results of the target organ tissue mask corresponding to each frame of medical image, the probability value of the target organ tissue mask corresponding to each frame of medical image being judged as an abnormality of the target organ tissue is obtained.
[0046] For each frame of a medical image, based on the probability value that the target organ tissue mask corresponding to that frame of the medical image is determined to be an abnormality of the target organ tissue and the confidence probability value of the region of interest of the target organ tissue extracted by the target detection model in that frame of the medical image, the second probability value corresponding to that frame of the medical image is calculated according to the following formula:
[0047] P segancls2 =P 2i *p i
[0048] In the formula, P segandcls2i P represents the second probability value corresponding to the i-th frame of the medical image. 2i p represents the probability value that the target organ tissue mask corresponding to the i-th frame of the medical image is identified as an abnormality of the target organ tissue. i This represents the confidence probability value of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model;
[0049] Based on the second probability value corresponding to each frame of medical image, obtain the second probability mean corresponding to the medical video;
[0050] If the mean of the second probability is greater than a second preset threshold, the classification result of the medical video is determined to be an abnormality of the target organ tissue; otherwise, the classification result of the medical video is determined to be a normality of the target organ tissue.
[0051] Optionally, before segmenting the region of interest (ROI) images of the target organ tissue corresponding to each frame of the medical image using a segmentation model, the method further includes:
[0052] For each frame of the medical image:
[0053] The length dimension of the region of interest corresponding to the target organ tissue in this frame of medical image is taken as the target side length.
[0054] The region of interest image of the target organ tissue is filled along the width direction to adjust the width dimension of the region of interest image of the target organ tissue to the target side length dimension;
[0055] The region of interest image of the target organ tissue is magnified or reduced by adjusting the width dimension to the target side length dimension, so as to adjust the size of the region of interest image of the target organ tissue to a preset size.
[0056] To achieve the above objectives, the present invention also provides an electronic device, including a processor and a memory, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, it implements the medical video classification method described above.
[0057] To achieve the above objectives, the present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, implements the medical video classification method described above.
[0058] Compared with existing technologies, the medical video classification method, electronic device, and storage medium provided by this invention have the following advantages: First, this invention employs a target detection model to extract the region of interest (ROI) of the target organ / tissue in each frame of the acquired medical video, obtaining the location information of the ROI corresponding to each frame. Then, curve fitting is performed based on the location information of the ROI of the target organ / tissue in each frame, and the location information of the ROI of the target organ / tissue in each frame is corrected based on the fitting result. Next, based on the corrected location information of the ROI of the target organ / tissue in each frame, the corresponding ROI of the target organ / tissue is cropped from each frame to obtain the corresponding ROI image. Then, a segmentation model is used to segment the ROI image of each frame to obtain the corresponding target organ / tissue mask. Then, a classification model is used to identify and classify the target organ / tissue mask corresponding to each frame to obtain the classification result of the target organ / tissue mask corresponding to each frame. Finally, based on the classification result of the target organ / tissue mask corresponding to each frame, the classification result of the medical video is obtained. Therefore, this invention corrects the location information of the target organ / tissue region of interest (ROI) extracted from each frame of medical images by the target detection model, and then crops the corresponding ROI image from the medical image based on the corrected ROI location information. This allows for the acquisition of more accurate ROI images, effectively improving the accuracy of the acquired target organ / tissue mask and laying a solid foundation for obtaining accurate classification results. Furthermore, since the medical video classification method provided by this invention comprehensively considers the classification results of the target organ / tissue mask corresponding to each frame of the medical video, it effectively improves the accuracy of medical video classification (i.e., accurately identifying whether the target organ / tissue in the medical video is normal or abnormal). This reduces potential discrepancies caused by human factors, better assisting doctors in improving diagnostic efficiency and reducing the risk of errors in organ / tissue anomaly analysis using medical videos. In addition, this invention enables an end-to-end algorithm flow, has strong versatility, and effectively improves the efficiency of medical video classification. Attached Figure Description
[0059] Figure 1 A flowchart illustrating a medical video classification method provided in one embodiment of the present invention;
[0060] Figure 2a A heartbeat image provided as a specific example of the present invention;
[0061] Figure 2b From Figure 2a A region of interest image of the aortic valve leaflets cropped from a cardiac image;
[0062] Figure 2c To Figure 2b Image of the region of interest of the aortic valve leaflets after filling;
[0063] Figure 3 This is a schematic diagram of the structure of a segmentation model provided as a specific example of the present invention;
[0064] Figure 4 A schematic diagram of the bottleneck layer provided as a specific example of the present invention;
[0065] Figure 5 This is a schematic diagram of the structure of a transition block provided in a specific example of the present invention;
[0066] Figure 6 This is a schematic diagram of the structure of an upward transition block provided in a specific example of the present invention;
[0067] Figure 7a This is a specific example of the invention, showing the region of interest of the aortic valve leaflets during ventricular diastole;
[0068] Figure 7b To Figure 7a The aortic valve leaflet image obtained by segmentation;
[0069] Figure 7c This is a specific example of the invention, showing the region of interest (ROI) image of the aortic valve leaflets during ventricular systole.
[0070] Figure 7d To Figure 7c The aortic valve leaflet image obtained by segmentation;
[0071] Figure 8 A schematic diagram showing the region of interest and aortic valve leaflets in an echocardiogram provided as a specific example of the present invention;
[0072] Figure 9 This is a block diagram of an electronic device according to one embodiment of the present invention.
[0073] The reference numerals in the attached figures are as follows:
[0074] Processor-101; Communication interface-102; Memory-103; Communication bus-104. Detailed Implementation
[0075] The medical video classification method, electronic device, and storage medium proposed in this invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The advantages and features of this invention will become clearer from the following description. It should be noted that the accompanying drawings are in a very simplified form and use non-precise proportions, used only to facilitate and clearly illustrate the embodiments of this invention. Please refer to the accompanying drawings to make the objectives, features, and advantages of this invention more apparent and understandable. It should be understood that the structures, proportions, sizes, etc., depicted in the accompanying drawings are only used to complement the content disclosed in the specification, for those skilled in the art to understand and read, and are not intended to limit the implementation conditions of this invention. Any modifications to the structure, changes in proportions, or adjustments to the size, provided that the effects and objectives achieved by this invention are the same or similar, should still fall within the scope of the technical content disclosed in this invention.
[0076] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0077] Furthermore, in the description of this specification, the reference to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., means that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0078] The core idea of this invention is to provide a medical video classification method, electronic device, and storage medium that can effectively identify whether the target organ tissue in the medical video is normal, thereby better assisting doctors in improving diagnostic efficiency.
[0079] It should be noted that the medical video classification method of this invention can be applied to the electronic devices described in this invention. These electronic devices can be personal computers, mobile terminals, etc., and the mobile terminals can be hardware devices with various operating systems, such as mobile phones and tablets. Furthermore, although this document uses echocardiography video as an example, as those skilled in the art will understand, the medical video can also be cardiac video acquired by devices other than ultrasound equipment (e.g., cardiac endoscopes); of course, the medical video can also be video of other organs besides cardiac video, and this invention does not limit this. Additionally, although this document uses aortic valve leaflets as the target organ tissue in the description, as those skilled in the art will understand, the target organ tissue can also be mitral valve leaflets, tricuspid valve leaflets, pulmonary valve leaflets, etc., and of course, the target organ tissue can also be other organ tissues besides heart valve leaflets, and this invention does not limit this. It should also be noted that in this document, the long side direction of the image is defined as the length direction, and the short side direction of the image is defined as the width direction.
[0080] To achieve the above-mentioned goals, this invention provides a medical video classification method, please refer to [the following text is missing]. Figure 1 The diagram illustrates a flowchart of a medical video classification method provided by an embodiment of the present invention. Figure 1 As shown, the medical video classification method includes the following steps:
[0081] Step S100: Use a target detection model to extract the region of interest of the target organ tissue in each frame of the acquired medical video, so as to obtain the location information of the region of interest of the target organ tissue corresponding to each frame of the medical video.
[0082] Step S200: Perform curve fitting based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical images, and correct the location information of the region of interest of the target organ tissue corresponding to each frame of medical images based on the fitting results.
[0083] Step S300: Based on the location information of the corrected target organ tissue region of interest corresponding to each frame of medical images, the corresponding target organ tissue region of interest is cropped from each frame of medical images to obtain the corresponding target organ tissue region of interest image.
[0084] Step S400: Use a segmentation model to segment the region of interest image of the target organ tissue corresponding to each frame of medical image to obtain the corresponding target organ tissue mask.
[0085] Step S500: Use a classification model to identify and classify the target organ tissue mask corresponding to each frame of medical images, so as to obtain the classification result of the target organ tissue mask corresponding to each frame of medical images.
[0086] Step S600: Obtain the classification result of the medical video based on the classification result of the target organ tissue mask corresponding to each frame of medical image.
[0087] Therefore, this invention corrects the location information of the target organ / tissue region of interest (ROI) extracted from each frame of medical images by the target detection model, and then crops the corresponding ROI image from the medical image based on the corrected ROI location information. This allows for the acquisition of more accurate ROI images, effectively improving the accuracy of the acquired target organ / tissue mask and laying a solid foundation for obtaining accurate classification results. Furthermore, since the medical video classification method provided by this invention comprehensively considers the classification results of the target organ / tissue mask corresponding to each frame of the medical video, it effectively improves the accuracy of medical video classification (i.e., accurately identifying whether the target organ / tissue in the medical video is normal or abnormal). This reduces potential discrepancies caused by human factors, better assisting doctors in improving diagnostic efficiency and reducing the risk of errors in organ / tissue anomaly analysis using medical videos. In addition, this invention enables an end-to-end algorithm flow, has strong versatility, and effectively improves the efficiency of medical video classification.
[0088] As an example, the medical video is an echocardiogram (each video contains multiple cardiac cycles). The resolution of the echocardiogram can be set according to specific circumstances, such as 600×800. Specifically, the echocardiogram is a PSAX-AV cross-sectional image acquired by an ultrasound device. Therefore, by first using a pre-trained target detection model to detect the region of interest (ROI) of each frame of the acquired echocardiogram video, the location information of the corresponding target heart valve leaflet (e.g., aortic valve leaflet) ROI can be obtained. Then, the obtained ROI location information is fitted and analyzed, and the location of the target heart valve leaflet (e.g., aortic valve leaflet) ROI obtained by the target detection model is corrected based on the fitting results. This corrects errors that are too large or incorrectly extracted target detection boxes (i.e., the ROI of the target heart valve leaflet (e.g., aortic valve leaflet)). Finally, based on the corrected location information of the target heart valve leaflet (e.g., aortic valve leaflet) ROI, a more accurate ROI of the target heart valve leaflet (e.g., aortic valve leaflet) can be obtained. The process involves: first, segmenting the target heart valve leaflets (e.g., aortic valve leaflets) in the region of interest (ROI) images of each frame of the acquired echocardiogram video using a pre-trained segmentation model; then, segmenting the target heart valve leaflets (e.g., aortic valve leaflets) using a pre-trained classification model; finally, classifying the target heart valve leaflets (e.g., aortic valve leaflets) in each frame of the echocardiogram video using a pre-trained classification model to determine whether the target heart valve leaflets (e.g., aortic valve leaflets) in each frame of the echocardiogram video are normal or not; and finally, considering the classification results of the target heart valve leaflets (e.g., aortic valve leaflets) in each frame of the echocardiogram video, the classification result of the echocardiogram video is obtained.
[0089] Therefore, this invention obtains the region of interest (ROI) image of the target organ / tissue by first cropping it from a medical image, and then segments the ROI image using a segmentation model. This further reduces the computational load of the segmentation model and thus improves computational efficiency. For details, please refer to... Figure 2a and Figure 2b ,in Figure 2a The cardiac image provided by a specific example of the present invention is illustrated. Figure 2b The illustration shows from Figure 2a The region of interest image of the aortic valve leaflets cropped from a medical image. For example... Figure 2a and Figure 2bAs shown, by using a target detection model to detect each frame of medical images (e.g., cardiac images), the location of the target organ tissue (e.g., aortic valve leaflets) in the region of interest can be accurately determined.
[0090] In one exemplary embodiment, the step of performing curve fitting based on the location information of the region of interest (ROI) of the target organ tissue corresponding to each frame of medical images, and correcting the location information of the ROI of the target organ tissue corresponding to each frame of medical images based on the fitting result, includes:
[0091] Based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical image extracted by the target detection model, curve fitting is performed to obtain the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue.
[0092] Based on the correspondence between the fitted image frames and the location information of the target organ tissue region of interest, the location information of the target organ tissue region of interest corresponding to each frame of medical image is corrected to obtain the corrected location information of the target organ tissue region of interest corresponding to each frame of medical image.
[0093] Therefore, by performing curve fitting on the location information of the target organ tissue region of interest corresponding to each frame of medical image extracted by the target detection model, the correspondence between the fitted image frame and the location information of the target organ tissue region of interest can be obtained. Thus, based on the correspondence between the fitted image frame and the location information of the target organ tissue region of interest, the location information of the fitted target organ tissue region of interest corresponding to each frame of medical image can be obtained.
[0094] Further, the step of performing curve fitting based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical image extracted by the target detection model, to obtain the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue, includes:
[0095] Based on the x-coordinate and y-coordinate information of the first corner point and the second corner point of the target organ tissue region of interest corresponding to each frame of medical image extracted by the target detection model, curve fitting is performed on the x-coordinate, y-coordinate, and y-coordinate of the first corner point of the target organ tissue region of interest, respectively. This is to obtain the correspondence between the fitted image frame and the x-coordinate of the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the second corner point of the target organ tissue region of interest, and the correspondence between the fitted image frame and the second corner point of the target organ tissue region of interest.
[0096] Specifically, the first corner point can be the upper left corner of the region of interest of the target organ tissue extracted by the target detection model in each frame of the medical image, and the second corner point can be the lower right corner of the region of interest of the target organ tissue extracted by the target detection model in each frame of the medical image. Thus, the position information of the first corner point and the position information of the second corner point can represent the position information of the region of interest of the target organ tissue. By fitting the abscissa of the first corner point of the region of interest (ROI) of the target organ / tissue corresponding to each frame of medical images extracted by the target detection model, a curve representing the correspondence between the fitted image frame and the first corner point of the ROI can be obtained. Similarly, by fitting the ordinate of the ordinate of the first corner point of the ROI of the target organ / tissue corresponding to each frame of medical images extracted by the target detection model, a curve representing the correspondence between the fitted image frame and the first corner point of the ROI can be obtained. Likewise, by fitting the abscissa of the second corner point of the ROI of the target organ / tissue corresponding to each frame of medical images extracted by the target detection model, a curve representing the correspondence between the fitted image frame and the second corner point of the ROI can be obtained. Finally, by fitting the ordinate of the second corner point of the ROI of the target organ / tissue corresponding to each frame of medical images extracted by the target detection model, a curve representing the correspondence between the fitted image frame and the second corner point of the ROI can be obtained. Therefore, based on these four fitted curves, we can obtain the x-coordinate and y-coordinate information of the first corner point and the second corner point of the fitted target organ tissue region of interest corresponding to any frame of medical image. In other words, we can obtain the position information of the fitted target organ tissue region of interest corresponding to any frame of medical image.
[0097] In one exemplary embodiment, the step of correcting the location information of the target organ / tissue region of interest corresponding to each frame of the medical image based on the correspondence between the fitted image frames and the location information of the target organ / tissue region of interest, to obtain the corrected location information of the target organ / tissue region of interest corresponding to each frame of the medical image, includes:
[0098] For each frame of the medical image:
[0099] Based on the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue, the location information of the fitted region of interest of the target organ tissue corresponding to the medical image frame is obtained.
[0100] The first positional deviation information corresponding to the medical image is obtained based on the absolute value of the difference between the positional information of the target organ tissue region of interest corresponding to the frame of medical image extracted by the target detection model and the positional information of the fitted target organ tissue region of interest corresponding to the frame of medical image.
[0101] Based on the first positional deviation information corresponding to the medical image frame and the confidence probability value of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model, the second positional deviation information corresponding to the medical image frame is obtained.
[0102] Based on the second positional deviation information corresponding to the medical image frame, it is determined whether the positional information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is accurate;
[0103] If so, the location information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is used as the corrected location information of the target organ tissue region of interest corresponding to the medical image frame.
[0104] If not, then based on the location information of the target organ tissue region of interest corresponding to the medical image in the previous frame with accurate location information and the location information of the target organ tissue region of interest corresponding to the medical image in the next frame with accurate location information, the corrected location information of the target organ tissue region of interest corresponding to the medical image in the current frame is obtained.
[0105] Therefore, for each frame of a medical image, based on the location information of the target organ / tissue region of interest extracted by the target detection model and the fitted location information of the target organ / tissue region of interest, the first location deviation information corresponding to that frame of the medical image is obtained. Then, based on the first location deviation information and the confidence probability value of the target organ / tissue region of interest extracted by the target detection model, the second location deviation information corresponding to that frame of the medical image is obtained. Furthermore, based on the second location deviation information, it is determined whether the location information of the target organ / tissue region of interest extracted by the target detection model for that frame of the medical image is accurate. This allows for a more accurate determination of the accuracy of the location information of the target organ / tissue region of interest extracted by the target detection model for that frame of the medical image, effectively avoiding misjudgments and thus effectively improving the correction effect of the location information of the target organ / tissue region of interest. This further lays a good foundation for obtaining accurate medical video classification results.
[0106] Furthermore, the step of obtaining the second positional deviation information corresponding to the medical image frame based on the first positional deviation information corresponding to the frame and the confidence probability value of the target organ tissue region of interest extracted by the target detection model, includes:
[0107] The second positional deviation information corresponding to this frame of medical image is obtained according to the following formula:
[0108] e i =E i *(1-p i )
[0109] In the formula, e i E represents the second positional deviation corresponding to the i-th frame of the medical image. i p represents the first positional deviation corresponding to the i-th frame of the medical image. i This represents the confidence probability value of the region of interest corresponding to the target organ tissue in the i-th frame of the medical image extracted by the target detection model.
[0110] Specifically, taking the i-th frame of a medical image as an example, assuming that the location information of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model is (w 1i ,h 1i ,w 2i ,h 2i ), where w 1i h 1i w 2i h 2iThese represent the x-coordinate, y-coordinate, and y-coordinate of the first corner point, the second corner point, and the third corner point, respectively, of the region of interest (ROI) corresponding to the target organ tissue in the i-th frame of the medical image extracted by the target detection model. The location information of the fitted ROI corresponding to the i-th frame of the medical image, obtained based on the fitting results, is (w'...). 1i ,h' 1i ,w' 2i ,h' 2i ), where w' 1i h' 1i w' 2i h' 2i Let |w| represent the x-coordinate of the first corner point, the y-coordinate of the first corner point, the x-coordinate of the second corner point, and the y-coordinate of the second corner point, respectively, of the fitted region of interest (ROI) corresponding to the i-th frame of the medical image. Then, the absolute value of the difference between the x-coordinate of the first corner point of the ROI extracted by the target detection model and the x-coordinate of the first corner point of the fitted ROI is |w|. 1i -w' 1i The absolute value of the difference between the ordinate of the first corner point of the target organ tissue region of interest extracted by the target detection model and the ordinate of the first corner point of the fitted target organ tissue region of interest is |h 1i -h' 1i The absolute value of the difference between the x-coordinate of the second corner point of the target organ tissue region of interest extracted by the target detection model and the x-coordinate of the second corner point of the fitted target organ tissue region of interest is |w 2i -w' 2i The absolute value of the difference between the ordinate of the second corner point of the target organ tissue region of interest extracted by the target detection model and the ordinate of the second corner point of the fitted target organ tissue region of interest is |h 2i -h' 2i |;that is, the first positional deviation E corresponding to the i-th frame of the medical image. i for:
[0111] E i =(|w 1i -w' 1i |,|h 1i -h' 1i |,|w 2i -w' 2i |,|h 2i -h' 2i |)
[0112] Then the second positional deviation e corresponding to the i-th frame of the medical image i for:
[0113] e i =(|w 1i -w' 1i |*(1-p i ),|h 1i -h' 1i |*(1-p i ),|w 2i -w' 2i |*(1-p i ),|h 2i -h' 2i |*(1-p i ))
[0114] In one exemplary embodiment, determining whether the location information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is accurate based on the second location deviation information corresponding to the medical image frame includes:
[0115] Based on the first position deviation information corresponding to each frame of medical image, the mean first position deviation information corresponding to the medical video is obtained;
[0116] The average confidence probability value of the target organ tissue region of interest corresponding to each frame of medical image is extracted based on the target detection model to obtain the average confidence probability value corresponding to the medical video.
[0117] Based on the mean first positional deviation information corresponding to the medical video and the mean confidence probability corresponding to the medical video, the mean second positional deviation information corresponding to the medical video is obtained.
[0118] The position judgment threshold is obtained based on the preset multiple threshold and the average value of the second position deviation corresponding to the medical video;
[0119] For each frame of the medical video, based on the second positional deviation information corresponding to that frame of the medical image and the positional judgment threshold, it is determined whether the positional information of the target organ tissue region of interest corresponding to that frame of the medical image extracted by the target detection model is accurate.
[0120] Specifically, the average first positional deviation of the medical video is obtained by averaging the first positional deviations corresponding to each frame of the medical image; the average confidence probability of the medical video is obtained by averaging the confidence probability values corresponding to each frame of the medical image; the product of the average first positional deviation and the average confidence probability of the medical video is the average second positional deviation of the medical video; the product of the average second positional deviation and a preset multiple threshold is the position judgment threshold. Since the position judgment threshold is obtained based on the average second positional deviation of the medical video and the preset multiple threshold, the position judgment threshold is different for different medical videos. That is, the position judgment threshold in this invention is dynamically changing, which can further improve the correction effect of the positional information of the region of interest of the target organ tissue, laying a good foundation for obtaining accurate positional information of the region of interest of the target organ tissue.
[0121] Assuming the medical video comprises n frames of medical images, then the average first positional deviation corresponding to the medical video is... It can be represented as:
[0122]
[0123] The mean confidence probability corresponding to the medical video It can be represented as:
[0124]
[0125] The mean of the second position deviation corresponding to the medical video It can be represented as:
[0126]
[0127] Assuming the preset magnification factor is m, then the position determination threshold T corresponding to the medical video is... th It can be represented as:
[0128]
[0129] It should be noted that, as those skilled in the art will understand, during location determination, the abscissa, ordinate, abscissa, and ordinate of the first corner point, the second corner point, and the ordinate of the second corner point of the region of interest of the target organ tissue are determined respectively, and the abscissa, ordinate, second corner point, and ordinate of the second corner point of the region of interest of the target organ tissue are corrected accordingly based on the corresponding determination results. Specifically, taking the i-th frame of the medical image as an example, if the abscissa of the first corner point of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image satisfies the following condition: |w1i -w' 1i |*(1-p i (greater than) This indicates that the x-coordinate error of the first corner point of the target organ tissue region of interest extracted by the target detection model in the i-th frame of the medical image is large. Therefore, the average of the x-coordinates of the first corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the x-coordinates of the first corner point are accurate, i.e., the x-coordinates of the first corner point of the region of interest extracted by the target detection model are accurate) and the x-coordinates of the first corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the x-coordinates of the first corner point are accurate, i.e., the x-coordinates of the first corner point of the region of interest extracted by the target detection model are accurate) is used as the corrected x-coordinate of the first corner point of the target organ tissue region of interest in the i-th frame of the medical image; if |w 1i -w' 1i |*(1-p i Less than or equal to This indicates that the x-coordinate of the first corner point of the region of interest of the target organ tissue in the i-th frame of the medical image extracted by the target detection model is accurate. Therefore, the x-coordinate of the first corner point of the region of interest of the target organ tissue in the i-th frame of the medical image extracted by the target detection model is directly used as the x-coordinate of the first corner point of the corrected region of interest of the target organ tissue in the i-th frame of the medical image.
[0130] If the ordinate of the first corner point of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image satisfies the following condition: |h 1i -h' 1i |*(1-p i (greater than) This indicates that the error in the ordinate of the first corner point of the target organ tissue region of interest extracted by the target detection model in the i-th frame of the medical image is large. Therefore, the average of the ordinates of the first corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the ordinates of the first corner point are accurate, i.e., the ordinates of the first corner point of the region of interest extracted by the target detection model are accurate) and the ordinates of the first corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the ordinates of the first corner point are accurate, i.e., the ordinates of the first corner point of the region of interest extracted by the target detection model are accurate) is taken as the corrected ordinate of the first corner point of the target organ tissue region of interest in the i-th frame of the medical image; if |h 1i -h'1i |*(1-p i Less than or equal to This indicates that the ordinate of the first corner point of the region of interest of the target organ tissue extracted by the target detection model in the i-th frame of the medical image is accurate. Therefore, the ordinate of the first corner point of the region of interest of the target organ tissue in the i-th frame of the medical image extracted by the target detection model is directly used as the ordinate of the first corner point of the corrected region of interest of the target organ tissue in the i-th frame of the medical image.
[0131] If the x-coordinate of the second corner point of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image satisfies the following condition: |w 2i -w' 2i |*(1-p i (greater than) This indicates that the x-coordinate error of the second corner point of the target organ tissue region of interest extracted by the target detection model in the i-th frame of the medical image is large. Therefore, the average of the x-coordinates of the second corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the x-coordinates of the second corner point are accurate, i.e., the x-coordinates of the second corner point of the region of interest extracted by the target detection model are accurate) and the average of the x-coordinates of the second corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the x-coordinates of the second corner point are accurate, i.e., the x-coordinates of the second corner point of the region of interest extracted by the target detection model are accurate) is used as the corrected x-coordinate of the second corner point of the target organ tissue region of interest in the i-th frame of the medical image; if |w 2i -w' 2i |*(1-p i Less than or equal to This indicates that the x-coordinate of the second corner point of the region of interest of the target organ tissue in the i-th frame of the medical image extracted by the target detection model is accurate. Therefore, the x-coordinate of the second corner point of the region of interest of the target organ tissue in the i-th frame of the medical image extracted by the target detection model is directly used as the x-coordinate of the second corner point of the corrected region of interest of the target organ tissue in the i-th frame of the medical image.
[0132] If the ordinate of the second corner point of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image satisfies the following condition: |h 2i -h' 2i |*(1-p i (greater than) This indicates that the error in the ordinate of the second corner point of the target organ tissue region of interest extracted by the target detection model in the i-th frame of the medical image is large. Therefore, the average of the ordinates of the second corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the ordinates of the second corner point are accurate, i.e., the ordinates of the second corner point of the region of interest extracted by the target detection model are accurate) and the ordinates of the second corner point of the target organ tissue region of interest in the medical image adjacent to the i-th frame (where the ordinates of the second corner point are accurate, i.e., the ordinates of the second corner point of the region of interest extracted by the target detection model are accurate) is used as the corrected ordinate of the second corner point of the target organ tissue region of interest in the i-th frame of the medical image; if |h 2i -h' 2i |*(1-p i Less than or equal to This indicates that the ordinate of the second corner point of the region of interest of the target organ tissue extracted by the target detection model in the i-th frame of the medical image is accurate. Therefore, the ordinate of the second corner point of the region of interest of the target organ tissue in the i-th frame of the medical image extracted by the target detection model is directly used as the ordinate of the second corner point of the corrected region of interest of the target organ tissue in the i-th frame of the medical image.
[0133] In one exemplary implementation, the object detection model is a ResNet50 neural network model. Because ResNet50 uses skip connections (or shortcuts), it directly transmits the activation values of one network layer to deeper layers. Furthermore, skip connections only transmit data; through skip connections, the signal can be transmitted without attenuation during backpropagation, without worrying about gradient changes, thus enabling the transmission of effective gradients to the next layer. Therefore, skip connections effectively alleviate the gradient vanishing problem caused by deepening network layers. By stacking residual blocks, very deep network models can be constructed, allowing for effective training even at deep network layers.
[0134] Further, the step of cropping the corresponding target organ / tissue region of interest from each frame of medical images based on the location information of the corrected region of interest for each frame of medical images to obtain the corresponding target organ / tissue region of interest image includes:
[0135] Based on the location information of the corrected region of interest of the target organ tissue corresponding to each frame of medical image, calculate the location information of the region of interest of the target organ tissue after magnification by a preset factor for each frame of medical image;
[0136] The location information of the region of interest of the target organ tissue after being magnified by a preset factor is used as the location information of the region of interest of the target organ tissue corresponding to the medical image frame.
[0137] While object detection models can identify regions of interest (ROIs) in medical images, providing initial localization for subsequent segmentation models, they also result in the loss of detailed information such as surrounding tissues. Therefore, this invention calculates the location of the ROI after magnification based on the corrected ROI location information for each frame of the medical image. Specifically, the bounding box of the corrected ROI is magnified by a predetermined factor, such as 1.3 times, to obtain the magnified bounding box. The area defined by this magnified bounding box is the final ROI. Since this magnified bounding box includes detailed information such as surrounding tissues, it further improves the segmentation accuracy of subsequent segmentation models. It should be noted that, as those skilled in the art will understand, the center position of the magnified bounding box is the same as the center position of the unmagnified bounding box.
[0138] In one exemplary embodiment, before segmenting the region of interest image of the target organ tissue corresponding to each frame of medical image using a segmentation model, the method further includes:
[0139] For each frame of the medical image:
[0140] The length dimension of the region of interest corresponding to the target organ tissue in this frame of medical image is taken as the target side length.
[0141] The region of interest image of the target organ tissue is filled along the width direction to adjust the width dimension of the region of interest image of the target organ tissue to the target side length dimension;
[0142] The region of interest image of the target organ tissue is magnified or reduced by adjusting the width dimension to the target side length dimension, so as to adjust the size of the region of interest image of the target organ tissue to a preset size.
[0143] When the segmentation model is a neural network model, since neural network models require images of a uniform size as input, adjusting the size of the region of interest image of the target organ tissue to a preset size can meet the input requirements of the segmentation model. Specifically, the preset size can be set according to specific circumstances. As a preferred embodiment, in the preset size, the length and width dimensions of the image are consistent, that is, the image after adjustment to the preset size is a square image, for example, the preset size is 320*320. Therefore, by setting the length and width dimensions in the preset size to be consistent, it is easier to adjust the size of the region of interest image of the target organ tissue to the preset size.
[0144] For details, please refer to Figure 2b and Figure 2c ,in Figure 2c The illustration shows the... Figure 2b Image of the region of interest of the aortic valve leaflets after filling. (e.g.) Figure 2b and Figure 2c As shown, the region of interest image of the target organ tissue (e.g., the region of interest image of the aortic valve leaflet) can be filled with black pixels (pixel value of 0) along the width direction to adjust the width dimension of the region of interest image of the target organ tissue (e.g., the region of interest image of the aortic valve leaflet) to be consistent with the length dimension, that is, to adjust the region of interest image of the target organ tissue (e.g., the region of interest image of the aortic valve leaflet) into a square image. Then, the squared region of interest image of the target organ tissue (e.g., the region of interest image of the aortic valve leaflet) can be enlarged or reduced by a certain factor to adjust the size of the region of interest image of the target organ tissue (e.g., the region of interest image of the aortic valve leaflet) to a preset size.
[0145] In one exemplary embodiment, the segmentation model employs a DenseNet neural network structure as its backbone network (for clarity, the backbone network of the segmentation model is referred to as the first backbone network, and the backbone network of the classification model as the second backbone network). Since the DenseNet neural network is a densely connected convolutional neural network, with each layer's input derived from the outputs of all preceding layers, this neural network structure enhances feature transfer and utilizes features more effectively. Furthermore, the DenseNet neural network model exhibits good resistance to overfitting, making it particularly suitable for applications with relatively scarce training data. Therefore, using the DenseNet neural network model as the segmentation model in this invention can effectively improve the segmentation efficiency and accuracy of target organ tissue (e.g., aortic valve leaflets) images. Specifically, the DenseNet neural network model consists of multiple densely connected blocks connected by transition blocks; that is, any two adjacent densely connected blocks are connected by a transition block, and the number of convolutional output channels within each densely connected block is consistent, facilitating the superposition of feature information from each layer.
[0146] One layer in a densely connected block is called a bottleneck layer. Dense connections in DenseNet connect each layer in a densely connected block to all subsequent layers, enabling feature reuse.
[0147] Please continue to refer to this. Figure 3 The diagram illustrates the structure of a segmentation model provided in a specific example of the present invention. Figure 3As shown in this example, the segmentation model includes a first backbone network and a segmentation network connected together. The first backbone network includes a first convolutional layer, a first pooling layer (preferably a max pooling layer), a first dense connection block, a first transition block, a second dense connection block, a second transition block, a third dense connection block, a third transition block, and a fourth dense connection block connected in sequence. The segmentation network includes a first upward transition block, a second upward transition block, and a second convolutional layer (with a kernel size of 1×1) connected in sequence. In this configuration, the first convolutional layer extracts features of the target organ tissue (e.g., aortic valve leaflets) from the input image; the first pooling layer performs pooling operations on the output of the first convolutional layer to remove unnecessary redundant information from the image; the first dense connection block extracts features of the target organ tissue (e.g., aortic valve leaflets) from the output of the first pooling layer; the first transition block compresses the output of the first dense connection block to reduce the size of the feature map output by the first dense connection block; the second dense connection block extracts features of the target organ tissue (e.g., aortic valve leaflets) from the output of the first transition block; the second transition block compresses the output of the second dense connection block to reduce the size of the feature map output by the second dense connection block; and the third dense connection block extracts features of the target organ tissue (e.g., aortic valve leaflets) from the output of the first transition block; the second transition block compresses the output of the second dense connection block to reduce the size of the feature map output by the second dense connection block; and the third dense connection block is used to extract features of the target organ tissue (e.g., aortic valve leaflets) from the input image. The output of the second transition block is used to extract features of the target organ tissue (e.g., aortic valve leaflets). The third transition block is used to compress the output of the third dense connection block to reduce the size of the feature map output by the third dense connection block. The fourth dense connection block is used to extract features of the target organ tissue (e.g., aortic valve leaflets) from the output of the third transition block. The first upward transition block is used to deconvolve the output of the fourth dense connection block to increase the size of the feature map output by the fourth dense connection block. The second upward transition block is used to deconvolve the output of the first upward transition block to increase the size of the feature map output by the first upward transition block. The second convolutional layer is used to perform nonlinear mapping regression on the output of the second upward transition block to obtain the segmentation result of the target organ tissue (e.g., aortic valve leaflets).
[0148] Specifically, the second convolutional layer A and the second convolutional layer B can perform nonlinear mapping regression on the output of the second upward transition block using the sigmoid function. The formula for the sigmoid function is as follows:
[0149]
[0150] As shown in the above equation, the Sigmoid function can map any input real number to the real number mapping interval (0,1). When the input value x is large, the output value g tends to 1, and when the input value x is small, the output value g tends to 0.
[0151] It should be noted that, as those skilled in the art will understand, the first dense connection block, the second dense connection block, the third dense connection block, and the fourth dense connection block all include multiple bottleneck layers, and the number of bottleneck layers in the first dense connection block, the second dense connection block, the third dense connection block, and the fourth dense connection block can be the same or different. The specific number can be set according to actual needs, and the present invention does not limit this. For example, the first dense connection block may have 6 bottleneck layers, the second dense connection block may have 12 bottleneck layers, the third dense connection block may have 24 bottleneck layers, and the fourth dense connection block may have 16 bottleneck layers.
[0152] Please continue to refer to this. Figure 4 The diagram illustrates the structure of the bottleneck layer provided in a specific example of the present invention. Figure 4 As shown, the bottleneck layer comprises a first batch normalization layer A, a first activation layer A, a third convolutional layer A, a first batch normalization layer B, a first activation layer B, and a third convolutional layer B connected in sequence. The kernel size of the third convolutional layer A is 1×1, and the kernel size of the third convolutional layer B is 3×3. Therefore, by adding a 1×1 convolution before the 3×3 convolution in the bottleneck layer, this invention reduces the number of feature maps and the dimensionality of each feature map, thereby reducing computational cost and fusing features from various channels. Furthermore, since the bottleneck layer performs batch normalization (BN) and ReLU activation operations before both the 1×1 and 3×3 convolution operations, training speed and convergence efficiency can be further improved.
[0153] Please continue to refer to this. Figure 5 The diagram illustrates the structure of a transition block provided in a specific example of the present invention. Figure 5 As shown, the first transition block, the second transition block, and the third transition block each include a second batch normalization layer, a second activation layer, a fourth convolutional layer, and a second pooling layer (preferably an average pooling layer) connected in sequence. The kernel size of the fourth convolutional layer is 1×1. Thus, the convolutional operation of the fourth convolutional layer can reduce the dimensionality of the feature map, and the average pooling operation of the second pooling layer can solve the problem of excessive channels in the feature map, preventing model complexity caused by too many densely connected blocks. Furthermore, since each transition block performs batch normalization (BN) and ReLU activation operations before the 1×1 convolutional operation, the number of parameters can be further compressed.
[0154] Please continue to refer to this. Figure 6 The diagram illustrates the structure of an upward transition block provided in a specific example of the present invention. Figure 6 As shown, both the first upward transition block and the second upward transition block include a third batch normalization layer A, a third activation layer A, a fifth convolutional layer A, a third batch normalization layer B, a third activation layer B, a fifth convolutional layer B, a third batch normalization layer C, a third activation layer C, and a first deconvolutional layer connected in sequence. The size of the convolutional kernels of the fifth convolutional layer A and the fifth convolutional layer B is 3×3.
[0155] Furthermore, the training samples used in the segmentation model training process include sample medical images (e.g., sample cardiac images) with labeled regions of interest (ROIs) of the target organ tissue (e.g., aortic valve leaflet ROIs) and corresponding target organ tissue mask images (e.g., aortic valve leaflet mask images). Specifically, the OpenCV contour extraction algorithm can be used to find the contours of the target organ tissue (e.g., aortic valve leaflets) within the ROIs of the sample medical images to segment the target organ tissue mask images (e.g., aortic valve leaflet mask images). It should be noted that, as those skilled in the art will understand, since the neural network model requires images of a uniform size as input, the sample medical images with labeled ROIs and their corresponding target organ tissue mask images need to be converted to a preset size, such as 320×320. Furthermore, to improve the robustness of the trained segmentation model, the present invention also amplifies the acquired samples, specifically by adjusting the contrast of the sample medical images, adding Gaussian noise, and performing transformations such as translation, rotation, and scaling.
[0156] In one exemplary implementation, the segmentation model uses a binary cross-entropy loss function during training, the formula of which is shown below:
[0157]
[0158]
[0159] In the formula, y i For real labels, This is the predicted result.
[0160] Furthermore, after training the segmentation model, this invention also uses the Dice coefficient formula to evaluate the algorithm accuracy of the segmentation model, as shown below:
[0161]
[0162] In the formula, X represents the prediction result, and Y represents the true label.
[0163] The value of Dice ranges from 0 to 1. The closer the Dice value is to 1, the higher the segmentation accuracy of the segmentation model.
[0164] As an example, during the training of the segmentation model, the learning rate is set to 1e-3 (i.e., 0.001), and Adam (adaptive moment estimation) is used as the optimizer. The learning rate of each parameter is dynamically adjusted using the first and second moment estimates of the gradient, and clipnorm = 0.001 is added to the optimizer parameters for gradient clipping. Furthermore, during the training of the segmentation model, the batch size is 16, the epochs are 100, and an early stopping strategy is adopted. If the loss function of the validation set does not decrease for 20 consecutive training epochs, the training is terminated early.
[0165] Please continue to refer to this. Figures 7a to 7d ,in Figure 7a The image of the region of interest (ROI) of the aortic valve leaflet during ventricular diastole (ROI image of the target organ / tissue) is schematically shown in a specific example of the present invention. Figure 7b The illustration shows the... Figure 7a The aortic valve leaflet image obtained by segmentation (target organ tissue image); Figure 7c The image of the region of interest (ROI) of the aortic valve leaflet during ventricular systole (ROI image of the target organ / tissue) is schematically shown in a specific example of the present invention. Figure 7d The illustration shows the... Figure 7c The aortic valve leaflet image obtained through segmentation (target organ / tissue image). For example... Figures 7a to 7d As shown, by using the segmentation model in this invention to segment the region of interest (e.g., region of interest of aortic valve leaflets) of the target organ tissue corresponding to each frame of medical image (e.g., cardiac image), the target organ tissue image (e.g., aortic valve leaflet image) corresponding to each frame of medical image (e.g., cardiac image) can be accurately obtained.
[0166] Please continue to refer to this. Figure 8 This illustration schematically shows the region of interest (ROI) of the aortic valve leaflets (target organ / tissue) and the aortic valve leaflets (target organ / tissue) in a cardiac image provided by a specific example of the present invention, wherein the area defined by the rectangular border is the ROI of the aortic valve leaflets. Figure 8As shown, in one exemplary embodiment, for visual demonstration, the bounding box of the region of interest (GIO) of the target organ (e.g., the GIO of the aortic valve leaflets) and the outline of the target organ (e.g., the aortic valve leaflets) can be drawn on each frame of medical images (e.g., cardiac images). Median filtering is then used to set the grayscale value of each pixel in each frame of medical images to the median of the grayscale values of all pixels within that pixel's neighborhood window. The size parameter of the filtering kernel can be set according to specific circumstances, for example, to 5×5. Thus, median filtering can effectively remove salt-and-pepper noise from each frame of medical images. It should be noted that, as those skilled in the art will understand, in other embodiments, other filtering methods besides median filtering can be used to filter each frame of medical images, and this invention does not limit this to such methods.
[0167] In one exemplary embodiment, obtaining the classification result of the medical video based on the classification result of the target organ tissue mask corresponding to each frame of medical image includes:
[0168] The classification results of the medical video are obtained by using the classification results of the target organ tissue mask corresponding to each frame of medical image and the confidence probability values of the region of interest of the target organ tissue in each frame of medical image extracted by the target detection model.
[0169] Therefore, by using the classification results of the target organ tissue mask corresponding to each frame of medical images and the confidence probability value of the region of interest of the target organ tissue in each frame of medical images extracted by the target detection model, the classification result of the medical video can be obtained. By comprehensively considering the confidence probability value of the region of interest of the target organ tissue in each frame of medical images and the classification results of the target organ tissue mask corresponding to each frame of medical images, the accuracy of medical video classification can be further improved.
[0170] Furthermore, in some embodiments, obtaining the classification result of the medical video based on the classification results of the target organ tissue mask corresponding to each frame of medical images and the confidence probability values of the regions of interest of the target organ tissue in each frame of medical images extracted by the target detection model includes:
[0171] Based on the classification results of the target organ tissue mask corresponding to each frame of medical image, the probability value of the target organ tissue mask corresponding to each frame of medical image being judged as normal target organ tissue is obtained;
[0172] For each frame of a medical image, based on the probability value that the target organ tissue mask corresponding to that frame of the medical image is determined to be normal, and the confidence probability value of the region of interest of the target organ tissue extracted by the target detection model for that frame of the medical image, the first probability value corresponding to that frame of the medical image is calculated according to the following formula:
[0173] P segancls1i =P 1i *p i
[0174] In the formula, P segandcls1i Let P represent the first probability value corresponding to the i-th frame of the medical image. 1i p represents the probability value that the target organ tissue mask corresponding to the i-th frame of the medical image is judged as normal. i This represents the confidence probability value of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model;
[0175] Based on the first probability value corresponding to each frame of medical image, obtain the first probability mean corresponding to the medical video;
[0176] If the mean of the first probability is greater than a first preset threshold, the classification result of the medical video is determined to be that the target organ tissue is normal; otherwise, the classification result of the medical video is determined to be that the target organ tissue is abnormal.
[0177] Specifically, the average first probability value corresponding to the medical video can be obtained by averaging the first probability values corresponding to each frame of the medical image. Assuming the medical video includes n frames of medical images, the average first probability value corresponding to the medical video is... It can be represented as:
[0178]
[0179] Furthermore, in other embodiments, the step of obtaining the classification result of the medical video based on the classification result of the target organ tissue mask corresponding to each frame of medical images and the confidence probability value of the region of interest of the target organ tissue in each frame of medical images extracted by the target detection model includes:
[0180] Based on the classification results of the target organ tissue mask corresponding to each frame of medical image, the probability value of the target organ tissue mask corresponding to each frame of medical image being judged as an abnormality of the target organ tissue is obtained.
[0181] For each frame of a medical image, based on the probability value that the target organ tissue mask corresponding to that frame of the medical image is determined to be an abnormality of the target organ tissue and the confidence probability value of the region of interest of the target organ tissue extracted by the target detection model in that frame of the medical image, the second probability value corresponding to that frame of the medical image is calculated according to the following formula:
[0182] P segancls2 =P 2i *p i
[0183] In the formula, P segandcls2i P represents the second probability value corresponding to the i-th frame of the medical image. 2i p represents the probability value that the target organ tissue mask corresponding to the i-th frame of the medical image is identified as an abnormality of the target organ tissue. i This represents the confidence probability value of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model;
[0184] Based on the second probability value corresponding to each frame of medical image, obtain the second probability mean corresponding to the medical video;
[0185] If the mean of the second probability is greater than a second preset threshold, the classification result of the medical video is determined to be an abnormality of the target organ tissue; otherwise, the classification result of the medical video is determined to be a normality of the target organ tissue.
[0186] Specifically, the average second probability value corresponding to the medical video can be obtained by averaging the second probability values corresponding to each frame of the medical image. Assuming the medical video includes n frames of medical images, the average second probability value corresponding to the medical video is... It can be represented as:
[0187]
[0188] In one exemplary embodiment, the classification model includes a connected second backbone network and a classification network. The second backbone network is used to extract target organ tissue features (e.g., aortic valve leaflet features) from the input target organ tissue image. The classification network is used to identify whether the target organ tissue (e.g., aortic valve leaflet) in the target organ tissue image (e.g., aortic valve leaflet image) is normal or abnormal based on the target organ tissue features (e.g., aortic valve leaflet features) extracted by the second backbone network. Specifically, the classification network includes a global average pooling layer, a fully connected layer, and a softmax activation function layer connected in sequence. The global average pooling layer is used to perform dimensionality reduction on the target organ tissue features (e.g., aortic valve leaflet features) extracted by the second backbone network to reduce the number of model parameters, thereby minimizing the overfitting effect. The fully connected layer is used to perform nonlinear mapping regression on the output of the global average pooling layer. The softmax activation function layer is used to normalize the output of the fully connected layer to generate a probability value for each classification category (normal and abnormal target organ tissue), thereby obtaining the classification result of the target organ tissue image (the classification category with the higher probability value is the classification result of the target organ tissue image, i.e., the classification result of the corresponding medical image).
[0189] Specifically, the second backbone network of the classification model also adopts the Densenet neural network model. The structure of the second backbone network of the classification model is largely the same as that of the first backbone network of the segmentation model, and will not be described in detail here.
[0190] Furthermore, the training samples used in the classification model training process include target organ tissue mask images (e.g., aortic valve leaflet mask images) obtained by segmenting each frame of medical training images (e.g., echocardiogram images) in the acquired medical training video (e.g., echocardiogram video) and their corresponding category labels (for each frame of medical training images in the same medical training video, the category label of the medical training video is used as the category label of the target organ tissue mask image corresponding to each frame of medical training image; that is, if the category label of the medical training video is "target organ tissue normal," then the category label of all target organ tissue mask images corresponding to that medical training video is "target organ tissue normal," and vice versa, they are all "target organ tissue abnormal"). The acquired training samples are divided into training set, validation set, and test set according to a certain ratio.
[0191] Based on the same inventive concept, the present invention also provides an electronic device, please refer to... Figure 9 The diagram illustrates a block structure of an electronic device according to an embodiment of the present invention. Figure 9As shown, the electronic device includes a processor 101 and a memory 103. The memory 103 stores a computer program, which, when executed by the processor 101, implements the medical video classification method described above.
[0192] like Figure 9 As shown, the electronic device also includes a communication interface 102 and a communication bus 104, wherein the processor 101, the communication interface 102, and the memory 103 communicate with each other via the communication bus 104. The communication bus 104 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus 104 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not indicate that there is only one bus or one type of bus. The communication interface 102 is used for communication between the aforementioned electronic device and other devices.
[0193] The processor 101 referred to in this invention can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 101 is the control center of the electronic device, connecting various parts of the electronic device through various interfaces and lines.
[0194] The memory 103 can be used to store the computer program. The processor 101 implements various functions of the electronic device by running or executing the computer program stored in the memory 103 and calling the data stored in the memory 103.
[0195] The memory 103 may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0196] The present invention also provides a readable storage medium storing a computer program, which, when executed by a processor, can implement the medical video classification method described above.
[0197] The readable storage medium of embodiments of the present invention can be any combination of one or more computer-readable media. The readable medium can be a computer-readable signal medium or a computer-readable storage medium. Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections having one or more wires, portable computer hard disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device.
[0198] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0199] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0200] In summary, compared with the prior art, the medical video classification method, electronic device, and storage medium provided by the present invention have the following advantages: The present invention first uses a target detection model to extract the region of interest (ROI) of the target organ / tissue in each frame of the acquired medical video, thereby obtaining the location information of the ROI corresponding to each frame of the medical video; then, curve fitting is performed based on the location information of the ROI of the target organ / tissue in each frame of the medical video, and the location information of the ROI of the target organ / tissue in each frame of the medical video is corrected based on the fitting result; then, based on the corrected location information of the ROI of the target organ / tissue in each frame of the medical video, the corresponding ROI of the target organ / tissue is cropped from each frame of the medical video, thereby obtaining the corresponding ROI image; next, a segmentation model is used to segment the ROI image of the target organ / tissue in each frame of the medical video, thereby obtaining the corresponding target organ / tissue mask; then, a classification model is used to identify and classify the target organ / tissue mask corresponding to each frame of the medical video, thereby obtaining the classification result of the target organ / tissue mask corresponding to each frame of the medical video; finally, based on the classification result of the target organ / tissue mask corresponding to each frame of the medical video, the classification result of the medical video is obtained. Therefore, this invention corrects the location information of the target organ / tissue region of interest (ROI) extracted from each frame of medical images by the target detection model, and then crops the corresponding ROI image from the medical image based on the corrected ROI location information. This allows for the acquisition of more accurate ROI images, effectively improving the accuracy of the acquired target organ / tissue mask and laying a solid foundation for obtaining accurate classification results. Furthermore, since the medical video classification method provided by this invention comprehensively considers the classification results of the target organ / tissue mask corresponding to each frame of the medical video, it effectively improves the accuracy of medical video classification (i.e., accurately identifying whether the target organ / tissue in the medical video is normal or abnormal). This reduces potential discrepancies caused by human factors, better assisting doctors in improving diagnostic efficiency and reducing the risk of errors in organ / tissue anomaly analysis using medical videos. In addition, this invention enables an end-to-end algorithm flow, has strong versatility, and effectively improves the efficiency of medical video classification.
[0201] It should be noted that computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0202] It should be noted that the apparatus and methods disclosed in the embodiments herein can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments herein. In this regard, each block in a flowchart or block diagram may represent a module, program, or part of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system to perform the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0203] In addition, the functional modules in the various embodiments of this article can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0204] The above description is merely a description of preferred embodiments of the present invention and is not intended to limit the scope of the invention in any way. Any changes or modifications made by those skilled in the art based on the above disclosure are within the protection scope of the present invention. Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the present invention and its equivalents, the present invention also intends to include these modifications and variations.
Claims
1. A medical video classification method, characterized in that, include: An object detection model is used to extract the region of interest (ROI) of the target organ or tissue in each frame of the acquired medical video, so as to obtain the location information of the ROI of the target organ or tissue corresponding to each frame of the medical video. Curve fitting is performed based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical images, and the location information of the region of interest of the target organ tissue corresponding to each frame of medical images is corrected based on the fitting results. Based on the location information of the corrected target organ tissue region of interest corresponding to each frame of medical image, the corresponding target organ tissue region of interest is cropped from each frame of medical image to obtain the corresponding target organ tissue region of interest image; A segmentation model is used to segment the region of interest (ROI) image of the target organ tissue corresponding to each frame of medical image in order to obtain the corresponding target organ tissue mask; A classification model is used to identify and classify the target organ and tissue masks corresponding to each frame of medical images in order to obtain the classification results of the target organ and tissue masks corresponding to each frame of medical images. The classification result of the medical video is obtained based on the classification result of the target organ tissue mask corresponding to each frame of medical image.
2. The medical video classification method according to claim 1, characterized in that, The step of performing curve fitting based on the location information of the region of interest (ROI) of the target organ / tissue corresponding to each frame of medical images, and correcting the location information of the ROI of the target organ / tissue corresponding to each frame of medical images based on the fitting result, includes: Based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical image extracted by the target detection model, curve fitting is performed to obtain the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue. Based on the correspondence between the fitted image frames and the location information of the target organ tissue region of interest, the location information of the target organ tissue region of interest corresponding to each frame of medical image is corrected to obtain the corrected location information of the target organ tissue region of interest corresponding to each frame of medical image.
3. The medical video classification method according to claim 2, characterized in that, The step of correcting the location information of the target organ's region of interest corresponding to each frame of the medical image based on the correspondence between the fitted image frames and the location information of the target organ's region of interest includes: For each frame of the medical image: Based on the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue, the location information of the fitted region of interest of the target organ tissue corresponding to the medical image frame is obtained. The first positional deviation information corresponding to the medical image is obtained based on the absolute value of the difference between the positional information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model and the positional information of the fitted target organ tissue region of interest corresponding to the medical image frame. Based on the first positional deviation information corresponding to the medical image frame and the confidence probability value of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model, the second positional deviation information corresponding to the medical image frame is obtained. Based on the second positional deviation information corresponding to the medical image frame, it is determined whether the positional information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is accurate; If so, the location information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is used as the corrected location information of the target organ tissue region of interest corresponding to the medical image frame. If not, then based on the location information of the target organ tissue region of interest corresponding to the medical image in the previous frame with accurate location information and the location information of the target organ tissue region of interest corresponding to the medical image in the next frame with accurate location information, the corrected location information of the target organ tissue region of interest corresponding to the medical image in the current frame is obtained.
4. The medical video classification method according to claim 3, characterized in that, The step of obtaining the second positional deviation information corresponding to the medical image frame based on the first positional deviation information corresponding to the frame and the confidence probability value of the target organ tissue region of interest extracted by the target detection model for the frame of medical image includes: The second positional deviation information corresponding to this frame of medical image is obtained according to the following formula: And i =And i *(1-p i ) In the formula, e i E represents the second positional deviation corresponding to the i-th frame of the medical image. i p represents the first positional deviation corresponding to the i-th frame of the medical image. i This represents the confidence probability value of the region of interest corresponding to the target organ tissue in the i-th frame of the medical image extracted by the target detection model.
5. The medical video classification method according to claim 3, characterized in that, The step of determining whether the location information of the target organ tissue region of interest corresponding to the medical image frame extracted by the target detection model is accurate based on the second positional deviation information corresponding to the medical image frame includes: Based on the first position deviation information corresponding to each frame of medical image, the mean first position deviation information corresponding to the medical video is obtained; The average confidence probability value of the target organ tissue region of interest corresponding to each frame of medical image is extracted based on the target detection model to obtain the average confidence probability value corresponding to the medical video. Based on the mean first positional deviation information corresponding to the medical video and the mean confidence probability corresponding to the medical video, the mean second positional deviation information corresponding to the medical video is obtained. The position judgment threshold is obtained based on the preset multiple threshold and the average value of the second position deviation corresponding to the medical video; For each frame of the medical video, based on the second positional deviation information corresponding to that frame of the medical image and the positional judgment threshold, it is determined whether the positional information of the target organ tissue region of interest corresponding to that frame of the medical image extracted by the target detection model is accurate.
6. The medical video classification method according to claim 2, characterized in that, The step of performing curve fitting based on the location information of the region of interest of the target organ tissue corresponding to each frame of medical image extracted by the target detection model, to obtain the correspondence between the fitted image frame and the location information of the region of interest of the target organ tissue, includes: Based on the x-coordinate and y-coordinate information of the first corner point and the second corner point of the target organ tissue region of interest corresponding to each frame of medical image extracted by the target detection model, curve fitting is performed on the x-coordinate, y-coordinate, and y-coordinate of the first corner point of the target organ tissue region of interest, respectively. This is to obtain the correspondence between the fitted image frame and the x-coordinate of the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the first corner point of the target organ tissue region of interest, the correspondence between the fitted image frame and the second corner point of the target organ tissue region of interest, and the correspondence between the fitted image frame and the second corner point of the target organ tissue region of interest.
7. The medical video classification method according to claim 1, characterized in that, The step of obtaining the classification result of the medical video based on the classification result of the target organ tissue mask corresponding to each frame of medical image includes: The classification results of the medical video are obtained by using the classification results of the target organ tissue mask corresponding to each frame of medical image and the confidence probability values of the region of interest of the target organ tissue in each frame of medical image extracted by the target detection model.
8. The medical video classification method according to claim 7, characterized in that, The step of obtaining the classification result of the medical video based on the classification results of the target organ tissue mask corresponding to each frame of medical images and the confidence probability value of the region of interest of the target organ tissue in each frame of medical images extracted by the target detection model includes: Based on the classification results of the target organ tissue mask corresponding to each frame of medical image, the probability value of the target organ tissue mask corresponding to each frame of medical image being judged as normal target organ tissue is obtained; For each frame of a medical image, based on the probability value that the target organ tissue mask corresponding to that frame of the medical image is determined to be normal, and the confidence probability value of the region of interest of the target organ tissue extracted by the target detection model for that frame of the medical image, the first probability value corresponding to that frame of the medical image is calculated according to the following formula: P segancls1i =P 1i *p i In the formula, P segandcls1i Let P represent the first probability value corresponding to the i-th frame of the medical image. 1i p represents the probability value that the target organ tissue mask corresponding to the i-th frame of the medical image is judged as normal. i This represents the confidence probability value of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model; Based on the first probability value corresponding to each frame of medical image, obtain the first probability mean corresponding to the medical video; If the mean of the first probability is greater than a first preset threshold, the classification result of the medical video is determined to be that the target organ tissue is normal; otherwise, the classification result of the medical video is determined to be that the target organ tissue is abnormal.
9. The medical video classification method according to claim 7, characterized in that, The step of obtaining the classification result of the medical video based on the classification results of the target organ tissue mask corresponding to each frame of medical images and the confidence probability value of the region of interest of the target organ tissue in each frame of medical images extracted by the target detection model includes: Based on the classification results of the target organ tissue mask corresponding to each frame of medical image, the probability value of the target organ tissue mask corresponding to each frame of medical image being judged as an abnormality of the target organ tissue is obtained. For each frame of a medical image, based on the probability value that the target organ tissue mask corresponding to that frame of the medical image is determined to be an abnormality of the target organ tissue and the confidence probability value of the region of interest of the target organ tissue extracted by the target detection model in that frame of the medical image, the second probability value corresponding to that frame of the medical image is calculated according to the following formula: P segancls2 =P 2i *p i In the formula, P segandcls2i P represents the second probability value corresponding to the i-th frame of the medical image. 2i p represents the probability value that the target organ tissue mask corresponding to the i-th frame of the medical image is identified as an abnormality of the target organ tissue. i This represents the confidence probability value of the region of interest of the target organ tissue corresponding to the i-th frame of the medical image extracted by the target detection model; Based on the second probability value corresponding to each frame of medical image, obtain the second probability mean corresponding to the medical video; If the mean of the second probability is greater than a second preset threshold, the classification result of the medical video is determined to be an abnormality of the target organ tissue; otherwise, the classification result of the medical video is determined to be a normality of the target organ tissue.
10. The medical video classification method according to claim 1, characterized in that, Before segmenting the region of interest (ROI) images of the target organ tissues corresponding to each frame of medical images using a segmentation model, the method further includes: For each frame of the medical image: The length dimension of the region of interest corresponding to the target organ tissue in this frame of medical image is taken as the target side length. The region of interest image of the target organ tissue is filled along the width direction to adjust the width dimension of the region of interest image of the target organ tissue to the target side length dimension; The region of interest image of the target organ tissue is magnified or reduced by adjusting the width dimension to the target side length dimension, so as to adjust the size of the region of interest image of the target organ tissue to a preset size.
11. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, implements the method of any one of claims 1 to 10.
12. A readable storage medium, characterized in that, The readable storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Cooperative intelligent security and protection method and device based on polymorphic fitting
CN113449663A
Blood vessel risk assessment method, computer equipment and storage medium
CN113744223A