Rapid Diagnosis Device for Polyp Conditions Based on Semantic Segmentation and Multimodality

Through a rapid diagnosis device for polyp status based on semantic segmentation and multimodality, the BEIT and VX2TEXT networks are used to process polyp scanning videos, which solves the problem of long and high cost of polyp diagnosis in the prior art, and achieves a fast and accurate health status assessment.

CN115631182BActive Publication Date: 2025-06-20FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211381144.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-01
Publication Date
2025-06-20
Estimated Expiration
2042-11-01

AI Technical Summary

Technical Problem

The existing polyp condition diagnosis technology takes time and is expensive, and depends on doctors' judgment, so efficiency and accuracy are difficult to guarantee.

Method used

A rapid diagnosis device for polyp status based on semantic segmentation and multimodality was used to process polyp scanning videos using improved BEIT and VX2TEXT networks to generate medical diagnosis and health status assessment scores.

Benefits of technology

Significantly reduce the time for doctors to judge, provide high reference health status assessment, optimize medical resources, and reduce the risk of medical accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631182B_ABST
    Figure CN115631182B_ABST
Patent Text Reader

Abstract

The present invention proposes a rapid diagnosis device for polyp conditions based on semantic segmentation and multi-modal, based on a computer system, including: a polyp segmentation model obtained by training a public polyp dataset using an improved BEIT network; a polyp video processing module for extracting frames from a polyp scan video, segmenting it using the polyp segmentation model to generate predicted frames, and then forming a segmented polyp video from the set of predicted frames through an optical flow method; an evaluation module for performing multi-modal training on the polyp video obtained by the polyp video processing module using an improved VX2TEXT network according to a polyp video dataset, so as to realize generating an evaluation of the polyp condition by inputting a polyp scan video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of machine learning, medical auxiliary equipment, etc., and particularly relates to a rapid polyp condition diagnosis device based on semantic segmentation and multi-modalities. Background Art

[0002] With the continuous development of technology and social progress, the combination of AI and medicine has become one of the relatively popular AI applications at present. Among them, for the diagnosis of traditional polyp conditions, doctors need to make a judgment through a series of examinations such as polyp scan videos and blood routine analysis indicators. The whole process takes a long time and costs a lot. Summary of the Invention

[0003] To solve the problems of defects and deficiencies existing in the prior art, the present invention proposes a rapid polyp condition diagnosis device based on semantic segmentation and multi-modalities, which realizes rapid auxiliary evaluation of polyp conditions based on visual technologies such as BEIT, VX2TEXT, and optical flow method. The shape, size, and number of polyps of a patient are obtained from a polyp scan video, a medical diagnosis is generated, and a health condition evaluation score with strong reference is output. It can greatly reduce the judgment time of doctors, provide a very good reference for doctors to make judgments, optimize medical resources, and reduce medical accidents.

[0004] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0005] A rapid polyp condition diagnosis device based on semantic segmentation and multi-modalities, characterized in that, based on a computer system, it includes:

[0006] A polyp segmentation model, obtained by training a public polyp dataset using an improved BEIT network;

[0007] A polyp video processing module, used to extract frames from a polyp scan video, perform segmentation using the polyp segmentation model to generate prediction frames, and then form a segmented polyp video from the set of prediction frames through the optical flow method;

[0008] An evaluation module, used to perform multi-modal training on the polyp video obtained by the polyp video processing module using an improved VX2TEXT network according to a polyp video dataset, so as to realize generating a polyp condition evaluation by inputting a polyp scan video;

[0009] The improved BEIT network is specifically as follows: The 15-layer BEIT is used as the feature extraction backbone network. Every three BEIT block form a stage. Before the input of each stage, a residual connection is made with the input of the previous stage. After each stage, downsampling is performed once to form multi-scale feature information. For the downsampling layer, mean removal normalization is used to reduce the dependence between different stages. In the BEIT backbone network, the local attention mechanism in the first stage is replaced by a global attention mechanism to increase the receptive field of low-scale features, and the local attention mechanisms in the second and third stages are replaced by deformable attention mechanisms to prevent information loss while reducing the amount of computation and the number of parameters. Finally, after the output of the last stage, a 1×1 convolutional kernel is connected to reduce the channel dimension and remove the feature noise channels;

[0010] The improved VX2TEXT network uses VX2TEXT for multi-modal training, extracts feature information from multi-modal inputs, and combines them into natural language texts. Position encoding is added to its text feature extraction backbone network to ensure the correct absolute position information of the text. The global attention mechanism is used to enhance the connection between texts. The text feature extraction backbone network is divided into three stages, and each stage extracts features from different aspects. The first stage extracts the description of the polyp shape in the text, the second stage extracts the description of the polyp size in the text, and the third stage extracts the description of the number of polyps in the text. In the decoding process, the autoregressive decoder mode is used to fuse multi-modal information and generate texts.

[0011] Further, the training process of the polyp segmentation model specifically includes the following steps:

[0012] Step S11: Three open-source semantic segmentation datasets, Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB, are used. For the data formats of different datasets, opencv is used to process them into the same png format, and the Monte Carlo method is used to divide the datasets into a training set, a validation set, and a test set according to the ratio of 8:1:1;

[0013] Step S12: The improved BEIT feature extraction backbone network is used to replace the feature extraction backbone network of the semantic segmentation model Segnet. The feature pyramid and adaptive feature pooling are used to integrate multi-scale features into each proposed region to avoid arbitrary allocation and resulting in feature interference, thereby obtaining four feature layers of different sizes (8, 8), (16, 16), (32, 32), and (64, 64);

[0014] Step S13: Obtain the segmentation loss function MDICE by considering the overlap between the true label and the predicted value and the boundary penalty for prediction errors Loss which is defined as follows:

[0015]

[0016] where TP is the intersection of the predicted region and the label region, ALL is the sum of the areas of the predicted region and the label region; N is the number of pixel points on the segmentation result boundary line, i represents the index value of the pixel point, and (a i , b i ) is the i-th pixel point on the segmentation boundary, is the pixel point on the label boundary closest to (a i , b i ), α is the weight coefficient for TP and ALL, and β is the weight coefficient for the boundary penalty for prediction errors;

[0017] Step S14: Use the improved BEIT network as the feature extraction backbone of the Segnet model, train it on the Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB datasets until the loss converges, take the highest value of the average DICE of the validation sets of the three datasets, and save it as the best model; where the definition of DICE is as follows:

[0018]

[0019] where X represents the set of target pixel points of the segmentation true value, Y represents the set of segmentation target pixel points in the prediction result, and |X∩Y| in the formula is the intersection between X and Y, and |X| and |Y| respectively represent the number of elements in X and Y.

[0020] Furthermore, the working process of the polyp video processing module specifically includes the following steps:

[0021] Step S21: Extract frames from the video segmentation polyp dataset SUN-SEG, save the extracted frames as pictures, and generate independent folders for each video;

[0022] Step S22: Use the best semantic segmentation model obtained in Step S14 to perform semantic segmentation on the pictures in different folders obtained in Step S21 to generate prediction pictures, where in the prediction pictures, the pixel values of the background part are 0, and the pixel values of the polyp part are 255;

[0023] Step S23: Split each video independent folder, and save the segmented prediction pictures in the folder. Generate videos according to each folder. Use the optical flow method to find the relationship between the front and back frames based on the changes and correlations in the time domain, and generate new pictures to achieve the effect of frame interpolation;

[0024] Step S24: When inputting a brand-new polyp video into the polyp video processing module, a segmented binary video is obtained, where black pixel points represent the background and white pixel points represent polyps; through frame extraction and segmentation processing of SUN-SEG, a binary video after SUN-SEG segmentation is obtained.

[0025] Furthermore, the evaluation construction of the evaluation module specifically includes the following steps:

[0026] Step S31: For the binary video after SUN-SEG segmentation obtained in Step S24, perform annotation for each video one by one. The annotation content includes at least: writing medical diagnosis information for each video, and the health condition score of the corresponding patient;

[0027] Step S32: In the diagnosis information part, describe from three aspects: the shape, size, and number of polyps;

[0028] Step S33: Obtain the health condition score Score healthy , where the definition is as follows:

[0029] Score healthy = 100 - Score shape - Score size - Score count

[0030] Among them, Score shape represents the influence score on the health degree judged according to the shape of the polyp in the polyp diagnosis information; among them, Score shape is defined as follows:

[0031]

[0032] Among them, Size base represents the diameter length of the polyp base, represents the area of the polyp base, Score MS represents the polyp stalk score, where if there is a stalk, Score MS = 1, if there is no stalk, Score MS = 0;

[0033] Score sizeIndicates the influence score of the size on the severity of the polyp in the polyp diagnosis information, where Score size is defined as follows:

[0034]

[0035] where π is the pi defined in mathematics, is the area size of the whole polyp;

[0036] Score count Indicates the influence score of the number of polyps on the severity in the polyp diagnosis information, where Score count is defined as follows:

[0037] Score count = 0.8 * Count 2

[0038] where Count represents the number of polyps, and Count 2 is the square of the number of polyps;

[0039] Step S34: Write down the diagnosis information and the health condition score for each video; put the diagnosis information and the health condition score into the corresponding json - formatted file according to the file name to form a complete dataset corresponding to videos and labels.

[0040] Furthermore, the construction process of the evaluation module specifically includes the following steps:

[0041] Step S41: Perform multi - modal training using the improved VX2TEXT;

[0042] Step S42: Use the traditional video - text dataset TVQA for pre - training. After obtaining the baseline model, randomly divide the dataset obtained in step S34 into a training set, a validation set, and a test set. Use the training set for training, save the best model on the validation set, check the results on the test set, and save the final model.

[0043] Compared with the prior art, the present invention and its preferred solutions have the following beneficial effects:

[0044] 1. Combining the BEIT structure with the deformable attention structure not only retains the global features but also enhances the extraction of local features, improving the reason for the blurred segmentation edge.

[0045] 2. Using adaptive feature pooling, multi - scale features can be integrated into each proposal region, avoiding the situation of arbitrary assignment and resulting in feature noise.

[0046] 3. A new segmentation accuracy loss function is defined, increasing the penalty for edge errors, strengthening the learning of edge features, and ensuring the correct shape of the polyp.

[0047] 4. Use global attention instead of local attention mechanism to increase the receptive field of low-scale features, improve the model accuracy, and accelerate the model convergence.

[0048] 5. Combine multiple computer vision algorithms, which have good detection accuracy for polyp videos, can automatically generate diagnostic information and health status scores in the medical field, and complete the end-to-end model architecture from video to generated text. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0050] Figure 1 It is a flowchart of the implementation method for constructing the device of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] To make the features and advantages of this patent more obvious and understandable, specific embodiments are given below for detailed description as follows:

[0052] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0054] As Figure 1 shown, based on the present invention, a rapid diagnosis device for polyp conditions based on semantic segmentation and multi-modal is provided. The construction and development process of this embodiment is summarized into the following steps:

[0055] Step S1: Use the improved BEIT network to train the publicly available polyp dataset to obtain a polyp segmentation model;

[0056] Step S11: The datasets used by the semantic segmentation network are Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB, which are three open-source semantic segmentation datasets. For the data formats of different datasets (such as nii format or DICOM format), use OpenCV to process them into the same png format, and use the Monte Carlo method to divide the dataset into a training set, a validation set, and a test set according to the ratio of 8:1:1;

[0057] In this embodiment, a 15-layer BEIT is used as the feature extraction backbone network. Among them, every 3 BEIT block blocks form a stage. Before the input of each stage, a residual connection is made with the input of the previous stage. After each stage, downsampling is performed once to form multi-scale feature information. For the downsampling layer, mean removal normalization is used to reduce the dependence between different stages. In the BEIT backbone network, the local attention mechanism in the first stage is replaced by a global attention mechanism to increase the receptive field of low-scale features. The local attention mechanisms in the second and third stages are replaced by deformable attention mechanisms to reduce the amount of calculation and the number of parameters while preventing information loss. Finally, after the output of the last stage, a 1×1 convolutional kernel is connected to reduce the channel dimension and remove the feature noise channels.

[0058] Step S12: Use the improved BEIT feature extraction backbone network in S12 to replace the feature extraction backbone network of the semantic segmentation model Segnet. Using a feature pyramid and adaptive feature pooling, multi-scale features can be integrated into each proposal region to avoid arbitrary allocation and resulting in feature interference, so that four different-sized feature layers (8, 8), (16, 16), (32, 32), and (64, 64) can be obtained.

[0059] Step S13: By considering the overlap between the true label and the predicted value and the boundary penalty for prediction errors, the segmentation loss function MDICE Loss is defined as follows:

[0060]

[0061] where TP is the intersection of the predicted region and the label region, ALL is the sum of the area of the predicted region and the area of the label region. N is the number of pixel points on the segmentation result edge line, i represents the index value of the pixel point, and (a i , b i ) is the i-th pixel point on the segmentation edge, is the distance on the label edge from (a i , b i)The nearest pixel points, α is the weight coefficient for TP and ALL, and β is the weight coefficient for the boundary penalty of prediction errors.

[0062] Step S14: Use the improved Segnet model with the BEIT network as the feature extraction backbone, train it on the Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB datasets until the loss converges, take the highest value of the average DICE of the validation sets of the three datasets, and save it as the best model. Among them, DICE is a classic metric for describing the accuracy of semantic segmentation, and its definition is as follows:

[0063]

[0064] Where X represents the set of target pixel points of the segmentation ground truth, Y represents the set of segmented target pixel points in the prediction result. In the formula, |X∩Y| is the intersection between X and Y, and |X| and |Y| respectively represent the number of elements in X and Y.

[0065] Step S2: Extract frames from the polyp scan video, use the medical image segmentation network obtained in S1 for segmentation to generate prediction frames, and form a segmented polyp video from the set of prediction frames through the optical flow method. Specifically, it includes the following steps:

[0066] Step S21: Extract frames from the publicly available video segmentation polyp dataset SUN-SEG, extract one frame every 17 frames, which can greatly reduce the workload of segmentation. Save the extracted frames as pictures and generate independent folders for each video;

[0067] Step S22: Use the best semantic segmentation model obtained in step S14 to perform semantic segmentation on the pictures in different folders obtained in step S21 to generate prediction pictures. In the prediction pictures, the pixel values of the background part are 0, and the pixel values of the polyp part are 255;

[0068] Step S23: Segment each video independent folder, save the segmented prediction pictures in the folder, generate videos according to each folder, use the optical flow method to find the relationship between the front and back frames based on the changes and correlations in the time domain of the front and back frames, generate new pictures, and achieve the effect of frame filling. While ensuring the highest possible accuracy, it greatly speeds up the model operation speed.

[0069] Step S24: Implement inputting a brand-new polyp video to obtain a segmented binary video, where black pixel points represent the background and white pixel points represent the polyp. Through the frame extraction and segmentation processing of SUN-SEG, a binary video after the segmentation of SUN-SEG is obtained.

[0070] Step S3: Create a dataset and write a diagnostic opinion and a corresponding health score for each video in the polyp video dataset obtained in S2. Specifically, it includes the following steps:

[0071] Step S31: For the binary video obtained by SUN-SEG segmentation in step S24, label each video one by one. The labeling content mainly includes writing medical diagnostic information for each video and the health score of the patient based on the video.

[0072] Step S32: In the diagnostic information section, describe from three aspects: the shape, size, and number of polyps. The shape includes flat, hill-shaped, cauliflower-shaped, etc., and includes information such as whether there is a pedicle and the base width. The size mainly describes the diameter of the polyp, and the number mainly describes whether the polyp is single or multiple.

[0073] Step S33: Obtain the health score Score through the diagnostic information healthy , where Score healthy is affected by the above three aspects. Score healthy is defined as follows:

[0074] Score healthy = 100 - Score shape - Score size - Score count

[0075] Among them, Score shape represents the influence score on the health degree judged according to the shape of the polyp in the polyp diagnostic information. Among them, Score shape is defined as follows:

[0076]

[0077] Among them, Size base represents the diameter length of the polyp base, represents the area of the polyp base, Score MS represents the polyp pedicle score, where if there is a pedicle, Score MS = 1, and if there is no pedicle, Score MS = 0.

[0078] Score size represents the influence score of the size on the severity of the polyp in the polyp diagnostic information. Among them, Score size is defined as follows:

[0079]

[0080] Among them, π is the pi defined in mathematics, is the area size of the whole polyp.

[0081] Score count represents the influence score of the number of polyps on the severity in the polyp diagnosis information, where Score count is defined as follows:

[0082] Score count = 0.8 * Count 2

[0083] where Count represents the number of polyps, and Count 2 is the square of the number of polyps.

[0084] Step S34: Write down the diagnosis information and the health condition score for each video. Put the diagnosis information and the health condition score into the corresponding json - formatted file according to the file name, forming a complete dataset of videos corresponding to labels.

[0085] Step S4: Use the improved VX2TEXT network to perform multi - modal training on the diagnosis information and the segmented polyp videos, so as to realize automatically generating the polyp condition diagnosis information and the corresponding health degree score by inputting the polyp scan video. Specifically, it includes the following steps:

[0086] Step S41: Adopt VX2TEXT for multi - modal training. It can extract feature information from multi - modal inputs (composed of different forms of inputs such as videos and texts), and combine them into natural language texts. Add positional encoding to its text feature extraction backbone network to ensure the correct absolute position information of the text, use the global attention mechanism to enhance the connection between texts, and divide the text feature extraction backbone network into 3 stages. Each stage extracts different aspects of features. The first stage focuses on extracting the description of the polyp shape in the text, the second stage focuses on extracting the description of the polyp size in the text, and the third stage focuses on extracting the description of the number of polyps in the text. At the same time, in the decoding process, use the autoregressive decoder mode to fuse multi - modal information and generate texts.

[0087] Step S42: Use the traditional video - text dataset TVQA for pre - training. After obtaining a more appropriate baseline model, randomly divide the dataset obtained in S34 into a training set, a validation set, and a test set. Use the training set for training, save the best model on the validation set, check the results on the test set, and save the final model.

[0088] Step S43: Realize quickly generating the rapid diagnosis information of the polyp condition and the health degree score required in the medical field by inputting the polyp examination video.

[0089] Based on the above design, a rapid polyp condition diagnosis device of the present invention is formed in a computer system in a coded form, including:

[0090] A polyp segmentation model, obtained by training a public polyp dataset using an improved BEIT network;

[0091] A polyp video processing module, used to extract frames from a polyp scan video, segment it using the polyp segmentation model to generate prediction frames, and then form a segmented polyp video from the set of prediction frames through an optical flow method;

[0092] An evaluation module, used to perform multi-modal training on the polyp video obtained by the polyp video processing module using an improved VX2TEXT network according to a polyp video dataset, and realize generating a polyp condition evaluation by inputting a polyp scan video.

[0093] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0094] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0095] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks. Figure 1 One process or a plurality of processes and / or blocks Figure 1 Steps for realizing the functions specified in one block or a plurality of blocks.

[0097] As mentioned above, it is only a preferred embodiment of the present invention, and is not a limitation on the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

[0098] This patent is not limited to the above best implementation mode. Anyone inspired by this patent can obtain various other forms of rapid diagnosis devices for polyp conditions based on semantic segmentation and multi-modal. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage of this patent.

Claims

1. A rapid diagnosis device for polyp conditions based on semantic segmentation and multi-modalities, characterized in that, Based on a computer system, including: A polyp segmentation model obtained by training an open polyp dataset using an improved BEIT network; A polyp video processing module for extracting frames from a polyp scan video, segmenting using the polyp segmentation model to generate predicted frames, and then forming a segmented polyp video from the set of predicted frames through an optical flow method; An evaluation module for performing multi-modal training on the polyp video obtained by the polyp video processing module using an improved VX2TEXT network according to a polyp video dataset to achieve generating a polyp condition evaluation from an input polyp scan video; The improved BEIT network specifically is: using a 15-layer BEIT as the feature extraction backbone network, where every 3 BEIT block blocks form a stage. Before the input of each stage, a residual connection is made with the input of the previous stage. After each stage, downsampling is performed once to form multi-scale feature information; de-mean normalization is used for the downsampling layer to reduce the dependence between different stages; in the BEIT backbone network, the local attention mechanism in the first stage is replaced with a global attention mechanism to increase the receptive field of low-scale features, and the local attention mechanisms in the second and third stages are replaced with deformable attention mechanisms to prevent information loss while reducing the amount of computation and the number of parameters; finally, after the output of the last stage, a 1×1 convolutional kernel is connected to reduce the channel dimension and remove feature noise channels; The improved VX2TEXT network uses VX2TEXT for multi-modal training, extracts feature information from multi-modal inputs and combines them into natural language text; and adds positional encoding to its text feature extraction backbone network to ensure the correct absolute position information of the text, uses a global attention mechanism to enhance the connection between texts, and divides the text feature extraction backbone network into 3 stages, where each stage extracts features from different aspects. The first stage extracts the description of the polyp shape in the text, the second stage extracts the description of the polyp size in the text, and the third stage extracts the description of the number of polyps in the text; during the decoding process, an autoregressive decoder mode is used to fuse multi-modal information and generate text.

2. The rapid diagnosis device for polyp conditions based on semantic segmentation and multi-modalities according to claim 1, characterized in that, The training process of the polyp segmentation model specifically includes the following steps: Step S11: Using three open-source semantic segmentation datasets: Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB, processing them into the same png format using opencv according to the data formats of different datasets, and using the Monte Carlo method to divide the datasets into a training set, a validation set, and a test set in a ratio of 8:1:1; Step S12: Use the improved BEIT feature extraction backbone network to replace the feature extraction backbone network of the semantic segmentation model Segnet. Use the feature pyramid and adaptive feature pooling to integrate multi-scale features into each proposed region, avoiding arbitrary assignment and resulting in feature interference, thereby obtaining four feature layers of different sizes (8, 8), (16, 16), (32, 32), and (64, 64). Step S13: The segmentation loss function MDICE is obtained by considering the overlap between the true label and the predicted value and the boundary penalty for prediction errors. Loss It is defined as follows: where TP is the intersection of the predicted region and the labeled region, ALL is the sum of the areas of the predicted region and the labeled region; N is the number of pixel points on the segmentation result boundary line, i represents the index value of the pixel point, (a i , b i ) is the i-th pixel point on the segmentation boundary, is the pixel point on the labeled boundary closest to (a i , b i ), α is the weight coefficient for TP and ALL, and β is the weight coefficient for the boundary penalty for prediction errors; Step S14: Use the Segnet model with the improved BEIT network as the feature extraction backbone and train it on the Kvasir-SEG, CVC-ClinicDB, and CVC-ColonDB datasets until the loss converges. Take the highest value of the average DICE of the validation sets of the three datasets and save it as the best model. The definition of DICE is as follows: Where X represents the set of target pixel points of the segmentation ground truth, and Y represents the set of segmentation target pixel points in the prediction result. In the formula, |X∩Y| is the intersection between X and Y, and |X| and |Y| represent the number of elements in X and Y respectively.

3. The rapid diagnosis device for polyp conditions based on semantic segmentation and multi-modalities according to claim 2, characterized in that, The working process of the polyp video processing module specifically includes the following steps: Step S21: Extract frames from the video segmentation polyp dataset SUN-SEG, save the extracted frames as pictures, and generate independent folders for each video. Step S22: Use the best semantic segmentation model obtained in Step S14 to perform semantic segmentation on the pictures in different folders obtained in Step S21 to generate prediction pictures. In the prediction pictures, the pixel values of the background part are 0, and the pixel values of the polyp part are 255. Step S23: Segment each video independent folder, save the segmented prediction pictures in the folder, generate videos according to each folder, use the optical flow method to find the relationship between the front and back frames based on the changes and correlations in the time domain of the front and back frames, and generate new pictures to achieve the effect of frame filling. Step S24: When a brand-new polyp video is input into the polyp video processing module, a segmented binary video is obtained, where the black pixel points represent the background and the white pixel points represent the polyp. Through the frame extraction and segmentation processing of SUN-SEG, a binary video after the segmentation of SUN-SEG is obtained.

4. The rapid diagnosis device for polyp conditions based on semantic segmentation and multi-modalities according to claim 3, characterized in that, The evaluation construction of the evaluation module specifically includes the following steps: Step S31: Label each video in the binary video after the segmentation of SUN-SEG obtained in Step S24. The labeling content at least includes: writing medical diagnosis information for each video and the health condition score of the corresponding patient. Step S32: In the diagnosis information part, describe it from three aspects: the shape, size, and number of polyps. Step S33: Obtain the health score Score based on the diagnostic information healthy , where the definitions are as follows: Score healthy = 100 - Score shape -Score size -Score count Among them, Score shape indicates the impact score on the health level judged according to the shape of the polyp in the polyp diagnosis information; among them, Score shape is defined as follows: Among them, Size base represents the diameter length of the polyp base, represents the area of the polyp base, Score MS represents the polyp stalk score, where if there is a stalk, then Score MS = 1, and if there is no stalk, then Score MS = 0; Score size Indicates the impact score of the size on the severity of the polyp in the polyp diagnosis information, where Score size is defined as follows: where π is the pi defined in mathematics, is the area of the whole polyp; Score count Indicates the impact score of the number of polyps on the severity in the polyp diagnosis information, where Score count is defined as follows: Score count = 0.8 * Count 2 where Count represents the number of polyps, and Count 2 is the square of the number of polyps; Step S34: Record the diagnosis information and health condition score for each video; put the diagnosis information and health condition score into the corresponding json format file according to the file name to form a complete dataset corresponding to videos and labels.

5. The polyp condition rapid diagnosis device based on semantic segmentation and multi-modal according to claim 4, characterized in that, The construction process of the evaluation module specifically includes the following steps: Step S41: Adopt the improved VX2TEXT for multi-modal training. Step S42: After pre-training using the traditional video text dataset TVQA to obtain a baseline model, randomly divide the dataset obtained in Step S34 into a training set, a validation set, and a test set. Train using the training set, save the model with the best performance on the validation set, check the results on the test set, and save the final model.

Citation Information

Patent Citations

  • Semantic segmentation method and device for polyp image

    CN113781489A

  • Polyp segmentation method based on combination of attention splitting network and TRW-S algorithm

    CN114359214A