Endoscopic lesion feature recognition system
The gastrointestinal endoscopic lesion feature recognition system utilizes a multi-task classification neural network to identify multiple feature groups and perform secondary judgments, solving the problems of high misdiagnosis rate and low efficiency of traditional AI models in diagnosing various diseases under endoscopy, and achieving efficient and accurate identification of multiple lesions and diseases.
Patent Information
- Application Number
- CN202311040222.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-17
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-08-17
AI Technical Summary
Traditional AI models struggle to identify multiple diseases simultaneously when used in endoscopic diagnosis of gastrointestinal diseases, resulting in high misdiagnosis rates, high computational demands, high model complexity, and high sample processing costs.
A digestive tract endoscopic lesion feature recognition system is adopted, which identifies multiple feature groups through a multi-task classification neural network and performs secondary judgment to achieve automatic identification of various lesions and diseases.
It improves the accuracy and efficiency of diagnosis, reduces the workload of manual annotation, provides truly usable auxiliary diagnostic tools, and expands the value of AI in clinical applications.
Smart Images

Figure CN117036815B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular, the present application relates to a digestive tract endoscope under lesion feature recognition system. BACKGROUND
[0002] The application of artificial intelligence (AI) in clinical medicine has become a research hotspot. The main method is to build a multi-layer deep learning model, collect a large number of digestive endoscopic lesion pictures from hospitals, and require doctors to label the lesions represented by the pictures as training data, so that AI learns the characteristics of the data, and ultimately achieves the purpose of identifying certain gastrointestinal diseases. Usually, in the form of software prompts, real-time assistance is provided to endoscopic doctors during examination, and a prompt signal is generated on the screen after the lesion is found, reminding the doctor to observe and operate and make a preliminary diagnosis. Such technology has its own research and application in esophageal cancer, gastric cancer, colorectal adenomatous polyps and other directions, and has proven to have clinical value, which can improve the efficiency of endoscopic examination and reduce missed diagnosis. In addition to localized lesions, AI-assisted diagnosis of chronic atrophic gastritis, ulcerative colitis and other lesions has also made progress. Although AI has reached the level of doctors in identifying some static pictures of lesions, and some technologies can even make a benign and malignant risk diagnosis, overall, such applications have not yet reached the level of mature products required by the clinic, with a high false positive rate, and there is still a distance to widespread application. The main achievements are scientific research and clinical trials.
[0003] Currently, the main AI technology basically identifies and diagnoses endoscopic features of a single disease, and does not have the function of identifying multiple diseases. If the disease diagnosis capability needs to be increased, a new model needs to be established, and after completing new sample training, it works in parallel with the original model to make a diagnosis. However, there are dozens of disease diagnoses in clinical digestive endoscopy, and to assist doctors, it needs to have the ability to diagnose and identify dozens of diseases at the same time, which puts higher requirements on the AI technology using single-disease models in terms of computing power and accuracy. With the current hardware capability of small computers, it is difficult to balance the needs of low latency and versatility. Secondly, AI technology needs a large number of samples as a prerequisite for training models, and each increase in disease diagnosis requires corresponding costs for sample collection and labeling, which also limits the rapid expansion of the diagnosis capability of such AI applications. Thirdly, in the current technical solution, for complex and variable types of diseases, there are great limitations in the diagnosis conclusion. Due to the limitations of the technical principle, the more variable and complex the manifestations of the same disease are, the more difficult it is for AI to reduce false positives and false negatives to a reasonable level.
[0004] In the field of diagnosing gastrointestinal diseases under endoscopy, there are currently various AI technologies and methods, such as deep learning, convolutional neural networks, support vector machines, Bayesian networks, etc. These technologies and methods can theoretically be used for diagnosing gastrointestinal diseases under endoscopy, but they have limitations or problems that are difficult to overcome in achieving the purpose of identifying multiple diseases, such as: 1) single disease or lesion recognition: some AI technologies can only recognize a single disease or lesion, but cannot simultaneously recognize multiple different lesions and diseases. 2) Large sample requirement: some AI technologies require a large amount of labeled data and the knowledge and experience of professional doctors to train high-quality models. 3) High model complexity: for some AI technologies, establishing multiple models or retraining models can increase the complexity and training time of the models. 4) High misdiagnosis rate: some AI technologies may not consider the mutual influence and similarity between different lesions, which can lead to misdiagnosis. Therefore, in order to solve the above technical problems, a new solution is needed to make AI-assisted digestive endoscopy diagnosis technology widely applicable. SUMMARY
[0005] Technical problems to be solved:
[0006] Traditional AI models are usually trained and optimized for a single disease or a single lesion, requiring a large amount of labeled data and the knowledge and experience of professional doctors. Moreover, in actual application, different lesions may have mutual influence and similarity, and traditional single-disease training methods cannot handle these problems well. In addition, establishing a model for each disease also increases the complexity of the model, the cost of sample processing, and the training time, limiting its efficiency and feasibility in actual application.
[0007] Specifically, one aspect of the present application aims to solve the problems of traditional AI models in assisting in diagnosing gastrointestinal diseases under endoscopy, and proposes a new system. That is, different lesion categories are analyzed for features, multiple feature points that meet the clinical diagnosis criteria are disassembled and classified, and the AI model is identified as the target. This way, different lesions and diseases can be automatically identified at the same time, avoiding the problem of establishing multiple models and retraining. Secondary judgment of the identified feature combinations can lead to correct disease diagnosis. For the unpredictable patient condition in clinical work, this method can improve the accuracy and efficiency of diagnosis, provide doctors with a truly usable auxiliary diagnosis tool, and make up for the shortcomings of traditional AI-assisted endoscopy technology, such as single diagnosis capability, poor anti-interference, large demand for computing power, and high misdiagnosis rate.
[0008] Technical scheme:
[0009] A lesion feature recognition system under a digestive tract scope, the system comprising:
[0010] 1) an image acquisition module; and
[0011] 2) a multi-level feature recognition and training module under the digestive tract scope, comprising a first multi-task classification neural network for recognizing a preset n different categories of digestive tract feature groups and label values in the digestive tract feature groups, wherein 2≤n≤10; and
[0012] 3) a feature result collection module for collecting the recognition results output by the multi-level feature recognition and training module under the digestive tract scope;
[0013] Wherein, the digestive tract feature group comprises a local lesion feature group, and the label values of the local lesion feature group include local concave lesions, local raised lesions or local flat lesions.
[0014] The downstream network of the local lesion feature group comprises a local lesion multi-level feature recognition and training module, and the local lesion multi-level feature recognition and training module comprises a second multi-task classification neural network for recognizing a preset m different categories of local lesion features and corresponding supplementary label values, wherein 1≤m≤5.
[0015] Another aspect of the present application is to provide an electronic device comprising an electronic device body and the above-mentioned lesion feature recognition system under the digestive tract scope, and the lesion feature recognition system under the digestive tract scope is installed on the electronic device body.
[0016] Advantages of the present application:
[0017] The present application adopts a new AI model based on feature point classification and summary, which can simultaneously recognize multiple different lesions and diseases, avoids the problem of establishing multiple models and retraining, and improves the accuracy and efficiency of endoscopic diagnosis of gastrointestinal diseases. Specifically:
[0018] 1) Improve diagnostic accuracy: the AI model used in the present application can automatically recognize and classify multiple different lesions and diseases, avoiding the problem of establishing multiple models and retraining. By making a second judgment on the identified feature combinations, more accurate disease diagnosis results can be obtained, reducing the misdiagnosis rate and improving the diagnostic accuracy.
[0019] 2) Improve diagnostic efficiency: traditional AI models need to establish a model for each disease, which increases the complexity of the model, the cost of sample processing and training time, and limits its efficiency and feasibility in practical applications. The AI model based on feature point classification and summary used in the present application can recognize multiple different lesions and diseases at the same time, avoiding the problem of establishing multiple models and retraining, and improving the efficiency of endoscopic diagnosis of gastrointestinal diseases.
[0020] 3) Reduce the workload of manual annotation: the AI model used in the present application can automatically analyze and classify the features of different lesions and diseases, reducing the workload of manual annotation.
[0021] 4) Improve clinical application value: the AI model of the present application can recognize multiple different lesions and diseases at the same time, and can provide doctors with a truly usable auxiliary diagnostic tool, improve the examination efficiency of endoscopists, reduce missed diagnosis, and expand the value of AI in clinical application, so that AI technology application can be popularized. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 A schematic diagram of a gastrointestinal endoscopic lesion feature recognition system constructed in the embodiments of the present application. DETAILED DESCRIPTION
[0023] The present application discloses a gastrointestinal endoscopic lesion feature recognition system, and those skilled in the art can refer to the content of this paper to improve the process parameters appropriately. It should be particularly pointed out that all similar substitutions and changes are obvious to those skilled in the art, and they are considered to be included in the present application, and the relevant personnel can obviously make changes or appropriate changes and combinations to the content described herein without departing from the content, spirit and scope of the present application, to realize and apply the present application technology.
[0024] In the present application, unless otherwise specified, the scientific and technical terms used in this paper have the meanings commonly understood by those skilled in the art. Unless otherwise specified, in the entire specification and claims, the term "includes" or its variants such as "contains" or "includes" and the like will be understood to include the stated element or component, without excluding other elements or components. The terms "such as", "for example" and the like are intended to indicate exemplary embodiments, and are not intended to limit the scope of the present disclosure.
[0025] In order for those skilled in the art to better understand the technical solutions of the present application, the present application will be further described in detail below in conjunction with specific embodiments.
[0026] As shown in Figure 1 The present embodiment discloses a gastrointestinal endoscopic lesion feature recognition system of the present application, which comprises:
[0027] S1. Image acquisition module.
[0028] The image acquisition module is used to acquire images in real time or pre-stored in a medium. The image acquisition module can include a video acquisition module and / or a picture acquisition module. In some embodiments of the present application, the image acquisition module can use a video capture card to transmit digital / analog video signals such as HDMI, DVI, SDI, S-Video, etc. into a computer, read the video signals through opencv, and convert them into RGB image format frame by frame.
[0029] In some embodiments of the present application, when the image acquisition module acquires the video real-time image signal output by the endoscope device, the image acquisition module performs interval sampling on the obtained frame-by-frame images, for example, at an interval sampling rate of 1 second / time, to obtain the images to be identified, and records the frame sequence number of the images to be identified.
[0030] S2. Multi-level feature recognition and training module under endoscopy.
[0031] The multi-level feature recognition and training module under endoscopy includes a first multi-task classification neural network for identifying a preset n different categories of endoscopic features and label values in the endoscopic feature groups, wherein 2≤n≤10. The endoscopic feature groups include a local lesion feature group. In some embodiments of the present application, n can be 2, 3, 4, 5, 6, 7, 8, 9 or 10. In one embodiment of the present application, n=5. That is, the endoscopic feature groups include five different categories, i.e., a mucosa feature group, a blood vessel texture feature group, a bleeding feature group, a mucus feature group and a local lesion feature group. In some other embodiments of the present application, at least one of the mucosa feature group, the blood vessel texture feature group, the bleeding feature group or the mucus feature group, or other related endoscopic feature groups, can be optionally selected. The above mucosa features, blood vessel texture features, bleeding features and mucus features all belong to diffuse lesion features.
[0032] In some embodiments of the present application, the label values corresponding to the different categories of endoscopic feature groups can include, for example, the label values of the mucosa feature group include mucosa thinning, mucosa strip redness, mucosa patch redness, mucosa edema, island mucosa, residual glandular mucosa, gastric mucosa, intestinal mucosa or intestinal metaplasia; the label values of the blood vessel texture feature group include blood vessel non-transparency, strip blood vessels, dendritic blood vessels or point-patch blood vessels; the label values of the bleeding feature group include no bleeding, point-patch bleeding, small amount of bleeding or large amount of bleeding; the label values of the mucus feature group include yellowish mucus, thick mucus, bile reflux mucus, clear mucus or mucus lake; and the label values of the local lesion feature group include local concave lesion, local raised lesion or local flat lesion.
[0033] To achieve one object of the present application, in some embodiments of the present application, the first multi-task classification neural network comprises a set of main feature extraction networks for identifying a preset n different categories of endoscopic feature groups and n sets of branch feature extraction networks respectively for identifying the label values in the endoscopic feature groups downstream of the main feature extraction networks. In one embodiment of the present application, n = 5. The construction method of the first multi-task classification neural network comprises:
[0034] Step A) The main feature extraction network is composed of at least 3 sets (for example, 4 sets, 6 sets, 8 sets, 9 sets, 10 sets) of multi-branch parallel feature extraction units, and after operation, a first main feature extraction network feature map is obtained;
[0035] Step B) At least 3 sets (for example, 4 sets, 5 sets, 6 sets, 8 sets) of multi-branch parallel feature extraction units are used to respectively construct the nth branch feature extraction network; the first main feature extraction network feature map is respectively input, and after multi-branch parallel convolution operation, the nth branch feature extraction network feature map is finally obtained respectively;
[0036] Wherein, the design method of the multi-branch parallel convolution operation unit is:
[0037] a) A slicing unit is constructed for dividing the original input feature map into four groups to obtain a first sliced feature group, a second sliced feature group, a third sliced feature group, and a fourth sliced feature group;
[0038] b) A grouped feature extraction unit is constructed by 3 depth separable convolutions for feature extraction of the first, second, and third sliced feature groups, wherein the sizes of the 3 depth separable convolutions are 3*3, 1*13, and 13*1 respectively for extracting spatial features, vertical channel features, and horizontal channel features. Finally, the extracted spatial features, vertical channel features, and horizontal channel features are spliced with the fourth sliced feature group, and after passing through a regularization layer, a first output is obtained;
[0039] c) A feature enhancement unit is constructed, which is composed of a 1*1 convolution layer, an activation function layer, a 1*1 convolution layer, and a residual connection. By connecting 1*1 convolution, activation function, and 1*1 convolution in series, a second output is obtained. The residual connection is used to splice the original input feature map and the second output in the horizontal direction to obtain a third output, which is the output result of the multi-branch parallel convolution operation unit;
[0040] Wherein, the 1*1 convolution layer is composed of conventional technical means in the art, and the activation function is GeLU;
[0041] Step C) constructing the first, second, third, fourth and fifth prediction layers using a full connection layer and a Sigmoid function as a combination unit, specifically, inputting the first, second, third, fourth and fifth branch feature maps into the full connection layer, the network layer compresses the high-dimensional feature map into a 1-dimensional vector form, and finally obtains the first, second, third, fourth and fifth feature group label recognition results through Sigmoid function operation, wherein the Sigmoid operation formula is:
[0042]
[0043] Wherein, p k refers to the probability of solving the class k, is a trainable hyperparameter for class k, b k is a trainable bias hyperparameter for class k.
[0044] In some embodiments of the present application, the training set is input into the first multi-task classification neural network for training in batches.
[0045] The training set is obtained by collecting gastrointestinal picture data according to a plurality of preset multi-label feature groups, performing label annotation on each picture to obtain a first feature label dataset, and splitting the first feature label dataset into a training set and a validation set according to a ratio of 8:2. The feature groups and corresponding label values are shown in Table 1.
[0046] Table 1
[0047]
[0048] In order to achieve one object of the present application, the method for training the first multi-task classification neural network comprises:
[0049] Step D) predicting the original label value corresponding to each of the gastrointestinal feature groups, calculating using a BCE cross-entropy loss function to obtain the nth loss function, and calculating the total network loss using a weighted summation method, and the specific formula is:
[0050]
[0051] Wherein, k is the number of feature group branches; L k is the loss derivative of the k feature branch; N is the number of pictures in the training batch; i is the corresponding class number; y i is the true label value of the i class in the picture, that is, if the picture is of the i class, then y i is 1, if not, then y i is 0; p i is the prediction confidence of the picture for the i class.
[0052] L = w1 x L1 + w2 x L2 +... + w k x L k
[0053] Wherein, L is total loss, w is loss function weight; total loss is obtained by summation;
[0054] The adjustment strategy for the loss function weight in the application is stage decay, that is, the fixed weight at the beginning of training is [2.5, 2.5, 2.5, 2.5, 2.5], and it is adjusted once every 10 rounds, and the adjustment basis is: if L t < L (t-10) , then w t = w t x θ; if L t > L (t-10) , then w t = w t x (2- θ). Wherein, the θ is the decay coefficient, θ = 0.85;
[0055] Step E) update the neural network model parameters using the back propagation technique according to the value of the total loss of the network, and stop training when the final loss function is stable to the preset value or reaches the set maximum iteration number, to obtain the pre-trained first multi-task classification neural network.
[0056] The constructed and trained first multi-task classification neural network is used to identify and predict the lesion features under the gastrointestinal endoscope. 1) The frame-by-frame images obtained by the image acquisition module are interval sampled at a sampling rate of 1 second / time to obtain the images to be identified, and the frame number of the images to be identified is recorded; 2) the first multi-task classification neural network is used to identify a plurality of groups of multi-label features, and the first mucosa feature group identification result, the first blood vessel texture feature group identification result, the first bleeding feature group identification result, the first mucus feature group identification result and the first local lesion feature group identification result are obtained.
[0057] The first mucosa feature group identification result, the first blood vessel texture feature group identification result, the first bleeding feature group identification result, the first mucus feature group identification result and the first local lesion feature group identification result obtained above are input into the feature result collection module for suitable purposes, including but not limited to, auxiliary diagnosis of digestive tract diseases.
[0058] To achieve one object of the present application, in one embodiment of the present application, the downstream network of the local lesion feature group can further comprise a local lesion multi-level feature recognition and training module, which comprises a second multi-task classification neural network for identifying m different categories of preset local lesion features and the corresponding supplementary identification values of the categories, wherein 1≤m≤5. In some embodiments of the present application, the m can be 1, 2, 3, 4 or 5. In one embodiment of the present application, the m=4. That is, the local lesion features comprise 4 different categories, such as lesion morphology features, lesion area features, lesion surface features or lesion edge features.
[0059] Further, S4. Local lesion multi-level feature recognition and training module.
[0060] The local lesion multi-level feature recognition and training module comprises a second multi-task classification neural network for identifying 4 different categories of preset local lesion features and the corresponding supplementary identification values of the categories. The second multi-task classification neural network can be obtained by using the same construction and training method as the first multi-task classification neural network. In some embodiments of the present application, the construction method of the second multi-task classification neural network comprises:
[0061] Step F) constructing a lesion segmentation artificial intelligence neural network based on a DeepLabV3+ algorithm to predict the images of the local lesion feature group and obtain a first lesion segmentation map. Specifically,
[0062] i) constructing a required data set. Contour line labeling is performed on lesions, and the targets include raised lesion targets and erosion lesion targets to obtain a first segmented lesion data set;
[0063] Secondly, the lesion multi-level labels are marked to construct a second label feature data set
[0064] ii) training a lesion segmentation model according to the first segmented lesion data set to obtain a first lesion segmentation model, wherein the input data is an image and the output is a mask map containing only a segmented area.
[0065] Step G) a second multi-task classification neural network algorithm for predicting lesion multi-level feature label prediction. The first lesion segmentation map is input to obtain first lesion morphology feature group, first lesion area feature group, first lesion surface feature group and first lesion edge feature group prediction results. Specifically,
[0066] i) using the same construction method as the first multi-task classification neural network to construct the network, and since there are only 4 feature groups, the number of branches is set to 4, and therefore the number of subsequent loss functions is also 4.
[0067] ii) training the pre-constructed second multi-task classification neural network algorithm to obtain the label recognition capability of the lesion feature group, that is,
[0068] When the recognition label in the first lesion feature recognition group contains "local erosive lesion" and "local raised lesion", the pre-constructed second multi-task classification neural network algorithm is used to recognize the supplementary label value of the feature group. The first supplementary lesion morphology feature group, the first supplementary lesion area feature group, the first supplementary lesion surface feature group, and the first supplementary lesion edge feature group are obtained.
[0069] 1) Among them, the first supplementary lesion morphology feature group, the label corresponding to "local erosive lesion" includes red swelling and non-red swelling; and the label corresponding to "local raised lesion" includes flat, lateral, pedunculated, and irregular;
[0070] 2) Among them, the first supplementary lesion area feature group, the label corresponding to "local erosive lesion" includes point, shallow small, and deep large; and the label corresponding to "local raised lesion" includes multiple and single;
[0071] 3) Among them, the first supplementary lesion surface feature group, the label corresponding to "local erosive lesion" includes no white fur, gray-white, and yellow-white; and the label corresponding to "local raised lesion" includes light white, light pink, and deep red;
[0072] 4) Among them, the first supplementary lesion edge feature group, the label corresponding to "local erosive lesion" includes smooth and burr; and the label corresponding to "local raised lesion" includes smooth and rough;
[0073] Step H) Constructing a label conversion module to fuse the supplementary label value obtained in step G) with the label value of the local lesion feature group to obtain the label value of the new local lesion feature group (second local lesion feature group).
[0074] The above-mentioned second local lesion feature group recognition result can be input into the feature result collection module together with the above-mentioned first mucosa feature group recognition result, first blood vessel texture feature group recognition result, first bleeding feature group recognition result, and first mucus feature group recognition result for suitable purposes, including but not limited to, auxiliary diagnosis of digestive tract diseases.
[0075] In order to achieve one of the purposes of the present application, in an embodiment of the present application, the above-mentioned digestive tract endoscopic lesion feature recognition system further comprises a mucosa position recognition module for converting the label value in the mucosa feature group according to the mucosa position and outputting the recognition result to the feature result collection module. Among them, the position includes esophagus, gastric body, gastric fundus, gastric antrum, duodenum, cecum, ascending colon, transverse colon, descending colon, sigmoid colon, or rectum.
[0076] Further, S3. Mucosa position recognition module.
[0077] The method for constructing the mucosal site recognition module comprises:
[0078] Step I) using a time sequence segment classification model to classify and recognize the site position of the current video / picture to obtain a position recognition result.
[0079] i) collecting any suitable time length video segment in a digestive tract endoscopy;
[0080] ii) data preprocessing, using interval sampling technology to reduce the dimension and scale the image size to a suitable size according to the total number of frames of the collected time length;
[0081] iii) using ResNet-3D to input and train the preprocessed data, updating the parameters using the back propagation technology during the process, obtaining a preset time sequence segment classification model, and making it have the position recognition ability.
[0082] Step J) converting the feature recognition result of the first mucosal feature group image obtained by the digestive tract endoscopic multi-level feature recognition and training module after sampling for recognition, specifically,
[0083] i) accumulating the images read frame by frame, and using the above preprocessing module to process the segment set after reaching a suitable time length;
[0084] ii) using the preset time sequence segment classification model to predict the position to obtain the frame number range and the corresponding prediction result of the position;
[0085] iii) according to the frame number range, searching the first mucosal feature group corresponding to the frame number in the interval in reverse, if the position corresponding to the frame number range interval is “gastric body”, “gastric fundus” and “gastric antrum”, and the first feature group in the interval contains the “intestinal mucosa” label, remove the “intestinal mucosa” label in the first mucosal feature group and add the “intestinal metaplasia” label.
[0086] iv) according to the frame segment position corresponding recognition result, adding the label mucosal site recognition module output result to the first mucosal feature group to obtain the second mucosal feature group recognition result.
[0087] The above-mentioned second mucosal feature group recognition result can be input into the feature result collection module together with the above-mentioned first blood vessel texture feature group recognition result, first bleeding feature group recognition result, first mucus feature group recognition result and first local lesion feature group recognition result for suitable purposes, including but not limited to, assisting in the diagnosis of digestive tract diseases.
[0088] In some embodiments of the present application, the recognition results outputted by the multi-level feature recognition and training module under the endoscope can include: the first / second mucosa feature group recognition result, the first blood vessel texture feature group recognition result, the first bleeding feature group recognition result, the first mucus feature group recognition result, and the first / second local lesion feature group recognition result, which are inputted into the feature result collection module for suitable purposes, including but not limited to, the auxiliary diagnosis of digestive tract diseases.
[0089] In order to achieve one of the objectives of the present application, in one embodiment of the present application, the lesion feature recognition system under the endoscope further comprises a disease diagnosis decision module for obtaining a disease diagnosis result by using the recognition result.
[0090] Further, S5. the disease diagnosis decision module.
[0091] The construction method of the disease diagnosis decision module comprises:
[0092] Step K) performing disease diagnosis annotation on the medical images in the first label data set to obtain a first disease data set, the input of which is, for example, the second mucosa feature group recognition result, the first blood vessel texture feature group recognition result, the first bleeding feature group recognition result, the first mucus feature group recognition result, and the second local lesion feature group recognition result (i.e. label value), and the output of which is the disease category. In some embodiments of the present application, the disease category includes but is not limited to reflux esophagitis, gastric esophageal varices, and ulcerative colitis. In some other embodiments of the present application, the output can also be the nature of the disease, for example, including but not limited to local high-risk lesions, local low-risk lesions, diffuse lesions, and typing.
[0093] Step L) constructing a pre-set disease diagnosis decision algorithm, specifically, annotating the diseases contained in the medical images and iteratively training the data in batches based on the XGBOOST algorithm to obtain a pre-constructed disease diagnosis decision algorithm.
[0094] Step M) the disease diagnosis decision module can further comprise a preprocessing module for aligning the second mucosa feature group recognition result, the first blood vessel texture feature group recognition result, the first bleeding feature group recognition result, the first mucus feature group recognition result, and the second local lesion feature group recognition result (i.e. label value) to obtain an aggregated feature set.
[0095] Step N) inputting the feature set into the pre-constructed disease diagnosis decision algorithm and performing algorithm reasoning to obtain a disease diagnosis result.
[0096] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, several improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as falling within the protection scope of the present application.
Claims
1. A system for recognizing features of lesions under a digestive tract endoscope, characterized by, The system comprises: 1) an image acquisition module; and 2) a multi-level feature recognition and training module under a digestive tract endoscope, comprising a first multi-task classification neural network for recognizing a preset n different categories of feature groups under a digestive tract endoscope and label values in the feature groups, wherein 1≤n≤10; and 3) a feature result collection module for collecting the recognition results output by the multi-level feature recognition and training module under the digestive tract endoscope; Wherein the feature groups under the digestive tract endoscope include a local lesion feature group, and the label values of the local lesion feature group include a local concave lesion, a local raised lesion or a local flat lesion; the feature groups under the digestive tract endoscope further include at least one selected from a mucosa feature group, a blood vessel texture feature group, a bleeding feature group or a mucus feature group; The downstream network of the local lesion feature group comprises a multi-level feature recognition and training module for local lesions, which comprises a second multi-task classification neural network for recognizing a preset m different categories of local lesion features and corresponding supplementary label values, wherein 1≤m≤5. 2.The system according to claim 1, wherein, The image acquisition module comprises a video acquisition module and / or a picture acquisition module; Wherein the video acquisition module is used to read the video signal output by the endoscope device to obtain video frame images; The video acquisition module performs interval sampling on the obtained frame-by-frame images to obtain images to be recognized and records the frame sequence numbers of the images to be recognized. 3.The system according to claim 1, characterized in that, The label values of the mucosa feature group include thinned mucosa, mucosa strip redness, mucosa sheet redness, mucosa edema, island-shaped mucosa, residual glandular mucosa, gastric mucosa, intestinal mucosa or intestinal metaplasia; the label values of the blood vessel texture feature group include blood vessel non-transparency, strip-shaped blood vessels, dendritic blood vessels or point and sheet-shaped blood vessels; the label values of the bleeding feature group include no bleeding, point and sheet-shaped bleeding, small amount of bleeding or large amount of bleeding; and the label values of the mucus feature group include yellowish mucus, thick mucus, bile reflux mucus, clear mucus or mucus lake. 4.The system of claim 1, wherein, The first multi-task classification neural network comprises a group of main feature extraction networks for recognizing a preset n different categories of feature groups under a digestive tract endoscope and n groups of branch feature extraction networks located downstream of the main feature extraction networks for recognizing label values in the feature groups under the digestive tract endoscope, and the construction method of the first multi-task classification neural network comprises: Step A) The main feature extraction network is composed of at least 3 groups of multi-branch parallel feature extraction units, and after operation, a first main feature extraction network feature map is obtained; Step B) The nth branch feature extraction network is composed of at least 3 groups of multi-branch parallel feature extraction units; the first main feature extraction network feature map is input respectively, and after multi-branch parallel convolution operation, the nth branch feature extraction network feature map is finally obtained respectively; Wherein, the design method of the multi-branch parallel convolution operation unit is: a) Construct a slicing unit for dividing the original input feature map into four groups to obtain a first sliced feature group, a second sliced feature group, a third sliced feature group and a fourth sliced feature group; b) constructing a group feature extraction unit by 3 depth separable convolutions, respectively extracting features of the first slice feature group, the second slice feature group and the third slice feature group, wherein the sizes of the 3 depth separable convolutions are 3 3, 1 13, 13 1, for extracting spatial features, vertical channel features and horizontal channel features, and finally splicing the extracted spatial features, vertical channel features and horizontal channel features with the fourth slice feature group, passing through a regularization layer to obtain a first output; c) a feature enhancement unit constructed by one 1 1 convolutional layer, an activation function layer, one 1 1 convolutional layer and one residual connection, by 1 1 convolution, an activation function, 1 1 convolution in series operation to obtain a second output, using the residual connection to splice the original input feature map and the second output in the horizontal direction feature map to obtain a third output, which is the output result of the multi-branch parallel convolution operation unit. Step C) constructing the nth prediction layer using a full connection layer and a Sigmoid function as a combination unit, and obtaining a label recognition result of the nth group of endoscopic features through a Sigmoid function operation, wherein the Sigmoid operation formula is: ; wherein, denotes the probability of solving the class k, is a trainable hyperparameter for class k, is a trainable bias hyperparameter for class k.
5. The system according to claim 1 or 4, wherein The training method of the first multi-task classification neural network comprises inputting a training set into the first multi-task classification neural network in batches for training, comprising: Step D) predicting the original label value corresponding to each of the groups of endoscopic features, calculating a BCE cross-entropy loss function, obtaining the nth group of loss functions, and calculating the total network loss by using a weighted summation method, and the specific formula is: ; wherein k is the number of characteristic component branches; is the loss derivative of k characteristic branches; N is the number of pictures in the training batch; i is the corresponding category number; is the real label value of the i category in the picture, that is, if the picture is the i category, is 1, not the i category, then is 0; is the prediction confidence of the picture on the i category; ......+ ; Wherein, L is the total loss, w is the loss function weight; the total loss is obtained by summation; Step E) updating the neural network model parameters by using the back propagation technology according to the value of the total network loss, and stopping training when the final loss function is stable to a preset value or reaches a set maximum iteration number, and obtaining the first multi-task classification neural network pre-trained. 6.The system of claim 3, wherein, The endoscopic lesion feature recognition system further comprises a mucosa position recognition module for converting the label values in the mucosa feature group according to the mucosa position and outputting the recognition result to the feature result collection module. Wherein, the positions include esophagus, stomach body, stomach bottom, gastric antrum, duodenum, cecum, ascending colon, transverse colon, descending colon, sigmoid colon or rectum. 7.The system of claim 1, wherein, The local lesion features include at least one selected from lesion morphology features, lesion area features, lesion surface features or lesion edge features.
8. The system according to claim 1 or 7, wherein The construction method of the second multi-task classification neural network comprises: Step F) constructing a local lesion segmentation artificial intelligence neural network, which is used for classifying and predicting the input image of the local lesion feature group to obtain a first lesion segmentation map; Step G) constructing a local feature multi-label recognition network based on segmentation region, which is used for identifying m different categories of the local lesion features and the corresponding supplementary identification values of the categories; Step H) constructing a label conversion module, which is used for fusing the supplementary identification values obtained in step G) with the label values of the local lesion feature group to obtain new label values of the local lesion feature group.
9. An electronic device, comprising: The electronic device comprises an electronic device body and an endoscopic lesion feature recognition system as claimed in any one of claims 1-8, and the endoscopic lesion feature recognition system is installed on the electronic device body.
Citation Information
Patent Citations
Digestive tract disease auxiliary diagnosis system based on deep learning
CN111128396A