Laryngoscope image recognition method and system based on multi-feature extraction
Through the laryngoscopic image recognition method based on multi-feature extraction, the transparency and interpretability problems of artificial intelligence-assisted diagnosis and treatment system in the diagnosis of laryngeal cancer are solved, and the diagnosis results with high accuracy and transparency are achieved, enhancing the credibility in clinical practice.
Patent Information
- Application Number
- CN202510107624.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the prior art, artificial intelligence-assisted diagnosis and treatment systems have problems with transparency and interpretability in the diagnosis of laryngeal cancer, resulting in low credibility and acceptability in clinical practice.
A laryngoscopic image recognition method based on multi-feature extraction is proposed. By acquiring the patient's throat image and preprocessing, morphological features and quantitative features are extracted using the deep learning feature extraction model, multi-scale feature fusion is performed, and human-computer interaction visualization is performed to construct an interpretable diagnostic result.
It improves the accuracy and transparency of laryngeal cancer diagnosis, provides a diagnostic basis through interpretable feature extraction and fusion, and enhances physicians' trust in the diagnosis results.
Smart Images

Figure CN120125942A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical image recognition, and in particular, to a laryngoscope image recognition method and system based on multi-feature extraction. Background Art
[0002] Electronic laryngoscopes are widely used in the detection of early laryngeal cancer. Narrow-band imaging technology can enhance the contrast of blood vessels on the mucosal surface, better describe tumor-specific neovascularization, and detect superficial and abnormal mucosal lesions that may be missed by conventional white-light imaging. However, there are significant skill differences among endoscopists in detecting suspicious lesions, resulting in a relatively high missed diagnosis rate of laryngeal cancer, seriously threatening the lives of patients. Improving the diagnostic accuracy of early laryngeal cancer under endoscopy has great clinical significance.
[0003] With the continuous progress of artificial intelligence technology, deep learning neural networks have shown great potential in assisting doctors in disease screening and diagnosis. They can not only effectively make up for the lack of experience but also significantly improve the accuracy and efficiency of diagnosis. Combining deep learning technology with an electronic laryngoscope system can provide more accurate and automated auxiliary diagnosis and treatment for laryngeal diseases. However, the current research on artificial intelligence-assisted diagnosis and treatment is mainly based on end-to-end deep learning algorithms, which only output diagnostic conclusions. The diagnostic process is an opaque and incomprehensible "black box" that does not explain the decision-making process and diagnostic basis, greatly affecting the credibility and acceptability of artificial intelligence systems in clinical practice.
[0004] In summary, the technical problems existing in the related technologies need to be improved. Summary of the Invention
[0005] The main purpose of the embodiments of this application is to propose a laryngoscope image recognition method and system based on multi-feature extraction, which can perform secondary judgment on the combined local lesion features identified, improving the accuracy and transparency of the lesion diagnosis results.
[0006] To achieve the above object, on the one hand, an embodiment of this application proposes a laryngoscope image recognition method based on multi-feature extraction, and the method includes:
[0007] Obtain a patient's laryngeal image and perform preprocessing on the image data to obtain a preprocessed patient's laryngeal image;
[0008] Perform multi-feature extraction processing on the preprocessed patient's laryngeal image based on a deep learning feature extraction model to obtain the morphological features and quantitative features of the patient's laryngeal image;
[0009] Perform multi-scale feature fusion processing on the morphological features and quantitative features of the patient's laryngeal image to obtain the key features of the patient's laryngeal image with feature weights;
[0010] Perform human-computer interaction visualization processing on the key features of the patient's larynx with feature weights to construct the recognition and diagnosis results of the patient's larynx.
[0011] In some embodiments, the obtaining of the patient's larynx image and performing preprocessing on the image data to obtain the preprocessed patient's larynx image includes:
[0012] Obtain the patient's larynx image through an electronic laryngoscope device;
[0013] Extract the clear key frames of the patient's larynx image and filter the blurred frames and non-target regions of the patient's larynx image to construct the coarsely screened patient's larynx image;
[0014] Perform accuracy detection on the coarsely screened patient's larynx image to obtain the standard patient's larynx image;
[0015] Perform image enhancement and formatting processing on the standard patient's larynx image to obtain the preprocessed patient's larynx image.
[0016] In some embodiments, the performing of accuracy detection on the coarsely screened patient's larynx image to obtain the standard patient's larynx image includes:
[0017] Construct an image sharpness evaluation function based on the Laplace operator, and perform image blur detection processing on the coarsely screened patient's larynx image to obtain a clear patient's larynx image;
[0018] Calculate the proportion of effective matching points of the clear patient's larynx image based on the ORB algorithm to obtain an image repeatability score;
[0019] Based on a preset score threshold, combine the image repeatability score to perform image repeatability detection on the clear patient's larynx image to obtain a filtered patient's larynx image;
[0020] Perform detection of key structural features of the larynx on the filtered patient's larynx image through a target detection algorithm to obtain the standard patient's larynx image.
[0021] In some embodiments, the performing of multi-feature extraction processing on the preprocessed patient's larynx image based on a deep learning feature extraction model to obtain the morphological features of the patient's larynx image and the quantitative features of the patient's larynx image includes:
[0022] Performing feature extraction processing on the preprocessed laryngeal images of the patient through a Transformer network model based on a cross-fusion encoder to obtain the morphological features of the first laryngeal images of the patient, where the morphological features of the first laryngeal images of the patient include texture features, color features, boundary features, blood vessel morphological features, blood vessel color features, blood vessel orientation features, and mucosal color features;
[0023] Performing feature extraction processing on the preprocessed laryngeal images of the patient through a ResNet-50 deep residual network model to obtain the morphological features of the second laryngeal images of the patient, where the morphological features of the second laryngeal images of the patient represent lesion location features;
[0024] Combining the morphological features of the first laryngeal images of the patient with the morphological features of the second laryngeal images of the patient to obtain the morphological features of the laryngeal images of the patient;
[0025] Performing lesion image analysis on the preprocessed laryngeal images of the patient to obtain the quantitative features of the laryngeal images of the patient, where the quantitative features of the laryngeal images of the patient include aspect ratio information, color spectrum information, S-channel image entropy information, texture information, histogram of oriented gradients information, and color moment information.
[0026] In some embodiments, the performing feature extraction processing on the preprocessed laryngeal images of the patient through a Transformer network model based on a cross-fusion encoder to obtain the morphological features of the first laryngeal images of the patient includes:
[0027] Inputting the preprocessed laryngeal images of the patient into the Transformer network model based on a cross-fusion encoder, where the Transformer network model based on a cross-fusion encoder includes a plurality of image feature extraction modules and a linear layer, and the plurality of image feature extraction modules include a patch embedding layer and a cross-fusion encoder module;
[0028] Based on the patch embedding layer of the Transformer network model based on a cross-fusion encoder, performing dimensionality reduction processing on the preprocessed laryngeal images of the patient to obtain the laryngeal images of the patient after dimensionality reduction;
[0029] Based on the cross-fusion encoder module of the Transformer network model based on a cross-fusion encoder, performing image feature extraction processing on the laryngeal images of the patient after dimensionality reduction to obtain the preliminary morphological features of the first laryngeal images of the patient;
[0030] Based on the linear layer of the Transformer network model based on a cross-fusion encoder, performing image feature category prediction processing on the preliminary morphological features of the first laryngeal images of the patient to obtain the morphological features of the first laryngeal images of the patient.
[0031] In some embodiments, the cross - fusion encoder module of the Transformer network model based on the cross - fusion encoder performs image feature extraction processing on the dimension - reduced patient laryngeal image to obtain the morphological features of the preliminary first patient laryngeal image, including:
[0032] Input the dimension - reduced patient laryngeal image into the cross - fusion encoder module of the Transformer network model of the cross - fusion encoder. The cross - fusion encoder module includes a first normalization layer, a parallel attention layer, a second normalization layer, and a multi - layer perceptron;
[0033] Based on the first normalization layer and the second normalization layer of the cross - fusion encoder module, perform feature extraction on the dimension - reduced patient laryngeal image to obtain a patient laryngeal feature image;
[0034] Based on the parallel attention layer of the cross - fusion encoder module, divide the patient laryngeal feature image into global feature information and fine - grained features and perform cross - splicing to obtain the spliced patient laryngeal image features;
[0035] Based on the multi - layer perceptron of the cross - fusion encoder module, perform perceptual output on the spliced patient laryngeal image features to obtain the morphological features of the preliminary first patient laryngeal image.
[0036] In some embodiments, the multi - scale feature fusion processing of the morphological features and the quantitative features of the patient laryngeal image to obtain the key features of the patient laryngeal with feature weights includes:
[0037] Fuse the morphological features and the quantitative features of the patient laryngeal image to obtain multi - scale feature data of the patient laryngeal image;
[0038] Based on the random forest model, use the method of sampling with replacement to repeatedly and randomly select the multi - scale feature data of the patient laryngeal image to construct a training sample subset;
[0039] Train multiple decision trees based on the training sample subset. Each tree randomly selects features for splitting during the training process to construct a random forest model;
[0040] Based on the random forest model, perform voting decision on the multi - scale feature data of the patient laryngeal image to obtain the key features of the patient laryngeal with feature weights.
[0041] In some embodiments, the expression of the voting decision of the random forest model is specifically as follows:
[0042]
[0043] In the above formula, represents the final classification result, and T m (x) represents the classification result of the m-th tree for the sample x, C represents the number of sample categories, c represents the sample category, and ∥(·) represents the indicator function.
[0044] In some embodiments, the man-machine interaction visualization processing of the key features of the patient's larynx with feature weights to construct the recognition and diagnosis result of the patient's larynx includes:
[0045] Outputting interpretable information according to the key features of the patient's larynx with feature weights, where the interpretable information includes texture information with weight ratio, mass boundary information with weight ratio, blood vessel morphology information with weight ratio, S-channel image entropy information with weight ratio, blood vessel color information with weight ratio, and blood vessel direction information with weight ratio;
[0046] Performing visual saliency annotation processing on the key features of the patient's larynx with feature weights according to the interpretable information to obtain the lesion area of the patient's larynx;
[0047] Performing confirmation feedback on the lesion area of the patient's larynx to obtain the recognition and diagnosis result of the patient's larynx.
[0048] To achieve the above object, on the other hand, an endoscope image recognition system based on multi-feature extraction is proposed in an embodiment of the present application. The system includes:
[0049] A first module for acquiring a patient's larynx image and performing preprocessing on the image data to obtain a preprocessed patient's larynx image;
[0050] A second module for performing multi-feature extraction processing on the preprocessed patient's larynx image based on a deep learning feature extraction model to obtain the morphological features and quantitative features of the patient's larynx image;
[0051] A third module for performing multi-scale feature fusion processing on the morphological features and quantitative features of the patient's larynx image to obtain the key features of the patient's larynx with feature weights;
[0052] A fourth module for performing man-machine interaction visualization processing on the key features of the patient's larynx with feature weights to construct the recognition and diagnosis result of the patient's larynx.
[0053] The embodiments of the present application at least include the following beneficial effects: The present application provides a laryngoscope image recognition method and system based on multi-feature extraction. This solution preprocesses the image data by acquiring the laryngeal image of the patient, extracts the local lesion area that meets clinical diagnosis, and further performs multi-feature extraction processing through a deep learning feature extraction model to obtain the morphological features and quantitative features of the patient's laryngeal image. Then, multi-scale feature fusion processing is performed to output the key features of the patient's larynx with feature weights. By making a secondary judgment on the identified combination of local lesion features, it can guide clinical decisions, obtain a more accurate lesion diagnosis result, reduce the misdiagnosis rate, and finally perform human-computer interaction visualization processing. An interpretable artificial intelligence system is used to diagnose early laryngeal cancer, increasing the transparency of the diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of a laryngoscope image recognition method based on multi-feature extraction provided by an embodiment of the present application;
[0055] Figure 2 is a schematic structural diagram of a laryngoscope image recognition system based on multi-feature extraction provided by an embodiment of the present application;
[0056] Figure 3 is a schematic diagram of preprocessing the image data of the patient's laryngeal image provided by an embodiment of the present application;
[0057] Figure 4 is a schematic diagram of performing multi-feature extraction processing on the preprocessed patient's laryngeal image provided by an embodiment of the present application;
[0058] Figure 5 is a schematic diagram of multi-scale feature fusion decision-making provided by an embodiment of the present application;
[0059] Figure 6 is a schematic diagram of human-computer interaction online learning optimization provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods that are consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0061] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, as used herein, the words "if", "when" may be interpreted as "when...", "when...", or "in response to determining".
[0062] The terms "at least one", "a plurality", "each", "any one", etc. used in this application, at least one includes one, two or more than two, a plurality includes two or more than two, each refers to each of the corresponding plurality, and any one refers to any one of the plurality.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0064] Refer to Figure 1 , Figure 1 is a flowchart of a laryngoscope image recognition method based on multi-feature extraction provided by an embodiment of the present invention. Refer to Figure 1 , the method includes the following steps:
[0065] S100. Obtain a patient's laryngeal image and perform preprocessing on the image data to obtain a preprocessed patient's laryngeal image;
[0066] Specifically, obtain a laryngeal image in real time through an electronic laryngoscope device, automatically extract clear key frames, filter out blurred frames and non-target areas, perform further enhancement processing on the filtered image, and finally convert the collected image into a standardized format.
[0067] Furthermore, it should be noted that in some embodiments, step S100 may include: S110. Obtain a patient's laryngeal image through an electronic laryngoscope device;
[0068] S120. Extract clear key frames of the patient's laryngeal image and filter out blurred frames and non-target areas of the patient's laryngeal image to construct a roughly screened patient's laryngeal image;
[0069] S130. Perform accuracy detection on the roughly screened patient's laryngeal image to obtain a standard patient's laryngeal image;
[0070] Further, it should also be noted that in some embodiments, step S130 may include: S131, constructing an image sharpness evaluation function based on the Laplace operator, performing image blur detection processing on the coarsely screened patient laryngeal images to obtain clear patient laryngeal images; S132, calculating the proportion of effective matching points of the clear patient laryngeal images based on the ORB algorithm to obtain an image repeatability score; S133, performing image repeatability detection on the clear patient laryngeal images in combination with the image repeatability score based on a preset score threshold to obtain filtered patient laryngeal images; S134, detecting the key structural features of the larynx on the filtered patient laryngeal images through a target detection algorithm to obtain standard patient laryngeal images.
[0071] S140, performing image enhancement and formatting processing on the standard patient laryngeal images to obtain preprocessed patient laryngeal images.
[0072] In some specific embodiments, as Figure 3 shown, real-time image information of the patient's larynx is acquired, clear key frames are automatically extracted, and blurred frames and non-target regions are filtered. Further enhancement processing is performed on the screened images, and finally the collected laryngeal images are standardized. The Laplace operator is used as the image sharpness evaluation function, and a blur determination interval (a, b) is set according to the quality requirements. Since the larger the gradient value, the clearer the image, only the images with a returned gradient value greater than b are retained. The image repeatability score - that is, the proportion of effective matching points, is calculated based on the ORB (Oriented FAST and Rotated BRIEF) algorithm. 0 indicates no effective matching points (completely non-repetitive), and 1 indicates that all matching points are verified as effective (completely repetitive). According to the score, a threshold is set. Above 0.6 indicates that the repetitive region is significant, and they are filtered. At the same time, a target detection algorithm is used to detect the existence of the key structures of the larynx to ensure that the target region is included in the image. Finally, the processed image is denoised, the brightness is adjusted, and the contrast is enhanced, and the image is standardized to a unified size and format as the input to adapt to the deep learning model.
[0073] S200, performing multi-feature extraction processing on the preprocessed patient laryngeal images based on a deep learning feature extraction model to obtain the morphological features and quantitative features of the patient laryngeal images;
[0074] Specifically, the original laryngoscope images are preprocessed and standardized to a unified size and format, and then input into a deep learning model for feature extraction. Fourteen features related to laryngeal diseases are extracted from the laryngoscope images based on in-depth literature research and expert experience, including eight deep learning-based features (texture, color, boundary, blood vessel morphology, blood vessel color, blood vessel orientation, mucosal color, lesion location) and six quantitative features (aspect ratio, color spectral information, entropy of the S-channel image, texture information, histogram of oriented gradients, color moment). The texture, color, boundary, blood vessel morphology, blood vessel color, blood vessel orientation, and mucosal color are obtained by constructing a multi-task classifier using a Transformer network based on cross-fusion encoders; the lesion location is extracted using a classic ResNet-50 deep residual network, which is used to identify laryngeal anatomical features and locate the lesion.
[0075] It should be noted that in some embodiments, step S200 may include: S210, performing feature extraction processing on the preprocessed laryngeal image of the patient through a Transformer network model based on cross-fusion encoders to obtain the morphological features of the first laryngeal image of the patient, and the morphological features of the first laryngeal image of the patient include texture features, color features, boundary features, blood vessel morphology features, blood vessel color features, blood vessel orientation features, and mucosal color features;
[0076] Furthermore, it should be noted that in some embodiments, step S210 may include: S211, inputting the preprocessed laryngeal image of the patient into a Transformer network model based on cross-fusion encoders, and the Transformer network model based on cross-fusion encoders includes several image feature extraction modules and a linear layer, and several image feature extraction modules include a patch embedding layer and a cross-fusion encoder module; S212, performing dimensionality reduction processing on the preprocessed laryngeal image of the patient based on the patch embedding layer of the Transformer network model based on cross-fusion encoders to obtain the laryngeal image of the patient after dimensionality reduction; S213, performing image feature extraction processing on the laryngeal image of the patient after dimensionality reduction based on the cross-fusion encoder module of the Transformer network model based on cross-fusion encoders to obtain the preliminary morphological features of the first laryngeal image of the patient;
[0077] Specifically, the dimension-reduced laryngeal image of the patient is input into the cross-fusion encoder module of the Transformer network model of the cross-fusion encoder. The cross-fusion encoder module includes a first normalization layer, a parallel attention layer, a second normalization layer, and a multi-layer perceptron. Based on the first normalization layer and the second normalization layer of the cross-fusion encoder module, feature extraction is performed on the dimension-reduced laryngeal image of the patient to obtain a laryngeal feature image of the patient. Based on the parallel attention layer of the cross-fusion encoder module, the laryngeal feature image of the patient is divided into global feature information and fine-grained features and cross-stitched to obtain the stitched laryngeal image feature of the patient. Based on the multi-layer perceptron of the cross-fusion encoder module, perceptual output is performed on the stitched laryngeal image feature of the patient to obtain the morphological feature of the initial first laryngeal image of the patient.
[0078] In this embodiment, as Figure 4 shown, based on the Transformer model, a feature extraction network with a cross-fusion encoder module is designed. This network contains 4 stages, and each stage stacks a patch embedding layer (PatchEmbed) and an encoder module. Finally, the image category is predicted through a linear layer (LN), which can be expressed as:
[0079]
[0080] In the above formula, represents the input image, and H, W, and C represent the width, height, and number of channels of the input image respectively. In the first stage, the embedding layer first reduces the resolution of the input image to H 1 = H / N and W 1 = W / N. Based on empirical analysis, N = 4 is set, and the number of channels is increased to C1. Attention calculation is performed through the encoder module to extract image features. In each subsequent stage, a non-overlapping 2×2 convolutional kernel is used for spatial-to-depth operation to increase the dimension of the feature map and reduce the resolution. Then, a cross-encoder module composed of two sub-networks, namely parallel attention and multi-layer perceptron, is used for feature extraction. The cross-fusion encoder uses two self-attention mechanisms placed in parallel for detailed and global feature extraction respectively. This layer is located between two layer normalizations. The features of the input image are divided into two subsets to process global information and fine-grained features respectively, which can be expressed as:
[0081] X i = Split(Z L ), i ∈ {0, 1}
[0082] In the above formula, X i represents the i-th feature subset obtained by cutting.
[0083] For the two feature subsets, window - to - window attention is used to learn the global information of the image, and self - attention within the window is used to learn fine - grained features. Then, cross - splicing is performed on the self - attention output to ensure that each pixel in the final output feature map has both global and local information, which can be expressed as:
[0084] Z′ L+1 =Conv 1×1 {Contact[MHSA(X 1 ),MHSA(X 0 )]}
[0085] Finally, a residual structure is obtained by combining LN and a multi - layer perceptron (MLP) to get the encoder output, which can be expressed as:
[0086] Z L+1 =MLP(LN(Z′ L+1 ))+Z′ L+1
[0087] In the above formula, Z L+1 represents the morphological features of the first patient's laryngeal image in the output.
[0088] S214. Based on the linear layer of the Transformer network model with the cross - fused encoder, perform image feature category prediction processing on the morphological features of the preliminary first patient's laryngeal image to obtain the morphological features of the first patient's laryngeal image.
[0089] S220. Use the ResNet - 50 deep residual network model to perform feature extraction processing on the pre - processed patient's laryngeal image to obtain the morphological features of the second patient's laryngeal image, and the morphological features of the second patient's laryngeal image represent the lesion location features;
[0090] S230. Combine the morphological features of the first patient's laryngeal image and the morphological features of the second patient's laryngeal image to obtain the morphological features of the patient's laryngeal image;
[0091] S240. Perform lesion image analysis on the pre - processed patient's laryngeal image to obtain the quantitative features of the patient's laryngeal image. The quantitative features of the patient's laryngeal image include aspect ratio information, color spectrum information, S - channel image entropy information, texture information, histogram of oriented gradients information, and color moment information.
[0092] In some specific embodiments, in addition to the above 8 deep learning-based features, 6 quantitative features are obtained by analyzing the lesion images. The first is the aspect ratio: the width-to-height ratio of the lesion area, which reflects the shape of the lesion. The second is the color spectral information: the main color components extracted after transforming the image color space. The third is the entropy of the S-channel image: in the HSI color space, the entropy value of the S-channel image reflects the color characteristics. The fourth is the texture information: the texture features of the lesion are extracted based on the local binary pattern (LBP) method. The fifth is the histogram of oriented gradients: the boundary and shape characteristics of the lesion are captured through the distribution information. The sixth is the color moment: it reflects the brightness, distribution area, and symmetry of the color. Finally, 14 features related to the laryngeal image and the lesion are extracted for multi-scale feature fusion decision-making.
[0093] S300. Perform multi-scale feature fusion processing on the morphological features and quantitative features of the patient's laryngeal image to obtain the key features of the patient's laryngeal image with feature weights.
[0094] It should be noted that in some embodiments, step S300 may include: S310. Fuse the morphological features and quantitative features of the patient's laryngeal image to obtain the multi-scale feature data of the patient's laryngeal image. S320. Based on the random forest model, use the sampling method with replacement to repeatedly and randomly select the multi-scale feature data of the patient's laryngeal image to construct a training sample subset. S330. Train multiple decision trees based on the training sample subset. Each tree randomly selects features for splitting during the training process to construct a random forest model. S340. Based on the random forest model, perform voting decision on the multi-scale feature data of the patient's laryngeal image to obtain the key features of the patient's laryngeal image with feature weights.
[0095] In some specific embodiments, as Figure 5 shown, the multi-scale feature fusion decision module uses the "sampling with replacement" method based on the random forest model to generate multiple training sample subsets from the image feature data, and each sample subset is used to train a decision tree. For each decision tree, during the splitting process of each node, a part of the features are randomly selected from all the features as candidate features for splitting. After multiple decision trees are independently trained, a random forest model is formed.
[0096] The related definitions are as follows:
[0097] Data set: where represents the feature vector of the sample, and y i ∈{1, 2,..., C} represents the category of the sample.
[0098] Decision tree model: T m (x) represents the classification result of the m-th tree for the sample x, where m = 1, 2,..., M, and M is the total number of decision trees.
[0099] Feature subset: At each splitting node, the randomly selected feature subset is denoted as where |F m | << d.
[0100] By integrating multiple decision trees, the fusion decision of multi-scale features of laryngeal images can be achieved, and the prediction result is determined by majority voting, which can be expressed as:
[0101]
[0102] where ∥(·) is the indicator function. When T m (x) = c, ∥ = 1, otherwise 0, representing the final classification result.
[0103] The final model will output a predicted class label, and 6 key features are extracted by calculating the importance of each feature in the classification task to explain the judgment basis of the model.
[0104] S400. Perform human-computer interaction visualization processing on the key features of the patient's larynx with feature weights to construct the recognition and diagnosis results of the patient's larynx;
[0105] It should be noted that in some embodiments, step S400 may include: S410. Output interpretable information according to the key features of the patient's larynx with feature weights. The interpretable information includes texture information with weight ratio, mass boundary information with weight ratio, blood vessel morphology information with weight ratio, S-channel image entropy information with weight ratio, blood vessel color information with weight ratio, and blood vessel orientation information with weight ratio; S420. Perform visual saliency annotation processing on the key features of the patient's larynx with feature weights according to the interpretable information to obtain the lesion area of the patient's larynx; S430. Perform confirmation feedback on the lesion area of the patient's larynx to obtain the recognition and diagnosis results of the patient's larynx.
[0106] In some specific embodiments, such as Figure 6As shown in the figure, the human-computer interaction online learning is written in Python language and consists of three parts: an interpretable disease diagnosis unit, a visual lesion display unit, and an online feedback optimization unit. The interpretable disease diagnosis unit can display the diagnosis results of the current case and the current image frame, and at the same time give interpretable information, including six key features and their proportions such as texture information, tumor boundary, blood vessel morphology, S-channel image entropy, blood vessel color, and blood vessel direction, to help doctors better understand the diagnosis effect. The visual lesion display unit can not only display the patient's laryngeal image in real time, but also perform significant annotation on the lesion area based on deep learning features. The online feedback optimization unit allows doctors to give feedback on the diagnosis results, mark the misdiagnosed features at the same time, dynamically update the training data set, and use the incremental learning algorithm to optimize the model parameters in real time to improve the generalization ability of the model.
[0107] In summary, in the embodiment of the present invention, the collected laryngeal images of the patient are standardized by using an electronic laryngoscope system, and then input into the trained deep learning models of Transformer and ResNet-50. Fourteen features related to laryngeal tumors are extracted based on prior knowledge. Then, a random forest model is used for multi-scale feature fusion decision-making to output an auxiliary diagnosis result containing interpretable information. Finally, an interactive platform based on a graphical interface is developed to display the diagnosis results and key features in real time. At the same time, an incremental learning mechanism is introduced to dynamically update the model data set according to the doctor's feedback data to improve the model accuracy.
[0108] The advantages of the embodiment of the present invention compared with the prior art are:
[0109] 1) Improve the diagnostic accuracy. The artificial intelligence model adopted is based on a lesion feature recognition system, which can automatically identify and classify lesion features. By making a secondary judgment on the combined local lesion features identified, it can guide clinical decisions, obtain more accurate lesion diagnosis results, reduce the misdiagnosis rate, and improve the diagnostic performance.
[0110] 2) Increase the transparency of the diagnosis. The previous artificial intelligence endoscopic diagnosis system can only return the final decision result without explanation. It is difficult for endoscopic doctors to learn from the model, and the black box model has poor interpretability, which seriously limits the clinical application of the artificial intelligence system. The embodiment of the present invention uses an interpretable artificial intelligence system to diagnose early laryngeal cancer. By feature extraction, the abstract diagnostic theory is concretized, providing the diagnosis result and diagnosis basis for endoscopic doctors, and increasing the transparency of the diagnosis.
[0111] Please refer to Figure 2 , the embodiment of the present application also provides a laryngoscope image recognition system based on multi-feature extraction, which can implement the above-mentioned laryngoscope image recognition method based on multi-feature extraction. The system includes:
[0112] The first module 201 is configured to obtain a patient's laryngeal image and perform preprocessing on the image data to obtain the preprocessed patient's laryngeal image;
[0113] The second module 202 is configured to perform multi-feature extraction processing on the preprocessed patient's laryngeal image based on a deep learning feature extraction model to obtain the morphological features and quantitative features of the patient's laryngeal image;
[0114] The third module 203 is configured to perform multi-scale feature fusion processing on the morphological features and quantitative features of the patient's laryngeal image to obtain the key features of the patient's laryngeal image with feature weights;
[0115] The fourth module 204 is configured to perform human-computer interaction visualization processing on the key features of the patient's laryngeal image with feature weights to construct the recognition and diagnosis results of the patient's larynx.
[0116] It can be understood that the content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented by the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0117] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A laryngoscope image recognition method based on multi-feature extraction, characterized in that: The method comprises the following steps: Acquire a laryngeal image of the patient and perform image data preprocessing to obtain a preprocessed laryngeal image of the patient; Performing multi-feature extraction processing on the preprocessed patient laryngeal image based on a deep learning feature extraction model to obtain morphological features of the patient laryngeal image and quantitative features of the patient laryngeal image; Performing multi-scale feature fusion processing on the morphological features of the patient's laryngeal image and the quantitative features of the patient's laryngeal image to obtain key features of the patient's laryngeal image with feature weights; The key features of the patient's larynx with feature weights are processed visually through human-computer interaction to construct a patient's larynx identification and diagnosis result.
2. The method according to claim 1, characterized in that The step of acquiring a laryngeal image of a patient and performing image data preprocessing to obtain a preprocessed laryngeal image of the patient includes: Acquire images of the patient's larynx using an electronic laryngoscope device; Extracting clear key frames of the patient's laryngeal image and filtering blurry frames and non-target areas of the patient's laryngeal image to construct a roughly screened patient's laryngeal image; Performing precision testing on the roughly screened patient laryngeal image to obtain a standard patient laryngeal image; The standard patient laryngeal image is subjected to image enhancement and formatting processing to obtain the pre-processed patient laryngeal image.
3. The method according to claim 2, characterized in that The step of performing precision detection on the roughly screened patient laryngeal image to obtain a standard patient laryngeal image includes: An image clarity evaluation function is constructed based on the Laplace operator, and image blur detection processing is performed on the coarsely screened patient laryngeal image to obtain a clear patient laryngeal image; Calculate the ratio of valid matching points of the clear patient laryngeal image based on the ORB algorithm to obtain an image repeatability score; Based on a preset scoring threshold and in combination with the image repeatability score, an image repeatability test is performed on the clear patient laryngeal image to obtain a filtered patient laryngeal image; The filtered patient laryngeal image is subjected to laryngeal key structural feature detection by a target detection algorithm to obtain the standard patient laryngeal image.
4. The method according to claim 1, characterized in that: The method of performing multi-feature extraction processing on the preprocessed patient laryngeal image based on the deep learning feature extraction model to obtain morphological features of the patient laryngeal image and quantitative features of the patient laryngeal image includes: Performing feature extraction processing on the preprocessed patient laryngeal image through a Transformer network model based on a cross-fusion encoder to obtain morphological features of the first patient laryngeal image, wherein the morphological features of the first patient laryngeal image include texture features, color features, boundary features, vascular morphology features, vascular color features, vascular direction features, and mucosal color features; Performing feature extraction processing on the preprocessed patient laryngeal image by using a ResNet-50 deep residual network model to obtain morphological features of the second patient laryngeal image, wherein the morphological features of the second patient laryngeal image represent lesion location features; Combining the morphological features of the first patient's laryngeal image with the morphological features of the second patient's laryngeal image to obtain the morphological features of the patient's laryngeal image; The preprocessed patient laryngeal image is subjected to lesion image analysis to obtain quantitative features of the patient laryngeal image, wherein the quantitative features of the patient laryngeal image include aspect ratio information, color spectrum information, S channel image entropy information, texture information, directional gradient histogram information, and color moment information.
5. The method according to claim 4, characterized in that The step of performing feature extraction processing on the preprocessed patient laryngeal image by using a Transformer network model based on a cross-fusion encoder to obtain morphological features of the first patient laryngeal image includes: Inputting the preprocessed patient laryngeal image into the Transformer network model based on the cross-fusion encoder, wherein the Transformer network model based on the cross-fusion encoder includes a plurality of image feature extraction modules and linear layers, wherein the plurality of image feature extraction modules include a patch embedding layer and a cross-fusion encoder module; Based on the patch embedding layer of the Transformer network model of the cross-fusion encoder, the preprocessed patient laryngeal image is subjected to dimensionality reduction processing to obtain a dimensionality reduced patient laryngeal image; Based on the cross-fusion encoder module of the Transformer network model of the cross-fusion encoder, image feature extraction processing is performed on the dimension-reduced patient laryngeal image to obtain preliminary morphological features of the first patient laryngeal image; Based on the linear layer of the Transformer network model of the cross-fusion encoder, image feature category prediction processing is performed on the morphological features of the preliminary first patient laryngeal image to obtain the morphological features of the first patient laryngeal image.
6. The method according to claim 5, characterized in that The cross-fusion encoder module of the Transformer network model based on the cross-fusion encoder performs image feature extraction processing on the dimension-reduced patient laryngeal image to obtain preliminary morphological features of the first patient laryngeal image, including: Inputting the dimensionally reduced patient laryngeal image into a cross-fusion encoder module of a Transformer network model of the cross-fusion encoder, wherein the cross-fusion encoder module includes a first normalization layer, a parallel attention layer, a second normalization layer and a multi-layer perceptron; Based on the first normalization layer and the second normalization layer of the cross-fusion encoder module, feature extraction is performed on the reduced-dimensional patient laryngeal image to obtain a patient laryngeal feature image; Based on the parallel attention layer of the cross-fusion encoder module, the patient's laryngeal feature image is divided into global feature information and fine-grained features and cross-joined to obtain a joined patient's laryngeal image feature; Based on the multi-layer perceptron of the cross-fusion encoder module, the features of the spliced patient laryngeal image are perceived and output to obtain the morphological features of the preliminary first patient laryngeal image.
7. The method according to claim 1, characterized in that The multi-scale feature fusion processing is performed on the morphological features of the patient's laryngeal image and the quantitative features of the patient's laryngeal image to obtain the key features of the patient's laryngeal image with feature weights, including: fusing the morphological features of the patient's laryngeal image with the quantitative features of the patient's laryngeal image to obtain multi-scale feature data of the patient's laryngeal image; Based on the random forest model, the multi-scale feature data of the patient's laryngeal images were repeatedly randomly selected using the sampling method with replacement to construct a training sample subset. Training multiple decision trees based on the training sample subset, each tree randomly selecting features to split during the training process, and constructing a random forest model; Voting decisions are made on the multi-scale feature data of the patient's laryngeal image based on the random forest model to obtain the key features of the patient's laryngeal image with feature weights.
8. The method according to claim 7, characterized in that The expression of the voting decision of the random forest model is specifically as follows: In the above formula, represents the final classification result, T m (x) represents the classification result of the mth tree for sample x, C represents the number of sample categories, c represents the sample category, and ∥(·) represents the indicator function.
9. The method according to claim 1, characterized in that: The step of performing human-computer interactive visualization processing on the key features of the patient's larynx with feature weights to construct a patient's larynx identification and diagnosis result includes: Outputting interpretable information according to the key features of the patient's larynx with feature weights, the interpretable information including texture information with weight proportions, tumor boundary information with weight proportions, vascular morphology information with weight proportions, S channel image entropy information with weight proportions, vascular color information with weight proportions, and vascular direction information with weight proportions; According to the explainable information, visually annotate the key features of the patient's larynx with the feature weights to obtain the lesion area of the patient's larynx; Confirmation feedback is provided for the lesion area of the patient's larynx to obtain a diagnosis result of the patient's larynx identification.
10. A laryngoscope image recognition system based on multi-feature extraction, characterized in that: The system comprises: The first module is used to obtain a laryngeal image of a patient and perform image data preprocessing to obtain a preprocessed laryngeal image of the patient; The second module is used to perform multi-feature extraction processing on the pre-processed patient laryngeal image based on a deep learning feature extraction model to obtain morphological features of the patient laryngeal image and quantitative features of the patient laryngeal image; A third module is used to perform multi-scale feature fusion processing on the morphological features of the patient's laryngeal image and the quantitative features of the patient's laryngeal image to obtain key features of the patient's laryngeal image with feature weights; The fourth module is used to perform human-computer interactive visualization processing on the key features of the patient's larynx with feature weights to construct the patient's larynx identification and diagnosis results.
Citation Information
Patent Citations
Laryngoscope image multi-attribute classification method based on multi-modal information fusion
CN116664929A
Chest medical image multi-label intelligent diagnosis algorithm based on multi-modal comparative learning
CN118136239A
Intelligent endoscope image feature processing method and device
CN118470481A
Endoscope image feature learning model training method and apparatus, and endoscope image classification model training method and apparatus
WO2023071680A1
CT pancreatic tumor automatic segmentation method and system, terminal and storage medium
WO2024000161A1