Epistaxis Recognition Method and Computer Device Based on Multi-Scale Feature Fusion
By using a full convolutional network and Transformer to extract features in nosebleed image recognition, and combining multi-scale feature fusion and lighting normalization algorithms, the problems of low accuracy of nosebleed image recognition and insufficient diversity of data sets are solved, and higher recognition accuracy and model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411693781.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The accuracy of nosebleed image recognition and insufficient diversity of data sets lead to poor accuracy and generalization of image analysis.
The local features of nosebleed images are extracted by a full convolutional network, combined with Transformer to extract global features, and the recognition accuracy is improved through multi-scale feature fusion, context attention mechanism and Focal Loss loss function. At the same time, the data set is expanded through image enhancement and generation adversarial networks, the generalization ability of the model is enhanced, and the adaptive lighting normalization algorithm is used to deal with the lighting problem.
It improves the accuracy of nosebleed recognition and generalization ability of the model, enhances the processing ability of small bleeding areas and light changes, and improves the accuracy and reliability of image analysis.
Smart Images

Figure CN119206421B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of epistaxis recognition, and in particular to an epistaxis recognition method and a computer device based on multi-scale feature fusion. Background Art
[0002] With the continuous progress of medical imaging technology, the application of computer vision in clinical practice has also developed significantly, especially in the analysis of diseases such as epistaxis. Automatic image analysis can help medical professionals accurately identify the bleeding source, evaluate the severity of the condition, and plan treatment strategies. However, despite the potential benefits, current challenges in epistaxis image processing involve multiple technical and clinical issues. One key problem is the variability in the presentation of epistaxis cases, especially in terms of severity and underlying causes. Secondary epistaxis may be related to trauma, medications (especially anticoagulants), or other health conditions (such as hypertension or blood diseases). This variability makes it difficult for automated systems to consistently and accurately process images for diagnosis and treatment. Coupled with the variations in image quality and lighting conditions during nasal endoscopy, as well as the limitations in the scale and diversity of the training datasets used for epistaxis detection and segmentation. The current epistaxis image analysis has the following problems:
[0003] The accuracy of epistaxis image recognition and the diversity of the dataset are limited. Current epistaxis image analysis methods have difficulties in accurate segmentation, especially for small or irregular bleeding areas. Further, the differences in image quality, lighting conditions, and limited datasets result in poor generalization, reducing the accuracy of epistaxis recognition. Summary of the Invention
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide an epistaxis recognition method and a computer device based on multi-scale feature fusion, which improve the accuracy of epistaxis recognition.
[0005] The present invention adopts the following technical solutions to achieve the above purpose. In the first aspect, the present invention provides an epistaxis recognition method based on multi-scale feature fusion, including:
[0006] S1. Extract local features of the epistaxis image using a fully convolutional network, in the following manner:
[0007] , where I represents the input image, represents the fully convolutional network, represents the local feature map;
[0008] S2. Extract global features using a Transformer to capture global context information;
[0009] Input the local feature map into the Transformer to obtain the global feature, , which represents the global feature map;
[0010] S3. Perform multi-scale feature fusion;
[0011] Apply different scale transformations to the local feature map and the global feature map to generate feature maps of different scales;
[0012] For the local feature map, generate a local multi-scale feature pyramid:
[0013] , where S represents different scales, represents the downsampling operation, which represents the local multi-scale feature map;
[0014] For the global feature map, generate a global multi-scale feature pyramid:
[0015] , which represents the global multi-scale feature map;
[0016] Fuse the local and global feature maps of different scales;
[0017] Concatenation fusion:
[0018] ;
[0019] Addition fusion:
[0020] , which represents the fused feature map;
[0021] S4. Perform context fusion and output the final feature;
[0022] Assign different weights to the local and global features at each scale through the context attention mechanism;
[0023] Attention weight calculation:
[0024] , where represents feature fusion, FC represents the fully connected layer, which represents the weight at each scale;
[0025] Finally, output the final feature:
[0026] , which represents the final feature map;
[0027] S5. Calculate the cross-entropy loss;
[0028] , represents the cross - entropy loss, represents the predicted probability of the model for the correct class, is the balance factor, used to handle class imbalance, is a tuning parameter, called the focusing parameter, used to adjust the model's attention to difficult samples;
[0029] S6. CNN model training;
[0030] Input the final feature map and calculate the predicted probability of each pixel in the final feature map;
[0031] Calculate the cross - entropy loss according to the cross - entropy loss calculation formula, then increase the loss for the set samples, and adjust and , and update the weights of the CNN model;
[0032] Finally, identify epistaxis through the trained CNN model.
[0033] Furthermore, before using the fully convolutional network to extract local features of the epistaxis image, the method further includes: enhancing the epistaxis image and augmenting the dataset, specifically including:
[0034] Geometric rotation of the epistaxis image is performed as follows:
[0035] ;
[0036] Translation of the epistaxis image is performed as follows:
[0037] , where is the translation distance;
[0038] Adjustment of the contrast and brightness of the epistaxis image is performed as follows:
[0039] , where controls the contrast, controls the brightness, is the rotation angle, respectively represent the coordinates of the original epistaxis image, respectively represent the coordinates of the enhanced epistaxis image, represents the enhanced epistaxis image;
[0040] Augment the epistaxis image dataset through the generative adversarial network, specifically including:
[0041] Generate a fake image from the random noise z through the generator of the generative adversarial network ;
[0042] At the same time, the real image x and the fake image are input into the discriminator of the generative adversarial network, and the judgment result is output and , represents the authenticity score of the real image x, represents the fake image 's authenticity score;
[0043] Optimize the loss function of the generative adversarial network. The goal of the generator is to minimize the following loss:
[0044] , represents the loss function of the generator;
[0045] The goal of the discriminator is to maximize the following loss:
[0046] , represents the loss function of the discriminator;
[0047] Finally, expand the epistaxis images through the trained generative adversarial network.
[0048] Furthermore, before using the fully convolutional network to extract the local features of the epistaxis images, the method further includes: performing illumination normalization processing on the epistaxis images, specifically including:
[0049] Performing illumination analysis on the input epistaxis images:
[0050] Calculating the brightness distribution of the epistaxis images, and converting the epistaxis images from the RGB color space to the LAB color space. The conversion formula is as follows:
[0051] , where L is the brightness channel, and A and B are the color channels respectively;
[0052] Calculating the brightness histogram:
[0053] Performing histogram analysis on the brightness channel L to obtain the illumination distribution in the epistaxis images, in the following way:
[0054] , where H(L) is the histogram of the brightness channel, represents the frequency of the brightness value in the image;
[0055] Detecting the uneven illumination areas:
[0056] Calculating the maximum value, minimum value and average value of the brightness of the epistaxis images, in the following way:
[0057] ;
[0058] Wherein is the maximum luminance value in the image, is the minimum luminance value, is the average luminance of the image;
[0059] Processing of over-bright and over-dark regions:
[0060] For over-bright regions, perform luminance reduction processing to make the difference from the average luminance within the set range. The method is as follows:
[0061] , wherein is the adjustment parameter that controls the amplitude of luminance adjustment, represents the adjusted luminance value;
[0062] For over-dark regions, perform luminance enhancement processing to make the difference from the average luminance within the set range. The method is as follows:
[0063] , wherein is another adjustment parameter used to enhance the luminance of dark regions;
[0064] Global illumination normalization processing:
[0065] Normalize the adjusted luminance channel. The processing method is as follows:
[0066] ;
[0067] Wherein and are the upper and lower limits of the luminance value;
[0068] Reconstructing the image:
[0069] Merge the processed luminance channel with the unchanged A channel and B channel to form a new LAB image. The method is as follows:
[0070] ;
[0071] Then convert the LAB image back to the RGB color space to obtain the normalized image. The method is as follows:
[0072] , represents the normalized image.
[0073] Furthermore, in step S1, when the fully convolutional network uses an adaptive convolution kernel to extract local features of the epistaxis image, the method is as follows:
[0074] , represents the size of the convolutional kernel at the current feature scale;
[0075] ;
[0076] and represent the minimum and maximum sizes of the convolutional kernel respectively, is the change amount of the gradient of the current feature map, is a tuning parameter used to control the sensitivity of the convolutional kernel adjustment, is the maximum value of the gradient change for normalization.
[0077] Furthermore, the input image includes endoscopic images, CT images, and MRI images. Before extracting local features, the method further includes:
[0078] Aligning the endoscopic images, CT images, and MRI images through rigid registration and non-rigid registration;
[0079] Rigid registration is used to solve rotation and translation errors and ensure the alignment of CT images or MRI images with endoscopic images. The rigid transformation matrix is as follows:
[0080] , where is the transformed coordinate, is the rotation matrix, is the translation vector, is the input image coordinate;
[0081] Non-rigid registration is used to handle situations other than rigid registration. The non-rigid registration formula is as follows:
[0082] , where is the weight, is the deformation kernel function, is the reference point.
[0083] In a second aspect, the present invention provides a computer device, including a memory, where the memory stores program instructions, and when the program instructions run, they execute the above-mentioned epistaxis recognition method based on multi-scale feature fusion.
[0084] The beneficial effects of the present invention are as follows:
[0085] The present invention uses a fully convolutional network to extract local feature information of epistaxis images and a Transformer to extract global feature information, improving the accuracy of feature information.
[0086] In view of the problems of unbalanced data sets and the easy neglect of small bleeding areas, the present invention uses the Focal Loss function to dynamically adjust the attention weights during the training process, focusing on difficult negative samples and small sample areas, thereby improving the accuracy.
[0087] In view of the problem of insufficient data set diversity, the present invention expands the data set scale through image enhancement techniques (such as rotation, translation, random cropping, etc.) and uses a generative adversarial network (GAN) to synthesize new images to generate more nasal bleeding images under different conditions, enhancing the generalization ability of the model.
[0088] Before the start of image analysis, in view of the problem of uneven illumination, the present invention applies an adaptive illumination normalization algorithm. This algorithm adjusts the illumination difference in the image to equalize over-dark or over-bright areas, avoiding the interference of illumination changes on the detection of bleeding sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 is a flowchart of a nosebleed recognition method based on multi-scale feature fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0090] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0091] The present invention provides a nosebleed recognition method based on multi-scale feature fusion, as Figure 1 shown, including:
[0092] S1. Use a fully convolutional network to extract local features of nosebleed images;
[0093] The fully convolutional network is mainly used to extract local detail information in the image. The bleeding source in the nasal cavity may be small and irregular in shape. Therefore, the fully convolutional network can extract details layer by layer from the image and capture local features such as small bleeding sources.
[0094] The specific method is as follows:
[0095] , I represents the input image, represents the fully convolutional network, represents the local feature map. In the present invention, represents the feature map with a larger size, and each pixel point represents the feature of the local area.
[0096] S2. Use a Transformer to extract global features;
[0097] The global feature extraction uses Transformer, which is good at capturing global context information and can understand the structure and semantics of the entire image. Through the self-attention mechanism, Transformer analyzes the correlation between different regions in the image, especially suitable for the overall understanding of complex anatomical structures.
[0098] The specific method is as follows:
[0099] Input the local feature map into Transformer to obtain global features, , denoted as the global feature map. The self-attention mechanism of Transformer can capture the correlation of distant regions in the image and enhance the global understanding of complex structures.
[0100] S3. Perform multi-scale feature fusion;
[0101] The core of multi-scale feature fusion is to combine local features and global features to ensure that the model can process both detailed information and overall structure simultaneously.
[0102] Feature pyramid construction:
[0103] The local features extracted from FCN and the global features generated by Transformer . After different scale transformations, feature maps of different scales are generated to capture different spatial information respectively.
[0104] For the local feature map, generate a local multi-scale feature pyramid:
[0105] , where S represents different scales, denotes the downsampling operation;
[0106] For the global feature map, generate a global multi-scale feature pyramid:
[0107] ;
[0108] Through the above method, global and local features at different levels can be obtained.
[0109] Feature fusion:
[0110] Fuse the local and global features of different scales. In the present invention, splicing fusion is mainly used and addition fusion is supplemented, so that richer and more significant features can be obtained.
[0111] Splicing fusion:
[0112] ;
[0113] After splicing, further feature extraction is performed through convolution operations to obtain richer fused features.
[0114] Additive fusion:
[0115] , represents the fused feature;
[0116] Fuse features of different scales by adding them pixel by pixel to enhance feature representation.
[0117] S4. Perform context fusion and output the final features;
[0118] The context attention mechanism is used to further enhance the feature fusion effect. The context attention mechanism assigns different weights to the local and global features at each scale to ensure that the features of the key regions (bleeding sources) are amplified.
[0119] Attention weight calculation:
[0120] , represents feature fusion, FC represents the fully connected layer, represents the weight at each scale;
[0121] Finally, output the final features:
[0122] , represents the final feature map;
[0123] S5. Calculate the cross-entropy loss;
[0124] Focal Loss is an improved version of the cross-entropy loss, aiming to reduce the contribution of easy samples to the loss and increase the model's attention to hard samples. The formula is as follows:
[0125] , represents the cross-entropy loss, represents the predicted probability of the model for the correct class, is the balance factor, used to handle class imbalance, is a tuning parameter, called the focusing parameter, used to adjust the model's attention to hard samples.
[0126] Handling class imbalance:
[0127] By introducing the balance factor , Focal Loss assigns a greater weight to the small sample (small bleeding source) class to ensure that they are not ignored during training.
[0128] Enhancing hard samples:
[0129] By adjusting the parameters , when the prediction probability of the model for a certain sample is low (i.e., a sample that is difficult to predict, often a small bleeding source), the loss value will be amplified, forcing the model to pay more attention to these difficult-to-detect areas. On the contrary, for those samples that are easy to detect, the loss value is reduced to reduce over-training on simple samples.
[0130] Flexibility to adapt to different bleeding areas:
[0131] As the training progresses, the model will gradually increase the weight for those bleeding areas that are difficult to detect (such as bleeding points with low contrast or irregular morphology), and finally enhance the sensitivity of the model to these areas.
[0132] S6. CNN model training:
[0133] Input the final feature map and calculate the prediction probability of each pixel in the final feature map;
[0134] Calculate the cross-entropy loss according to the cross-entropy loss calculation formula, then increase the loss for the set samples, and adjust and , update the weights of the CNN model;
[0135] Finally, identify epistaxis through the trained CNN model.
[0136] In the present invention, the epistaxis images are enhanced by performing operations such as geometric transformation and color adjustment on the existing data to generate more diverse training data. This helps to expand the scale of the dataset and improve the generalization ability of the model. The specific steps are as follows:
[0137] Perform geometric rotation on the epistaxis images in the following manner:
[0138] ;
[0139] Perform translation on the epistaxis images in the following manner:
[0140] , where is the translation distance, is the rotation angle, respectively represent the coordinates of the original epistaxis image, respectively represent the coordinates of the enhanced epistaxis image, represents the enhanced epistaxis image;
[0141] Perform contrast and brightness adjustment on the epistaxis images in the following manner:
[0142] , where Control the contrast, control the brightness;
[0143] A generative adversarial network is a generative model that mainly generates high-quality synthetic images through the adversarial training of two networks (a generator and a discriminator).
[0144] Generator: Responsible for generating nasal bleeding images similar to real images from random noise.
[0145] Discriminator: Responsible for distinguishing between generated images and real images.
[0146] The goal of the generative adversarial network is to make the images generated by the generator more and more real, and ultimately be able to deceive the discriminator.
[0147] Expand the nasal bleeding image dataset through the generative adversarial network, specifically including:
[0148] Generate fake images from random noise z through the generator of the generative adversarial network ;
[0149] At the same time, input the real image x and the fake image into the discriminator of the generative adversarial network, and output the judgment result and , represents the authenticity score of the real image x, represents the authenticity score of the fake image ;
[0150] Optimize the loss function of the generative adversarial network. The goal of the generator is to minimize the following loss:
[0151] , represents the loss function of the generator;
[0152] The goal of the discriminator is to maximize the following loss:
[0153] , represents the loss function of the discriminator;
[0154] Finally, expand the nasal bleeding images through the trained generative adversarial network.
[0155] The present invention applies an adaptive illumination normalization algorithm to solve the problem of uneven illumination. By adjusting the illumination difference in the image, the too dark or too bright areas are made uniform, avoiding the interference of illumination changes on the detection of the bleeding source.
[0156] The illumination normalization algorithm is used to address the issue of uneven illumination in images. Especially in nasal endoscope images, due to complex lighting conditions, there are often over-bright or over-dark regions, which can affect the subsequent detection of bleeding sources. The purpose of illumination normalization is to adjust the illumination differences in the image to make the overall illumination of the image uniform, thereby improving the image quality and enhancing the detection accuracy.
[0157] The illumination normalization process for nosebleed images specifically includes:
[0158] Perform illumination analysis on the input nosebleed image:
[0159] Calculate the brightness distribution of the nosebleed image. Convert the nosebleed image from the RGB color space to the LAB color space. The conversion formula is as follows:
[0160] , where L is the brightness channel, and A and B are the color channels respectively;
[0161] Calculate the brightness histogram:
[0162] Perform histogram analysis on the brightness channel L to obtain the illumination distribution in the image. The histogram represents the frequency of different brightness values in the image, which helps to identify over-bright and over-dark regions in the image.
[0163] The specific method is as follows:
[0164] , where H(L) is the histogram of the brightness channel, represents the frequency of the brightness value in the image;
[0165] Detect regions with uneven illumination:
[0166] Through the histogram, it is possible to identify which parts of the image are over-bright and which are over-dark. To further determine the degree of uneven illumination, calculate the maximum, minimum, and average values of the image brightness.
[0167] Calculate the maximum, minimum, and average values of the nosebleed image brightness as follows:
[0168] ;
[0169] where is the maximum brightness value in the image, is the minimum brightness value, is the average brightness of the image;
[0170] Treatment of over-bright and over-dark regions:
[0171] For overly bright regions, perform brightness reduction processing to make the difference between their brightness and the average brightness within the set range, i.e., make it close to the average brightness, in the following way:
[0172] , where is an adjustment parameter that controls the amplitude of brightness adjustment;
[0173] For overly dark regions, perform brightness enhancement processing to make the difference between their brightness and the average brightness within the set range, i.e., make it close to the average brightness, in the following way:
[0174] , where is another adjustment parameter used to enhance the brightness of dark regions;
[0175] Global illumination normalization processing:
[0176] Normalize the adjusted brightness channel to ensure that the brightness values are within the specified range (set to 0 to 255). This can eliminate the influence of illumination differences and make the brightness distribution of the image more uniform.
[0177] Normalize the adjusted brightness channel in the following way:
[0178] ;
[0179] where and are the upper and lower limits of the brightness values;
[0180] Image reconstruction:
[0181] Merge the processed brightness channel with the unchanged A channel and B channel to form a new LAB image in the following way:
[0182] ;
[0183] Then convert the LAB image back to the RGB color space to obtain the normalized image in the following way:
[0184] , represents the normalized image.
[0185] The present invention designs an adaptive convolution kernel by combining FCN and Transformer, aiming to improve the feature extraction ability of the model in processing different regions of the complex nasal cavity structure. For narrow regions and local details, the adaptive convolution kernel will shrink to capture fine features; for spacious regions, the adaptive convolution kernel will expand to obtain more context information. With this model, local details and global semantics can be taken into account, improving the positioning and detection accuracy of the bleeding source.
[0186] Adaptive Convolution Kernel Adjustment
[0187] The core idea of the adaptive convolution kernel is to automatically adjust the size of the convolution kernel according to different regions. By analyzing local features of the image (such as gradient changes), it can be determined whether the size of the convolution kernel needs to be adjusted. The formula for the adaptive size of the convolution kernel is as follows:
[0188]
[0189] Where:
[0190] is the size of the convolution kernel at the current feature scale;
[0191] and represent the minimum and maximum sizes of the convolution kernel respectively (set to to );
[0192] is the change amount of the gradient of the current feature map;
[0193] is the adjustment parameter, set to 0.5, used to control the sensitivity of the convolution kernel adjustment;
[0194] is the maximum value of the gradient change, used for normalization.
[0195] The present invention also provides a method for localizing nasal bleeding sources based on multi-modal information fusion.
[0196] Multi-modal information fusion technology aims to combine information from different imaging modalities (such as endoscopic images, CT, MRI, etc.) to make up for the deficiencies of a single modality in capturing anatomical structures and detailed information. By combining the local details of endoscopic images with the global anatomical information of CT and MRI, the recognition accuracy of internal nasal bleeding sources can be improved. Multi-modal information fusion technology can help doctors understand the nasal cavity structure more comprehensively, especially in complex surgical environments.
[0197] Specifically include:
[0198] 1. Data preprocessing
[0199] Before performing multi-modal information fusion, it is necessary to align images of different modalities (endoscopic images, CT, MRI) to ensure that they are analyzed in the same spatial coordinate system. This is called image registration. The registration process usually requires spatially aligning the global anatomical information of CT or MRI with the local information of endoscopic images.
[0200] Rigid registration: Used to solve rotation and translation errors and ensure the alignment of CT / MRI images with endoscopic images. The rotation angle range is set to . The rigid transformation matrix is as follows:
[0201] , where is the transformed coordinate, is the rotation matrix, is the translation vector set to 3px, is the input image coordinate.
[0202] Non-rigid registration: Used to handle more complex deformation situations. The present invention uses non-rigid deformation algorithms such as B-spline or thin plate spline (TPS) to solve the elastic differences in anatomical structures. The non-rigid registration formula is as follows:
[0203]
[0204] where is the weight, is the deformation kernel function, is the reference point.
[0205] 2. Multi-modal feature extraction
[0206] Extract useful features from different modalities (endoscopic images, CT, MRI). The features of endoscopic images are mainly detailed features, capturing local bleeding areas; the features of CT and MRI are mainly global anatomical structure information.
[0207] Endoscopic images: Use a convolutional neural network to extract local features and generate a feature vector , and the convolutional kernel size is set to . The formula is as follows:
[0208]
[0209] where is the endoscopic image, and CNN extracts features.
[0210] CT / MRI images: Extract anatomical structure features through a global feature extraction network based on Transformer and generate a feature vector , and the number of attention heads , ensure effective capture of context information. The formula is as follows:
[0211]
[0212] Among them, is a CT or MRI image, Extract global features.
[0213] 3. Multimodal Feature Alignment
[0214] Extract useful features from different modalities (endoscopic images, CT, MRI). The features of endoscopic images are mainly detailed features, capturing local bleeding areas; the features of CT and MRI are mainly global anatomical structure information.
[0215] Endoscopic images: Use a convolutional neural network to extract local features and generate feature vectors , and the convolutional kernel size is set to . The formula is as follows:
[0216]
[0217] Among them, is an endoscopic image, and CNN extracts features.
[0218] CT / MRI images: Extract anatomical structure features through a global feature extraction network based on Transformer and generate feature vectors , the number of attention heads , ensure effective capture of context information. The formula is as follows:
[0219]
[0220] Among them, is a CT or MRI image, Extract global features.
[0221] Since the features of endoscopic images and CT / MRI images come from different modalities and spatial resolutions, they need to be feature-aligned before fusion. The goal of feature alignment is to ensure that features from different modalities can be compared and fused in the same scale and space.
[0222] Feature alignment: Align the local features of endoscopic images with the global features of CT / MRI. Align the feature maps through scale-space transformation and convolution operations, and the downsampling step size is 2 to ensure that feature maps with different resolutions can be aligned. The formula is as follows:
[0223]
[0224] Among them, It is an alignment operation, including feature downsampling or upsampling.
[0225] Fusion strategy: The present invention adopts addition fusion and concatenation fusion.
[0226] Addition fusion: Add the endoscopic features and CT / MRI features element by element to combine local and global features.
[0227]
[0228] Concatenation fusion: Concatenate the features together. The formula is as follows:
[0229]
[0230] Then it is further processed through a convolutional layer, where the convolutional kernel size is set to . The formula is as follows:
[0231]
[0232] 4. Modal Weighting and Attention Mechanism
[0233] After feature fusion, in order to ensure that key modalities (such as anatomical information of CT / MRI or local detail information of the endoscope) obtain higher weights, the feature weights of different modalities can be automatically assigned through the attention mechanism.
[0234] Modal weighting mechanism: Assign weights to the features of different modalities through the attention mechanism to emphasize key feature regions.
[0235]
[0236] Among them, and are the weights of CT / MRI and endoscopic features respectively. FC is a fully connected layer, and Softmax is used to normalize the weights to ensure that the weights of each modality are between 0 and 1.
[0237] Self-attention mechanism: Through the self-attention mechanism, the model can automatically focus on key anatomical structures or local details in a specific modality. The number of heads in the self-attention mechanism is set to 8 to ensure that the model can focus on the details and global information of different modalities when fusing features.
[0238]
[0239] Among them, Q, K, and V are the query matrix, key matrix, and value matrix respectively, is the dimension of the key matrix.
[0240] 5. Bleeding Source Location and Final Feature Analysis
[0241] Fused features It contains local detail features from the endoscope and global anatomical structure information from CT / MRI. Through further processing by convolutional layers, it can be used to locate the bleeding source inside the nasal cavity. Through a classification or regression network, the specific location of the bleeding source is determined. The formula is as follows:
[0242]
[0243] Among them, Classifier is the network for classification or localization, which can output the position coordinates of the bleeding source. The classifier used is a simple fully connected layer, and the output dimension of the classification network is the coordinates related to the bleeding source localization (for example, coordinates).
[0244] The present invention also provides a method for nasal bleeding localization based on time series analysis.
[0245] When the time series analysis algorithm processes consecutive frame images, it can effectively detect dynamic changes, especially in dynamic scenes in the field of medical images (such as in nasal bleeding source localization). By performing time series modeling on consecutive image frames, the flow changes of blood can be detected, and dynamic information such as the position of the bleeding source and the blood flow direction can be inferred. Specifically, in consecutive image frames (such as the dynamic video captured by the endoscope), the present invention hopes to identify the dynamic changes in the images through the time series analysis algorithm, so as to determine the blood flow trajectory and the position of the bleeding source.
[0246] Specifically, it includes:
[0247] 1. Image preprocessing and feature extraction
[0248] First, extract consecutive frames of the sequence from the video frames. Assume the consecutive frames are . After preprocessing each frame image, feature extraction is performed.
[0249] The preprocessing includes:
[0250] Denoising: Use Gaussian filtering or bilateral filtering on each frame image to remove noise.
[0251] Image segmentation: Use FCN or other segmentation networks to extract possible bleeding areas.
[0252]
[0253] Among them, is the image after preprocessing.
[0254] Feature extraction:
[0255] Extract key image features from each frame. The present invention uses a fully convolutional network (FCN) to extract local features, where the convolutional kernel size is set to . The feature vector represents the feature representation of the image at a certain time point. The formula is as follows:
[0256]
[0257] where, is the feature vector of the th frame.
[0258] 2. Dynamic change detection
[0259] By performing time series modeling on the image features of consecutive frames, the dynamic changes in blood flow can be detected. The present invention adopts three commonly used time series models, namely the autoregressive model, the moving average model, and the autoregressive moving average model. In the case of complex time-dependent relationships, the present invention will also adopt the long short-term memory network, which is also a common deep learning time series.
[0260] Ⅰ. Autoregressive model:
[0261] In the autoregressive model, the image feature at the current moment can be predicted by the features at the previous several moments:
[0262]
[0263] where: is the coefficient of the autoregressive model, indicating the influence of the past moments on the current moment;
[0264] is the autoregressive order, usually adjusted according to the data and set to 2;
[0265] is the noise term, indicating the error.
[0266] Ⅱ. Moving average model:
[0267] The moving average model considers that the image feature at the current moment is affected by the past error terms:
[0268]
[0269] where: is the moving average coefficient.
[0270] is the past error term, indicating the influence of the previous prediction error on the current prediction.
[0271] is the order of the moving average, set to 2.
[0272] III. Autoregressive Moving Average Model:
[0273] Combining the autoregressive model and the moving average model, the feature at the current moment is affected by the features at the previous several moments and the errors at the previous several moments:
[0274]
[0275] IV. LSTM (Long Short-Term Memory Network) Model:
[0276] When the time dependence relationship is relatively complex, the LSTM model can be used, and its structure is suitable for long-term and short-term dependencies.
[0277]
[0278] Among them, is the hidden state of the LSTM, representing the memory of the current moment for all past moments, and is set to 128 in the present invention.
[0279] 3. Detection of Dynamic Changes
[0280] Through time series modeling, the difference between the features of the current frame and the previous frame can be calculated, and then the dynamic changes in blood flow can be detected.
[0281]
[0282] Among them, is the feature change value between the current frame and the previous frame. If the feature change value exceeds a certain threshold, it is considered that an obvious dynamic change (i.e., nasal blood flow) has occurred in the current frame. The change detection formula is as follows:
[0283]
[0284] Among them, represents the amplitude of the feature change, and the threshold needs to be set according to the actual data, and is usually set to 0.1 or 0.2 in the present invention.
[0285] 4. Bleeding Source Location and Trajectory Prediction
[0286] By analyzing the change sequence over a period of time, the direction and speed of bleeding can be further determined. Assuming that obvious changes are detected between and , the direction of blood flow can be inferred by estimating the change trajectory.
[0287] The direction formula is as follows:
[0288]
[0289] Among them, and are respectively the eigenvalue changes of the th frame in the and directions.
[0290] The velocity formula is as follows:
[0291]
[0292] Among them, is the time frame interval, representing the velocity of blood flow.
[0293] 5. Bleeding source localization
[0294] Based on the detected blood flow direction and velocity, the position of the bleeding source can be inferred in reverse, assuming that the position of the bleeding source is the starting point of dynamic change in consecutive frames.
[0295] The above are only the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be changed within the scope of the concept described herein through the above teachings or the techniques or knowledge in the relevant field. And the changes and alterations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.
Claims
1. A nosebleed recognition method based on multi-scale feature fusion, characterized in that: include: S1. Use a fully convolutional network to extract local features of nosebleed images in the following way: , I represents the input image, represents a fully convolutional network, Represents a local feature map; S2, using Transformer to extract global features and capture global context information; Input the local feature map into Transformer to obtain the global feature. , Represents the global feature map; S3, perform multi-scale feature fusion; The local feature map and the global feature map are transformed into different scales to generate feature maps of different scales; For the local feature map, generate a local multi-scale feature pyramid: , S represents different scales, represents the downsampling operation, Represents a local multi-scale feature map; For the global feature map, generate a global multi-scale feature pyramid: , Represents a global multi-scale feature map; Fuse local and global feature maps of different scales; Splicing and fusion: ; Additive Fusion: , Represents the fused feature map; S4, perform context fusion and output the final feature map; Different weights are assigned to local feature maps and global feature maps at each scale through the contextual attention mechanism; Attention weight calculation: , represents feature fusion, FC represents the fully connected layer, Represents the weight at each scale; Finally, the final feature map is output: , represents the final feature map; S5, cross entropy loss calculation; , represents the cross entropy loss, represents the model's predicted probability for the correct category, is a balancing factor used to deal with class imbalance. is a tuning parameter, called the focus parameter, which is used to adjust the model's focus on difficult samples; S6, CNN model training; Input the final feature map and calculate the predicted probability of each pixel in the final feature map; Calculate the cross entropy loss according to the cross entropy loss calculation formula, then increase the loss for the set sample and adjust and , update the CNN model weights; Finally, the trained CNN model is used to identify nosebleeds.
2. The nosebleed recognition method based on multi-scale feature fusion according to claim 1 is characterized in that: Before using the full convolutional network to extract the local features of the nosebleed image, the method also includes: enhancing the nosebleed image and expanding the data set, specifically including: Perform geometric rotation on the nosebleed image as follows: ; Translate the nosebleed image as follows: ,in is the translation distance; Adjust the contrast and brightness of the nosebleed image as follows: ,in Control contrast, Control brightness, is the rotation angle, Respectively represent the coordinates of the original nosebleed image, Respectively represent the enhanced nosebleed image coordinates, represents the enhanced image of nose bleeding; Expand the nosebleed image dataset by generating adversarial networks, including: Generate fake images from random noise z through the generator of the GAN ; At the same time, the real image x and the fake image Input the discriminator of the generated adversarial network and output the judgment result and , represents the authenticity score of the real image x, Represents a fake image Authenticity rating; Optimizing the loss function of the generative adversarial network, the goal of the generator is to minimize the following loss: , represents the loss function of the generator; The goal of the discriminator is to maximize the following loss: , represents the loss function of the discriminator; Finally, the nosebleed image is expanded through the trained generative adversarial network.
3. The nosebleed recognition method based on multi-scale feature fusion according to claim 1 is characterized in that: Before using the full convolutional network to extract the local features of the nosebleed image, the method further includes: performing illumination normalization processing on the nosebleed image, specifically including: Perform lighting analysis on the input nosebleed image: Calculate the brightness distribution of the nosebleed image and convert the nosebleed image from RGB color space to LAB color space. The conversion formula is as follows: , L is the brightness channel, A and B are the color channels respectively; Calculate the brightness histogram: Perform histogram analysis on the brightness channel L to obtain the illumination distribution in the nosebleed image in the following way: , where H(L) is the histogram of the brightness channel, Represents the frequency of brightness values in the image; Detecting areas of uneven lighting: Calculate the maximum, minimum, and mean brightness of the nosebleed image as follows: ; in is the maximum brightness value in the image, is the minimum brightness value, is the average brightness of the image; Processing of overly bright and dark areas: For overly bright areas, the brightness is reduced so that the difference between it and the average brightness is within the set range, as follows: ,in It is an adjustment parameter that controls the amplitude of brightness adjustment. Indicates the adjusted brightness value; For dark areas, the brightness is increased so that the difference between the area and the average brightness is within the set range, as follows: ,in is another adjustment parameter used to enhance the brightness of dark areas; Global illumination normalization: The adjusted brightness channel is normalized as follows: ; in and are the upper and lower limits of the brightness value; Reconstruct the image: The processed brightness channel The A and B channels are recombined with the unchanged ones to form a new LAB image as follows: ; Then convert the LAB image back to RGB color space to get the normalized image as follows: 。 4. The nosebleed recognition method based on multi-scale feature fusion according to claim 1, characterized in that: In step S1, when the fully convolutional network uses an adaptive convolution kernel to extract local features of the nosebleed image, the method is as follows: , Indicates the convolution kernel size at the current feature scale; ; and Respectively represent the minimum and maximum sizes of the convolution kernel, is the change in the current feature map gradient, It is a tuning parameter used to control the sensitivity of the convolution kernel adjustment. is the maximum value of the gradient change and is used for normalization.
5. The nosebleed recognition method based on multi-scale feature fusion according to claim 1, characterized in that: The input image includes an endoscopic image, a CT image, and an MRI image. Before extracting the local features, the method further includes: Align endoscopic images, CT images, and MRI images through rigid and non-rigid registration; Rigid registration is used to solve rotation and translation errors to ensure that the CT image or MRI image is aligned with the endoscopic image. The rigid transformation matrix is as follows: ,in, are the transformed coordinates, is the rotation matrix, is the translation vector, are the input image coordinates; Non-rigid registration is used to handle situations other than rigid registration. The non-rigid registration formula is as follows: ,in, is the weight, is the deformation kernel function, It is a reference point.
6. A computer device comprising a memory, wherein the memory stores program instructions, characterized in that: When the program instructions are executed, the nosebleed recognition method based on multi-scale feature fusion as described in any one of claims 1 to 5 is executed.
Citation Information
Patent Citations
Medical image small target segmentation method based on double-branch feature fusion attention
CN116681679A