Inner ear malformation detection method
By combining the pulsed neural network with the YOLOv7 framework, using the biologically inspired LIF neuron model and trainable membrane potential threshold parameters, the SpikeYOLO model was constructed, which solved the problems of insufficient detection accuracy and high computational overhead in inner ear malformation detection, and achieved efficient and low-power detection effect.
Patent Information
- Application Number
- CN202510197228.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems in the detection of inner ear deformity, large computing overhead and strong dependence on large-scale labeled data in the detection of inner ear deformity. The training process of traditional deep learning models requires a large amount of computing resources and energy consumption, which is difficult to meet the energy-saving and environmental protection requirements of medical equipment.
By organically combining the advantages of pulsed neural networks (SNNs) with the YOLOv7 framework, the biologically inspired LIF neuron model is used for time sequence information encoding, and combined with trainable adaptive membrane potential threshold parameters, the SpikeYOLO model is constructed. This model simplifies feature extraction and timing processing processes by deeply integrating YOLOv7 and SNN, reducing computational complexity and energy consumption.
It realizes efficient inner ear malformation detection, improves detection accuracy and speed, reduces calculation volume and energy consumption, and is suitable for the deployment and application of clinical practice.
Smart Images

Figure CN120070395A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for detecting inner ear malformations. Background Art
[0002] Currently, the detection of inner ear malformations mainly relies on physicians' visual interpretation of temporal bone CT / MRI images. This method is not only time-consuming but also easily affected by physicians' experience and subjective judgment. With the development of artificial intelligence technology, deep learning methods have shown great potential in the field of medical image detection and recognition. Among them, the YOLO (You Only Look Once) series of algorithms have been widely used in medical image analysis due to their excellent real-time performance and detection accuracy. However, there are still some key problems to be solved in practical applications.
[0003] Although the traditional YOLOv7 framework performs well in the field of object detection, it faces several major challenges when dealing with fine medical images such as inner ear malformations: First, the inner ear structure is delicate and complex, and traditional convolutional neural networks have a large computational overhead when processing high-resolution medical images and are prone to overfitting; second, there are subtle differences between different types of inner ear malformations, which requires the model to have stronger feature extraction and classification capabilities; finally, the difficulty in obtaining medical image data leads to limited training samples, which restricts the performance of deep learning models.
[0004] The research community has proposed various improvement schemes for the above problems. Some researchers have tried to improve the model's feature extraction ability by optimizing the network structure and introducing attention mechanisms. There are also methods that use transfer learning and data augmentation techniques to alleviate the problem of insufficient samples. However, although these methods have made improvements in some aspects, they often lead to an increase in computational complexity, which is not conducive to the deployment and application of the model in clinical practice. At the same time, the training process of traditional deep learning models requires a large amount of computational resources and energy consumption, which does not meet the energy conservation and environmental protection requirements of medical devices.
[0005] Inspired by the working mechanism of the biological nervous system, spiking neural networks (SNNs) have shown unique advantages in computer vision tasks due to their low power consumption and high efficiency. Compared with traditional artificial neural networks, spiking neural networks process and transmit information by simulating the spike generation mechanism of biological neurons, and have stronger biological interpretability and computational efficiency. However, there are relatively few studies on applying spiking neural networks to medical image detection and recognition at present. The main reason is that the training of spiking neural networks is relatively difficult, and how to effectively combine the advantages of the YOLO framework remains an unsolved problem.
[0006] Therefore, how to organically combine the advantages of spiking neural networks with the YOLOv7 framework to develop an inner ear malformation image detection and recognition method that can ensure detection accuracy while also having the characteristics of low power consumption and high efficiency is a key scientific issue that needs to be solved in this field. This is not only of great significance for improving the accuracy and efficiency of inner ear malformation image detection, but will also provide a useful reference for the application of spiking neural networks in other medical image analysis tasks. Summary of the invention
[0007] In view of the above-mentioned deficiencies in the prior art, the present invention provides an inner ear malformation detection method, which solves the problems existing in the prior art, such as insufficient detection accuracy, high computational overhead, and strong dependence on large-scale labeled data.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is: a method for detecting inner ear malformation, comprising the following steps: S1, obtaining inner ear image sample data, preprocessing the inner ear image samples, and performing segmentation and labeling; S2, improve the YOLOv7 backbone network and integrate the spike neural module SNN, and use the inner ear image sample data processed by S1 to train the improved YOLOv7 network to obtain the SpikeYOLO model; S3. Use the SpikeYOLO model to analyze the inner ear image sample data; S4, determine whether the analysis result meets the preset conditions, if so, go to S5, otherwise, return to S1; S5. Use the SpikeYOLO model to analyze the inner ear image to be detected and obtain the detection result.
[0009] The beneficial effects of the present invention are as follows: the present invention deeply integrates the pulse neural module SNN with biological neuron characteristics with the YOLOv7 framework, adopts the biologically inspired LIF neuron model for temporal information encoding, and combines the trainable adaptive membrane potential threshold parameters to realize the information processing mechanism of brain-like computing; the present invention fully utilizes the advantages of the pulse neural module SNN in low power consumption and high efficiency, and solves the shortcomings of traditional methods in terms of computational complexity and real-time performance. In summary, the present invention aims to improve the detection accuracy and speed of inner ear malformations, and solves the shortcomings of traditional methods in terms of computational complexity and real-time performance.
[0010] Furthermore, the S1 comprises the following steps: S101, obtaining inner ear image sample data; S102, preprocessing the inner ear image samples, and standardizing and normalizing the preprocessed inner ear image samples; S103. Label the anatomical structure features and malformed parts of the inner ear based on the normalized inner ear images; S104. Generate object detection labels for YOLOv7 network training based on the labeled inner ear images, and divide them into training set, validation set and test set to obtain the preprocessed training sample data, completing the segmentation and labeling process.
[0011] The beneficial effects of the above further solution are: By standardizing the data preprocessing and labeling process, high-quality training samples are provided for the SpikeYOLO model; at the same time, the standardization and consistency of the data are ensured, which is beneficial to the training and convergence of the model, and improves the performance and reliability of the overall detection system.
[0012] Furthermore, the S2 includes the following steps: S201. Improve the YOLOv7 backbone network and set up a lightweight feature pyramid network; S202. Set up and fuse the spiking neural module SNN; S203. Optimize the loss function and perform end-to-end training on the improved YOLOv7 network through the gradient backpropagation method to complete the improvement of the YOLOv7 network; S204. Train the improved YOLOv7 network using the training set and verify it using the validation set to obtain the SpikeYOLO model.
[0013] The beneficial effects of the above further solution are: The present invention constructs an efficient SpikeYOLO model through the improvement of the YOLOv7 structure and the fusion of the spiking neural module SNN; through end-to-end training, the training process is simplified, the training efficiency is improved, and the collaborative optimization of feature extraction and temporal processing is ensured at the same time; the advantages of traditional CNN and spiking neural module SSNN are complementary.
[0014] Furthermore, the S201 includes the following steps: S2011. Lightweight improvement of the YOLOv7 backbone network; S2012. Optimize the YOLOv7 network feature extraction module, retain and simplify the multi-branch stacking structure Multi_Concat_Block; S2013. Improve the connection method of the transition module Transition_Block, and adjust the receptive field size and feature map resolution according to the inner ear structure features; S2014. Simplify the structure of the spatial pyramid pooling cross-stage partial network module SPPCSPC.
[0015] The beneficial effects of the above further solution are as follows: While retaining the core detection ability of the YOLO architecture, the backbone network is lightweighted by optimizing the network structures of the Backbone and Neck parts, reducing the network complexity; the optimization of the feature extraction module can maintain the diverse expression of features, enhance the feature transmission effect, reduce redundant operations, and improve the feature extraction efficiency; simplifying the structure of the Spatial Pyramid Pooling Cross-Stage Partial Network module (SPPCSPC) can retain the core spatial pyramid pooling function, reduce the computational overhead, and improve the feature processing speed.
[0016] Furthermore, S2012 includes the following steps: A1. Optimize the YOLOv7 network feature extraction module; A2. Retain the feature transmission path, simplify the four-branch stacking structure into a two-branch stacking module, and configure a set of convolution normalization activation functions and three sets of convolution normalization activation functions for the two-branch stacking module; A3. Implement channel number compression through 1×1 convolution operations and introduce a residual connection mechanism.
[0017] The beneficial effects of the above further solution are as follows: Through the above design, the YOLOv7 network structure is more concise and efficient, the feature transmission path is clearer, the computational complexity is significantly reduced, and the number of model parameters is effectively controlled. Furthermore, S2014 includes the following steps: B1. Reconstruct the original spatial pyramid pooling structure: Simplify the original five-level pyramid pooling layers of 1×1, 3×3, 5×5, 7×7, and 13×13 into a three-level pyramid pooling layer structure of 3×3, 5×5, and 7×7. Among them, each pyramid pooling layer uses the maximum pooling operation; B2. Optimize the feature fusion method: Use the form of weighted summation to replace the original concatenation operation. The expression for feature fusion is as follows: ; where represents the output inner ear image after fusion, N represents the number of pooling branches, represents the weight coefficient, represents the i th pooling operation, represents the input inner ear image.
[0018] The beneficial effects of the above further solution are as follows: By reconstructing the original pyramid structure, the pooling levels can be simplified, the computational overhead can be reduced, and the feature extraction ability of key scales can be retained; by replacing concatenation with weighted summation, the memory occupancy can be reduced, and at the same time, simplifying the feature fusion calculation can improve the processing speed.
[0019] Furthermore, S202 includes the following steps: C1. Construct a spiking neuron based on the LIF neuron model, and the expression of the LIF neuron model is as follows: ; ; ; where represents the membrane potential integrating temporal information and spatial information , represents the resting voltage, represents the actual spike output, represents the decay coefficient, represents the current membrane voltage, represents the Heaviside step function; C2. Set a trainable membrane potential threshold parameter; C3. Introduce a spiking neuron in the feature extraction layer to replace the scaled exponential linear unit (SiLU) activation function in the original YOLO network framework, and set a temporal domain joint optimization strategy to fuse the spiking neural module (SNN) with the YOLO network framework. Among them, the temporal domain optimization includes dynamically adjusting the neuron firing threshold; the spatial domain optimization includes adjusting the feature map and the receptive field size; C4. Set an adaptive membrane potential reset mechanism and introduce depthwise separable convolution. The expression of the depthwise separable convolution is as follows: ; where represents the value of the output feature map at the n th output channel and position , h and w represent the height and width positions of the feature map respectively, c represents the number of channels, C represents the total number of channels, i, j represents the sliding window position index in the convolution operation, and represent the height and width of the convolution kernel respectively, represents the value of the input feature map at the c th input channel and position , represents the weight of the depthwise convolution kernel of the c th channel at position , represents the weight of the pointwise convolution kernel at the n th output channel and the c th input channel.
[0020] The beneficial effects of the above further solution are as follows: The construction of the LIF neuron module introduces the biological neuron mechanism, enhances the biological inspiration of the model, integrates spatio-temporal information, and realizes the temporal coding of information; setting the trainable threshold parameter can improve the adaptability and flexibility of the network; the fusion of the spiking neural module SNN improves the feature extraction effect through joint optimization in the spatio-temporal domain; the optimization of the spiking neural module SNN can reduce the number of parameters and improve the efficiency of information transmission. Compared with the standard convolution operation, the number of parameters of the separable depth convolution is reduced from to , significantly reducing the computational complexity.
[0021] Furthermore, the expression of the loss function in S203 is as follows: ; ; ; ; ; ; ; where, represents the loss function of the improved YOLOv7 network, and represent the balance coefficients, represents the detection loss, represents the spike loss, represents the bounding box regression loss, represents the confidence loss for having a target, represents the confidence loss for not having a target, represents the class prediction loss, represents the time step, t represents the moment, represents the actual spike output, represents the expected spike output, represents the weight coefficient of the coordinate prediction, represents the L2 norm, represents the center coordinate of the predicted bounding box, respectively represent the width and height of the predicted bounding box, represents that there is a target in the grid cell, represents the predicted confidence score, represents the true value of the target confidence, represents the weight coefficient of the non-target area, represents only calculating the grid cells with targets, represents the predicted class probability, represents the true probability value of the class prediction, represents the center coordinate of the true bounding box, represents the width and height of the true bounding box.
[0022] The beneficial effects of the above further solution are as follows: This solution combines the object detection task and the optimization of the output of the spiking neural network, and can simultaneously process the spatial position prediction (bounding box regression), the class probability prediction (object classification), and the output of the time series in the spiking neural network; this multi-task loss function design helps to improve the overall performance by balancing the importance of different tasks; A polynomial loss function is adopted to ensure that the model can accurately locate the bounding box of the target. and Ensure that the model can correctly judge whether there is an object in the target box and correctly identify the object and the background; By minimizing the difference between the predicted class probability and the true class probability, it helps to improve the accuracy of the model for object classification; It is used to optimize the time series output of the spiking neural network, making the timing behavior of the model more accurate.
[0023] Furthermore, the step S204 includes the following steps: Set the threshold parameter for object detection, and define the detection box with an intersection over union greater than the threshold as a valid target box; During the training of the improved YOLOv7 network on the test set, introduce the non-maximum suppression mechanism to eliminate duplicate detection boxes; Construct a heatmap visualization module, and the heatmap visualization module is used to characterize the attention of the SpikeYOLO model to different regions of the inner ear image; Set a two-stage detection strategy, and the two-stage detection strategy is as follows: perform inner ear region localization on the input inner ear image, perform malformation feature analysis in the localization region, compare the localization result with the true lesion position, and according to the comparison result, evaluate the performance of the SpikeYOLO model through the coincidence degree between the brightest region of the heatmap and the true lesion position, and complete the construction of the SpikeYOLO model.
[0024] The beneficial effects of the above further solution are as follows: The present invention designs a multi-level optimization strategy, including a lightweight feature pyramid network, a simplified SPPCSPC module structure, and a two-stage detection strategy. At the same time, a heatmap visualization module and a non-maximum suppression mechanism are introduced, so that the model can provide accurate detection results (it can reach the feature extraction layer, and the effect of using the lightweight improved YOLOv7 backbone network > 90%, Accuracy > 90%, Precision > 90%, Recall > 90%), effectively improving the detection accuracy.
[0025] Furthermore, the SpikeYOLO model includes: The feature extraction layer is used to extract features by adopting the lightweight-improved YOLOv7 backbone network and through the simplified multi-branch stacking structure Multi_Concat_Block and the improved transition module Transition_Block; The feature-pulse adaptation layer is used to set up a feature normalization module to map CNN features to the interval [0, 1]; construct a pulse coding module, adopt a frequency coding method, and set a fixed time window T; The temporal processing layer is used to take the features processed by the feature-pulse adaptation layer as input, utilize the spiking neural module SNN of the LIF neuron model, and construct a multi-layer spiking neural module SNN structure to maintain spatial structure information; The feature fusion layer is used to perform multi-scale feature fusion processing according to the temporal processing result by using the simplified spatial pyramid pooling cross-stage partial network module SPPCSPC structure; The object detection layer is used to locate the inner ear malformation area according to the fusion result by combining the two-stage detection strategy and the heat map mechanism.
[0026] The beneficial effects of the above further solution are as follows: The SpikeYOLO model demonstrates superior performance in object detection tasks by integrating the YOLOv7 backbone network, SNN temporal processing, multi-scale feature fusion, heat map mechanism, and two-stage detection strategy. It can not only effectively process spatial information but also enhance the perception of temporal changes, and is particularly suitable for tasks that require considering both spatial and temporal information, such as the location of inner ear malformation areas. Through this innovative design, the SpikeYOLO model shows significant advantages in terms of accuracy, efficiency, and robustness. Brief Description of the Drawings
[0027] Figure 1 It is the flowchart of the method of the present invention.
[0028] Figure 2 It is the schematic diagram of the binary classification result.
[0029] Figure 3 It is the schematic diagram of the LIF neuron model in the SNN module. Detailed Embodiments
[0030] The following describes the detailed embodiments of the present invention to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0031] Embodiment Before describing the present invention, the following terms are described first: AUC: The area enclosed by the ROC curve and the coordinate axes, which is between 0 and 1. It is usually used to evaluate the performance of a binary classification model. The larger the AUC value, the better the model performance, as Figure 2 shown; Accuracy = Number of correct predictions / Total number = TP + TN / (TP + TN + FT + FN); Recall = TP / (TP + FN); Precision = TP / (TP + FP), where TP (True Positive), FP (False Positive), FN (False Negative), and TN (True Negative).
[0032] SpikeYOLO model: A YOLO model based on spikes.
[0033] SNN module (Spiking Neural Network): Spiking neural network.
[0034] Transition module Transition_Block: A module in the YOLO7 source code.
[0035] LIF neuron model: Leaky Integrate-and-Fire neuron, that is, a leaky integration and firing neuron.
[0036] As Figure 1 shown, the present invention provides a method for detecting inner ear malformations, and its implementation method is as follows: S1. Obtain inner ear image sample data, preprocess the inner ear image samples, and perform segmentation and annotation. Its implementation method is as follows: S101. Obtain inner ear image sample data; S102. Preprocess the inner ear image samples, and perform standardization and normalization on the preprocessed inner ear image samples; S103. Based on the normalized inner ear images, annotate the anatomical structure features and malformed parts of the inner ear; S104. Based on the annotated inner ear images, generate object detection labels for YOLOv7 network training, and divide them into training set, validation set, and test set to obtain preprocessed training sample data, completing the segmentation and annotation process.
[0037] In this embodiment, the relevant sample data of the inner ear images includes different temporal bone CT / MRI images of different patients or normal people, as well as temporal bone CT / MRI images of the same patient or normal people at different angles.
[0038] In this embodiment, data preprocessing is performed on the relevant sample data, including image size adjustment and data augmentation; the preprocessed training sample data is standardized and normalized to meet the network input requirements; professional physicians are relied on to label the enhanced image data, marking the anatomical structure features and malformed parts of the inner ear; target detection labels for network training are generated, including class labels, bounding box regression labels, and confidence labels; the labeled image data is divided into a training set, a validation set, and a test set to obtain the preprocessed training sample data.
[0039] S2. Improve the YOLOv7 backbone network and fuse the spiking neural module SNN, and use the inner ear image sample data processed by S1 to train the improved YOLOv7 network to obtain the SpikeYOLO model. The implementation method is as follows: S201. Improve the YOLOv7 backbone network and set up a lightweight feature pyramid network. The implementation method is as follows: S2011. Lightweight improvement of the YOLOv7 backbone network; S2012. Optimize the feature extraction module of the YOLOv7 network, retain and simplify the multi-branch stacked structure Multi_Concat_Block. The implementation method is as follows: A1. Optimize the feature extraction module of the YOLOv7 network; A2. Retain the feature transfer path, simplify the four-branch stacked structure into a two-branch stacked module, and configure a set of convolution normalization activation functions and three sets of convolution normalization activation functions for the two-branch stacked module; A3. Implement channel number compression through 1×1 convolution operations and introduce a residual connection machine; S2013. Improve the connection method of the transition module Transition_Block, and adjust the receptive field size and feature map resolution according to the inner ear structure features; S2014. Simplify the structure of the spatial pyramid pooling cross-stage partial network module SPPCSPC. The implementation method is as follows: B1. Reconstruct the original spatial pyramid pooling structure: simplify the original five-level pyramid pooling layers of 1×1, 3×3, 5×5, 7×7, and 13×13 into a three-level pyramid pooling layer structure of 3×3, 5×5, and 7×7. Among them, each pyramid pooling layer uses the maximum pooling operation; B2. Optimize the feature fusion method: use the form of weighted summation to replace the original concatenation operation; S202. Set up and integrate the spiking neural module SNN, and its implementation method is as follows: C1. Construct spiking neurons based on the LIF neuron model; C2. Set the trainable membrane potential threshold parameter; C3. Introduce spiking neurons in the feature extraction layer to replace the scaled exponential linear unit SiLU activation function in the original YOLO network framework, and set the time-domain joint optimization strategy to integrate the spiking neural module SNN with the YOLO network framework. Among them, time-domain optimization includes dynamically adjusting the neuron firing threshold; spatial-domain optimization includes adjusting the feature map and receptive field size; In this embodiment, the present invention combines the spiking neural module SNN and the yolov7 framework to form a new network, which involves some contents such as spike coding and the change of the neural network activation function. Among them, the firing threshold of the spiking neural module SNN is used as the activation function.
[0040] C4. Set the adaptive membrane potential reset mechanism and introduce separable depth convolution; S203. Optimize the loss function and perform end-to-end training on the improved YOLOv7 network by the gradient backpropagation method to complete the improvement of the YOLOv7 network; S204. Use the training set to train the improved YOLOv7 network and use the validation set for validation to obtain the SpikeYOLO model, and its implementation method is as follows: Set the threshold parameter for object detection, and define the detection box with an intersection over union greater than the threshold as a valid target box; During the training of the improved YOLOv7 network on the test set, introduce the non-maximum suppression mechanism to eliminate duplicate detection boxes; Construct a heatmap visualization module, which is used to represent the attention of the SpikeYOLO model to different regions of the inner ear image; Set a two-stage detection strategy, and the two-stage detection strategy is: perform inner ear region localization on the input inner ear image, and perform malformation feature analysis in the localization region, compare the localization result with the true lesion position, and according to the comparison result, evaluate the performance of the SpikeYOLO model through the coincidence degree between the brightest region of the heatmap and the true lesion position to complete the construction of the SpikeYOLO model. Among them, the result of the malformation feature analysis can output the malformation type judgment result and generate a quantitative description of the malformation feature; it can also be used in the model to adjust the detection strategy of the model and guide the parameter update of the model.
[0041] In this embodiment, the SpikeYOLO model includes: Feature extraction layer, which is used to adopt the lightweight-improved YOLOv7 backbone network and perform feature extraction through the simplified multi-branch stacked structure Multi_Concat_Block and the improved transition module Transition_Block; Feature-spike adaptation layer, which is used to set a feature normalization module to map the CNN features to the interval [0, 1]; construct a spike coding module, adopt a frequency coding method, and set a fixed time window T. In this embodiment, in the SpikeYOLO model, the processing results of the feature-spike adaptation layer are mainly applied to the temporal processing layer. In the temporal processing layer, the spiking neural network (SNN) module will use the features processed by the feature-spike adaptation layer as input to perform temporal spiking neural processing. Through the spiking neural network model, the original features will be transmitted in the form of spikes and processed by LIF neurons to retain spatial and temporal structure information.
[0042] Temporal processing layer, which is used to take the features processed by the feature-spike adaptation layer as input, use the spiking neural module SNN of the LIF neuron model to construct a multi-layer spiking neural module SNN structure, and retain spatial structure information; Feature fusion layer, which is used to perform multi-scale feature fusion processing according to the temporal processing results by using the simplified spatial pyramid pooling cross-stage partial network module SPPCSPC structure; Object detection layer, which is used to locate the inner ear malformation area according to the fusion result by combining the two-stage detection strategy and the heat map mechanism.
[0043] In this embodiment, the present invention first optimizes the Backbone and Neck networks in the YOLOv7 network, that is, performs lightweight improvement on the YOLOv7 backbone network. On the premise of retaining the core detection ability of the YOLO architecture, the network structures of the Backbone and Neck parts are optimized to reduce the network complexity. In addition, in order to adapt to the model, the present invention also needs to optimize the specific modules in the YOLOv7 network, that is, optimize the network feature extraction module, retain and simplify the multi-branch stacked structure Multi_Concat_Block structure, and strengthen the feature fusion ability; improve the connection method of the transition module Transition_Block, and adjust the receptive field size and feature map resolution according to the inner ear structure characteristics to improve the feature transfer efficiency. Finally, improve the spatial pyramid pooling cross-stage partial network module SPPCSPC structure in the original network, that is, simplify the spatial pyramid pooling cross-stage partial network module SPPCSPC structure, and reduce the computational overhead while maintaining the multi-scale feature extraction ability. Specifically, the feature fusion calculation formula of the simplified SPPCSPC module is as follows: ; where represents the output inner ear image after fusion,N Indicates the number of pooling branches, which is simplified to 3, Indicates the weight coefficient, Indicates the i th pooling operation, Indicates the input inner ear image.
[0044] In this embodiment, as Figure 3 shown, a pulsed neuron based on the Leaky Integrate-and-Fire (LIF) model is constructed to achieve temporal coding of information; specifically, the LIF neuron model is represented by the following formula: ; ; ; where, Indicates the integrated time information and spatial information of the membrane potential, Indicates the resting voltage, Indicates the actual pulse output, Indicates the decay coefficient, Indicates the current membrane voltage, Indicates the Heaviside step function, which is 1 when and 0 otherwise. When exceeds the firing threshold , the pulsed neuron emits a pulse , and the output is reset to the resting voltage . If no pulse is emitted, directly decays to .
[0045] In this embodiment, a pulsed neuron based on the Leaky Integrate-and-Fire (LIF) model is constructed to achieve temporal coding of information; a trainable membrane potential threshold parameter is designed to improve the adaptability of the network; the pulsed neural module SNN is fused with the YOLO framework, and pulsed neurons are introduced in the feature extraction layer to replace the scaled exponential linear unit SiLU activation function in the original framework; a spatio-temporal domain joint optimization strategy is designed to balance the detection accuracy and computational efficiency; the performance of the pulsed neural module SNN is optimized, and an adaptive membrane potential reset mechanism is designed to avoid the problem of neuron inactivation; separable depth convolution is introduced to reduce the number of parameters and improve the efficiency of information transmission. Specifically, the calculation method of separable depth convolution can be represented by the following formula, where is the input feature map, is the depth convolution kernel, is the pointwise convolution kernel, and * represents the convolution operation: ; where, indicates the output feature map, Denotes depth convolution, Denotes pointwise convolution.
[0046] Among them, the mathematical expression of the detailed separable depth convolution is as follows: ; Among them, Denotes the value of the output feature map at the n th output channel and position . h And w Denote the height and width positions of the feature map respectively, c Denotes the number of channels, C Denotes the total number of channels, i, j Denotes the sliding window position index in the convolution operation, And Denote the height and width of the convolution kernel respectively, Denotes the value of the input feature map at the c th input channel and position . Denotes the c th channel's depth convolution kernel's weight at position , Denotes the pointwise convolution kernel's weight at the n th output channel and the c th input channel.
[0047] Compared with the standard convolution operation, the number of parameters of the separable depth convolution is reduced from to , significantly reducing the computational complexity.
[0048] In this embodiment, the expression of the loss function is as follows: ; ; ; ; ; ; ; Among them, Denotes the loss function of the improved YOLOv7 network, And Denote the balance coefficients, Denotes the detection loss, Denotes the impulse loss, Denotes the bounding box regression loss, Denotes the confidence loss for having a target, Denotes the confidence loss for not having a target, Denotes the class prediction loss, Denotes the time step, t Denotes the moment, Denotes the actual impulse output, Indicates the expected pulse output, Indicates the weight coefficient for coordinate prediction, Indicates the L2 norm, Indicates the center coordinates of the predicted bounding box, Respectively indicate the width and height of the predicted bounding box, Indicates the presence of a target in the grid cell, Indicates the predicted confidence score, Indicates the true value of the target confidence, Indicates the weight coefficient for the target-free region, Indicates to only calculate the grid cells with targets, Indicates the predicted class probability, Indicates the true probability value of the class prediction, Indicates the center coordinates of the true bounding box, Indicates the width and height of the true bounding box.
[0049] In this embodiment, the preprocessed relevant sample data is input into the network for SpikeYOLO model training. The process is as follows: Set the threshold parameter for object detection, and define the detection box with an intersection over union (IOU) greater than 0.5 as a valid target box; Introduce the non-maximum suppression (NMS) mechanism during training to eliminate duplicate detection boxes; First, sort the detection boxes in descending order according to the confidence score, select the detection box with the highest score, calculate its IOU with other detection boxes, and remove the redundant boxes with an IOU greater than the set threshold. Finally, repeat the above process until all detection boxes are processed; Construct a heatmap visualization module to characterize the attention of the model to different regions of the image; Adopt a two-stage detection strategy; The model will perform inner ear region localization on the input image, and then perform malformation feature analysis within the localized region. The localization result is compared and verified with the true lesion position (marked with a red box) annotated by experts. The performance of the model is evaluated by the coincidence degree between the brightest region of the heatmap and the true lesion position.
[0050] S3. Analyze the inner ear image sample data using the SpikeYOLO model; S4. Determine whether the analysis result meets the preset conditions. If so, enter S5; otherwise, return to S1; In this embodiment, the present invention uses the AUC score to determine whether the analysis result meets the preset conditions. Specifically, in the present invention, meets the preset conditions.
[0051] S5. Analyze the inner ear image to be detected using the SpikeYOLO model to obtain the detection result.
[0052] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting inner ear malformation, characterized in that: The following steps are involved: S1, obtaining inner ear image sample data, preprocessing the inner ear image samples, and performing segmentation and labeling; S2, improve the YOLOv7 backbone network and integrate the spike neural module SNN, and use the inner ear image sample data processed by S1 to train the improved YOLOv7 network to obtain the SpikeYOLO model; S3. Use the SpikeYOLO model to analyze the inner ear image sample data; S4, determine whether the analysis result meets the preset conditions, if so, go to S5, otherwise, return to S1; S5. Use the SpikeYOLO model to analyze the inner ear image to be detected and obtain the detection result.
2. The inner ear malformation detection method according to claim 1, characterized in that: The S1 comprises the following steps: S101, obtaining inner ear image sample data; S102, preprocessing the inner ear image samples, and standardizing and normalizing the preprocessed inner ear image samples; S103, based on the normalized inner ear image, marking the anatomical structure features and deformity parts of the inner ear; S104: Based on the annotated inner ear image, generate target detection labels for YOLOv7 network training, and divide them into training set, validation set and test set to obtain preprocessed training sample data, and complete segmentation and labeling processing.
3. The inner ear malformation detection method according to claim 2, characterized in that: The S2 comprises the following steps: S201. Improve the YOLOv7 backbone network and set up a lightweight feature pyramid network; S202, setting and fusing a spike neural module SNN; S203, optimizing the loss function, and performing end-to-end training on the improved YOLOv7 network through the gradient back propagation method to complete the improvement of the YOLOv7 network; S204, training the improved YOLOv7 network using the training set and verifying it using the verification set to obtain a SpikeYOLO model.
4. The inner ear malformation detection method according to claim 3, characterized in that: The S201 includes the following steps: S2011, lightweight improvement of YOLOv7 backbone network; S2012, optimize the YOLOv7 network feature extraction module, retain and simplify the multi-branch stacking structure Multi_Concat_Block; S2013. Improve the connection mode of the transition module Transition_Block, and adjust the receptive field size and feature map resolution according to the inner ear structure characteristics; S2014, simplified spatial pyramid pooling cross-stage local network module SPPCSPC structure.
5. The inner ear malformation detection method according to claim 4, characterized in that: The S2012 comprises the following steps: A1. Optimize the YOLOv7 network feature extraction module; A2. Retain the feature transfer path, simplify the four-branch stacking structure into a two-branch stacking module, and configure one set of convolutional normalization activation functions and three sets of convolutional normalization activation functions for the two-branch stacking module; A3. Channel number compression is achieved through 1×1 convolution operation, and the residual connection mechanism is introduced.
6. The inner ear malformation detection method according to claim 4, characterized in that: The S2014 comprises the following steps: B1. Reconstruct the original spatial pyramid pooling structure: simplify the original 1×1, 3×3, 5×5, 7×7, 13×13 five-level pyramid pooling layer into a 3×3, 5×5, 7×7 three-level pyramid pooling layer structure, where each pyramid pooling layer uses the maximum pooling operation; B2. Optimize the feature fusion method: Use weighted summation to replace the original series operation. The expression of feature fusion is as follows: in, represents the inner ear image output after fusion, N represents the number of pooling branches, represents the weight coefficient, Indicates i A pooling operation, Represents the input inner ear image.
7. The inner ear malformation detection method according to claim 3, characterized in that: The S202 comprises the following steps: C1. Construct a spiking neuron based on the LIF neuron model. The expression of the LIF neuron model is as follows: in, Indicates integrated time information and spatial information The membrane potential, represents the resting voltage, Indicates the actual pulse output, represents the decay coefficient, represents the current membrane voltage, represents the Heaviside step function; C2, set trainable membrane potential threshold parameters; C3. Introduce spike neurons in the feature extraction layer to replace the scaled exponential linear unit SiLU activation function in the original YOLO network framework, and set a time domain joint optimization strategy to fuse the spike neural module SNN and the YOLO network framework. The time domain optimization includes dynamically adjusting the neuron firing threshold; the spatial domain optimization includes adjusting the feature map and the receptive field size. C4. Setting an adaptive membrane potential reset mechanism and introducing separable depthwise convolution, wherein the expression of the separable depthwise convolution is as follows: in, Indicates that the output feature map is n Output channels, positions The value at h and w Respectively represent the height and width position of the feature map, c Indicates the number of channels, C Indicates the total number of channels, i, j Represents the sliding window position index in the convolution operation, and Represent the height and width of the convolution kernel respectively, Indicates that the input feature map is c Input channels, positions The value of Indicates c The depth convolution kernel of the channel is at position The weight of Represents the point-by-point convolution kernel in n output channels and c The weights of the input channels.
8. The inner ear malformation detection method according to claim 3, characterized in that: The expression of the loss function in S203 is as follows: in, Represents the loss function of the improved YOLOv7 network, and represents the balance coefficient, represents the detection loss, represents the pulse loss, represents the bounding box regression loss, Indicates that there is target confidence loss, represents no target confidence loss, represents the category prediction loss, represents the time step, t Indicates the time, Indicates the actual pulse output, represents the expected pulse output, represents the weight coefficient of coordinate prediction, represents the L2 norm, represents the center coordinates of the predicted bounding box, Represent the width and height of the predicted bounding box respectively, Indicates that there is a target in the grid cell. represents the prediction confidence score, represents the true value of the target confidence, represents the weight coefficient of the target-free area, Indicates that only the grid cells with targets are calculated. represents the predicted class probability, Represents the true probability value of the category prediction, represents the center coordinates of the true bounding box, Represents the true bounding box width and height.
9. The inner ear malformation detection method according to claim 2, characterized in that: The step S204 includes the following steps: Set the threshold parameter for target detection and define the detection box with an intersection-over-union ratio greater than the threshold as a valid target box; In the process of training the improved YOLOv7 network on the test set, the non-maximum suppression mechanism is introduced to eliminate repeated detection boxes; Constructing a heat map visualization module, which is used to characterize the attention paid by the SpikeYOLO model to different regions of the inner ear image; A two-stage detection strategy is set up, and the two-stage detection strategy is as follows: the inner ear area of the input inner ear image is located, and the deformity feature analysis is performed in the located area, and the located result is compared with the actual lesion position. According to the comparison result, the performance of the SpikeYOLO model is evaluated by the overlap between the brightest area of the heat map and the actual lesion position, and the construction of the SpikeYOLO model is completed.
10. The inner ear malformation detection method according to claim 9, characterized in that: The SpikeYOLO model includes: The feature extraction layer is used to extract features using a lightweight and improved YOLOv7 backbone network through a simplified multi-branch stacking structure Multi_Concat_Block and an improved transition module Transition_Block; The feature-pulse adaptation layer is used to set the feature normalization module and map the CNN features to the [0,1] interval; construct the pulse coding module, use the frequency coding method, and set a fixed time window T; The time series processing layer is used to take the features processed by the feature-pulse adaptation layer as input, and use the pulse neural module SNN of the LIF neuron model to build a multi-layer pulse neural module SNN structure to maintain the spatial structure information; The feature fusion layer is used to perform multi-scale feature fusion processing based on the time series processing results using the simplified spatial pyramid pooling cross-stage local network module SPPCSPC structure; The target detection layer is used to locate the inner ear malformation area based on the fusion results, combining the two-stage detection strategy and the heat map mechanism.