Neural fiber twist grading system based on multi-example learning and channel attention

Through a multi-example learning and channel attention, the subjectivity and noise impact of corneal nerve fiber twist assessment in the prior art is solved, and high accuracy and generalization ability are improved.

CN120337988APending Publication Date: 2025-07-18YANGTZE DELTA REGION INST OF UNIV OF ELECTRONICS SCI & TECH OF CHINE (HUZHOU)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510265748.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing corneal nerve fiber twisting evaluation method is subjective, time-consuming and labor-intensive, and is not suitable for large-scale clinical applications. The deep learning model ignores the sequence dependence of the entire nerve fiber and is susceptible to imaging noise, resulting in poor classification accuracy and generalization.

Method used

A neural fiber twist hierarchy system based on multi-example learning and channel attention is used to generate neural fiber sequences through skeletonization, segmentation, depth-first search and differential coding, and a one-dimensional convolutional neural network, channel attention and gated recurrent unit network are combined to extract multi-scale features for hierarchical prediction.

Benefits of technology

It improves the accuracy and generalization ability of nerve fiber twist degree classification, effectively captures the sequence dependence of nerve fibers and reduces the impact of noise, and improves the prediction performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337988A_ABST
    Figure CN120337988A_ABST
Patent Text Reader

Abstract

The invention discloses a nerve fiber twist grading system based on multi-instance learning and channel attention, which comprises a curve sequence conversion module and a multi-scale feature extraction module, and is characterized in that the curve sequence conversion module comprises a skeletonization module, a segmented nerve fiber module and a conversion module; the multi-scale feature extraction module comprises a multi-instance learning module and a group of neural network modules, and the neural network modules comprise a feature extraction module, a feature enhancement module, a global feature extraction module and a gating circulation unit network module; according to the method, the convolutional neural network, the channel attention module and the gating circulation unit network are combined, so that accurate grading of the nerve fiber torsion degree is realized; through curve coordinate differential coding and multi-scale feature splicing, morphological information of a main nerve fiber structure can be effectively captured, and meanwhile, a channel attention module and an example feature aggregation strategy improve the robustness and generalization ability of classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a nerve fiber distortion grading system based on multi-instance learning and channel attention. Background Art

[0002] In medical imaging, the confocal corneal microscope (CCM) is a non-invasive corneal imaging technology widely used in the fields of ophthalmology and neurology. The CCM technology can image the sub-basal nerve plexus of the cornea quickly and with high resolution, so as to clearly observe the morphological features of the nerve fibers in the cornea. Based on the CCM technology, doctors can observe the nerve fiber structure features that cannot be presented by traditional fundus color photos, especially in terms of the distortion, branching, density, etc. of the corneal nerve fibers. Observing and qualitatively analyzing the nerve parameters under the cornea is of great clinical value for detecting diseases such as diabetic neuropathy, dry eye disease, and fungal keratitis. These diseases usually show changes in the distortion of corneal nerve fibers. Therefore, grading and classifying the distortion of corneal nerves in normal and diseased images helps doctors initially judge the degree of nerve lesions in patients and provides support for clinical decision-making. Research shows that dividing the distortion of corneal nerves into 4 grades can significantly reveal the changes in corneal nerve fibers in patients with diabetic neuropathy and dry eye disease.

[0003] The methods for analyzing the distortion of corneal nerve images have evolved from traditional manual measurement and visual assessment to automatic analysis through machine learning and neural networks. Earlier distortion assessment methods usually relied on manual marking and analysis, such as measuring the distortion of corneal nerve fibers by manually tracing them. Although these methods are effective in small-scale analysis, they are not suitable for large-scale clinical applications because the evaluation process is highly subjective, time-consuming and laborious, and the consistency among different observers is poor.

[0004] Therefore, automated nerve distortion assessment methods have gradually become mainstream. The early steps for automatically calculating blood vessel distortion included preprocessing, segmentation, blood vessel network separation, and curvature calculation. However, the errors introduced in these processing steps may accumulate and lead to information loss, and the algorithm and parameters need to be frequently adjusted when facing different image qualities and devices, with a cumbersome process and poor generalization ability.

[0005] In recent years, the rise of deep learning methods has provided new solutions for the automated assessment of corneal nerve tortuosity. Deep learning models based on convolutional neural networks can learn features from images end-to-end, automatically extract the morphological features and spatial information of nerve fibers, and perform classification. However, due to the limited receptive field of convolutional neural networks, these methods usually ignore the sequential features of the entire corneal nerve fiber structure. Modeling only based on local spatial features may be difficult to fully capture the sequential dependencies between different nerve fibers. In addition, the imaging noise in CCM images and the complex morphological changes of nerve fibers are likely to cause problems of overfitting and boundary uncertainty, affecting the accuracy and generalization of the model for tortuosity classification. Summary of the Invention

[0006] To solve the deficiencies of the prior art and achieve the purpose of improving the accuracy of tortuosity classification and generalization for different datasets, the present invention adopts the following technical solutions:

[0007] A nerve fiber tortuosity grading system based on multi-instance learning and channel attention, including a curve sequence conversion module and a multi-scale feature extraction module. The curve sequence conversion module includes a skeletonization module, a segmented nerve fiber module, and a conversion module. The multi-scale feature extraction module includes a multi-instance learning module and a group of neural network modules. The neural network modules include a feature extraction module, a feature enhancement module, a global feature extraction module, and a gated recurrent unit network module;

[0008] The skeletonization module obtains the segmented nerve fiber image and skeletonizes it. The segmented nerve fiber module uses the characteristic that the width of the nerve fiber in the skeletonized segmentation map is a single pixel, and uses a convolutional matrix to identify and remove the branch points of the nerve in the skeletonized segmentation map to obtain independent nerve fiber curve segments. Then, the conversion module performs a depth-first search on the segmented nerve fiber map, extracts the coordinates of each single-pixel curve, and converts the extracted coordinates using differential coding, representing each coordinate point of the curve as the difference between it and the previous coordinate point to generate a nerve fiber sequence;

[0009] The feature extraction module takes a sequence as an example, extracts the morphological features of a nerve fiber sequence based on a single example, and then the feature enhancement module enhances the features. The global feature extraction module extracts global features through the enhanced features. The gated recurrent unit network module processes the enhanced features curve by curve and extracts long-time series features based on time steps. The enhanced features are also input into the next neural network module for feature extraction. Based on the global features and long-time series features generated by each neural network module, multi-scale example features are obtained, which represent the spatial information and time-dependent information of the input curve sequence at different scales, providing a richer sequence feature representation. The multi-instance learning module aggregates the multi-scale example features along the sequence dimension. The aggregated bag-level features not only integrate the features of the same example at different learning scales but also fuse the feature information across examples, forming a global feature representation that can effectively predict the bag-level label. The multi-scale feature extraction module constructs a loss based on the predicted and true nerve distortion levels for backpropagation, updating the parameters in the feature extraction module, feature enhancement module, and gated recurrent unit network module, so that the trained system can be used for the grading prediction of nerve distortion.

[0010] Further, the segmented nerve fiber module obtains the binary skeleton map of the nerve fiber. After filling the boundary of the binary image with 0, the convolution convolve function is used to calculate the sum of all pixels in the 3*3 neighborhood (out-of-bounds is regarded as 0) of each pixel, and then the coordinates where the sum of the 3*3 neighborhood pixels is greater than the threshold 3 are judged. This pixel point is judged as the branch point of the curve, and a position mark is made on the binary image, and then an inversion operation is performed after binarization to remove the branch points in the skeleton nerve fiber. After multiplying each pixel by the skeletonized image, the segmented nerve fiber map is obtained.

[0011] Further, the conversion module performs a depth-first search on the segmented nerve fiber map, recursively traverses the neighborhood of each pixel point to ensure that all connected coordinate points can be completely extracted along the entire path of the curve. After outputting the coordinate sequence of each segment of the nerve fiber curve, starting from the head coordinate, differential encoding is performed on the sequence and the same sequence length and sequence number are defined for each set of sequences.

[0012] Further, the feature extraction module includes a convolutional layer, a normalization layer, and an activation layer. After the nerve fiber sequence is input into the convolutional layer, it passes through a normalization layer with the same output dimension as the convolutional layer, and finally obtains morphological features through the activation layer.

[0013] Further, the feature enhancement module is a channel attention module for enhancing features. The morphological feature F1 is used to calculate the attention weight w through the channel attention module c to obtain the enhanced feature F1 ′ to strengthen the attention to the key coordinates in each sequence. The process is as follows:

[0014] Perform global average pooling on each input channel to obtain the global feature representation z of each channel c ;

[0015] Execute the channel recalibration operation through a two-layer fully connected network to process the global feature representation z of each channel c and generate the attention weights w between channels c ;

[0016] After adjusting the attention weights w of each channel to the same size as the input morphological feature F1, perform element-wise multiplication on each channel to obtain the enhanced feature F1 c ′ .

[0017] Furthermore, the generation of the attention weights w c is achieved by reducing the dimension through the first-layer fully connected network, activating with the ReLU activation function, increasing the dimension through the second-layer fully connected network, and activating with the Sigmoid activation function. The formula is as follows:

[0018] w c = σ(W2δ(W1z c ))

[0019] where and respectively represent the weight matrices of the two-layer fully connected layers, δ represents the ReLU activation function, σ represents the Sigmoid activation function, and r represents the dimensionality reduction ratio.

[0020] Furthermore, the global feature extraction module obtains the global feature after performing adaptive average pooling on the enhanced feature.

[0021] Furthermore, the gated recurrent unit network module processes the enhanced feature curve by curve, dynamically updates the hidden state, and at each time step, based on the current input curve feature and the hidden state of the previous time step, generates a new hidden state through the update gate and reset gate mechanisms and passes it to the next time step. After multiple time steps of iteration, the gated recurrent unit network extracts the long time series feature.

[0022] Furthermore, the multi-scale feature extraction module inputs the enhanced feature in each neural network module into the next neural network module. After training with multiple neural network modules, the global feature and the long time series feature are obtained, and then through the concatenation operation, a set of global features and long time series features are gradually concatenated in sequence to obtain the multi-scale example feature.

[0023] Furthermore, the multi-instance learning module processes the multi-scale example feature x bag ​Perform overall average aggregation along the sequence dimension. The feature extraction module extracts features in the form of sequences as examples and images as bags.

[0024]

[0025] Where m represents the number of global features and long-time series features included in the multi-scale example features, and n represents the number of examples contained in each bag.

[0026] The advantages and beneficial effects of the present invention are as follows:

[0027] In the serialization stage of the nerve fiber curve, by segmenting the skeletonized nerve fiber segmentation map, the present invention removes the influence of different types of noise existing in the CCM image on the generalization of the model, and at the same time reduces the scale of the input sequence, making up for the problem of insufficient receptive field of the one-dimensional convolutional neural network; in the design of the deep neural network structure, the present invention integrates the advantages of the one-dimensional convolutional neural network, the channel attention module and the gated recurrent unit network, makes full use of the excellent performance of the convolutional neural network in local feature extraction, and at the same time combines the advantages of the gated recurrent unit network in capturing long-distance dependencies and the improvement of computational efficiency brought by fewer parameters. Through the channel attention module, the feature channels are dynamically weighted according to importance, enabling the network to classify and predict the input image more efficiently; the present invention utilizes the multi-instance learning framework to capture the global information of different examples and the same example at different scales, effectively aggregates the features of a single curve sequence into the bag-level features of the entire image, enriches the feature representation, and improves the accuracy and generalization ability of model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the system structure of an embodiment of the present invention.

[0029] Figure 2a It is the segmented nerve fiber map in an embodiment of the present invention.

[0030] Figure 2b It is the visual diagram of curve serialization in an embodiment of the present invention.

[0031] Figure 3 It is the display diagram of the distortion grading result on the CCM data set in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0032] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.

[0033] Such as Figure 1As shown in the figure, a nerve fiber distortion grading system based on multi-instance learning and channel attention includes an input module, a curve sequence conversion module, and a multi-scale feature extraction module. The curve sequence conversion module includes a skeletonization module, a segmented nerve fiber module, and a conversion module. The multi-scale feature extraction module includes a multi-instance learning module and a set of neural network modules. The neural network module includes a feature extraction module, a feature enhancement module, a global feature extraction module, and a gated recurrent unit network module.

[0034] The dataset uses the existing open-source corneal confocal microscopy image dataset CORN (COrneal Nerve Database): CORN-3, CORN 1500 And set the training set and validation set of CORN-3 in a ratio of 3:1. At the same time, set the training set and validation set of CORN 1500 in a ratio of 4:1, and directly use the images with a size of 384*384 in the dataset.

[0035] The input module extracts CCM rectangular images with a size of 384*384 and binary segmentation maps from the open-source corneal confocal microscopy (CCM) image dataset.

[0036] The skeletonization module uses the skeletonization function in the skimage library to process the segmentation map. Through the skeletonization operation, the pre-segmented CCM image is processed to obtain a binary skeleton map of nerve fibers; the skeletonization operation ensures that the width of a single curve is 1 at any position.

[0037] The segmented nerve fiber module uses the characteristic that the width of the nerve fibers in the skeletonized segmentation map is single-pixel, and uses a convolutional matrix to identify and remove the branch points in the nerve to obtain independent nerve fiber curve segments; the specific process of segmenting the nerve fiber curve is as follows:

[0038] After filling the boundary of the binary image with "0", use the convolution convolve function to calculate the sum of all pixels in the 3*3 neighborhood (out-of-bounds is regarded as 0) of each pixel, and then judge the coordinates of those 3*3 neighborhood pixel sums greater than 3, and perform a binary inversion operation on them, and multiply each pixel by the skeletonized image to obtain the segmented nerve fiber map as shown in Figure 2a In the figure. In the present invention, by constructing a 3*3 detection matrix, the number of adjacent points of each pixel is counted. When the number of adjacent points plus the pixel itself is greater than 3, this pixel point is judged as a branch point of the curve, and a position mark is made on the binary image. According to the binary image of the branch point, the branch points in the skeleton nerve fibers are removed, so that the curve is segmented.

[0039] In the conversion module, the segmented nerve fiber map is subjected to depth-first search (DFS) to extract the coordinates of each single-pixel curve, and the coordinates are saved in a set to form the complete path of the curve. The extracted coordinate data is converted using differential encoding, and each coordinate point of the curve is represented as the difference from the previous point to obtain the nerve fiber sequence. The original coordinate data is compressed into a sequence of relative coordinates, effectively retaining its spatial structure information. The process of converting the nerve fiber curve to a sequence is as follows:

[0040] As Figure 2b shown, the segmented nerve fiber map is input into a custom depth-first search module, which recursively traverses the neighborhood of each pixel point to ensure that all connected coordinate points can be completely extracted along the entire path of the curve. After outputting the coordinate sequence of each segment of the nerve fiber curve, starting from the "head" coordinate, differential encoding is performed on the sequence and the same sequence length and sequence number are defined for each set of sequences. In this embodiment, differential encoding saves the "head" coordinate of each segment of the curve, and each coordinate is represented as an integer difference from the previous coordinate within the range of [-1, 0, 1]. The sequence length and sequence number are 256 and 64 respectively.

[0041] For the feature extraction module, the nerve fiber sequence is input into a one-dimensional convolutional network for feature extraction. The sequence is input into the feature extraction module in the form of an example and an image as a package, that is, the Block module in this embodiment, to obtain the morphological feature F1 based on a single example.

[0042] In this embodiment, the preliminary extraction network for curve features is integrated in a Block module. In the Block module, the differentially encoded curve sequence is input into a one-dimensional convolutional network with a kernel size of 3*3 in the first layer, then passed through a BatchNorm normalization layer with the same output dimension as the convolutional layer, and finally passed through a LeakyReLU activation function to obtain the preliminary morphological feature F1. The output size of the hidden layer is 32.

[0043] The feature enhancement module is a channel attention module for enhancing features. The morphological feature F1 is used to calculate the weight w c through the channel attention module to obtain the enhanced feature F1 ′ , strengthening the attention to the key coordinates in each sequence. The global feature extraction module performs adaptive average pooling on the enhanced feature F1 ′ to obtain the global feature F1 ″ .

[0044] In this embodiment, the number of channels C of the morphological feature F1 is 32, and the length L of the sequence is 256. Feature enhancement includes the following steps:

[0045] First, perform global average pooling on each input channel to obtain the global feature representation of each channel; morphological features are obtained through global average pooling to obtain the global feature representation of each channel where, x ic represents the value at the i-th position of the c-th channel in the sequence feature, and z c represents the global average pooling result of this channel;

[0046] Second, perform channel recalibration operations through two small fully connected networks to process the global feature representation z of each channel c to generate the attention weights w between channels c ; the first fully connected network reduces the dimension and uses ReLU activation; the second network increases the dimension and uses Sigmoid activation. The formula is as follows:

[0047] w c = σ(W2δ(W1z c ))

[0048] where, and respectively represent the weight matrices of the two fully connected layers, δ represents the ReLU activation function, σ is the Sigmoid activation function, r represents the reduction ratio, and r = 16 at this time;

[0049] Third, after adjusting the weight w of each channel to the same size as the input morphological feature F1, perform element-wise multiplication on each channel to obtain the enhanced feature F1 c . ′ .

[0050] Enhanced feature F1 ′ obtains the global feature description F1 through the adaptive average pooling layer ″ .

[0051] The gated recurrent unit network module uses the gated recurrent unit network to model sequence dependencies. The obtained enhanced feature F1 ′ is input into the gated recurrent unit (GRU) function in the torch library to model the dynamic changes and dependencies of each curve sequence, and a long time series feature r1 containing the hidden state of each time step and the final hidden state of the last time step are obtained.

[0052] Specifically, the gated recurrent unit network module will process the enhanced feature F1 curve by curve ′, Dynamically update the hidden state. At each time step, the gated recurrent unit network generates a new hidden state based on the current input curve features and the hidden state of the previous time step through the update gate and reset gate mechanisms, and passes it to the next time step. After multiple time steps of iteration, the gated recurrent unit network extracts the long time series feature r1.

[0053] The multi-scale feature extraction module inputs the enhanced feature F output by the channel attention module in each layer of the neural network { ′ 1,2,3+ into the Block module of the next neural network layer jointly composed of a one-dimensional convolutional network and a gated recurrent unit network. After training through four layers of the same neural network layer, the enhanced feature F after pooling is obtained { ′ 1 ′ ,2,3,4+ and the sequence feature output r of the gated recurrent unit network module {1,2,3,4+ , and then use the cat function to concatenate F { ′ 1 ′ ,2,3,4+ and r {1,2,3,4+ step by step in order for a total of 8 features to obtain the multi-scale example feature x bag = concat(F1 ″ , F2 ″ , F3 ″ , F4 ″ , r1, r2, r3, r4). These features represent the spatial information and temporal dependence information of the input curve sequence at different scales, providing a more abundant sequence feature representation.

[0054] The multi-instance learning module introduces the multiple instance learning (MIL) framework to perform overall average aggregation on the multi-scale example feature x bag along the sequence dimension n = 64 represents the number of examples contained in each bag in the multi-instance learning framework. The aggregated bag-level feature x ′ bag not only integrates the features of the same example at different learning scales, but also fuses the feature information across examples, forming a global feature representation that can effectively predict the bag-level label.

[0055] Update the model parameters based on backpropagation: Perform gradient backpropagation based on the classification loss, input the obtained x ′ bag into a linear fully connected layer with an input-output dimension of 4, and use the cross-entropy algorithm according to the logistic regression value output of the fully connected layer after softmax conversion Calculate the loss L between the predicted label and the true label cel , perform gradient backpropagation, and calculate L cel the gradient with respect to the model parameters, and then update the learnable parameters in the three key modules of the one-dimensional convolutional network, channel attention module, and gated recurrent unit network module. Stop after training for 640 rounds, where B represents the batch size, and at this time B = 128, y b,class represents the true label (0 or 1) of sample b on 4 classes class, p b,class represents the predicted probability of the model for class class, that is, the softmax probability output of sample b on class class.

[0056] Finally, obtain the prediction result of the CCM image distortion grading: Use the trained model to predict the nerve distortion level of the CCM image in the target dataset, extract and encode the curve sequence, and successively complete feature extraction, feature aggregation, and grading prediction, and finally obtain the prediction result of the corneal nerve distortion grading.

[0057] In this embodiment, the cross-dataset of the CCM image is predicted according to the trained model. The CCM images in different datasets are input into the trained model and the above operations are repeated. The argmax function is used to process the logistic regression value output of the fully connected layer, and the class corresponding to the maximum logistic regression value obtained is used as the predicted label. The accuracy, sensitivity, and specificity of the three evaluation indicators are jointly calculated by the predicted label and the true label.

[0058] The test results are as Figure 3 shown, where each column from left to right is the result of the present invention (Ours) based on the CORN-3 dataset, the joint training result (Ours(cross-dataset)) based on the CORN-3 and CORN 1500 datasets, and other methods (the multi-instance learning method TransMIL based on the transformer structure, BANet based on bilinear attention, and the classic CNN structure and its bidirectional gated recurrent unit network variant CNN-BiGRU). Figure 3This is a comparison of the predicted distortion level results of the present invention on the CCM dataset. It can be seen from the figure that the accuracy of predicting the distortion level of CCM images using the method of the present invention is significantly higher than that of TransMIL and CNN. The use of the channel attention module in the method of the present invention also makes the prediction performance exceed that of CNN-BiGRU. BANet is a prediction method based on two-dimensional images. The method of the present invention effectively reduces the influence of noise in the image background and improves the prediction accuracy. The joint training results of the dual datasets prove that the method of the present invention has a certain robustness in cross-dataset prediction. In summary, the prediction performance of the method of the present invention for the distortion level is significantly better than other methods.

[0059] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A nerve fiber distortion grading system based on multi-instance learning and channel attention, including a curve sequence conversion module and a multi-scale feature extraction module, characterized in that: The curve sequence conversion module includes a skeletonization module, a segmented nerve fiber module, and a conversion module. The multi-scale feature extraction module includes a multi-instance learning module and a group of neural network modules. The neural network module includes a feature extraction module, a feature enhancement module, a global feature extraction module, and a gated recurrent unit network module; The skeletonization module obtains the segmented nerve fiber image and skeletonizes it. The segmented nerve fiber module uses the characteristic that the width of the nerve fiber in the skeletonized segmentation map is a single pixel to identify and remove the branch points of the nerve in the skeletonized segmentation map, obtaining independent nerve fiber curve segments. Then, the conversion module performs a depth-first search on the segmented nerve fiber map, extracts the coordinates of each single-pixel curve, and converts the extracted coordinates using differential encoding, representing each coordinate point of the curve as the difference from the previous coordinate point to generate a nerve fiber sequence; The feature extraction module takes the sequence as an example and extracts the morphological features of the nerve fiber sequence based on a single example. Then, the feature enhancement module enhances the features. The global feature extraction module extracts global features through the enhanced features. The gated recurrent unit network module processes the enhanced features curve by curve and extracts long-time sequence features based on time steps. The enhanced features are also input into the next neural network module for feature extraction. Based on the global features and long-time sequence features generated by each neural network module, multi-scale example features are obtained. The multi-instance learning module aggregates the multi-scale example features along the sequence dimension. The multi-scale feature extraction module constructs a loss based on the predicted and true nerve distortion levels for backpropagation, updating the parameters in the feature extraction module, the feature enhancement module, and the gated recurrent unit network module, so that the trained system is used for the grading prediction of nerve distortion.

2. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, characterized in that: The segmented nerve fiber module obtains the binary skeleton map of the nerve fiber. After filling the boundary of the binary image with 0, it calculates the sum of all pixels in the neighborhood of each pixel, then determines the coordinates where the neighborhood pixel sum is greater than the threshold, and performs a binary inversion operation on them. After multiplying each pixel by the skeletonized image, the segmented nerve fiber map is obtained.

3. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, characterized in that: The conversion module performs a depth-first search on the segmented nerve fiber map, extracts all connected coordinate points, and after outputting the coordinate sequence of each nerve fiber curve segment, starting from the head coordinate, performs differential encoding on the sequence and defines the same sequence length and sequence number for each set of sequences.

4. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, characterized in that: The feature extraction module includes a convolutional layer, a normalization layer, and an activation layer. After the nerve fiber sequence is input into the convolutional layer, it passes through a normalization layer with the same output dimension as the convolutional layer, and finally obtains morphological features through the activation layer.

5. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, characterized in that: The feature enhancement module is a channel attention module for enhancing features. The morphological feature F1 is used to calculate the attention weight w through the channel attention module c to obtain the enhanced feature F1 ′ The process is as follows: Global average pooling is performed on each input channel to obtain the global feature representation z for each channel c ; The channel recalibration operation is performed through a two-layer fully connected network to process the global feature representation z of each channel c to generate the attention weights w between channels c ; Adjust the attention weight w of each channel c to the same size as the morphological feature F1 of the input, and then perform element-wise multiplication on each channel to obtain the enhanced feature F1 ′ .

6. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 5, characterized in that: The generation of the attention weight w c is achieved by reducing the dimension through the first fully connected network, activating it with the ReLU activation function, then increasing the dimension through the second fully connected network and activating it with the Sigmoid activation function. The formula is as follows: w c = σ(W2δ(W1z c )) Among them, and respectively represent the weight matrices of two fully connected layers, δ represents the ReLU activation function, σ represents the Sigmoid activation function, and r represents the dimensionality reduction ratio.

7. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, characterized in that: The global feature extraction module performs adaptive average pooling on the enhanced features to obtain global features.

8. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, characterized in that: The gated recurrent unit network module processes the enhanced features curve by curve, dynamically updates the hidden state, and at each time step, generates a new hidden state based on the current input curve feature and the hidden state of the previous time step through the update gate and reset gate mechanisms, and passes it to the next time step. After multiple time steps of iteration, the gated recurrent unit network extracts long time series features.

9. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 1, wherein: The multi-scale feature extraction module inputs the enhanced features in each neural network module into the next neural network module. After training with multiple neural network modules, global features and long time series features are obtained, and then through a concatenation operation, a set of global features and long time series features are gradually concatenated in sequence to obtain multi-scale example features.

10. The nerve fiber distortion grading system based on multi-instance learning and channel attention according to claim 9, wherein: The multi-instance learning module aggregates the overall average of the multi-scale instance features x bag along the sequence dimension. The feature extraction module extracts features in the form of sequences as instances and images as bags. Among them, m represents the number of global features and long time series features included in the multi-scale example features, and n represents the number of examples contained in each said package.