A three-dimensional medical image recognition method and system

Feature extraction of three-dimensional medical images through time-series convolution network and two-dimensional convolution leakage integral distribution model solves the problem of real-time and low accuracy of three-dimensional medical image recognition, and achieves efficient lesion recognition.

CN113554581BActive Publication Date: 2025-07-25LYNXI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010316264.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-21
Publication Date
2025-07-25
Estimated Expiration
2040-04-21

AI Technical Summary

Technical Problem

In the prior art, there are problems such as poor real-time recognition of three-dimensional medical images and low recognition accuracy.

Method used

The time-series convolution network is used to extract features of three-dimensional medical images, including time-series convolution processing, integration processing, dimension reduction processing and full connection processing. Combined with the two-dimensional convolution leakage integral distribution model, feature extraction is performed through membrane potential accumulation and integration distribution mechanisms to reduce the computational complexity.

Benefits of technology

It improves the real-time and accuracy of three-dimensional medical image recognition, and reduces the number of network parameters and operation complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113554581B_ABST
    Figure CN113554581B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional medical image recognition method and system, including: preprocessing the medical image to be recognized to obtain a three-dimensional image block to be recognized; extracting features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result. The beneficial effects of the present invention are: it can reduce the number of parameters and the computational complexity, reduce the computational time, and improve the real-time performance and accuracy of recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image recognition, and more particularly, to a three-dimensional medical image recognition method and system. Background Art

[0002] With the development of computer-aided medical image diagnosis technology, medical images can be used to assist doctors in disease diagnosis. However, in related technologies, when identifying lesions in three-dimensional medical images, there are problems of poor real-time performance and low recognition accuracy. Summary of the Invention

[0003] To solve the above problems, the purpose of the present invention is to provide a three-dimensional medical image recognition method and system, which can reduce the number of parameters and computational complexity, reduce the computational time, and improve the real-time performance and accuracy of recognition.

[0004] The present invention provides a three-dimensional medical image recognition method, including:

[0005] Preprocess the medical image to be recognized to obtain a three-dimensional image block to be recognized;

[0006] Extract features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result.

[0007] As a further improvement of the present invention, extracting features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result includes:

[0008] Perform at least one temporal convolution process on the three-dimensional image block to be recognized to obtain a first feature vector;

[0009] Integrate the first feature vector to obtain a second feature vector;

[0010] Reduce the dimension of the second feature vector to obtain a third feature vector;

[0011] Perform at least one fully connected process on the third feature vector to obtain a fourth feature vector;

[0012] Determine the lesion recognition result according to the fourth feature vector.

[0013] As a further improvement of the present invention, performing at least one temporal convolution process on the three-dimensional image block to be recognized to obtain a first feature vector includes:

[0014] Perform a first temporal convolution process on the three-dimensional image block to be recognized to obtain a first temporal convolution vector;

[0015] When performing a temporal convolution process once, determine the first temporal convolution vector as the first feature vector;

[0016] When performing n temporal convolution processes, perform pooling on the first temporal convolution vector to obtain a first intermediate feature vector, perform temporal convolution on the first intermediate feature vector to obtain a second temporal convolution vector, perform pooling on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, determine the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

[0017] As a further improvement of the present invention, perform at least one fully connected process on the third feature vector to obtain a fourth feature vector, including:

[0018] Perform a first fully connected process on the third feature vector to obtain a first intermediate vector;

[0019] When performing a fully connected process once, determine the first intermediate vector as the fourth feature vector;

[0020] When performing m fully connected processes, perform a second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on, determine the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

[0021] As a further improvement of the present invention, the method further includes: training the temporal convolution network according to a training set.

[0022] As a further improvement of the present invention, the temporal convolution network includes: at least one cascaded temporal convolution layer, a three-dimensional pooling layer, a dimensionality transformation layer, and at least one fully connected layer.

[0023] As a further improvement of the present invention, performing feature extraction on the three-dimensional image block to be recognized through the temporal convolution network includes:

[0024] Performing at least one temporal convolution process on the three-dimensional image block to be recognized through the at least one temporal convolution layer to obtain a first feature vector;

[0025] Performing integration processing on the first feature vector through the three-dimensional pooling layer to obtain a second feature vector;

[0026] Performing dimensionality reduction processing on the second feature vector through the dimensionality transformation layer to obtain a third feature vector;

[0027] Performing at least one fully connected process on the third feature vector through the at least one fully connected layer to obtain a fourth feature vector;

[0028] Determine the lesion recognition result according to the fourth feature vector.

[0029] As a further improvement of the present invention, at least one temporal convolution process is performed on the three-dimensional image block to be recognized through the at least one temporal convolution layer, and a first feature vector is obtained, including:

[0030] Performing a first temporal convolution process on the three-dimensional image block to be recognized through a first temporal convolution layer to obtain a first temporal convolution vector;

[0031] When performing one temporal convolution process, determining the first temporal convolution vector as the first feature vector;

[0032] When performing n temporal convolution processes, performing a pooling process on the first temporal convolution vector to obtain a first intermediate feature vector, performing a temporal convolution process on the first intermediate feature vector through a second temporal convolution layer to obtain a second temporal convolution vector, performing a pooling process on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, and determining the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

[0033] As a further improvement of the present invention, at least one fully connected process is performed on the third feature vector through the at least one fully connected layer to obtain a fourth feature vector, including:

[0034] Performing a first fully connected process on the third feature vector through a first fully connected layer to obtain a first intermediate vector;

[0035] When performing one fully connected process, determining the first intermediate vector as the fourth feature vector;

[0036] When performing m fully connected processes, performing a second fully connected process on the first intermediate vector through a second fully connected layer to obtain a second intermediate vector, and so on, and determining the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

[0037] As a further improvement of the present invention, the temporal convolution layer is implemented by using a two-dimensional convolutional spiking neuron model.

[0038] As a further improvement of the present invention, the temporal convolution layer is implemented by using a two-dimensional convolutional leaky integrate-and-fire model.

[0039] As a further improvement of the present invention, the two-dimensional convolutional leaky integrate-and-fire model is used for:

[0040] According to the input value X at time t of the two-dimensional convolutional leaky integrate-and-fire model t After convolution operation with the weight W, the value I is obtained t , after adding to the biological voltage value at time t-1 , the membrane potential value at time t is obtained

[0041] According to the membrane potential value at time t and the firing threshold V th , determine the output value F at time t t ;

[0042] According to the output value F at time t t determine whether to reset the membrane potential, and according to the reset voltage value V reset determine the reset membrane potential value Among them,

[0043] According to the reset membrane potential value determine the biological voltage value at time t Among them, α and β are the leakage factors of the Leak activation function;

[0044] Among them, the output value F at time t t serves as the input to the next layer cascaded with the two-dimensional convolutional leaky integrate-and-fire model, and the biological voltage value at time t serves as the input for calculating the membrane potential value at time t + 1.

[0045] As a further improvement of the present invention, the determining the output value F at time t according to the membrane potential value at time t and the firing threshold V th , includes: t If the membrane potential value at time t

[0046] is greater than or equal to the firing threshold V th , then determine the output value F at time t t to be 1;

[0047] If the membrane potential value at time tis less than the firing threshold V th , then determine the output value F at time t t to be 0.

[0048] As a further improvement of the present invention, training the temporal convolutional network according to the training set includes: obtaining a sample data set;

[0049] Among them, obtaining a sample data set includes:

[0050] Preprocess the obtained medical images to obtain three-dimensional image blocks;

[0051] Construct a sample data set based on the three-dimensional image blocks;

[0052] ​Perform data augmentation on the samples in the sample dataset, and expand the processed data to the sample dataset. Among them, performing data augmentation on the samples in the sample dataset includes: randomly flipping the samples in 3 directions, randomly adjusting the size of the samples, rotating the samples at an arbitrary angle in three-dimensional space, and randomly moving the samples in 3 directions, with at least one of these operations.

[0053] As a further improvement of the present invention, preprocess the medical image to obtain a three-dimensional image block, including:

[0054] Extract features from the medical image to obtain an image of the lesion area;

[0055] Convert the format of the image of the lesion area to obtain a three-dimensional image block in npy format.

[0056] As a further improvement of the present invention, in the sample dataset, 70% of the three-dimensional image blocks contain at least one nodule, and 30% of the three-dimensional image blocks do not contain nodules.

[0057] As a further improvement of the present invention, train the temporal convolutional network according to the training set, including: training the temporal convolutional network through the stochastic gradient descent algorithm and the step learning rate adjustment method.

[0058] The present invention also provides a three-dimensional medical image recognition system, which includes:

[0059] A preprocessing module, which is used to preprocess the medical image to be recognized to obtain a three-dimensional image block to be recognized;

[0060] A feature extraction module, which is used to extract features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result.

[0061] As a further improvement of the present invention, extracting features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result includes:

[0062] Perform at least one temporal convolution process on the three-dimensional image block to be recognized to obtain a first feature vector;

[0063] Integrate the first feature vector to obtain a second feature vector;

[0064] Perform dimensionality reduction on the second feature vector to obtain a third feature vector;

[0065] Perform at least one fully connected process on the third feature vector to obtain a fourth feature vector;

[0066] Determine the lesion recognition result according to the fourth feature vector.

[0067] As a further improvement of the present invention, at least one temporal convolution process is performed on the three-dimensional image block to be recognized to obtain a first feature vector, including:

[0068] Perform a first temporal convolution process on the three-dimensional image block to be recognized to obtain a first temporal convolution vector;

[0069] When performing one temporal convolution process, determine the first temporal convolution vector as the first feature vector;

[0070] When performing n temporal convolution processes, perform pooling processing on the first temporal convolution vector to obtain a first intermediate feature vector, perform temporal convolution processing on the first intermediate feature vector to obtain a second temporal convolution vector, perform pooling processing on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, and determine the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

[0071] As a further improvement of the present invention, at least one fully connected process is performed on the third feature vector to obtain a fourth feature vector, including:

[0072] Perform a first fully connected process on the third feature vector to obtain a first intermediate vector;

[0073] When performing one fully connected process, determine the first intermediate vector as the fourth feature vector;

[0074] When performing m fully connected processes, perform a second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on, and determine the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

[0075] As a further improvement of the present invention, the system further includes: training the temporal convolution network according to a training set.

[0076] As a further improvement of the present invention, the temporal convolution network includes: at least one cascaded temporal convolution layer, a three-dimensional pooling layer, a dimensional transformation layer, and at least one fully connected layer.

[0077] As a further improvement of the present invention, feature extraction is performed on the three-dimensional image block to be recognized through the temporal convolution network, including:

[0078] Perform at least one temporal convolution process on the three-dimensional image block to be recognized through the at least one temporal convolution layer to obtain a first feature vector;

[0079] Perform integration processing on the first feature vector through the three-dimensional pooling layer to obtain a second feature vector;

[0080] The dimensionality reduction process is performed on the second feature vector through the dimensionality transformation layer to obtain a third feature vector;

[0081] The third feature vector is subjected to at least one fully connected process through the at least one fully connected layer to obtain a fourth feature vector;

[0082] The lesion recognition result is determined according to the fourth feature vector.

[0083] As a further improvement of the present invention, the at least one temporal convolutional layer performs at least one temporal convolutional process on the three-dimensional image block to be recognized to obtain a first feature vector, including:

[0084] The first temporal convolutional layer performs a first temporal convolutional process on the three-dimensional image block to be recognized to obtain a first temporal convolutional vector;

[0085] When performing one temporal convolutional process, the first temporal convolutional vector is determined as the first feature vector;

[0086] When performing n temporal convolutional processes, pooling is performed on the first temporal convolutional vector to obtain a first intermediate feature vector, the second temporal convolutional layer performs a temporal convolutional process on the first intermediate feature vector to obtain a second temporal convolutional vector, pooling is performed on the second temporal convolutional vector to obtain a second intermediate feature vector, and so on, and the nth temporal convolutional vector is determined as the first feature vector, where n is an integer greater than 1.

[0087] As a further improvement of the present invention, the at least one fully connected layer performs at least one fully connected process on the third feature vector to obtain a fourth feature vector, including:

[0088] The first fully connected layer performs a first fully connected process on the third feature vector to obtain a first intermediate vector;

[0089] When performing one fully connected process, the first intermediate vector is determined as the fourth feature vector;

[0090] When performing m fully connected processes, the second fully connected layer performs a second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on, and the mth intermediate vector is determined as the fourth feature vector, where m is an integer greater than 1.

[0091] As a further improvement of the present invention, the temporal convolutional layer is implemented by a two-dimensional convolutional spiking neuron model.

[0092] As a further improvement of the present invention, the temporal convolutional layer is implemented by a two-dimensional convolutional leaky integrate-and-fire model.

[0093] As a further improvement of the present invention, the two-dimensional convolutional leaky integrate-and-fire model is used for:

[0094] obtaining a value I through a convolutional operation on the input value X at time t according to the two-dimensional convolutional leaky integrate-and-fire model t and the weight W; t adding it to the biological voltage value at time t-1 to obtain the membrane potential value at time t

[0095] determining the output value F at time t according to the membrane potential value at time t and the firing threshold V th ; t

[0096] determining whether to reset the membrane potential according to the output value F at time t t and determining the reset membrane potential value according to the reset voltage value V reset ; wherein,

[0097] determining the biological voltage value at time t according to the reset membrane potential value ; wherein, α and β are leakage factors of the Leak activation function;

[0098] wherein, the output value F at time t t serves as the input to the next layer cascaded with the two-dimensional convolutional leaky integrate-and-fire model, and the biological voltage value at time t serves as the input for calculating the membrane potential value at time t+1.

[0099] As a further improvement of the present invention, determining the output value F at time t according to the membrane potential value at time t and the firing threshold V th includes: t

[0100] if the membrane potential value at time t is greater than or equal to the firing threshold V th then determining that the output value F at time t t is 1;

[0101] if the membrane potential value at time t is less than the firing threshold V th then determining that the output value F at time t t is 0.

[0102] As a further improvement of the present invention, training the temporal convolutional network according to a training set includes: obtaining a sample data set;

[0103] Among them, obtaining a sample data set includes:

[0104] Preprocessing the obtained medical images to obtain three-dimensional image blocks;

[0105] Constructing a sample data set based on the three-dimensional image blocks;

[0106] Performing data augmentation processing on the samples in the sample data set, and expanding the processed data to the sample data set. Among them, performing data augmentation processing on the samples in the sample data set includes: randomly flipping the samples in 3 directions, randomly adjusting the size of the samples, rotating the samples at an arbitrary angle in three-dimensional space, and randomly moving the samples in 3 directions, at least one of these processes.

[0107] As a further improvement of the present invention, preprocessing the medical images to obtain three-dimensional image blocks includes:

[0108] Performing feature extraction on the medical images to obtain images of lesion regions;

[0109] Performing format conversion on the images of the lesion regions to obtain three-dimensional image blocks in npy format.

[0110] As a further improvement of the present invention, in the sample data set, 70% of the three-dimensional image blocks contain at least one nodule, and 30% of the three-dimensional image blocks do not contain nodules.

[0111] As a further improvement of the present invention, training the temporal convolutional network according to the training set includes: training the temporal convolutional network by the stochastic gradient descent algorithm and the step learning rate adjustment method.

[0112] The present invention also provides an electronic device, including a memory and a processor. The memory is used to store one or more computer instructions. Among them, the one or more computer instructions are executed by the processor to implement the described method.

[0113] The present invention also provides a computer-readable storage medium, on which a computer program is stored. It is characterized in that the computer program is executed by the processor to implement the described method.

[0114] The beneficial effects of the present invention are: utilizing the sparsity of neuron firing, significantly reducing the number of network parameters and the operation complexity, and improving the real-time performance and accuracy of recognition. Description of the Drawings

[0115] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0116] Figure 1 It is a schematic flowchart of a three-dimensional medical image recognition method according to an exemplary embodiment of the present disclosure. Specific embodiments

[0117] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.

[0118] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present disclosure, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0119] In addition, in the description of the present disclosure, the terms used are only for the purpose of illustration and are not intended to limit the scope of the present disclosure. The terms "including" and / or "comprising" are used to specify the existence of the described elements, steps, operations, and / or components, but do not exclude the existence or addition of one or more other elements, steps, operations, and / or components. The terms "first", "second", etc. may be used to describe various elements, do not represent an order, and do not limit these elements. In addition, in the description of the present disclosure, unless otherwise specified, "a plurality of" means two or more. These terms are only used to distinguish one element from another. With the following accompanying drawings, these and / or other aspects become obvious, and it is easier for those of ordinary skill in the art to understand the description of the embodiments of the present disclosure. The accompanying drawings are only used to depict the embodiments of the present disclosure for the purpose of illustration. Those skilled in the art will easily recognize from the following description that alternative embodiments of the structure and method shown in the present disclosure can be adopted without departing from the principles described in the present disclosure.

[0120] A three-dimensional medical image recognition method according to an embodiment of the present disclosure, as Figure 1 shown, includes:

[0121] Preprocess the medical image to be recognized to obtain a three-dimensional image block to be recognized;

[0122] Extract features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result;

[0123] Among them, the temporal convolutional network includes at least one temporal convolutional layer.

[0124] The medical image described in this disclosure can be a three-dimensional medical image such as chest CT or brain CT. This disclosure does not limit the type of medical image. Among them, the temporal convolutional layer can be a convolutional layer with temporal information and can be used for convolutional processing with temporal information on the feature map. The temporal convolutional layer can extract the temporal information between multiple pictures.

[0125] In an alternative implementation, preprocessing the medical image to obtain a three-dimensional image block includes: extracting features from the medical image to obtain an image of the lesion area; performing format conversion on the image of the lesion area to obtain a three-dimensional image block in npy format. The medical image is generally in RAW format, MHD format, or DICOM format, and is converted into a three-dimensional image block in npy format of a digital matrix, which is convenient for training using a temporal convolutional network.

[0126] For example, this disclosure can select the LUNA16 dataset, which contains 888 cases and labels 1186 pulmonary nodules. When extracting features from chest CT, the image can be subjected to mask extraction, convex hull and dilation, and gray-scale normalization processing.

[0127] (1) First, convert all the original data into HU values;

[0128] (2) Mask extraction: On the 2D slice, use a Gaussian filter with a standard deviation of 1 and a threshold of -600 to obtain a mask of the lungs and the darker surrounding parts, and then perform connectivity analysis to remove connected components smaller than 30mm 2 and parts with an eccentricity greater than 0.99. Then calculate all 3D connected components in the binary 3D matrix and only retain the non-edge parts (used to remove the darker surrounding parts of the lungs) and parts with a volume between 0.68 and 7.5L;

[0129] (3) Convex hull and dilation processing: If a nodule is connected to the outer wall of the lung, it will not appear in the above-extracted mask. For this case, first divide the lung into two parts, namely the left lung and the right lung, then perform convex hull processing on the left and right lungs respectively and dilate them outward by 10 pixels. However, for some 2D slices, the bottom of the lung is similar to a crescent shape. If the area after convex hull processing is more than 1.5 times the original, the convex hull is abandoned to avoid introducing too much other tissue;

[0130] (4) Gray-scale normalization processing: Linearly transform the HU values ([-1200, 600]) to gray-scale values within 0 - 255, and set the gray-scale values of pixels outside the mask to 170. Also, if the gray-scale values of pixels in the dilated area are higher than 210, they are also set to 170.

[0131] The entire preprocessing process of the chest CT can be achieved, for example, through the following method: First, use the load_itk_image function to load the original data and the mask; obtain the boundary of the mask, that is, the edge of the non-zero part, and get a box. Then use the function resample to apply a new resolution to it, that is, resampling, to unify the resolution; then use the function lum_trans to clip the data to -1200 - 600, and set the data outside this range to -1200 or 600, and then normalize the data to 0 - 255; use the function process_mask to perform a dilation operation on the mask to remove small holes in the lung, then apply the new mask to the original data, and fill the data values outside the mask with 170 (i.e., the new value after normalization of the HU value of water); resample the original data and then intercept the data within the box; read the label, use the function worldToVoxelCoord to convert it to voxel coordinates, and then apply a new resolution to it; store the processed data and label in the.npy format.

[0132] In an alternative embodiment, a temporal convolutional network is used to extract features from the three-dimensional image block to be recognized, and the lesion recognition result is obtained, including:

[0133] Perform at least one temporal convolution process on the three-dimensional image block to be recognized to obtain a first feature vector;

[0134] Integrate the first feature vector to obtain a second feature vector;

[0135] Reduce the dimension of the second feature vector to obtain a third feature vector;

[0136] Perform at least one fully connected process on the third feature vector to obtain a fourth feature vector;

[0137] Determine the lesion recognition result according to the fourth feature vector.

[0138] The present disclosure uses temporal convolution processing. Through the membrane potential accumulation and integral firing mechanism, the computational complexity is reduced, and the images at the previous and subsequent time steps can be connected for convolution processing, realizing the extraction of features between and within frames with high accuracy, low power consumption, and low computational volume, and improving the recognition accuracy.

[0139] In an alternative embodiment, performing at least one temporal convolution processing on the three-dimensional image block to be recognized to obtain a first feature vector includes:

[0140] Performing a first temporal convolution processing on the three-dimensional image block to be recognized to obtain a first temporal convolution vector;

[0141] When performing one temporal convolution processing, determining the first temporal convolution vector as the first feature vector;

[0142] When performing n temporal convolution processings, performing pooling processing on the first temporal convolution vector to obtain a first intermediate feature vector, performing temporal convolution processing on the first intermediate feature vector to obtain a second temporal convolution vector, performing pooling processing on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, and determining the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

[0143] Among them, the temporal convolution processing may refer to performing convolution processing with temporal information on the three-dimensional image block. For example, the three-dimensional image block can be convolved through a convolutional layer with temporal information. In this way, multiple pictures at the previous and subsequent time steps can be associated, and the temporal information between the pictures can be processed.

[0144] For example, when performing 2 temporal convolution processings, performing a first temporal convolution processing on the three-dimensional image block to be recognized to obtain a first temporal convolution vector, performing pooling processing on the first temporal convolution vector to obtain a first intermediate feature vector, performing a second temporal convolution processing on the first intermediate feature vector to obtain a second temporal convolution vector, and determining the second temporal convolution vector as the first feature vector. The present disclosure does not limit the number of times of temporal convolution processing.

[0145] In an alternative embodiment, performing at least one fully connected processing on the third feature vector to obtain a fourth feature vector includes:

[0146] Performing a first fully connected processing on the third feature vector to obtain a first intermediate vector;

[0147] When performing one fully connected processing, determining the first intermediate vector as the fourth feature vector;

[0148] When performing the fully connected processing for m times, the second fully connected processing is performed on the first intermediate vector to obtain a second intermediate vector, and so on, and the m-th intermediate vector is determined as the fourth feature vector, where m is an integer greater than 1.

[0149] Among them, the fully connected processing may refer to performing a non-linear transformation on the features extracted in the foregoing steps, extracting the associations between these features, and finally mapping them to the output space.

[0150] For example, when performing 4 times of fully connected processing, the first fully connected processing is performed on the third feature vector to obtain a first intermediate vector, the second fully connected processing is performed on the first intermediate vector to obtain a second intermediate vector, the third fully connected processing is performed on the second intermediate vector to obtain a third intermediate vector, the third fully connected processing is performed on the third intermediate vector to obtain a fourth intermediate vector, and the fourth intermediate vector is determined as the fourth feature vector. The present disclosure does not limit the number of times of fully connected processing.

[0151] In an optional implementation manner, the method further includes: training the temporal convolutional network according to a training set.

[0152] In an optional implementation manner, the temporal convolutional network includes: at least one cascaded temporal convolutional layer, a three-dimensional pooling layer, a dimensionality transformation layer, and at least one fully connected layer.

[0153] In an optional implementation manner, feature extraction is performed on the three-dimensional image block to be recognized through a temporal convolutional network, including:

[0154] Performing at least one temporal convolutional processing on the three-dimensional image block to be recognized through at least one temporal convolutional layer to obtain a first feature vector;

[0155] Performing an integration processing on the first feature vector through a three-dimensional pooling layer to obtain a second feature vector;

[0156] Performing a dimensionality reduction processing on the second feature vector through a dimensionality transformation layer to obtain a third feature vector;

[0157] Performing at least one fully connected processing on the third feature vector through at least one fully connected layer to obtain a fourth feature vector;

[0158] Determining a lesion recognition result according to the fourth feature vector.

[0159] In an optional implementation manner, performing at least one temporal convolutional processing on the three-dimensional image block to be recognized through at least one temporal convolutional layer to obtain a first feature vector, including:

[0160] Performing a first temporal convolutional processing on the three-dimensional image block to be recognized through a first temporal convolutional layer to obtain a first temporal convolutional vector;

[0161] When performing a temporal convolution process once, the first temporal convolution vector is determined as the first feature vector;

[0162] When performing n temporal convolution processes, the first temporal convolution vector is pooled to obtain a first intermediate feature vector. The first intermediate feature vector is subjected to a temporal convolution process through a second temporal convolution layer to obtain a second temporal convolution vector. The second temporal convolution vector is pooled to obtain a second intermediate feature vector, and so on. The nth temporal convolution vector is determined as the first feature vector, where n is an integer greater than 1.

[0163] Among them, the temporal convolution process may refer to performing a convolution process with temporal information on a three-dimensional image block. For example, a three-dimensional image block can be subjected to a convolution process through a convolution layer with temporal information. In this way, multiple pictures at adjacent time steps can be related, and the temporal information between the pictures can be processed.

[0164] For example, the temporal convolution network includes 2 temporal convolution layers. When performing 2 temporal convolution processes, the first temporal convolution layer performs a first temporal convolution process on the three-dimensional image block to be recognized to obtain a first temporal convolution vector. The first temporal convolution vector is pooled to obtain a first intermediate feature vector. The second temporal convolution layer performs a second temporal convolution process on the first intermediate feature vector to obtain a second temporal convolution vector. The second temporal convolution vector is determined as the first feature vector. The present disclosure does not limit the number of temporal convolution layers.

[0165] In an alternative embodiment, at least one fully connected layer performs at least one fully connected process on the third feature vector, and obtaining a fourth feature vector includes:

[0166] The first fully connected layer performs a first fully connected process on the third feature vector to obtain a first intermediate vector;

[0167] When performing a fully connected process once, the first intermediate vector is determined as the fourth feature vector;

[0168] When performing m fully connected processes, the second fully connected layer performs a second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on. The mth intermediate vector is determined as the fourth feature vector, where m is an integer greater than 1.

[0169] Among them, the fully connected process may refer to performing a non-linear transformation on the features extracted in the foregoing steps, extracting the associations between these features, and finally mapping them to the output space.

[0170] For example, it includes 4 fully connected layers. When performing 4 times of fully connected processing, the first fully connected layer performs the first fully connected processing on the third feature vector to obtain the first intermediate vector. The second fully connected layer performs the second fully connected processing on the first intermediate vector to obtain the second intermediate vector. The third fully connected layer performs the third fully connected processing on the second intermediate vector to obtain the third intermediate vector. The fourth fully connected layer performs the third fully connected processing on the third intermediate vector to obtain the fourth intermediate vector, and the fourth intermediate vector is determined as the fourth feature vector. The present disclosure does not limit the number of fully connected layers.

[0171] In an alternative embodiment, the temporal convolutional layer can be implemented using a two-dimensional convolutional spiking neuron model.

[0172] The method of SNN (Spiking Neural Network) generally constructs a multi-layer spiking neural network based on biological rules. The neuron model of SNN is different from the more abstract neuron model in the deep neural network. Generally, it uses the biological neuron model commonly used in computational neurology to generate spikes or event firings through the threshold effect. In the medical scenario, the individual samples vary greatly, the annotation is difficult, and the scale of the sample set is small. Therefore, the SNN with few parameters and strong learning ability has advantages in medical image processing. Structurally speaking, the neuron structure in the biological brain is much more complex than that in the ANN (Artificial Neural Network). Comparing the signal models of brain neurons, the neurons in the ANN only need to sum the signals and then directly output all of them through a simple activation function, while the brain neurons directly affect the membrane potential after receiving the signals, and when the membrane potential is large enough, a spike signal is emitted. Functionally speaking, humans can learn the features of a new category through very few samples, while the ANN requires a large number of labeled samples for learning.

[0173] The neuron model of SNN can be, for example, the LIF (Leaky Integrity-Fire) model, the IF (Integrity-Fire) model, etc. The present disclosure does not specifically limit the neuron model of SNN.

[0174] In an alternative embodiment, the temporal convolutional layer is implemented using a two-dimensional convolutional leaky integrate-and-fire model (ConvLIF2D). ConvLIF2D adopts the LIF neuron model.

[0175] Among them, the temporal convolutional layer can be a ConvLIF2D layer, which further extracts features from the feature map through convolution, and realizes the extraction of multi-frame temporal information through the accumulation of the biological membrane potential of the LIF neurons within the layer, so as to realize the convolutional processing with time series on the feature map. The temporal convolutional layer can extract the temporal information between multiple pictures.

[0176] In an alternative embodiment, the two-dimensional convolutional leaky integrate-and-fire model layer is configured to:

[0177] Obtain a value I through a convolution operation on the input value X at time t of the two-dimensional convolutional leaky integrate-and-fire model t and the weight W; t Add it to the biological voltage value at time t - 1 to obtain the membrane potential value at time t

[0178] Determine the output value F at time t based on the membrane potential value at time t and the firing threshold V th ; t ;

[0179] Determine whether to reset the membrane potential based on the output value F at time t t and determine the reset membrane potential value according to the reset voltage value V reset ; wherein,

[0180] Determine the biological voltage value at time t according to the reset membrane potential value ; wherein, α and β are leakage factors of the Leak activation function;

[0181] wherein, the output value F at time t t serves as the input to the next layer cascaded with the two-dimensional convolutional leaky integrate-and-fire model, and the biological voltage value at time t serves as the input for calculating the membrane potential value at time t + 1.

[0182] In an alternative embodiment, the determining the output value F at time t based on the membrane potential value at time t and the firing threshold V th includes: t If the membrane potential value at time t

[0183] is greater than or equal to the firing threshold V th , then determine that the output value F at time t t is 1;

[0184] If the membrane potential value at time t is less than the firing threshold V th , then determine that the output value F at time t t is 0.

[0185] ​This disclosure uses a ConvLIF2D model for temporal convolution processing and spike pulses for information encoding. The spike signal is a binary signal. A value of 1 represents that the neuron is activated (generates an event) and is transmitted to the neurons connected to its backend, that is, a pulse event is sent to the backend neurons; a value of 0 represents that the neuron is not activated and has no impact on the backend neurons, that is, the above-mentioned pulse event is not generated. Through this event-driven method (i.e., it only works when an event is received), compared with the fully convolutional network, higher energy efficiency can be obtained. Utilizing the sparsity of neuron firing significantly reduces the number of network parameters and the computational complexity. In addition, temporal convolution processing can connect the content of the previous and subsequent time steps (corresponding to the z-axis of the medical image). When performing convolution processing on a three-dimensional image block, the accuracy of the network can be improved.

[0186] In an alternative embodiment, training a temporal convolutional network based on a training set includes: obtaining a sample data set;

[0187] Among them, obtaining a sample data set includes:

[0188] Preprocessing the obtained medical images to obtain three-dimensional image blocks;

[0189] Constructing a sample data set based on the three-dimensional image blocks;

[0190] Performing data augmentation processing on the samples in the sample data set and expanding the processed data to the sample data set. Among them, performing data augmentation processing on the samples in the sample data set includes: randomly flipping the samples in 3 directions, randomly adjusting the size of the samples, rotating the samples at an arbitrary angle in three-dimensional space, and randomly moving the samples in 3 directions with at least one of the above processes.

[0191] For example, the samples can be randomly flipped in the x, y, and z directions, the samples can be randomly resized between 0.75 - 1.25 mm, the samples can be rotated at an angle of 0.5 degrees in three-dimensional space, and the samples can be randomly moved in the x, y, and z directions with the moving distance less than 15% of the nodule radius.

[0192] In an alternative embodiment, the positive and negative samples of the training set in the sample data set are unfolded according to time steps and visualized. The positive and negative samples are divided according to the level of disease severity. For example, in the LUNA16 data set, those with levels 0 - 2 are determined as negative samples, and those with levels 3 - 4 are determined as positive samples. After visualizing the positive and negative samples, it is convenient to display the samples.

[0193] In an alternative embodiment, the preprocessing process of the above sample data set is as described above and will not be elaborated here.

[0194] In an alternative embodiment, in the sample dataset, 70% of the three-dimensional image patches contain at least one nodule, and 30% of the three-dimensional image patches do not contain nodules. By adding data without nodules, data augmentation is performed on the medical imaging scenario with fewer samples to reduce overfitting.

[0195] In an alternative embodiment, training the temporal convolutional network according to the training set includes: training the temporal convolutional network by the Stochastic Gradient Descent (SGD) algorithm and the step-wise learning rate schedule method. The stochastic gradient descent method randomly selects a group from the samples, updates the gradient once after training, then selects another group and updates again. In the case of a large sample size, a trained network model with a loss value within an acceptable range can be obtained without training all the samples. For the step-wise learning rate adjustment method, when the loss value no longer decreases, the learning rate is reduced by 1 / 10 until the network converges, and then a trained network model can be obtained.

[0196] A three-dimensional medical image recognition system according to an embodiment of the present disclosure, the system includes:

[0197] A preprocessing module, which is used to preprocess the medical image to be recognized to obtain a three-dimensional image patch to be recognized;

[0198] A feature extraction module, which is used to extract features from the three-dimensional image patch to be recognized through a temporal convolutional network to obtain a lesion recognition result;

[0199] Wherein, the temporal convolutional network includes at least one temporal convolutional layer.

[0200] The medical image described in the present disclosure can be a three-dimensional medical image such as a chest CT or a brain CT. The present disclosure does not limit the type of medical image. Among them, the temporal convolutional layer can be a convolutional layer with temporal information, which can be used to perform temporal convolution processing on the feature map. The temporal convolutional layer can extract the temporal information between multiple pictures.

[0201] In an alternative embodiment, preprocessing the medical image to obtain a three-dimensional image patch includes: extracting features from the medical image to obtain an image of the lesion area; performing format conversion on the image of the lesion area to obtain a three-dimensional image patch in npy format. The medical image is generally in RAW format, MHD format, or DICOM format, and is converted into a three-dimensional image patch in the digital matrix npy format, which is convenient for training using the temporal convolutional network.

[0202] For example, the present disclosure may select the LUNA16 dataset, which contains 888 cases with 1,186 pulmonary nodules marked. When extracting features from chest CT, the image can be subjected to mask extraction, convex hull and dilation, and gray-scale normalization processing.

[0203] (1) First, convert all the original data into HU values;

[0204] (2) Mask extraction: On the 2D slice, use Gaussian filtering with a standard deviation of 1 and a threshold of -600 to obtain the mask of the lungs and the darker surrounding parts, and then perform connectivity analysis to remove the connected components smaller than 30 mm 2 and the parts with an eccentricity greater than 0.99. Then calculate all the 3D connected components in the binary 3D matrix and only retain the non-edge parts (for removing the darker parts around the lungs) and the parts with a volume between 0.68 and 7.5 L;

[0205] (3) Convex hull and dilation processing: If the nodule is connected to the outer wall of the lung, it will not appear in the mask extracted above. For this situation, first divide the lung into two parts, namely the left lung and the right lung, and then perform convex hull processing on the left and right lungs respectively and dilate them by 10 pixels. However, for some 2D slices, the bottom of the lung is similar to a crescent shape. After performing convex hull processing on this type, if the area is more than 1.5 times the original, the convex hull is abandoned to avoid introducing too much other tissue;

[0206] (4) Gray-scale normalization processing: Linearly transform the HU values ([-1,200, 600]) to gray-scale values within 0 to 255, and set the gray-scale values of the pixels outside the mask to 170, and also set the gray-scale values of the pixels in the dilated area higher than 210 to 170.

[0207] The entire preprocessing process of chest CT can be achieved, for example, through the following method: First, use the load_itk_image function to load the original data and the mask; obtain the boundary of the mask, that is, the edge of the non-zero part, and obtain a box. Then, use the resample function to apply a new resolution to it, that is, resampling, to unify the resolution; then use the lum_trans function to clip the data to -1200 to 600, and set the data outside this range to -1200 or 600, and then normalize the data to 0 to 255; use the process_mask function to perform a dilation operation on the mask to remove small holes in the lungs, and then apply the new mask to the original data and fill the data values outside the mask with 170 (that is, the new value after normalization of the HU value of water); resample the original data and then intercept the data within the box; read the label, use the worldToVoxelCoord function to convert it to voxel coordinates, and then apply a new resolution to it; store the processed data and label in the.npy format.

[0208] In an alternative embodiment, the three-dimensional image block to be recognized is subjected to feature extraction through a temporal convolutional network to obtain a lesion recognition result, including:

[0209] Perform at least one temporal convolutional process on the three-dimensional image block to be recognized to obtain a first feature vector;

[0210] Integrate the first feature vector to obtain a second feature vector;

[0211] Reduce the dimension of the second feature vector to obtain a third feature vector;

[0212] Perform at least one fully connected process on the third feature vector to obtain a fourth feature vector;

[0213] Determine the lesion recognition result according to the fourth feature vector.

[0214] The present disclosure uses temporal convolutional processing. Through the membrane potential accumulation and integral firing mechanism, the computational complexity is reduced, and the images at the previous and subsequent time steps can be connected for convolutional processing to achieve high-accuracy, low-power, and low-computation extraction of features between and within frames, improving the recognition accuracy.

[0215] In an alternative embodiment, performing at least one temporal convolutional process on the three-dimensional image block to be recognized to obtain a first feature vector includes:

[0216] Perform the first temporal convolutional process on the three-dimensional image block to be recognized to obtain a first temporal convolutional vector;

[0217] When performing one temporal convolutional process, determine the first temporal convolutional vector as the first feature vector;

[0218] When performing the time series convolution process for n times, perform pooling on the first time series convolution vector to obtain a first intermediate feature vector, perform time series convolution on the first intermediate feature vector to obtain a second time series convolution vector, perform pooling on the second time series convolution vector to obtain a second intermediate feature vector, and so on. Determine the nth time series convolution vector as the first feature vector, where n is an integer greater than 1.

[0219] Among them, the time series convolution process may refer to performing convolution on a three-dimensional image block with time series information. For example, convolution can be performed on a three-dimensional image block through a convolution layer with time series information. In this way, multiple pictures at consecutive time steps can be related, and the time series information between the pictures can be processed.

[0220] For example, when performing the time series convolution process for 2 times, perform the first time series convolution on the three-dimensional image block to be recognized to obtain a first time series convolution vector, perform pooling on the first time series convolution vector to obtain a first intermediate feature vector, perform the second time series convolution on the first intermediate feature vector to obtain a second time series convolution vector, and determine the second time series convolution vector as the first feature vector. The present disclosure does not limit the number of times of the time series convolution process.

[0221] In an alternative embodiment, perform at least one fully connected process on the third feature vector to obtain a fourth feature vector, including:

[0222] Perform the first fully connected process on the third feature vector to obtain a first intermediate vector;

[0223] When performing one fully connected process, determine the first intermediate vector as the fourth feature vector;

[0224] When performing m fully connected processes, perform the second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on. Determine the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

[0225] Among them, the fully connected process may refer to performing a non-linear transformation on the features extracted in the foregoing steps, extracting the associations between these features, and finally mapping them to the output space.

[0226] For example, when performing 4 fully connected processes, perform the first fully connected process on the third feature vector to obtain a first intermediate vector, perform the second fully connected process on the first intermediate vector to obtain a second intermediate vector, perform the third fully connected process on the second intermediate vector to obtain a third intermediate vector, perform the third fully connected process on the third intermediate vector to obtain a fourth intermediate vector, and determine the fourth intermediate vector as the fourth feature vector. The present disclosure does not limit the number of times of the fully connected process.

[0227] In an alternative embodiment, the system further includes: training a temporal convolutional network according to a training set.

[0228] In an alternative embodiment, the temporal convolutional network includes: at least one cascaded temporal convolutional layer, a three-dimensional pooling layer, a dimensionality transformation layer, and at least one fully connected layer.

[0229] In an alternative embodiment, feature extraction is performed on the three-dimensional image block to be recognized through the temporal convolutional network, including:

[0230] Performing at least one temporal convolution process on the three-dimensional image block to be recognized through at least one temporal convolutional layer to obtain a first feature vector;

[0231] Integrating the first feature vector through the three-dimensional pooling layer to obtain a second feature vector;

[0232] Reducing the dimensionality of the second feature vector through the dimensionality transformation layer to obtain a third feature vector;

[0233] Performing at least one fully connected process on the third feature vector through at least one fully connected layer to obtain a fourth feature vector;

[0234] Determining a lesion recognition result according to the fourth feature vector.

[0235] In an alternative embodiment, performing at least one temporal convolution process on the three-dimensional image block to be recognized through at least one temporal convolutional layer to obtain a first feature vector includes:

[0236] Performing a first temporal convolution process on the three-dimensional image block to be recognized through the first temporal convolutional layer to obtain a first temporal convolution vector;

[0237] When performing one temporal convolution process, determining the first temporal convolution vector as the first feature vector;

[0238] When performing n temporal convolution processes, performing a pooling process on the first temporal convolution vector to obtain a first intermediate feature vector, performing a temporal convolution process on the first intermediate feature vector through the second temporal convolutional layer to obtain a second temporal convolution vector, performing a pooling process on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, determining the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

[0239] Among them, the temporal convolution process may refer to performing a convolution process with temporal information on the three-dimensional image block. For example, the three-dimensional image block can be convolved through a convolutional layer with temporal information. In this way, multiple pictures at previous and subsequent time steps can be associated, and the temporal information between pictures can be processed.

[0240] For example, the temporal convolutional network includes two temporal convolutional layers. When performing two temporal convolutional processes, the first temporal convolutional layer performs the first temporal convolutional process on the three-dimensional image block to be recognized to obtain the first temporal convolutional vector, performs pooling processing on the first temporal convolutional vector to obtain the first intermediate feature vector, and the second temporal convolutional layer performs the second temporal convolutional process on the first intermediate feature vector to obtain the second temporal convolutional vector, and determines the second temporal convolutional vector as the first feature vector. The present disclosure does not limit the number of temporal convolutional layers.

[0241] In an alternative embodiment, at least one fully connected layer performs at least one fully connected process on the third feature vector, and obtaining a fourth feature vector includes:

[0242] The first fully connected layer performs the first fully connected process on the third feature vector to obtain the first intermediate vector;

[0243] When performing one fully connected process, the first intermediate vector is determined as the fourth feature vector;

[0244] When performing m fully connected processes, the second fully connected layer performs the second fully connected process on the first intermediate vector to obtain the second intermediate vector, and so on, and the m-th intermediate vector is determined as the fourth feature vector, where m is an integer greater than 1.

[0245] Among them, the fully connected process may refer to performing a non-linear transformation on the features extracted in the foregoing steps, extracting the associations between these features, and finally mapping them to the output space.

[0246] For example, it includes four fully connected layers. When performing four fully connected processes, the first fully connected layer performs the first fully connected process on the third feature vector to obtain the first intermediate vector, the second fully connected layer performs the second fully connected process on the first intermediate vector to obtain the second intermediate vector, the third fully connected layer performs the third fully connected process on the second intermediate vector to obtain the third intermediate vector, the fourth fully connected layer performs the third fully connected process on the third intermediate vector to obtain the fourth intermediate vector, and the fourth intermediate vector is determined as the fourth feature vector. The present disclosure does not limit the number of fully connected layers.

[0247] In an alternative embodiment, the temporal convolutional layer can be implemented using a two-dimensional convolutional spiking neuron model.

[0248] The method of SNN (Spiking Neural Network) generally constructs a multi-layer spiking neural network based on biological rules. The neuron model of SNN is different from the more abstract neuron model in the deep neural network. Generally, the biological neuron model commonly used in computational neurology is adopted, and spikes or event firings are generated through the threshold effect. In the medical scenario, the individual samples vary greatly, the annotation is difficult, and the scale of the sample set is small. Therefore, the SNN with fewer parameters and strong learning ability has advantages in medical image processing. Structurally, the neuron structure in the biological brain is far more complex than that in the ANN (Artificial Neural Network). Comparing the signal models of brain neurons, the neurons in the ANN only need to sum the signals and then directly output all of them through a simple activation function, while the brain neurons directly affect the membrane potential after receiving signals, and when the membrane potential is large enough, a spike signal is emitted. Functionally, humans can learn the features of a new category through very few samples, while the ANN requires a large number of labeled samples for learning.

[0249] The neuron model of SNN can be, for example, the LIF model, the IF model, etc. The present disclosure does not make specific limitations on the neuron model of SNN.

[0250] In an alternative embodiment, the temporal convolutional layer is implemented by a two-dimensional convolutional leaky integrate-and-fire model (ConvLIF2D). ConvLIF2D adopts the LIF neuron model.

[0251] Among them, the temporal convolutional layer can be a ConvLIF2D layer, which further extracts features from the feature map through convolution, and realizes the extraction of multi-frame temporal information through the accumulation of the biological membrane potential of the LIF neurons within the layer, so as to realize the convolutional processing with time series on the feature map. The temporal convolutional layer can extract the temporal information between multiple pictures.

[0252] In an alternative embodiment, the two-dimensional convolutional leaky integrate-and-fire model is used for:

[0253] According to the input value X at time t of the two-dimensional convolutional leaky integrate-and-fire model t After convolution operation with the weight W, the value I is obtained t , added to the biological voltage value at time t - 1 , to obtain the membrane potential value at time t

[0254] According to the membrane potential value at time t And the firing threshold V th , determine the output value F at time t t ;

[0255] According to the output value F at time t t Determine whether to reset the membrane potential, and according to the reset voltage value Vreset Determine the reset membrane potential value wherein

[0256] According to the reset membrane potential value Determine the biological voltage value at time t wherein α and β are the leakage factors of the Leak activation function;

[0257] wherein, the output value F at time t t As the input of the next layer cascaded with the two-dimensional convolutional leaky integrate-and-fire model, the biological voltage value at time t As the input for calculating the membrane potential value at time t + 1.

[0258] In an alternative embodiment, the membrane potential value according to time t and the firing threshold V th , determine the output value F at time t t , including:

[0259] If the membrane potential value at time t is greater than or equal to the firing threshold V th , then determine the output value F at time t t is 1;

[0260] If the membrane potential value at time t is less than the firing threshold V th , then determine the output value F at time t t is 0.

[0261] The present disclosure uses the ConvLIF2D model for temporal convolutional processing and uses spike pulses for information encoding. The spike signal is a binary signal. 1 represents that the neuron is activated (generates an event) and is transmitted to the neurons connected to its backend, that is, a pulse event is sent to the backend neurons; 0 represents that the neuron is not activated and has no impact on the backend neurons, that is, the above pulse event is not generated. Through this event-driven method (i.e., it works only when an event is received), compared with the fully convolutional network, higher energy efficiency can be obtained. Utilizing the sparsity of neuron firing significantly reduces the network parameter quantity and operation complexity. In addition, temporal convolutional processing can connect the content of the previous and subsequent time steps (corresponding to the z-axis of the medical image). When performing convolutional processing on a three-dimensional image block, the accuracy of the network can be improved.

[0262] In an alternative embodiment, training the temporal convolutional network according to the training set includes: obtaining a sample data set;

[0263] wherein, obtaining the sample data set includes:

[0264] Preprocess the acquired medical images to obtain three-dimensional image blocks;

[0265] Construct a sample data set based on the three-dimensional image blocks;

[0266] Perform data augmentation on the samples in the sample data set, and expand the processed data to the sample data set. Among them, performing data augmentation on the samples in the sample data set includes: randomly flipping the samples in 3 directions, randomly adjusting the size of the samples, rotating the samples at any angle in three-dimensional space, and randomly moving the samples in 3 directions, with at least one of these operations.

[0267] For example, the samples can be randomly flipped in the x, y, and z directions, the samples can be randomly resized between 0.75 - 1.25 mm, the samples can be rotated at an angle of 0.5 degrees in three-dimensional space, and the samples can be randomly moved in the x, y, and z directions, with the moving distance less than 15% of the nodule radius.

[0268] In an alternative embodiment, the positive and negative samples of the training set in the sample data set are unfolded according to time steps and visualized. The positive and negative samples are divided according to the level of disease severity. For example, in the LUNA16 data set, those with levels 0 - 2 are determined as negative samples, and those with levels 3 - 4 are determined as positive samples. After visualizing the positive and negative samples, it is convenient to display the samples.

[0269] In an alternative embodiment, the preprocessing process of the above sample data set is as described above and will not be elaborated here.

[0270] In an alternative embodiment, in the sample data set, 70% of the three-dimensional image blocks contain at least one nodule, and 30% of the three-dimensional image blocks do not contain nodules. By adding data without nodules, data augmentation is performed on the medical image scenario with fewer samples, reducing overfitting.

[0271] In an alternative embodiment, training the temporal convolutional network according to the training set includes: training the temporal convolutional network by using the Stochastic Gradient Descent (SGD) algorithm and the step-wise learning rate schedule method. The Stochastic Gradient Descent method randomly selects a group from the samples, updates the gradient once after training, then selects another group and updates again. In the case of a large sample size, a trained network model with a loss value within an acceptable range can be obtained without training all the samples. For the step-wise learning rate adjustment method, when the loss value no longer decreases, the learning rate is reduced by 1 / 10 until the network converges, and then a trained network model can be obtained.

[0272] The present disclosure also relates to an electronic device, including a server, a terminal, etc. The electronic device includes: at least one processor; a memory communicatively connected to the at least one processor; and a communication component communicatively connected to a storage medium, where the communication component receives and sends data under the control of the processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement the three-dimensional medical image recognition method in the above embodiments.

[0273] In an alternative embodiment, the memory, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory, that is, to implement the three-dimensional medical image recognition method.

[0274] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store an option list, etc. In addition, the memory may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to an external device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0275] One or more modules are stored in the memory and, when executed by one or more processors, implement the three-dimensional medical image recognition method in any of the above method embodiments.

[0276] The above-mentioned product can execute the method provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For the technical details not described in detail in this embodiment, reference can be made to the three-dimensional medical image recognition method provided by the embodiments of the present application.

[0277] The present disclosure also relates to a computer-readable storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute some or all of the embodiments of the above-mentioned three-dimensional medical image recognition method.

[0278] That is, those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0279] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present disclosure can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0280] In addition, those of ordinary skill in the art can understand that although some of the embodiments described herein include certain features included in other embodiments but not other features, the combination of the features of different embodiments means that it is within the scope of the present disclosure and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0281] Those skilled in the art should understand that although the present disclosure has been described with reference to exemplary embodiments, various changes can be made and equivalents can be substituted for its elements without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the present disclosure without departing from the essential scope of the present disclosure. Therefore, the present disclosure is not limited to the specific embodiments disclosed, but the present disclosure will include all embodiments falling within the scope of the appended claims.

Claims

1. A three-dimensional medical image recognition method, characterized in that, Including: Preprocessing the medical image to be recognized to obtain a three-dimensional image block to be recognized; Performing feature extraction on the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result; Wherein, the temporal convolutional network includes at least one temporal convolutional layer, the temporal convolutional layer is implemented by a two-dimensional convolutional leaky integrate-and-fire model, and the two-dimensional convolutional leaky integrate-and-fire model encodes information using binary signal pulses; the two-dimensional convolutional leaky integrate-and-fire model is used for: According to the input value X at time t of the two-dimensional convolutional leaky integrate-and-fire model t and the weight W, the value I is obtained through a convolution operation t , and added to the biological voltage value at time t-1 to obtain the membrane potential value at time t According to the membrane potential value at time t and the emission threshold V th , determine the output value F at time t t ; Based on the output value F at time t t Determine whether to reset the membrane potential, and based on the reset voltage value V reset Determine the reset membrane potential value wherein According to the reset membrane potential value Determine the biological voltage value at time t Wherein α and β are the leakage factors of the Leak activation function; Among them, the output value F at time t t As the input of the next layer cascaded with the two-dimensional convolutional leaky integrate-and-fire model, the biological voltage value at time t As the input for calculating the membrane potential value at time t + 1.

2. The method according to claim 1, characterized in that Performing feature extraction on the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result, including: Performing at least one temporal convolutional process on the three-dimensional image block to be recognized to obtain a first feature vector; Performing an integration process on the first feature vector to obtain a second feature vector; Performing a dimensionality reduction process on the second feature vector to obtain a third feature vector; Performing at least one fully connected process on the third feature vector to obtain a fourth feature vector; Determining the lesion recognition result according to the fourth feature vector.

3. The method according to claim 2, wherein Performing at least one temporal convolutional process on the three-dimensional image block to be recognized to obtain a first feature vector, including: Performing a first temporal convolutional process on the three-dimensional image block to be recognized to obtain a first temporal convolutional vector; When performing one temporal convolutional process, determining the first temporal convolutional vector as the first feature vector; When performing n temporal convolutional processes, performing a pooling process on the first temporal convolutional vector to obtain a first intermediate feature vector, performing a temporal convolutional process on the first intermediate feature vector to obtain a second temporal convolutional vector, performing a pooling process on the second temporal convolutional vector to obtain a second intermediate feature vector, and so on, determining the nth temporal convolutional vector as the first feature vector, where n is an integer greater than 1.

4. The method according to claim 2, characterized in that Performing at least one fully connected process on the third feature vector to obtain a fourth feature vector, including: Performing a first fully connected process on the third feature vector to obtain a first intermediate vector; When performing one fully connected process, determining the first intermediate vector as the fourth feature vector; When performing m fully connected processes, performing a second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on, determining the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: training the temporal convolutional network according to a training set.

6. The method according to claim 5, wherein The temporal convolutional network includes: at least one cascaded temporal convolutional layer, a three-dimensional pooling layer, a dimensionality transformation layer, and at least one fully connected layer.

7. The method according to claim 6, wherein Performing feature extraction on the three-dimensional image block to be recognized through the temporal convolutional network, including: Performing at least one temporal convolutional process on the three-dimensional image block to be recognized through the at least one temporal convolutional layer to obtain a first feature vector; Performing an integration process on the first feature vector through the three-dimensional pooling layer to obtain a second feature vector; Performing a dimensionality reduction process on the second feature vector through the dimensionality transformation layer to obtain a third feature vector; Performing at least one fully connected process on the third feature vector through the at least one fully connected layer to obtain a fourth feature vector; Determine the lesion recognition result according to the fourth eigenvector.

8. The method according to claim 7, wherein Perform at least one temporal convolution process on the three-dimensional image block to be recognized through the at least one temporal convolution layer, and obtain a first eigenvector, including: Perform a first temporal convolution process on the three-dimensional image block to be recognized through a first temporal convolution layer, and obtain a first temporal convolution vector; When performing one temporal convolution process, determine the first temporal convolution vector as the first eigenvector; When performing n temporal convolution processes, perform a pooling process on the first temporal convolution vector to obtain a first intermediate eigenvector, perform a temporal convolution process on the first intermediate eigenvector through a second temporal convolution layer to obtain a second temporal convolution vector, perform a pooling process on the second temporal convolution vector to obtain a second intermediate eigenvector, and so on, and determine the nth temporal convolution vector as the first eigenvector, where n is an integer greater than 1.

9. The method according to claim 7, characterized in that, Perform at least one fully connected process on the third eigenvector through the at least one fully connected layer to obtain a fourth eigenvector, including: Perform a first fully connected process on the third eigenvector through a first fully connected layer to obtain a first intermediate vector; When performing one fully connected process, determine the first intermediate vector as the fourth eigenvector; When performing m fully connected processes, perform a second fully connected process on the first intermediate vector through a second fully connected layer to obtain a second intermediate vector, and so on, and determine the mth intermediate vector as the fourth eigenvector, where m is an integer greater than 1.

10. The method according to claim 1 or 6, characterized in that, The temporal convolution layer is implemented by using a two-dimensional convolutional spiking neuron model.

11. The method according to claim 1, wherein Said according to the membrane potential value at time t and the emission threshold V th , determining the output value F at time t t , including: If the membrane potential value at time t is greater than or equal to the emission threshold V th , then determine the output value F at time t t to be 1; If the membrane potential value at time t is less than the emission threshold V th , then determine the output value F at time t t to be 0.

12. The method according to claim 5, characterized in that Train the temporal convolution network according to the training set, including: obtaining a sample data set; Among them, obtaining a sample data set includes: Preprocess the obtained medical image to obtain a three-dimensional image block; Construct a sample data set based on the three-dimensional image block; Perform data augmentation processing on the samples in the sample data set, and expand the processed data to the sample data set. Among them, performing data augmentation processing on the samples in the sample data set includes: randomly flipping the sample in 3 directions, randomly adjusting the size of the sample, rotating the sample at an arbitrary angle in three-dimensional space, and randomly moving the sample in 3 directions, at least one of the above processes.

13. The method according to claim 1 or 12, characterized in that, Preprocess the medical image to obtain a three-dimensional image block, including: Extract features from the medical image to obtain an image of the lesion area; Convert the format of the image of the lesion area to obtain a three-dimensional image block in npy format.

14. The method according to claim 12, wherein In the sample data set, 70% of the three-dimensional image blocks contain at least one nodule, and 30% of the three-dimensional image blocks do not contain nodules.

15. The method according to claim 5, wherein Train the temporal convolution network according to the training set, including: training the temporal convolution network through a random gradient descent algorithm and a step learning rate adjustment method.

16. A three-dimensional medical image recognition system, characterized in that, The system includes: A preprocessing module, which is used to preprocess the medical image to be recognized to obtain a three-dimensional image block to be recognized; A feature extraction module, which is used to extract features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result. Among them, the temporal convolutional network includes at least one temporal convolutional layer, and the temporal convolutional layer is implemented by a two-dimensional convolutional leaky integrate-and-fire model, and the two-dimensional convolutional leaky integrate-and-fire model uses binary signal pulses for information encoding; the two-dimensional convolutional leaky integrate-and-fire model is used for: According to the input value X at time t of the two-dimensional convolutional leaky integrate-and-fire model t and the weight W, the value I is obtained after convolution operation t , and added to the biological voltage value at time t-1 to obtain the membrane potential value at time t According to the membrane potential value at time t and the firing threshold V th , determine the output value F at time t t ; According to the output value F at time t t Determine whether to reset the membrane potential, and according to the reset voltage value V reset Determine the reset membrane potential value Wherein According to the reset membrane potential value Determine the biological voltage value at time t where α and β are the leakage factors of the Leak activation function; Among them, the output value F at time t t As the input of the next layer cascaded with the two-dimensional convolutional leaky integrate-and-fire model, the biological voltage value at time t As the input for calculating the membrane potential value at time t+1.

17. The system according to claim 16, wherein Extracting features from the three-dimensional image block to be recognized through a temporal convolutional network to obtain a lesion recognition result, including: Performing at least one temporal convolution process on the three-dimensional image block to be recognized to obtain a first feature vector; Integrating the first feature vector to obtain a second feature vector; Reducing the dimension of the second feature vector to obtain a third feature vector; Performing at least one fully connected process on the third feature vector to obtain a fourth feature vector; Determining the lesion recognition result according to the fourth feature vector.

18. The system according to claim 17, wherein Performing at least one temporal convolution process on the three-dimensional image block to be recognized to obtain a first feature vector, including: Performing a first temporal convolution process on the three-dimensional image block to be recognized to obtain a first temporal convolution vector; When performing one temporal convolution process, determining the first temporal convolution vector as the first feature vector; When performing n temporal convolution processes, performing pooling on the first temporal convolution vector to obtain a first intermediate feature vector, performing temporal convolution on the first intermediate feature vector to obtain a second temporal convolution vector, performing pooling on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, determining the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

19. The system according to claim 17, wherein Performing at least one fully connected process on the third feature vector to obtain a fourth feature vector, including: Performing a first fully connected process on the third feature vector to obtain a first intermediate vector; When performing one fully connected process, determining the first intermediate vector as the fourth feature vector; When performing m fully connected processes, performing a second fully connected process on the first intermediate vector to obtain a second intermediate vector, and so on, determining the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

20. The system according to any one of claims 16-19, characterized in that, The system further includes: training the temporal convolutional network according to a training set.

21. The system according to claim 20, wherein The temporal convolutional network includes: at least one cascaded temporal convolutional layer, a three-dimensional pooling layer, a dimension transformation layer, and at least one fully connected layer.

22. The system according to claim 21, wherein, Extracting features from the three-dimensional image block to be recognized through the temporal convolutional network, including: Performing at least one temporal convolution process on the three-dimensional image block to be recognized through the at least one temporal convolutional layer to obtain a first feature vector; Integrating the first feature vector through the three-dimensional pooling layer to obtain a second feature vector; Reducing the dimension of the second feature vector through the dimension transformation layer to obtain a third feature vector; Performing at least one fully connected process on the third feature vector through the at least one fully connected layer to obtain a fourth feature vector; Determining the lesion recognition result according to the fourth feature vector.

23. The system according to claim 22, wherein Performing at least one temporal convolution process on the three-dimensional image block to be recognized through the at least one temporal convolution layer to obtain a first feature vector, including: Performing a first temporal convolution process on the three-dimensional image block to be recognized through a first temporal convolution layer to obtain a first temporal convolution vector; When performing one temporal convolution process, determining the first temporal convolution vector as the first feature vector; When performing n temporal convolution processes, performing a pooling process on the first temporal convolution vector to obtain a first intermediate feature vector, performing a temporal convolution process on the first intermediate feature vector through a second temporal convolution layer to obtain a second temporal convolution vector, performing a pooling process on the second temporal convolution vector to obtain a second intermediate feature vector, and so on, determining the nth temporal convolution vector as the first feature vector, where n is an integer greater than 1.

24. The system according to claim 22, wherein Performing at least one fully connected process on the third feature vector through the at least one fully connected layer to obtain a fourth feature vector, including: Performing a first fully connected process on the third feature vector through a first fully connected layer to obtain a first intermediate vector; When performing one fully connected process, determining the first intermediate vector as the fourth feature vector; When performing m fully connected processes, performing a second fully connected process on the first intermediate vector through a second fully connected layer to obtain a second intermediate vector, and so on, determining the mth intermediate vector as the fourth feature vector, where m is an integer greater than 1.

25. The system according to claim 16 or 21, characterized in that, The temporal convolution layer is implemented using a two-dimensional convolutional spiking neuron model.

26. The system according to claim 16, wherein Said according to the membrane potential value at time t and the emission threshold V th , determine the output value F at time t t , including: If the membrane potential value at time t is greater than or equal to the emission threshold V th , then determine the output value F at time t t to be 1; If the membrane potential value at time t is less than the emission threshold V th , then determine that the output value F at time t t is 0.

27. The system according to claim 20, wherein, Training the temporal convolution network according to a training set, including: obtaining a sample data set; Among them, obtaining a sample data set includes: Preprocessing the obtained medical image to obtain a three-dimensional image block; Constructing a sample data set based on the three-dimensional image block; Performing data augmentation processing on the samples in the sample data set and expanding the processed data to the sample data set. Among them, performing data augmentation processing on the samples in the sample data set includes: performing at least one of randomly flipping the sample in 3 directions, randomly adjusting the size of the sample, rotating the sample at an arbitrary angle in three-dimensional space, and randomly moving the sample in 3 directions.

28. The system according to claim 16 or 27, characterized in that Preprocessing the medical image to obtain a three-dimensional image block, including: Performing feature extraction on the medical image to obtain an image of the lesion area; Performing format conversion on the image of the lesion area to obtain a three-dimensional image block in npy format.

29. The system according to claim 27, wherein, In the sample data set, 70% of the three-dimensional image blocks contain at least one nodule, and 30% of the three-dimensional image blocks do not contain nodules.

30. The system according to claim 20, wherein Training the temporal convolution network according to a training set, including: training the temporal convolution network through a stochastic gradient descent algorithm and a step learning rate adjustment method.

31. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer instructions, where the one or more computer instructions are executed by a processor to implement the method according to any one of claims 1-15.

32. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor to implement the method according to any one of claims 1-15.

Citation Information

Patent Citations

  • CT image pulmonary nodule detection method based on 3D residual neural network

    CN107590797A

  • Method and device for training and predicting prediction model of transient stability of power grid system

    CN110838075A