An STRes3dRNN model, an evaluation system, and an evaluation method

Through the STRes3dRNN model, the 3D residual network and hybrid attention mechanism are used, combined with the BiLSTM network, the dependence and subjectivity problems of existing lung function evaluation methods are solved, and the radiation-free and low-cost automated lung function evaluation is achieved, and the evaluation accuracy and efficiency are improved.

CN119027722BActive Publication Date: 2025-07-04NANJING FORESTRY UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411002226.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2025-07-04
Estimated Expiration
2044-07-25

AI Technical Summary

Technical Problem

The existing lung function assessment methods rely on active cooperation of patients, which have problems such as radiation exposure and expensive equipment, and the interpretation of EIT images is highly subjective, making it difficult to accurately judge the lung function status.

Method used

The STRes3dRNN model is adopted to capture the spatiotemporal features and dependencies in the EIT image sequence by stacking 3D residual network and hybrid attention mechanism, combined with the BiLSTM network, and realize automated evaluation.

Benefits of technology

It realizes automated lung function assessment without active cooperation, improves the accuracy and efficiency of the assessment, reduces radiation exposure, reduces equipment costs, and enhances the ability to identify lung function status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119027722B_ABST
    Figure CN119027722B_ABST
Patent Text Reader

Abstract

The present invention proposes an STRes3dRNN model, an evaluation system, and an evaluation method, belonging to the field of medical imaging. The technical key points are as follows: S100, obtaining continuous voltage data to generate an EIT image sequence; S200, using a three-dimensional residual network to extract feature information and output a spatio-temporal coupled feature sequence map; S300, optimizing the weight distribution through a hybrid attention mechanism composed of spatio-temporal, channel, and motion attention modules to improve performance; S400, using a bidirectional long short-term memory network to capture the dependencies in the EIT image sequence and enhance the ability to capture spatio-temporal feature correlations; S500, combining a fully connected layer and a Softmax function to predict spatio-temporal classification and identify and classify diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical imaging and deep learning, and particularly relates to an STRes3dRNN model, an evaluation system, and an evaluation method. Background Art

[0002] Respiratory diseases are the third leading cause of death globally. Early screening for lung diseases is crucial to delay disease progression in a timely manner and improve the prognosis of patients. Electrical Impedance Tomography (EIT) is a non-invasive medical imaging technique that can visualize the conductivity distribution caused by lung ventilation in real time and provide continuous dynamic information on lung ventilation. Although studies have used EIT technology to confirm the heterogeneity of lung ventilation function in patients with Chronic Obstructive Pulmonary Disease (COPD), no one has yet used EIT image sequences for the pulmonary function evaluation of clinical patients to distinguish normal, obstructive, or damaged states.

[0003] To overcome the above problems, researchers have evaluated pulmonary function by extracting and processing impedance changes at different time series. After obtaining the time series EIT images of the subject to be tested, an area impedance distribution map containing time characteristics is analyzed, and distribution characteristic parameters are obtained based on this. By comparing with the standard distribution characteristic parameters, it is determined whether the regional ventilation function is damaged (e.g., CN118121183A, CN118121184A).

[0004] Based on the research of the above prior art, it is found that the following problems still exist in the application of EIT in related technologies for etiological diagnosis:

[0005] (1) Existing pulmonary function evaluation methods include pulmonary function tests and medical imaging examinations. The above methods all rely on the active cooperation of patients and cumbersome breathing movements, have limitations, and are difficult to implement in actual clinical applications.

[0006] (2) Chest X-rays and CT scans are usually accompanied by problems such as radiation exposure and expensive equipment, resulting in shortcomings in regular examinations and treatment monitoring, and limiting their application scenarios and usage frequencies.

[0007] (3) Manual interpretation of images obtained by EIT technology is subjective, and misassessment is prone to occur in early case screening. Simple time series image processing cannot accurately judge the pulmonary function status. Summary of the Invention

[0008] The object of the present invention is to solve the problems existing in the above-mentioned prior art, and provide a pulmonary function evaluation method and system for electrical impedance tomography of an STRes3dRNN (3D Spatiotemporal Residual Recurrent Neural Network) model, an evaluation system, and an evaluation method.

[0009] The technical solution of this application lies in:

[0010] An STRes3dRNN model, including, connected in sequence: a network input layer, an initial stacking layer, an intermediate stacking layer, an end stacking layer, and a network output layer;

[0011] Among them, the network input layer is used to read the EIT image sequence;

[0012] Among them, the initial stacking layer is sequentially stacked by a 3D residual initial unit, an intermediate 3D residual standard unit, and an end 3D residual standard unit; the 3D residual initial unit is used to initially extract high-dimensional image features, the intermediate 3D residual standard unit is used to reduce the dimension of the high-dimensional image features, and the end 3D residual standard unit is used to reduce the dimension and deepen the feature information;

[0013] Among them, the intermediate stacking layer includes: n sequentially connected single intermediate stacking layers; the single intermediate stacking layer is stacked by a HAM unit, a 3D residual upsampling unit, and an end 3D residual standard unit; n is a natural number greater than or equal to 1; the HAM unit optimizes the weight of the allocated information, eliminates redundant information, and strengthens the feature information; the 3D residual upsampling unit upsamples the low-dimensional image features; the end 3D residual standard unit is used to reduce the dimension and deepen the feature information;

[0014] Among them, the end stacking layer is sequentially stacked by a 3D global average pooling layer, a BiLSTM layer, and a fully connected layer; the 3D global average pooling layer reads the image sequence information obtained by the intermediate stacking layer and extracts its feature vector; the BiLSTM layer processes and outputs the feature vector; the fully connected layer maps the input feature vector to the target classification quantity and obtains the mapped score value, and the target classification is: healthy, obstructive, damaged.

[0015] Furthermore, the 3D residual initial unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, and a 3D max pooling layer connected in sequence; the three-dimensional convolutional layer is used to convolve and extract the image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data of the previous step through a non-linear activation function to improve the representation ability of the features; the 3D max pooling layer compresses the features using different channels to achieve invariance, and extracts and outputs the maximum value.

[0016] Furthermore, both the middle 3D residual standard unit and the end 3D residual standard unit include a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence. The 3D convolutional layer is used to convolve and extract image sequence features. The batch normalization layer is used to normalize the sequence features to accelerate the convergence speed. The activation layer processes the data given by the batch normalization layer through a non-linear activation function to improve the feature representation ability. The additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation.

[0017] Furthermore, the HAM unit includes a spatio-temporal attention module, a channel attention module, and a motion attention module. The above three modules are applied to enhance the feature information in the spatio-temporal, channel, and motion aspects respectively, and the enhanced information is weighted to obtain a fused and enhanced data of spatio-temporal, channel, and motion information with the same number of layers and dimensions remaining unchanged.

[0018] Furthermore, the 3D residual upsampling unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence. The 3D convolutional layer is used to convolve and extract image sequence features. The batch normalization layer normalizes the sequence features to accelerate the convergence speed. The activation layer processes the data of the previous step through a non-linear activation function to improve the feature representation ability. The additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation.

[0019] Furthermore, the HAM unit includes a spatio-temporal attention module, a channel attention module, and a motion attention module. The above three modules are applied to enhance the feature information in the spatio-temporal, channel, and motion aspects respectively, and the enhanced information is weighted to obtain a fused and enhanced data of spatio-temporal, channel, and motion information with the same number of layers and dimensions remaining unchanged;

[0020] Spatio-temporal attention module (Spatial-Temporal Excitation, STE): After performing average pooling on the channel C dimension, this module uses a 3D convolutional kernel of size 3×3×3 to model the information in the time and space dimensions to capture the spatio-temporal features in the EIT image sequence;

[0021] Channel attention module (Channel Excitation, CE): This module performs average pooling on the height H and width W dimensions to reduce the number of channels to 1 / 16 of the original, and then uses a 1D convolutional kernel of size 3 for operations between two fully connected layers (FC) to represent the channel features in the time dimension;

[0022] Motion Excitation (ME) module: This module performs tensor separation on the number of frames T, processes it using a 2D convolutional kernel of size 3×3, and simultaneously adds differential processing of adjacent frames to simulate the motion information of EIT lung ventilation changes.

[0023] Furthermore, the BiLSTM layer consists of two channels: forward propagation and backward propagation. Forward and backward processing are performed simultaneously. The processing time step is determined by the input image sequence, and it automatically cycles through the same number of time steps as the sequence number; the forward and backward processing results are concatenated to obtain a high-dimensional feature vector.

[0024] An evaluation system, which includes:

[0025] A storage unit for storing continuous voltage data measured by an EIT device;

[0026] An EIT image sequence generation unit that reads the continuous voltage data from the storage unit to generate an EIT image sequence;

[0027] An STRes3dRNN model that reads the EIT image sequence from the EIT image sequence generation unit, and its result output is passed to the probability calculation unit;

[0028] A probability calculation unit for giving the probability percentages of health, blockage, and injury.

[0029] Furthermore, the working method of the probability calculation unit is:

[0030]

[0031] Softmax(z i ) represents the predicted probability of the i-th category, represents the raw score of the i-th category.

[0032] An evaluation method that does not directly aim at disease diagnosis, which includes the following steps:

[0033] S100, obtaining continuous voltage data to generate an EIT image sequence; the input of the EIT network is an EIT image sequence of 16 frames of 64×64×3, where 64×64 is the image size and 3 is a three-channel RGB image;

[0034] S200, transmitting the image sequence I to the initial stacking layer; the initial stacking layer is stacked by 3D residual initial units, intermediate 3D residual standard units, and end 3D residual standard units;

[0035] S200 includes the following sub-steps:

[0036] S201, The image sequence I is input into the 3D residual initial unit to obtain the high-dimensional features of the 7×7×7 image sequence with 64 layers;

[0037] S202, The high-dimensional features obtained in step S201 are input into the intermediate 3D residual standard unit to obtain the image features after dimensionality reduction;

[0038] S203, The image features obtained in step S202 are input into the final 3D residual standard unit to obtain the image sequence after dimensionality reduction and deepening of feature information;

[0039] S300, The result obtained by the initial stacking layer calculation is input to the intermediate stacking layer; The intermediate stacking layer includes: 3 sequentially connected single intermediate stacking layers; The single intermediate stacking layer is stacked by a HAM unit, a 3D residual upsampling unit, and a 3D residual standard unit:

[0040] S300 includes the following sub-steps:

[0041] S301, The HAM unit of the first single intermediate stacking layer reads the 1×1×1 image sequence with 128 layers, optimizes the weight of the allocated information, eliminates redundant information, strengthens the feature information, and outputs the 1×1×1 image sequence with 128 layers;

[0042] S302, The 3D residual upsampling unit of the first single intermediate stacking layer reads the 1×1×1 image sequence with 128 layers and outputs the 3×3×3 image sequence with 128 layers;

[0043] S303, The final 3D residual standard unit of the first single intermediate stacking layer reads the 3×3×3 image sequence with 128 layers and outputs the 1×1×1 image sequence with 256 layers;

[0044] S304, The HAM unit of the second single intermediate stacking layer reads the image sequence of the 1×1×1 image sequence with 256 layers, optimizes the weight of the allocated information, eliminates redundant information, strengthens the feature information, and outputs the image sequence of the 1×1×1 image sequence with 256 layers;

[0045] S305, The 3D residual upsampling unit of the second single intermediate stacking layer reads the 1×1×1 image sequence with 256 layers and outputs the 3×3×3 image sequence with 256 layers;

[0046] S306, The final 3D residual standard unit of the second single intermediate stacking layer reads the image sequence of the 3×3×3 image sequence with 256 layers and outputs the image sequence of the 1×1×1 image sequence with 512 layers;

[0047] S307, the HAM unit in the middle stacking layer of the third monomer reads the image sequence of a 512-layer 1×1×1, and by optimizing the weight of the allocated information, eliminating redundant information, and strengthening feature information, outputs the image sequence of a 512-layer 1×1×1;

[0048] S308, the 3D residual upsampling unit in the middle stacking layer of the third monomer reads the image sequence of a 512-layer 1×1×1 and outputs the image sequence of a 512-layer 3×3×3;

[0049] S309, the final 3D residual standard unit in the middle stacking layer of the third monomer reads the image sequence of the image sequence of a 512-layer 3×3×3 and outputs the image sequence of a 512-layer 1×1×1;

[0050] S400, the result calculated by the middle stacking layer is input to the final stacking layer, and the final stacking layer includes a three-dimensional average pooling layer, a BiLSTM layer, and a fully connected layer connected in sequence; Step S400 further includes the following sub-steps:

[0051] S401, the three-dimensional average pooling layer reduces the dimension of the image sequence of a 512-layer 1×1×1 and extracts its feature vector;

[0052] S402, the BiLSTM layer processes and outputs a 128-layer feature vector;

[0053] S403, the fully connected layer maps the input feature vector to the target classification and obtains the mapped score value; among them, the number of target classifications is 3, which respectively represent healthy, blocked, and damaged.

[0054] Further, it also includes: S500, the result of the final stacking layer is input to the Softmax function, and the predicted probability is output, that is, the probability percentage of healthy, blocked, and damaged:

[0055]

[0056] where, Softmax(z i ) represents the predicted probability of the i-th category, represents the original score of the i-th category.

[0057] The advantages of the technical solution of the present invention are mainly reflected in:

[0058] (1) This application has developed an STRes3dRNN model, which can effectively extract spatio-temporal features from the EIT image sequence and realize the evaluation of the functional state of patient organs (such as the brain, lungs).

[0059] (2) For a three-dimensional residual network, network degradation problems are prone to occur as the number of network layers increases during its application. Therefore, this application introduces three-dimensional residual learning to optimize the training process and alleviate the degradation problem. Residual connections allow the input to skip the 3D convolutional units of the main branch, simplify the gradient transmission path, provide a simpler network flow path, and thus improve network performance and training efficiency.

[0060] The output function of the residual unit in this application is:

[0061]

[0062] Among them, represents the input of the 3D residual, f Res3D represents the residual function, and δ represents the learnable parameter. Residual connections allow the input to skip the 3D convolutional units of the main branch, simplify the gradient transmission path, provide a simpler network flow path, and thus improve network performance and training efficiency.

[0063] (3) The hybrid attention mechanism can reallocate the weights between spatio-temporal features, enhance the attention to key features, reduce the adverse effects of redundant features on spatio-temporal classification, and provide more accurate model support for EIT pulmonary function assessment.

[0064] The output function of the hybrid attention mechanism determined in this application is:

[0065]

[0066] Among them, represents the input of the 3D residual, is the hybrid activation function, and θ represents all learnable parameters.

[0067] (4) The bidirectional long short-term memory network can capture the long-term dependencies in the EIT image sequence and can also capture information from two different time directions, greatly enhancing the ability to capture the correlations in spatio-temporal features.

[0068] The function of the bidirectional long short-term memory network determined in this application is:

[0069]

[0070] Among them, ⊙ represents the Hadamard product, i, f, g, and o respectively represent the input gate, forget gate, memory cell, and output gate, and σ c is the hyperbolic tangent function (tanh) used for state activation. These gating mechanisms enable BiLSTM to effectively manage the storage and forgetting of information, which is crucial for capturing long-term dependencies.

[0071] (5) Spatiotemporal classification prediction includes a fully connected layer (FC) and a Softmax function, which can integrate the spatiotemporal feature information extracted in the previous steps, summarize and form a global representation, and effectively extract and learn the spatiotemporal feature information in EIT lung ventilation changes by combining a three-dimensional residual network, a hybrid attention mechanism, and a bidirectional long short-term memory network. Description of the Drawings

[0072] The following further describes the present application in detail with reference to the embodiments in the drawings, but does not constitute any limitation to the present application.

[0073] Figure 1 It is the architecture design diagram of the STRes3dRNN model in the present invention.

[0074] Figure 2 It is the schematic diagram of lung function evaluation in the present invention.

[0075] Figure 3 It is the EIT index diagram in the present invention.

[0076] Figure 4 It is the actual test diagram in the present invention.

[0077] Figure 5 It is the training progress of the STRes3dRNN model in the present invention.

[0078] Figure 6 It is the classification result of the STRes3dRNN confusion matrix in the present invention.

[0079] Figure 7 It is the parameter optimization experiment of the STRes3dRNN in the present invention.

[0080] Figure 8 It is the comparison result of the EIT spatiotemporal classification performance indexes based on different network models in the present invention. Detailed Embodiments

[0081] The objectives, advantages, and features of the present invention will be explained through the non-limiting description of the following preferred embodiments. These embodiments are only typical examples of applying the technical solutions of the present invention, and any technical solutions formed by equivalent replacement or equivalent transformation fall within the scope of protection required by the present invention.

[0082] <1: A method for evaluating lung function by electrical impedance tomography based on a three-dimensional spatiotemporal residual recurrent neural network>

[0083] <1.1. Technical problems to be solved>

[0084] Although previous studies have used EIT technology to confirm the heterogeneity of lung ventilation function in COPD patients, no one has yet used EIT image sequences for clinical lung function assessment to distinguish between normal, obstructive or damaged states.

[0085] <1.2、Technical ideas>

[0086] The technical idea of ​​the pulmonary function assessment method based on electrical impedance tomography of three-dimensional spatiotemporal residual recurrent neural network proposed in this application is:

[0087] (1) Res3d is constructed by stacking multiple 3D convolution kernels and 3D residual units to capture various spatiotemporal correlations in the EIT image sequence and output a feature sequence graph rich in spatiotemporal coupling information;

[0088] (2) Using HAM to redistribute the weights of spatiotemporal features, reduce the adverse effects of redundant features on spatiotemporal classification, and improve accuracy;

[0089] (3) BiLSTM is used to capture long-term dependencies in EIT image sequences, enhancing sensitivity to temporal dynamics and pulmonary ventilation change trends;

[0090] (4) Combining the three-dimensional residual network, hybrid attention mechanism and bidirectional long short-term memory network, the spatiotemporal feature information of EIT pulmonary ventilation changes can be effectively extracted and learned to accurately identify and classify lung diseases.

[0091] <1.3. Overall technical solution>

[0092] The evaluation principle of the method of this application is as follows Figure 2 As shown, the EIT technology applies current excitation on the body surface and measures voltage changes, combines the finite element physical field model to reconstruct the lung impedance distribution image, and uses the functional evaluation method of electrical impedance tomography based on three-dimensional spatiotemporal residual recurrent neural network to perform functional evaluation.

[0093] <1.4、STRes3dRNN architecture design>

[0094] Table 1 STRes3dRNN architecture table of this application

[0095]

[0096]

[0097] It should be noted that the "layer" in "64 layers, 128 layers, 256 layers" in Table 1 refers to the depth of the convolutional neural network, that is, the number of network layers.

[0098] It should be noted that: The 3D residual initial unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, and a 3D max pooling layer connected in sequence; the 3D convolutional layer is used to convolve and extract image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data of the previous step through a non-linear activation function to improve the feature representation ability; the 3D max pooling layer compresses the features using different channels to achieve invariance, and extracts and outputs the maximum value.

[0099] The 3D residual standard unit (middle, end) includes (which belongs to the prior art, see: https: / / mp.weixin.qq.com / s / Isv9JZp9qjwRgnNqJyjfxg) a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence; the 3D convolutional layer is used to convolve and extract image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data given by the batch normalization layer through a non-linear activation function to improve the feature representation ability; the additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation.

[0100] It should be noted that: The HAM unit includes a spatio-temporal attention module, a channel attention module, and a motion attention module; the above three modules are applied to enhance the feature information in the spatio-temporal, channel, and motion aspects respectively, and the enhanced information is weighted to obtain a fusion-enhanced data with the same number of layers and dimensions for spatio-temporal, channel, and motion information.

[0101] It should be noted that: The 3D residual upsampling unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence; the 3D convolutional layer is used to convolve and extract image sequence features; the batch normalization layer normalizes the sequence features to accelerate the convergence speed; the activation layer processes the data of the previous step through a non-linear activation function to improve the feature representation ability; the additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation.

[0102] It should be noted that: The BiLSTM layer is composed of two channels of forward propagation and backward propagation, and the forward and backward processes are carried out simultaneously. The processing time step is determined by the input image sequence, and it automatically loops for the same number of time steps as the sequence number; the forward and backward processing results are concatenated to obtain a high-dimensional feature vector.

[0103] <1.5. Algorithm>

[0104] A working method of an STRes3dRNN model includes the following steps:

[0105] S100. Obtain continuous voltage data to generate an EIT image sequence I. The input of the EIT network is an EIT image sequence of 16 frames with a size of 64×64×3. 64×64 is the image size, and 3 represents a three-channel RGB image;

[0106] S200. Transmit the image sequence I to the initial stacking layer (stack1). The initial stacking layer is sequentially stacked by 3D residual initial units, 3D residual standard units, and 3D residual standard units: S200 includes the following sub-steps:

[0107] S201. Input the image sequence I into the 3D residual initial unit to obtain high-dimensional features of a 64-layer 7×7×7 image sequence;

[0108] S202. Input the high-dimensional features obtained in step S201 into the intermediate 3D residual standard unit to obtain the image features after dimensionality reduction;

[0109] S203. Input the image features obtained in step S202 into the final 3D residual standard unit to obtain an image sequence with dimensionality reduction and enhanced feature information;

[0110] S300. Input the result of the initial stacking layer (stack1) into the intermediate stacking layers (stack2 - stack4);

[0111] The intermediate stacking layers include: 3 sequentially connected single intermediate stacking layers; The single intermediate stacking layer is stacked by HAM units, 3D residual upsampling units, and 3D residual standard units:

[0112] The working steps of the single intermediate stacking layer are:

[0113] S301. The HAM unit reads the image sequence, eliminates redundant information, and outputs the image sequence after enhancing the feature information;

[0114] S302. The 3D residual upsampling unit upsamples the low-dimensional features and outputs the upsampled image sequence;

[0115] S303. The 3D residual standard unit downsamples the high-dimensional image sequence to obtain an image sequence with enhanced feature information;

[0116] S400. Input the result of the intermediate stacking layer into the final stacking layer (stack5), the final stacking layer; The final stacking layer includes sequentially connected: 3D global average pooling layer, BiLSTM layer, and fully connected layer;

[0117] S400 includes the following sub-steps:

[0118] S401. The 3D global average pooling layer reads the 512-layer 1×1×1 image sequence for dimensionality reduction and extracts its feature vector;

[0119] In S402, the BiLSTM layer processes and outputs a feature vector of 128 layers;

[0120] In S403, the fully connected layer maps the input feature vector to the number of target classifications (the number is 3, healthy, blocked, damaged), and obtains the mapped score value;

[0121] In S500, the result of the final stacked layer is input into the Softmax function, and the predicted probability is output, that is, the probability percentage of healthy, blocked, and damaged, and the sum of the three is equal to 1:

[0122]

[0123] where Softmax(z i ) represents the predicted probability of the i-th category, represents the original score (i.e., the score value) of the i-th category.

[0124] When the network is trained, the loss is calculated through the cross-entropy loss function. For each training, the Softmax function will output the prediction result, and the difference between the prediction result and the actual result is calculated through the loss function formula. The training stops when the maximum number of training times or the minimum error reaches the set threshold. The loss function formula:

[0125]

[0126] where, y i is the one-hot encoded value of the i-th category in the true label, is the classification probability of the i-th category in the model prediction.

[0127] <2: Clinical trial design>

[0128] <2.1. Subject description>

[0129] From June 2022 to September 2023, a total of 192 subjects were recruited, and their written informed consents were obtained. Before the test, 6 subjects with contraindications to PFT or EIT measurement were excluded. During the test, 9 subjects could not effectively complete the pulmonary function examination, and another 12 subjects had poor contact with the EIT measurement system. After these exclusions, the respiratory data of 165 subjects were used for subsequent clinical analysis.

[0130] The baseline physical characteristics, PFT values, and EIT indicators of different subjects are listed in Table I and Figure 3Among them. Statistically, there was no significant difference among the three groups (P>0.05): 60 subjects with healthy lungs (aged 54.68±8.37 years, height 170.98±8.36 cm, weight 67.50±9.08 kg) were classified into the healthy group; 55 subjects with chronic obstructive pulmonary disease (aged 69.57±9.09 years, height 163.26±8.57 cm, weight 63.07±9.27 kg) were classified into the obstructive group; subjects with diseases such as pneumonia, pneumothorax, and pleural effusion (aged 65.21±10.37 years, height 168.34±9.26 cm, weight 64.29±11.05 kg) were classified into the injury group.

[0131] Subjects were required to sit for EIT measurement. After the measurement was stable, they took calm breaths for about 1 to 3 minutes, and the EIT data was recorded synchronously. This process was repeated three times to collect multiple sets of data for analysis. If the subject felt uncomfortable, the measurement was terminated immediately. Both EIT and PFT measurements were carried out after approval by the Medical Ethics Committee of Jinan University (JNUKY-2022-005) and clinical trial registration.

[0132] Table 2 Baseline physical characteristics of different subjects

[0133]

[0134]

[0135] <2.2. EIT data acquisition>

[0136] The EIT signal was collected in real time by a thoracic electrical impedance tomography device (EIT-1000, Nanjing Jilun Medical Intelligence Technology Co., Ltd., Jiangsu, China). As Figure 4 shown, the system consists of a wearable electrode sensor, data acquisition hardware, and image reconstruction software. To accurately obtain lung ventilation activity information, the system designed a fourth-order low-pass digital filter to effectively remove cardiac signals and measurement noise.

[0137] For each respiratory cycle, taking the end-expiratory V0 as the reference point, the relevant voltage signal ΔV i =(V i -V0) / V0 (i = 1, 2, …, T) was used as the input of the GN algorithm for image reconstruction. To enhance lung visualization, the EIT image sequence D={Δσ i} (i = 1, 2,..., T) was normalized to the range [0, 1]. Assuming D min and D max represent the minimum and maximum pixel values of the EIT image respectively, the normalized image sequence is expressed as:

[0138]

[0139] Due to the differences in the breathing times of the subjects, the inconsistency of the EIT image sequences in the time dimension poses a challenge to the spatio-temporal classification task based on STRes3dRNN. To address this issue, 16 consecutive EIT images are extracted at equal intervals as the network input. For cases where the breathing time is insufficient, a zero-padding strategy is adopted to ensure the consistency of the time dimension of the EIT image sequences of all patients.

[0140] <2.3. Model Training>

[0141] Clinically, it is based on the EIT data of 165 subjects during calm breathing. The consecutive EIT images extracted in each breathing cycle form a sample. The dataset contains a total of N = 7654 EIT image sequences, where the size of the input images is X = 64×64×16×3 (height × width × depth × channels). The pulmonary function assessment involves three diagnostic categories: Y = healthy, obstructive, and damaged, which increases the complexity of model prediction.

[0142] During the training process, we adopted the Adam optimizer, set the mini-batch size to B = 128, the number of iterations to E = 50, and the initial learning rate to α = 0.001. After every 10 training epochs, the learning rate was gradually reduced by a decay factor of β = 0.9, and other parameters were set to default. In addition, to enhance the generalization ability of the model, we performed data augmentation on the dataset after each training epoch, including rotation, translation, scaling, flipping, and cropping, to generate new training samples.

[0143] The computer hardware environment used in the experiment includes: an Intel Core i9-14900K CPU@3.2GHz processor, an NVIDIA RTX A4000 graphics card, and 4800MHz DDR5@64GB RAM memory. The network model was implemented using the Python language and the TensorFlow framework, and image reconstruction was completed through MATLAB. The ten-fold cross-validation method was adopted, randomly extracting 10% of the data as the test set and the remaining part as the training set.

[0144] <2.4. Evaluation Criteria>

[0145] To evaluate the classification performance of the STRes3dRNN model, the following metrics are used in this application: Accuracy (A), Precision (P), Recall (R), and F1 Score (F1). Accuracy refers to the proportion of the number of samples predicted correctly by the model to the total number of samples, which is an intuitive indicator to measure classification accuracy. Precision reflects the proportion of samples actually being positive among those predicted as positive by the model. Recall represents the proportion of positive samples classified by the model to all actual positive samples. The F1 Score is the harmonic mean of Precision and Recall, comprehensively reflecting the performance of these two metrics.

[0146] The calculation formulas are as follows:

[0147]

[0148] Among them, TP (True Positive), TN (True Negative), FP (False Positive), and FN (False Negative) represent the quantities of different types of prediction results respectively.

[0149] <2.5, Parameter Optimization Experiment>

[0150] (a) Initial learning rate α and image sequence depth T: To explore the influence of different initial learning rates α and image sequence depths T on the classification performance of STRes3dRNN, we selected combinations of initial learning rates α of 0.1, 0.01, 0.001, and 0.0001, and image sequence depths T of 8, 16, 32, and 64 for combined testing.

[0151] (b) Training sample size N: The training sample size N is an important parameter in the EIT health diagnosis task. More training samples can provide richer prior information, thus improving the classification accuracy to a certain extent. To verify the robustness of the STRes3dRNN model, we compared the classification accuracy and loss values under different training sample ratios. We randomly selected 60%, 65%, 70%, 75%, and 80% of the samples from the EIT image sequence as the training set, and the remaining 20% for validation.

[0152] <2.6, Ablation Comparison Experiment>

[0153] To test the classification performance of the method proposed in this application in pulmonary function assessment, we designed an ablation comparison experiment. This experiment verified the effectiveness of each module in the network by removing a certain module one by one and comparing the classification effects. We mainly compared the following three network configurations: the basic three-dimensional residual network Res3d, STRes3d combined with the HAM module, and STRes3dRNN (our method) that integrates the BiLSTM module. In addition, we also tested three mainstream network models, including C3D, R(2+1)D, and Variational Auto-Encoders (VAE), for performance comparison.

[0154] <3: Results and Analysis>

[0155] <3.1 Training Progress and Classification Results of the STRes3dRNN Model>

[0156] The training progress of the STRes3dRNN model is as Figure 5 shown. The entire training process involves nearly 5,000 iterations. Each iteration includes the estimation of network gradients and the update of weight parameters. The alternating background colors in the figure are used to mark each round of training, where one round of training represents a complete traversal of the entire dataset.

[0157] In the initial stage (0 - 1,000 iterations), the model has a shallow understanding of EIT data, and the classification accuracy and loss value fluctuate greatly, but overall show a trend of rapid increase and decrease, indicating that the model is learning; in the fluctuating stage (1,000 - 3,000 iterations), the model's understanding of EIT data deepens, the parameter adjustment is more stable, the classification accuracy gradually increases, and the decline of the loss value slows down, indicating that the model performance is gradually maturing; in the stable stage (after 3,000 iterations), the classification accuracy of the model reaches 97.40%, and the loss value is lower than 0.14, showing that the model has become relatively mature; generally speaking, the training progress process demonstrates the gradual improvement of the performance of STRes3dRNN and finally tends to be stable. In the left figure, the training accuracy (blue) is smoothed to show the performance improvement trend, and the validation accuracy (black) reflects the performance of the model on unseen data. In the right figure, the training loss (yellow) and the validation loss (black) respectively reflect the loss of each mini-batch and the entire validation set.

[0158] In the trained STRes3dRNN model, a total of 765 EIT cases were used for the health diagnosis task, and the classification results of the confusion matrix are shown in Figure 6The figure clearly reveals the relationship between the prediction results and the true labels. The overall accuracy is 97.40%, shown in the cell in the lower right corner of the figure, and the error rate is only 2.6%. The right column and bottom row of the figure show the precision and recall of each category. In the healthy class, 98.4% of the 251 predicted samples were correctly classified and 98.8% of the 250 actual samples were correctly predicted. In the obstruction class, 97.4% of the 268 predicted samples were correctly classified and 96.3% of the 271 actual samples were correctly predicted. In the injury class, 96.3% of the 246 predicted samples were correctly classified and 97.1% of the 271 actual samples were correctly predicted. These classification results show that the STRes3dRNN model has high accuracy and reliability in health diagnosis tasks, and can effectively distinguish individuals under different lung ventilation conditions, thereby providing more favorable technical support for clinical diagnosis and treatment.

[0159] <3.2. Experimental results of STRes3dRNN model parameter optimization>

[0160] The classification accuracy results of the initial learning rate α and the image sequence depth T are as follows Figure 7 As shown in the left figure. When the initial learning rate α is fixed, as the image sequence depth T increases, the classification accuracy usually improves, indicating that the network can learn more complex spatiotemporal features. However, too deep an image sequence may lead to overfitting and affect the accuracy. For example, when α=0.001, T increases from 16 to 64, and the accuracy drops from 97.40% to 95.78%. When the image sequence depth T is fixed, a proper reduction in the initial learning rate α helps the network learn faster, but too low an initial learning rate may cause the network to fall into a local minimum, affecting the classification performance. Among all the test combinations, the network model performs best with an accuracy of 97.40% when the initial learning rate α is 0.001 and the image sequence depth T is 16. This shows that this parameter combination has achieved a good balance between the network's learning ability and generalization ability, effectively avoiding overfitting while making full use of computing resources.

[0161] The comparison results of training sample size N are as follows Figure 7 As shown in the right figure. When the proportion of training samples is only 60%, despite the small number of training samples, our proposed method can effectively extract deep spatiotemporal features, thereby achieving a high classification accuracy (Accuracy = 83.58 ± 0.61% and Loss = 1.04 ± 0.04). As the proportion of training samples increases, the classification accuracy gradually improves and tends to stabilize. When the sample size is sufficient, the classification accuracy can reach the highest. Therefore, especially when the sample size is scarce, our method can also show good classification robustness.

[0162] <3.3. Ablation comparison test results>

[0163] All six network models were classified and tested through the ten-fold cross-validation method, and the quantification results are as Figure 8 shown. Among the four evaluation indicators, STRes3dRNN scored the highest, specifically Pre = 97.37 ± 0.86%, Rec = 97.40 ± 1.04%, F1 = 97.38 ± 0.86%, and Acc = 97.40%. Among them, for the recognition of normal subjects, the F1 score reached the highest of 98.40%, which indicates that the STRes3dRNN model has almost no errors in the judgment of health conditions and is very reliable in the diagnosis of lung function of healthy people.

[0164] The above-mentioned embodiments are the preferred embodiments of the present invention, which are only used to conveniently illustrate the present invention and do not impose any formal restrictions on the present invention. Any person with ordinary knowledge in the technical field to which the present invention pertains, if without departing from the technical features of the present invention, makes local modifications or equivalent embodiments by using the technical content disclosed in the present invention, and without departing from the technical feature content of the present invention, still belongs to the scope of the technical features of the present invention.

Claims

1. An evaluation system, characterized in that, Including: A storage unit for storing continuous voltage data measured by the EIT device; An EIT image sequence generation unit that reads the continuous voltage data of the storage unit to generate an EIT image sequence; An STRes3dRNN model that reads the EIT image sequence of the EIT image sequence generation unit, and the result output is passed to the probability calculation unit; A probability calculation unit for giving the probability percentages of health, blockage, and injury; The STRes3dRNN model includes a network input layer, an initial stacking layer, an intermediate stacking layer, a final stacking layer, and a network output layer connected in sequence; Among them, the network input layer is used to read the EIT image sequence; Among them, the initial stacking layer is sequentially stacked by a 3D residual initial unit, an intermediate 3D residual standard unit, and a final 3D residual standard unit; the 3D residual initial unit is used to initially extract high-dimensional image features, the intermediate 3D residual standard unit is used to reduce the dimension of the high-dimensional image features, and the final 3D residual standard unit is used to reduce the dimension and deepen the feature information; Among them, the intermediate stacking layer includes: n sequentially connected single-body intermediate stacking layers; the single-body intermediate stacking layer is stacked by a HAM unit, a 3D residual upsampling unit, and a final 3D residual standard unit; n is a natural number greater than or equal to 1; the HAM unit optimizes the weight of the allocated information, eliminates redundant information, and strengthens the feature information; the 3D residual upsampling unit upsamples the low-dimensional image features; the final 3D residual standard unit is used to reduce the dimension and deepen the feature information; Among them, the final stacking layer is sequentially stacked by a 3D global average pooling layer, a BiLSTM layer, and a fully connected layer; the 3D global average pooling layer reads the image sequence information obtained by the intermediate stacking layer and extracts its feature vector; the BiLSTM layer processes and outputs the feature vector; the fully connected layer maps the input feature vector to the number of target classifications and obtains the mapped score value, and the target classifications are: health, blockage, and injury; The 3D residual initial unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, and a 3D max pooling layer connected in sequence; the three-dimensional convolutional layer is used to convolve and extract the image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data of the previous step through a non-linear activation function to improve the representation ability of the features; the 3D max pooling layer compresses the features using different channels to achieve invariance, extracts and outputs the maximum value; Both the intermediate 3D residual standard unit and the final 3D residual standard unit include a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence; the 3D convolutional layer is used to convolve and extract the image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data given by the batch normalization layer through a non-linear activation function to improve the representation ability of the features; the additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation; The 3D residual upsampling unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence. The 3D convolutional layer is used to convolve and extract image sequence features. The batch normalization layer normalizes the sequence features to accelerate the convergence speed. The activation layer processes the data of the previous step through a non-linear activation function to improve the feature representation ability. The additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation.

2. An evaluation system according to claim 1, characterized in that, The HAM unit includes a spatio-temporal attention module, a channel attention module, and a motion attention module. The spatio-temporal attention module, the channel attention module, and the motion attention module are applied to enhance the feature information in the spatio-temporal, channel, and motion aspects respectively, and the enhanced information is weighted to obtain a fused and enhanced data with the same number of layers and dimensions for spatio-temporal, channel, and motion information.

3. An evaluation system according to claim 1, characterized in that, The BiLSTM layer is composed of two channels: forward propagation and backward propagation. Forward and backward processing are carried out simultaneously, and the processing time step is determined by the input image sequence, automatically cycling the same number of time steps as the number of sequences. The forward and backward processing results are concatenated to obtain a high-dimensional feature vector.

4. An evaluation system according to claim 1, characterized in that The working method of the probability calculation unit is as follows: Softmax(z i ) represents the predicted probability of the i-th class, representing the raw score of the i-th class.

5. An evaluation method that does not directly aim at disease diagnosis, characterized in that, It includes the following steps: S100, obtaining continuous voltage data to generate an EIT image sequence I. The image sequence I is a 16-frame EIT image sequence of 64×64×3, where 64×64 is the image size and 3 is a three-channel RGB image. S200, transmitting the image sequence I to the initial stacking layer. The initial stacking layer is stacked by 3D residual initial units, intermediate 3D residual standard units, and end 3D residual standard units. S200 includes the following sub-steps: S201, inputting the image sequence I into the 3D residual initial unit to obtain high-dimensional features of a 64-layer 7×7×7 image sequence. S202, inputting the high-dimensional features obtained in step S201 into the intermediate 3D residual standard unit to obtain downsampled image features. S203, inputting the image features obtained in step S202 into the end 3D residual standard unit to obtain a downsampled and feature-information deepened image sequence. S300, inputting the result calculated by the initial stacking layer into the intermediate stacking layer. The intermediate stacking layer includes: 3 sequentially connected single-body intermediate stacking layers. The single-body intermediate stacking layer is stacked by HAM units, 3D residual upsampling units, and 3D residual standard units. S300 includes the following sub-steps: S301, the HAM unit of the first single-body intermediate stacking layer reads a 128-layer 1×1×1 image sequence, optimizes the allocation of information weights, eliminates redundant information, strengthens feature information, and outputs a 128-layer 1×1×1 image sequence. S302, the 3D residual upsampling unit of the first single-body intermediate stacking layer reads a 128-layer 1×1×1 image sequence and outputs a 128-layer 3×3×3 image sequence. S303, the end 3D residual standard unit of the first single-body intermediate stacking layer reads a 128-layer 3×3×3 image sequence and outputs a 256-layer 1×1×1 image sequence. S304. The HAM unit of the second monomer's middle stacked layer reads the image sequence of a 256-layer 1×1×1 image sequence, optimizes the weight allocation of information, eliminates redundant information, strengthens feature information, and outputs the image sequence of a 256-layer 1×1×1 image sequence; S305. The 3D residual upsampling unit of the second monomer's middle stacked layer reads the 256-layer 1×1×1 image sequence and outputs a 256-layer 3×3×3 image sequence; S306. The final 3D residual standard unit of the second monomer's middle stacked layer reads the image sequence of a 256-layer 3×3×3 image sequence and outputs the image sequence of a 512-layer 1×1×1 image sequence; S307. The HAM unit of the third monomer's middle stacked layer reads the image sequence of a 512-layer 1×1×1 image sequence, optimizes the weight allocation of information, eliminates redundant information, strengthens feature information, and outputs the image sequence of a 512-layer 1×1×1 image sequence; S308. The 3D residual upsampling unit of the third monomer's middle stacked layer reads the 512-layer 1×1×1 image sequence and outputs a 512-layer 3×3×3 image sequence; S309. The final 3D residual standard unit of the third monomer's middle stacked layer reads the image sequence of a 512-layer 3×3×3 image sequence and outputs a 512-layer 1×1×1 image sequence; S400. The result calculated by the middle stacked layer is input into the final stacked layer, and the final stacked layer includes a three-dimensional average pooling layer, a BiLSTM layer, and a fully connected layer connected in sequence; Step S400 further includes the following sub-steps: S401. The three-dimensional average pooling layer reduces the dimension of the 512-layer 1×1×1 image sequence and extracts its feature vector; S402. The BiLSTM layer processes and outputs a 128-layer feature vector; S403. The fully connected layer maps the input feature vector to the target classification and obtains the mapped score value; among them, the number of target classifications is 3, which respectively represent healthy, blocked, and damaged; The 3D residual initial unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, and a 3D max pooling layer connected in sequence; the 3D convolutional layer is used to convolve and extract the image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data of the previous step through a non-linear activation function to improve the feature representation ability; the 3D max pooling layer compresses the features using different channels to achieve invariance, extracts and outputs the maximum value; Both the middle 3D residual standard unit and the final 3D residual standard unit include a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence; the 3D convolutional layer is used to convolve and extract the image sequence features; the batch normalization layer is used to normalize the sequence features to accelerate the convergence speed; the activation layer processes the data given by the batch normalization layer through a non-linear activation function to improve the feature representation ability; the additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation; The 3D residual upsampling unit includes a 3D convolutional layer, a batch normalization layer, an activation layer, a 3D convolutional layer, a batch normalization layer, an additional layer, and an activation layer connected in sequence. The 3D convolutional layer is used to convolve and extract image sequence features. The batch normalization layer normalizes the sequence features to accelerate the convergence speed. The activation layer processes the data of the previous step through a non-linear activation function to improve the feature representation ability. The additional layer is used to randomly turn off some nodes to prevent overfitting and avoid network degradation.

6. The evaluation method according to claim 5, wherein It also includes: S500, the result of the last stacking layer is input into the Softmax function to output the predicted probability, that is, the probability percentages of healthy, blocked, and damaged: where Softmax(z i ) represents the predicted probability of the i-th class, represents the raw score of the i-th class.

Citation Information

Patent Citations

  • Lung function evaluation method and system based on time characteristics of lung electrical impedance image

    CN118121183A

  • Lung function evaluation method, system and equipment based on respiratory impedance change

    CN118121184A

  • CXR image classification method and system based on residual convolution and multi-head self-attention

    CN115995015A

  • Motor imagery electroencephalogram signal classification method based on multi-scale convolution and self-attention

    CN118349906A