A Functional Near-Infrared Spectroscopic Signal Enhancement and Classification Method Based on Spatiotemporal Autocorrelation and Encoded Attention

CN122310192APending Publication Date: 2026-06-30NORTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610181474.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing fNIRS data augmentation methods cannot effectively replicate the spatiotemporal topological and temporal dynamic characteristics of signals, resulting in poor data authenticity. Furthermore, traditional classification models struggle to effectively integrate spatiotemporal features, leading to insufficient classification accuracy and generalization ability.

Method used

An augmented data equivalent to the original data is generated by using a spatiotemporal autocorrelation model, and combined with the STEAFNet deep learning model, the spatiotemporal features of the signal are captured and fused through three-level module collaborative processing.

Benefits of technology

It significantly improves the classification accuracy and generalization ability of fNIRS signals, solves the problem of insufficient sample size, and enhances the fidelity and diversity of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122310192A_ABST
    Figure CN122310192A_ABST
Patent Text Reader

Abstract

A functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and attention encoding is proposed to address the limitations of limited sample size and insufficient generalization ability of classification models in functional near-infrared spectral signal data. This method is applicable to the auxiliary diagnosis of neurodevelopmental disorders such as attention deficit hyperactivity disorder (ADHD). The method first preprocesses the raw near-infrared spectral data, extracting oxyhemoglobin signals and calculating spatial autocorrelation parameters, channel-level temporal autocorrelation parameters, and the eigenvalue distribution of the functional connectivity matrix. Then, enhanced data is generated based on the spatiotemporal autocorrelation model. Finally, the raw and enhanced data are input into the STEAFNet deep learning model for accurate classification. This invention generates high-quality enhanced data, expands the sample size, and maintains high fidelity. Combined with deep learning, it improves classification accuracy and generalization ability, providing technical support for clinical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of neuroimaging signal processing and clinical auxiliary diagnosis technology, specifically to a data enhancement and classification method for functional near-infrared spectral signals, which is particularly suitable for the auxiliary diagnosis of neurodevelopmental disorders such as attention deficit hyperactivity disorder. Background Technology

[0002] Attention Deficit Hyperactivity Disorder (ADHD) is one of the most common neurodevelopmental disorders in childhood. Its core symptoms include inattention, hyperactivity, and impulsivity, which severely affect children's learning, social interactions, and daily lives. Traditional ADHD diagnosis mainly relies on behavioral observation and scale assessments, which are highly subjective, easily influenced by the assessor's experience, and lack objective neurobiological evidence, making early and accurate diagnosis difficult.

[0003] Functional near-infrared spectroscopy (fNIRS), a non-invasive neuroimaging technique, reflects brain region functional activity by detecting changes in the concentrations of oxyhemoglobin (HbO2) and deoxyhemoglobin (HbR) in the cerebral cortex. It offers advantages such as portability, resistance to motion interference, and applicability to pediatric populations, making it an important tool for researching the neural mechanisms of ADHD and aiding in diagnosis. However, fNIRS data faces significant bottlenecks in clinical applications: on the one hand, recruiting pediatric subjects is difficult and the data collection cycle is long, resulting in generally limited sample sizes; on the other hand, fNIRS signals have a low signal-to-noise ratio and significant inter-individual physiological differences, making them prone to overfitting and insufficient generalization ability when directly used to train deep learning classification models, thus failing to meet the reliability requirements of clinical diagnosis.

[0004] To address the issue of insufficient sample size, data augmentation techniques have become crucial. Existing fNIRS data augmentation methods often employ simple transformations such as signal noise addition, time shifting, and scaling, which only weakly expand sample diversity and fail to replicate the inherent spatiotemporal topological characteristics and dynamic patterns of fNIRS signals. This results in poor data realism and limited effectiveness in improving model classification performance. In recent years, research has proposed signal generation models based on spatiotemporal autocorrelation characteristics. By capturing the core statistical laws of spatial autocorrelation (SA) and temporal autocorrelation (TA) of neural signals, these models generate alternative time series with the same structure as the original data, providing a new approach for high-quality data augmentation. The effectiveness of this model has been validated in functional magnetic resonance imaging (fMRI) data, but its adaptation and application in fNIRS data have not yet been fully explored.

[0005] Regarding classification models, existing fNIRS signal classification methods rely on manual feature extraction in traditional machine learning, making it difficult to capture the complex spatiotemporal relationships of the signals. Conventional deep learning models often focus solely on temporal dynamics or spatial topology, lacking effective fusion of spatiotemporal features. The STEAFNet deep learning model simultaneously captures the spatial topological relationships and long-range temporal dependencies of fNIRS signals and adaptively fuses features through an attention mechanism, demonstrating excellent classification potential. However, the performance of this model is still limited by the scale and quality of the training data.

[0006] In summary, there is an urgent need for a high-quality data augmentation method that can accurately replicate the core spatiotemporal characteristics of fNIRS signals and combine it with an efficient deep learning classification model to solve the problems of insufficient sample size, poor realism of augmented data, and insufficient generalization ability of classification models in existing technologies, and promote the application of fNIRS technology in the clinical diagnosis of neurodevelopmental disorders such as ADHD. Summary of the Invention

[0007] To overcome the shortcomings of the prior art, the purpose of this invention is to provide a functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention. This method addresses the problems of existing enhancement methods that can only simply expand the sample size and cannot replicate the inherent spatial topological and temporal dynamic characteristics of fNIRS signals, resulting in poor data fidelity. It also addresses the technical problems of traditional models being unable to effectively integrate the spatiotemporal features of fNIRS signals and thus failing to fully utilize the information in the enhanced data, leading to limited classification accuracy.

[0008] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0009] A functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention, characterized by comprising the following steps:

[0010] Step 1: Perform motion artifact removal, filtering, and standardization on the raw near-infrared spectral data to extract the HbO2 signal;

[0011] Step 2: Based on the preprocessed oxyhemoglobin signal, calculate the spatial autocorrelation parameters, channel-level temporal autocorrelation parameters, and eigenvalue distribution of the functional connectivity matrix;

[0012] Step 3: Generate augmented data based on the spatiotemporal autocorrelation model. By adjusting the random seed to control the randomness of the sequence, multiple statistically equivalent but different augmented data are generated for a single subject. The augmented data replicates the spatial topological characteristics and temporal dynamic characteristics of the original near-infrared spectral signal.

[0013] Step 4: Input the raw data and augmented data into the STEAFNet deep learning model. Through the collaborative processing of three-level modules, the classification results are output.

[0014] The detailed steps are as follows:

[0015] Step 1: Perform motion artifact removal, bandpass filtering, and z-score normalization on the raw fNIRS data to eliminate individual baseline differences, and extract oxyhemoglobin. The signal is used to obtain preprocessed data with dimensions C×T (C is the number of channels, and T is the length of the time series).

[0016] Step 2: Based on the preprocessed Signal, calculate three types of key parameters:

[0017] SA parameters: including spatial decay rate ( ) and long-range spatial stability ( By calculating the Euclidean distance (based on MNI coordinates) and Pearson correlation coefficient for all channel pairs, an exponential decay curve of "distance-correlation coefficient" is fitted, and the curve formula is as follows:

[0018] (1)

[0019] In the formula, Let Pearson correlation coefficient be the correlation coefficient between channel i and channel j. Let be the Euclidean distance between channel i and channel j. This is obtained by fitting the curve. and ;

[0020] Channel-level TA parameters: The Pearson correlation coefficient between adjacent time points in the full time series of a single channel is calculated using the following formula:

[0021] (2)

[0022] In the formula, For a single channel at time point t Signal value, Let T be the signal mean of the entire time series of this channel, and T be the length of the time series. The C values ​​can be obtained using this formula. value;

[0023] The eigenvalue distribution of the functional connective (FC) matrix: Calculate the C×C dimension... The Pearson correlation matrix of the signal is subjected to eigenvalue decomposition, and the sorted eigenvalue distribution is saved.

[0024] Step 3: Based on the spatiotemporal autocorrelation model, by constraining spatial topology and temporal dynamics, generate augmented data that is statistically equivalent to the original data;

[0025] Step 3 includes the following sub-steps:

[0026] Step S31: Using the correlation spectrum sampling algorithm, multiple independent C-line symbols are generated by setting different random seeds (e.g., seed=100, 200, 300). The long-memory Brown noise sequence (of length T) of the power spectrum is filtered by a 0.01Hz high-pass filter to ensure that the sequence has the long-memory smoothness characteristics of neural signals;

[0027] Step S32, based on the results obtained in step 2 , Based on the Euclidean distance between channels, construct a C×C spatial covariance matrix. The formula for calculating the matrix elements is as follows:

[0028] (3)

[0029] In the formula, The values ​​of channels i and j in the spatial covariance matrix are given. By constraining the multivariate normal distribution of the correlation spectrum sampling, the correlation between channels of the basic sequence is matched with the covariance matrix, thus embedding the spatial topological association of the original data.

[0030] Step S33: Add channel-specific white noise to each base sequence: based on the power spectrum of the original sequence and the target... The noise variance is calculated using the following formula:

[0031] (4)

[0032] In the formula, For the goal (i.e., the original channel) value), The power spectral amplitude of the original sequence is given by T, and the sequence length is given by T. Based on the random seed corresponding to step S31, independent white noise corresponding to this variance is generated and added to the base sequence channels of the corresponding groups, so that the power spectral amplitude of each generated sequence is... All are consistent with the original channel, but the details of signal fluctuations differ due to the difference in random seeds;

[0033] Step S34: Optimize using differential evolution algorithm and The objective is to minimize the mean squared error of the eigenvalues ​​of the generated sequence and the original data's FC matrix. The parameters are iteratively adjusted (population size 50, maximum iterations 1000 generations). The formula for calculating the mean squared error of the eigenvalues ​​is as follows:

[0034] (5)

[0035] In the formula, The k-th eigenvalue of the original FC matrix. To generate the k-th eigenvalue of the sequence FC matrix, where C is the number of channels; the above optimization process is executed for each group of basic sequences, and finally multiple augmented data (each augmented data dimension is C×T) are generated for a single subject that are highly matched with the spatial characteristics of the original data, statistically equivalent but with different signals.

[0036] Step 4: Input the original fNIRS data and the augmented data into the STEAFNet deep learning model for classification. This model fully utilizes the spatiotemporal information of the augmented data to achieve high-precision classification through the collaborative work of three-level modules: "spatiotemporal feature extraction - temporal feature capture - spatiotemporal feature fusion".

[0037] Step 4 includes the following sub-steps:

[0038] Step S41: Based on the spatial topological characteristics and local spatiotemporal correlation of fNIRS signals, the 3D-STRE module first extracts the coupling features of the channel dimension and the time dimension simultaneously through 3D convolution operation, and then assigns differentiated weights to the features of different brain regions through the spatial attention mechanism to enhance the feature contribution of key functional brain regions, suppress the interference of noise channels, and output a local spatiotemporal feature map focusing on the core region.

[0039] Step S42: In order to capture the long-range time dynamic dependence of the fNIRS signal, the TAWE module uses BiLSTM to model the forward and reverse time correlation of the signal, and then uses the time attention mechanism to adaptively weight the features of different time steps, highlighting the effective features of the task-related time period, weakening the redundant information of the rest or interference period, and outputting the sequence features that enhance the key time dynamics.

[0040] Step S43: To achieve efficient complementarity of spatial and temporal features, the FAFC module uses a branch attention mechanism to perform fine-grained weight allocation on the spatiotemporal features of the 3D-STRE module and the temporal features of the TAWE module, adaptively fuses the two types of core features, and then maps and classifies the fused features through a fully connected network, finally outputting the classification result.

[0041] The Euclidean distance between channels is calculated using the MNI coordinates of the fNIRS channels; the variance of the white noise is obtained by comparing the power spectrum of the original sequence with the target... The determination is made by reverse reasoning; the "statistical equivalence" is reflected in the spatial autocorrelation parameters of the enhanced data and the original data. The eigenvalue distributions of the functional connectivity matrix are consistent;

[0042] In step 4, the STEAFNet deep learning model's 3D spatiotemporal representation and encoding (3D-STRE) module extracts local spatiotemporal features through 3D convolution and spatial attention mechanisms; the temporal attention weighted encoding (TAWE) module captures long-range temporal dependencies through a bidirectional long short-term memory network (BiLSTM) and temporal attention mechanisms; and the fine-grained attention fusion and classification (FAFC) module outputs classification results after fusing spatial and temporal features through branch attention.

[0043] The classification results are used to differentiate ADHD from healthy controls, and as an auxiliary diagnostic tool for other neurodevelopmental disorders.

[0044] The beneficial effects of this invention are:

[0045] 1. This invention is applicable to the processing of fNIRS signals, especially to the auxiliary diagnosis of neurodevelopmental disorders such as ADHD, and is compatible with data acquired by common fNIRS devices.

[0046] 2. This invention employs a spatiotemporal autocorrelation model for data augmentation, replicating spatial topological characteristics through SA parameter constraints. Matching ensures dynamic consistency over time, and combining it with differential evolution algorithms to optimize spatial parameters not only solves the problem that traditional augmentation methods cannot restore the inherent statistical regularity of signals, but also significantly improves the fidelity and diversity of augmented data.

[0047] 3. This invention combines high-quality augmented data with the STEAFNet deep learning model, and captures spatiotemporal features through a three-level module collaboration and adaptive fusion. Compared with traditional classification methods, it effectively alleviates the overfitting problem caused by insufficient sample size and significantly improves the accuracy and generalization ability of fNIRS signal classification. Attached Figure Description

[0048] Figure 1This is an overall flowchart of a functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention according to the present invention;

[0049] Figure 2 This is a schematic diagram illustrating the data augmentation generation principle based on the spatiotemporal autocorrelation model of the present invention.

[0050] Figure 3 This is a diagram of the STEAFNet deep learning model architecture. Detailed Implementation

[0051] The invention will be further described below with reference to the accompanying drawings.

[0052] A functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention, such as Figure 1 As shown, follow these steps:

[0053] Step 1: fNIRS signal acquisition is performed on the subject. After acquisition, preprocessing is performed to extract oxyhemoglobin (…). )Signal;

[0054] In this embodiment, the ETG-One fNIRS system (Hitachi Medical Corporation) was used for data acquisition, with a sampling rate of 10Hz. The device emits near-infrared light at wavelengths of 695nm and 830nm. Subjects were required to complete a verbal fluency task (VFT) to obtain raw fNIRS data. Preprocessing was then performed on the raw fNIRS data.

[0055] 1. Motion artifact removal: Anomaly segments of the signal are identified using the sliding window standard deviation method, and corrected using linear interpolation. Simultaneously, a time-domain derivative distribution repair algorithm is used to eliminate peak artifacts and baseline drift.

[0056] 2. Bandpass filtering: A Butterworth bandpass filter of 0.01-0.2Hz is applied to filter out physiological interferences such as heartbeat and respiration, as well as instrument noise, while retaining the effective signal of cerebral hemodynamics;

[0057] 3. z-score normalization: The signal of each channel is standardized to eliminate baseline differences between individuals, so that the signal mean is 0 and the standard deviation is 1; after preprocessing, a C×T dimension is obtained. Signal matrix (C is the number of channels, T is the length of the time series).

[0058] Step 2: Based on the preprocessed For the signal, calculate the SA parameters, channel-level TA parameters, and FC matrix eigenvalue distribution:

[0059] 1. SA parameter calculation: Calculate the Euclidean distance of all channel pairs. (Based on channel-specific MNI coordinates, channel positioning references the international 10 / 20 system) and Pearson correlation coefficient A genetic algorithm was used to fit the exponential decay curve of the "distance-correlation coefficient" relationship. The curve formula is as follows:

[0060] (1)

[0061] In the formula, Let Pearson correlation coefficient be the correlation coefficient between channel i and channel j. To find the Euclidean distance between channel i and channel j, the genetic algorithm population size is set to 50, the maximum number of iterations is 1000 generations, and the function tolerance is... The parameter boundary is , After fitting, the scores of each subject were obtained. and ;

[0062] 2. TA parameter calculation: For all frames of each channel. For the signal, calculate the Pearson correlation coefficient between adjacent time points in the full time series of a single channel. The formula is as follows:

[0063] (2)

[0064] In the formula, For a single channel at time point t Signal value, Let T be the signal mean of the entire time series of this channel, and T be the length of the time series. The C values ​​can be obtained using this formula. value;

[0065] 3. Calculation of eigenvalue distribution of FC matrix:

[0066] Calculate the C×C dimension The signal Pearson correlation matrix (FC matrix) is used to perform eigenvalue decomposition on the FC matrix, and the eigenvalues ​​are arranged in descending order and the distribution results are saved.

[0067] Step 3: Based on the core parameters calculated in Step 2, pseudo-data equivalent to the original data statistics is generated by constraining spatial topology and temporal dynamics. This specifically includes the following sub-steps:

[0068] Step S31: Use the correlation spectrum sampling algorithm to generate C symbols. The power spectrum is a long-memory Brownian noise sequence with a sequence length of T; the generated sequence is subjected to a 0.01Hz high-pass filter to ensure that the sequence has the long-memory smoothing characteristics of neural signals;

[0069] Step S32, based on the results obtained in step 2 , Euclidean distance between the passage Construct a C×C spatial covariance matrix. The formula for calculating the matrix elements is as follows:

[0070] (3)

[0071] In the formula, The values ​​of channels i and j in the spatial covariance matrix are given. By constraining the multivariate normal distribution of the correlation spectrum sampling, the correlation between channels of the basic sequence is matched with the covariance matrix, thus embedding the spatial topological association of the original data.

[0072] Step S33: Add channel-specific white noise to each base sequence to improve the quality of the generated sequence. Consistent with the original channel; based on the power spectrum of the original sequence and the target (i.e., the TA-Δ1 value of the original channel), the noise variance is calculated using the following formula:

[0073] (4)

[0074] In the formula, For the goal (i.e., the original channel) value), The power spectral amplitude of the original sequence is given by T, where T is the sequence length. Different random seeds (e.g., seed=100, 200, 300) are used to generate white noise corresponding to this variance (following a mean of 0 and a variance of...). (Gaussian distribution), white noise generated by different random seeds is added to the corresponding channels respectively, so that each generated sequence is Consistent with the original channel, but the specific noise fluctuation details differ depending on the random seed;

[0075] Step S34: Optimize using differential evolution algorithm and With the objective of minimizing the mean square error of the eigenvalues ​​of the generated sequence and the original data's FC matrix, the formula for calculating the mean square error of the eigenvalues ​​is as follows:

[0076] (5)

[0077] In the formula, The k-th eigenvalue of the original FC matrix. To generate the k-th eigenvalue of the sequence FC matrix, where C is the number of channels; the population size of the differential evolution algorithm is set to 50 and the maximum number of iterations is 1000 generations. For the same subject, the parameter optimization process is executed based on the base sequences generated by different random seeds in step S31. After iteratively adjusting the parameters, multiple statistically equivalent but signal-different augmented data are finally generated for each subject. The dimension of the augmented data is still C×T.

[0078] Step 4: Input the raw data and augmented data into the STEAFNet deep learning model to distinguish between ADHD and healthy subjects. This includes the following sub-steps:

[0079] Step S41: Hierarchical Construction and Strict Partitioning of the Dataset:

[0080] 1. Core principles of data splitting: Augmented data is only used to expand the training set, while the validation and test sets use original real data throughout. Furthermore, the training, validation, and test sets adopt a "subject-independent" splitting strategy to prevent data leakage from the source.

[0081] 2. Specific division process:

[0082] Raw data pool: Contains raw, real fNIRS data from all children with ADHD and all healthy subjects;

[0083] Augmented data generation: Multiple augmented data are generated only from the subject data in the original data pool that are to be included in the training set, according to different random seeds. The augmented data does not participate in the construction of the validation set and test set.

[0084] Five-fold cross-validation split: Each fold splits the original data pool into three parts: 70% training set, 10% validation set, and 20% test set. Then, augmented data corresponding to the original training data is added to the training set to form a mixed training set of "original training data + augmented data". The validation set and test set retain the original real data attributes.

[0085] Step S42: The STEAFNet model is implemented using PyTorch 2.5.0 and Torchvision 0.20.0, running on an NVIDIA GeForce GTX 1050 GPU (CUDA Compute Capability 6.1). The core parameters of the model were determined through 60 iterations of training using the Optuna Bayesian optimization algorithm. The optimization objective was to maximize the classification accuracy using 5-fold cross-validation. Parameters to be optimized included: the number of output channels in the 3D convolutional layers, the number of hidden units in the BiLSTM layers, the dropout rate of each module, the activation function type, the channel reduction ratio, the output dimension of the fully connected layers, and the initial learning rate and weight decay coefficient of the AdamW optimizer. The model adopts a dual-branch deep learning architecture. Through the synergistic effect of the 3D spatiotemporal representation and encoding (3D-STRE) module, the temporal attention weighted encoding (TAWE) module, and the fine-grained attention fusion and classification (FAFC) module, it effectively captures and fuses the spatial topology and temporal dynamic features of the fNIRS signal. The specific structure is as follows:

[0086] 1. 3D-STRE module:

[0087] This module focuses on "preserving the true anatomical topology of fNIRS channels and extracting local spatiotemporal correlations." It first processes the pre-processed... The signal (dimension C×T) is projected onto a fixed 5×9 anatomical grid, filling only the positions corresponding to the 22 true channels, and padding the remaining positions with zeros to avoid interpolation noise. This is then stacked along the time dimension to form a 3D spatiotemporal tensor (dimensions B×C×T×H×W, where B is the batch size, and H and W are the spatial grid dimensions). The module internally consists of three stacked 3D convolutional attention blocks (3D-CABs), each containing a 3D convolutional layer, a spatial attention submodule, and residual connections.

[0088] 1) 3D Convolutional Layer: Convolutional operations are performed on 3D tensors using multi-scale convolutional kernels (using kernels of different sizes along the time dimension), simultaneously capturing channel correlations in the spatial dimension and local dynamic changes in the temporal dimension; then, an activation function layer is introduced to introduce non-linear transformations, enhancing the model's ability to express complex features; finally, a max pooling layer is used to downsample in the spatiotemporal dimension, preserving key features while compressing data dimensions, improving computational efficiency and feature robustness;

[0089] 2) Spatial attention submodule: First, the 3D convolution output feature map is convolved with 1×1×1 to achieve local linear projection. Then, the polarity cancellation problem of fNIRS signal is alleviated by absolute value operation. After channel reduction, a spatial attention map is generated by 1×1×1 convolution and sigmoid activation. This attention map is multiplied element-wise with the original feature map to highlight the features of key brain regions related to ADHD pathology, such as the prefrontal cortex, and suppress interference from irrelevant channels.

[0090] 3) Residual connection: When the dimensions of the input features and the output features after attention weighting do not match, the dimensions are aligned by 1×1×1 convolution and residuals are added to ensure smooth gradient propagation and accelerate model convergence.

[0091] 2. TAWE module:

[0092] This module focuses on capturing the long-range temporal dependencies of fNIRS signals, overcoming the limitations of 3D convolution in long-term sequence modeling. Its core consists of stacked Temporal Attention Blocks (TABs), each containing a Bidirectional Long Short-Term Memory (BiLSTM) network, a Temporal Branch Attention (TBAM) mechanism, and residual connections.

[0093] 1) BiLSTM layer: Receives raw fNIRS time series data, models the time correlation between "past to present" and "future to present" through forward LSTM and backward LSTM respectively, fully captures bidirectional long-range time dependencies, and outputs hidden state sequences containing complete time series information;

[0094] 2) TBAM mechanism: First, perform element-level absolute value operation on the hidden state sequence output by BiLSTM, and then generate time step attention weights through linear transformation and nonlinear activation. These weights are multiplied element by element with the hidden state sequence to adaptively highlight the effective features of key task periods (such as the task execution period in speech fluency tasks) and weaken redundant information in rest periods or interference periods.

[0095] 3) Residual connection: Align the input and output feature dimensions through linear transformation and perform residual addition to avoid the gradient vanishing problem caused by the increase of model depth.

[0096] 3. FAFC module:

[0097] This module is responsible for adaptively fusing the spatial features output by the 3D-STRE module and the temporal features output by the TAWE module, thus solving the problem of low information utilization caused by direct stitching of heterogeneous features.

[0098] 1) Feature mapping: First, spatial features and temporal features are mapped to the same dimensional space through fully connected layers to form feature vectors of the same dimension;

[0099] 2) Branch Attention Weighting: Reuse the core structure of TBAM to calculate fine-grained attention weights for the two types of features after mapping. The weight allocation covers each dimension of the feature, realizing adaptive feature recalibration for features of different branches and different dimensions, strengthening the contribution of discriminative features and suppressing the interference of noisy features;

[0100] 3) Classification layer: The weighted and fused joint features are input into the fully connected network, and after nonlinear transformation, the classification probability is output through Softmax activation to complete the identification of ADHD and healthy control groups.

[0101] Example 2

[0102] A functional near-infrared spectral signal processing device based on spatiotemporal autocorrelation and coded attention, characterized in that it includes a preprocessing module, a data augmentation module, and a classification module;

[0103] Preprocessing module: Used to perform motion artifact removal, bandpass filtering, and z-score normalization on the raw fNIRS data, and to extract oxyhemoglobin ( )Signal;

[0104] Data augmentation module: Based on the spatiotemporal autocorrelation model, it generates augmented data that is statistically equivalent to the original data through correlation spectrum sampling, spatial covariance matrix construction, channel-specific white noise addition, and parameter optimization.

[0105] The classification module is used to input the raw data and augmented data into the STEAFNet deep learning model, and output the classification results through the collaborative processing of the 3D-STRE module, TAWE module and FAFC module.

Claims

1. A functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention, characterized in that, Includes the following steps: Step 1: Perform motion artifact removal, bandpass filtering, and z-score normalization on the raw fNIRS data to eliminate individual baseline differences, and extract oxyhemoglobin. The signal is used to obtain preprocessed data with dimensions C×T (C is the number of channels, and T is the length of the time series). Step 2: Based on the preprocessed Signal, calculate three types of key parameters: SA parameters: including spatial decay rate ( ) and long-range spatial stability ( By calculating the Euclidean distance (based on MNI coordinates) and Pearson correlation coefficient for all channel pairs, an exponential decay curve of "distance-correlation coefficient" is fitted, and the curve formula is as follows: (1) In the formula, Let Pearson correlation coefficient be the correlation coefficient between channel i and channel j. Let be the Euclidean distance between channel i and channel j. This is obtained by fitting the curve. and ; Channel-level TA parameters: The Pearson correlation coefficient between adjacent time points in the full time series of a single channel is calculated using the following formula: (2) In the formula, For a single channel at time point t Signal value, Let T be the signal mean of the entire time series of this channel, and T be the length of the time series. The C values ​​can be obtained using this formula. value; The eigenvalue distribution of the functional connective (FC) matrix: Calculate the C×C dimension... The Pearson correlation matrix of the signal is subjected to eigenvalue decomposition, and the sorted eigenvalue distribution is saved. Step 3: Based on the spatiotemporal autocorrelation model, by constraining spatial topology and temporal dynamics, generate augmented data that is statistically equivalent to the original data. The steps are as follows: Step S31: Using the correlation spectrum sampling algorithm, multiple independent C-line symbols are generated by setting different random seeds (e.g., seed=100, 200, 300). The long-memory Brown noise sequence (of length T) of the power spectrum is filtered by a 0.01Hz high-pass filter to ensure that the sequence has the long-memory smoothness characteristics of neural signals; Step S32, based on the results obtained in step 2 , Based on the Euclidean distance between channels, construct a C×C spatial covariance matrix. The formula for calculating the matrix elements is as follows: (3) In the formula, The values ​​of channels i and j in the spatial covariance matrix are given. By constraining the multivariate normal distribution of the correlation spectrum sampling, the correlation between channels of the basic sequence is matched with the covariance matrix, thus embedding the spatial topological association of the original data. Step S33: Add channel-specific white noise to each base sequence: based on the power spectrum of the original sequence and the target... The noise variance is calculated using the following formula: (4) In the formula, For the goal (i.e., the original channel) value), The power spectral amplitude of the original sequence is given by T, and the sequence length is given by T. Based on the random seed corresponding to step S31, independent white noise corresponding to this variance is generated and added to the base sequence channels of the corresponding groups, so that the power spectral amplitude of each generated sequence is... All are consistent with the original channel, but the details of signal fluctuations differ due to the difference in random seeds; Step S34: Optimize using differential evolution algorithm and The objective is to minimize the mean squared error of the eigenvalues ​​of the generated sequence and the original data's FC matrix. The parameters are iteratively adjusted (population size 50, maximum iterations 1000 generations). The formula for calculating the mean squared error of the eigenvalues ​​is as follows: (5) In the formula, The k-th eigenvalue of the original FC matrix. To generate the k-th eigenvalue of the sequence FC matrix, where C is the number of channels; the above optimization process is performed on each set of basic sequences, and finally multiple augmented data (each augmented data dimension is C×T) are generated for a single subject, which are highly matched with the spatial characteristics of the original data, statistically equivalent but with different signals. Step 4: Input the original fNIRS data and the augmented data into the STEAFNet deep learning model for classification. This model fully utilizes the spatiotemporal information of the augmented data to achieve high-precision classification through the collaborative work of three-level modules: "spatiotemporal feature extraction - temporal feature capture - spatiotemporal feature fusion". The steps are as follows: Step S41: Based on the spatial topological characteristics and local spatiotemporal correlation of fNIRS signals, the 3D-STRE module first extracts the coupling features of the channel dimension and the time dimension simultaneously through 3D convolution operation, and then assigns differentiated weights to the features of different brain regions through the spatial attention mechanism to enhance the feature contribution of key functional brain regions, suppress the interference of noise channels, and output a local spatiotemporal feature map focusing on the core region. Step S42: In order to capture the long-range time dynamic dependence of the fNIRS signal, the TAWE module uses BiLSTM to model the forward and reverse time correlation of the signal, and then uses the time attention mechanism to adaptively weight the features of different time steps, highlighting the effective features of the task-related time period, weakening the redundant information of the rest or interference period, and outputting the sequence features that enhance the key time dynamics. Step S43: To achieve efficient complementarity of spatial and temporal features, the FAFC module uses a branch attention mechanism to perform fine-grained weight allocation on the spatiotemporal features of the 3D-STRE module and the temporal features of the TAWE module, adaptively fuses the two types of core features, and then maps and classifies the fused features through a fully connected network, finally outputting the classification result.

2. The functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention as described in claim 1, characterized in that, The Euclidean distance between channels is calculated using the MNI coordinates of the fNIRS channels; the variance of the white noise is obtained by comparing the power spectrum of the original sequence with the target... The "statistical equivalence" is determined by reverse deduction; it is reflected in the spatial autocorrelation parameters of the enhanced data and the original data. The eigenvalue distributions of the functional connectivity matrix are consistent.

3. The functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention as described in claim 1, characterized in that, In step 4, the STEAFNet deep learning model's 3D spatiotemporal representation and encoding (3D-STRE) module extracts local spatiotemporal features through 3D convolution and spatial attention mechanisms; the temporal attention weighted encoding (TAWE) module captures long-range temporal dependencies through a bidirectional long short-term memory network (BiLSTM) and temporal attention mechanisms; and the fine-grained attention fusion and classification (FAFC) module outputs classification results after fusing spatial and temporal features through branch attention.

4. The functional near-infrared spectral signal enhancement and classification method based on spatiotemporal autocorrelation and coded attention as described in claim 1, characterized in that, The classification results are used to differentiate ADHD from healthy controls, and as an auxiliary diagnostic tool for other neurodevelopmental disorders.