A functional near-infrared spectroscopy signal decoding method, device, equipment and medium
By integrating neurodynamics and temporal features into the fNIRS signal decoding method, the problem of parameter redundancy in existing technologies is solved, achieving efficient physiological state recognition and accurate neural network decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGCHUN UNIV
- Filing Date
- 2025-07-08
- Publication Date
- 2026-06-12
Smart Images

Figure CN120850032B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interfaces, and more particularly to a method, apparatus, device, and medium for decoding functional near-infrared spectral signals. Background Technology
[0002] Brain-computer interface (BCI) is a technology that establishes a direct communication channel between the human brain and external devices. It aims to translate the brain's intentions into recognizable control signals in real time without relying on peripheral nerve and muscle pathways. This technology is widely used in motor function reconstruction, neurorehabilitation, assisted communication, and human-computer interaction, demonstrating significant value, particularly in stroke rehabilitation, motor disorder assistance, silent communication, and intelligent control systems.
[0003] Functional near-infrared spectroscopy (fNIRS) is a non-invasive neuroimaging technique used to quantify hemodynamic responses associated with brain neural activity. Based on the optical properties of blood, fNIRS uses near-infrared light in the wavelength range of 650 nm to 950 nm to measure changes in the concentrations of oxyhemoglobin (HbO) and deoxyhemoglobin (HbR) in the cerebral cortex.
[0004] The existing deep learning methods for fNIRS signal decoding have not fully utilized the features extracted and applied. Although many existing methods have achieved high decoding accuracy, their parameters are too redundant, making it difficult to achieve efficient decoding. Summary of the Invention
[0005] In view of this, the present invention provides a functional near-infrared spectral signal decoding method, apparatus, device and computer-readable storage medium, which can solve the problem that the feature extraction and application of deep learning methods based on fNIRS signal decoding are not sufficient. Although many existing methods have high decoding accuracy, the parameters are too redundant and it is difficult to achieve efficient decoding.
[0006] Some embodiments of the present invention provide a method for decoding functional near-infrared spectral signals. The present invention is described below from multiple aspects, and the embodiments and beneficial effects described below can be referred to mutually.
[0007] In a first aspect, the present invention provides a method for decoding functional near-infrared spectral signals, comprising:
[0008] Acquire functional near-infrared spectral signals;
[0009] The changes in oxyhemoglobin concentration and deoxyhemoglobin concentration were obtained based on functional near-infrared spectral signals. The changes in oxyhemoglobin concentration were preprocessed to obtain a first signal segment sequence based on time, and the changes in deoxyhemoglobin concentration were preprocessed to obtain a second signal segment sequence based on time.
[0010] The first signal segment sequence and the second signal segment sequence are paired to obtain the homologous signal segment sequence. Based on the homologous signal segment sequence, the homologous signal segment subsequence is determined. Based on the feature extraction of the homologous signal segment subsequence, the neurodynamic features and temporal features are obtained. The homologous signal segment subsequence is a sequence obtained based on the difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration.
[0011] A fusion feature input sequence is constructed based on neurodynamic features and temporal features. The fusion feature input sequence is then input into a neural network model to obtain multimodal collaborative representation data. This multimodal collaborative representation data is used to characterize the spatial correlation patterns between each signal segment.
[0012] Based on the mean pooling of multimodal collaborative representation data, global representation data is obtained, and physiological state recognition results are obtained by identifying the global representation data.
[0013] In one possible implementation of the first aspect above, obtaining the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration based on functional near-infrared spectral signals includes:
[0014] Based on the modified Beer-Lambert law, the functional near-infrared spectral signal was converted to obtain the changes in oxyhemoglobin concentration and the changes in deoxyhemoglobin concentration.
[0015] In one possible implementation of the first aspect described above, the step of preprocessing the change in oxyhemoglobin concentration includes:
[0016] Based on filtering of changes in oxyhemoglobin concentration and baseline calibration;
[0017] A sliding window based on time length is preset, and the change in oxyhemoglobin concentration is truncated based on the sliding window to obtain the first single task whole signal for physiological state recognition.
[0018] Based on a preset time window, the entire signal of the first single task is extracted to obtain the first signal segment sequence;
[0019] The steps for preprocessing changes in deoxyhemoglobin concentration include:
[0020] Based on filtering of deoxyhemoglobin concentration changes and baseline calibration;
[0021] Based on the sliding window to extract the change in deoxyhemoglobin concentration, a second single task whole signal is obtained for physiological state recognition.
[0022] The second signal segment sequence is obtained by extracting the entire signal of the second single task based on the time window.
[0023] In one possible implementation of the first aspect above, the homologous signal segment subsequence includes three signal segments extracted from the homologous signal segment sequence based on a time window. The three signal segments include: a first signal segment, and a second and a third signal segment determined based on the first signal segment.
[0024] In one possible implementation of the first aspect above, determining the subsequence of the same source signal segment based on the sequence of same source signal segments includes any one of the following:
[0025] Identify the first signal segment, which is the signal segment where the maximum difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration is located;
[0026] The two signal segments adjacent to and before the first signal segment are designated as the second and third signal segments; or, the second signal segment is separated from the first signal segment by one signal segment, and the third signal segment is adjacent to the first signal segment; or, the second signal segment is adjacent to the first signal segment, and the third signal segment is separated from the first signal segment by one signal segment.
[0027] In one possible implementation of the first aspect described above, neurodynamic features include the maximum Lyapunov index and the phase-locked value.
[0028] In one possible implementation of the first aspect above, temporal features are obtained based on the feature extraction of subsequences of signals from the same source, including:
[0029] Based on time-series modeling of subsequences of signals from the same source, the predicted output results are obtained;
[0030] The predicted output results are aggregated by time-dimensional mean to obtain statistical representation data;
[0031] Based on the cross-pair projection of statistical characterization data, projected characterization data is obtained;
[0032] Temporal characteristics are obtained based on the nonlinear transformation of the projection representation data.
[0033] In one possible implementation of the first aspect above, a fused feature input sequence is constructed based on neurodynamic features and temporal features, including:
[0034] By aligning the neurodynamic features and temporal features according to their spatial dimensions, we obtain the neurodynamic projection features and the temporal projection features.
[0035] The neurodynamic projection features and temporal projection features are concatenated to obtain the fused feature input sequence.
[0036] In one possible implementation of the first aspect described above, the fused feature input sequence is input into a neural network model to obtain multimodal collaborative representation data, including:
[0037] Encoded data is obtained by encoding the position of the input sequence based on fused features;
[0038] The encoded data is normalized, and then diverse features are captured from the normalized data to obtain diverse features;
[0039] Based on the diverse features and coded data residual connections, residual connection output data is obtained;
[0040] The residual connection output data is normalized, and the normalized residual connection output data is then fed forward to obtain the complex relationship fitting data.
[0041] Multimodal collaborative characterization data is obtained by fitting data with complex relationships and residual connection output data.
[0042] In one possible implementation of the first aspect above, the physiological state identification result is obtained by identifying global representation data, including:
[0043] Based on the nonlinear mapping of global representation data, nonlinear mapping representation data is obtained;
[0044] Regularized representation data is obtained by regularizing the representation data based on nonlinear mapping.
[0045] Based on the linear transformation of the regularized representation data, linear transformation representation data is obtained;
[0046] Physiological state identification results are obtained by classifying data based on linear transformation representation.
[0047] Secondly, this application provides a functional near-infrared spectral signal decoding device, comprising:
[0048] The acquisition module is used to acquire functional near-infrared spectral signals;
[0049] The preprocessing module is used to obtain the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration based on the functional near-infrared spectral signal, preprocess the change in oxyhemoglobin concentration to obtain a first signal segment sequence based on time, preprocess the change in deoxyhemoglobin concentration to obtain a second signal segment sequence based on time, pair the first signal segment sequence and the second signal segment sequence to obtain a homologous signal segment sequence, and determine a homologous signal segment subsequence based on the homologous signal segment sequence, wherein the homologous signal segment subsequence is a sequence obtained based on the difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration;
[0050] The feature extraction module is used to extract features from subsequences of homologous signal segments to obtain neurodynamic and temporal features;
[0051] The fusion module is used to construct a fusion feature input sequence based on the neurodynamic features and the temporal features, and input the fusion feature input sequence into the neural network model to obtain multimodal collaborative representation data. The multimodal collaborative representation data is used to represent the spatial correlation pattern between each signal segment.
[0052] The analysis module is used to obtain global representation data based on the mean pooling of the multimodal collaborative representation data, and to identify the physiological state recognition result by recognizing the global representation data.
[0053] Thirdly, the present invention provides an electronic device, comprising: a processor; and a memory storing computer program instructions.
[0054] Among them, the near-infrared spectral signal decoding method enables the processor to perform functions when computer program instructions are executed by the processor.
[0055] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform a near-infrared spectral signal decoding method.
[0056] Fifthly, the present invention discloses an apparatus comprising:
[0057] The memory is used to store instructions executed by one or more processors of the device, and the processor is one of the processors of the device for performing the methods disclosed in any of the first to fourth aspects above.
[0058] This invention extracts the neurodynamic and temporal features of subsequences of homologous signal segments, fuses and identifies these features, and achieves multi-angle extraction of fNIRS signal features, effectively improving the neural network's recognition of fNIRS signals. By limiting the number and length of time windows, signal features are effectively preserved while avoiding redundancy in signal data, which would reduce computational efficiency. Through temporal feature extraction and high-level logical reasoning, the amount of input data in the inference stage is reduced, while avoiding redundant modeling. Using only a lightweight encoder, core requirements can be focused on while discarding redundant configurations. Attached Figure Description
[0059] Figure 1 A flowchart of a functional near-infrared spectral signal decoding method according to an embodiment of the present invention is shown;
[0060] Figure 2 A model diagram of the functional near-infrared spectral signal decoding method according to an embodiment of the present invention is shown;
[0061] Figure 3 A structural diagram of the fNIRS signal acquisition device is shown;
[0062] Figure 4 A flowchart of step S120 of an embodiment of the present invention is shown;
[0063] Figure 5 A flowchart of step S130 of an embodiment of the present invention is shown;
[0064] Figure 6 A flowchart of temporal feature extraction according to an embodiment of the present invention is shown;
[0065] Figure 7 A graph showing the performance comparison of the models is provided.
[0066] Figure 8 A flowchart of step S140 of an embodiment of the present invention is shown;
[0067] Figure 9 A flowchart of Transformer model data processing according to an embodiment of the present invention is shown;
[0068] Figure 10 A flowchart of step S150 of an embodiment of the present invention is shown;
[0069] Figure 11 This is a block diagram of the device according to an embodiment of the present invention;
[0070] Figure 12 This is a block diagram of a SoC (System on Chip) according to an embodiment of the present invention.
[0071] Figure labels: 1. Detector; 2. Light source; 3. fNIRS signal acquisition channel. Detailed Implementation
[0072] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0073] To facilitate understanding of the technical solution of this invention, the technical problem to be solved by this invention will be explained first.
[0074] Functional near-infrared spectroscopy (fNIRS) is a non-invasive neuroimaging technique used to quantify hemodynamic responses associated with brain neural activity. Based on the optical properties of blood, fNIRS uses near-infrared light in the wavelength range of 650 nm to 950 nm to measure changes in the concentrations of oxyhemoglobin (HbO) and deoxyhemoglobin (HbR) in the cerebral cortex.
[0075] In the existing field of brain-computer interfaces, direct connection between the human brain and external devices is achieved by decoding and recognizing fNIRS signals. The aim is to convert human brain intentions into recognizable control signals in real time without relying on peripheral nerve and muscle pathways. This technology is widely used in motor function reconstruction, neurorehabilitation, assisted communication, and human-computer interaction technologies.
[0076] With the development of technology, the decoding of fNIRS signals has evolved from initial statistical analysis methods to linear model-based methods, and now to deep learning algorithms. Decoding fNIRS signals using deep learning algorithms can better reflect the dynamic changes in brain activity, thereby improving the real-time performance and accuracy of the decoding. However, existing deep learning methods for fNIRS signal decoding have the following problems:
[0077] (1) Feature extraction and application are not yet fully utilized;
[0078] (2) Although many existing methods have high decoding accuracy, the parameters are too redundant and it is difficult to achieve efficient decoding.
[0079] To address the aforementioned technical problems, this invention provides a functional near-infrared spectral signal decoding method. This method extracts features from subsequences of homologous signal segments to obtain the neurodynamic features and temporal features of each homologous signal segment, enabling deep learning algorithms to extract more complete features from homologous signal segments. By limiting the number and length of time windows, signal features are effectively preserved while avoiding redundancy in signal data, which would reduce computational efficiency. Through temporal feature extraction and high-level logical reasoning, the amount of input data in the inference stage is reduced, while avoiding redundant modeling. Using only a lightweight encoder allows focusing on core requirements while discarding redundant configurations.
[0080] The functional near-infrared spectral signal decoding method of the present invention will be described in detail below with reference to the accompanying drawings.
[0081] refer to Figure 1 and Figure 2 , Figure 1 A flowchart of a functional near-infrared spectral signal decoding method according to an embodiment of the present invention is shown. Figure 2 A model diagram of the functional near-infrared spectral signal decoding method according to an embodiment of the present invention is shown.
[0082] The method includes steps S110-S150.
[0083] S110, acquires functional near-infrared spectral signals.
[0084] like Figure 2 As shown, in the implementation process of this invention, step S110 and Figure 2 The fNIRS signal is obtained in step a. The following describes the acquisition method and equipment for the fNIRS signal.
[0085] refer to Figure 3 , Figure 3 The diagram shows the structure of the fNIRS signal acquisition device. In an embodiment of the present invention, the fNIRS signal acquisition device consists of 14 light sources and 16 detectors. Adjacent detectors and light sources form an fNIRS signal acquisition channel, for a total of 36 fNIRS signal acquisition channels. The 36 fNIRS signal acquisition channels acquire a total of 36 fNIRS signals.
[0086] S120 obtains the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration based on functional near-infrared spectral signals. The changes in oxyhemoglobin concentration are preprocessed to obtain a first signal segment sequence based on time, and the changes in deoxyhemoglobin concentration are preprocessed to obtain a second signal segment sequence based on time.
[0087] like Figure 2As shown, in the implementation process of this invention, step S120 and Figure 2 Step b in the data preprocessing section corresponds to the conversion, de-citrusing, and truncation of the functional near-infrared spectral signal. This effectively preserves the signal characteristics while avoiding redundancy in the signal data, thus reducing computational efficiency.
[0088] S130, the first signal segment sequence and the second signal segment sequence are paired to obtain a homologous signal segment sequence. Based on the homologous signal segment sequence, a homologous signal segment sub-sequence is determined. Based on the feature extraction of the homologous signal segment sub-sequence, the neurodynamic features and temporal features are obtained. The homologous signal segment sub-sequence is a sequence obtained based on the difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration.
[0089] like Figure 2 As shown, in the implementation process of this invention, step S130 and Figure 2 Steps c through f correspond to extracting the neurodynamic and temporal features of the homologous signal segment subsequences, which can effectively preserve the multifaceted features of the homologous signal segment subsequences.
[0090] It should be noted that the subsequence of the same source signal segment includes three signal segments extracted from the sequence of the same source signal segment based on a time window. The three signal segments include: a first signal segment, and a second and a third signal segment determined based on the first signal segment.
[0091] In embodiments of this application, neurodynamic features may include the maximum Lyapunov index and phase-locked value.
[0092] In the above steps, the expression for the maximum Lyapunov exponent is:
[0093]
[0094] Where LLE is the maximum Lyapunov exponent, K is the total number of points in the phase space, representing the state points of hemoglobin concentration change in the phase space, and k represents the time step. This represents the initial distance between the i-th phase point and its nearest neighbors. This represents the distance between the i-th phase point and its neighboring points after k steps.
[0095] In the above, the expression for the phase-locked value is:
[0096]
[0097] Wherein, PLV is the phase lock value, T is the time point index, t is the variable parameter, j is the imaginary unit, φ1(t)-φ2(t) represents the instantaneous phase difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration at time t, φ1(t) represents the instantaneous phase of the change in oxyhemoglobin concentration at time t, and φ2(t) represents the instantaneous phase of the change in deoxyhemoglobin concentration at time t.
[0098] S140, a fusion feature input sequence is constructed based on neurodynamic features and temporal features. The fusion feature input sequence is then input into the neural network model to obtain multimodal collaborative representation data. The multimodal collaborative representation data is used to represent the spatial correlation pattern between each signal segment.
[0099] like Figure 2 As shown, in the implementation of this invention, step S140 is equivalent to steps g and h. The neurodynamic characteristics can not only reflect the complex dynamic behavior of the nervous system, such as stability, chaotic characteristics and phase synchronization relationship between different neural signal sources, but also contain different information dimensions. For example, the maximum Lyapunov index (LLE) can reflect the divergence or convergence characteristics of the nervous system.
[0100] For time-series features, not only can the changing patterns of signals in the time dimension be captured, such as the trend of signal amplitude changes and frequency changes, but the frequency changes of signals in time-series features can also reflect the speed and rhythm of neural activity.
[0101] By fusing temporal and neurodynamic features, signals can be described simultaneously from both dynamic and temporal perspectives, avoiding the limitations of single-feature descriptions. Furthermore, this fusion creates a richer feature space, more accurately representing the essential characteristics of neural signals and providing a stronger information foundation for subsequent analysis. In addition, the fused features help reveal the complex interactions between neural signals in terms of time and dynamic features, leading to a deeper understanding of the physiological processes of the brain and other nervous systems. Moreover, by using a neural network model to learn features from the input sequence of the fused features and output multimodal collaborative representation data, not only can deep features be automatically learned, but multimodal collaborative representation data can also be effectively processed, improving the performance of spatial correlation pattern recognition.
[0102] S150 uses mean pooling based on multimodal collaborative representation data to obtain global representation data, and then identifies the physiological state recognition results from the global representation data.
[0103] like Figure 2 As shown, in the implementation process of this invention, step S150 and Figure 2The output of step I corresponds to the physiological state recognition result, which is used to enable the processor to control the external device to perform actions based on the physiological state recognition result.
[0104] It should be noted that multimodal collaborative representation data has high dimensionality and contains a large amount of detailed information about the spatial correlation patterns between signal segments. By performing mean pooling on the multimodal collaborative representation data, the high-dimensional multimodal collaborative representation data is reduced to a simpler global representation data. This not only retains the main information in the multimodal collaborative representation data, but also reduces the complexity of features. As a result, the input data of the model is reduced in the physiological state recognition result step, and the recognition efficiency of the model is improved.
[0105] The following describes specific embodiments. Figure 1 The steps of S120-S150 are further explained in detail.
[0106] refer to Figure 4 , Figure 4 A flowchart of step S120 of an embodiment of the present invention is shown. Step S120 includes:
[0107] S121, based on functional near-infrared spectral signals, obtains the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration.
[0108] It should be noted that the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration based on functional near-infrared spectral signals can be obtained using photoacoustic tomography, simulation and fitting methods based on Monte Carlo methods, time derivative distribution repair methods, modified Beer-Lambert law, and methods based on machine learning or deep learning.
[0109] Preferably, the change in hemoglobin concentration is obtained by converting the functional near-infrared spectral signal based on a modified Beer-Lambert law.
[0110] It should be noted that, compared to other methods, the modified Beer-Lambert law has the following advantages:
[0111] (1) The modified Beer-Lamber law is based on the original Beer-Lamber law and is modified by taking into account factors such as the light scattering effect in biological tissues. It is closely related to the physical process of functional near-infrared spectroscopy signal measurement and can intuitively reflect the quantitative relationship between light intensity attenuation and hemoglobin concentration change, which is easy to understand and interpret.
[0112] (2) Compared with some complex simulation methods or machine learning-based methods, the modified Beer-Lambert law formula is relatively simple and the calculation process is relatively easy. It does not require a lot of computing resources and a complex model training process. It can complete the calculation of hemoglobin concentration changes in a short time and is suitable for real-time or online application scenarios.
[0113] (3) By measuring at two different wavelengths, the changes in the concentrations of oxyhemoglobin and deoxyhemoglobin can be calculated separately, providing comprehensive information for studying brain function and hemodynamic changes, and helping to understand oxygen metabolism in physiological processes more deeply.
[0114] The modified Beer-Lambert law is expressed as follows:
[0115]
[0116] in, This represents the change in oxyhemoglobin concentration. This represents the change in deoxyhemoglobin concentration. The first wavelength of the light source in the acquisition device. Let λ be the second wavelength of the light source in the acquisition device, t be time, d be the differential path length factor, and l be the distance between the light source and the detector. Let be the change in optical density at the first wavelength at time t. Let be the change in optical density at the second wavelength at time t.
[0117] In this embodiment, since human tissue absorbs and scatters light relatively little in the near-infrared wavelength (700-900nm), near-infrared light can penetrate well into the cerebral cortex and be detected. Moreover, there is a significant difference in the absorption coefficients of oxyhemoglobin and deoxyhemoglobin near 760nm and 850nm. Therefore, the first wavelength is set to 760nm and the second wavelength is set to 850nm.
[0118] S122, based on filtering and baseline calibration of changes in oxyhemoglobin concentration; based on filtering and baseline calibration of changes in deoxyhemoglobin concentration.
[0119] It should be noted that filtering the changes in oxyhemoglobin and deoxyhemoglobin concentrations can effectively remove noise interference, retain effective signal components, and improve the signal-to-noise ratio. Baseline calibration of these changes can eliminate baseline drift and provide an accurate reference level. Currently, methods for filtering these changes include finite impulse response filters, infinite impulse filters, wavelet transform filters, moving average filters, and third-order Butterworth filters.
[0120] In this embodiment, a third-order Butterworth filter is used to perform bandpass filtering of 0.01 to 0.1 Hz on the changes in oxyhemoglobin concentration and the changes in deoxyhemoglobin concentration.
[0121] It should be noted that, compared with other filters, the advantages of the third-order Butterworth filter are its flat passband, good stopband attenuation characteristics, simple design, good linear phase characteristics, and high stability and reliability. Using a third-order Butterworth filter to perform bandpass filtering of 0.01 to 0.1 Hz on changes in hemoglobin concentration can not only effectively remove noise interference, but also preserve the slow cortical blood oxygen dynamics response.
[0122] By performing baseline calibration on changes in oxyhemoglobin and deoxyhemoglobin concentrations, the effects of baseline drift can be effectively eliminated, highlighting changes in blood oxygen response.
[0123] S123, a sliding window based on time length is preset, and the change in oxyhemoglobin concentration is truncated based on the sliding window to obtain the first single task whole signal for physiological state recognition; the change in deoxyhemoglobin concentration is truncated based on the sliding window to obtain the second single task whole signal for physiological state recognition.
[0124] In this embodiment, the sliding window duration is set to 10 seconds, and the movement time step is set to 1 second. By setting the sliding window size to 10 seconds, the feature signals used for physiological state recognition can be effectively captured. This design avoids data distortion caused by excessively short feature signals, while also effectively preventing data redundancy caused by excessively long feature signals. Setting the movement time step to 1 second ensures that the sliding window moves with sufficient intervals along the signal time axis, thereby precisely capturing signal changes. This parameter setting of the sliding window optimizes the feature signal capture process and provides a high-quality data foundation for accurate physiological state recognition.
[0125] S124, based on a preset time window, the entire signal of the first single task is truncated to obtain a first signal segment sequence; based on a time window, the entire signal of the second single task is truncated to obtain a second signal segment sequence.
[0126] In this embodiment, the time window size is 1 second. The entire signal of a single task can be divided into 10 equal-length signal segments by using a 1-second time window. This not only effectively balances the temporal and frequency resolution of the signal segments, but also effectively reduces data redundancy and noise impact, and improves the accuracy of feature extraction in subsequent steps.
[0127] refer to Figure 5 , Figure 5 A flowchart of step S130 of an embodiment of the present invention is shown. Step S130 includes:
[0128] S131, Pair the first signal segment sequence and the second signal segment sequence to obtain a sequence of signal segments from the same source.
[0129] It should be noted that the signal segments of oxyhemoglobin concentration change and deoxyhemoglobin concentration change acquired by each acquisition channel can be paired according to the wavelength of the light source of the acquisition device.
[0130] In the above steps, the expression for the sequence of signal segments from the same source is:
[0131]
[0132] in, For the i-th homologous signal segment, For the i-th signal segment in the first signal segment sequence, Let be the i-th signal segment in the second signal segment sequence, where i is a variable parameter and p is the number of channels of the fNIRS signal acquisition device.
[0133] S132, Determine the subsequence of the same source signal segment based on the sequence of the same source signal segment.
[0134] It should be noted that determining subsequences of the same source signal segment based on the sequence of the same source signal segment includes any one of the following:
[0135] Identify the first signal segment, which is the signal segment where the maximum difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration is located;
[0136] The two signal segments adjacent to and before the first signal segment are designated as the second and third signal segments; or, the second signal segment is separated from the first signal segment by one signal segment, and the third signal segment is adjacent to the first signal segment; or, the second signal segment is adjacent to the first signal segment, and the third signal segment is separated from the first signal segment by one signal segment.
[0137] S133, based on the feature extraction of subsequences of homologous signal segments, obtains neurodynamic features and temporal features.
[0138] refer to Figure 6 , Figure 6 A flowchart of time-series feature extraction according to an embodiment of the present invention is shown. The time-series feature extraction steps are as follows:
[0139] S133-1, based on time-series modeling of subsequences of signals from the same source, yields the predicted output.
[0140] It should be noted that when modeling time series based on subsequences of signals from the same source, Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), fNIRS-T model, fNIRSNet model, FCS-TPNet model, and Mamba model can be used.
[0141] Preferably, the predicted output is obtained by performing time series modeling based on the subsequences of the same source signal segments using the Mamba model.
[0142] In the above, the expression for the Mamba model is:
[0143]
[0144] Where t represents time, s t Let A represent the hidden state vector at time t, B be the state transition matrix, C be the input matrix, and D be the direct input mapping matrix. and Let y be the additive noise at time t. p,t This is the time sequence input of the p-th pair of homologous signal segments at time t.
[0145] refer to Figure 7 , Figure 7 A graph comparing the model performance is shown. For example... Figure 7 As shown, the performance of each model—including the publicly available datasets MA, WG, and UFFT—was obtained by using them as inputs to Long Short-Term Memory (LSTM), CNN, fNIRS-T, fNIRSNet, FCS-TPNet, and Mamba models, respectively. The results included accuracy, precision, recall, F1 score, and Kappa score. These five key metrics were used to quantitatively evaluate each model. The results show that the Mamba model exhibits superior performance across all metrics compared to the other models.
[0146] Compared to other methods, the Mamba model has the following advantages:
[0147] (1) Linear time complexity: Mamba is based on the State-Space Model (SSM) and adopts a selective state-space mechanism, which can efficiently process sequential data in different modes with linear time complexity. In contrast, the Transformer has a quadratic time complexity, which has high computational cost and resource consumption when processing long sequences; the gradient vanishing or exploding problem of LSTM also makes it difficult to effectively process long sequences.
[0148] (2) Long-distance dependency capture: Mamba can effectively capture long-distance dependencies, selectively propagate or discard information, and highlight tokens with high relevance. When processing long sequences, it can retain key information and remove redundancy.
[0149] (3) Combination of parallelism and sequential processing: Mamba can achieve parallel computation through convolution during the training phase, thereby improving the training speed; during the inference phase, it can use recursion to generate data sequentially while maintaining constant memory consumption. In contrast, the sequential computation characteristics of LSTM are difficult to parallelize, resulting in low computational efficiency.
[0150] (4) Selective State Space Layer: The core of Mamba is the selective state space layer, which can selectively propagate or suppress information based on the input at each step. This selective mechanism makes Mamba more flexible in processing different sequences and can dynamically adjust the focus according to the characteristics of the input data.
[0151] (5) Dynamic parameter adjustment: Mamba’s SSM parameters can be dynamically adjusted according to the input, so that it can better adapt to different input sequences and improve the generalization ability and adaptability of the model.
[0152] S133-2, perform time-dimensional mean aggregation on the predicted output results to obtain statistical representation data.
[0153] The statistical representation of the above data is expressed as follows:
[0154]
[0155] Among them, h p The statistical characterization data reflects the average coupling pattern of HbO and HbR over time, where T is the time length of each segment of the same source signal, t is a variable parameter, and y... p,t Referencing step S133-1.
[0156] S133-3, based on cross-pair projection of statistical characterization data, to obtain projected characterization data.
[0157] In the above steps, the two-dimensional statistical representation can be mapped to the model dimension through a linear projection layer to obtain the projected representation data.
[0158] In this embodiment, the model dimension is 16.
[0159] It should be noted that the expression for the projected representation data is:
[0160]
[0161] Among them, z p For projection representation data, W proj Let b be the projection weight matrix. proj Let d be the bias vector. model For model dimensions, h p Refer to step S133-2.
[0162] S133-4, based on the nonlinear transformation of the projection representation data, obtains the time series characteristics.
[0163] In the above steps, the nonlinear transformation uses the mean method, which obtains the time series characteristics by taking the mean of the projected representation data.
[0164] It should be noted that the expression for the time series characteristics is:
[0165]
[0166] Where, m p For timing characteristics, P is the number of channels in the fNIRS signal acquisition device, p is a variable parameter, and z p With d model All of these are referenced in step S133-3.
[0167] refer to Figure 8 Step S140 includes S141-S143.
[0168] S141, the neurodynamic features and temporal features are aligned in spatial dimension to obtain the neurodynamic projection features and temporal projection features.
[0169] It should be noted that aligning neurodynamic features and temporal features to a single spatial dimension ensures that all features are on the same scale and within a given range. This not only effectively avoids processing complexity caused by feature differences but also allows for direct comparison and comprehensive analysis of the two types of features. This helps to discover correlations and differences between features, providing more valuable information for feature fusion and facilitating more accurate identification of physiological states in subsequent steps.
[0170] In the above steps, the characteristic expression of neurodynamic projection is:
[0171]
[0172] Among them, H proj As a neurodynamic projection feature, W h Let b be the projection matrix. h Here, H is the bias vector, W is the neurodynamic feature, and d is the window number. fusion To unify dimensions.
[0173] In this embodiment, since the spatial dimensions of both the neurodynamic features and the temporal features are close to 32, the unified dimension is set to 32.
[0174] The temporal projection feature expression is:
[0175]
[0176] Among them, Mproj For temporal projection features, W m Let b be the projection matrix. m Let m be the bias vector. p Refer to step S133-4.
[0177] S142, the neurodynamic projection features and temporal projection features are concatenated to obtain the fused feature input sequence.
[0178] It should be noted that concatenating the neurodynamic projection features and the temporal projection features to obtain the fused feature input sequence not only covers all the information contained in the original two features, avoiding information omissions due to the use of a single feature, but also helps to discover potential nonlinear relationships between features, generate new features with discriminative power, and improve the expressive power of the neural network in subsequent steps.
[0179] In the above steps, the expression for the fused feature input sequence is:
[0180]
[0181] Among them, Z cat Represented as the fused feature input sequence, M proj H proj 、W、d fusion All of these are referenced in step S141.
[0182] S143, input the data corresponding to the fused feature input sequence into the neural network model, and output multimodal collaborative representation data.
[0183] It should be noted that in the above steps, multimodal collaborative representation data can be output through Long Short-Term Memory Network (LSTM), CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), and Transformer model.
[0184] Preferably, the fused feature input sequence is input into the neural network model to obtain multimodal collaborative representation data.
[0185] It should be noted that the Transformer model has advantages over other models in the following ways:
[0186] (1) The self-attention mechanism of Transformer can dynamically weight elements at different positions in the fused feature input sequence, and can capture global dependencies. This mechanism enables the model to give more weight to data that are farther away in the fused feature input sequence to mimic the behavior of LSTM, and to give more weight to data that are closer in topology to replicate the behavior of 1D-CNN, thereby capturing the spatial correlation patterns between the fused feature input sequences more flexibly.
[0187] (2) Transformer breaks through the loop structure limitations of RNN and LSTM, and can process data in parallel, thus speeding up the processing speed.
[0188] (3) When capturing spatial correlation patterns in the fusion feature input sequence, Transformer can learn the implicit distribution of the fusion feature input sequence and capture more detailed dependencies, thereby improving the accuracy of physiological state identification in subsequent steps.
[0189] In this embodiment, the Transformer model contains only a lightweight encoder, which can effectively extract multimodal representation data from the fused feature input sequence while focusing on core requirements and discarding redundant configurations.
[0190] refer to Figure 9 , Figure 9 A flowchart of the Transformer model data processing according to an embodiment of the present invention is shown. The Transformer model calculation steps are as follows:
[0191] S143-1, based on the data position encoding corresponding to the fusion feature input sequence, the encoded data is obtained.
[0192] In the above steps, the expression for the encoded data is:
[0193]
[0194] Among them, Z pos Represented as encoded data, POS(w) is the position code, Z cat Refer to step S142.
[0195] S143-2, normalize the encoded data, and capture diverse features from the normalized data to obtain diverse features;
[0196] In the above steps, the normalized data expression is:
[0197]
[0198] Among them, Z norml For normalized data, Z pos Refer to step S143-1.
[0199] It should be noted that in the above steps, a multi-head attention mechanism is used to capture diverse features from the normalized data, resulting in diverse features.
[0200]
[0201] in, For diverse characteristics, head hLet W0 be the output of the h-th head attention mechanism, and W0 be the learned weight matrix.
[0202] In this embodiment, since the length of the fused feature input sequence is divisible by 4, a 4-head attention mechanism is selected, which can not only capture the diverse features of the normalized data better, but also prevent data overfitting.
[0203] 143-3, based on diverse features and residual connections of encoded data, residual connection output data is obtained.
[0204] In the above steps, the residual connection output data expression is:
[0205]
[0206] Among them, Z att Z is the output data for the residual connection. pos Referencing step S143-1, Refer to step S143-2.
[0207] S143-4 normalizes the residual connection output data and performs feedforward propagation on the normalized residual connection output data to obtain complex relationship fitting data.
[0208] In the above steps, the expression for the normalized residual connection output data is:
[0209]
[0210] Among them, Z norm2 Z represents the normalized residual connection output data. att Refer to step S143-4.
[0211] The expression for fitting complex relationship data is as follows:
[0212]
[0213] Among them, Z ffn To fit data to complex relationships, W1 is the weight matrix and b2 is the bias parameter.
[0214] S143-5, based on the fitting data of complex relationships and the output data of residual connection, residual connection is performed to obtain multimodal collaborative characterization data.
[0215] In the above steps, the expression for the multimodal collaborative representation data is:
[0216]
[0217] Where Z represents the multimodal collaborative representation data, Z att Referencing step S143-3, Zffn Refer to step S143-4.
[0218] refer to Figure 10 , Figure 10 A flowchart of step S150 of an embodiment of the present invention is shown. Figure 10 As shown, step S150 further includes S151-S155.
[0219] S151, global representation data is obtained based on mean pooling of multimodal collaborative representation data.
[0220] In the above steps, the expression for the global representation data is:
[0221]
[0222] Among them, Z agg For global representation data, M is the number of time windows, m is the variable parameter, and Z is the global representation data. m For the multimodal collaborative representation data of the fused feature input sequence in the m-th time window, d fusion For the time window dimension.
[0223] S152, based on the nonlinear mapping of the global representation data, obtain the nonlinear mapping representation data.
[0224] In the above steps, the expression representing the nonlinear mapping data is:
[0225]
[0226] Where h1 is the nonlinear mapping representation data, Z agg Referring to step S151, W3 is the weight matrix, b3 is the bias vector, and d hid For the hidden layer dimension.
[0227] S153, regularization is performed on the nonlinear mapping representation data to obtain regularized representation data.
[0228] In the above steps, the expression for regularizing the data is:
[0229]
[0230] in, Represented as regularized representation data, h1 and d hid All of these are referenced in step S152.
[0231] S154, based on the regularized representation data linear transformation, obtains the linear transformation representation data.
[0232] In the above steps, the expression representing the data by the linear transformation is:
[0233]
[0234] Where h2 represents the linear transformation data. Referring to step S153, W4 is the weight matrix, and b4 is the bias vector. This represents the number of task categories.
[0235] S155, based on linear transformation to characterize data classification, obtains physiological state recognition results.
[0236]
[0237] in, For the physiological state recognition results, h2 and All of these are referenced in step S154.
[0238] It should be noted that the softmax function is chosen for linear transformation to represent classification primarily because the physiological state recognition results described above represent the probabilities of each task category. The softmax function can convert any real value in the model output into a probability distribution. It maps the score of each task category to the interval (0, 1), and the sum of the probabilities of all task categories is 1. This allows the output to be intuitively represented as the probability of each task category belonging to that task category, facilitating classification decisions.
[0239] Now for reference Figure 11 The diagram shows a block diagram of a device 1200 according to one embodiment of this application. Device 1200 may include one or more processors 1201 coupled to a controller hub 1203. In at least one embodiment, the controller hub 1203 communicates with the processor 1201 via a multi-branch bus such as a Front Side Bus (FSB), a point-to-point interface such as a Quick Path Interconnect (QPI), or a similar connection 1206. The processor 1201 executes instructions controlling general-type data processing operations. In one embodiment, the controller hub 1203 includes, but is not limited to, a Graphics Memory Controller Hub (GMCH) (not shown) and an Input / Output Hub (IOH) (which may be on a separate chip) (not shown), wherein the GMCH includes memory and a graphics controller and is coupled to the IOH.
[0240] Device 1200 may also include a coprocessor 1202 and a memory 1204 coupled to a controller hub 1203. Alternatively, one or both of the memory and the GMCH may be integrated within the processor (as described in this application), with memory 1204 and coprocessor 1202 directly coupled to processor 1201 and controller hub 1203, which resides on a single chip with the IOH. Memory 1204 may be, for example, Dynamic Random Access Memory (DRAM), Phase Change Memory (PCM), or a combination of both. In one embodiment, coprocessor 1202 is a dedicated processor, such as, for example, a high-throughput MIC (Many Integerated Core) processor, a network or communication processor, a compression engine, a graphics processor, a general-purpose graphics processor (GPGPU), or an embedded processor, etc. Optional properties of coprocessor 1202 are indicated by dashed lines. Figure 9 middle.
[0241] As a computer-readable storage medium, memory 1204 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. For example, memory 1204 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device such as one or more hard-disk drives (HDD(s)), one or more compact disc (CD) drives, and / or one or more digital versatile disc (DVD) drives.
[0242] In one embodiment, device 1200 may further include a network interface controller (NIC) 1206. Network interface 1206 may include a transceiver for providing a radio interface for device 1200 to communicate with any other suitable device, such as a front-end module, antenna, etc. In various embodiments, network interface 1206 may be integrated with other components of device 1200. Network interface 1206 can implement the functionality of the communication unit in the above embodiments.
[0243] Device 1200 may further include input / output (I / O) device 1205. I / O 1205 may include: a user interface designed to enable a user to interact with device 1200; a peripheral component interface designed to enable peripheral components to also interact with device 1200; and / or sensors designed to determine environmental conditions and / or location information related to device 1200.
[0244] It is worth noting that, Figure 9 This is merely an example. That is, although... Figure 9 The diagram shows that device 1200 includes multiple devices such as processor 1201, controller hub 1203, and memory 1204. However, in actual applications, devices using the methods of this application may include only a portion of the devices in device 1200. For example, it may include only processor 1201 and NIC 1206. Figure 9 The properties of the optional devices are shown in dashed lines. According to some embodiments of this application, the memory 1204, which is a computer-readable storage medium, stores instructions that, when executed on a computer, cause the system 1200 to perform a functional near-infrared spectral signal decoding method according to the above embodiments. Specific details can be found in the methods described in the above embodiments, and will not be repeated here.
[0245] Now for reference Figure 12 The diagram shown is a block diagram of a SoC (System on Chip) 1300 according to an embodiment of this application. Figure 12 In the diagram, similar components share the same reference numerals. Additionally, dashed boxes are an optional feature for more advanced SoCs. Figure 10 In this embodiment, SoC 1300 includes: an interconnect unit 1350 coupled to an application processor 1310; a system proxy unit 1380; a bus controller unit 1390; an integrated memory controller unit 1340; a group or one or more coprocessors 1320, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random access memory (SRAM) unit 1330; and a direct memory access (DMA) unit 1360. In one embodiment, the coprocessor 1320 includes a dedicated processor, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high-throughput MIC processor, or an embedded processor.
[0246] The static random access memory (SRAM) cell 1330 may include one or more computer-readable media for storing data and / or instructions. The computer-readable storage medium may store instructions, specifically, temporary and permanent copies of those instructions. These instructions may include, when executed by at least one unit in the processor, causing the SoC 1300 to perform a functional near-infrared spectral signal decoding method according to the above embodiments. Specific details can be found in the methods of the above embodiments and will not be repeated here.
[0247] Various embodiments of the mechanisms disclosed in this application can be implemented in hardware, software, firmware, or combinations of these implementation methods. Embodiments of this application can be implemented as computer programs or program code executable on a programmable system, the programmable system including at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0248] Program code can be applied to input instructions to execute the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, the processing system includes any system having a processor such as, for example, a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.
[0249] The program code can be implemented using a high-level procedural language or an object-oriented programming language to communicate with the processing system. Assembly language or machine language can also be used when needed. In fact, the mechanisms described in this application are not limited to any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0250] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored thereon on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine-readable (e.g., computer-readable) form, including but not limited to floppy disks, optical disks, CD-ROMs, compact disc read-only memory (CD-ROMs), magneto-optical disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic cards or optical cards, flash memory, or tangible machine-readable storage for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in the form of electrical, optical, acoustic, or other forms of propagated signals. Therefore, machine-readable media includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a machine-readable (e.g., computer-readable) form.
[0251] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order may not be necessary. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the accompanying drawings. Furthermore, including structural or methodological features in a particular figure does not imply that such features are required in all embodiments, and in some embodiments, these features may be omitted or may be combined with other features.
[0252] It should be noted that all units / modules mentioned in the device embodiments of this application are logical units / modules. Physically, a logical unit / module can be a physical unit / module, a part of a physical unit / module, or a combination of multiple physical units / modules. The physical implementation of these logical units / modules themselves is not the most important factor; the combination of functions implemented by these logical units / modules is the key to solving the technical problems proposed in this application. Furthermore, to highlight the innovative aspects of this application, the above-described device embodiments of this application have not introduced units / modules that are not closely related to solving the technical problems proposed in this application. This does not mean that the above-described device embodiments do not contain other units / modules.
[0253] It should be noted that in the examples and description of this patent, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0254] Although this application has been illustrated and described with reference to certain preferred embodiments thereof, those skilled in the art should understand that various changes in form and detail may be made thereto without departing from the spirit and scope of this application.
Claims
1. A method for decoding functional near-infrared spectral signals, characterized in that, include: Acquire functional near-infrared spectral signals; Based on the functional near-infrared spectral signal, the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration are obtained. The changes in oxyhemoglobin concentration are preprocessed to obtain a first signal segment sequence based on time, and the changes in deoxyhemoglobin concentration are preprocessed to obtain a second signal segment sequence based on time. Pair the first signal segment sequence and the second signal segment sequence to obtain a sequence of signal segments from the same source; Based on the homologous signal fragment sequence, a homologous signal fragment subsequence is determined, wherein the homologous signal fragment subsequence is a sequence obtained based on the difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration. The homologous signal segment subsequence includes three signal segments extracted from the homologous signal segment sequence based on a time window. The three signal segments include a first signal segment, and a second and a third signal segment determined based on the first signal segment. The step of determining the subsequence of the same source signal segment based on the sequence of the same source signal segments includes any one of the following: The first signal segment is determined as the signal segment containing the maximum difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration. The two signal segments adjacent to and before the first signal segment are designated as the second and third signal segments; or, the second signal segment is separated from the first signal segment by one signal segment, and the third signal segment is adjacent to the first signal segment; or, the second signal segment is adjacent to the first signal segment, and the third signal segment is separated from the first signal segment by one signal segment. Feature extraction is performed on the subsequences of the homologous signal segments to obtain neurodynamic features and temporal features; Based on the neurodynamic features and the temporal features, a fusion feature input sequence is constructed. The fusion feature input sequence is then input into a neural network model to obtain multimodal collaborative representation data. The multimodal collaborative representation data is used to represent the spatial correlation patterns between each signal segment. The multimodal collaborative representation data is averaged and pooled to obtain global representation data, and the global representation data is then used to identify physiological state recognition results.
2. The method according to claim 1, characterized in that, The method of obtaining the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration based on the functional near-infrared spectral signal includes: Based on the modified Beer-Lambert law, the functional near-infrared spectral signal is converted to obtain the changes in oxyhemoglobin concentration and the changes in deoxyhemoglobin concentration.
3. The method according to claim 1, characterized in that, The step of preprocessing the change in oxyhemoglobin concentration includes: Based on the filtering and baseline calibration of the changes in oxyhemoglobin concentration; A sliding window based on time length is preset, and the change in oxyhemoglobin concentration is truncated based on the sliding window to obtain the first single task whole signal for physiological state recognition; Based on a preset time window, the entire signal segment of the first single task is extracted to obtain the first signal segment sequence; The step of preprocessing the change in deoxyhemoglobin concentration includes: Based on the filtering and baseline calibration of the deoxyhemoglobin concentration change; Based on the sliding window, the change in deoxyhemoglobin concentration is truncated to obtain a second single task whole signal for physiological state recognition; Based on the time window, the entire signal segment of the second single task is extracted to obtain the second signal segment sequence.
4. The method according to claim 1, characterized in that, The neurodynamic characteristics include the maximum Lyapunov index and the phase-locked value.
5. The method according to claim 1, characterized in that, The extraction of temporal features based on the subsequence features of the homologous signal segments includes: Based on the time-series modeling of the aforementioned signal segments from the same source, the predicted output results are obtained; The predicted output results are aggregated by time-dimensional mean to obtain statistical representation data; Based on the cross-pair projection of the statistical characterization data, projected characterization data is obtained; The time-series features are obtained based on the nonlinear transformation of the projection representation data.
6. The method according to claim 1, characterized in that, The construction of the fusion feature input sequence based on the neurodynamic features and the temporal features includes: The neurodynamic features and the temporal features are aligned spatially to obtain neurodynamic projection features and temporal projection features; The neurodynamic projection features and the temporal projection features are concatenated to obtain the fused feature input sequence.
7. The method according to claim 1, characterized in that, The step of inputting the fused feature input sequence into a neural network model to obtain multimodal collaborative representation data includes: Encoded data is obtained based on the data position encoding corresponding to the fused feature input sequence; The encoded data is normalized, and the normalized data is then subjected to diversified feature capture to obtain diversified features; Based on the aforementioned diverse features and the residual concatenation of the encoded data, residual concatenation output data is obtained; The residual connection output data is normalized, and the normalized residual connection output data is then fed forward to obtain complex relationship fitting data. Based on the complex relationship fitting data and the residual connection output data residual connection, the multimodal collaborative characterization data is obtained.
8. The method according to claim 1, characterized in that, The process of identifying the physiological state from the global representation data includes: Based on the nonlinear mapping of the global representation data, nonlinear mapping representation data is obtained; Based on the regularization of the nonlinear mapping representation data, regularized representation data is obtained; Based on the linear transformation of the regularized characterization data, linear transformation characterization data is obtained; The physiological state identification result is obtained based on the classification of the linear transformation characterization data.
9. A functional near-infrared spectral signal decoding device, characterized in that, include: The acquisition module is used to acquire functional near-infrared spectral signals; The preprocessing module is used to obtain the changes in oxyhemoglobin concentration and deoxyhemoglobin concentration based on the functional near-infrared spectral signal, preprocess the changes in oxyhemoglobin concentration to obtain a first signal segment sequence based on time, and preprocess the changes in deoxyhemoglobin concentration to obtain a second signal segment sequence based on time. The preprocessing module is further configured to pair the first signal fragment sequence and the second signal fragment sequence to obtain a homologous signal fragment sequence, and determine a homologous signal fragment sub-sequence based on the homologous signal fragment sequence, wherein the homologous signal fragment sub-sequence is a sequence obtained based on the difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration; The homologous signal segment subsequence includes three signal segments extracted from the homologous signal segment sequence based on a time window. The three signal segments include a first signal segment, and a second and a third signal segment determined based on the first signal segment. The step of determining the subsequence of the same source signal segment based on the sequence of the same source signal segments includes any one of the following: The first signal segment is determined as the signal segment containing the maximum difference between the change in oxyhemoglobin concentration and the change in deoxyhemoglobin concentration. The two signal segments adjacent to and before the first signal segment are designated as the second and third signal segments; or, the second signal segment is separated from the first signal segment by one signal segment, and the third signal segment is adjacent to the first signal segment; or, the second signal segment is adjacent to the first signal segment, and the third signal segment is separated from the first signal segment by one signal segment. The feature extraction module is used to extract features based on the subsequence of the homologous signal segments to obtain neurodynamic features and temporal features; The fusion module is used to construct a fusion feature input sequence based on the neurodynamic features and the temporal features, and input the fusion feature input sequence into the neural network model to obtain multimodal collaborative representation data. The multimodal collaborative representation data is used to represent the spatial correlation pattern between each signal segment. The analysis module is used to perform mean pooling on the multimodal collaborative representation data to obtain global representation data, and to identify the physiological state recognition result from the global representation data.
10. An electronic device, characterized in that, include: processor; and a memory, in which computer program instructions are stored. Wherein, when the computer program instructions are executed by the processor, the processor causes the processor to perform the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the method according to any one of claims 1-8.