Auditory attention decoding method and system based on lightweight space-time network

By using a lightweight spatiotemporal neural network model and phase synchronization analysis, the problems of high computational load and insufficient decoding robustness in existing technologies are solved, enabling low-power real-time auditory attention decoding in devices such as hearing aids, and improving the accuracy and stability of decoding.

CN121901706APending Publication Date: 2026-04-21ANHUI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI UNIV
Filing Date
2026-01-07
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing auditory attention decoding models have high computational loads, making it difficult to meet the power consumption and real-time response requirements of embedded devices such as hearing aids. Furthermore, they fail to fully utilize the dynamic connectivity information of brain networks, resulting in insufficient decoding robustness.

Method used

A lightweight spatiotemporal neural network model is used, combined with phase synchronization analysis of multi-channel EEG signals, to generate a functional connectivity topology graph sequence, extract the temporal stability features of graph attributes, and improve the accuracy and robustness of the decoding results through a calibration mechanism.

Benefits of technology

It enables low-power real-time processing on resource-constrained devices, enhances the system's adaptability to individual differences and complex scene changes, and improves the accuracy and overall robustness of auditory attention decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901706A_ABST
    Figure CN121901706A_ABST
Patent Text Reader

Abstract

The invention discloses an auditory attention decoding method and system based on a lightweight space-time network, and relates to the technical field of electroencephalogram signals. The method comprises the following steps: collecting a multi-channel electroencephalogram signal when a user listens to an auditory scene containing at least two competitive sound sources; based on the lightweight space-time neural network model, outputting initial attention distribution probability distribution; calculating phase synchronization indexes among channels in a sliding window mode based on the electroencephalogram signals, generating a function connection topological graph sequence, and extracting graph attribute time sequence characteristics reflecting the dynamic stability of the network from the function connection topological graph sequence; and performing adaptive numerical calibration on the initial probability distribution to obtain calibrated attention distribution probability distribution, and finally determining a target sound source concerned by the user. According to the method, the problem of edge equipment deployment is solved through lightweight model design, the decoding precision and robustness are improved by fusing brain network dynamic information, and an effective technical scheme is provided for intelligent auditory enhancement of intelligent auditory auxiliary equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electroencephalogram (EEG) signal technology, specifically to an auditory attention decoding method and system based on lightweight spatiotemporal networks. Background Technology

[0002] The fundamental goal of auditory attention decoding technology is to accurately identify the direction of an individual's attention in a complex acoustic environment by analyzing the user's neurophysiological signals; that is, to determine the specific target the user is actually listening to among multiple competing sound sources presented simultaneously. This technology has significant value for the evolution of intelligent hearing aids, cochlear implants, and other hearing assistive devices, as well as advanced human-computer interaction systems.

[0003] However, existing technologies face two main challenges in practical applications. On the one hand, while deep learning-based decoding models can achieve high recognition accuracy, their large number of parameters and high computational complexity make it difficult to meet the constraints of power consumption, size, and real-time response capabilities in embedded devices such as hearing aids. On the other hand, conventional decoding methods mostly focus on the spatiotemporal characteristics of the EEG signals themselves, failing to fully explore the dynamic reorganization characteristics of functional connectivity networks between different cortical regions during selective auditory attention tasks. This directly restricts the robustness and final recognition accuracy of the decoding system in the face of individual differences and scene changes. Summary of the Invention

[0004] This invention addresses the technical problems of high computational load in existing auditory attention decoding models and insufficient robustness due to the failure of the decoding process to fully utilize the dynamic connectivity information of brain networks. It provides an auditory attention decoding method and system based on lightweight spatiotemporal networks.

[0005] The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides an auditory attention decoding method based on a lightweight spatiotemporal network, comprising: Collect multichannel EEG signals from target users while listening to an auditory scene containing at least two competing sound sources; The multi-channel EEG signals are input into a pre-trained lightweight spatiotemporal neural network model for decoding, and the output is a preliminary attention allocation probability distribution. Based on the multi-channel EEG signals, the process is continuously advanced in a sliding window manner according to a preset step size and window length. The phase synchronization index between the EEG signals of each channel within each time window is calculated to generate a functional connectivity topology graph sequence. The graph attribute temporal stability features are then extracted from the functional connectivity topology graph sequence. The initial attention allocation probability distribution is numerically calibrated using the temporal stability features of the graph attributes to generate a calibrated attention allocation probability distribution. The target sound source that the target user pays attention to in the auditory scene is determined based on the attention allocation probability distribution.

[0006] Secondly, this invention provides an auditory attention decoding system based on a lightweight spatiotemporal network, comprising: The signal acquisition module is used to acquire multi-channel EEG signals of the target user when listening to an auditory scene containing at least two competing sound sources; The spatiotemporal decoding module is used to input the multi-channel EEG signals into a pre-trained lightweight spatiotemporal neural network model for decoding processing, and output a preliminary attention allocation probability distribution. The functional connectivity analysis module is used to continuously advance in a sliding window manner based on the multi-channel EEG signals according to a preset step size and window length, calculate the phase synchronization index between the EEG signals of each channel within each time window, generate a functional connectivity topology graph sequence, and extract graph attribute temporal stability features from the functional connectivity topology graph sequence. The probability calibration module is used to numerically calibrate the initial attention allocation probability distribution using the temporal stability features of the graph attributes, and generate a calibrated attention allocation probability distribution. An attention determination module is used to determine the target sound source that the target user is paying attention to in the auditory scene based on the attention allocation probability distribution.

[0007] The beneficial effects of this invention are: Compared to existing technologies, this invention first reduces the computational complexity and number of parameters by constructing and applying a lightweight spatiotemporal neural network model. This allows the high-performance auditory attention decoding algorithm to be deployed on resource-constrained edge computing devices such as hearing aids, achieving localized, low-power real-time processing. Secondly, it innovatively introduces EEG functional connectivity topology analysis based on phase synchronization, extracting graph attribute temporal features reflecting attention stability from a dynamic brain network perspective. Thirdly, it uses these features to adaptively calibrate the preliminary decoding results of the neural network model, effectively integrating the spatiotemporal patterns of the signal with the collaborative interaction information of brain regions. Finally, the calibration mechanism enhances the system's adaptability to individual differences and complex scene changes, thereby improving the accuracy and overall robustness of auditory attention decoding while ensuring the model's lightweight nature. Attached Figure Description

[0008] Figure 1 A flowchart illustrating the auditory attention decoding method based on a lightweight spatiotemporal network provided by this invention; Figure 2 This is a schematic diagram of the structure of the auditory attention decoding system based on a lightweight spatiotemporal network provided by the present invention.

[0009] In the attached diagram, the components represented by each number are as follows: Signal acquisition module 11, spatiotemporal decoding module 12, functional connection analysis module 13, probability calibration module 14, attention determination module 15. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0012] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.

[0013] Example 1, as Figure 1 As shown, this embodiment of the invention provides an auditory attention decoding method based on a lightweight spatiotemporal network, including: S10: Collect multichannel EEG signals from the target user while listening to an auditory scene containing at least two competing sound sources; First, multichannel EEG signals were collected from target users while they listened to an auditory scene containing at least two competing sound sources. Target users refer to individuals requiring auditory attention state analysis and decoding. Examples include hearing aid wearers with hearing impairments or ordinary individuals needing to communicate verbally in noisy environments. An auditory scene containing at least two competing sound sources is an acoustic environment presenting multiple sound stimuli simultaneously, where these stimuli overlap or compete in time, frequency, or semantic content, thus requiring the listener to actively allocate selective attention, such as a two-speaker environment where two speakers' voices are played simultaneously.

[0014] Specifically, multichannel EEG signals are collected from the target user while listening to an auditory scene containing at least two competing sound sources, including: Multiple EEG acquisition electrodes are worn on the target user's scalp at predetermined brain regions according to a pre-defined layout. In an auditory scene containing at least two competing sound sources, signals from each EEG acquisition electrode are continuously recorded at a preset sampling rate to obtain multi-channel EEG signals.

[0015] Specifically, the signal acquisition process comprises two ordered operational stages. The first stage is electrode placement, which requires precisely placing multiple EEG signal acquisition electrodes on the target user's scalp at predetermined brain regions according to a pre-defined layout. This pre-defined layout is based on general EEG electrode placement standards or a spatial arrangement designed for specific research and application needs. It involves pre-defining the coordinates and relative relationships of each EEG acquisition electrode to anatomical landmarks in the head. This layout typically covers pre-defined brain regions related to auditory and attentional processing to ensure effective capture of neurophysiological signals reflecting auditory and attentional activities. Furthermore, the pre-defined brain regions on the scalp refer to cortical areas closely related to auditory perception and selective attention, determined based on knowledge of brain functional anatomy, such as the superior temporal gyrus, transverse temporal gyrus, and frontoparietal network regions. These regions are set based on prior knowledge of brain activity areas using standard EEG lead systems or specific experimental paradigms.

[0016] The second stage is signal recording, which begins after the target user starts listening to a specific auditory scene. Specifically, this auditory scene contains at least two competing sound sources. During this process, each EEG acquisition electrode continuously and synchronously records the potential changes from its contact point on the scalp at a preset sampling rate. The signals from each EEG acquisition electrode are synchronously collected to form a multi-channel EEG signal. The preset sampling rate refers to the number of discrete sampling points per second of the analog potential signal for each electrode channel during digital EEG signal acquisition. It is set according to the Nyquist sampling theorem and the analysis of the main effective frequency components of the EEG signal. For example, to fully preserve the typical frequency band components related to auditory attention in the EEG signal, the preset sampling rate is usually set to no less than 250 Hz; common settings include 256 Hz, 512 Hz, or 1000 Hz.

[0017] The resulting multichannel EEG signals fully recorded the trajectory of electrical activity in multiple regions of the user's cerebral cortex over time during the presentation of competitive auditory stimuli.

[0018] S20: The multi-channel EEG signal is input into a pre-trained lightweight spatiotemporal neural network model for decoding and processing, and a preliminary attention allocation probability distribution is output. Secondly, the aforementioned multi-channel EEG signals are input into a pre-trained lightweight spatiotemporal neural network model for decoding. This pre-trained lightweight spatiotemporal neural network model is a deep learning network structure optimized through model compression. Its model parameter count and computational complexity are significantly lower than the original model. This lightweight spatiotemporal neural network model learns through training to map the complex nonlinear relationship between the spatiotemporal pattern features of multi-channel EEG signals and the user's auditory attention target. It can efficiently extract key features from the input multi-channel EEG signals and perform classification inference to output the user's preliminary attention allocation probability distribution for each competing sound source in the auditory scene.

[0019] Specifically, the construction process of the lightweight spatiotemporal neural network model includes: Multiple multi-channel EEG signal samples were collected from historical users under known attentional target sound source conditions to form an EEG signal training sample set; Obtain the target sound source label representing the user's attention corresponding to each of the above-mentioned EEG signal training samples, and form an attention label sample set; Based on the principle of depthwise separable convolution, an initial spatiotemporal neural network model containing spatial feature extraction module and temporal feature extraction module is constructed; The initial spatiotemporal neural network model is trained in a supervised manner using the EEG signal training sample set and the attention label sample set until the verification convergence is obtained, thus obtaining the basic spatiotemporal neural network model. Redundant neuron channels in the basic spatiotemporal neural network model are pruned, and the network parameters are quantized with low bit width to obtain a lightweight spatiotemporal neural network model.

[0020] First, a training dataset needs to be constructed. This is done by collecting multi-channel EEG signal samples from multiple historical users under clearly known target sound sources, creating an EEG signal training sample set. Simultaneously, each of these EEG signal training samples is labeled with a corresponding target sound source label that clearly represents the user's attention state, forming an attention label sample set. This provides a realistic basis for supervised learning.

[0021] Next, in the model building phase, based on the principle of depthwise separable convolution, an initial spatiotemporal neural network model is constructed. This model includes a spatial feature extraction module specifically for extracting spatial distribution patterns of EEG signals, and a temporal feature extraction module specifically for capturing the temporal dynamic changes of EEG signals. The principle of depthwise separable convolution is a lightweight network design concept that decomposes the standard convolution operation into two independent steps, aiming to reduce the computational complexity and number of parameters of the model. Specifically, this involves first performing independent spatial convolution on each channel of the input data to extract features, and then linearly combining all the feature channels extracted in the previous step through pointwise convolution. This ensures feature extraction capability while reducing the computational burden of the model.

[0022] Furthermore, the initial spatiotemporal neural network model is subjected to supervised training using the pre-constructed EEG signal training sample set and attention label sample set. This training process is continuously iterated until validation convergence is achieved, reaching a stable state. At this point, a basic spatiotemporal neural network model with good decoding performance is obtained. The convergence condition, for example, is that the improvement in attention decoding accuracy on the validation set is less than 0.5‰ for 10 consecutive training cycles, or the change in the validation set loss function value is lower than a preset minimum threshold, such as 0.0001, for 10 consecutive cycles, indicating that the training has reached convergence.

[0023] An exemplary spatiotemporal neural network model based on the principle of deep separable convolution aims to efficiently capture the complex spatiotemporal patterns contained in multi-channel EEG signals to map their nonlinear correlation with auditory attention. The specific model construction includes two core modules. Specifically, the spatial feature extraction module is primarily responsible for capturing the spatial distribution patterns of brain activity across different electrode channels at the same time point. This module can employ the spatial convolution component of deep separable convolution, where the convolution kernel spans 1 in the time dimension and operates independently on each input channel in the channel dimension, achieving spatial filtering and feature extraction of signals from different brain regions. The temporal feature extraction module is responsible for capturing the dynamic sequence patterns of EEG signals evolving over time on each electrode channel. This module can employ a one-dimensional temporal convolution or long short-term memory network structure, where the convolution kernel or recursive unit slides along the time axis to extract local and long-term dependency features of the signal in the time dimension. The output features of the two modules are integrated by a fusion layer and finally connected to a fully connected classification layer. This fully connected classification layer uses the Softmax activation function and outputs the preliminary attention allocation probability distribution corresponding to each competing sound source.

[0024] Furthermore, the aforementioned EEG signal training sample set is used as input, and the corresponding attention label sample set is used as supervision signal. During training, the overall data is divided into training, validation, and test sets according to a preset ratio, for example, a 7:2:1 ratio. Key training hyperparameters include a learning rate of 0.001, a total of 200 training epochs, and a batch size of 32. The learning rate is set to balance the speed and stability of parameter updates, the number of training epochs ensures that the model can fully learn the spatiotemporal attention patterns in the data, and the batch size balances training efficiency and gradient estimation stability.

[0025] In further training and optimization, the backpropagation algorithm and Adam optimizer are used to iteratively update the network weight parameters. The cross-entropy loss function, commonly used in classification tasks, is selected to measure the difference between the probability distribution of the model's output and the true attention labels. The training process is monitored by a validation set. When the attention decoding accuracy on the validation set no longer improves within 20 consecutive training epochs, or when the change in the validation set loss function value tends to stabilize (e.g., the change is less than 0.0001 over 20 consecutive epochs), the model training is considered to have reached convergence. At this point, a basic spatiotemporal neural network model with good generalization ability is obtained. This model can effectively learn and represent the complex mapping relationship between the spatiotemporal features of EEG signals and auditory attention states.

[0026] Since basic spatiotemporal neural network models typically have a large number of parameters and contain a certain amount of computational redundancy after training and convergence, and target deployment platforms such as hearing aids and cochlear implants have strict constraints on memory, power consumption and real-time performance, it is necessary to perform special compression and optimization on basic spatiotemporal neural network models to obtain lightweight spatiotemporal neural network models with significantly reduced number of parameters and greatly improved computational efficiency, making them suitable for deployment in resource-constrained embedded devices.

[0027] Specifically, redundant neuron channels in the basic spatiotemporal neural network model are pruned, and the network parameters are quantized with low bit width to obtain a lightweight spatiotemporal neural network model, including: Calculate the sum of the absolute values ​​of the channel weights of each convolutional layer in the basic spatiotemporal neural network model; Identify and remove neuron channels whose sum of absolute channel weights is lower than a preset threshold, and generate a neural network model with channel pruning. The parameters of the neural network model after channel pruning are represented by low-bit-width quantization and fine-tuned to obtain a compressed and lightweight spatiotemporal neural network model.

[0028] Specifically, the lightweighting process begins with structural pruning. This step first requires evaluating the importance of each channel within each convolutional layer of the underlying spatiotemporal neural network model. Specifically, by calculating the sum of the absolute values ​​of the channel weights in each convolutional layer, a quantitative metric can be obtained to measure the contribution of that channel to the overall network output.

[0029] Secondly, based on a preset threshold, neuronal channels whose sum of absolute weights is lower than the preset threshold are identified. These neuronal channels are deemed redundant or have limited contribution due to their low activation intensity, and are removed from the network structure, resulting in a more compact, channel-pruned neural network model. The preset threshold is a boundary value used to distinguish the importance of channels, characterizing the strictness of the model pruning. It is set according to the target compression ratio, model performance tolerance, and the constraints of computational resources in the specific application scenario. For example, the preset threshold can be set to 30% of the average sum of the absolute values ​​of all channel weights in the corresponding convolutional layer. This channel pruning operation directly reduces the number of convolutional kernels in the basic spatiotemporal neural network model and the number of input channels in subsequent neural network layers, effectively reducing the overall computational graph complexity.

[0030] Secondly, after structural pruning, the parameter representation of the channel-pruned neural network model is further compressed. Specifically, low-bit-width quantization is performed on all weight parameters retained in the channel-pruned neural network model, converting the network parameters originally stored in 32-bit floating-point format into a lower-bit-width integer or fixed-point format, such as 8-bit integers. This quantization operation significantly reduces the model's storage space requirements.

[0031] Meanwhile, to compensate for potential model accuracy loss due to channel pruning and parameter quantization, fine-tuning training is required after quantization. This fine-tuning training uses a portion of the original training data to iteratively update the quantized network parameters a limited number of times with a significantly reduced learning rate. This allows the model weights to readjust to the task objective under the constraints of the quantized representation, thereby restoring and stabilizing its decoding performance. Finally, through structured channel pruning, parameterized low-bit-width quantization, and targeted fine-tuning training, a compressed, lightweight spatiotemporal neural network model is obtained, which significantly reduces the number of model parameters and greatly improves computational and storage efficiency while maintaining the core decoding functionality.

[0032] S30: Based on the multi-channel EEG signals, the process is continuously advanced in a sliding window manner according to a preset step size and window length. The phase synchronization index between the EEG signals of each channel within each time window is calculated, a functional connectivity topology graph sequence is generated, and graph attribute temporal stability features are extracted from the functional connectivity topology graph sequence. Specifically, based on the multi-channel EEG signals, the process proceeds continuously in a sliding window manner according to a preset step size and window length, calculating the phase synchronization index between the EEG signals of each channel within each time window, and generating a functional connectivity topology sequence, including: Set the window length and sliding step size of the sliding time window; Starting from the beginning of the multi-channel EEG signal, the time window is slid sequentially according to the sliding step size, and the multi-channel EEG signal segment corresponding to each time window is extracted. For each time window, calculate the phase synchronization index between any two channels of EEG signal segments. Each EEG acquisition channel is defined as a network node, and the phase synchronization index between any two network nodes is defined as the weight of the connection edge. The functional connection topology graph of the current time window is constructed. Arrange the functional connection topology diagrams of all time windows according to the chronological order of the corresponding time windows, and combine them to form a sequence of functional connection topology diagrams.

[0033] Specifically, generating a sequence of functional connectivity topology maps is a systematic and dynamic brain network construction process designed to characterize the patterns of brain functional connectivity evolution over time. This sequence is an ordered set of weighted network maps arranged chronologically, where each network map represents the strength and pattern of functional connectivity between different brain regions within a corresponding time segment. This sequence can be used to analyze the dynamic reorganization characteristics and stability of brain functional networks under auditory attentional states.

[0034] First, the window length and sliding step size of the sliding time window need to be predefined. The sliding time window refers to a fixed time interval for continuously extracting data segments along the time axis of the multi-channel EEG signal for independent analysis. Its window length determines the time range of the EEG signal covered in each analysis, while the sliding step size controls the degree of overlap or spacing between adjacent analysis windows on the time axis. Specifically, the window length and sliding step size are set based on the typical timescale of neural oscillations related to auditory attention in the EEG signal and the required temporal resolution. For example, the window length can be set to 2 seconds, and the sliding step size can be set to 0.1 seconds to achieve a balance between temporal continuity and analytical granularity.

[0035] Secondly, after setting the parameters, starting from the beginning of the multi-channel EEG signal, the time window is slid forward along the time axis according to the set sliding step size. After each slide, the EEG signals of all channels within the time period covered by the current time window are captured to obtain a multi-channel EEG signal segment corresponding to the current time window.

[0036] Furthermore, phase synchronization analysis is performed on each extracted multichannel EEG signal segment. Specifically, phase synchronization analysis is a computational method to assess the degree of phase coordination of neural oscillations in different brain regions. By calculating the phase synchronization index between any two different channels of the EEG signal within the multichannel EEG signal segment, the degree of phase coordination between the neural oscillations of the two channels is quantified.

[0037] Specifically, the phase synchronization index between any two channels of EEG signals is calculated, including: The instantaneous phase sequence of the EEG signals from the two channels within the time window is extracted respectively; Calculate the instantaneous phase difference between the two channels at each moment within the time window to obtain the instantaneous phase difference sequence; The modulus of the complex exponential average of the phase difference sequence is calculated as an index of phase synchronization between the EEG signals of any two channels.

[0038] First, the instantaneous phase sequences of the EEG signals from both channels within the current analysis time window need to be extracted separately. For example, methods such as Hilbert transform or wavelet transform can be used to convert the time-domain signal of each channel into a sequence of instantaneous phase values ​​that vary over time, thereby stripping away the amplitude information of the signal and focusing on its periodic phase dynamics.

[0039] Secondly, based on the two extracted instantaneous phase sequences, the difference between the two phase values ​​at each moment within the same time window is calculated, thus obtaining an instantaneous phase difference sequence that varies with time. This instantaneous phase difference sequence directly reflects the real-time dynamic evolution of the relative phase relationship between the neural oscillatory activities of the two brain regions, and can be used to accurately assess the instantaneous intensity and stability of the functional coupling between the two brain regions.

[0040] Finally, the obtained instantaneous phase difference sequence is comprehensively quantized. Specifically, the phase difference value at each moment in the instantaneous phase difference sequence is converted into a unit vector on the complex plane, i.e., its complex exponent value is calculated. Next, the average value of the complex exponent value over the entire time window is calculated. The magnitude of this average complex exponent value is defined as the phase synchronization index between the two EEG signals. The value of this phase synchronization index ranges from 0 to 1, where 0 indicates that the phases of the two signals are completely independent and randomly distributed, and 1 indicates that the phases of the two signals maintain a completely consistent phase relationship throughout the entire time window. A higher value indicates a stronger functional connection between the two brain regions in terms of phase synchronization during that time period.

[0041] Furthermore, after calculating the indicators, network modeling is performed. Each EEG acquisition channel is abstracted as a network node, and the total number of nodes equals the number of EEG acquisition channels. For any two different network nodes, the calculated phase synchronization index between them is defined as the weight of the connection edge connecting the two nodes. In this way, a complete functional connectivity topology can be constructed for each time window. This functional connectivity topology, in the form of a weighted network, intuitively represents the strength pattern of functional connectivity between different brain regions within that time window.

[0042] Finally, each functional connectivity topology map generated in chronological order is arranged and combined. All functional connectivity topologies corresponding to all time windows are arranged and combined into an ordered sequence according to the chronological order of the original EEG signals they represent, thus forming a functional connectivity topology map sequence. This sequence dynamically records the continuous changes in the brain's functional connectivity network over time throughout the entire analysis period.

[0043] Further, from the functional connection topology graph sequence, the temporal stability features of graph attributes are extracted, including: Randomly select one functional connection topology graph from the sequence of functional connection topology graphs as the first functional connection topology graph; For each connecting edge in the first functional connection topology graph, calculate the reciprocal of the connecting edge weight as the distance; For each pair of non-repeating network nodes in the first functional connection topology graph, calculate the sum of the calculated distances of all possible paths between each pair of non-repeating network nodes, and select the minimum value from all the sums of calculated distances as the shortest path length between each pair of non-repeating network nodes. Calculate the reciprocal of the shortest path length as the transmission efficiency value between each pair of non-repeating network nodes in the first functional connection topology graph; Calculate the arithmetic mean of the transmission efficiency values ​​of all non-repeating network node pairs in the first functional connection topology graph, and use it as the global efficiency attribute value of the first functional connection topology graph. Traverse the sequence of functional connection topology graphs and calculate the global efficiency attribute value corresponding to each functional connection topology graph in the sequence; The global efficiency attribute values ​​are arranged according to the chronological order of the functional connection topology diagram to obtain a sequence of global efficiency attribute values. The variance of the global efficiency attribute value sequence is calculated and used as a temporal stability feature of the graph attribute.

[0044] Specifically, extracting graph attribute temporal stability features from functional connectivity topology sequences is an analytical process aimed at quantifying the dynamic characteristics of brain functional networks. These graph attribute temporal stability features characterize the stability of global information transmission efficiency fluctuations in functional connectivity networks over time; smaller fluctuations indicate a more temporally stable network state.

[0045] First, a functional connectivity topology graph is randomly selected from the sequence of functional connectivity topologies as the starting point for analysis, referred to as the first functional connectivity topology graph. Network attributes are calculated for this first functional connectivity topology graph. Specifically, a distance metric between network nodes is defined. For each connection edge in the first functional connectivity topology graph, the reciprocal of the edge weight is calculated as the calculated distance. This calculated distance characterizes the difficulty of information transmission corresponding to the strength of functional connectivity between two brain regions, transforming the ease of transmission represented by high synchronization, i.e., high connection strength, into a short-distance concept in the network graph.

[0046] Furthermore, for each pair of unique network nodes in the first functional connection topology graph (i.e., all possible node combinations), the sum of the calculated distances of all possible paths between each pair of unique network nodes is calculated, and the minimum value is selected from all the summations. This minimum value is defined as the shortest path length between that pair of nodes. This shortest path length reflects the theoretically optimal efficiency of information transmission between each pair of unique network nodes in the network.

[0047] Furthermore, the reciprocal of the shortest path length is calculated and defined as the transmission efficiency value between the pair of nodes. Specifically, a higher transmission efficiency value indicates a higher theoretical efficiency in information transmission between the pair of nodes. Next, the arithmetic mean of the transmission efficiency values ​​of all non-repeating node pairs in the first functional connection topology is calculated. This average value is the global efficiency attribute value of the functional connection topology. The global efficiency attribute value is a comprehensive indicator describing the overall information transmission capability of the entire network. A higher global efficiency attribute value indicates a stronger overall integration and information flow capability of the network, suggesting a more efficient and closely coordinated working mode of the brain-like functional connection network within the corresponding time window.

[0048] Next, the functional connection topology sequence is traversed, and the above steps are repeated for each functional connection topology in the sequence to calculate the global efficiency attribute value for each functional connection topology. The global efficiency attribute values ​​are then arranged into a time series according to the original chronological order of the functional connection topologies, thus obtaining the global efficiency attribute value sequence.

[0049] Finally, statistical analysis is performed on the global efficiency attribute value sequence to calculate its variance. The calculated variance is defined as the graph attribute temporal stability feature. This graph attribute temporal stability feature characterizes the stability of the global information transmission efficiency of the functionally connected network over time. Specifically, the magnitude of the variance directly quantifies the amplitude of global efficiency fluctuations over time; the smaller the variance, the more stable the network's information transmission efficiency is in the time dimension, and vice versa.

[0050] S40: Numerically calibrate the preliminary attention allocation probability distribution using the temporal stability features of the graph attributes to generate a calibrated attention allocation probability distribution; Specifically, the initial attention allocation probability distribution is numerically calibrated using the temporal stability features of the graph attributes to generate a calibrated attention allocation probability distribution, including: Obtain the historical graph attribute temporal stability feature sequence generated by the lightweight spatiotemporal neural network model in a historical verification process similar to the current auditory scene, and calculate the preset high-order statistical quantile of the historical graph attribute temporal stability feature sequence as a stability discrimination threshold; Calculate the ratio of the stability discrimination threshold to the temporal stability feature of the graph attribute, and use the natural logarithm of the ratio as the calibration intensity coefficient value; Based on the calibration intensity coefficient value, the initial attention allocation probability distribution is numerically transformed to generate a numerically adjusted intermediate probability distribution; The intermediate probability distribution is normalized to obtain the calibrated attention allocation probability distribution.

[0051] Specifically, using the temporal stability features of graph attributes to numerically calibrate the probability distribution of the initial attention allocation is a process of adjusting the decoding confidence based on the dynamic stability of the brain's functional connectivity network, aiming to improve the robustness and accuracy of the decoding results.

[0052] Specifically, the first step is to obtain the graph attribute temporal stability feature sequence generated during the historical validation process of the lightweight spatiotemporal neural network model. This historical data comes from experiments or applications with similar acoustic characteristics to the current auditory scenario. Statistical analysis is then performed on this historical feature sequence to calculate its preset higher-order statistical quantiles, such as the 75th or 90th quantile, as a stability threshold. Specifically, this preset higher-order statistical quantile is a statistic used to divide the data distribution, representing the numerical level of a specific proportion of data points in the historical feature sequence. This higher-order statistical quantile can be dynamically determined based on the historical data distribution and the stability requirements of the actual application, representing a typical high-level boundary of the network's stability characteristics under good historical conditions.

[0053] Then, the obtained stability threshold is compared with the currently calculated real-time temporal stability features of graph attributes, and the ratio between the two is calculated. The natural logarithm of this ratio is taken, and the result is defined as the calibration strength coefficient value. The sign and magnitude of this calibration strength coefficient value directly reflect the stability of the current brain functional connectivity network compared to the historical baseline. Specifically, a positive calibration strength coefficient value indicates that the current network stability is better than the historical baseline, while a negative calibration strength coefficient value indicates that the current stability is insufficient.

[0054] Furthermore, based on the sign and magnitude of the calibration intensity coefficient values, a corresponding numerical transformation is performed on the initial attention allocation probability distribution output by the lightweight spatiotemporal neural network model. This transformation aims to adjust the distribution shape of the decoding probability according to network stability, thereby generating a numerically adjusted intermediate probability distribution.

[0055] Specifically, based on the calibration intensity coefficient value, the initial attention allocation probability distribution is numerically transformed to generate a numerically adjusted intermediate probability distribution, including: If the calibration intensity coefficient value is greater than zero, the value is transformed according to a first preset scheme, wherein the first preset scheme includes: The sound source with the highest probability value is selected from the preliminary attention allocation probability distribution and used as the dominant sound source. Keep the probability values ​​of all non-dominant sound sources in the initial attention allocation probability distribution unchanged; Calculate the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient as the exponent, and use it as the probability adjustment coefficient of the dominant sound source; The probability value of the dominant sound source in the initial attention allocation probability distribution is multiplied by the probability adjustment coefficient to obtain the adjusted probability value of the dominant sound source. The probability values ​​of all non-dominant sound sources are integrated with the adjusted probability values ​​of the dominant sound sources to form an intermediate probability distribution; If the calibration intensity coefficient value is less than or equal to zero, the numerical transformation is performed according to the second preset scheme, wherein the second preset scheme includes: Calculate the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient as the exponent, and use it as the first weight; Calculate the difference between the value 1 and the first weight, and use it as the second weight; Construct a uniform probability distribution with the same number of sound sources as the initial attention allocation probability distribution, wherein the probability value of each sound source in the uniform probability distribution is equal; The probability value of each sound source in the initial attention allocation probability distribution is multiplied by the first weight, and the probability value of the corresponding sound source in the uniform probability distribution is multiplied by the second weight. The two sets of product results are then added together to form an intermediate probability distribution.

[0056] Specifically, the numerical transformation of the initial attention allocation probability distribution based on the calibration intensity coefficient value employs two different preset schemes, the selection of which depends on the sign of the calibration intensity coefficient value. The core objective is to dynamically adjust the decoding confidence based on the stability of the brain's functional connectivity network.

[0057] First, when the calibration intensity coefficient is greater than zero, it indicates that the temporal stability of the current brain functional connectivity network is better than the historical baseline, and the network state is reliable. At this point, a first preset scheme is used for numerical transformation. Specifically, this first preset scheme first identifies the sound source with the highest probability value from the initial attention allocation probability distribution and determines it as the dominant sound source. Second, the probability values ​​of all non-dominant sound sources remain unchanged. An exponential function value with the natural constant e as the base and the calibration intensity coefficient as the exponent is calculated, and this calculation result is used as the probability adjustment coefficient for the dominant sound source. The original probability value of the dominant sound source in the initial attention allocation probability distribution is multiplied by this probability adjustment coefficient to obtain the enhanced adjusted probability value of the dominant sound source.

[0058] Finally, the unchanged probability values ​​of non-dominant sound sources are integrated with the enhanced adjusted probability values ​​of dominant sound sources to form an intermediate probability distribution after numerical adjustment. This first preset scheme aims to further enhance the model's confidence in judging the most likely target sound source when the network state is stable, thereby improving the recognition accuracy and decision certainty of the decoding system under ideal neural state.

[0059] Furthermore, when the calibration intensity coefficient value is less than or equal to zero, it indicates that the current network stability has not reached or is lower than the historical benchmark, and the network state may be fluctuating or unreliable. In this case, a second preset scheme is used for numerical transformation. This second preset scheme first calculates the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient value as the exponent, and uses this as the first weight. The difference between the value 1 and the first weight is calculated to obtain the second weight. Simultaneously, a uniform probability distribution with the same number of sound sources as the initial attention allocation probability distribution is constructed, in which each sound source is assigned an equal probability value.

[0060] Then, a weighted fusion is performed, multiplying the probability value of each sound source in the initial attention allocation probability distribution by the first weight, and simultaneously multiplying the probability value of the corresponding sound source in the uniform probability distribution by the second weight. Finally, the two sets of product results for each sound source are added together to form a new intermediate probability distribution. This second pre-set scheme aims to smooth the probability distribution by mixing the original decoding result with a completely uncertain uniform distribution when the network state is unstable, thereby reducing over-reliance on a single dominant sound source, improving the robustness of decision-making, and preventing overly certain erroneous judgments when confidence is insufficient.

[0061] Finally, the obtained intermediate probability distribution is normalized to ensure that the sum of the probabilities of all competing sound sources is 1. The probability distribution obtained after normalization is the final attention allocation probability distribution calibrated with graph attribute temporal stability features.

[0062] Specifically, this calibration process organically integrates the dynamic stability information of the brain network into the decoding decision, so that the final output attention allocation probability distribution not only depends on the spatiotemporal characteristics of the signal, but also takes into account the robustness of the brain functional state, thereby improving the accuracy and reliability of the auditory attention decoding system in the face of complex scenarios and individual differences.

[0063] S50: Determine the target sound source that the target user is paying attention to in the auditory scene based on the attention allocation probability distribution.

[0064] Specifically, from the aforementioned calibrated attention allocation probability distribution, the sound source with the highest probability value is selected, and this sound source is determined to be the target sound source that the user is currently focusing on. This target sound source is the final output of the auditory attention decoding system.

[0065] In summary, the embodiments of this application have at least the following technical effects: Compared to existing technologies, this application first constructs a lightweight spatiotemporal network using depthwise separable convolution and model pruning quantization techniques, effectively solving the problem of high computational load and large parameter count in traditional decoding models, making real-time local deployment difficult in resource-constrained edge devices such as hearing aids. Secondly, it innovatively introduces dynamic functional connectivity topology analysis based on phase synchronization, extracting graph attribute temporal features reflecting attentional stability from a brain network perspective, thus overcoming the shortcomings of existing methods that rely solely on spatiotemporal features while ignoring brain region collaboration information. Thirdly, it designs an adaptive probability calibration mechanism based on network stability features, which dynamically adjusts the confidence distribution of the decoding output according to the reliability of the current neural state, thereby improving the system's robustness and final recognition accuracy under complex scenarios and individual differences.

[0066] In summary, the technical solution provided in this application achieves an effective balance between high-performance decoding algorithms and low-power embedded platform deployment requirements, providing core technical support for the practical development of next-generation intelligent hearing assistive devices.

[0067] Example 2, as Figure 2 As shown, based on the same inventive concept as the auditory attention decoding method based on lightweight spatiotemporal networks provided in Embodiment 1, this embodiment of the invention also provides an auditory attention decoding system based on lightweight spatiotemporal networks, including: The signal acquisition module 11 is used to acquire multi-channel EEG signals of the target user when listening to an auditory scene containing at least two competing sound sources; The spatiotemporal decoding module 12 is used to input the multi-channel EEG signal into a pre-trained lightweight spatiotemporal neural network model for decoding processing and output a preliminary attention allocation probability distribution. The functional connectivity analysis module 13 is used to continuously advance in a sliding window manner based on the multi-channel EEG signals according to a preset step size and window length, calculate the phase synchronization index between the EEG signals of each channel in each time window, generate a functional connectivity topology graph sequence, and extract graph attribute temporal stability features from the functional connectivity topology graph sequence. The probability calibration module 14 is used to perform numerical calibration on the preliminary attention allocation probability distribution using the temporal stability characteristics of the graph attributes, and generate a calibrated attention allocation probability distribution. The attention determination module 15 is used to determine the target sound source that the target user is paying attention to in the auditory scene based on the attention allocation probability distribution.

[0068] The signal acquisition module 11 is specifically used for: Specifically, multichannel EEG signals are collected from the target user while listening to an auditory scene containing at least two competing sound sources, including: Multiple EEG acquisition electrodes are worn on the target user's scalp at predetermined brain regions according to a pre-defined layout. In an auditory scene containing at least two competing sound sources, signals from each EEG acquisition electrode are continuously recorded at a preset sampling rate to obtain multi-channel EEG signals.

[0069] The spatiotemporal decoding module 12 is specifically used for: Specifically, the construction process of the lightweight spatiotemporal neural network model includes: Multiple multi-channel EEG signal samples were collected from historical users under known attentional target sound source conditions to form an EEG signal training sample set; Obtain the target sound source label representing the user's attention corresponding to each of the above-mentioned EEG signal training samples, and form an attention label sample set; Based on the principle of depthwise separable convolution, an initial spatiotemporal neural network model containing spatial feature extraction module and temporal feature extraction module is constructed; The initial spatiotemporal neural network model is trained in a supervised manner using the EEG signal training sample set and the attention label sample set until the verification convergence is obtained, thus obtaining the basic spatiotemporal neural network model. Redundant neuron channels in the basic spatiotemporal neural network model are pruned, and the network parameters are quantized with low bit width to obtain a lightweight spatiotemporal neural network model.

[0070] Specifically, redundant neuron channels in the basic spatiotemporal neural network model are pruned, and the network parameters are quantized with low bit width to obtain a lightweight spatiotemporal neural network model, including: Calculate the sum of the absolute values ​​of the channel weights of each convolutional layer in the basic spatiotemporal neural network model; Identify and remove neuron channels whose sum of absolute channel weights is lower than a preset threshold, and generate a neural network model with channel pruning. The parameters of the neural network model after channel pruning are represented by low-bit-width quantization and fine-tuned to obtain a compressed and lightweight spatiotemporal neural network model.

[0071] Specifically, the functional connection analysis module 13 is used for: Specifically, based on the multi-channel EEG signals, the process proceeds continuously in a sliding window manner according to a preset step size and window length, calculating the phase synchronization index between the EEG signals of each channel within each time window, and generating a functional connectivity topology sequence, including: Set the window length and sliding step size of the sliding time window; Starting from the beginning of the multi-channel EEG signal, the time window is slid sequentially according to the sliding step size, and the multi-channel EEG signal segment corresponding to each time window is extracted. For each time window, calculate the phase synchronization index between any two channels of EEG signal segments. Each EEG acquisition channel is defined as a network node, and the phase synchronization index between any two network nodes is defined as the weight of the connection edge. The functional connection topology graph of the current time window is constructed. Arrange the functional connection topology diagrams of all time windows according to the chronological order of the corresponding time windows, and combine them to form a sequence of functional connection topology diagrams.

[0072] Specifically, the phase synchronization index between any two channels of EEG signals is calculated, including: The instantaneous phase sequence of the EEG signals from the two channels within the time window is extracted respectively; Calculate the instantaneous phase difference between the two channels at each moment within the time window to obtain the instantaneous phase difference sequence; The modulus of the complex exponential average of the phase difference sequence is calculated as an index of phase synchronization between the EEG signals of any two channels.

[0073] Further, from the functional connection topology graph sequence, the temporal stability features of graph attributes are extracted, including: Randomly select one functional connection topology graph from the sequence of functional connection topology graphs as the first functional connection topology graph; For each connecting edge in the first functional connection topology graph, calculate the reciprocal of the connecting edge weight as the distance; For each pair of non-repeating network nodes in the first functional connection topology graph, calculate the sum of the calculated distances of all possible paths between each pair of non-repeating network nodes, and select the minimum value from all the sums of calculated distances as the shortest path length between each pair of non-repeating network nodes. Calculate the reciprocal of the shortest path length as the transmission efficiency value between each pair of non-repeating network nodes in the first functional connection topology graph; Calculate the arithmetic mean of the transmission efficiency values ​​of all non-repeating network node pairs in the first functional connection topology graph, and use it as the global efficiency attribute value of the first functional connection topology graph. Traverse the sequence of functional connection topology graphs and calculate the global efficiency attribute value corresponding to each functional connection topology graph in the sequence; The global efficiency attribute values ​​are arranged according to the chronological order of the functional connection topology diagram to obtain a sequence of global efficiency attribute values. The variance of the global efficiency attribute value sequence is calculated and used as a temporal stability feature of the graph attribute.

[0074] Specifically, the probability calibration module 14 is used for: Specifically, the initial attention allocation probability distribution is numerically calibrated using the temporal stability features of the graph attributes to generate a calibrated attention allocation probability distribution, including: Obtain the historical graph attribute temporal stability feature sequence generated by the lightweight spatiotemporal neural network model in a historical verification process similar to the current auditory scene, and calculate the preset high-order statistical quantile of the historical graph attribute temporal stability feature sequence as a stability discrimination threshold; Calculate the ratio of the stability discrimination threshold to the temporal stability feature of the graph attribute, and use the natural logarithm of the ratio as the calibration intensity coefficient value; Based on the calibration intensity coefficient value, the initial attention allocation probability distribution is numerically transformed to generate a numerically adjusted intermediate probability distribution; The intermediate probability distribution is normalized to obtain the calibrated attention allocation probability distribution.

[0075] Specifically, based on the calibration intensity coefficient value, the initial attention allocation probability distribution is numerically transformed to generate a numerically adjusted intermediate probability distribution, including: If the calibration intensity coefficient value is greater than zero, the value is transformed according to a first preset scheme, wherein the first preset scheme includes: The sound source with the highest probability value is selected from the preliminary attention allocation probability distribution and used as the dominant sound source. Keep the probability values ​​of all non-dominant sound sources in the initial attention allocation probability distribution unchanged; Calculate the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient as the exponent, and use it as the probability adjustment coefficient of the dominant sound source; The probability value of the dominant sound source in the initial attention allocation probability distribution is multiplied by the probability adjustment coefficient to obtain the adjusted probability value of the dominant sound source. The probability values ​​of all non-dominant sound sources are integrated with the adjusted probability values ​​of the dominant sound sources to form an intermediate probability distribution; If the calibration intensity coefficient value is less than or equal to zero, the numerical transformation is performed according to the second preset scheme, wherein the second preset scheme includes: Calculate the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient as the exponent, and use it as the first weight; Calculate the difference between the value 1 and the first weight, and use it as the second weight; Construct a uniform probability distribution with the same number of sound sources as the initial attention allocation probability distribution, wherein the probability value of each sound source in the uniform probability distribution is equal; The probability value of each sound source in the initial attention allocation probability distribution is multiplied by the first weight, and the probability value of the corresponding sound source in the uniform probability distribution is multiplied by the second weight. The two sets of product results are then added together to form an intermediate probability distribution.

[0076] The attention determination module 15 is specifically used for: The target sound source that the target user pays attention to in the auditory scene is determined based on the attention allocation probability distribution.

[0077] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0078] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0079] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and modifications fall within the scope of this application and its equivalents, this application intends to include such modifications and modifications.

Claims

1. An auditory attention decoding method based on lightweight spatiotemporal networks, characterized in that, The method includes: Collect multichannel EEG signals from target users while listening to an auditory scene containing at least two competing sound sources; The multi-channel EEG signals are input into a pre-trained lightweight spatiotemporal neural network model for decoding, and the output is a preliminary attention allocation probability distribution. Based on the multi-channel EEG signals, the process is continuously advanced in a sliding window manner according to a preset step size and window length. The phase synchronization index between the EEG signals of each channel within each time window is calculated to generate a functional connectivity topology graph sequence. The graph attribute temporal stability features are then extracted from the functional connectivity topology graph sequence. The initial attention allocation probability distribution is numerically calibrated using the temporal stability features of the graph attributes to generate a calibrated attention allocation probability distribution. The target sound source that the target user pays attention to in the auditory scene is determined based on the attention allocation probability distribution.

2. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 1, characterized in that, Collect multichannel EEG signals from the target user while listening to an auditory scene containing at least two competing sound sources, including: Multiple EEG acquisition electrodes are worn on the target user's scalp at predetermined brain regions according to a pre-defined layout. In an auditory scene containing at least two competing sound sources, signals from each EEG acquisition electrode are continuously recorded at a preset sampling rate to obtain multi-channel EEG signals.

3. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 1, characterized in that, The construction process of a lightweight spatiotemporal neural network model includes: Multiple multi-channel EEG signal samples were collected from historical users under known attentional target sound source conditions to form an EEG signal training sample set; Obtain the target sound source label representing the user's attention corresponding to each of the above-mentioned EEG signal training samples, and form an attention label sample set; Based on the principle of depthwise separable convolution, an initial spatiotemporal neural network model containing spatial feature extraction module and temporal feature extraction module is constructed; The initial spatiotemporal neural network model is supervisedly trained using the EEG signal training sample set and the attention label sample set until the verification convergence is obtained, thus obtaining the basic spatiotemporal neural network model. Redundant neuron channels in the basic spatiotemporal neural network model are pruned, and the network parameters are quantized with low bit width to obtain a lightweight spatiotemporal neural network model.

4. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 3, characterized in that, Redundant neuron channels in the basic spatiotemporal neural network model are pruned, and the network parameters are quantized with low bit width to obtain a lightweight spatiotemporal neural network model, including: Calculate the sum of the absolute values ​​of the channel weights of each convolutional layer in the basic spatiotemporal neural network model; Identify and remove neuron channels whose sum of absolute channel weights is lower than a preset threshold, and generate a neural network model with channel pruning. The parameters of the neural network model after channel pruning are represented by low-bit-width quantization and fine-tuned to obtain a compressed and lightweight spatiotemporal neural network model.

5. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 1, characterized in that, Based on the multi-channel EEG signals, the process proceeds continuously in a sliding window manner according to a preset step size and window length. The phase synchronization index between the EEG signals of each channel within each time window is calculated, generating a functional connectivity topology sequence, including: Set the window length and sliding step of the sliding time window; Starting from the beginning of the multi-channel EEG signal, the time window is slid sequentially according to the sliding step size, and the multi-channel EEG signal segment corresponding to each time window is extracted. For each time window, calculate the phase synchronization index between any two channels of EEG signal segments. Each EEG acquisition channel is defined as a network node, and the phase synchronization index between any two network nodes is defined as the weight of the connection edge. The functional connection topology graph of the current time window is constructed. Arrange the functional connection topology diagrams of all time windows according to the chronological order of the corresponding time windows, and combine them to form a sequence of functional connection topology diagrams.

6. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 5, characterized in that, Calculate the phase synchronization index between any two channels of EEG signals, including: The instantaneous phase sequence of the EEG signals from the two channels within the time window is extracted respectively; Calculate the instantaneous phase difference between the two channels at each moment within the time window to obtain the instantaneous phase difference sequence; The modulus of the complex exponential average of the phase difference sequence is calculated as an indicator of phase synchronization between the EEG signals of any two channels.

7. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 1, characterized in that, From the functional connection topology graph sequence, extract graph attribute temporal stability features, including: Randomly select one functional connection topology graph from the sequence of functional connection topology graphs as the first functional connection topology graph; For each connecting edge in the first functional connection topology graph, calculate the reciprocal of the connecting edge weight as the distance; For each pair of non-repeating network nodes in the first functional connection topology graph, calculate the sum of the calculated distances of all possible paths between each pair of non-repeating network nodes, and select the minimum value from all the sums of calculated distances as the shortest path length between each pair of non-repeating network nodes. Calculate the reciprocal of the shortest path length as the transmission efficiency value between each pair of non-repeating network nodes in the first functional connection topology graph; Calculate the arithmetic mean of the transmission efficiency values ​​of all non-repeating network node pairs in the first functional connection topology graph, and use it as the global efficiency attribute value of the first functional connection topology graph. Traverse the sequence of functional connection topology graphs and calculate the global efficiency attribute value corresponding to each functional connection topology graph in the sequence; The global efficiency attribute values ​​are arranged according to the chronological order of the functional connection topology diagram to obtain a sequence of global efficiency attribute values. The variance of the global efficiency attribute value sequence is calculated and used as a temporal stability feature of the graph attribute.

8. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 1, characterized in that, Using the temporal stability features of the graph attributes, the initial attention allocation probability distribution is numerically calibrated to generate a calibrated attention allocation probability distribution, including: Obtain the historical graph attribute temporal stability feature sequence generated by the lightweight spatiotemporal neural network model in a historical verification process similar to the current auditory scene, and calculate the preset high-order statistical quantile of the historical graph attribute temporal stability feature sequence as a stability discrimination threshold; Calculate the ratio of the stability discrimination threshold to the temporal stability feature of the graph attribute, and use the natural logarithm of the ratio as the calibration intensity coefficient value; Based on the calibration intensity coefficient value, the initial attention allocation probability distribution is numerically transformed to generate a numerically adjusted intermediate probability distribution; The intermediate probability distribution is normalized to obtain the calibrated attention allocation probability distribution.

9. The auditory attention decoding method based on lightweight spatiotemporal networks according to claim 8, characterized in that, Based on the calibration intensity coefficient value, the initial attention allocation probability distribution is numerically transformed to generate a numerically adjusted intermediate probability distribution, including: If the calibration intensity coefficient value is greater than zero, the value is transformed according to a first preset scheme, wherein the first preset scheme includes: The sound source with the highest probability value is selected from the preliminary attention allocation probability distribution and used as the dominant sound source. Keep the probability values ​​of all non-dominant sound sources in the initial attention allocation probability distribution unchanged; Calculate the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient as the exponent, and use it as the probability adjustment coefficient of the dominant sound source; The probability value of the dominant sound source in the initial attention allocation probability distribution is multiplied by the probability adjustment coefficient to obtain the adjusted probability value of the dominant sound source. The probability values ​​of all non-dominant sound sources are integrated with the adjusted probability values ​​of the dominant sound sources to form an intermediate probability distribution; If the calibration intensity coefficient value is less than or equal to zero, the numerical transformation is performed according to the second preset scheme, wherein the second preset scheme includes: Calculate the value of an exponential function with the natural constant e as the base and the calibration intensity coefficient as the exponent, and use it as the first weight; Calculate the difference between the value 1 and the first weight, and use it as the second weight; Construct a uniform probability distribution with the same number of sound sources as the initial attention allocation probability distribution, wherein the probability value of each sound source in the uniform probability distribution is equal; The probability value of each sound source in the initial attention allocation probability distribution is multiplied by the first weight, and the probability value of the corresponding sound source in the uniform probability distribution is multiplied by the second weight. The two sets of product results are then added together to form an intermediate probability distribution.

10. An auditory attention decoding system based on lightweight spatiotemporal networks, characterized in that, An auditory attention decoding method based on a lightweight spatiotemporal network as described in any one of claims 1-9, comprising: The signal acquisition module is used to acquire multi-channel EEG signals of the target user when listening to an auditory scene containing at least two competing sound sources; The spatiotemporal decoding module is used to input the multi-channel EEG signals into a pre-trained lightweight spatiotemporal neural network model for decoding processing, and output a preliminary attention allocation probability distribution. The functional connectivity analysis module is used to continuously advance in a sliding window manner based on the multi-channel EEG signals according to a preset step size and window length, calculate the phase synchronization index between the EEG signals of each channel within each time window, generate a functional connectivity topology graph sequence, and extract graph attribute temporal stability features from the functional connectivity topology graph sequence. The probability calibration module is used to numerically calibrate the initial attention allocation probability distribution using the temporal stability features of the graph attributes, and generate a calibrated attention allocation probability distribution. An attention determination module is used to determine the target sound source that the target user is paying attention to in the auditory scene based on the attention allocation probability distribution.