Lightweight auditory attention detection method and system based on spatial-temporal feature enhancement

Through the lightweight auditory attention detection method and system based on spatial and temporal feature enhancement, the problems of low decoding accuracy, weak generalization and high calculation cost in the prior art are solved, and the auditory attention detection effect with higher accuracy and lower calculation cost is achieved.

CN120167952APending Publication Date: 2025-06-20ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510254351.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing methods of using EEG signals for auditory attention detection have limited decoding accuracy and generalization capabilities, high calculation costs, and are difficult to achieve practical application in hearing aids and earphone devices.

Method used

Lightweight auditory attention detection methods and systems based on spatial and temporal feature enhancement are used to process EEG data through spatial and temporal dependence coding, multi-scale temporal feature enhancement, spatial and temporal cross attention mechanism and classification prediction unit to extract auditory attention detection features.

Benefits of technology

It improves the accuracy of auditory attention detection tasks and reduces the calculation cost, making this method more suitable for practical applications and has high accuracy and strong generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120167952A_ABST
    Figure CN120167952A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight auditory attention detection method and system based on spatial-temporal feature enhancement, and relates to the technical field of auditory attention detection processing. According to the method, multiple sections of aligned electroencephalogram data are obtained through preprocessing, and a lightweight auditory attention detection model is designed to process the aligned electroencephalogram data so as to obtain an auditory attention detection result. A lightweight auditory attention detection model used in the method comprises a space-time dependent coding part, a multi-scale time feature enhancement part, a space-time cross attention part and a classification prediction part. According to the method, finer auditory attention spatio-temporal characteristics with better discrimination in the electroencephalogram signals can be captured, and the performance with better precision than that of an existing algorithm is obtained; and the model realizes light weight, the calculation cost is reduced, and the method is more suitable for practical application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of auditory attention detection and processing, and more specifically, to: 1. A lightweight auditory attention detection method based on spatio-temporal feature enhancement; 2. A lightweight auditory attention detection system based on spatio-temporal feature enhancement. Background Art

[0002] Auditory attention is an important neural mechanism for the brain to process sound information in a complex acoustic environment, enabling people to have the ability to focus on a specific speaker in a noisy environment - that is, to identify the direction of the target speaker.

[0003] Electroencephalogram (EEG) signals have been widely used in various auditory attention detection methods due to their non-invasive and high temporal resolution characteristics - by processing EEG signals through a network model to decode the attention selection activities in the brain and identify the speaker being focused on in a multi-speaker environment.

[0004] Although existing methods for auditory attention detection using EEG signals have achieved certain results, there are still the following problems:

[0005] 1. The decoding accuracy and generalization ability of existing models are limited.

[0006] 2. The existing models have a relatively high computational cost, making it difficult to implement for practical applications such as hearing aids and earphones. Summary of the Invention

[0007] Based on this, in view of the problems of low decoding accuracy, weak generalization, and high computational cost of existing models, it is necessary to provide a lightweight auditory attention detection method and system based on spatio-temporal feature enhancement.

[0008] The present invention is implemented by the following technical solutions:

[0009] In a first aspect, the present invention discloses a lightweight auditory attention detection method based on spatio-temporal feature enhancement, including:

[0010] Step 1: Obtain the EEG data X generated by being stimulated in a multi-speaker scenario original ;

[0011] Step 2: First, preprocess X original to remove artifacts to obtain the EEG data X pre , then window-slice X pre to obtain N segments of equally long EEG data x1 to x N , and then perform Euclidean alignment on x1 to x N to eliminate the data spatial distribution differences to obtain N segments of aligned EEG data

[0012] Step 3: Take Input the trained lightweight auditory attention detection model for processing to obtain N auditory attention detection results Predict1 to Predict that characterize the azimuth of the target speaker N ;

[0013] Among them, the lightweight auditory attention detection model includes: a spatio-temporal dependence encoding part, a multi-scale time feature enhancement part, a spatio-temporal cross-attention part, and a classification prediction part.

[0014] The spatio-temporal dependence encoding part is used to first perform temporal convolution on to obtain a temporal feature F t n , and then perform spatial convolution on F t n to obtain a spatial feature The multi-scale time feature enhancement part is used to perform feature extraction on F t n in multiple time dimensions and fuse them into a multi-scale time feature U n ; The spatio-temporal cross-attention part is used to first fuse U n and to obtain a shallow spatio-temporal feature F' n , and then process F' n and F t n through a low-parameter spatio-temporal cross-attention mechanism to obtain a deep spatio-temporal feature E n ; The classification prediction part is used to perform classification prediction on E n to obtain Predict n ; n ∈ [1, N].

[0015] This lightweight auditory attention detection method based on spatio-temporal feature enhancement implements the method or process according to the embodiments of the present disclosure.

[0016] In a second aspect, the present invention discloses a lightweight auditory attention detection system based on spatio-temporal feature enhancement, which uses the lightweight auditory attention detection method based on spatio-temporal feature enhancement disclosed in the first aspect.

[0017] The lightweight auditory attention detection system based on spatio-temporal feature enhancement includes: a data acquisition module, a preprocessing module, and an auditory attention detection module.

[0018] The data acquisition module is used to acquire electroencephalogram data X generated by stimulation in a multi-speaker scenario original ; The preprocessing module is used to first preprocess X original to remove artifacts to obtain electroencephalogram data X pre , and then perform window slicing on X pre to obtain N segments of equal-length electroencephalogram data x1 to xN , then perform Euclidean alignment on x1 to x N to eliminate the difference in the data space distribution and obtain N segments of aligned EEG data The auditory attention detection module is used to input the trained lightweight auditory attention detection model for processing to obtain N auditory attention detection results Predict1 to Predict that characterize the azimuth of the target speaker N .

[0019] This lightweight auditory attention detection system based on spatio-temporal feature enhancement implements the method or process according to the embodiments of the present disclosure.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] 1. The present invention obtains multiple segments of aligned EEG data through preprocessing, and designs a lightweight auditory attention detection model to process the aligned EEG data to obtain auditory attention detection results; the present invention can capture finer and more discriminative spatio-temporal features of auditory attention in EEG signals, and thus obtain better performance than existing algorithms in terms of accuracy; and the model is lightweight, reducing the computational cost and being more suitable for practical applications.

[0022] 2. The present invention designs a lightweight auditory attention detection model, which first captures continuous time steps and multi-channel features through a spatio-temporal dependence encoding part - first performs temporal convolution within each channel to capture temporal dependence, and then extracts cross-channel spatial features through spatial convolution to enhance the spatio-temporal representation ability, thereby ensuring high accuracy and strong generalization ability; then extracts multi-scale, multi-time-step multi-channel dependence relationships through a multi-scale temporal feature enhancement part while ensuring low parameters to integrate multi-level features and enhance and supplement robust temporal dynamic patterns; then extracts spatio-temporal feature context information through a novel spatio-temporal cross-attention part in a group-parallel manner, and recalibrates the weights by encoding global information, thereby enhancing the deep spatio-temporal correlation; finally, performs classification prediction through a classification prediction part to obtain the detection result. The lightweight auditory attention detection model used in the present invention accurately extracts auditory attention detection features and improves the accuracy of the auditory attention detection task. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0024] Figure 1Flow chart of the lightweight auditory attention detection method based on spatio-temporal feature enhancement proposed in Embodiment 1 of the present invention;

[0025] Figure 2 For Figure 1 Structural diagram of the spatio-temporal dependence encoding part in

[0026] Figure 3 For Figure 1 Structural diagram of the multi-scale temporal feature enhancement part in

[0027] Figure 4 For Figure 1 Structural diagram of the spatio-temporal cross-attention part in

[0028] Figure 5 For Figure 4 Structural diagram of the spatio-temporal dual-path processing layer in

[0029] Figure 6 For Figure 4 Structural diagram of the spatio-temporal cross-attention processing layer in

[0030] Figure 7 For Figure 1 Structural diagram of the classification prediction part in Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0032] It should be noted that when a component is referred to as being "installed on" another component, it can be directly on the other component or there may also be an intermediate component. When a component is considered to be "disposed on" another component, it can be directly disposed on the other component or there may be an intermediate component at the same time. When a component is considered to be "fixed to" another component, it can be directly fixed to the other component or there may be an intermediate component at the same time.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this invention belongs. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The term "or / and" used herein includes any and all combinations of one or more of the related listed items.

[0034] Embodiment 1

[0035] First, regarding the two problems mentioned in the background art: Through research, it is found that the former problem is due to the neglect of the cross-dependence of continuous time between different brain regions, the failure to effectively capture multi-scale dynamic time patterns and comprehensive spatio-temporal features, resulting in insufficient utilization of spatio-temporal features and limiting the decoding accuracy and generalization ability; the latter problem is that improving the decoding performance requires complex feature extraction and attention mechanisms with high computational complexity, leading to a large number of parameters and high computational complexity.

[0036] After understanding the above reasons, Embodiment 1 of the present invention specifically provides a lightweight auditory attention detection method based on spatio-temporal feature enhancement. Refer to Figure 1 , which shows the data flow diagram of this method, and actually also shows the process of this method.

[0037] As Figure 1 shown, the lightweight auditory attention detection method based on spatio-temporal feature enhancement includes the following steps:

[0038] Step 1, obtain the electroencephalogram data X generated by stimulation in a multi-speaker scenario original .

[0039] It should be noted that the stimulation mentioned here is the speech in a multi-speaker scenario.

[0040] Step 2, first preprocess X original to remove artifacts to obtain the electroencephalogram data X pre , then perform windowing segmentation on X pre to obtain N segments of equally long electroencephalogram data x1 to x N , and then perform Euclidean alignment on x1 to x N to eliminate the data spatial distribution difference to obtain N segments of aligned electroencephalogram data

[0041] Step 2 is actually a series of preprocessing operations on X original , which can also be written as including the following steps:

[0042] S201, preprocess X original to remove artifacts to obtain the electroencephalogram data X pre ;

[0043] S202, perform windowing segmentation on X pre to obtain N segments of equally long electroencephalogram data x1 to x N ;

[0044] S203, perform Euclidean alignment on x1 to x N to eliminate the data spatial distribution difference to obtain N segments of aligned electroencephalogram data

[0045] Ⅰ. For S201, the preprocessing method mainly includes the following operations:

[0046] ① Filter out line noise and power frequency interference at 50 Hz;

[0047] ② Perform band-pass filtering from 0.1 to 50 Hz using Fourier transform to remove non-relevant frequencies;

[0048] ③ Downsample to 128 Hz to reduce the amount of data and retain the main frequencies;

[0049] ④ Normalize each channel to standardize the mean and variance of the EEG data;

[0050] ⑤ Re-reference the EEG data to the common average.

[0051] Ⅱ. For S202, the method of window segmentation includes:

[0052] Determine the sliding window length L and step size Step according to the decision time;

[0053] Start the sliding window from the starting position of X pre and slide it with equal step size according to Step until reaching the end position of X pre ; Each step size corresponds to a segment of EEG data, and a total of x1 to x N are obtained.

[0054] It should be noted that Step should be less than L so that there is partial overlap between adjacent segments of EEG data. In this Embodiment 1, Step = L / 2. For easy understanding, an example is given below: If the decision time is 1 second, then L is equal to the sampling rate (128 Hz) and Step is equal to half of the sampling rate (64 Hz).

[0055] Ⅲ. For S303, the method of Euclidean alignment includes:

[0056] First normalize x1 to x N to obtain the normalized EEG data

[0057] Calculate the average covariance matrix S X ;

[0058] Perform eigenvalue decomposition on S X through inverse square root transformation and multiply it corresponding to to align the data to the same space to obtain

[0059] Of course, the process of S303 can also be expressed formulaically as:

[0060]

[0061] In the formula, |.| represents the absolute value operation, and MAX(.) represents the maximum value operation; T represents inversion.

[0062] Step 3: Input the trained lightweight auditory attention detection model for processing to obtain N auditory attention detection results Predict1 to Predict representing the azimuth of the target speaker N .

[0063] One of the cores of the present invention lies in providing a newly designed lightweight auditory attention detection model (abbreviated as Light-Listen-Net). It should be noted that the present invention needs to use the trained Light-Listen-Net with its network parameters in the optimal state. For the training process of Light-Listen-Net, the conventional training method of neural networks can be referred to, which will not be elaborated here.

[0064] Refer to Figure 1 , functionally, Light-Listen-Net includes: a spatio-temporal dependence encoding part, a multi-scale temporal feature enhancement part, a spatio-temporal cross-attention part, and a classification prediction part.

[0065] The following is a detailed introduction to each part:

[0066] Ⅰ. The spatio-temporal dependence encoding part is used to first perform temporal convolution on t n to obtain temporal feature F t n and then perform spatial convolution on F

[0067] Refer to Figure 2 , the spatio-temporal dependence encoding part includes: 1 temporal convolution layer and 1 spatial convolution layer.

[0068] Generally speaking: the spatio-temporal dependence encoding part first performs depthwise separable convolution operation on along the time dimension, significantly accelerating the calculation process while reducing the amount of calculation and the number of parameters; then perform depthwise separable convolution operation on F t n along the spatial dimension, further extracting the spatial features in the EEG signals while ensuring a low-parameter design, thereby modeling the cross-channel brain region cooperation mode.

[0069] Specifically:

[0070] 1. The temporal convolutional layer includes: 1 depthwise separable convolutional layer in the temporal dimension and 1 GELU activation function layer. In the temporal convolutional layer: The depthwise separable convolutional layer in the temporal dimension is used to perform depthwise separable convolution along the temporal dimension, and the GELU activation function layer is used to process the output of the depthwise separable convolutional layer in the temporal dimension through the GELU activation function to obtain F along the temporal dimension, and the GELU activation function layer is used to process the output of the depthwise separable convolutional layer in the temporal dimension through the GELU activation function to obtain F t n .

[0071] Of course, the processing process of the temporal convolutional layer can be formulated as:

[0072]

[0073] In the formula, TemporalDWConv2d(.) represents the processing process of the depthwise separable convolutional layer in the temporal dimension; GELU(.) represents the processing process of the GELU activation function layer.

[0074] 2. The spatial convolutional layer includes: 1 depthwise separable convolutional layer in the spatial dimension and 1 GELU activation function layer. In the spatial convolutional layer: The depthwise separable convolutional layer in the spatial dimension is used to perform depthwise separable convolution along the spatial dimension, and the GELU activation function layer is used to process the output of the depthwise separable convolutional layer in the spatial dimension through the GELU activation function to obtain t n along the spatial dimension, and the GELU activation function layer is used to process the output of the depthwise separable convolutional layer in the spatial dimension through the GELU activation function to obtain

[0075] Of course, the processing process of the spatial convolutional layer can be formulated as:

[0076]

[0077] In the formula, SpatialDWConv2d(.) represents the processing process of the depthwise separable convolutional layer in the spatial dimension; GELU(.) represents the processing process of the GELU activation function layer.

[0078] II. The multi-scale temporal feature enhancement part is used to perform feature extraction on F t n in multiple temporal dimensions and fuse them into a multi-scale temporal feature U n .

[0079] Referring to Figure 3 , the multi-scale temporal feature enhancement part includes: 4 dilated convolutional layers, 1 temporal dimension alignment layer, and 1 splicing layer.

[0080] In summary, the multi-scale temporal feature enhancement unit first performs four groups of dilated convolutions with different scales to extract temporal dimension features in parallel, expanding the receptive field while maintaining the number of parameters for efficient processing, and achieving the coordinated capture of fine-grained and long-range temporal dependencies; then, after temporal dimension alignment, a concatenation layer is performed to obtain U n It can comprehensively reflect the temporal patterns of EEG signals and enhance the model's ability to capture complex temporal dependencies.

[0081] In this Embodiment 1, the specifications of the convolutional kernels of the first dilated convolutional layer, the second dilated convolutional layer, the third dilated convolutional layer, and the fourth dilated convolutional layer are set to 1, 2, 3, and 5 respectively.

[0082] Specifically, in the multi-scale temporal feature enhancement unit:

[0083] The first dilated convolutional layer is used to extract ultra-short temporal scale features from F t n Extract ultra-short temporal scale features

[0084] The second dilated convolutional layer is used to extract shorter temporal scale features from F t n Extract shorter temporal scale features

[0085] The third dilated convolutional layer is used to extract medium temporal scale features from F t n Extract medium temporal scale features

[0086] The fourth dilated convolutional layer is used to extract longer temporal scale features from F t n Extract longer temporal scale features

[0087] The temporal dimension alignment layer is used to align Align the minimum time length in the temporal dimension;

[0088] The concatenation layer is used to concatenate the outputs of the alignment layer along the depth to obtain U n 。

[0089] It should be noted that: the temporal dimension alignment layer will Perform end truncation along the temporal dimension and dynamically align to the minimum time length - that is The corresponding ultra-short time specification.

[0090] Of course, the processing process of the multi-scale temporal feature enhancement unit can be formulated as:

[0091]

[0092] In the formula, DilatedConvk (.) represents the processing of the k-th dilated convolutional layer; AlightT min (.) represents the processing of the time dimension alignment layer; Concat(.) represents the processing of the concatenation layer.

[0093] III. The spatio-temporal cross-attention part is used to first fuse the features of U n and to obtain the shallow spatio-temporal feature F'. n Then, F' n and F t n are processed through a low-parameter spatio-temporal cross-attention mechanism to obtain the deep spatio-temporal feature E n .

[0094] See Figure 4 , the spatio-temporal cross-attention part includes: 1 feature fusion layer, 1 spatio-temporal cross-attention layer.

[0095] Generally speaking, the spatio-temporal cross-attention part first performs the feature fusion of U n and , then performs spatio-temporal dual-path decomposition and interaction enhancement through the designed spatio-temporal cross-attention mechanism, and obtains E through attention weighting processing n .

[0096] Specifically:

[0097] 1. The feature fusion layer includes: 1 skip convolutional layer, 1 stacking layer.

[0098] The skip convolutional layer is used to align the feature dimensions of U n and through skip convolution;

[0099] The stacking layer is used to add the outputs of the skip convolutional layer to obtain the shallow spatio-temporal feature F' n .

[0100] Of course, the processing process of the feature fusion layer can be formulated as:

[0101]

[0102] In the formula, SkipConv(.) represents the processing of the skip convolutional layer.

[0103] 2. The spatio-temporal cross-attention layer includes: 1 depth dimension alignment layer, 1 dimension adjustment layer, 1 spatio-temporal dual-path processing layer, 1 cross-attention processing layer.

[0104] ① The depth dimension alignment layer is used to align according to F' n and F t nThe relationship in the depth dimension to F t n Perform adaptive alignment to obtain the aligned temporal features

[0105] It should be noted that: If F t n has more input channels than F' n , then perform dimensionality reduction on F through one-dimensional convolution t n ; If F t n has fewer input channels than F' n , then pad with 0 to increase the number of channels of F t n .

[0106] Of course, the processing process of the depth dimension alignment layer can be formulated as:

[0107]

[0108] In the formula, APA(.) represents the processing process of the depth dimension alignment layer

[0109] ② The dimension adjustment layer is used to divide and reshape F' n , respectively to obtain 2 batch processing tensors

[0110]

[0111] It should be noted that the dimension adjustment layer will divide F' n , along the depth dimension into G groups, and then reshape the features of each group to obtain a batch processing tensor with a batch dimension of B×G; B represents the batch size

[0112] Of course, the processing process of the dimension adjustment layer can be formulated as:

[0113] GroupF1 n = Reshape(F n ′);

[0114]

[0115] In the formula, Reshape(.) represents the processing process of the dimension adjustment layer

[0116] ③ The spatio-temporal dual-path processing layer is used to perform spatio-temporal dual-path decomposition and interaction enhancement processing on to obtain 2 enhanced features

[0117] See Figure 5, the spatio-temporal dual-path processing layer includes: 2 spatial adaptive average pooling layers, 2 temporal adaptive average pooling layers, 2 product layers, 4 sigmoid activation function layers, and 2 group normalization layers.

[0118] Specifically, in the spatio-temporal dual-path processing layer:

[0119] The first spatial adaptive average pooling layer is used to perform adaptive average pooling processing on GroupF1 n in the spatial dimension;

[0120] The first sigmoid activation function layer is used to process the output of the first spatial adaptive average pooling layer through the sigmoid activation function;

[0121] The first temporal adaptive average pooling layer is used to perform adaptive average pooling processing on GroupF1 n in the temporal dimension;

[0122] The second sigmoid activation function layer is used to process the output of the first temporal adaptive average pooling layer through the sigmoid activation function;

[0123] The first product layer is used to multiply the output of the first sigmoid activation function layer, the output of the second sigmoid activation function layer, and GroupF1 n ;

[0124] The first group normalization layer is used to perform group normalization on the output of the first product layer to obtain EF1 n ;

[0125] The second spatial adaptive average pooling layer is used to perform adaptive average pooling processing in the spatial dimension;

[0126] The third sigmoid activation function layer is used to process the output of the second spatial adaptive average pooling layer through the sigmoid activation function;

[0127] The second temporal adaptive average pooling layer is used to perform adaptive average pooling processing in the temporal dimension;

[0128] The fourth sigmoid activation function layer is used to process the output of the second temporal adaptive average pooling layer through the sigmoid activation function;

[0129] The second product layer is used to multiply the output of the third sigmoid activation function layer, the output of the fourth sigmoid activation function layer, and ;

[0130] The second group normalization layer is used to perform group normalization on the output of the second product layer to obtain

[0131] Of course, the processing process of the spatio-temporal dual-path processing layer can be formulated as:

[0132] EF1n = GN(GroupF1n ⊙ σ(AAPs(GroupF1n)) ⊙ σ(AAP T (GroupF1n)));

[0133]

[0134] In the formula, GN(.) represents the processing process of the group normalization layer; AAP s (.) represents the processing process of the spatial adaptive average pooling layer; AAP T (.) represents the processing process of the temporal adaptive average pooling layer; σ(.) represents the processing process of the sigmoid activation function layer; ⊙ represents the processing process of the product layer.

[0135] ④ The cross-attention processing layer is used to perform low-parameter spatio-temporal cross-attention processing to obtain E n .

[0136] Referring to Figure 6 , the cross-attention processing layer includes: two adaptive global pooling layers, two matrix multiplication layers, one splicing layer, one one-dimensional convolutional layer, one product layer, and one dimensional reshaping layer.

[0137] Specifically, in the cross-attention processing layer:

[0138] The first adaptive global pooling layer is used to perform adaptive global pooling on EF1 n ;

[0139] The first matrix multiplication layer is used to process the output of the first adaptive global pooling layer through matrix multiplication ;

[0140] The second adaptive global pooling layer is used to perform adaptive global pooling on ;

[0141] The second matrix multiplication layer is used to process the output of EF1 n , the output of the second adaptive global pooling layer through matrix multiplication;

[0142] The splicing layer is used to splice the output of the first matrix multiplication layer and the output of the second matrix multiplication layer;

[0143] The one-dimensional convolutional layer is used to perform one-dimensional convolution on the output of the concatenation layer to obtain the attention weight W n ;

[0144] The product layer is used to multiply W n by EF1 n ;

[0145] The dimensionality reshaping layer is used to reshape the output of the product layer to obtain E n .

[0146] Of course, the processing process of the cross-attention processing layer can be formulated as:

[0147]

[0148] E n = Reshape(EF1 n ⊙W n );

[0149] In the formula, Conv1d(.) represents the processing process of the one-dimensional convolutional layer; AAP(.) represents the processing process of the adaptive average pooling layer; Mul(.) represents the processing process of the matrix multiplication layer; Concat(.) represents the processing process of the concatenation layer; Reshape(.) represents the processing process of the dimensionality reshaping layer; ⊙ represents the processing process of the product layer.

[0150] Ⅳ. The classification and prediction unit is used to perform classification and prediction on E n to obtain Predict n ; n ∈ [1, N].[[]END]]

[0151] See Figure 7 , the classification and prediction unit includes: 1 global average pooling layer, 1 linear layer, and 1 prediction function layer.

[0152] Generally speaking, the classification and prediction unit will first reduce the dimension of E n through global average pooling operation, then through linear transformation, and finally calculate the detection result through the prediction function.

[0153] Specifically, in the classification and prediction unit:

[0154] The global average pooling layer is used to perform global average pooling on E n ;

[0155] The linear layer is used to process the output of the global average pooling layer to obtain the auditory attention representation F n ;

[0156] The prediction function layer is used to calculate F n through the softmax function to obtain Predict n .

[0157] Of course, the processing process of the classification prediction unit can be formulated as:

[0158] F n = Linear(GAP(E n ));

[0159] Predict n = softmax(W·F n + b);

[0160] In the formula, GAP(.) represents the processing process of the global average pooling layer; Linear(.) represents the processing process of the linear layer; softmax(.) represents the processing process of the prediction function layer; b represents the bias term; W represents the weight matrix of the linear layer.

[0161] Then, using the above-mentioned YFD network model that has been trained, it is possible to process to obtain accurate Predict1 to Predict N .

[0162] Finally, it should be noted that:

[0163] In the background technology, it is mentioned that there are two problems with the existing model: The first problem is only the result appearance, and it is not easy to discover its internal reason that the spatio-temporal features are not fully utilized; Although the second problem analyzes that the attention mechanism is the main reason for the high computational cost, it is also not easy to solve it while ensuring or even improving the accuracy. Therefore, it is even more difficult to design the present method that takes into account solving the two problems.

[0164] Simulation verification

[0165] In this Embodiment 1, the above-mentioned proposed method is verified by simulation:

[0166] Build Light-Listen-Net, and introduce the existing network models MBSSFCC, DBPNet, and DARNet for comparison of effects, model parameters, and computational complexity. Among them, in order to quantitatively evaluate the effect of auditory attention detection, the classification accuracy rate is used as the evaluation index.

[0167] It should be noted that the above 4 network models are trained using the same publicly available dataset (such as the KUL EEG dataset). It should be noted that the data of the publicly available dataset will also be preprocessed in step two before being used for training the network model.

[0168] The comparison results of the 4 network models are shown in Table 1.

[0169] Table 1 Comparison Results of 4 Network Models

[0170] Network model Accuracy rate (%) Number of model parameters (million) Computational complexity (million) Light-Listen-Net 96.9 0.01 12.42 MBSSFCC 86.5 83.91 89.15 DBPNet 94.4 0.91 96.55 DARNet 94.8 0.08 16.36

[0171] As can be seen from Table 1, the accuracy of Light-Listen-Net reaches 96.9%. Compared with MBSSFCC, DBPNet, and DARNet, while Light-Listen-Net improves by 10.4%, 2.5%, and 2.1% respectively, the number of model parameters is reduced by 8391 times, 91 times, and 8 times respectively, and the computational complexity is reduced by 86.1%, 87.1%, and 24.1% respectively. Therefore, Light-Listen-Net realizes the lightweight of the model and the reduction of computational complexity while improving the accuracy, having great advantages.

[0172] Example 2

[0173] This Example 2 provides a lightweight auditory attention detection system based on spatio-temporal feature enhancement, which uses the lightweight auditory attention detection method based on spatio-temporal feature enhancement disclosed in Example 1.

[0174] The lightweight auditory attention detection system based on spatio-temporal feature enhancement includes: a data acquisition module, a preprocessing module, and an auditory attention detection module.

[0175] The data acquisition module is configured to: acquire the electroencephalogram data X generated by stimulation in a multi-speaker scenario original ;

[0176] The preprocessing module is configured to: first preprocess X original to remove artifacts to obtain the electroencephalogram data X pre , then perform windowing segmentation on X pre to obtain N segments of equally long electroencephalogram data x1 to x N , then perform Euclidean alignment on x1 to x N to eliminate the data spatial distribution difference to obtain N segments of aligned electroencephalogram data

[0177] The auditory attention detection module is configured to: input into the trained lightweight auditory attention detection model for processing to obtain N auditory attention detection results Predict1 to Predict N .

[0178] Since this system uses the lightweight auditory attention detection method based on spatio-temporal feature enhancement in Example 1, it also has the same effect and will not be repeated here.

[0179] Example 3

[0180] Embodiment 3 discloses a computer device, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps of the lightweight auditory attention detection method based on spatio-temporal feature enhancement disclosed in Embodiment 1 are implemented.

[0181] Embodiment 3 also discloses a readable storage medium. Computer program instructions are stored in the readable storage medium. When the computer program instructions are read and run by a processor, the steps of the lightweight auditory attention detection method based on spatio-temporal feature enhancement disclosed in Embodiment 1 are executed.

[0182] Embodiment 3 also discloses a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the lightweight auditory attention detection method based on spatio-temporal feature enhancement disclosed in Embodiment 1 are implemented.

[0183] The above-described embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

Claims

1. A lightweight auditory attention detection method based on spatiotemporal feature enhancement, characterized in that: include: Step 1: Obtain EEG data X generated by stimulation in a multi-speaker scene original ; Step 2: First, original Preprocessing is performed to remove artifacts to obtain EEG data X pre , and then X pre Perform window segmentation to obtain N segments of EEG data x1 to x N , then for x1~x N Perform Euclidean alignment to eliminate data spatial distribution differences and obtain N segments of aligned EEG data Step 3: Input the trained lightweight auditory attention detection model for processing to obtain N auditory attention detection results Predict1~Predict N ; Among them, the lightweight auditory attention detection model includes: The spatiotemporal dependency encoding unit is used to first Perform temporal convolution to obtain temporal features Again Perform spatial convolution to obtain spatial features The multi-scale temporal feature enhancement unit is used to Extract features from multiple time dimensions and fuse them into a multi-scale time feature U n ; The spatiotemporal cross attention part is used to first n and Perform feature fusion to obtain shallow spatiotemporal features F' n , then to F' n and The deep spatiotemporal features E are obtained by processing them through a low-parameter spatiotemporal cross attention mechanism. n ; as well as The classification prediction part is used to n Perform classification prediction to get Predict n ; n∈[1,N].

2. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 1 is characterized in that: In step 2, the method of performing window segmentation includes: Determine the sliding window length L and step length Step according to the decision time; Move the sliding window from X pre Start from the starting position, and slide in equal steps according to Step until you reach X pre The end position of each step length corresponds to a section of EEG data, and a total of x1~x N .

3. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 1 is characterized in that: In step 2, the method of performing Euclidean alignment includes: For x1~x N Normalize first to get the normalized EEG data calculate The average covariance matrix S X ; For S X The eigendecomposition is performed by inverse square root transformation and Multiply the corresponding data to align them to the same space to get 4. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 1 is characterized in that: The spatiotemporal dependency encoding unit includes: 1 temporal convolution layer and 1 spatial convolution layer; Among them, the time convolution layer includes: 1 time dimension depth separable convolution layer, 1 GELU activation function layer; in the time convolution layer: the time dimension depth separable convolution layer is used to A depth-wise separable convolution is performed along the time dimension. The GELU activation function layer is used to process the output of the depth-wise separable convolution layer along the time dimension through the GELU activation function to obtain The spatial convolution layer includes: 1 spatial dimension depth separable convolution layer and 1 GELU activation function layer; in the spatial convolution layer: the spatial dimension depth separable convolution layer is used to The GELU activation function layer is used to process the output of the spatial dimension depth-separable convolution layer through the GELU activation function to obtain 5. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 1 is characterized in that: The multi-scale temporal feature enhancement unit includes: 4 dilated convolutional layers, 1 temporal dimension alignment layer, and 1 splicing layer; The first dilated convolutional layer is used to Extracting ultra-short time scale features The second dilated convolutional layer is used to Extracting shorter time scale features The third dilated convolutional layer is used to Extracting medium-time scale features The fourth dilated convolutional layer is used to Extracting longer time scale features The time dimension alignment layer is used to Align the minimum time length in the time dimension; The concatenation layer is used to concatenate the output of the alignment layer along the depth to obtain U n .

6. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 2 is characterized in that: The spatiotemporal cross attention part includes: 1 feature fusion layer, 1 spatiotemporal cross attention layer; The feature fusion layer includes: 1 skip convolution layer and 1 stacking layer; the skip convolution layer is used to transfer U n and The feature dimension is aligned; the superposition layer is used to add the output of the jump convolution layer to obtain the shallow spatiotemporal feature F' n ; The spatiotemporal cross attention layer includes: 1 depth dimension alignment layer, 1 dimension adjustment layer, 1 spatiotemporal dual-path processing layer, and 1 cross attention processing layer; the depth dimension alignment layer is used to adjust the F' n and The relationship between the depth dimension Perform adaptive alignment to obtain aligned temporal features The Dimension Adjustment layer is used to transform F' n , Divide and reshape to get 2 batch tensors The spatiotemporal dual-path processing layer is used to Perform spatiotemporal dual-path decomposition and interactive enhancement processing to obtain two enhanced features The cross attention processing layer is used to Perform low-parameter spatiotemporal cross attention processing to obtain E n .

7. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 6 is characterized in that: The spatiotemporal dual-path processing layers include: 2 spatial adaptive average pooling layers, 2 temporal adaptive average pooling layers, 2 product layers, 4 sigmoid activation function layers, and 2 group normalization layers; The first spatial adaptive average pooling layer is used to Perform adaptive average pooling in the spatial dimension; The first sigmoid activation function layer is used to process the output of the first spatial adaptive average pooling layer through the sigmoid activation function; The first time-adaptive average pooling layer is used to Perform adaptive average pooling in the time dimension; The second sigmoid activation function layer is used to process the output of the first time-adaptive average pooling layer through the sigmoid activation function; The first product layer is used to combine the output of the first sigmoid activation function layer, the output of the second sigmoid activation function layer, To multiply; The first group normalization layer is used to group normalize the output of the first product layer to obtain The second spatial adaptive average pooling layer is used to Perform adaptive average pooling in the spatial dimension; The third sigmoid activation function layer is used to process the output of the second spatial adaptive average pooling layer through the sigmoid activation function; The second time-adaptive average pooling layer is used to Perform adaptive average pooling in the time dimension; The fourth sigmoid activation function layer is used to process the output of the second time-adaptive average pooling layer through the sigmoid activation function; The second product layer is used to multiply the output of the third sigmoid activation function layer, the output of the fourth sigmoid activation function layer, To multiply; The second group normalization layer is used to group normalize the output of the second product layer to obtain 8. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 6 is characterized in that: The cross-attention processing layer includes: 2 adaptive global pooling layers, 2 matrix multiplication layers, 1 concatenation layer, 1 one-dimensional convolution layer, 1 product layer, and 1 dimension reshaping layer; The first adaptive global pooling layer is used to Perform adaptive global pooling; The first matrix multiplication layer is used to multiply the The output of the first adaptive global pooling layer is processed; The second adaptive global pooling layer is used to Perform adaptive global pooling; The second matrix multiplication layer is used to multiply the The output of the second adaptive global pooling layer is processed; The concatenation layer is used to concatenate the output of the first matrix multiplication layer and the output of the second matrix multiplication layer; The one-dimensional convolution layer is used to perform one-dimensional convolution on the output of the concatenation layer to obtain the attention weight W. n ; The product layer is used to transform W n and To multiply; The dimension reshaping layer is used to reshape the output of the product layer to obtain E n .

9. The lightweight auditory attention detection method based on spatiotemporal feature enhancement according to claim 1 is characterized in that: The classification prediction part includes: 1 global average pooling layer, 1 linear layer, and 1 prediction function layer; In the classification prediction section: The global average pooling layer is used to n Perform global average pooling; The linear layer is used to process the output of the global average pooling layer to obtain the auditory attention representation F n ; The prediction function layer is used to predict F through the softmax function. n Calculate to get Predict n .

10. A lightweight auditory attention detection system based on spatiotemporal feature enhancement, characterized in that: It uses a lightweight auditory attention detection method based on spatiotemporal feature enhancement as described in any one of claims 1-8; The lightweight auditory attention detection system based on spatiotemporal feature enhancement includes: Data acquisition module, which is used to obtain EEG data generated by stimulation in a multi-speaker scene original ; The pre-processing module is used to first original Preprocessing is performed to remove artifacts to obtain EEG data X pre , and then X pre Perform window segmentation to obtain N segments of EEG data x1 to x N , then for x1~x N Perform Euclidean alignment to eliminate data spatial distribution differences and obtain N segments of aligned EEG data as well as The auditory attention detection module is used to Input the trained lightweight auditory attention detection model for processing to obtain N auditory attention detection results Predict1~Predict N .