Lightweight sleep stage classification method and device

The EEG signal characteristics are extracted through the depth separation convolution and parameterless attention module, and combined with the KAN classifier and simulated annealing algorithm to optimize the model parameters, the problem of complexity and low accuracy of the sleep stage classification algorithm in the existing technology is solved, and a lightweight and efficient sleep stage classification method is realized.

CN120105050APending Publication Date: 2025-06-06NORTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510034937.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, the sleep stage classification algorithm is complex and has low accuracy, and traditional deep learning models have challenges in computing complexity and parameter quantity, making it difficult to effectively port on mobile devices and embedded devices.

Method used

The lightweight deep learning method is adopted to extract the static characteristics of the EEG signal through deep separation convolution, and weighted processing is combined with the parameterless attention module to capture the contextual features between sleep stages. Optimize model parameters using KAN's classifier and gradient descent algorithm that simulates annealing to reduce model complexity and computational cost.

Benefits of technology

It realizes efficient sleep stage classification, improves classification accuracy, reduces model parameters and calculation costs, and is suitable for porting mobile devices and embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105050A_ABST
    Figure CN120105050A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight sleep stage classification method and device, and the method comprises the steps: 1, carrying out the time-invariant static feature extraction of electroencephalogram signal data during the sleep period of a subject through depth separable convolution; step 2, inputting the static features into a parameter-free attention module, and performing weighting processing to obtain weighted features; 3, taking the weighted features as input of a recurrent neural network, capturing context features among different sleep stages, and extracting time sequence information in the electroencephalogram signal data; 4, time sequence information in the electroencephalogram signal data is processed based on a KAN classifier in the prediction model, a corresponding sleep stage prediction result is output, and model parameters of the prediction model are updated based on a simulated annealing gradient descent algorithm. The method is used for solving the technical problems of complex classification algorithm and low classification precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of intelligent medical computer-aided diagnosis, and in particular to a lightweight sleep stage classification method and device. Background Art

[0002] Sleep staging plays an important role in improving sleep quality and treating diseases related to sleep disorders. In order to determine the sleep stage, experts obtain physiological information of the human body in various ways to make a judgment. Polysomnography (PSG) is considered the gold standard for sleep scoring and is widely used to treat typical sleep disorders. PSG includes biological signals related to physical activity, such as brain activity (electroencephalogram, EEG), eye movements (electrooculogram, EOG), heart rhythm (electrocardiogram, and jaw, facial or limb muscle activity (electromyogram, EMG). Among these biological signals, EEG signals perform best in sleep staging tasks.

[0003] Generally, the recorded physiological information is divided into 30-second time segments, which are then classified by experts into one of five sleep stages based on the criteria established by the American Academy of Sleep Medicine (AASM). However, this manual, lightweight sleep stage classification method is labor-intensive and time-consuming, as the EEG signals must be collected for several nights while the subjects are sleeping and then handed over to clinicians for processing. In addition, manual sleep staging is also affected by human subjective evaluation and data presentation. Therefore, many researchers have tried to develop a method that can automatically classify sleep stages.

[0004] With the rapid development of computer hardware, especially CPU and GPU, the computing power of hardware has been greatly improved. Therefore, deep learning methods have been widely developed. However, researchers generally pursue improving performance by adding components, while ignoring the demand for computational complexity in practical applications. For example, XSleepNet adopts a two-branch parallel processing structure, and considers the prediction results of the two branches during back propagation, but this also leads to an increase in the number of parameters and computational costs. In addition, the success of Transformer has also made researchers begin to pay attention to the ability of self-attention mechanisms, such as TransSleep, but the premise of the self-attention mechanism is a large-scale rating matrix, which inevitably leads to an increase in the number of model parameters and an increase in the amount of computation. Correspondingly, more training data and longer training time are required. However, in some specific areas, such as the use of daily mobile devices or large-scale censuses, miniaturized models are required. On the other hand, in traditional training methods, the size of the model is positively correlated with performance. How to adjust the training strategy so that lightweight models perform well on sleep staging tasks is also a problem. Summary of the invention

[0005] The embodiments of the present application provide a lightweight sleep stage classification method and device to solve the technical problems of complex classification algorithms and low classification accuracy.

[0006] On the one hand, an embodiment of the present application provides a lightweight sleep stage classification method, including:

[0007] Step 1: Use deep separable convolution to extract time-invariant static features from the EEG signal data of the subjects during sleep;

[0008] Step 2: Input the static features into the parameter-free attention module for weighted processing to obtain weighted features;

[0009] Step 3: Use the weighted features as the input of the recurrent neural network to capture the contextual features between different sleep stages and extract the timing information in the EEG signal data;

[0010] Step 4: Process the timing information in the EEG signal data based on the KAN classifier in the prediction model and output the corresponding sleep stage prediction result, wherein the prediction model is based on the simulated annealing gradient descent algorithm to update the model parameters.

[0011] Optionally, the method includes: the prediction model uses a gradient descent algorithm of simulated annealing to adjust the update optimization strategy of the parameters.

[0012] Optionally, in step 2, the static features are input into a parameter-free attention module for weighted processing to obtain weighted features, including:

[0013] Calculate the mean and variance between static features;

[0014] Calculate the corresponding mean and variance based on the mean and variance;

[0015] The similarity between different features is calculated based on the mean and variance, and the weights corresponding to different features in the classification stage are output.

[0016] Optionally, the prediction model uses a simulated annealing gradient descent algorithm to adjust the update optimization strategy of the parameters, including:

[0017] The comparison results of the loss functions of two adjacent rounds of prediction are calculated, and according to the comparison results, it is determined whether to directly update the parameters of the prediction model or to execute the gradient descent algorithm of simulated annealing to update the parameters of the prediction model.

[0018] Optionally, the prediction model uses a simulated annealing gradient descent algorithm to adjust the update optimization strategy of the parameters, and also includes:

[0019] If the loss function of the current round is greater than or equal to the loss function of the previous round, the parameters of the prediction model are updated according to the difference between the current "temperature" and the loss function.

[0020] On the other hand, the present application also provides a lightweight sleep stage classification device, comprising:

[0021] The static feature extraction module is used to extract time-invariant static features from the EEG signal data of the subjects during sleep using deep separable convolution;

[0022] The feature weighting module is used to input static features into the parameter-free attention module for weighted processing to obtain weighted features;

[0023] The timing information extraction module is used to use the weighted features as the input of the recurrent neural network to capture the contextual features between different sleep stages and extract the timing information in the EEG signal data;

[0024] The output module is used to process the time series information in the EEG signal data based on the KAN classifier in the prediction model and output the corresponding sleep stage prediction results, wherein the prediction model is based on the gradient descent algorithm of simulated annealing to update the model parameters.

[0025] The present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing any one of the above methods.

[0026] The present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a computer, any of the above methods is implemented.

[0027] The present application has the following advantages:

[0028] The prediction model used in this application combines the prior knowledge of professionals in sleep staging and uses the corresponding lightweight component KAN classifier to complete the subtasks, which reduces the number of model parameters while improving the interpretability of the model.

[0029] Secondly, during the training process, the prediction model combines the frequency principle of deep learning and uses the simulated annealing gradient descent algorithm to adjust the parameter update strategy to avoid erroneous feature information, thereby improving the accuracy of the model classification results.

[0030] Experimental results show that compared with the methods proposed by other existing models, the proposed method has achieved the most advanced level in the classification of sleep stages. In addition, compared with the existing deep neural networks, our model has the smallest parameters, and the difficulty of transplantation on mobile devices and embedded devices is greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0032] Figure 1 A step diagram of a lightweight sleep stage classification method provided in an embodiment of the present application.

[0033] Figure 2 A diagram of a lightweight sleep stage classification network framework based on KAN-net provided in an embodiment of the present application.

[0034] Figure 3 A schematic diagram of a lightweight sleep stage classification device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0036] Figure 1 A sleep stage classification method is provided for the embodiment of the present application. It should be noted that the sleep stage classification method provided in the present application is based on a new framework called SASleepNet. Figure 2 As shown in the figure, the prediction model combines the prior knowledge of professionals in sleep staging and uses the corresponding lightweight component KAN classifier to complete the subtask, which reduces the number of model parameters and improves the interpretability of the model. The specific method includes:

[0037] Step 1: Use deep separable convolution to extract time-invariant static features from the EEG signal data of the subjects during sleep.

[0038] In one embodiment, a one-dimensional convolutional neural network is used to learn representations from the input EEG signal data of the subject during sleep. In order to minimize parameters, the model initially uses a one-dimensional deep separable convolution to extract time-invariant static features from the original EEG signal through a feature extractor. Compared with standard convolution, the model uses fewer parameters. Specifically, the module consists of four convolutional layers, each of which contains batch normalization and a rectified linear unit (ReLU) activation function. However, the inherent limitation of deep separable convolution, the limited inter-channel information transmission may have a negative impact on sleep staging performance. To alleviate it, standard convolution and deep separable convolution are applied in sequence to obtain representations. This method can learn to discriminate time-invariant static features from EEG data (Electroencephalogram, EEG) with fewer parameters while retaining the interaction between channels. As shown in formula (1), the calculation method is as follows:

[0039]

[0040] Among them, CNN is standard convolution, DCNN is depth-separable convolution, and both CNN and DCNN are functions that convert single-channel EEG signals into time-invariant features. d and θ r are the parameters of CNN and DCNN respectively, e i is the extracted time-invariant static feature, x i is the input EEG signal data (EEG signal) of the subject during sleep, and i refers to the label of different sleep stages.

[0041] Step 2: Input the static features into the parameter-free attention module for weighted processing to obtain weighted features;

[0042] In one embodiment, in order to learn the distinguishing information of different sleep stages, the static features are input into the parameter-free attention module, and the parameter adaptive attention mechanism is used to enhance the static features. Among them, SimAM uses energy functions to generate independent attention weights for each neuron and achieves state-of-the-art (SOTA, State-of-the-Art Performance) performance in various visual tasks. The model used in this embodiment can calculate the feature e obtained by deep separable convolution in each period i The purpose of this layer is to weight the features extracted during representation learning according to their relative contribution to the final classification, providing weighted feature representations for downstream modules.

[0043] Optionally, in step 2, the static features are input into a parameter-free attention module for weighted processing to obtain weighted features, including:

[0044] Calculate the mean and variance between static features;

[0045] Calculate the corresponding mean and variance based on the mean and variance;

[0046] The similarity between different features is calculated based on the mean and variance, and the weights corresponding to different features in the classification stage are output.

[0047] In one embodiment, the specific data processing process of step 2 is as follows:

[0048]

[0049]

[0050]

[0051] in, is the feature obtained by the above depth-wise separable convolution, λ is the regularization coefficient, M is the number of static features, μ t are the mean and variance between static features, is the mean and variance, w j are the weights corresponding to different features in the classification stage.

[0052] It should be noted that the weight w j The lower, the characteristics The more different it is from the surrounding neurons, the more important its processing is. Then, the feature With the corresponding weight w j Multiply them together to get the weighted features.

[0053] Step 3: Use the weighted features as the input of the recurrent neural network to capture the contextual features between different sleep stages and extract the timing information in the EEG signal data.

[0054] In one embodiment, it should be noted that sleep stage transitions exhibit time dependence, and clinicians consider previous stages when determining later sleep stages. For example, REM is usually after stage N2, and less common in stage W or N1. In order to capture these temporal dynamics, this embodiment uses an RNN architecture, combined with a recurrent neural network consisting of LSTM layers and Dropout layers to reduce overfitting. This stage is intended to simulate the transition dynamic relationship between different sleep stages.

[0055] The specific processing of weighted features is shown in formula (5):

[0056]

[0057] Among them, RNN represents the LSTM model used to process weighted features, θs is the learnable parameter of the model, o j Respectively represent LSTM processing input The vector obtained later is are the characteristics of different stages, (o 1 ,o 2 ,…,o M ) is the timing information in the EEG signal data.

[0058] Step 4: Process the timing information in the EEG signal data based on the KAN classifier in the prediction model and output the corresponding sleep stage prediction result, wherein the prediction model is based on the simulated annealing gradient descent algorithm to update the model parameters.

[0059] In one embodiment, in order to solve the problem that traditional multi-layer perceptron classifiers (MLPs) have poor ability to fit nonlinear features, this embodiment proposes a solution of integrating KAN-net into the classifier. In the prediction model, the KAN-based classifier processes the output of the temporal context learning module (LSTM) to determine the sleep stage of each period. Unlike traditional multi-layer perceptrons, which rely heavily on linear transformations and may discard important nonlinear relationships between features and classification results. Here, the prediction model is centered on three layers of learnable nonlinear functions, each using a different B-spline function. The middle layer summarizes the B-spline output without further nonlinear transformation. The flexibility of the spline enables it to adaptively simulate complex data relationships by adjusting its shape, thereby minimizing approximation errors and enhancing the network's ability to learn subtle patterns from high-dimensional data sets. This solves the limitations of MLPs and improves the accuracy of sleep stage classification. On the other hand, replacing linear transformations with B-splines helps reduce the total number of model parameters. This part is shown in formula (6):

[0060] f(H)=KAN(o 1 ,o 2 ,...,o M )

[0061] =(Φ 3 Φ 2 Φ 1 )(o 1 ,o 2 ,...,o M ) (6)

[0062] Among them, (o 1 ,o 2 ,…,o M ) is the output variable of LSTM, Φ qis a continuous function, q = 1, 2, 3..., and can be transformed nonlinearly. Here, the B-spline function is used to complete Φ q The output is the corresponding predicted sleep stage result.

[0063] It should be noted that the KAN-based classifier effectively captures nonlinear features in the time series and achieves accurate many-to-many sleep stage classification. After training, network pruning and sparsification methods informed by established sleep stage principles can be used to improve the network in order to further reduce the complexity of the model and enhance interpretability.

[0064] Optionally, the method includes: the prediction model uses a gradient descent algorithm of simulated annealing to adjust the update optimization strategy of the parameters.

[0065] Optionally, the prediction model uses a simulated annealing gradient descent algorithm to adjust the update optimization strategy of the parameters, including:

[0066] The comparison results of the loss functions of two adjacent rounds of prediction are calculated, and according to the comparison results, it is determined whether to directly update the parameters of the prediction model or to execute the gradient descent algorithm of simulated annealing to update the parameters of the prediction model.

[0067] Optionally, the prediction model uses a simulated annealing gradient descent algorithm to adjust the update optimization strategy of the parameters, and also includes:

[0068] If the loss function of the current round is greater than or equal to the loss function of the previous round, the parameters of the prediction model are updated according to the difference between the current "temperature" and the loss function.

[0069] In one embodiment, in the early stage of training, the model focuses on learning low-frequency features and is less affected by noise. At the same time, the judgment made by the model at this time is less reliable, so the traditional batch gradient descent algorithm is used to fit the contour information of the data distribution, and the comparison result of the loss function of two adjacent rounds of prediction is calculated. According to the comparison result, it is determined whether to directly update the parameters of the prediction model or to execute the simulated annealing gradient descent algorithm to update the parameters of the prediction model, as shown in formula (7):

[0070]

[0071] Among them, Loss t is the loss function of the previous round, Loss t-1 is the loss function of the previous round, and x is the threshold.

[0072] It should be noted that according to the two comparison results Loss t <Loss t-1 and Loss t ≥Loss t-1, respectively determine which x threshold the generated random number meets, to determine the parameter update method of the prediction model. t <Loss t-1 This means that if the loss function of the current batch is less than the loss function of the previous batch, then the acceptance is 1, that is, the parameters are updated directly. t ≥Loss t-1 This means that if the loss function of the current batch is greater than or equal to the loss function of the previous batch, then the acceptance will be adjusted according to the difference between the current "temperature" and the loss function, that is, the gradient descent algorithm of simulated annealing is used to adjust the parameters.

[0073] In the later stages of training, the model can make more accurate judgments. At the same time, when learning high-frequency detail information, it is easily affected by the noise inherent in the data. At this time, the model will judge whether it can learn useful information based on the loss function value of the model on the current batch of data, and then update the parameters.

[0074]

[0075] As shown in formula (8), the acceptance probability of parameter updates is controlled by an exponential function, which is similar to the temperature in simulated annealing. As training proceeds, the "temperature" will gradually decrease according to the formula.

[0076] Based on the above-mentioned lightweight sleep stage classification method, this application also conducts certain simulation experiments on this method, and the simulation results are as follows:

[0077] This application mainly deals with five classification tasks. In the classification task, there are four different combinations between the predicted label (Predicted label) and the correct label (True label), which constitute the confusion matrix. Three indicators can be calculated through the confusion matrix, namely the pre-classification F1 score (F1), the macro average F1 score (MF1) and the overall accuracy (ACC).

[0078] Table 1 compares the results of SASleepNet with other models on three datasets. The results show that compared with other sleep staging models, SASleepNet achieves the highest accuracy, with an accuracy of 88.1% for SleepEDF20, 85.2% for Slee compared to pEDF78, and 86.7% for SHHS.

[0079] Among the five sleep stages, they are weak, N1, N2, N3, and REM. Among them, the N1 stage is more difficult to classify because the amount of data in the N1 stage is relatively small and it belongs to the transition period, resulting in the presence of characteristics of different sleep stages in the same epoch. From the results, it can be seen that the method provided in this application performs better in the classification of the N1 stage, especially in Sleep-EDF78 and SHHS. The accuracy of the N1 stage is about 10% to 20% higher than other models, which is mainly due to the use of the simulated annealing algorithm. This model avoids the problem that the characteristics of different sleep stages exist in the same period, thereby affecting the correct learning of the model. For such wrong information, the simulated annealing training strategy will selectively ignore it to avoid misleading the direction of model update.

[0080] In addition to the improvement in results, the variance range of SASleepNet output after using the simulated annealing training method (taking ACC accuracy as an example) is rapidly reduced, and its variance range is significantly smaller than that of other models. This is because the simulated annealing algorithm selectively updates parameters according to the changes in the loss function, reducing the impact of the error gradient on the model, allowing the model to reach a stable optimal solution position. In addition, the smaller the accuracy variance range of the test set, the more stable the performance of the model. Compared with selecting the optimal solution to store in the validation set, this method reduces the possibility of misjudging the model performance due to random errors. Furthermore, the interpretability of deep learning models has been criticized by the medical industry, mainly due to the short board of credibility. Deep learning is a probabilistic problem. The variance of the model evaluation results affects the credibility of the model. The experimental results show that the variance of the evaluation index after the simulated annealing method is only 0.3%, indicating that the evaluation results of the model are very stable, providing a solid foundation for practical applications.

[0081] Table 1: Comparison of performance indicators of this algorithm and different classifiers

[0082]

[0083] On the other hand, the present application also provides a lightweight sleep stage classification device, such as Figure 3 As shown, including:

[0084] The static feature extraction module is used to extract time-invariant static features from the EEG signal data of the subjects during sleep using deep separable convolution;

[0085] The feature weighting module is used to input static features into the parameter-free attention module for weighted processing to obtain weighted features;

[0086] The timing information extraction module is used to use the weighted features as the input of the recurrent neural network to capture the contextual features between different sleep stages and extract the timing information in the EEG signal data;

[0087] The output module is used to process the timing information in the EEG signal data based on the KAN classifier in the prediction model and output the corresponding sleep stage prediction result.

[0088] It should be noted that the lightweight sleep stage classification device provided in this embodiment can implement method steps that are completely consistent with the above-mentioned lightweight sleep stage classification method, which will not be repeated here.

[0089] The present application provides a terminal device, which includes a processor and a memory, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is used to execute the program instructions stored in the computer storage medium. The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function; the processor described in the embodiment of the present application can be used for the operation of a lightweight sleep stage classification method.

[0090] The present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, any one of the above method embodiments, a lightweight sleep stage classification method, is implemented.

[0091] In one embodiment, the present application also provides a storage medium, specifically a computer-readable storage medium (Memory), which is a memory device in a terminal device for storing programs and data. It is understandable that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and the extended storage medium supported by the terminal device. The computer-readable storage medium provides a storage space, which stores the operating device of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement a lightweight sleep stage classification method in the above embodiment.

[0092] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, devices, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0093] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices, and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0094] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1A function specified in one or more boxes.

[0095] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0096] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0097] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A lightweight sleep stage classification method, characterized in that: include: Step 1: Use deep separable convolution to extract time-invariant static features from the EEG signal data of the subjects during sleep; Step 2, inputting the static features into the parameter-free attention module, performing weighted processing, and obtaining weighted features; Step 3, using the weighted features as input to a recurrent neural network to capture contextual features between different sleep stages and extract timing information from the EEG signal data; Step 4: Process the timing information in the EEG signal data based on the KAN classifier in the prediction model, and output the corresponding sleep stage prediction result, wherein the prediction model updates the model parameters based on the gradient descent algorithm of simulated annealing.

2. A lightweight sleep stage classification method as claimed in claim 1, characterized in that: include: The prediction model adopts a simulated annealing gradient descent algorithm to adjust the update optimization strategy of parameters.

3. A lightweight sleep stage classification method as claimed in claim 1, characterized in that: The step 2, inputting the static features into the parameter-free attention module, performing weighted processing to obtain weighted features, includes: Calculating the mean and variance between the static features; Calculate the corresponding mean and variance based on the mean and variance; The similarities between different features are calculated based on the mean and variance and the mean and variance, and the weights corresponding to the different features in the classification stage are output.

4. A lightweight sleep stage classification method as claimed in claim 2, characterized in that: The prediction model uses a simulated annealing gradient descent algorithm to adjust the update optimization strategy of the parameters, including: Compare the loss functions of two adjacent rounds of prediction and determine whether to directly update the parameters of the prediction model or to execute a gradient descent algorithm of simulated annealing to update the parameters of the prediction model according to the comparison result.

5. A lightweight sleep stage classification method as claimed in claim 4, characterized in that: The prediction model uses a simulated annealing gradient descent algorithm to adjust the update optimization strategy of the parameters, and also includes: If the loss function of the current round is greater than or equal to the loss function of the previous round, the parameters of the prediction model are updated according to the difference between the current "temperature" and the loss function.

6. A lightweight sleep stage classification device, characterized in that: include: The static feature extraction module is used to extract time-invariant static features from the EEG signal data of the subjects during sleep using deep separable convolution; A feature weighting module, used for inputting the static features into the parameter-free attention module for weighted processing to obtain weighted features; A timing information extraction module, used to use the weighted features as input to a recurrent neural network, capture contextual features between different sleep stages, and extract timing information from the EEG signal data; The output module is used to process the timing information in the EEG signal data based on the KAN classifier in the prediction model and output the corresponding sleep stage prediction result, wherein the prediction model updates the model parameters based on the gradient descent algorithm of simulated annealing.

7. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 5 is implemented.