Explanatable sleep staging method and device based on training after visual language model supervision fine tuning

By rendering multi-channel sleep data into two-dimensional images and using a visual language model for supervised fine-tuning training, structured JSON results are generated. This solves the problems of interpretability and integration of clinical rules in existing sleep staging technologies, and achieves transparent and highly trustworthy automated sleep staging.

CN121542800APending Publication Date: 2026-02-17ZHEJIANG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511663265.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing technologies lack interpretability in the sleep staging process, fail to make the interpretation process transparent, and are not explicitly integrated into clinical rules, resulting in low trust among clinicians and making it difficult to trace the interpretation logic of automated tools.

Method used

By rendering multi-channel sleep data into two-dimensional images and using a visual language model for supervised fine-tuning training, structured JSON results are generated, including sleep stage labels, judgment reasoning text, and related interpretation rule numbers, simulating the expert's thought process.

Benefits of technology

It achieves transparency in sleep staging results and provides data support that conforms to clinical rules, enhancing clinicians' trust in automated staging results and maintaining consistency between the model's reasoning logic and recognized clinical guidelines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542800A_ABST
    Figure CN121542800A_ABST
Patent Text Reader

Abstract

The invention discloses an interpretable sleep staging method based on training after visual language model supervision fine tuning, which comprises the following steps: acquiring multi-channel sleep data and a corresponding sleep reasoning text to construct a data set; performing numbering based on interpretation rules related to sleep in clinical rule knowledge to construct a corresponding rule base; taking a pre-trained visual language model as a basic model, and performing supervision fine tuning training on the basic model through the data set to obtain a visual language model; multi-channel sleep data to be analyzed and preset cue words are input into the visual language model to generate a structured JSON result, and the JSON result comprises a sleep stage label, a judgment reason text and a rule number list of related interpretation rules. The invention further provides a device capable of explaining sleep staging. According to the method provided by the invention, the thinking process of experts can be simulated, so that comprehensive data support conforming to clinical rules is provided for each sleep staging result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence-assisted medical technology, and in particular relates to an interpretable sleep staging method and device based on supervised fine-tuning training using a visual language model. Background Technology

[0002] Clinically, sleep staging using full-night sleep data (including physiological signals such as electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG)) recorded by polysomnography (PSG) is the gold standard for sleep monitoring and diagnosis of related diseases. According to the American Academy of Sleep Medicine (AASM) standards, sleep is divided into five stages: wakefulness (W), non-rapid eye movement (NREM) stage 1 (N1), NREM stage 2 (N2), NREM stage 3 (N3), and rapid eye movement (REM) sleep.

[0003] Traditional sleep staging relies on trained sleep technicians or doctors to manually interpret hours of polysomnography data. This process requires not only extensive professional knowledge and clinical experience from the interpreter, but is also extremely time-consuming and labor-intensive, typically taking one to two hours to process a full night's sleep data. Furthermore, the results of manual interpretation are subject to subjectivity, with consistency among different experts usually ranging from 80% to 90%, which to some extent affects the standardization and objectivity of diagnosis.

[0004] To overcome the limitations of manual interpretation, academia and industry have proposed various automated sleep staging methods based on machine learning and deep learning. Early research mainly focused on extracting time-domain, frequency-domain, or nonlinear dynamic features from polysomnography signals, and then using traditional machine learning models such as support vector machines and random forests for classification. In recent years, models based on convolutional neural networks, recurrent neural networks, and their variants have been widely used in end-to-end sleep staging. These models can automatically learn discriminative features from the raw signals.

[0005] Patent document CN119548095A discloses a sleep stage detection method based on a multimodal and visual transformation network. This method first preprocesses and frames the EEG and EEG signals, then extracts the time-frequency features of the EEG signals and the time-domain features of the EEG signals, respectively. Next, it uses a visual transformation network to fuse the multimodal features, and finally outputs the sleep stage and confidence score through a classifier. This method improves classification performance through multimodal feature fusion.

[0006] Patent document CN120217200A discloses a sleep staging method based on spatiotemporal feature encoding and multi-source fusion. This method constructs a graph structure representation of multi-source physiological signals, uses a graph spatial encoder to capture the spatial dependencies between channels, then models the long-range dependencies of time series through a transformer architecture, and finally integrates spatiotemporal features through a fusion module to output sleep stage labels. This method innovates in its model architecture and can effectively utilize the spatiotemporal correlation of multi-source signals.

[0007] However, the aforementioned existing technologies have the following shortcomings: First, the output information is limited, only providing sleep stage classification results or confidence levels, lacking explanation of the interpretation process, and failing to inform clinicians of the basis for judgment, making the interpretation process opaque and difficult to gain the full trust of clinicians; Second, there is insufficient integration of clinical knowledge, failing to explicitly incorporate the valuable expert knowledge of the clinical interpretation rules established by the American Academy of Sleep Medicine into the model, and the model's learning process is entirely data-driven, which may lead to deviations between its interpretation logic and clinical standards; Third, the lack of interpretability severely limits the application value of automated tools in assisted diagnosis, teaching, and research. When using such tools, if doctors have doubts about a certain interpretation result, they cannot trace the judgment logic. Summary of the Invention

[0008] This invention discloses an interpretable sleep staging method and apparatus based on supervised fine-tuning training using a visual language model. This method can mimic the thought process of experts, thereby providing comprehensive and clinically compliant data support for each sleep staging result.

[0009] To achieve the objectives of this invention, the following technical solution is provided: an interpretable sleep staging method based on supervised fine-tuning training using a visual language model, comprising the following steps: Acquire multi-channel sleep data and corresponding sleep inference text, draw corresponding waveform images from the multi-channel sleep data, and use the waveform images of the target frame and the corresponding preceding and following frames as an image sequence, and label the image sequence according to the sleep stage, and combine the labels, image sequence and sleep inference text into a dataset. Based on the interpretation rules related to sleep in clinical rule knowledge, we number them to build a corresponding rule base; A pre-trained visual language model is used as the base model, and the base model is fine-tuned and trained in a supervised manner using a dataset to obtain a visual language model. The multi-channel sleep data to be analyzed and the preset prompt words are input into the visual language model to generate a structured JSON result. The JSON result includes sleep stage labels, judgment reason text, and a list of rule numbers for related judgment rules.

[0010] This invention renders multi-channel multi-channel sleep recording signals into two-dimensional images and applies a powerful visual language model to this task. Leveraging its superior multimodal understanding and generation capabilities, it identifies key visual features from the images and generates explanatory text.

[0011] Specifically, the multi-channel sleep data includes electroencephalogram (EEG) channels, electrooculogram (EOG) channels, and electromyogram (EMG) channels.

[0012] Specifically, the waveform image is plotted as a waveform image by aligning multi-channel sleep data along the time axis to draw it on the same chart, and rendered using a fixed amplitude range.

[0013] Specifically, the fixed amplitude range of the rendering is: EEG / EOG ±50 μV, EMG ±40 μV.

[0014] Specifically, the image sequence selects the target frame and its adjacent previous and next frames at 30-second intervals.

[0015] Specifically, the list of referenced rule numbers is obtained from the American Academy of Sleep Medicine's Handbook for Interpreting Sleep and Related Events, which includes interpretation rules for the wakefulness phase, non-rapid eye movement (NREM) stage 1, NREM stage 2, NREM stage 3, and rapid eye movement (REM) phase.

[0016] Specifically, the visual language model is output in a structured JSON format.

[0017] Specifically, the system prompts include role and task definitions, image rendering parameter descriptions, clinical sleep staging rule base, input data descriptions, tasks and instructions, output format requirements, and examples.

[0018] To achieve the second objective of this invention, the following technical solution is provided: Steps for performing the above-described interpretable sleep staging method based on supervised fine-tuning training using a visual language model include: The data acquisition module is used to acquire multi-channel sleep data; The preprocessing module is used to preprocess multi-channel sleep data; The image rendering module renders the preprocessed multi-channel sleep data into a two-dimensional waveform image within a fixed amplitude range, and outputs an image sequence consisting of the target frame and its adjacent preceding / following frames. The inference module, based on the input image sequence and preset prompts, calls a supervised and fine-tuned visual language model to generate structured JSON results; The results management and presentation module parses and presents the structured JSON results.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: By outputting detailed textual reasons for the judgment and the cited clinical rule numbers, the model's decision-making process becomes transparent, greatly enhancing clinicians' trust in the automated staging results. By constructing a comprehensive set of system prompts, the American Academy of Sleep Medicine's clinical interpretation criteria are explicitly and structurally encoded into the model. This ensures that the model's reasoning logic is consistent with recognized clinical guidelines, avoiding the learning that might occur with a purely data-driven model and not conform to clinical logic. By rendering polysomnography signals into images with fixed scales, not only is all the information of the signal preserved, but it is also transformed into a form that is more suitable for processing by modern powerful visual models. By using visual language models to process these images, it is possible to better capture complex visual features such as the shape, amplitude, frequency and duration of key events such as K-complex waves, spindle waves and slow waves, which has potential performance advantages compared with traditional methods. Attached Figure Description

[0020] Figure 1 This is a flowchart of the interpretable sleep staging method provided in this embodiment; Figure 2 This is a schematic diagram of the rendered multi-channel sleep data provided in this embodiment; Figure 3 This is a schematic diagram of a dataset sample provided in this embodiment; Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0022] The method provided in this embodiment transforms the traditional one-dimensional polysomnography time-series signal analysis task into a multimodal image-to-text generation task. By rendering the polysomnography signal waveform into an information-rich image, and leveraging the powerful image and text understanding and reasoning capabilities of the visual language model, it achieves accurate judgment of sleep stages and detailed explanation of the judgment process under the guidance of clinical rules.

[0023] like Figure 1The diagram shown is a schematic representation of the method provided in this embodiment.

[0024] Acquire multi-channel sleep data and corresponding sleep inference text, draw corresponding waveform images from the multi-channel sleep data, and use the waveform images of the target frame and the corresponding preceding and following frames as an image sequence. Label the image sequence according to the sleep stage, and combine the labels, image sequence and sleep inference text into a dataset.

[0025] More specifically, to enable the visual language model to understand polysomnography signals, this invention converts the original multi-channel time-series signals into two-dimensional images. This step transforms abstract physiological electrical signals into visual patterns that the model can process.

[0026] Based on the recommendations of the American Academy of Sleep Medicine, six signal channels crucial for sleep stage identification were selected for rendering, including three EEG channels (F4-M1, C4-M1, O2-M1), two OEO channels (LOC, ROC), and one MEG channel (ChinEMG). These six channels reflect activity in the frontal lobe, central region, occipital lobe, left and right eye movements, and mental muscle tone levels, respectively, and are essential signal sources for identifying key features of each sleep stage.

[0027] Before rendering, the raw signal is preprocessed to eliminate noise and standardize the format. Bandpass filters are applied according to the characteristics of different signals, notch filters are applied to all channels to eliminate power frequency interference, and all channels are resampled uniformly. The continuous signal is divided into standard 30-second frames, with each frame corresponding to one image to be rendered.

[0028] The core technical feature lies in the use of fixed amplitude range rendering: the display amplitude range for EEG and EOS channels is fixed at ±50 microvolts, and for EMG channels at ±40 microvolts. This means that the vertical display area of ​​each channel on the image corresponds to a constant voltage range. This design allows the model to directly and quantitatively estimate the true voltage amplitude through the visual height of the waveform, which is crucial for identifying key events such as slow waves.

[0029] The image uses a black background to maximize the signal-to-noise ratio, and finely drawn temporal grid lines facilitate the model's time localization and frequency estimation. The six channels are arranged from top to bottom, and each channel is assigned a unique high-contrast color to help the model distinguish between different channels.

[0030] like Figure 2The image shown is a rendered schematic diagram of the waveform image provided in this embodiment, taking a duration of 30 seconds as an example. This image corresponds to the N2 sleep stage. The waveforms of six physiological signal channels are displayed from top to bottom: F4-M1, C4-M1, O2-M1, LOC, ROC, and Chin EMG. The image background is black with white grid lines. The vertical display range of each channel is normalized, with ±50 microvolts for EEG and EOS channels and ±40 microvolts for EMG channels.

[0031] We used publicly available or acquired polysomnography datasets, which contain a large number of high-quality overnight polysomnography recordings and expert-annotated sleep staging results. We processed the polysomnography data from all subjects in the dataset, generating corresponding images for each 30-second frame.

[0032] like Figure 3 As shown, each training sample simulates a complete dialogue with the model, consisting of the following parts: System prompts, i.e., the fixed instructions constructed above; The user message contains an image sequence, namely three images of the target frame and its context; The assistant message is the standard answer that the model needs to learn to generate. It is a JSON object in which the sleep stage label comes from expert annotations in the dataset, while the judgment reason text and the reference rule number are constructed based on the correspondence between expert annotations and clinical rules.

[0033] The core technical feature is that the dataset partitioning must be carried out at the subject level to ensure that all data of the same person belongs to the same set, thereby avoiding data leakage and ensuring the objectivity of model evaluation.

[0034] We will number the sleep-related interpretation rules in clinical rule knowledge to build a corresponding rule base.

[0035] To enable the model to think like an expert, this embodiment constructs a comprehensive set of system prompts, injecting clinical knowledge into the model.

[0036] First, based on the latest Sleep and Related Events Interpretation Manual from the American Academy of Sleep Medicine, all the interpretation rules for the five sleep stages—W, N1, N2, N3, and R—were summarized and extracted, and each rule was clearly numbered. For example: the wakefulness stage rule includes the presence of occipital alpha rhythm and active eye movements; the N2 stage rule includes the presence of K-complex waves and sleep spindle waves; the N3 stage rule includes slow-wave activity exceeding 20%; and the REM stage rule includes the presence of rapid eye movements and the lowest level of mental muscle tone.

[0037] System prompts are carefully crafted text instructions that serve as fixed commands for interacting with the model. Their structure includes: Role and task definition: The instruction model plays the role of a senior sleep technician, whose task is to analyze a sequence of three consecutive frames and to stage and interpret the target frame in the middle. Image rendering parameter description: Clearly specify the rendering standards for the model image, especially the fixed amplitude scale, channel order, and color mapping; Clinical sleep staging rule base: This base lists all the numbered clinical rules summarized above, which must be referenced and cited when interpreting the model. Input data format description: Inform the model that the input will be three images, representing the preceding frame, the target frame, and the following frame; Task instructions: The model is required to carefully analyze the context of the three images, identify key waveforms and events in the target frame, reference applicable rule numbers from the rule base, generate detailed judgment reasons, and finally give a sleep staging conclusion; Output format requirements: The model must output results in a structured JSON format, including three fields: the text of the judgment reason, the list of referenced rule numbers, and the sleep stage label; Example: Provide high-quality input and output examples to demonstrate the expected reasoning process and output format to the model.

[0038] The core technical feature lies in explicitly encoding clinical rules as system prompts, rather than having the model learn them implicitly from data. This ensures that the model's reasoning logic is consistent with accepted clinical guidelines, avoiding the learning that might occur with a purely data-driven model and be inconsistent with clinical logic.

[0039] A basic model is constructed based on the rule base and preset system prompts, and the basic model is then trained under supervision using a dataset to obtain a visual language model.

[0040] This embodiment uses supervised fine-tuning to adapt a general, pre-trained visual language model to the specific task of sleep staging.

[0041] Choose a high-performance open-source visual language model as the base model. These models have been pre-trained on large-scale image and text data and have powerful capabilities in visual feature extraction, language understanding, and text generation.

[0042] The basic visual language model is fine-tuned using a prepared training dataset. During training, the model receives system prompts and user messages and is instructed to generate outputs that are as consistent as possible with the assistant's messages. The goal of training is to minimize the difference between the model's generated content and the standard answer, typically using the cross-entropy loss function.

[0043] During training, the classification accuracy on the validation set is monitored in real time to prevent model overfitting, and the best-performing model version is saved.

[0044] Once the model is trained, it can be used to perform interpretable sleep staging on new polysomnography data.

[0045] For any 30-second frame to be analyzed, it and the two frames before and after it are first rendered into three images. Then, these three images, along with a fixed system prompt, are fed as input to the fine-tuned model.

[0046] After receiving input, the model's internal visual encoder analyzes the image content, while the language model performs inference under the guidance of system prompts, ultimately generating a JSON-formatted text output.

[0047] The sleep stages, rationale, and cited clinical rules are parsed from the JSON string output by the model. This information can be integrated into the user interface and presented to clinicians in a clear and intuitive way.

[0048] This embodiment also provides an interpretable sleep staging device for performing the steps of the interpretable sleep staging method based on visual language model-supervised fine-tuning training provided in the above embodiments, including: The data acquisition module is used to acquire multi-channel sleep data; The preprocessing module is used to preprocess multi-channel sleep data; The image rendering module renders the preprocessed multi-channel sleep data into a two-dimensional waveform image within a fixed amplitude range, and outputs an image sequence consisting of the target frame and its adjacent preceding / following frames. The inference module, based on the input image sequence and preset prompts, calls a supervised and fine-tuned visual language model to generate structured JSON results; The results management and presentation module parses and presents the structured JSON results.

[0049] To better illustrate the process of the method provided in this embodiment, overnight polysomnography data and corresponding expert annotation files of 62 subjects were obtained from the MASS-SS3 dataset. This dataset stores physiological signals in EDF+ format and sleep stage annotations in XML format.

[0050] The EDF+ file was read using a signal processing library, and the six target channel signals required for this invention were extracted according to the channel definitions in the annotation file. For EEG signals, channels named "EEG F4-CLE", "EEG C4-CLE", and "EEG O2-CLE" were selected. For EOS signals, channels named "EOG Left Horiz" and "EOG Right Horiz" were selected. For EMG signals, the EMG signal of the mentalis muscle was calculated by differential calculation using two electrodes ("EMG Chin1" and "EMG Chin2").

[0051] A preprocessing procedure was applied to the extracted 6-channel signals. The EEG and EOS signals were bandpass filtered from 0.3 to 35 Hz, the EMG signals were bandpass filtered from 10 to 100 Hz, and all channels were notched with a 50 Hz filter to eliminate power frequency interference. All channels were then resampled to 100 Hz.

[0052] The preprocessed continuous signal is segmented into 30-second frames. An image rendering algorithm is implemented to render each frame as a 448×224 pixel PNG image. The following parameters are used during rendering: the image background is black; the six channels are arranged from top to bottom and drawn using yellow, green, red, cyan, magenta, and blue respectively; the display amplitude range of the EEG and EOS channels is fixed at ±50 microvolts, and the EMG channel is fixed at ±40 microvolts; time grid lines are drawn, with one thin gray line per second and one thick gray line every 5 seconds; horizontal lines are drawn to separate the channels.

[0053] Based on the American Academy of Sleep Medicine's Manual of Sleep and Related Events, 3rd Edition, interpretation rules for each stage (W, N1, N2, N3, and R) were compiled and numbered. A complete system cue word file was written, which defines in detail the model's roles, tasks, input / output formats, and includes a complete clinical rule base.

[0054] The system prompts the model to act as a senior sleep technician, tasked with analyzing three consecutive 30-second polysomnography (PSG) frame sequences. The image rendering parameter description clearly informs the model of key information such as the fixed amplitude scale, channel color correspondence, and channel order. The clinical sleep staging rule base comprehensively lists the interpretation rules and numbers for each sleep stage. For example, W.1 indicates that occipital alpha rhythms are visible for more than 50% of the frame time; N2.1 indicates the presence of at least one K-complex; N2.2 indicates the presence of at least one sleep spindle; N3.1 indicates that more than 20% of the frame time is occupied by slow-wave activity with a frequency of 0.5–2 Hz and a peak-to-peak value greater than 75 μV; R.1 indicates the presence of rapid eye movements; and R.2 indicates that mental muscle tone is at its lowest level among all sleep stages.

[0055] The system prompts and task instructions require the model to carefully analyze the context, identify key features, reference rules, and generate judgment criteria. The output format requirements specify JSON output, including a `reasoning_text` field storing the judgment reasoning text, an `applicable_rules` field storing the list of referenced rule numbers, and a `sleep_stage` field storing the sleep stage label. The example section provides high-quality input and output examples to demonstrate the expected reasoning process and output format to the model.

[0056] Construct training samples. For each valid frame of each subject, i.e., a frame with preceding and following frames, construct a training sample. For the target frame N, locate its corresponding image N.png, as well as the preceding image N-1.png and the following image N+1.png, and use the paths of these three images as user messages.

[0057] The helper message is constructed, representing the standard answer the model needs to learn to generate. The helper message is a JSON object where the `sleep_stage` field value comes from expert annotations in the dataset. The `reasoning_text` field value is constructed based on the correspondence between sleep stages and clinical rules in the expert annotations; for example, for stage N2, the constructed reasoning text would describe the observation of K-complexes and sleep spindles. The `applicable_rules` field value maps the features mentioned in the reasoning to the corresponding rule numbers; for example, stage N2 would include N2.1 and N2.2.

[0058] System prompts, user messages, and assistant messages are combined to form training samples. Each sample contains a complete dialogue structure.

[0059] The 62 participants were randomly divided into a training set of 38 participants, a validation set of 12 participants, and a test set of 12 participants. This division was ensured to be done at the participant level, meaning all frames from the same person belonged to the same set, to prevent the model from accessing any information from the test set participants during training.

[0060] Qwen2.5-VL-3B-Instruct was chosen as the base visual language model. This model has been pre-trained on large-scale image and text data and possesses powerful capabilities in visual feature extraction, language understanding, and text generation.

[0061] The base visual language model is fine-tuned with full parameter supervision using a prepared training dataset. During training, the model receives system prompts and user messages and is instructed to generate outputs that are as consistent as possible with the assistant's messages. The goal of training is to minimize the cross-entropy loss between the model's generated content and the standard answer.

[0062] The training hyperparameters are set as follows: learning rate 1e-5, batch size per device 1, gradient accumulation steps 16, training epochs 10, weight decay 0.1, warm-up ratio 0.03, maximum gradient norm 1.0, optimizer AdamW, and a learning rate scheduling strategy of linear warm-up plus linear decay.

[0063] During training, the sleep staging accuracy of the model is evaluated on the validation set every 400 update steps. The model checkpoint with the highest accuracy on the validation set is saved as the final best model.

[0064] Training was conducted on a server equipped with four NVIDIA H100 (80GB) GPUs.

[0065] After training, performance is evaluated on a separate test set. For each frame in the test set, an inference instance is constructed, containing the system prompt word and three images, and the trained model is invoked to generate JSON output.

[0066] The model's output JSON is parsed to extract the predicted sleep_stage field value, which is then compared with the expert-annotated ground truth labels. Based on this, the model's overall accuracy, macro-average F1 score, Kappa coefficient, and F1 score for each sleep stage are calculated.

[0067] To verify the effectiveness of the method of this invention, comparative experiments were designed. Comparison Method 1 uses the same rendered image but employs a ResNet-18 convolutional neural network for classification; this method does not output a reason for the judgment. Comparison Method 2 uses the original signal and employs the classic DeepSleepNet deep learning network for classification; this method also does not output a reason for the judgment. All methods use the same dataset and data partitioning. The experimental results are shown in Table 1 below. The experimental results show that the method of this invention achieved an overall sleep stage accuracy of 87.83% on the MASS-SS3 test set, with a macro-average F1 score of 0.8192 and a Kappa coefficient of 0.8126, which is comparable to the performance of existing non-interpretable methods. The F1 scores for each sleep stage are as follows: W stage 0.9158, N1 stage 0.5184, N2 stage 0.9202, N3 stage 0.8570, and R stage 0.8847.

[0068] The core advantage of this invention lies in its interpretability. The model not only outputs sleep stage classification results but also detailed reasoning texts and cited clinical rule numbers—a feature completely absent in comparative methods. This interpretability allows the invention to maintain high-accuracy staging while providing reliable decision support for clinical applications.

[0069] Based on the above method, an interpretable sleep staging device was implemented. This device is built on a high-performance workstation, with hardware configuration including an Intel(R) Xeon(R) Gold 6330 processor, 1 TB DDR4 memory, and an NVIDIA GeForce A800 graphics card. The operating system is Ubuntu 20.04 LTS.

[0070] The device is equipped with all the necessary software and libraries, including Python 3.10, PyTorch 2.6, Transformers library, MNE signal processing library, Matplotlib plotting library, and a pre-trained sleep staging model.

[0071] The device interacts with users via a web interface. Users access the device's web service through a browser and upload a standard polysomnography (PSG) data file. Upon receiving the file, the backend service automatically triggers the processing flow. The data acquisition module calls the signal processing library to read the file, while the preprocessing and image rendering modules execute sequentially, converting the overnight data into a series of PNG images and temporarily storing them. The inference module invokes a finely tuned visual language model via a sliding window, outputting a JSON result for each frame. The results management and presentation module collects the JSON results from all frames and integrates them into an interactive visualization report.

[0072] The report presents the sleep structure throughout the night in the form of a sleep stage diagram. When the user hovers the mouse over or clicks on any frame on the diagram, the interface dynamically displays the polysomnography waveform image of that frame, the staging results given by the model, detailed textual explanations of the judgment, and the cited clinical rule number.

[0073] This device provides clinicians with a one-stop solution from raw data to interpretable reports, greatly simplifying the process and presenting automated analysis results in an intuitive and credible manner.

Claims

1. An interpretable sleep staging method based on supervised fine-tuning training using a visual language model, characterized in that, Includes the following steps: Acquire multi-channel sleep data and corresponding sleep inference text, draw corresponding waveform images from the multi-channel sleep data, and use the waveform images of the target frame and the corresponding preceding and following frames as an image sequence, and label the image sequence according to the sleep stage, and combine the labels, image sequence and sleep inference text into a dataset. Based on the interpretation rules related to sleep in clinical rule knowledge, we number them to build a corresponding rule base; A pre-trained visual language model is used as the base model, and the base model is fine-tuned and trained in a supervised manner using a dataset to obtain a visual language model. The multi-channel sleep data to be analyzed and the preset prompt words are input into the visual language model to generate a structured JSON result. The JSON result includes sleep stage labels, judgment reason text, and a list of rule numbers for related judgment rules.

2. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model as described in claim 1, characterized in that, The multichannel sleep data includes EEG channels, EOG channels, and EMG channels.

3. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model according to claim 1, characterized in that, The waveform image is obtained by aligning multi-channel sleep data along the time axis to plot it on the same chart, and is rendered using a fixed amplitude range.

4. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model as described in claim 3, characterized in that, The fixed amplitude range of the rendering is: EEG / EOG ±50 μV, EMG ±40 μV.

5. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model according to claim 1, characterized in that, The image sequence selects the target frame and its adjacent preceding and following frames at 30-second intervals.

6. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model according to claim 1, characterized in that, The list of referenced rule numbers was obtained from the American Academy of Sleep Medicine's Handbook for the Interpretation of Sleep and Related Events, and includes interpretation rules for the wakefulness phase, non-rapid eye movement (NREM) stage 1, NREM stage 2, NREM stage 3, and rapid eye movement (REM) phase.

7. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model according to claim 1, characterized in that, The visual language model is output in structured JSON format.

8. The interpretable sleep staging method based on supervised fine-tuning training using a visual language model according to claim 1, characterized in that, The system prompts include role and task definitions, image rendering parameter descriptions, clinical sleep staging rule base, input data descriptions, tasks and instructions, output format requirements, and examples.

9. A sleep staging device that can explain sleep stages, characterized in that, The steps for performing the interpretable sleep staging method based on supervised fine-tuning training using a visual language model as described in any one of claims 1 to 8 include: The data acquisition module is used to acquire multi-channel sleep data; The preprocessing module is used to preprocess multi-channel sleep data; The image rendering module renders the preprocessed multi-channel sleep data into a two-dimensional waveform image within a fixed amplitude range, and outputs an image sequence consisting of the target frame and its adjacent preceding / following frames. The inference module, based on the input image sequence and preset prompts, calls a supervised and fine-tuned visual language model to generate structured JSON results; The results management and presentation module parses and presents the structured JSON results.

Citation Information

Patent Citations

  • Sleep stage detection method based on multi-mode and visual transformation network

    CN119548095A

  • Sleep staging method based on spatial-temporal feature coding and multi-source fusion

    CN120217200A