Neural electrophysiological signal analysis method based on music perception

By preprocessing and feature extraction of music and neural signals, combined with wavelet transform and deep learning, the problems of insufficient music signal modeling and low synchronization accuracy in existing technologies are solved, achieving high-precision neurophysiological signal analysis, which is suitable for multi-task emotion recognition and personalized extension.

CN120837091APending Publication Date: 2025-10-28WUHAN CONSERVATORY OF MUSIC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510923564.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing methods lack sufficient dimensions in music signal modeling, fail to capture high-order features such as rhythm, melody, and emotional structure, make it difficult to accurately synchronize neural signals with music events, employ crude feature extraction methods that are difficult to generalize in multi-frequency and multi-channel scenarios, and lack a unified and efficient signal processing and analysis workflow.

Method used

By combining music signal preprocessing, neural signal preprocessing, music event synchronization, and time-frequency feature extraction, and integrating wavelet transform and deep learning, a unified neurophysiological signal analysis method is constructed. This method includes music signal sampling, framing and windowing, short-time Fourier transform, feature extraction, bandpass filtering, independent component analysis, standardization, multi-band energy calculation, and feature fusion. Convolutional neural networks and bidirectional long short-term memory networks are used for recognition.

Benefits of technology

It improves the synchronization accuracy of music signals and neural signals, enhances the ability to extract multi-band features, and improves the ability to distinguish and model cognitive states. It is suitable for multi-task and multi-category emotion recognition, has good versatility and engineering applicability, and supports multimodal adaptation and individualized expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120837091A_ABST
    Figure CN120837091A_ABST
Patent Text Reader

Abstract

The invention provides a neural electrophysiological signal analysis method based on music perception. The neural electrophysiological signal analysis method comprises the following steps of (S1) music signal preprocessing, (S2) neural signal preprocessing, (S3) music event synchronization, (S4) time-frequency feature extraction and (S5) feature construction and recognition. The method has the advantages that the synchronization precision of music signals and neural signals is improved, the multi-band feature extraction capability of the neural signals is enhanced, multi-band electroencephalogram energy is combined with music high-dimensional features for joint modeling, a unified perception-nerve-recognition analysis closed-loop path is constructed, the expandability is high, and the method is easy to implement. And multi-modal adaptation and individualized expansion can be satisfied.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of music perception and neurocognitive research technology, specifically to a method for analyzing neuroelectrophysiological signals in intelligent systems such as neurocognitive research, music therapy, emotion recognition, and brain-computer interfaces. Background Technology

[0002] With the development of neuroscience, brain-computer interfaces, and artificial intelligence, music, as a complex time-emotional stimulus, has been widely used in cognitive function research and neurorehabilitation. Neurophysiological signals (such as EEG, ECG, and EDA) can reflect the neural responses generated by humans during music perception, and have extremely high research and application value.

[0003] However, existing methods have the following technical problems: insufficient dimensions in music signal modeling, failing to capture high-order features such as rhythm, melody, and emotional structure; difficulty in accurately synchronizing neural signals with music events, affecting stimulus-response pattern recognition; crude feature extraction methods, making it difficult to generalize in multi-band and multi-channel scenarios; and a lack of unified and efficient signal processing and analysis workflows. Summary of the Invention

[0004] The purpose of this invention is to provide a complete process for music perception-based neurophysiological signal analysis, from data preprocessing and feature modeling to response construction. Accordingly, in this invention, "channel" refers to the recording path corresponding to each electrode when acquiring neural electrical signals (such as EEG). Typically, EEG headsets have multiple electrodes arranged on the scalp (e.g., 8 channels, 32 channels, 64 channels, 128 channels, etc.). "Frequency band" refers to several sub-bands divided in the frequency domain of neural signals, with each frequency band corresponding to a specific physiological / psychological state of the brain.

[0005] Specifically, the present invention is achieved through the following technical solutions: A method for analyzing neurophysiological signals based on music perception includes the following steps: S1) Music signal preprocessing The original music signal is preprocessed by discretization, framing, time-frequency conversion and feature extraction to provide basic stimulus information for subsequent event synchronization of neural signals.

[0006] S11) Music signal sampling The analog audio signal is converted into a discrete sequence, providing a foundation for subsequent digital processing. The formula is x(t) = x[{n{T}_{s}}], n = 0, 1, ..., N-1 ,in, The discretized music signal values, Where n is the sampling time interval, n is the index of the sampling point, and N is the total number of sampling points.

[0007] S12) Frame division and windowing The discrete music signal is divided into several frames according to a time window, and windowing is applied to reduce boundary effects. The formula is as follows: ,in, Let T be the signal of the nth frame, and T be the frame shift (the offset between frames). The function is the window function (Hamming window), and L is the frame length.

[0008] S13) Short-Time Fourier Transform (STFT) The time-frequency analysis method is used to map each frame of the music signal to the frequency domain, capturing instantaneous frequency components. The formula is as follows: ,in, Let f be the spectrum of the nth frame at frequency f. The sampling rate.

[0009] S14) Feature Extraction Higher-order perceptual features of music are extracted and used with MFCC and rhythmic spectra for subsequent stimulus matching and synchronization modeling. The formula is as follows: ,in, The cepstral coefficients of the k-th order Mel frequency are... Let i be the center frequency of the i-th Mel filter. Let be the i-th Mel filter, and DCT be the Discrete Cosine Transform.

[0010] S2) Neural signal preprocessing The acquired raw neural signals are preprocessed by filtering, artifact removal, and normalization to eliminate noise interference and unify the dimensions, thereby improving the accuracy of subsequent feature extraction.

[0011] S21) Neural signal acquisition Real-time acquisition of neural electrical signals from multiple channels on the subject's scalp forms a multidimensional time series, as shown in the formula. ,in, denoted as the raw EEG signal of the i-th channel, and N is the total number of channels.

[0012] S22) Bandpass filter Filters are used to suppress power line interference and non-brain-derived noise, retaining only the dominant EEG frequency components within the 1-40Hz range. The formula is as follows: Where BPF is the bandpass filter function. For low and high cutoff frequencies (such as 1Hz and 40Hz).

[0013] S23) ICA Independent Component Analysis Artifact removal is achieved by extracting neural components unrelated to electrooculography (EOG) and electromyography (EMG) using a source signal separation method. The formula is as follows: ,in, Let A be the observed signal matrix, and A be the mixing matrix. The source signal matrix, This is the unmixing matrix.

[0014] S24) Standardization Processing The signal is converted to a standard normal distribution with a mean of 0 and a standard deviation of 1, and the numerical scale of each channel is standardized. The formula is as follows: ,in, Let be the mean of the i-th channel. The standard deviation is denoted as .

[0015] S3) Music Event Synchronization Based on algorithms such as spectral energy change and beat detection, the system automatically identifies music time nodes related to rhythm or emotion and aligns them with the timeline of neural signals to construct a "stimulus-response" analysis window.

[0016] S31) Music Event Extraction Based on algorithms such as spectral energy change and beat detection, music time nodes related to rhythm or emotion are automatically identified. The formula is as follows: ,in, Let k be the time point of the music event.

[0017] S32) Synchronization Window Expand the window before and after each music event point, extract neural signals as analysis samples, and establish temporal alignment. The formula is: {W}_{k}=\left [ {{t}_{k}-\Delta ,{t}_{k}+\Delta} \right ] ,in, Half width of the time window , representing the neural response segment under the k-th music event.

[0018] S4) Time-frequency feature extraction Multi-band power energy features are extracted from synchronized neural signal segments to characterize the brain's response patterns to musical stimulation at different rhythmic levels.

[0019] S41) Wavelet Transform The time signal is decomposed into different scales using continuous wavelet transform to characterize the intensity of multi-frequency activity. The formula is as follows: ,in, For continuous wavelet coefficients across the entire scale and time, 'a' represents the scale inversely proportional to the frequency, and 'b' represents the time shift. The parent wavelet is represented by *, and the complex conjugate is represented by *. This represents the original neural signal of the i-th channel.

[0020] S42) Band Energy Calculation Integrating the wavelet transform result within each specific frequency band quantifies the intensity of each rhythmic component; the formula is as follows: ,in, For channel i in frequency band energy, For use in the i-band Frequency band energy analysis, \left [ {{t}_{k}-\Delta ,{t}_{k}+\Delta} \right ] Let db be the neural response time window for the k-th music event, and db be the time window for processing wavelet coefficients. The points.

[0021] S5) Feature Construction and Recognition Finally, the extracted feature vectors are input into the classifier to identify the emotion or cognitive category to which the current neural state belongs, and it can support modeling and training of various model architectures.

[0022] S51) Construction of neural feature vectors Integrating all channels In a specific frequency band Based on the wavelet energy integral results, a channel-band two-dimensional energy matrix is ​​constructed and compressed and mapped into a unified vector, with the formula: {v}_{k}=\left [ {{E}^{\theta}_{1}+{E}^{\alpha}_{1}+{E}^{\beta}_{1},.,{E}^{\gamma}_{1},...,{E}^{\gamma}_{N}} \right ] ,in, This is the multi-channel energy feature vector corresponding to the k-th music event. For the i-th channel in the frequency band The energy on the frequency band is N, where N is the number of channels and F is the number of frequency bands. This vector serves as the response representation of the nervous system under the k-th music event and is input into the subsequent recognition network. It is encoded into spatial features by a convolutional neural network (CNN), as shown in the formula. ; S52) Construction of Music Feature Vectors We introduce perceptual feature vectors and co-model them with neural signal features, using the formula: {m}_{k}=\left [ {{r}_{k},{p}_{k},{t}_{k},{e}_{k},{s}_{k},{v}_{k},{d}_{k}} \right ] ,in Let be the music perception feature vector of the k-th music event. Rhythmic characteristics represent the number of beats per unit of time and the regularity of the rhythm. The pitch characteristics represent the frequency distribution of the main melody and its curve variations. Tonal characteristics are measured using MFCC coefficients to assess sound quality. This represents the sound pressure level within each frame's time window, as a volume energy characteristic. Meter intensity characteristics indicate the degree of accent misalignment and syllable density. The rate of change of pitch represents how quickly pitch changes and the amplitude of the fluctuation. The dynamic feature represents the temporal variation of volume; it is then encoded into a higher-order feature by a fully connected network (FC), as shown in the formula. ; S53) Multimodal fusion The joint vector is obtained by concatenating the feature vectors of the two modes, as shown in the formula: {u}_{k}=\left [ {{h}_{v},{h}_{m}} \right ]\in {R}^{{d}_{v}+{d}_{m}} An attention mechanism is introduced to adjust the weight distribution of the feature dimensions, as shown in the formula. ;in, For attention weights, , representing attention score, The value of the j-th dimension of the joint feature vector (from the features concatenated from neural and musical data). W is the learnable weight matrix, and b is the learnable bias term. The projection vector for attention determines whether the current dimension is "activated". Rate the importance of the j-th dimension; the higher the value, the more important the dimension.

[0023] S54) Timing Modeling The weighted joint features are input into a bidirectional long short-term memory (Bi-LSTM) network to capture the neural response context patterns before and after musical stimuli, as shown in the formula. .

[0024] S55) Confidence Calculation The distribution of the output class is predicted using softmax, with the formula: {^{∧}_{{P}_{k}}}=softmax\left ( {{o}_{k}} \right )=\left [ {{p}_{k,1},{p}_{k,2}...,{p}_{k,c}} \right ] The confidence level of the current identification is then defined as the maximum probability value, as shown in the formula: ;in, Let be the classification probability vector of the k-th sample. Let k be the predicted probability that the k-th sample belongs to the i-th emotion / cognitive category, satisfying C represents the total number of categories (including pleasure, relaxation, tension, panic, frustration, depression, sadness, irritability, anxiety, and other emotions). The maximum class confidence level is used as the overall judgment criterion.

[0025] S56) Judgment, Identification and Feedback The final identified label is the category with the highest predicted probability, and a confidence threshold is set. (e.g., 0.85), the formula is as follows: If Accept the classification results ;like If the current classification is deemed unreliable, the system automatically triggers a feedback mechanism, returning to step S4 to re-extract the time-frequency features of the neural signal and construct a new one. Then, the identification process is repeated.

[0026] To prevent entering an infinite loop, set a maximum number of retries. (3 times), after which the current result is output and marked as "low confidence".

[0027] Compared with the prior art, the present invention has the following advantages: This invention presents a neurophysiological signal analysis method based on music perception, which improves the synchronization accuracy between music signals and neural signals, enhances the multi-band feature extraction capability of neural signals, and employs joint modeling of multi-band EEG energy and high-dimensional music features to construct a unified closed-loop path for perception-neural-recognition analysis. This method is highly scalable and can meet the requirements of multimodal adaptation and individualized expansion. Specifically, through precise identification of music events such as beat points and pitch changes, it achieves time window alignment of neural responses, significantly improving the stimulus-response coupling quality. Furthermore, it refines the analysis of music signals based on wavelet transform combined with bandpass filtering. Energy extraction and analysis across multiple frequency bands improves the ability to distinguish differences in cognitive states. The construction of feature vector combinations followed by deep learning classifier modeling enhances the ability to model complex cognitive states, making it suitable for multi-task, multi-category emotion recognition applications. The method standardizes the entire process from music input and neural signal acquisition to synchronous analysis and emotion recognition, demonstrating good versatility and engineering feasibility. This method can be flexibly extended to other modal signals (such as EDA / ECG) and customized with individual music preference parameters, exhibiting strong adaptability and supporting a wide range of user scenarios. Attached Figure Description

[0028] Figure 1 This is a schematic diagram of the analysis method of the present invention. Detailed Implementation

[0029] The embodiments of the present invention will now be described in further detail with reference to the accompanying drawings.

[0030] like Figure 1 As shown, a neurophysiological signal analysis method based on music perception includes the following steps: S1) Music signal preprocessing The original music signal is preprocessed by discretization, framing, time-frequency conversion and feature extraction to provide basic stimulus information for subsequent event synchronization of neural signals.

[0031] S11) Music signal sampling The analog audio signal is converted into a discrete sequence, providing a foundation for subsequent digital processing. The formula is x(t) = x[{n{T}_{s}}], n = 0, 1, ..., N-1 ,in, The discretized music signal values, Where n is the sampling time interval, n is the index of the sampling point, and N is the total number of sampling points.

[0032] S12) Frame division and windowing The discrete music signal is divided into several frames according to a time window, and windowing is applied to reduce boundary effects. The formula is as follows: ,in, Let T be the signal of the nth frame, and T be the frame shift (the offset between frames). The function is the window function (Hamming window), and L is the frame length.

[0033] S13) Short-Time Fourier Transform (STFT) The time-frequency analysis method is used to map each frame of the music signal to the frequency domain, capturing instantaneous frequency components. The formula is as follows: ,in, Let f be the spectrum of the nth frame at frequency f. The sampling rate.

[0034] S14) Feature Extraction Higher-order perceptual features of music are extracted and used with MFCC and rhythmic spectra for subsequent stimulus matching and synchronization modeling. The formula is as follows: ,in, The cepstral coefficients of the k-th order Mel frequency are... Let i be the center frequency of the i-th Mel filter. Let be the i-th Mel filter, and DCT be the Discrete Cosine Transform.

[0035] S2) Neural signal preprocessing The acquired raw neural signals are preprocessed by filtering, artifact removal, and normalization to eliminate noise interference and unify the dimensions, thereby improving the accuracy of subsequent feature extraction.

[0036] S21) Neural signal acquisition Real-time acquisition of neural electrical signals from multiple channels on the subject's scalp forms a multidimensional time series, as shown in the formula. ,in, denoted as the raw EEG signal of the i-th channel, and N is the total number of channels.

[0037] S22) Bandpass filter Filters are used to suppress power line interference and non-brain-derived noise, retaining only the dominant EEG frequency components within the 1-40Hz range. The formula is as follows: Where BPF is the bandpass filter function. For low and high cutoff frequencies (such as 1Hz and 40Hz).

[0038] S23) ICA Independent Component Analysis Artifact removal is achieved by extracting neural components unrelated to electrooculography (EOG) and electromyography (EMG) using a source signal separation method. The formula is as follows: ,in, Let A be the observed signal matrix, and A be the mixing matrix. The source signal matrix, This is the unmixing matrix.

[0039] S24) Standardization Processing The signal is converted to a standard normal distribution with a mean of 0 and a standard deviation of 1, and the numerical scale of each channel is standardized. The formula is as follows: ,in, Let be the mean of the i-th channel. The standard deviation is denoted as .

[0040] S3) Music Event Synchronization Based on algorithms such as spectral energy change and beat detection, the system automatically identifies music time nodes related to rhythm or emotion and aligns them with the timeline of neural signals to construct a "stimulus-response" analysis window.

[0041] S31) Music Event Extraction Based on algorithms such as spectral energy change and beat detection, music time nodes related to rhythm or emotion are automatically identified. The formula is as follows: ,in, Let k be the time point of the music event.

[0042] S32) Synchronization Window Expand the window before and after each music event point, extract neural signals as analysis samples, and establish temporal alignment. The formula is: {W}_{k}=\left [ {{t}_{k}-\Delta ,{t}_{k}+\Delta} \right ] ,in, Half width of the time window , representing the neural response segment under the k-th music event.

[0043] S4) Time-frequency feature extraction Multi-band power energy features are extracted from synchronized neural signal segments to characterize the brain's response patterns to musical stimulation at different rhythmic levels.

[0044] S41) Wavelet Transform The time signal is decomposed into different scales using continuous wavelet transform to characterize the intensity of multi-frequency activity. The formula is as follows: ,in, For continuous wavelet coefficients across the entire scale and time, 'a' represents the scale inversely proportional to the frequency, and 'b' represents the time shift. The parent wavelet is represented by *, and the complex conjugate is represented by *. This represents the original neural signal of the i-th channel.

[0045] S42) Band Energy Calculation Integrating the wavelet transform result within each specific frequency band quantifies the intensity of each rhythmic component; the formula is as follows: ,in, For channel i in frequency band energy, For use in the i-band Frequency band energy analysis, \left [ {{t}_{k}-\Delta ,{t}_{k}+\Delta} \right ] Let db be the neural response time window for the k-th music event, and db be the time window for processing wavelet coefficients. The points.

[0046] S5) Feature Construction and Recognition Finally, the extracted feature vectors are input into the classifier to identify the emotion or cognitive category to which the current neural state belongs, and it can support modeling and training of various model architectures.

[0047] S51) Construction of neural feature vectors Integrating all channels In a specific frequency band Based on the wavelet energy integral results, a channel-band two-dimensional energy matrix is ​​constructed and compressed and mapped into a unified vector, with the formula: {v}_{k}=\left [ {{E}^{\theta}_{1}+{E}^{\alpha}_{1}+{E}^{\beta}_{1},.,{E}^{\gamma}_{1},...,{E}^{\gamma}_{N}} \right ] ,in, This is the multi-channel energy feature vector corresponding to the k-th music event. For the i-th channel in the frequency band The energy on the frequency band is N, where N is the number of channels and F is the number of frequency bands. This vector serves as the response representation of the nervous system under the k-th music event and is input into the subsequent recognition network. It is encoded into spatial features by a convolutional neural network (CNN), as shown in the formula. ; S52) Construction of Music Feature Vectors We introduce perceptual feature vectors and co-model them with neural signal features, using the formula: {m}_{k}=\left [ {{r}_{k},{p}_{k},{t}_{k},{e}_{k},{s}_{k},{v}_{k},{d}_{k}} \right ] ,in Let be the music perception feature vector of the k-th music event. Rhythmic characteristics represent the number of beats per unit of time and the regularity of the rhythm. The pitch characteristics represent the frequency distribution of the main melody and its curve variations. Tonal characteristics are measured using MFCC coefficients to assess sound quality. This represents the sound pressure level within each frame's time window, as a volume energy characteristic. Meter intensity characteristics indicate the degree of accent misalignment and syllable density. The rate of change of pitch represents how quickly pitch changes and the amplitude of the fluctuation. The dynamic feature represents the temporal variation of volume; it is then encoded into a higher-order feature by a fully connected network (FC), as shown in the formula. ; S53) Multimodal fusion The joint vector is obtained by concatenating the feature vectors of the two modes, as shown in the formula: {u}_{k}=\left [ {{h}_{v},{h}_{m}} \right ]\in {R}^{{d}_{v}+{d}_{m}} An attention mechanism is introduced to adjust the weight distribution of the feature dimensions, as shown in the formula. ;in, For attention weights, , representing attention score, The value of the j-th dimension of the joint feature vector (from the features concatenated from neural and musical data). W is the learnable weight matrix, and b is the learnable bias term. The projection vector for attention determines whether the current dimension is "activated". Rate the importance of the j-th dimension; the higher the value, the more important the dimension.

[0048] S54) Timing Modeling The weighted joint features are input into a bidirectional long short-term memory (Bi-LSTM) network to capture the neural response context patterns before and after musical stimuli, as shown in the formula. .

[0049] S55) Confidence Calculation The distribution of the output class is predicted using softmax, with the formula: {^{∧}_{{P}_{k}}}=softmax\left ( {{o}_{k}} \right )=\left [ {{p}_{k1},{p}_{k2}...,{p}_{kc}} \right ] The confidence level of the current identification is then defined as the maximum probability value, as shown in the formula: ;in, Let be the classification probability vector of the k-th sample. Let k be the predicted probability that the k-th sample belongs to the i-th emotion / cognitive category, satisfying C represents the total number of categories (including pleasure, relaxation, tension, panic, frustration, depression, sadness, irritability, anxiety, and other emotions). The maximum class confidence level is used as the overall judgment criterion.

[0050] S56) Judgment, Identification and Feedback The final identified label is the category with the highest predicted probability, and a confidence threshold is set. (e.g., 0.85), the formula is as follows: If Accept the classification results ;like If the current classification is deemed unreliable, the system automatically triggers a feedback mechanism, returning to step S4 to re-extract the time-frequency features of the neural signal and construct a new one. Then, the identification process is repeated.

[0051] To prevent entering an infinite loop, set a maximum number of retries. (3 times), after which the current result is output and marked as "low confidence".

[0052] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the concept of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for analyzing neurophysiological signals based on music perception, characterized in that... The steps include: S1) Music signal preprocessing The original music signal undergoes preprocessing, including discretization, framing, time-frequency conversion, and feature extraction. S2) Neural signal preprocessing The acquired raw neural signals are preprocessed by filtering, artifact removal, and normalization. S3) Music Event Synchronization Based on algorithms such as spectral energy change and beat detection, the system automatically identifies music time nodes related to rhythm or emotion and aligns them with the timeline of neural signals to construct a "stimulus-response" analysis window. S4) Time-frequency feature extraction Multi-band power energy features are extracted from synchronized neural signal segments to characterize the brain's response patterns to musical stimulation at different rhythmic levels. S5) Feature Construction and Recognition S51) Construction of neural feature vectors Integrating all channels In a specific frequency band Based on the wavelet energy integral results, a channel-band two-dimensional energy matrix is ​​constructed and compressed and mapped into a unified vector, as shown in the formula: ,in, This is the multi-channel energy feature vector corresponding to the k-th music event. For the i-th channel in the frequency band The energy on the surface, where N is the number of channels and F is the number of frequency bands; is then encoded into spatial features by a convolutional neural network (CNN), as shown in the formula: ; S52) Construction of Music Feature Vectors Introducing perceptual feature vectors and co-modeling them with neural signal features, the formula is as follows: ,in Let be the music perception feature vector of the k-th music event. As a rhythmic feature, Pitch characteristics, For timbre characteristics, For volume energy characteristics, For beat intensity characteristics, The rate of change of pitch. These are dynamic features; they are then encoded into higher-order features via a fully connected network (FC), as shown in the formula. ; S53) Multimodal fusion The joint vector is obtained by concatenating the feature vectors of the two modes, as shown in the formula. An attention mechanism is introduced to adjust the weight distribution of the feature dimensions, as shown in the formula. ; S54) Classification and Recognition The weighted joint features are input into a bidirectional long short-term memory (Bi-LSTM) network to capture the neural response context patterns before and after musical stimuli, as shown in the formula. ; The distribution of categories is predicted using the softmax output, and the formula is as follows: The formula for defining the confidence level of the current identification as the maximum probability value is as follows: ;in, Let be the classification probability vector of the k-th sample. Let k be the predicted probability that the k-th sample belongs to the i-th emotion / cognitive category, satisfying C represents the total number of categories. This represents the maximum class confidence level.

2. The method for analyzing neurophysiological signals based on music perception according to claim 1, characterized in that: In step S1), the original music signal is first converted into a discrete sequence. Then, the discrete music signal is divided into several frames according to the time window and windowed to reduce the boundary effect. Then, the time-frequency analysis method is used to map each frame of music signal to the frequency domain to capture the instantaneous frequency components. Finally, the higher-order perceptual features of the music are extracted.

3. The method for analyzing neurophysiological signals based on music perception according to claim 2, characterized in that: Step S1) involves extracting higher-order perceptual features of music using MFCC and rhythmic spectrograms, with the following formula: ,in The cepstral coefficients of the k-th order Mel frequency are... Let i be the center frequency of the i-th Mel filter. For the nth frame at frequency The spectrum on For the i-th Mel filter, It is the discrete cosine transform.

4. The method for analyzing neurophysiological signals based on music perception according to claim 1, characterized in that: Step S4 includes: S41) Wavelet Transform The time signal is decomposed into different scales using continuous wavelet transform to characterize the intensity of multi-frequency activity. The formula is as follows: ,in, For continuous wavelet coefficients across the entire scale and time, 'a' represents the scale inversely proportional to the frequency, and 'b' represents the time shift. The parent wavelet is represented by *, and the complex conjugate is represented by *. This represents the original neural signal of the i-th channel; S42) Band Energy Calculation Integrating the wavelet transform result within each specific frequency band quantifies the intensity of each rhythmic component; the formula is as follows: ,in, For channel i in frequency band energy, For use in the i-band Frequency band energy analysis, Let db be the neural response time window for the k-th music event, and db be the time window for processing wavelet coefficients. The points.

5. The method for analyzing neurophysiological signals based on music perception according to claim 1, characterized in that: In step S54), the total number of categories C includes pleasure, relaxation, tension, panic, frustration, depression, sadness, irritability, anxiety, and other emotions.

6. A method for analyzing neurophysiological signals based on music perception according to claim 1 or 4, characterized in that: In step S54), for the maximum class confidence... Set confidence threshold The formula is as follows: If Accept the classification results ;like If the current classification is deemed unreliable, the system automatically triggers a feedback mechanism, returning to step S4 to re-extract the time-frequency features of the neural signal and construct a new one. Then, the identification process is repeated.

Citation Information

Patent Citations

  • Neuropsychological assessment and intervention method, system and device based on music creation

    CN115363587A

  • Electroencephalogram signal emotion recognition method and device based on adaptive fusion network

    CN118452919A

  • Self-adaptive music intervention system based on multi-modal physiological feedback

    CN119499507A

  • Personalized communication system fused with music emotion perception

    CN120108432A

  • Attention state recognition neural network modeling and reasoning method based on electroencephalogram sequence

    CN120123752A