Multi-dimensional timbre perception space model based on electroencephalogram features, modeling method, device and storage medium
By constructing a multidimensional timbre perception spatial model based on EEG features, and combining acoustic and psychological perception features, the problem of quantifying the multidimensional attributes of timbre is solved, realizing a comprehensive representation and differentiated evaluation of timbre auditory perception, and supporting performers' understanding and mastery of timbre expression.
Patent Information
- Application Number
- CN202410646663.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-23
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2044-05-23
AI Technical Summary
Existing technologies lack objective and quantitative dimensions of auditory perception, making it difficult to fully characterize the multidimensional attributes of timbre. In particular, the mechanisms by which the human ear and brain participate in timbre due to its complexity have not been fully understood.
By collecting single-tone samples of musical instruments, acoustic time-frequency domain features and psychological perception features are extracted. Combined with EEG experiments, Euclidean distance is calculated to construct a multi-dimensional timbre perception space model. Using EEG features as a new evaluation dimension, a mapping relationship between timbre acoustic features, psychological perception features and EEG response is established.
It achieves an objective and quantitative assessment of the multidimensional attributes of timbre, integrating the acoustic physical characteristics, psychological perception characteristics, and auditory EEG characteristics of timbre, and provides a more comprehensive method for representing the differences in auditory perception of timbre, thus providing theoretical support for performers to understand and master timbre expression.
Smart Images

Figure CN118568466B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for timbre auditory perception, and particularly to a multidimensional timbre perception spatial model, modeling method, device, and storage medium based on electroencephalogram (EEG) features. Background Technology
[0002] At present, timbre, as an acoustic feature with multidimensional physical attributes in the process of music transmission, is an indispensable part of music auditory perception research. The physical characteristics of music are mainly reflected in pitch (Hz), duration (s), intensity (dB) and timbre. The acoustic physical attributes of timbre mainly include overtones, spectral envelope and harmonic structure [1]. Compared with other elements, the physical attributes of timbre are more complex and cannot be reflected by a single physical parameter. Existing studies have found that there are differences in the auditory perception of different timbres by the human ear. The discrimination of timbre is related to physical parameters such as overtone composition, number of harmonics, harmonic position and relative intensity of harmonics [2][3], but the specific correlation is still unknown. At the same time, some scholars have also conducted research on the psychological perception law of timbre in the field of psychology [4]. For example, the psychological perception of the timbre of common musical instruments is measured, and a characteristic curve of the psychological perception of timbre is constructed; subjective evaluation of different timbres is carried out using descriptive adjectives, and the psychological evaluation index of timbre is visualized and analyzed [5]; a terminology system for evaluating the timbre of musical instruments is proposed, and a subjective evaluation dataset of the timbre of musical instruments is established [6]; the correlation and differences of different timbres under various descriptor sets are explored, and descriptors that can more accurately represent the psychological perception characteristics of timbre are selected [7]. In recent years, with the development of brain imaging technology and neurocognitive science, researchers have found that multiple areas of the brain are involved in the perception of musical features such as pitch, rhythm and melody [8][9]
[10] , while there are relatively few studies on auditory EEG of timbre. This is because, due to the complexity of timbre characteristics, the human ear's perception of timbre involves not only ordinary auditory perception areas such as the temporal region, but also rich higher cognitive attributes.
[0003] Researchers have long attempted to establish a connection between timbre characteristics and auditory perception, and have proposed some methods that combine timbre psychological cognitive attributes and acoustic attributes
[11] . For example, they have compared and analyzed the dissimilarity rating of timbre acoustic features and timbre psychological perception feature descriptors based on triangulation, thereby defining timbre and achieving timbre classification
[12]
[13] . Such studies only start from the two dimensions of physical and psychological perception structure, and lack an objective and quantitative auditory perception dimension. Summary of the Invention
[0004] This invention provides a multidimensional timbre perception spatial model, modeling method, device, and storage medium based on electroencephalogram (EEG) features to solve the technical problems existing in the prior art.
[0005] The technical solution adopted by this invention to solve the technical problems existing in the prior art is as follows:
[0006] A method for modeling a multidimensional timbre perception spatial model based on electroencephalogram (EEG) features, comprising the following steps:
[0007] Step 1: Collect single-timbre samples from various musical instruments and preprocess the collected timbre samples;
[0008] Step 2: Extract acoustic time-frequency domain features from the preprocessed timbre samples, and extract psychological perception features through behavioral psychology experiments;
[0009] Step 3: Use samples of various musical instrument timbres as auditory stimuli to conduct EEG experiments, and extract the corresponding event-related potential (ERP) signals as EEG features.
[0010] Step 4: Use the Euclidean distance between different timbre points to represent the dissimilarity of different timbre features; calculate the distance between different timbre points;
[0011] Step 5: Based on the Euclidean distance matrix between sample timbre points, fit a low-dimensional space with mutually orthogonal dimensions, transform the dissimilarity of timbre feature vectors into the low-dimensional space, directly map the timbre feature vectors into the low-dimensional space to form a point set, and use the scattered points in the low-dimensional space to represent the timbre objects of each instrument to reflect the mapping relationship between the acoustic features, psychological perception features and EEG responses of various instrument timbres; and intuitively display the similarity and dissimilarity of the timbres of each instrument.
[0012] Furthermore, in step 1, the collected timbre samples are processed as follows: unifying audio duration and sampling rate, pre-emphasis, Gammatone filtering, frame segmentation, and windowing.
[0013] Further, in step 2, the following acoustic time-frequency domain features are extracted: low energy features, spectral descent, brightness, spectral irregularity features, spectral centroid features, spectral flatness, spectral skewness, spectral kurtosis, and spectral width; time domain features include zero-crossing rate, start time, steady-state time, and RMS energy.
[0014] Furthermore, in step 2, the method for extracting psychological perception features through behavioral psychology experiments includes the following steps:
[0015] The timbre samples were scored according to the following evaluation parameters: likability, loudness, warmth / coolness, brightness / darkness, soothingness, and fullness.
[0016] Furthermore, in step 3, the EEG experiment adopted the oddball paradigm subtype, with standard and deviated stimuli appearing in a random order; the standard stimulus was a 262Hz sine wave pure tone, with an occurrence probability of 70%; there were 17 deviated stimuli in total; among them, the target stimulus was the sound of a frog; and the distraction stimuli were 16 different instrument timbres, with each deviated stimulus having an equal occurrence probability, accounting for 30% in total.
[0017] Furthermore, in step 5, the formula for calculating the distance between different timbre points is as follows:
[0018]
[0019] In the formula:
[0020] d ij The spatial distance between timbre point i and timbre point j;
[0021] x ir Let i be the coordinates of the timbre point i in the dimension of psychological perception features;
[0022] x jr Let j be the coordinates of the timbre point j in the dimension of psychological perception features;
[0023] y is Let i be the coordinates of the timbre point i in the physical acoustic feature dimension;
[0024] y js Let j be the coordinates of the timbre point j in the physical acoustic feature dimension;
[0025] z it Let i be the coordinates of the timbre point i in the EEG feature dimension;
[0026] z jt Let j be the coordinates of the timbre point j in the EEG feature dimension;
[0027] R1 represents the total number of psychologically perceived characteristics of musical instrument timbre.
[0028] R2 represents the total number of physical acoustic characteristics of the instrument's timbre;
[0029] R3 represents the total number of EEG characteristics of musical instrument timbre;
[0030] w1 represents the weights of the psychological perception features;
[0031] w2 is the weight of the physical acoustic features;
[0032] w3 represents the weights of the EEG features.
[0033] Furthermore, in step 4, the method for reducing the dimensionality of the timbre feature vector using a multi-dimensional scaling transformation method includes the following steps:
[0034] Let D represent the original Euclidean space of the timbre feature vectors;
[0035] Let Q represent the low-dimensional space of the timbre feature vectors after dimensionality reduction;
[0036] Let d a d b These represent two timbre feature vectors in the original Euclidean space D;
[0037] Let q a q b Indicates the corresponding d a d b Two timbre feature vectors in the low-dimensional space Q;
[0038] dist ab Represents two timbre feature vectors d a d b Euclidean distance in primitive Euclidean space;
[0039] Let B denote the inner product matrix of the timbre feature vectors;
[0040] According to D and dist 2 ab Calculate B, perform eigenvalue and eigenvector decomposition on B, and fit and construct the following low-dimensional space Q:
[0041]
[0042] in Let B be a diagonal matrix formed by the largest eigenvalues of matrix B.
[0043] λ k The elements are on the diagonal; where k = 1, 2, ..., d′; d′ is the dimension of the chosen low-dimensional space;
[0044] The eigenvector matrix;
[0045] N is the dimension of the primitive Euclidean space;
[0046] Take the diagonal matrix and eigenvector matrix formed by the maximum values of the eigenvalues in the low-dimensional space Q, such that for any two timbre eigenvectors q in the low-dimensional space Q... a q b The distance is consistent with the distance in the original Euclidean space D, i.e., ||q a -q b ||≈dist ab ;
[0047] The stress coefficient is used as a measure of the goodness of fit in low-dimensional space. The smaller the stress coefficient, the better the model fit.
[0048] The present invention also provides a multidimensional timbre perception space model based on EEG features, which is constructed using the above-mentioned multidimensional timbre perception space model modeling method based on EEG features.
[0049] The present invention also provides an apparatus for a multidimensional timbre perception spatial model modeling method based on EEG features, comprising a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program and, when executing the computer program, implement the steps of the multidimensional timbre perception spatial model modeling method based on EEG features as described above.
[0050] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the multidimensional timbre perception spatial model modeling method based on EEG features as described above.
[0051] The advantages and positive effects of this invention are:
[0052] This invention is based on a comprehensive evaluation of the psychological perception characteristics of musical instrument timbre, the auditory nerve response characteristics of the human ear, and multidimensional acoustic characteristics. It determines the position in each dimension by calculating Euclidean distance and uses a dissimilarity index to standardize individual cases into scores. This invention also refers to previous research results to determine the dimension of the timbre space and uses a stress function to construct objective multidimensional scale characteristics of timbre. Through these works, this invention can characterize the mapping relationship between timbre acoustic characteristics, psychological perception characteristics, and electroencephalogram (EEG) responses.
[0053] This invention incorporates the brain's electroencephalogram (EEG) characteristics under timbre stimulation into a timbre perception space model, effectively filling the gap in the lack of objective and quantitative auditory perception dimensions. Therefore, based on a multidimensional understanding of timbre attributes, this invention effectively integrates the acoustic and psychological characteristics of timbre, and uses auditory EEG characteristics as a new evaluation dimension, realizing the construction of a multidimensional timbre perception space based on EEG characteristics. This establishes a more comprehensive and multidimensional method for differentiated representation of timbre auditory perception.
[0054] This invention also has many applications in actual performance scenarios. Through fundamental research on timbre audio descriptors, it quantitatively describes the perceptual characteristics of instrument sound sources, providing more objective theoretical support for performers to understand and master the expression of timbre. Attached Figure Description
[0055] Figure 1 This is a flowchart of the modeling method for a multidimensional timbre perception spatial model based on EEG features according to the present invention.
[0056] Figure 2 This is a flowchart of a method for extracting acoustic time-frequency domain features according to the present invention.
[0057] Figure 3 This is a diagram illustrating an electroencephalogram (EEG) experimental paradigm of the present invention.
[0058] Figure 4 This is a subjective evaluation data chart of musical instrument timbre according to the present invention.
[0059] Figure 5 This is a line graph of event-related potential (N400) and LPC amplitude data from an electroencephalogram (EEG) experiment according to the present invention.
[0060] Figure 6 This is a bar chart of event-related potentials (N400) and LPC latency data from an electroencephalogram (EEG) experiment according to the present invention.
[0061] Figure 7 This is a schematic diagram of a multidimensional timbre perception spatial model based on EEG features according to the present invention. Detailed Implementation
[0062] The present invention will now be described in detail with reference to the accompanying drawings and embodiments. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0063] The following are the Chinese definitions of English words, phrases, and abbreviations:
[0064] Oddball paradigm: A commonly used experimental paradigm in neurocognitive science.
[0065] RMS energy: Root mean square energy.
[0066] Gammatone filtering: A filter model used to simulate the frequency decomposition characteristics of the cochlea.
[0067] Nation-Instruments: National musical instruments.
[0068] West-Instruments: Western musical instruments.
[0069] Strings: Stringed instruments.
[0070] Winds: wind instruments.
[0071] Bowed Strings: Strings that have been bowed.
[0072] Pitched Percussion: The act of striking the string.
[0073] Plucked Strings: Strings are plucked.
[0074] Woodwinds: wooden pipes.
[0075] Brass: copper pipe.
[0076] Nation-Erhu: Erhu.
[0077] Nation-Dulcimer: Yangqin.
[0078] Nation-Lute: Pipa.
[0079] Nation-Zither: Guzheng.
[0080] Nation-Sheng: Sheng.
[0081] Nation-Xun: Xun.
[0082] Nation-Xiao: Xiao.
[0083] Nation-Suona: Suona.
[0084] Violin: Violin.
[0085] Piano: Piano.
[0086] Guitar: Guitar.
[0087] Harp: Harp.
[0088] Oboe: Oboe.
[0089] Clarinet: Clarinet. <�
[0090] Flute: Flute.
[0091] Trumpet: Trumpet.
[0092] MIRtoolbox 1.8.1: Acoustic feature extraction toolbox.
[0093] Neuroscan Synamps2 system: EEG device amplifier.
[0094] Curry8: EEG data acquisition software.
[0095] EEGLAB: EEG data preprocessing toolbox.
[0096] EEG: Electroencephalogram.
[0097] epochs: Time window.
[0098] Please refer to Figures 1 to 7 , a method for modeling a multi-dimensional timbre perception space model based on EEG features, the method comprising the following method steps:
[0099] Step 1, collect single timbre samples of multiple musical instruments, and preprocess the collected timbre samples;
[0100] Step 2: Extract acoustic time-frequency domain features from the preprocessed timbre samples, and extract psychological perception features through behavioral psychology experiments;
[0101] Step 3: Use samples of various musical instrument timbres as auditory stimuli to conduct EEG experiments, and extract the corresponding event-related potential (ERP) signals as EEG features.
[0102] Step 4: Use the Euclidean distance between different timbre points to represent the dissimilarity of different timbre features; calculate the distance between different timbre points;
[0103] Step 5: Based on the Euclidean distance matrix between sample timbre points, fit a low-dimensional space with mutually orthogonal dimensions, transform the dissimilarity of timbre feature vectors into the low-dimensional space, directly map the timbre feature vectors into the low-dimensional space to form a point set, and use the scattered points in the low-dimensional space to represent the timbre objects of each instrument to reflect the mapping relationship between the acoustic features, psychological perception features and EEG responses of various instrument timbres; and intuitively display the similarity and dissimilarity of the timbres of each instrument.
[0104] Preferably, in step 1, the collected timbre samples can be processed as follows: unify audio duration and sampling rate, pre-emphasize, apply Gammatone filtering, frame segmentation, and windowing.
[0105] Preferably, in step 2, the following acoustic time-frequency domain features can be extracted: low energy features, spectral descent, brightness, spectral irregularity (SI), spectral centroid features, spectral flatness, spectral skewness, spectral kurtosis, and spectral width; the time domain features include zero-crossing rate, start time, steady-state time, and RMS energy.
[0106] Preferably, in step 2, the method for extracting psychological perception features through behavioral psychology experiments may include the following steps:
[0107] The timbre samples were scored according to the following evaluation parameters: likability, loudness, warmth / coolness, brightness / darkness, soothingness, and fullness.
[0108] Preferably, in step 3, the EEG experiment can adopt the oddball paradigm subtype, where the standard stimulus and the deviated stimulus appear in a random order; the standard stimulus can be a 262Hz sine wave pure tone with an occurrence probability of 70%; there are 17 deviated stimuli in total; among them, the target stimulus is the sound of a frog; the distraction stimulus can be 16 different instrument timbres, and each deviated stimulus has an equal occurrence probability, accounting for 30% in total.
[0109] Preferably, in step 5, the formula for calculating the distance between different timbre points can be as follows:
[0110]
[0111] In the formula:
[0112] d ij The spatial distance between timbre point i and timbre point j;
[0113] x ir Let i be the coordinates of the timbre point i in the dimension of psychological perception features;
[0114] x jr Let j be the coordinates of the timbre point j in the dimension of psychological perception features;
[0115] y is Let i be the coordinates of the timbre point i in the physical acoustic feature dimension;
[0116] y js Let j be the coordinates of the timbre point j in the physical acoustic feature dimension;
[0117] z it Let i be the coordinates of the timbre point i in the EEG feature dimension;
[0118] z jt Let j be the coordinates of the timbre point j in the EEG feature dimension;
[0119] R1 represents the total number of psychologically perceived characteristics of musical instrument timbre.
[0120] R2 represents the total number of physical acoustic characteristics of the instrument's timbre;
[0121] r3 represents the total number of EEG characteristics of musical instrument timbre;
[0122] w1 represents the weights of the psychological perception features;
[0123] w2 is the weight of the physical acoustic features;
[0124] w3 represents the weights of the EEG features.
[0125] Preferably, in step 4, the method for reducing the dimensionality of the timbre feature vector using a multi-dimensional scaling transformation method may include the following steps:
[0126] Let D represent the original Euclidean space of the timbre feature vectors;
[0127] Let Q represent the low-dimensional space of the timbre feature vectors after dimensionality reduction;
[0128] Let d a d b These represent two timbre feature vectors in the original Euclidean space D;
[0129] Let q a q b Indicates the corresponding d a d bTwo timbre feature vectors in the low-dimensional space Q;
[0130] dist ab Represents two timbre feature vectors d a d b Euclidean distance in primitive Euclidean space;
[0131] Let B denote the inner product matrix of the timbre feature vectors;
[0132] According to D and dist 2 ab Calculate B, perform eigenvalue and eigenvector decomposition on B, and fit and construct the following low-dimensional space Q:
[0133]
[0134] in Let B be a diagonal matrix formed by the largest eigenvalues of matrix B.
[0135] λ k The elements are on the diagonal; where k = 1, 2, ..., d′; d′ is the dimension of the chosen low-dimensional space;
[0136] The eigenvector matrix;
[0137] N is the dimension of the primitive Euclidean space;
[0138] Take the diagonal matrix and eigenvector matrix formed by the maximum values of the eigenvalues in the low-dimensional space Q, such that for any two timbre eigenvectors q in the low-dimensional space Q... a q b The distance is consistent with the distance in the original Euclidean space D, i.e., ||q a -q b ||≈dist ab ;
[0139] The stress coefficient is used as a measure of the goodness of fit in low-dimensional space. The smaller the stress coefficient, the better the model fit.
[0140] The present invention also provides a multidimensional timbre perception space model based on EEG features, which is constructed using the above-mentioned multidimensional timbre perception space model modeling method based on EEG features.
[0141] The present invention also provides an apparatus for a multidimensional timbre perception spatial model modeling method based on EEG features, comprising a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program and, when executing the computer program, implement the steps of the multidimensional timbre perception spatial model modeling method based on EEG features as described above.
[0142] The present invention also provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the multidimensional timbre perception spatial model modeling method based on EEG features as described above.
[0143] The workflow and working principle of the present invention will be further described below with reference to a preferred embodiment:
[0144] This invention, based on multidimensional scaling analysis, discovers that acoustic features and psychological responses cannot fully characterize the multidimensional perceptual spatial attributes of timbre. This discovery inspires a successful investigation into the influence of EEG characteristics on timbre attributes. Multidimensional scaling analysis is heuristic; it objectively reflects differences in listener perception and evoked effects without requiring assumptions about spatial dimensions, thus helping to identify potential factors influencing the similarity between timbres.
[0145] like Figure 1 As shown, a multidimensional timbre perception spatial modeling method based on EEG features includes five stages: instrument timbre preprocessing, subjective evaluation, time-frequency domain feature extraction, EEG experiment, and establishment of timbre perception spatial model.
[0146] The first stage requires preprocessing the acquired audio signals, including: unifying audio duration and sampling rate, pre-emphasis, Gammatone filtering, frame segmentation, and windowing.
[0147] In the second stage, six sets of evaluation criteria (likability, loudness, warmth / coolness, brightness / darkness, soothingness, and fullness) were used to rate the 16 timbres, with scores ranging from -5 to 5, or from 0 to 10. Psychological perception characteristics were extracted through behavioral psychology experiments.
[0148] The third stage involves extracting acoustic time-frequency domain features. Frequency domain features include low energy characteristics, spectral descent (SRO), brightness, spectral irregularity (SI), spectral centroid (SC), spectral flatness (SFM), spectral skewness, spectral kurtosis, and spectral width. Time domain features include zero-crossing rate, start time, steady-state time, and RMS energy.
[0149] In the fourth stage, an timbre-based EEG experiment was conducted to extract EEG features. The EEG experiment used the oddball paradigm subtype, with standard and biased stimuli appearing in a random order. The standard stimulus was a 262Hz pure sinusoidal tone, with an occurrence probability of 70%. There were 17 biased stimuli in total, including a target stimulus of a frog's croak and 16 different instrument timbres as distraction stimuli, each with an equal occurrence probability, accounting for 30% of the total.
[0150] The fifth stage involves establishing a three-dimensional timbre space model.
[0151] Features correlated with EEG neural responses from two categories of features—psychological perception features and acoustic time-frequency domain features—are used as inputs to the two dimensions of the model. Various instrumental timbres are used as auditory stimuli in EEG experiments to extract event-related potentials (ERPs). The amplitude and latency of the extracted ERP components are used as the third dimension of the model.
[0152] The workflow for each stage is detailed below:
[0153] The first stage is the preprocessing of instrument timbre.
[0154] The final sound of any musical instrument is produced by the instrument's sound source after being processed by the instrument itself; timbre determines the unique sonic characteristics and expressive capabilities of each instrument. Taking into account current standards for classifying musical instruments, their physical characteristics, and the features of traditional and Western instruments, this invention focuses on string instruments and wind instruments. String instruments refer to all instruments that use the vibration of strings as their sound source, and can be further subdivided into bowed, plucked, and struck string instruments. Wind instruments are a general term for instruments that use air as their vibrating medium, and can be classified into brass and woodwind instruments according to their materials.
[0155] The workstation used to create the timbre samples for 8 Western musical instruments and 8 traditional Chinese musical instruments was Pro Tools Ultimate 2021.12.0, the calibration software was Melodyne 5.1.1.003, the loudness was -10 LUFS, and the format was lossless audio WAV. A 600ms audio file was selected and resampled at 44.1kHz. Temporal feature parameters were extracted from the individual notes of each instrument at a pitch of 262Hz (note name "C", solfège "do").
[0156] Table 1. Classification of Instrument Timbre
[0157]
[0158] Preprocessing of instrument timbre is achieved using a first-order FIR filter, with a pre-emphasis value of 0.97. A Gammatone filter bank, simulating human auditory perception, is used for filtering, with a center frequency range of 50Hz to 8000Hz. The filter function is used to cascade four second-order filters. Since the audio signals are all single tones, frame segmentation does not require consideration of whether a particular note has finished playing; therefore, this invention uses 20ms as one frame and a frame shift length of 10ms. A Hanning window is used to reduce spectral leakage.
[0159] The second stage involves extracting psychological perception characteristics through behavioral psychology experiments.
[0160] This invention first investigated nine sets of commonly used sound evaluation terms in subjective evaluations of sound quality. Through clustering and differential analysis, this invention selected six evaluation indicators that showed the greatest differences in participants' perceptual evaluations of 16 timbres during the scoring process. These included liking, loudness, warmth / coolness, brightness-depth, soothingness-urgency, and fullness-hollowness. This invention argues that these six evaluation indicators better characterize the psychological features of timbres. The scoring rules are as follows: For example, regarding liking, participants rated their personal preference for the instrumental timbres they heard, with scores ranging from "like" (+5 points) to "dislike" (-5 points), a total of 11 levels. Other physical indicators were scored similarly.
[0161] Based on the subjective evaluation data of the participants, this invention identified six groups of indicators (likability, loudness, brightness, soothingness, warmth / coolness, and fullness) that better describe the subjective feeling of timbre in order to examine the psychological characteristics of the participants. Figure 4 These are subjective evaluation values for the timbre of musical instruments.
[0162] The third stage involves extracting acoustic time-frequency domain features.
[0163] The timbre acoustic features utilize time-frequency domain features from MIRtoolbox 1.8.1. These frequency domain features include low-energy features, spectral descent (SRO), brightness, spectral irregularity (SI), spectral centroid (SC), spectral flatness (SFM), spectral skewness, spectral kurtosis, and spectral width. The time domain features include zero-crossing rate, start time, steady-state time, and RMS energy. The acoustic feature extraction framework of this invention is as follows: Figure 2 As shown.
[0164] Based on timbre acoustic descriptors, this invention extracts 13 time-frequency domain acoustic features of timbre as shown in the table below:
[0165] Table 2-1 Timbre Time-Frequency Domain Characteristics
[0166]
[0167] Table 2-2 (continued) Timbre Time-Frequency Domain Characteristics
[0168]
[0169]
[0170] In the fourth stage, an EEG experiment on vocal timbre was conducted to extract EEG characteristics.
[0171] Using the oddball paradigm subtype as an experimental paradigm for timbre EEG experiments, such as Figure 3 As shown in the figure. In this invention, the frog croak is used as the standard stimulus, which has a high probability of occurrence; the target stimuli are the 16 monotone timbres provided in the experiment, which have a low probability of occurrence.
[0172] The voice timbre EEG experiment involved 38 participants aged 19 to 37, 18 of whom were male. All participants were classified as right-handed according to the Edinburgh Hand Dominance Scale.
[0173] EEG signals for the timbre experiment were recorded using a Neuroscan Synamps2 system (64 electrodes) at a sampling rate of 1 kHz, with impedance maintained below 10 kΩ. Data were recorded using a Curry8 and analyzed offline in EEGLAB, with a 50 Hz notch filter applied to the amplifier. Continuous EEG data were divided into 1000 ms epochs, with a baseline of 200 ms before the deviance stimulus, and reference electrode M2 was set. For the timbre experiment, a 1 Hz high-pass filter and a 30 Hz low-pass filter were selected, and a ±75 μV extreme value standard was used to remove noisy segments, with a re-reference.
[0174] After preprocessing, 60 channels remained in the EEG experimental data. The ERP voltage values for different channels were obtained by averaging all segments of the subject's data. Observing the ERP waveforms, the waveforms of N400 and LPC were clearly observed. This invention averaged the peak times of the 60 channels, obtaining an average peak time of 422ms for N400 and 505ms for LPC. Therefore, time periods (412-432ms) and (495-515ms) were selected, and the N400 amplitude and latency, and the LPC amplitude and latency were extracted using the STUDY module in EEGLAB, respectively. Figure 5 and Figure 6 As shown. Among them, the late waveform of the trumpet exhibits more complex characteristics, making it difficult to effectively extract the amplitude and latency of its LPC.
[0175] The fifth stage involves establishing a three-dimensional timbre space model.
[0176] The timbre space, based on the distance matrix between timbres, establishes a multidimensional space. All instrument timbre objects are represented by scattered points in this space, and the similarity and dissimilarity of timbres are explored. This invention combines six sets of psychological perception features of instrument timbres, four sets of human auditory EEG features, and ten sets of acoustic spectrum features for evaluation. The dissimilarity of a total of 20 sets of timbre feature vectors is transformed into a low-dimensional space with mutually orthogonal dimensions. The analyzed column vectors are directly mapped to this low-dimensional space to form a point set, where the spatial distance between timbre points represents the perceptual dissimilarity, revealing the mapping relationship between timbre acoustic features, psychological perception features, and EEG responses.
[0177] The distance between the 16 single-tone timbre points is:
[0178]
[0179] In the formula:
[0180] d ij The spatial distance between timbre point i and timbre point j;
[0181] x ir Let i be the coordinates of the timbre point i in the dimension of psychological perception features;
[0182] x jr Let j be the coordinates of the timbre point j in the dimension of psychological perception features;
[0183] y is Let i be the coordinates of the timbre point i in the physical acoustic feature dimension;
[0184] y js Let j be the coordinates of the timbre point j in the physical acoustic feature dimension;
[0185] z it Let i be the coordinates of the timbre point i in the EEG feature dimension;
[0186] z jt Let j be the coordinates of the timbre point j in the EEG feature dimension;
[0187] R1 represents the total number of psychologically perceived characteristics of musical instrument timbre.
[0188] R2 represents the total number of physical acoustic characteristics of the instrument's timbre;
[0189] R3 represents the total number of EEG characteristics of musical instrument timbre;
[0190] w1 represents the weights of the psychological perception features;
[0191] w2 is the weight of the physical acoustic features;
[0192] w3 represents the weights of the EEG features.
[0193] After calculating the distance between single-tone timbre points, this invention reduces the 16×20-dimensional data to a 16×16-dimensional Euclidean timbre space. Let:
[0194] D represents the original Euclidean space of the timbre feature vector; the spatial dimension of D is N.
[0195] Q represents the low-dimensional space of the reduced timbre feature vector;
[0196] d a d b These represent two timbre feature vectors in the original Euclidean space D;
[0197] q a q b Indicates the corresponding d a d bTwo timbre feature vectors in the low-dimensional space Q;
[0198] dist ab Represents two timbre feature vectors d a d b Euclidean distance in primitive Euclidean space;
[0199] dist 2 ab Represents two timbre feature vectors d a d b The square of the Euclidean distance in the original Euclidean space;
[0200] B represents the inner product matrix of the timbre feature vectors;
[0201] The N-dimensional timbre primitive Euclidean space D is composed of a set of single-tone timbre points, and the Euclidean distance matrix between the N×N timbre points is... as follows:
[0202]
[0203] h ab Represents two timbre feature vectors d a d b Euclidean distance in primitive Euclidean space; a = 1, 2, ..., N; b = 1, 2, ..., N;
[0204] To facilitate intuitive observation of the differences between single-tone timbres, this invention utilizes the Multidimensional Scaling (MDS) method to reduce the dimensionality of timbre data. Based on the original Euclidean space D and dist... ab Calculate the inner product matrix B, perform eigenvalue and eigenvector decomposition on matrix B, and finally select the diagonal matrix and eigenvector matrix formed by the maximum values of the eigenvalues in the low-dimensional space. This ensures that the distance between any two timbre samples in the low-dimensional space Q is consistent with the distance in the original Euclidean space D. In other words, select the diagonal matrix and eigenvector matrix formed by the maximum values of the eigenvalues in the low-dimensional space Q such that the distance between any two timbre eigenvectors q in the low-dimensional space Q is equal to the distance between them. a q b The distance is consistent with the distance in the original Euclidean space D, such that ||q a -q b ||≈dist ab ;
[0205] According to D and dist 2 ab Calculate B, perform eigenvalue and eigenvector decomposition on B, and fit and construct the following low-dimensional space Q:
[0206]
[0207] In the formula:
[0208] Let B be a diagonal matrix formed by the largest eigenvalues of matrix B.
[0209] λ k The elements are on the diagonal; where k = 1, 2, ..., d′; d′ is the dimension of the chosen low-dimensional space;
[0210] The eigenvector matrix;
[0211] N is the dimension of the primitive Euclidean space. The dimension of the primitive Euclidean space equals the number of timbres per tone, and here N = 16.
[0212] This invention uses the stress coefficient as a measure of the goodness of fit in low-dimensional space; the smaller the stress coefficient, the better the model fit. Through comparison of stress coefficient evaluation indicators, the two-dimensional S-stress is 0.069, and the three-dimensional S-stress is 0.036. Compared to the two-dimensional timbre space, the three-dimensional timbre space can more effectively characterize the classification of instrument timbres. When the dimension is higher, the S-stress does not change much, but the spatial complexity increases. Considering the intuitiveness of the timbre space, the three-dimensional space is chosen, i.e., d′ = 3 in the timbre space.
[0213] Table 3 Stress coefficients S-stress in different dimensions
[0214]
[0215] Reduce the dimensionality of the subspace to a 3×N matrix. The formula is defined as follows:
[0216]
[0217] Among them, the elements in column e After dimensionality reduction by multidimensional scaling variance analysis, the relative coordinates of e = 1, 2, ..., N in the three-dimensional Cartesian coordinate system are given, where N represents the 16 monotone timbres.
[0218] A three-dimensional timbre perception space was constructed using multi-dimensional scales, which more intuitively represents the distribution of Chinese and Western musical instruments. For example... Figure 7 As shown, in three-dimensional space (Stress = 0.051, RSQ = 0.994), dashed lines map each dimension to visually observe the structural relationships between the various instruments. Specifically, Dim1 (Median: -0.086; IQR: 1.380), Dim2 (Median: 0.088; IQR: 1.071), and Dim3 (Median: 0.065; IQR: 0.911).
[0219] Observing the weighted coefficient values in the three-dimensional space, it was found that the feature attributes were affected relatively evenly by the three dimensions, indicating that all three dimensions contribute to the interpretability of timbre. In the first dimension, acoustic and psychological attributes have a significant impact; in the second dimension, EEG attributes have a significant impact; and in the third dimension, acoustic attributes have a significant impact. Furthermore, based on overall importance, the three dimensions explain 36.23%, 38.83%, and 24.58% of the similarity of 16 timbres, respectively, with a total determination coefficient of 0.99644. The weight value of psychological feature attributes is 0.1865, the weight value of acoustic feature attributes is 0.2464, and the weight value of EEG feature attributes is 0.5671, indicating that EEG feature attributes have high importance in the three-dimensional timbre space.
[0220] Table 4 Weighting coefficient values for three-dimensional space
[0221]
[0222] All bowed and struck string instruments (Erhu, Violin, Dulcimer, Piano, and Zither) have relatively close median distances (Median: 1.550; IQR: 2.225), and the Pipa (Lute) and Oboe (Oboe) have relatively close median distances (Median: 1.579), indicating a high degree of similarity in perceived timbre. In contrast, the Guitar, Harp, Sheng, and Flute have relatively large median distances from other timbres (Median: 3.869; IQR: 1.510), indicating greater differences in perceived timbre.
[0223] Guitar, harp, and sheng are typically performed as solos and are not commonly found in traditional symphony orchestra ensembles. This means that the three-dimensional timbre space proposed in this invention can effectively characterize the application of instrument timbre in a scene and can comprehensively and accurately explore the brain's electroencephalographic characteristics of timbre auditory perception.
[0224] This invention contains the following innovations: (1) The acoustic features of recorded instrument timbre samples are extracted, and the psychological perception features are extracted through behavioral psychology experiments. The features that are related to the EEG neural response in the two types of features are used as the inputs of the two dimensions of the model. (2) Various instrument timbres are used as auditory stimuli for EEG experiments to extract event-related potentials (ERPs). The amplitude and latency of the extracted ERP components are used as the third dimension of the model to characterize the differences in neural response to timbre auditory perception. (3) Six sets of timbre psychological perception evaluation indicators, 13 sets of acoustic evaluation indicators and four sets of EEG response features are applied to the construction of a three-dimensional timbre space, and the accuracy of the model is verified.
[0225] The embodiments described above are only used to illustrate the technical ideas and features of the present invention. Their purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The patent scope of the present invention should not be limited by these embodiments. That is, any equivalent changes or modifications made in accordance with the spirit disclosed in the present invention still fall within the patent scope of the present invention.
[0226] References cited in the background technology:
[0227] [1]Stumpf C.Die Sprachlaute:experimentell-phonetische Untersuchungennebst einem Anhangüber [M]. J. Springer, 1926.
[0228] [2]Helmholtz HL F.On the Sensations of Tone as a Physiological Basisfor the Theory of Music[M].Cambridge University Press,2009.
[0229] [3]Seashore HL F.On the Sensations of Tone as a Physiological Basis for the Theory of Music[M].Cambridge University Press,2009.
[0230] [4]Wessel D L.Psychoacoustics and music:A report from Michigan State University[J].PACE:Bulletin of the Computer Arts Society,1973,30:1-2.
[0231] [5]Reymore L.Characterizing prototypical musical instrument timbreswith Timbre Trait Profiles[J].Musicae Scientiae,2022,26(3):648-674.
[0232] [6]Jiang W,Liu J,Zhang X,et al.Analysis and modeling of timbreperception features in musicalsounds[J].Applied Sciences,2020,10(3):789.
[0233] [7]Jiang W,Liu J,Li Z,et al.Analysis and modeling of timbreperception features of chinese musicalinstruments[C] / / 2019IEEE / ACIS18thInternational Conference on Computer and Information Science(ICIS).IEEE,2019:191-195.
[0234] [8]Plomp R.Timbre as a multi-dimensional attribute of complex tones[J].Frequency analysis andperiodicity detection in hearing,1970:397-414.
[0235] [9]Siedenburg K,Jones-Mollerup K,McAdams S.Acoustic and categoricaldissimilarity of musical timbre:
[0236] Evidence from asymmetries between acoustic and chimeric sounds[J].Frontiers in Psychology,2016,6: 1977.
[0238]
[10] Saitis C,Siedenburg K.Brightness perception for musicalinstrument sounds:Relation to timbredissimilarity and source-cause categories[J].The Journal of the Acoustical Society of America,2020,
[0239] 148(4):2256-2266.
[0240]
[11] McAdams S.The perceptual representation of timbre[J].Timbre:Acoustics,perception,and cognition,
[0241] 2019:23-57.
[0242]
[12] Thoret E,Caramiaux B,Depalle P,et al.Learning metrics onspectrotemporal modulations reveals theperception of musical instrumenttimbre[J].Nature Human Behaviour,2021,5(3):369-377.
[0243]
[13] Tsekoura K,Foka A.Classification of EEG signals produced bymusical notes as stimuli[J].ExpertSystems with Applications,2020,159:113507.
Claims
1. A method for modeling a multidimensional timbre perception spatial model based on electroencephalogram (EEG) features, characterized in that, This method includes the following steps: Step 1: Collect single-timbre samples from various musical instruments and preprocess the collected timbre samples; Step 2: Extract acoustic time-frequency domain features from the preprocessed timbre samples, and extract psychological perception features through behavioral psychology experiments; Step 3: Use samples of various musical instrument timbres as auditory stimuli to conduct EEG experiments, and extract the corresponding event-related potential (ERP) signals as EEG features. Step 4: Use the Euclidean distance between different timbre points to represent the dissimilarity of different timbre features; calculate the distance between different timbre points; Step 5: Based on the Euclidean distance matrix between sample timbre points, fit a low-dimensional space with mutually orthogonal dimensions, transform the dissimilarity of timbre feature vectors into the low-dimensional space, directly map the timbre feature vectors into the low-dimensional space to form a point set, and use the scattered points in the low-dimensional space to represent the timbre objects of each instrument to reflect the mapping relationship between the acoustic features, psychological perception features and EEG responses of various instrument timbres; and intuitively display the similarity and dissimilarity of each instrument timbre.
2. The method for modeling a multidimensional timbre perception space model based on EEG features according to claim 1, characterized in that, In step 1, the collected timbre samples are processed as follows: unified audio duration and sampling rate, pre-emphasis, Gammatone filtering, frame segmentation, and windowing.
3. The method for modeling a multidimensional timbre perception space model based on EEG features according to claim 1, characterized in that, In step 2, the following acoustic time-frequency domain features are extracted: low energy features, spectral descent, brightness, spectral irregularity features, spectral centroid features, spectral flatness, spectral skewness, spectral kurtosis, and spectral width. The time-domain characteristics include zero-crossing rate, start time, steady-state time, and RMS energy.
4. The method for modeling a multidimensional timbre perception space model based on EEG features according to claim 1, characterized in that, Step 2, the method for extracting psychological perception features through behavioral psychology experiments includes the following steps: The timbre samples were scored according to the following evaluation parameters: likability, loudness, warmth / coolness, brightness / darkness, soothingness, and fullness.
5. The method for modeling a multidimensional timbre perception space model based on EEG features according to claim 1, characterized in that, In step 3, the EEG experiment used the oddball paradigm subtype, with standard and biased stimuli appearing in a random order; the standard stimulus was a 262Hz sine wave pure tone, with an occurrence probability of 70%; there were 17 types of biased stimuli, among which the target stimulus was the sound of a frog; the distraction stimuli were 16 different instrument timbres, and each biased stimulus had an equal occurrence probability, accounting for 30% in total.
6. The method for modeling a multidimensional timbre perception space model based on EEG features according to claim 1, characterized in that, In step 5, the formula for calculating the distance between different timbre points is as follows: In the formula: d ij The spatial distance between timbre point i and timbre point j; x ir Let i be the coordinates of the timbre point i in the dimension of psychological perception features; x jr Let j be the coordinates of the timbre point j in the dimension of psychological perception features; y is Let i be the coordinates of the timbre point i in the physical acoustic feature dimension; y js Let j be the coordinates of the timbre point j in the physical acoustic feature dimension; z it Let i be the coordinates of the timbre point i in the EEG feature dimension; z jt Let j be the coordinates of the timbre point j in the EEG feature dimension; R1 represents the total number of psychologically perceived characteristics of musical instrument timbre. R2 represents the total number of physical acoustic characteristics of the instrument's timbre; R3 represents the total number of EEG characteristics of musical instrument timbre; w1 represents the weights of the psychological perception features; w2 is the weight of the physical acoustic features; w3 represents the weights of the EEG features.
7. The method for modeling a multidimensional timbre perception space model based on EEG features according to claim 1, characterized in that, Step 4, the method for reducing the dimensionality of the timbre feature vector using a multi-dimensional scaling transformation method, includes the following steps: Let D represent the original Euclidean space of the timbre feature vectors; Let Q represent the low-dimensional space of the timbre feature vectors after dimensionality reduction; Let d a d b These represent two timbre feature vectors in the original Euclidean space D; Let q a q b Indicates the corresponding d a d b Two timbre feature vectors in the low-dimensional space Q; dist ab Represents two timbre feature vectors d a d b Euclidean distance in primitive Euclidean space; Let B denote the inner product matrix of the timbre feature vectors; According to D and dist 2 ab Calculate B, perform eigenvalue and eigenvector decomposition on B, and fit and construct the following low-dimensional space Q: in Let B be a diagonal matrix formed by the largest eigenvalues of matrix B. λ k The elements are on the diagonal; where k = 1, 2, ..., d′; d′ is the dimension of the chosen low-dimensional space; The eigenvector matrix; N is the dimension of the primitive Euclidean space; Take the diagonal matrix and eigenvector matrix formed by the maximum values of the eigenvalues in the low-dimensional space Q, such that for any two timbre eigenvectors q in the low-dimensional space Q... a q b The distance is consistent with the distance in the original Euclidean space D, i.e., ||q a -q b ||≈dist ab ; The stress coefficient is used as a measure of the goodness of fit in low-dimensional space. The smaller the stress coefficient, the better the model fit.
8. A device for modeling a multidimensional timbre perception spatial model based on electroencephalogram (EEG) features, comprising a memory and a processor, characterized in that, The memory is used to store a computer program; the processor is used to execute the computer program and, when executing the computer program, implement the steps of the multidimensional timbre perception space modeling method based on EEG features as described in any one of claims 1 to 7.
9. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the multidimensional timbre perception space modeling method based on EEG features as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Latent class multi-dimensional scale analysis-based sound quality modeling method
CN108596217A
Singing timbre similarity evaluation method based on two-dimensional singing timbre model
CN114067835A