Traditional folk song singing voice acoustics and physiology multi-modal analysis method and system

By synchronizing and calibrating the voice recording equipment and sensors using a unified clock source, and combining Fourier transform and correlation analysis, a quantitative correlation between acoustic features and physiological mechanisms is established. This solves the problem of time synchronization in traditional folk song research and enables scientific analysis and teaching guidance of traditional folk song techniques.

CN121237105APending Publication Date: 2025-12-30NORTHWEST UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511590597.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In existing technologies, voice recording devices, voice monitoring devices, and breathing sensors operate independently, lacking unified time reference control. This leads to time synchronization issues, affecting the accuracy of acoustic and physiological parameters and cross-modal data correlation, and making it impossible to accurately identify the intrinsic connection between physiological mechanisms and acoustic phenomena in traditional folk song singing techniques.

Method used

By synchronizing the voice recording device, voice monitoring device, and breathing sensor using a unified clock source, time-aligned multimodal signal data streams are obtained. Acoustic feature vectors are extracted using Fourier transform, and Pearson coefficients are calculated through correlation analysis to establish a quantitative correlation between acoustic features and physiological mechanisms. A linear regression model is then used to fit the mapping relationship.

Benefits of technology

It has achieved precise quantification and objective evaluation of traditional folk song singing techniques, accurately reproduced the physiological mechanisms of folk song techniques such as "rapid inhalation and slow exhalation", and output a quantitative correlation model for teaching and guiding traditional folk songs, thus realizing the scientific inheritance of national vocal techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237105A_ABST
    Figure CN121237105A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional folk song singing voice acoustics and physiology multi-modal analysis method and system, and the method comprises the steps: carrying out the synchronous calibration of a voice recording device, a voice monitoring device and a respiration sensor through a unified clock source, obtaining a multi-modal signal data flow, and obtaining a time-aligned original data set; according to the original data set subjected to time alignment, Fourier transform is adopted to process the voice signal part, acoustic feature vectors are determined, and a physiological signal sequence is extracted from the original data set subjected to time alignment; if the acoustic feature vectors are matched with the timestamps of the physiological signal sequences, Pearson coefficients between the acoustic feature vectors and the timestamps are calculated through correlation analysis, and quantitative correlation strength is judged; for the part of which the quantitative correlation strength is higher than a fourth preset threshold value, fitting a mapping relationship between the acoustic characteristics and the physiological mechanism by adopting a linear regression model to obtain a parameterized singing technique model; and according to the parameterized singing technique model, analyzing the relationship between traditional singing voice acoustics and physiological multiple modes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound multi-modal analysis, and particularly relates to a traditional folk song singing voice acoustic and physiological multi-modal analysis method and system. BACKGROUND

[0002] As an important carrier of cultural heritage of various nationalities, the scientific research on singing techniques of traditional folk songs is of great significance for the protection of oral culture and vocal music teaching. Current research mainly relies on spectral analysis technology to obtain the acoustic characteristics of singing voice, and at the same time, professional equipment is used to monitor the physiological signal changes in the singing process. These researches lay a foundation for understanding the singing mechanism of folk songs.

[0003] However, the existing research method has key defects in technical implementation. In the current signal collection process, the voice recording device, the voice monitoring device and the respiration sensor work independently, and there is a lack of unified time reference control, resulting in a time deviation of milliseconds between different devices. This time asynchronization problem directly affects the accuracy of subsequent feature extraction, so that the acoustic parameters and physiological parameters cannot be accurately corresponded. More seriously, due to the lack of effective cross-modal data correlation mechanism, researchers cannot establish the quantitative relationship between acoustic phenomena and physiological mechanisms, so that the analysis results can only stay at the description level of a single dimension.

[0004] This technical limitation has caused significant business obstacles in practical application. For example, taking the typical "quick inhale and slow exhale" singing technique in the Yugur folk song as an example, when researchers try to analyze this technique, they cannot accurately identify the internal relationship between the breathing rhythm and the tone performance because they cannot synchronously obtain the physiological data of respiratory control and the corresponding acoustic data of tone change. This technical defect not only limits the scientific understanding of traditional singing techniques, but also directly affects the cultural protection work based on parameter reconstruction and the objective guidance of vocal music teaching.

[0005] Therefore, how to realize accurate synchronous collection of multiple signals and establish a quantitative correlation model between acoustic characteristics and physiological mechanisms has become a key problem for promoting the transformation of traditional folk song research from qualitative description to quantitative analysis. SUMMARY

[0006] In order to solve the above technical problems, the present application provides a traditional folk song singing voice acoustic and physiological multi-modal analysis method, comprising the following steps:

[0007] Synchronize and calibrate the voice recording device, the voice monitoring device and the respiration sensor by a unified clock source, obtain a multi-modal signal data stream, and obtain a time-aligned original data set;

[0008] According to the time-aligned original data set, a Fourier transform is performed on a speech signal part to determine an acoustic feature vector, and a physiological signal sequence is extracted from the time-aligned original data set;

[0009] If the acoustic feature vector matches the timestamp of the physiological signal sequence, a Pearson coefficient between the two is calculated through correlation analysis to determine the quantitative correlation strength;

[0010] For the part where the quantitative correlation strength is higher than a fourth preset threshold, a linear regression model is used to fit the mapping relationship between the acoustic feature and the physiological mechanism to obtain a parameterized singing technique model;

[0011] Different folk song singing voices are obtained, and the relationship between the acoustic and physiological multi-modalities of the folk song singing voice is analyzed according to the parameterized singing technique model.

[0012] Preferably, the method for obtaining the time-aligned original data set comprises:

[0013] A unified clock source reference signal is obtained, and a synchronization clock pulse is sent to a voice recording device, a voice monitoring device, and a respiration sensor through clock distribution to obtain a clock calibration state of each device;

[0014] According to the clock calibration state, a multi-modal data acquisition process is started, the voice recording device acquires an audio signal, the voice monitoring device obtains an acoustic parameter, and the respiration sensor captures a respiration waveform to obtain three parallel data streams;

[0015] The three parallel data streams are marked using a timestamp marking technology, and if the data packet timestamp difference exceeds a first preset threshold, a resynchronization mechanism is triggered to obtain multi-modal signal data with accurate time marking;

[0016] The multi-modal signal data with accurate time marking is temporarily stored through a buffer management mechanism, each modal data packet is arranged according to the timestamp order to obtain the time-aligned original data set.

[0017] Preferably, the method for obtaining the acoustic feature vector comprises:

[0018] The time-synchronized original data set is obtained, a multi-channel speech signal is processed through a data alignment algorithm to obtain a speech data sequence with a unified time reference, and the speech data sequence is de-noised and de-silenced to obtain a purified speech signal segment;

[0019] According to the purified speech signal segment, a Fourier transform algorithm is used to convert to a frequency domain space to obtain frequency spectrum distribution data;

[0020] Based on the spectral distribution data, the main frequency components are identified using the peak detection method to determine the fundamental frequency candidate point set. If there are multiple peaks in the fundamental frequency candidate point set, the true fundamental frequency is screened through harmonic relationship verification to obtain fundamental frequency sequence data.

[0021] Based on the spectral distribution data, the speech resonance characteristics are analyzed using a linear predictive coding method to obtain the frequency positions of the formant peaks;

[0022] The acoustic feature vector is constructed using the fundamental frequency sequence data and the formant frequency positions.

[0023] Preferably, the method for extracting the physiological signal sequence includes:

[0024] Obtain the original dataset for time synchronization, parse the format of the original data, and obtain a mixed data stream containing signals from multiple sensors;

[0025] Physiological signal data is separated from the mixed data stream according to a preset signal identifier. If the signal identifier matches the respiratory sensor marker, the corresponding respiratory-related data segment is extracted. If the signal identifier matches the acoustic sensor marker, the corresponding voice-related data segment is extracted.

[0026] The separated physiological signal data is filtered using a bandpass filter. By setting a frequency range, high-frequency noise interference and low-frequency baseline drift are removed to obtain purified signal data.

[0027] The respiratory cycle markers in the purified signal data are identified by a peak detection algorithm, and the respiratory frequency is calculated based on the time interval between adjacent peaks to obtain a respiratory rhythm sequence.

[0028] The vocal tension state is analyzed based on the amplitude change pattern of the purified signal data. If the amplitude change exceeds the second preset threshold, it is judged as an abnormal tension state. If the amplitude change is within the normal range, it is judged as a normal tension state, and the vocal tension sequence is determined.

[0029] The breathing rhythm sequence and the vocal tension sequence are synchronized and aligned using a time window sliding method. The correspondence between the two sequences is established by timestamp matching to obtain the physiological signal sequence with consistent time.

[0030] Preferably, the method for determining the quantitative correlation strength includes:

[0031] The data streams of the acoustic feature vectors and the physiological signal sequences are acquired, and their respective time stamp information is extracted.

[0032] Based on the extracted time stamp information, the acoustic feature vector and the physiological signal sequence are time-aligned using a time window sliding matching method to obtain a synchronized data pair sequence.

[0033] If the time difference in the synchronized data pair sequence is less than the third preset threshold, it is determined to be a valid aligned data pair, and a time-matched feature-signal pairing set is obtained;

[0034] The feature-signal pairing set is processed by the Pearson correlation coefficient calculation method to obtain the linear correlation values ​​between each pairing.

[0035] Based on the distribution of the linear correlation values, statistical analysis methods are used to calculate the mean and variance of the overall correlation strength, determine the comprehensive correlation degree between acoustic features and physiological signals, and obtain the quantified correlation strength.

[0036] Preferably, the method for obtaining the parametric singing technique model includes:

[0037] The acoustic feature and physiological mechanism data pairs with quantitative correlation strength higher than the fourth preset threshold are filtered out by data screening to obtain the filtered dataset.

[0038] Feature vectors of acoustic features and physiological mechanisms are extracted from the selected dataset, and dimensionality reduction is performed using principal component analysis to obtain a dimensionality-reduced feature set.

[0039] For the dimensionality-reduced feature set, if the feature dimension is lower than a preset dimension, the missing features are filled by mean to obtain a standardized feature set;

[0040] The standardized feature set is trained using a linear regression model, and the regression coefficients are optimized using the least squares method to obtain a fitted mapping model;

[0041] Based on the fitted mapping model, the mapping parameters from acoustic features to physiological mechanisms are calculated to obtain the parameterized singing technique model.

[0042] This invention also provides a multimodal analysis system for the acoustic and physiological aspects of traditional folk song singing. The system applies the above-mentioned method and includes: a data acquisition module, a data extraction module, a correlation strength judgment module, a model construction module, and an analysis module.

[0043] The data acquisition module synchronously calibrates the voice recording device, voice monitoring device, and breathing sensor using a unified clock source to acquire multimodal signal data streams and obtain time-aligned raw datasets.

[0044] The data extraction module uses Fourier transform to process the speech signal portion based on the time-aligned original dataset, determines the acoustic feature vector, and extracts the physiological signal sequence from the time-aligned original dataset.

[0045] In the correlation strength determination module, if the acoustic feature vector matches the timestamp of the physiological signal sequence, the Pearson coefficient between the two is calculated through correlation analysis to determine the quantitative correlation strength.

[0046] The model building module uses a linear regression model to fit the mapping relationship between acoustic features and physiological mechanisms for the portion of the quantized correlation strength that is higher than the fourth preset threshold, thereby obtaining a parametric singing technique model.

[0047] The analysis module is used to acquire different folk song singing voices and, based on the parameterized singing technique model, analyze the relationship between the acoustics and physiological multimodalities of folk song singing voices.

[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0049] This invention addresses the challenge of accurately quantifying and objectively evaluating singing techniques in traditional vocal music teaching. It achieves this by synchronously calibrating voice recording equipment, voice monitoring devices, and breathing sensors using a unified clock source to acquire time-aligned multimodal signal data streams. Fourier transform is used to extract acoustic feature vectors, and filtering algorithms are employed to process physiological signals and obtain breathing rhythm and vocal tension sequences. Correlation analysis is used to calculate the Pearson coefficient between acoustic features and physiological signals. For correlation strengths exceeding a preset threshold, a linear regression model is used to fit the mapping relationship, establishing a parametric singing technique model. This invention accurately reproduces the physiological mechanism of the "rapid inhalation and slow exhalation" technique in traditional folk songs, such as those of the Yugur people. Through simulated signal sequence verification and mean square error optimization, a quantitative correlation model is ultimately output for use in teaching traditional folk songs, achieving the scientific inheritance of ethnic vocal techniques. Attached Figure Description

[0050] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0054] Example 1

[0055] In this embodiment, as Figure 1 As shown, a method for analyzing the acoustic and physiological multimodal characteristics of traditional folk song singing includes the following steps:

[0056] S1. Synchronously calibrate the voice recording device, voice monitoring device, and breathing sensor using a unified clock source to obtain multimodal signal data streams and obtain time-aligned raw datasets.

[0057] The method for obtaining the time-aligned raw dataset includes: acquiring a unified clock source reference signal; sending synchronous clock pulses to the voice recording device, voice monitoring device, and breathing sensor through clock allocation to obtain the clock calibration status of each device; based on the clock calibration status, initiating a multimodal data acquisition process, with the voice recording device acquiring audio signals, the voice monitoring device acquiring acoustic parameters, and the breathing sensor capturing breathing waveforms to obtain three parallel data streams; using timestamp marking technology to mark the three parallel data streams, and triggering a resynchronization mechanism if the timestamp difference of the data packets exceeds a first preset threshold to obtain multimodal signal data with precise timestamps; temporarily storing the multimodal signal data with precise timestamps through a buffer management mechanism, and arranging the data packets of each modality according to the timestamp order to obtain the time-aligned raw dataset.

[0058] In this embodiment, to achieve correlation analysis between acoustic and physiological signals at the microscopic level, the primary task of data acquisition is to ensure that all data streams are under a unified and accurate time reference. A dedicated synchronous clock source based on the IEEE 1588 precision clock protocol is used as the core time base of the entire data acquisition network. This clock source continuously sends high-precision synchronous clock pulse signals to all data acquisition nodes in the network via Ethernet interface in a multicast manner, including high-fidelity voice recording devices (such as professional sound cards and microphones), voice monitoring devices (such as laryngeal transducers or electroglottic transducers), and respiratory sensors (such as respiratory plethysmography or piezoelectric abdominal belts). The PTP slave clock built into each acquisition node dynamically corrects its own clock drift through continuous message exchange with the master clock, ultimately enabling all nodes to achieve microsecond-level clock synchronization accuracy and providing real-time feedback on their clock calibration status. After monitoring that the calibration status of all nodes is stable within the preset tolerance range, the acquisition process of all devices is atomically started according to a unified global clock, either by software triggering or hardware triggering, thereby ensuring that data streams of all modes are output simultaneously from the first data sample.

[0059] Upon receiving the synchronization start command, the multimodal data acquisition process begins. The voice recording device acquires raw audio waveforms at a high sampling rate, generating a mono or stereo PCM data stream. The voice monitoring device simultaneously captures specific acoustic parameters reflecting vocal cord vibration characteristics, such as fundamental frequency perturbations, amplitude perturbations, and voice spectrum features, forming a parameter data stream. The respiratory sensor continuously records respiratory motion waveforms in the chest or abdominal cavity at a relatively low sampling rate, generating a respiratory pressure or volume change data stream. These three parallel data streams are physically transmitted in parallel to the data acquisition server. To address potential minor jitter during transmission, a hardware and software combined timestamp technology is used: a hardware timestamp is applied when the data packet leaves the acquisition device, and a software timestamp is applied when it enters the receiving buffer. The system continuously monitors the timestamp difference between adjacent data packets in these two data streams. If the timestamp difference between packets in either data stream exceeds a first preset threshold (for audio streams, the first preset threshold is ±50 microseconds beyond the sampling period), a resynchronization mechanism is immediately triggered. This mechanism may include discarding time-scrambled packets, resetting buffers, or, in severe cases, reinitializing the data stream channel to ensure that each frame of data that is eventually aggregated carries an accurate and reliable time stamp.

[0060] After precise timestamping, data packets from the three independent channels are temporarily stored in a circular buffer with multi-priority management. The core management mechanism of this buffer is to sort all data packets based on their global timestamps, rather than their arrival order at the server. Because different devices have different sampling rates and data packet sizes, the buffer manager uses a unified time sequence as a benchmark to "align and encapsulate" all modal data packets arriving within each time unit, ultimately resulting in a strictly time-aligned, structured original dataset.

[0061] S2. Based on the time-aligned original dataset, Fourier transform is used to process the speech signal portion to determine the acoustic feature vector, and physiological signal sequences are extracted from the time-aligned original dataset.

[0062] S2.1. The method for obtaining acoustic feature vectors includes: acquiring a time-synchronized original dataset; processing multi-channel speech signals using a data alignment algorithm to obtain a speech data sequence with a unified time reference; removing noise interference and silence segments from the speech data sequence to obtain purified speech signal segments; converting the purified speech signal segments to the frequency domain using a Fourier transform algorithm to obtain spectral distribution data; identifying the main frequency components using a peak detection method based on the spectral distribution data to determine the fundamental frequency candidate point set; if multiple peaks exist in the fundamental frequency candidate point set, verifying and filtering the true fundamental frequency through harmonic relationships to obtain fundamental frequency sequence data; analyzing speech resonance characteristics using a linear predictive coding method based on the spectral distribution data to obtain the formant frequency positions; and constructing acoustic feature vectors using the fundamental frequency sequence data and formant frequency positions.

[0063] In this embodiment, firstly, after acquiring the original time-synchronized dataset, the multi-channel speech signal needs to be processed using a data alignment algorithm to eliminate time offsets caused by device latency or sampling rate differences. Specifically, an algorithm based on Dynamic Time Warping (DTW) is used to align the multi-channel speech signals (such as the left and right channels from a stereo microphone) to a unified time reference, generating a speech data sequence with a consistent time axis. Next, the sequence is preprocessed to remove noise interference and silence segments. A Least Mean Square (LMS) filter is used to remove noise interference from the sequence, and the filter parameters are dynamically adjusted according to the statistical characteristics of background noise to effectively suppress environmental noise and circuit noise. Silence segment detection is performed based on short-time energy and zero-crossing rate analysis: the signal energy and zero-crossing rate within each time window are calculated, and if the energy is below a preset threshold and the zero-crossing rate changes smoothly, it is determined to be a silence segment and removed. Finally, the purified speech signal segment is obtained. Secondly, a Fourier transform algorithm is applied to the purified speech signal segments to transform them from the time domain to the frequency domain. Specifically, a Short-Time Fourier Transform (STFT) is used with a Hamming window as the window function, a window length of 25 milliseconds, and an overlap rate of 50% to ensure high-resolution spectral distribution data in the time-frequency domain. Spectral calculation is performed using a Fast Fourier Transform (FFT) to generate the amplitude and phase spectra for each time frame. The amplitude spectrum is further used to calculate the power spectral density (PSD), and a logarithmic transform is used to enhance low-frequency details, thereby obtaining clear spectral distribution data. Based on the spectral distribution data, a peak detection method is used to identify the main frequency components to determine the set of candidate fundamental frequencies. Peak detection involves scanning the power spectrum to find local maxima and setting an amplitude threshold (in this embodiment, only points higher than the spectral mean by 10 dB are considered candidates). If multiple peaks are detected, the true fundamental frequency is screened by verifying harmonic relationships: Candidate peaks are checked for significant energy at integer multiples of frequency. If the sum of harmonic energy at a candidate point exceeds that of other candidate points, it is determined to be the true fundamental frequency; otherwise, an autocorrelation function is used for verification, calculating the signal's autocorrelation to confirm the fundamental frequency period. This process is performed frame-by-frame, ultimately generating fundamental frequency sequence data, where each time point corresponds to a fundamental frequency value used to characterize the fundamental frequency characteristics of the voice. Finally, based on the spectral distribution data, the linear predictive coding (LPC) method is used to analyze speech resonance characteristics to extract formant frequency positions: The LPC model calculates prediction coefficients by fitting an all-pole model of the speech signal, and then obtains the formant frequencies by solving the polynomial roots. Specifically, the LPC coefficients are rooted using a polynomial, and roots with positive imaginary parts are selected and converted into frequency values. These values ​​correspond to key positions such as the first formant (F1), the second formant (F2), etc.Simultaneously, an acoustic feature vector is constructed by combining fundamental frequency sequence data: this vector includes the fundamental frequency (F0), formant frequencies (F1, F2, F3), and derived features such as fundamental frequency perturbation and amplitude perturbation. These features are combined into an acoustic feature vector after standardization.

[0064] S2.2. The method for extracting physiological signal sequences includes: acquiring a time-synchronized raw dataset, parsing the format of the raw data to obtain a mixed data stream containing signals from multiple sensors; separating physiological signal data from the mixed data stream according to preset signal identifiers; if the signal identifier matches a respiratory sensor marker, extracting the corresponding respiratory-related data segment; if the signal identifier matches an acoustic sensor marker, extracting the corresponding voice-related data segment; filtering the separated physiological signal data using a bandpass filter, removing high-frequency noise interference and low-frequency baseline drift by setting a frequency range to obtain purified signal data; identifying respiratory cycle markers in the purified signal data using a peak detection algorithm, calculating the respiratory frequency based on the time interval between adjacent peaks to obtain a respiratory rhythm sequence; analyzing the voice tension state based on the amplitude change pattern of the purified signal data; if the amplitude change exceeds a second preset threshold, it is judged as an abnormal tension state; if the amplitude change is within the normal range, it is judged as a normal tension state, determining the voice tension sequence; synchronizing and aligning the respiratory rhythm sequence and the voice tension sequence using a time window sliding method, establishing the correspondence between the two sequences through timestamp matching to obtain a time-consistent physiological signal sequence.

[0065] In this embodiment, firstly, the mixed data stream is read and decomposed from the time-aligned raw dataset generated in step S1, and the metadata in the data stream is parsed, including the data source ID, timestamp, data length, and encoding format. For example, the data packet of a multichannel physiological recorder (PowerLab) respiratory signal may contain differential voltage values ​​of two channels (thoracic respiratory lead and abdominal respiratory lead), while the data packet of an electroglottic recorder contains an impedance change sequence reflecting the glottic contact area. The parser demultiplexes the mixed data stream according to a preset signal identifier, accurately separating the target physiological signal data segment, laying the foundation for subsequent modality-specific processing. Secondly, a modality-specific digital filter is used to preprocess the separated physiological signal data to remove various noise interferences. For respiratory signals, since their effective frequency components are usually concentrated between 0.1 Hz and 1.0 Hz, this embodiment uses a second-order Butterworth bandpass filter for noise removal, with a low cutoff frequency set to 0.05 Hz to suppress baseline drift due to body movement or slow breathing, and a high cutoff frequency set to 2 Hz to eliminate power line interference and high-frequency electromyographic noise. For voice-related physiological signals from an electroglottogram (EGG) analyzer, key information (such as the physiological origin of fundamental frequency perturbations) resides in the time-domain details of the waveform. Therefore, this embodiment uses a bandpass filter from 0.1 Hz to 50 Hz to preserve the basic EGG waveform morphology while removing DC offset and high-frequency noise. Subsequently, physiologically significant feature sequences are extracted from the purified signal data. Specifically, a peak detection algorithm based on amplitude thresholds is used to extract feature sequences from the filtered respiratory waveform. The algorithm first locates all potential peaks by finding the zero-crossing points of the first derivative of the signal, and then sets an adaptive dynamic threshold to filter out effective inspiratory peaks. The time interval between adjacent inspiratory peaks is defined as a complete respiratory cycle, and its reciprocal is the instantaneous respiratory frequency, ultimately generating a continuous respiratory rhythm sequence. For the voice tension sequence, the amplitude variation and waveform stability of the EGG signal are analyzed: by calculating the standard deviation and peak value of the EGG signal within each time window, the stability and strength of glottal closure can be quantified. If the amplitude change exceeds a second preset threshold (in this embodiment, based on 2.5 times the standard deviation of the singer's quiet vocal baseline), it is judged as an abnormal vocal tension state (possibly indicating excessive tension in the laryngeal muscles); if it is within the normal range, it is judged as a normal tension state, thus forming a discrete vocal tension state sequence corresponding to a timestamp. Finally, the breathing rhythm sequence and vocal tension sequence are synchronized and aligned using a time window sliding method, and the correspondence between the two sequences is established through timestamp matching to obtain a time-consistent physiological signal sequence.

[0066] S3. If the acoustic feature vector matches the timestamp of the physiological signal sequence, then the Pearson coefficient between the two is calculated through correlation analysis to determine the strength of the quantitative correlation.

[0067] The method for determining the quantitative correlation strength includes: acquiring the data streams of acoustic feature vectors and physiological signal sequences, and extracting their respective time stamp information; based on the extracted time stamp information, using a time window sliding matching method to perform time alignment processing on the acoustic feature vectors and physiological signal sequences to obtain synchronized data pair sequences; if the time difference in the synchronized data pair sequences is less than a third preset threshold, it is judged as an effective aligned data pair, and a time-matched feature-signal pair set is obtained; the feature-signal pair set is processed using the Pearson correlation coefficient calculation method to obtain the linear correlation values ​​between each pair; based on the distribution of the linear correlation values, a statistical analysis method is used to calculate the mean and variance of the overall correlation strength, determine the comprehensive correlation degree between acoustic features and physiological signals, and obtain the quantitative correlation strength.

[0068] In this embodiment, the acoustic feature vector and physiological signal sequence are first time-aligned to ensure data synchronization at a microscopic time scale. Specifically, time stamp information is extracted from the acoustic feature vector data stream and physiological signal sequence data stream obtained in step S2. These stamps are generated based on the unified IEEE 1588 precision clock protocol in step S1, with a precision of microseconds. Alignment is performed using a time window sliding matching method: a dynamic time window is defined, dividing the acoustic feature vector and physiological signal sequence into overlapping window units along the time axis. Within each window, the difference between the timestamp of the acoustic feature vector and the timestamp of the physiological signal sequence is calculated. If this difference is less than a third preset threshold (set to 5 milliseconds in this embodiment to match the half-frame period of the audio sampling rate), it is determined to be a valid aligned data pair, forming a synchronized data pair sequence. This process is implemented through a circular buffer. The buffer manager continuously monitors the timestamp sequence and automatically discards timed-out or misaligned data packets, ensuring that the final time-matched feature-signal pair set has high temporal consistency, laying the foundation for subsequent correlation analysis. Then, Pearson correlation coefficients were calculated for the time-matched feature-signal pair sets to quantify the linear correlation strength between acoustic features and physiological signals. The calculation process involved iteratively traversing all paired data: first, each feature-signal pair was standardized to eliminate the influence of dimensions; then, the covariance and standard deviation were calculated to obtain the linear correlation values ​​for each pair, ranging from -1 to 1. Finally, based on the calculated distribution of linear correlation values, statistical analysis methods were used to determine the overall quantitative correlation strength. First, descriptive statistics were performed on all paired correlation coefficients, calculating their mean and variance. If the mean was higher than 0.5 and the variance was low, a strong comprehensive correlation between acoustic features and physiological signals was determined; conversely, if the mean was close to 0 and the variance was high, the correlation was weak. Furthermore, hypothesis testing was used to assess the significance of the correlation coefficients; a p-value less than 0.05 was considered statistically significant. Finally, the quantitative correlation strength was obtained through a combination of the mean and variance.

[0069] S4. For the portion where the quantitative correlation strength is higher than the fourth preset threshold, a linear regression model is used to fit the mapping relationship between acoustic features and physiological mechanisms to obtain a parametric singing technique model.

[0070] The method for obtaining the parametric singing technique model includes: filtering out acoustic feature and physiological mechanism data pairs with a quantitative correlation strength higher than a fourth preset threshold through data screening to obtain a screened dataset; extracting feature vectors of acoustic features and physiological mechanisms from the screened dataset, and performing dimensionality reduction processing using principal component analysis to obtain a dimensionality-reduced feature set; for the dimensionality-reduced feature set, if the feature dimension is lower than a preset dimension, supplementing missing features by filling in the mean to obtain a standardized feature set; training the standardized feature set using a linear regression model, optimizing the regression coefficients using the least squares method to obtain a fitted mapping model; and calculating the mapping parameters from acoustic features to physiological mechanisms based on the fitted mapping model to obtain the parametric singing technique model.

[0071] In this embodiment, for data pairs with a quantitative correlation strength higher than a fourth preset threshold, rigorous data screening and feature optimization are first performed. Specifically, the fourth preset threshold is set to an absolute value of the Pearson correlation coefficient greater than 0.7. Data pairs with significant correlations are selected from the feature-signal pairing set using this threshold, forming a high-quality screening dataset. Subsequently, principal component analysis is performed to reduce the dimensionality of the multidimensional features in the screening dataset: the covariance matrix of acoustic and physiological features is calculated, eigenvalue decomposition is used to obtain eigenvectors, and the top k principal components with a cumulative contribution rate exceeding 85% are selected. The original high-dimensional features are projected into a low-dimensional space to obtain a dimensionality-reduced feature set. If the feature dimension after dimensionality reduction is still lower than a preset dimension (8 dimensions in this embodiment), the mean of that feature dimension in the training set is used to fill the gap, ultimately forming a standardized feature set. Based on the standardized feature set, a multiple linear regression model is used to establish the mapping relationship between acoustic features and physiological mechanisms. The model form is Y = Xβ + ε, where Y is the physiological mechanism index matrix, X is the acoustic feature design matrix, and β is the regression coefficient matrix to be determined. The objective function min‖Y-Xβ‖² was optimized using the least squares method, and the normal equation XᵀXβ=XᵀY was solved using QR decomposition to obtain the optimal regression coefficient estimates. To enhance model robustness, ridge regression regularization was introduced by adding a λ term (λ=0.01) to the diagonal of the XᵀX matrix to address multicollinearity. Model training employed 5-fold cross-validation, using mean squared error as the evaluation metric to ensure model generalization ability. The final parametric singing technique model contains a complete mapping parameter system: the regression coefficient matrix β quantitatively describes the contribution of each acoustic feature to the physiological mechanism, the intercept term represents the basic physiological level, and the model coefficients of determination R² reflect the overall explanatory power.

[0072] S5. Obtain different folk song singing voices, and analyze the relationship between the acoustics and physiological multimodalities of folk song singing voices based on the parametric singing technique model.

[0073] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.

[0074] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, it should be understood that the sequence number of each step in the above embodiments does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. The actions or steps recorded in the claims can be performed in a different order than that in the above embodiments and can still achieve the desired result. In addition, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0075] Example 2

[0076] In this embodiment, a multimodal analysis system for the acoustics and physiology of traditional folk song singing includes: a data acquisition module, a data extraction module, a correlation strength judgment module, a model construction module, and an analysis module.

[0077] The data acquisition module synchronizes and calibrates the voice recording device, voice monitoring device, and breathing sensor using a unified clock source to acquire multimodal signal data streams and obtain time-aligned raw datasets. The data extraction module processes the speech signal portion using Fourier transform based on the time-aligned raw datasets to determine acoustic feature vectors and extracts physiological signal sequences from the time-aligned raw datasets. In the correlation strength judgment module, if the timestamps of the acoustic feature vectors and physiological signal sequences match, the Pearson coefficient between them is calculated through correlation analysis to determine the quantized correlation strength. For the portion where the quantized correlation strength is higher than a fourth preset threshold, the model construction module uses a linear regression model to fit the mapping relationship between acoustic features and physiological mechanisms to obtain a parametric singing technique model. The analysis module is used to acquire different folk song singing voices and, based on the parametric singing technique model, analyzes the relationship between the acoustic and physiological multimodalities of folk song singing voices.

[0078] The system described in the above embodiments is used to implement the corresponding traditional folk song singing voice acoustic and physiological multimodal analysis method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0079] It should be noted that the aforementioned multimodal analysis system for the acoustics and physiology of traditional folk song singing is embodied in the form of functional units. The term "module" here can be implemented in software and / or hardware, without specific limitations.

[0080] For example, a "module" can be a software program, hardware circuit, or a combination of both that implements the above functions. Hardware circuits may include application-specific integrated circuits (ASICs), electronic circuits, processors (e.g., shared processors, proprietary processors, or group processors) and memory for executing one or more software or firmware programs, combined logic circuits, and / or other suitable components that support the described functions.

[0081] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A traditional folk song singing voice acoustic and physiological multi-modal analysis method, characterized in that, The method comprises the following steps: Synchronizing and calibrating the voice recording device, the voice monitoring device and the respiration sensor through a unified clock source, obtaining a multi-modal signal data stream, and obtaining a time-aligned original data set; According to the time-aligned original data set, processing the voice signal part by Fourier transform, determining the acoustic feature vector, and extracting the physiological signal sequence from the time-aligned original data set; If the timestamps of the acoustic feature vector and the physiological signal sequence match, calculating the Pearson coefficient between them through correlation analysis to determine the quantitative correlation strength; For the part where the quantitative correlation strength is higher than the fourth preset threshold, using a linear regression model to fit the mapping relationship between the acoustic features and the physiological mechanism to obtain a parameterized singing technique model; Obtaining different folk song singing voices, analyzing the relationship between the acoustic and physiological multi-modalities of the folk song singing voices according to the parameterized singing technique model.

2. The traditional folk song singing voice acoustic and physiological multi-modal analysis method according to claim 1, characterized in that, The method for obtaining the time-aligned original data set comprises: Obtaining a unified clock source reference signal, sending a synchronization clock pulse to the voice recording device, the voice monitoring device and the respiration sensor through clock distribution, obtaining the clock calibration state of each device; According to the clock calibration state, starting a multi-modal data acquisition process, the voice recording device acquires an audio signal, the voice monitoring device acquires acoustic parameters, and the respiration sensor captures a respiration waveform to obtain three parallel data streams; Using timestamp marking technology to mark the three parallel data streams, if the data packet timestamp difference exceeds the first preset threshold, triggering a resynchronization mechanism to obtain multi-modal signal data with accurate time marking; Through a buffer management mechanism, the multi-modal signal data with accurate time marking is temporarily stored, and each modal data packet is arranged according to the timestamp sequence to obtain the time-aligned original data set.

3. The traditional folk song singing voice acoustic and physiological multi-modal analysis method according to claim 1, characterized in that, The method for obtaining the acoustic feature vector comprises: Obtaining the time-synchronized original data set, processing the multi-channel voice signal through a data alignment algorithm to obtain a voice data sequence with a unified time reference, and removing noise interference and silent segments from the voice data sequence to obtain a purified voice signal segment; According to the purified voice signal segment, converting to a frequency domain space through a Fourier transform algorithm to obtain frequency spectrum distribution data; Based on the frequency spectrum distribution data, using a peak detection method to identify the main frequency components, determining a candidate set of fundamental frequency points, if there are multiple peaks in the candidate set of fundamental frequency points, filtering the real fundamental frequency through harmonic relationship verification to obtain a fundamental frequency sequence data; According to the frequency spectrum distribution data, using a linear predictive coding method to analyze the voice resonance characteristics to obtain the resonance peak frequency position; Through the fundamental frequency sequence data and the resonance peak frequency position, the acoustic feature vector is constructed.

4. The traditional folk song singing voice acoustic and physiological multi-modal analysis method according to claim 1, characterized in that, The method for extracting the physiological signal sequence comprises: Obtaining a time-synchronized original data set, performing format analysis on the original data to obtain a mixed data stream containing multiple sensor signals; According to the preset signal identifier, physiological signal data is separated from the mixed data stream, if the signal identifier matches the respiratory sensor marker, the corresponding respiratory-related data segment is extracted, and if the signal identifier matches the acoustic sensor marker, the corresponding voice-related data segment is extracted; The separated physiological signal data is filtered by a band-pass filter, high-frequency noise interference and low-frequency baseline drift are removed by setting a frequency range, and purified signal data is obtained; Respiratory cycle marker points in the purified signal data are identified by a peak detection algorithm, the respiratory frequency is calculated according to the time interval between adjacent peak values, and a respiratory rhythm sequence is obtained; The amplitude variation pattern of the purified signal data is analyzed to analyze the voice tension state, if the amplitude variation exceeds a second preset threshold, it is judged as an abnormal tension state, if the amplitude variation is within a normal range, it is judged as a normal tension state, and a voice tension sequence is determined; The respiratory rhythm sequence and the voice tension sequence are synchronized and aligned in a time window sliding manner, the corresponding relationship between the two sequences is established through timestamp matching, and the physiological signal sequence with consistent time is obtained.

5. The traditional folk song singing voice acoustic and physiological multi-modal analysis method according to claim 1, characterized in that, The method for judging the quantitative correlation strength comprises: Obtaining the data stream of the acoustic feature vector and the data stream of the physiological signal sequence, and extracting the respective time marker information; According to the extracted time marker information, the acoustic feature vector and the physiological signal sequence are time-aligned by a time window sliding matching method, and a synchronized data pair sequence is obtained; If the time difference in the synchronized data pair sequence is less than a third preset threshold, it is judged as an effective aligned data pair, and a time-matched feature-signal pairing set is obtained; The feature-signal pairing set is processed by a Pearson correlation coefficient calculation method to obtain the linear correlation values between each pairing; According to the distribution of the linear correlation values, the mean and variance of the overall correlation strength are calculated by a statistical analysis method, the comprehensive correlation degree of acoustic features and physiological signals is determined, and the quantitative correlation strength is obtained.

6. The traditional folk song singing voice acoustic and physiological multi-modal analysis method according to claim 1, characterized in that, The method for obtaining the parameterized singing technique model comprises: Filtering acoustic features and physiological mechanism data pairs with a quantitative correlation strength higher than a fourth preset threshold through data screening to obtain a screening data set; Extracting feature vectors of acoustic features and physiological mechanisms from the screening data set, and obtaining a reduced dimension feature set by principal component analysis dimension reduction processing; For the reduced dimension feature set, if the feature dimension is lower than a preset dimension, missing features are supplemented by mean filling to obtain a standardized feature set; The standardized feature set is trained by a linear regression model, and the regression coefficients are optimized by the least square method to obtain a fitting mapping model; According to the fitting mapping model, the mapping parameters of acoustic features to physiological mechanisms are calculated, and the parameterized singing technique model is obtained.

7. A traditional folk song singing voice acoustic and physiological multi-modal analysis system, the system applies the method of any one of claims 1-6, characterized in that, It comprises: a data acquisition module, a data extraction module, a correlation strength judgment module, a model construction module, and an analysis module; The data acquisition module synchronously calibrates the voice recording device, the voice monitoring device and the respiration sensor through a unified clock source, acquires a multi-modal signal data stream, and obtains a time-aligned original data set; The data extraction module processes a voice signal part by using Fourier transform according to the time-aligned original data set, determines an acoustic feature vector, and extracts a physiological signal sequence from the time-aligned original data set; In the correlation strength judgment module, if the timestamps of the acoustic feature vector and the physiological signal sequence match, a Pearson coefficient between the two is calculated through correlation analysis to judge the quantitative correlation strength; The model construction module adopts a linear regression model to fit a mapping relationship between the acoustic features and the physiological mechanism for the part with the quantitative correlation strength higher than a fourth preset threshold, and obtains a parameterized singing skill model; The analysis module is configured to acquire different folk song singing voices, analyze the relationship between the acoustic and physiological multi-modal of the folk song singing voices according to the parameterized singing skill model.

Citation Information

Cited By

  • Lung disease acoustic recognition method and device based on artificial intelligence

    CN121647647A