Immersive traditional cultural language audio feature extraction method based on artificial intelligence
By combining frequency phase residual matrix and spectral coupling tensor with perturbation guidance mechanism, the problem of extracting rhythmic and prosodic features in traditional cultural language audio is solved, and high-precision feature extraction and emotion recognition are achieved.
Patent Information
- Application Number
- CN202511121347.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-04
AI Technical Summary
Existing technologies struggle to accurately capture the expressive rhythmic and melodic features of traditional cultural languages, especially in the processes of emotion recognition and semantic understanding, where spectrograms fail to effectively integrate frequency and phase information.
By constructing a frequency-phase residual matrix and a spectral coupling tensor, and combining nonlinear reconstruction and sliding window variability assessment, a perturbation guidance mechanism is introduced to achieve joint modeling of spectral amplitude and phase information and enhancement of structural guidance features.
It improves the accuracy and modeling stability of audio feature extraction for traditional cultural languages, enhances the ability to model complex speech rhythms and cultural context changes, and improves the accuracy and adaptability of emotion recognition.
Smart Images

Figure CN120895025A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of feature extraction technology, and in particular to an artificial intelligence-based method for extracting audio features of immersive traditional cultural languages. Background Technology
[0002] Currently, AI-driven speech signal processing technology has been widely applied in multiple fields such as speech recognition, emotion recognition, and semantic understanding, becoming a key module in human-computer interaction systems. As the trend of digital culture development becomes increasingly apparent, the demand for digital expression of traditional cultural languages in education, cultural tourism, and intangible cultural heritage dissemination scenarios is also growing. In order to achieve an immersive human-computer interaction experience, high-quality extraction of emotional, stylistic, and rhythmic features from traditional cultural language audio has become a key issue that urgently needs to be addressed.
[0003] Patent CN109875832A proposes a speech emotion recognition method based on logarithmic Mel spectrogram and convolutional neural network. This method extracts the acoustic features of the audio signal, performs feature fusion, and then inputs the feature into a deep learning model to classify emotions.
[0004] However, this method does not take into account the dynamic changes in phase information, making it difficult to capture the expressive rhythmic and melodic features of traditional cultural languages. Summary of the Invention
[0005] This invention provides an AI-based immersive traditional cultural language audio feature extraction method, aiming to achieve joint modeling of spectral amplitude and frequency phase information and structure-guided feature enhancement. The method first acquires traditional cultural language audio materials and emotion tag information, generating a structured audio dataset containing a logarithmic Mel spectrogram and a frequency phase spectrogram. Then, it constructs an amplitude-aware phase residual matrix through phase difference calculation and nonlinear reconstruction, and establishes a frequency phase traction factor matrix based on residual variability. A guided tensor coupling coherent reconstruction mechanism is introduced to fuse the spectrogram and phase information to construct a spectral coupling tensor. The spectral coupling tensor is input to the feature extraction module, completing nonlinear transformation and dynamic response convolution to generate a multi-dimensional deep feature map. Finally, a perturbation-guided mechanism is used to complete feature weighting aggregation and category mapping, improving the accuracy and modeling stability of traditional cultural language audio feature extraction.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an immersive traditional cultural language audio feature extraction method based on artificial intelligence, the specific steps of which are as follows: S1. Obtain audio materials of traditional cultural languages and corresponding emotion tags, preprocess the audio materials, extract logarithmic Mel spectrograms and frequency-phase spectrograms, and construct a structured audio dataset of traditional cultural languages; S2. Calculate the phase difference between adjacent frames based on the frequency phase spectrum to obtain the frequency phase residual information; use a nonlinear reconstruction function combined with the frequency band amplitude variability to modulate the frequency phase residual information, and convert it into a phase residual matrix with the same dimension as the logarithmic Mel spectrum through a structural coupling mapping mechanism. S3. Define a sliding window region on the phase residual matrix, extract local variation features, construct a frequency phase residual traction factor matrix, and normalize the matrix to obtain a normalized traction factor matrix. S4. Calculate the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term based on the logarithmic Mel spectrum, and construct the spectral stress excitation tensor; modulate and enhance the spectral stress excitation tensor based on the normalized traction factor matrix, and construct a multi-channel spectral coupling tensor by combining the channel mapping function; S5. Couple the multi-channel spectrum tensor into the feature extraction module, perform nonlinear transformation and dynamic response convolution to obtain a multi-dimensional deep feature map; based on the perturbation guidance mechanism, perform weighted aggregation on the multi-dimensional deep feature map to obtain the emotion tags corresponding to the audio materials of traditional cultural language. S6. Construct an AI-based model for extracting features from traditional cultural language audio. Input the traditional cultural language audio dataset, use cross-entropy as the loss function, and execute steps S2 to S5 sequentially. Use emotion labels for supervised training until the training process converges, thus completing the feature extraction of traditional cultural language audio.
[0007] Preferably, in step S1, the construction of the traditional cultural language audio dataset involves first acquiring language audio materials with traditional cultural expression characteristics, including language data categorized as recitation, reading aloud, storytelling, and reading aloud; then, according to preset classification criteria, assigning emotion tags to the audio materials, with the tags corresponding to common emotion expression types in traditional cultural contexts, including six categories: happy, calm, solemn, dignified, sorrowful, and passionate; next, performing format unification and signal cleaning on the audio materials to obtain clear and structurally stable signal data; then, performing frame segmentation and window function weighting operations on the processed speech signal, and performing short-time Fourier transform processing to extract the amplitude and phase spectrum features of the speech in the time-frequency domain; in terms of spectral amplitude processing, weighted calculations are performed with a set of preset Mel filters to generate multi-channel Mel frequency response results, and logarithmic operations are performed on them to construct a logarithmic Mel. The spectrum diagram is generated; in terms of frequency and phase processing, the phase angles after Fourier transform are extracted to form a frequency and phase spectrum diagram; finally, the above spectrum diagram results are bound one by one with the corresponding emotion tag information, and indexed and organized in a unified format to construct a structured traditional cultural language audio dataset for subsequent modeling tasks.
[0008] Preferably, traditional cultural language audio has the characteristics of obvious rhythm jumps and unstable phase perturbations during the expression process, which often leads to drastic fluctuations in frequency information between time frames, making it difficult to accurately capture through traditional amplitude feature modeling. To address this problem, this invention introduces a nonlinear reconstruction mechanism to perform centered modulation processing on the inter-frame phase difference to enhance the ability to express the frequency perturbation structure. The reconstructed phase residual features are then mapped to the Mel spectrum dimension through structural coupling, providing perturbation-aware support for subsequent tensor construction.
[0009] Furthermore, in step S2, the specific method for establishing the phase residual matrix is as follows: S21. Calculate the phase difference between adjacent frames based on the frequency phase spectrum to obtain the inter-frame frequency phase residual information; S22. Based on the frequency phase residual information, a nonlinear reconstruction function combined with the frequency band amplitude variability is used to modulate the frequency phase residual. The mathematical model of the reconstruction function is as follows: ; in, Indicates the index of the time frame. Indicates the frequency channel index. The average phase value within the frequency band. The standard deviation of the residuals in the frequency direction. For the first Frame, First Phase residual information at the frequency, This refers to the frequency phase residual information after modulation. S23. The frequency-phase residual information reconstructed nonlinearly is mapped to the frequency dimension based on a structural coupling mapping mechanism. A weighted coupling model is used to construct a phase residual matrix with the same dimension as the logarithmic Mel spectrum. The mathematical model is as follows: ; in, This represents the Mel spectrum dimension channel number corresponding to the structure mapping. For frequency aisle Mapping weights between them This is the phase residual matrix.
[0010] Furthermore, in step S2, by calculating the phase difference between adjacent time frames in the frequency-phase spectrum, the inter-frame frequency-phase residual information is extracted. This helps to characterize the subtle rhythmic jumps and phase perturbation characteristics of speech flow boundaries during the pronunciation of traditional cultural languages, enhancing the model's response sensitivity to the prosodic structure of traditional cultural languages. For the frequency-phase residual signal, a nonlinear reconstruction function is introduced for reconstruction processing. Normalization enhancement is performed by combining the phase mean and amplitude variability of the frequency band, improving the expression accuracy of the phase residual in the amplitude-sensitive region and enhancing the feature extraction capability for the highly fluctuating audio intonation of traditional cultural languages. Finally, the modulated phase residual signal is mapped to the Mel spectrum using a structural coupling mapping mechanism. Figure One By employing a channel structure coupling mapping mechanism to achieve frequency alignment, the structural dimensions of the spectrum amplitude and phase perturbation are homogeneously fused, laying a structurally unified feature foundation for subsequent spectrum stress response modeling.
[0011] Preferably, traditional cultural language audio often exhibits nonlinear changes, rhythmic abrupt changes, and regional frequency perturbations during semantic expression. This characteristic leads to poor local stability of the frequency domain structure, making it difficult to extract important regional information of phase perturbations through fixed window modeling. To address this issue, this invention defines a sliding window region on the phase residual matrix, extracts statistical variability features within local frequency segments, and establishes a spectral variation coefficient to quantify the discreteness of frequency phase perturbations. This facilitates the dynamic identification of frequency bands with significant perturbations in speech segments, improving the sensitivity and structural discriminability of subsequent traction factor modeling.
[0012] Furthermore, in step S3, the specific method for constructing the frequency-phase residual traction factor matrix is as follows: S31. Define a sliding window local region on the phase residual matrix, and calculate the variation coefficient for the frequency phase residual signal within each local region. The mathematical model is as follows: ; in, Position of the sliding window The phase residual block inside, For variance, The mean, A small, positive constant. S32. Based on the coefficient of variation and the frequency sensitivity rate adjustment mechanism, construct the frequency phase residual traction factor matrix. The mathematical model for the frequency-phase residual traction factor is: ; in, For the response sensitivity coefficient, The average intensity of the frequency gradient, This is a frequency disturbance suppression term; S33, Regarding the traction factor matrix After normalization, the normalized frequency-phase residual traction factor matrix is obtained. .
[0013] Furthermore, in step S3, by defining a local region of a sliding window on the phase residual matrix, the frequency perturbation variation features within the local region are extracted, and a frequency-phase residual traction factor matrix is constructed. This helps to quantify the perturbation sensitivity of different frequency bands in traditional cultural language audio. First, the variation coefficient of the phase residual signal within the window is calculated to measure the discreteness of local perturbations in the frequency domain, thereby identifying high-variation regions of rhythmic abrupt changes and accent jumps in traditional cultural language audio segments. Subsequently, combined with the gradient response information in the frequency direction, a traction factor matrix is constructed through a frequency sensitivity rate adjustment mechanism to achieve responsive weighting of perturbation intensity in different frequency bands, improving the model's ability to focus on the structure of key intonation changes. Finally, the traction factor matrix is normalized to unify the dimensions of perturbation intensity, ensuring numerical stability and modeling consistency in the subsequent spectral coupling tensor modulation process, thereby enhancing the structural perception effect of spectrum-phase joint modeling on the prosodic style of traditional cultural language.
[0014] Preferably, traditional cultural language audio often exhibits rhythmic abrupt changes and drastic phase perturbation fluctuations during expression, leading to a structural disconnect between spectral amplitude and phase information. To address this issue, this invention proposes a spectral coupling tensor construction mechanism. By introducing a modulation enhancement strategy, it effectively integrates the spectrogram and phase traction factor, solving the problem of insensitivity of spectral features to phase perturbation response and improving the model's ability to model complex speech rhythms and cultural context changes.
[0015] Furthermore, in step S4, the specific method for constructing the spectral coupling tensor is as follows: S41. Based on the Mel spectrum diagram, calculate the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term to construct the initial spectral stress excitation tensor. The mathematical model of the spectral stress excitation tensor is as follows: ; in, For the logarithmic Mel spectrum, the first... The value at each position, For the principal stress terms of the local spectrum, For time-oriented variables, For spectral interaction coupling terms, For time-oriented local derivatives, For element-wise multiplication, , , This is the adjustment coefficient for the spectral stress tensor. For the spectral stress excitation tensor; S42. Based on the frequency-phase residual traction factor matrix, the spectral stress excitation tensor is modulated and enhanced to obtain the modulated spectral response tensor. The mathematical model is as follows: ; in, , For intensity adjustment parameters, This represents the energy value of the phase residual block. The modulated spectral response tensor; S43. Input the modulated spectral response tensor into the channel mapping function, project it from the frequency dimension to the structural domain dimension, and construct a multi-channel spectral coupling tensor.
[0016] Furthermore, in step S4, the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term are calculated using a log-Melbourne spectrogram to establish a spectral stress excitation tensor. This helps to accurately characterize the energy concentration trend and modal variation structure of traditional cultural language audio in the local time-frequency region, enhancing the model's ability to model traditional rhythmic levels and semantic style fluctuations. Subsequently, based on the frequency-phase residual traction factor matrix, a modulation enhancement mechanism is introduced, using the response intensity of the phase perturbation sensitive region as an adjustment factor to dynamically modulate the spectral stress excitation tensor, improving the adaptability and stability of the spectral response structure to frequency domain perturbations. Combining the amplitude response information of the phase residual signal, the strength of the spectral response tensor is controlled through dual adjustment parameters, enhancing the structural focus on highly variable segments in complex speech streams. Finally, the modulated spectral response tensor is projected from the frequency dimension to the structural domain dimension to construct a multi-channel spectral coupling tensor, realizing the synergistic fusion input of spectral and phase information, providing a structurally consistent and responsively sensitive time-frequency coupled representation basis for the subsequent feature extraction module.
[0017] Preferably, audio recordings of traditional cultural languages often exhibit significant differences in semantic style, frequent rhythmic jumps, and complex temporal and frequency modal structures during expression. Traditional static or equalization feature extraction methods struggle to accurately capture these internal perturbation changes. To address these issues, this invention introduces a perturbation-guided weighted aggregation mechanism during feature extraction. This mechanism utilizes structural perturbation-aware features to enhance the response accuracy to nonlinear intonation jumps and improves the sensitivity of feature compression expression to emotion discrimination through a channel-guided aggregation strategy. This achieves adaptability and robustness in modeling highly variable regions of audio recordings of traditional cultural languages.
[0018] Furthermore, in step S5, the specific method for weighted aggregation of multidimensional deep feature maps based on the perturbation-guided mechanism is as follows: S51. Input the multi-channel spectrum coupling tensor into the feature extraction module, perform a nonlinear transformation on it, and generate the first intermediate response tensor. S52. Perform dynamic response convolution operation on the first intermediate response tensor to obtain the structural perturbation sensing feature map; S53. Stack the structural perturbation-aware feature maps according to the channel dimension to construct a multi-dimensional deep feature map, and perform perturbation-guided weighted aggregation operation on the multi-dimensional deep feature map to obtain the channel aggregation output vector. The mathematical model is as follows: ; in, for The characteristic response compression function, where c is the graph channel index. For the c-th channel, the first... Location-based structural perturbation sensing features; The disturbance sensing factor for channel c; H represents the local structural fuzzy response adjustment term; H and W represent the height and width of the feature map. Furthermore, in step S53, the specific method for constructing the local structural fuzzy response adjustment term is as follows: S531. Based on the multi-dimensional deep feature map, perform directional gradient magnitude extraction operation within its local spatial neighborhood to obtain the spatial location. Disturbance structure strength value ; S532, The strength value of the disturbed structure The input is fed into the response smoothing adjustment function to obtain the local structural fuzzy response adjustment term at the corresponding location. The mathematical model of the response regulation function is as follows: ; in, Indicates position The strength value of the disturbed structure; This is the slope factor of the adjustment function. The threshold for structural fuzziness discrimination; Furthermore, in step S53, by constructing a channel perturbation perception factor and a local structural fuzzy response adjustment term, a weighted fusion of high-frequency perturbation regions and structurally ambiguous regions in the multidimensional deep feature map is achieved, which is beneficial to enhancing the ability to focus on key emotional expression positions in traditional cultural language audio. Among them, the channel perturbation perception factor is used to measure the significance of each channel in response intensity, improving the modeling accuracy of high-abrupt intonation; the local structural fuzzy response adjustment term is used to describe the degree of fuzziness of the perturbation direction gradient in the spatial neighborhood. Through a nonlinear smooth adjustment function, it guides the moderate suppression of structurally uncertain regions, effectively avoiding overfitting to the noise response of non-critical perturbation regions, thereby ensuring the stability of the overall feature expression and the accuracy of the emotion classification results. S54. Perform a normalized exponential mapping operation on each element of the channel aggregation output vector to form a normalized output probability vector. and to Take the maximum index to obtain the emotion tag corresponding to the audio material of traditional cultural language. .
[0019] Furthermore, in step S5, by performing nonlinear transformation and dynamic response convolution operations on the spectral coupling tensor, a structural perturbation-aware feature map is extracted. This is beneficial for mining the time-frequency response features of highly variable regions in traditional cultural language audio, enhancing the model's structural perception ability of rhythmic jumps and semantic prosodic differences. The structural perturbation-aware feature map is used as input to a multi-dimensional deep feature map, and a weighted aggregation operation is performed by combining the channel perturbation-aware factor and the local structural fuzzy response adjustment term to highlight the salient features of key emotional segments in traditional cultural language, improving the model's discriminative ability and structural stability in the emotion classification process. Finally, the corresponding emotion label is output through the maximum value index method to ensure the accuracy of the emotion recognition results.
[0020] Preferably, in S6 In this process, addressing the complex emotional expression dimensions, significant rhythmic style differences, and highly nonlinear feature structures in traditional cultural language audio, an immersive feature extraction model based on artificial intelligence is constructed. The model sequentially executes training stages including perturbation structure encoding modeling, spectrum-phase co-fusion, and response aggregation classification mapping. First, a dataset of traditional cultural language audio containing emotion labels is input. In the perturbation structure encoding modeling stage, multimodal feature representations are extracted through frequency-phase residual modeling and spectral coupling tensor construction, forming structurally consistent feature inputs. Subsequently, based on the structural coupling relationship between spectral amplitude and phase perturbation, a co-fusion operation is performed to enhance the structural representation ability of key rhythmic and emotional transition segments. In the classification response stage, a perturbation-guided weighting mechanism is used to complete channel aggregation of deep feature maps, and emotion labels corresponding to the traditional cultural language audio materials are obtained through a category mapping function. Finally, using cross-entropy as the loss function, supervised training is conducted using emotion labels. Under an end-to-end training framework, combined with iterative optimization strategies and perturbation structure guidance mechanisms, the model achieves robust extraction of traditional cultural language audio features and high-precision recognition of emotional states, improving the model's adaptability and discriminative performance in multi-style traditional language expression environments.
[0021] Compared with existing technologies, this invention constructs an immersive traditional cultural language audio feature extraction method based on artificial intelligence, proposing a spectrum-phase joint modeling and structural perturbation guidance mechanism. Since traditional modeling relying solely on amplitude spectrum features for emotion recognition has limitations, this invention proposes to introduce a collaborative construction mechanism of frequency phase perturbation and spectral coupling tensor into the traditional cultural speech and audio feature extraction task. Frequency phase residual information is constructed through the phase difference between adjacent frames, and nonlinear transformation combined with frequency band amplitude variability is introduced to perform reconstruction modulation, establishing a phase residual matrix under structural constraints, providing phase perturbation information support for subsequent modeling with minute rhythmic jumps. A sliding window variability evaluation strategy is introduced to quantize the frequency... To enhance the model's responsiveness to highly variable rhythmic regions, a frequency-phase traction factor matrix is constructed and normalized to improve domain perturbation sensitivity. A spectral coupling tensor is constructed by combining the logarithmic Mel spectrum and the frequency-phase traction factor to achieve deep spectrum-phase fusion, improving feature uniformity and response structure consistency at the tensor level. A perturbation-aware factor and a structural fuzzy response adjustment mechanism are introduced to implement response-weighted aggregation at the multi-channel feature map level, highlighting the significant expressive features of high-jump regions in traditional languages and suppressing non-critical perturbation interference. Finally, an end-to-end model framework is constructed, combining cross-entropy supervision and structural perturbation feedback strategies to optimize model parameters and output emotional labels for traditional cultural language audio features. This invention effectively improves the accuracy of structural feature extraction caused by nonlinear rhythms and prosodic style differences in traditional cultural language audio, while enhancing the model's adaptability and classification stability to highly variable expressions in cultural languages. Attached Figure Description
[0022] Figure 1 This is a flowchart of the immersive traditional cultural language audio feature extraction method based on artificial intelligence provided by the present invention.
[0023] Figure 2 This is a structural diagram of establishing the phase residual matrix provided by the present invention.
[0024] Figure 3 This is a structural diagram of the frequency-phase residual traction factor matrix provided by the present invention.
[0025] Figure 4 This is a structural diagram of the modeling of spectral coupling tensors provided by the present invention.
[0026] Figure 5 This invention provides a structural diagram for weighted aggregation of multidimensional deep feature maps based on a perturbation-guided mechanism.
[0027] Figure 6 This invention provides an emotion classification confusion matrix diagram based on artificial intelligence for immersive traditional cultural language audio feature extraction.
[0028] Figure 7This is a graph showing the label prediction accuracy results of the immersive traditional cultural language audio feature extraction based on artificial intelligence provided by the present invention. Detailed Implementation
[0029] This invention provides an AI-based immersive traditional cultural language audio feature extraction method, aiming to achieve joint modeling of spectral amplitude and frequency phase information and structure-guided feature enhancement. The method first acquires traditional cultural language audio materials and emotion tag information, generating a structured audio dataset containing a logarithmic Mel-level spectrogram and a frequency phase spectrogram. Then, it constructs an amplitude-aware phase residual matrix through phase difference calculation and nonlinear reconstruction, and establishes a frequency phase traction factor matrix based on residual variability. A guided tensor coupling coherent reconstruction mechanism is introduced to fuse the spectrogram and phase information to construct a spectral coupling tensor. The spectral coupling tensor is input to the feature extraction module, completing nonlinear transformation and dynamic response convolution to generate a multi-dimensional deep feature map. Finally, a perturbation-guided mechanism is used to complete feature weighting aggregation and category mapping, improving the accuracy and modeling stability of traditional cultural language audio feature extraction.
[0030] Please see Figure 1 As shown in the embodiments of this application, the specific steps of the immersive traditional cultural language audio feature extraction method based on artificial intelligence are as follows.
[0031] S1. Obtain audio materials of traditional cultural languages and corresponding emotion tags, preprocess the audio materials, extract logarithmic Mel spectrograms and frequency-phase spectrograms, and construct a structured audio dataset of traditional cultural languages.
[0032] Furthermore, the construction process of the traditional cultural language audio dataset includes three stages: audio acquisition, preprocessing, and structured organization. In the data acquisition stage, audio materials of traditional cultural languages, including chanting, recitation, storytelling, and reading aloud, were selected. All materials were acquired in a quiet indoor environment, and the recording used single-channel linear PCM format with a sampling precision of 16 bits and an original sampling rate of 44.1 kHz. In the preprocessing stage, the data was first uniformly resampled to 16 kHz to ensure consistent temporal resolution of the model input. Then, pre-emphasis processing and frame segmentation were performed, with a frame length of 25 ms and a frame shift of 10 ms. A Hamming window function was applied to each frame to reduce edge effects. Afterward, a short-time Fourier transform was performed to extract the complex spectrum of each frame. In the amplitude spectrum processing, the number of Mel filter banks was set to 40, and the filtering frequency band was from 300 Hz to 3400 Hz. The spectrum was weighted by Mel and then logarithmically compressed to construct a log-Mel spectrogram. During phase spectrum processing, the phase angles in the Fourier transform results were extracted to construct a frequency-phase spectrogram, ensuring alignment with the amplitude spectrogram in both time and frequency dimensions. Emotional labels were independently labeled by three experts with language background knowledge, and the label categories included six types: happy, calm, solemn, dignified, sorrowful, and excited. The processed spectrogram data and corresponding labels were organized in tensor form to construct a structured traditional cultural language audio dataset, which served as the input basis for subsequent model training.
[0033] S2. Calculate the phase difference between adjacent frames based on the frequency phase spectrum to obtain the frequency phase residual information; use a nonlinear reconstruction function combined with the frequency band amplitude variability to modulate the frequency phase residual information, and convert it into a phase residual matrix with the same dimension as the logarithmic Mel spectrum through a structural coupling mapping mechanism.
[0034] Furthermore, in step S2, the phase residual matrix is established, and the process is as follows: Figure 2 As shown, the specific steps for establishing the phase residual matrix are as follows.
[0035] S21. Calculate the phase difference between adjacent frames based on the frequency phase spectrum to obtain the inter-frame frequency phase residual information; In this embodiment, adjacent time frames of the frequency phase spectrum are... and The formula for calculating the inter-frame frequency phase residual information is as follows: Phase difference calculation is performed between frames. ; in, For the first Frame number The phase angle at each frequency channel is expressed in radians (rad) and ranges from -π to π.
[0036] S22. Based on the frequency phase residual information, a nonlinear reconstruction function combined with the frequency band amplitude variability is used to modulate the phase residual. The mathematical model of the reconstruction function is as follows: ; in, In this embodiment, the index representing the time frame is... ,in The total number of time frames. In this embodiment, the frequency channel index is indicated. F=40 The average phase value within the frequency band. The standard deviation of the residuals in the frequency direction. For the first Frame, First Frequency phase residual information at the frequency channel, This refers to the frequency phase residual information after modulation. In this embodiment, the total number of time frames Total length of input speech The value is determined by the frame shift size, and the calculation formula is as follows: ; in, In this embodiment, the total length of the input speech is... =2s, In this embodiment, the number of points per frame is... =400, In this embodiment, the frame shift point is the number of points. =160; In this embodiment, the average phase value within the frequency band The calculation formula is: ; In this embodiment, the residual standard deviation in the frequency direction The calculation formula is: .
[0037] S23. The frequency-phase residual information reconstructed nonlinearly is mapped to the frequency dimension based on a structural coupling mapping mechanism. A weighted coupling model is used to construct a phase residual matrix with the same dimension as the logarithmic Mel spectrum. The mathematical model is as follows: ; in, This represents the Mel spectrum dimension channel number corresponding to the structure mapping, in this embodiment... M=40 For frequency aisle Mapping weights between them The phase residual matrix; In this embodiment, mapping weights The weighting is given by the triangular weighting of the standard Mel filter bank. The calculation formula is: ; in Mel indicates the number of The center frequency of each filter is calculated using the following formula: ; in, In this embodiment, the number of FFT points is... =512, The sampling rate is used in this embodiment. =16000 Hz, B is the number of filters, in this embodiment, B=40. The minimum value of the center frequency is, in this embodiment... = 300 Hz, = 3400 Hz.
[0038] S3. Define a sliding window region on the phase residual matrix, extract local variation features, construct a frequency phase residual traction factor matrix, and normalize the matrix to obtain a normalized traction factor matrix.
[0039] Furthermore, in step S3, the frequency-phase residual traction factor matrix is constructed, and the process is as follows: Figure 3 As shown, the specific steps for constructing the frequency-phase residual traction factor matrix are as follows.
[0040] S31. Define a sliding window local region on the phase residual matrix, and calculate the variation coefficient for the frequency phase residual signal within each local region. The mathematical model is as follows: ; in, Position of the sliding window In this embodiment, the phase residual block within the circuit... For variance, The mean, A small positive constant is set in this embodiment. ; In this embodiment, the phase residual block The value can be: ; in, This represents the phase residual value of the nth pixel, where N is the window size. , Set the size of the sliding window in the frequency direction. = 5、 Set the size of the sliding window in the time direction. = 3; In this embodiment, mean The calculation formula is: ; In this embodiment, variance The calculation formula is: .
[0041] S32. Based on the coefficient of variation and the frequency sensitivity rate adjustment mechanism, construct the frequency phase residual traction factor matrix. The mathematical model for the frequency phase residual traction factor is as follows: ; in, In this embodiment, the response sensitivity coefficient is... =1.5, The average intensity of the frequency gradient, As a frequency disturbance suppression term, in this embodiment, it is set to... =2; In this embodiment, the average intensity of the frequency gradient The calculation formula is: ; Where k is the local offset step size within the sliding frequency window, ranging from 1 to... , For the first The frequency channel, the first Frame frequency phase residual information, For the first The frequency channel, the first Frame frequency phase residual information.
[0042] S33, Regarding the traction factor matrix After normalization, the normalized frequency-phase residual traction factor matrix is obtained. ; In this embodiment, to unify the numerical range, the traction factor matrix is subjected to max-min normalization, and the calculation formula is as follows: ; in, It is the minimum value in the entire frequency phase residual traction factor matrix. This is the maximum value in the entire frequency phase residual traction factor matrix. The normalized frequency-phase residual traction factor matrix. .
[0043] S4. Calculate the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term based on the logarithmic Mel spectrum to construct the spectral stress excitation tensor; modulate and enhance the spectral stress excitation tensor based on the normalized traction factor matrix, and construct a multi-channel spectral coupling tensor by combining the channel mapping function.
[0044] Furthermore, in step S4, the frequency-phase residual traction factor matrix is constructed, and the process is as follows: Figure 4 As shown, the specific steps for constructing the spectral coupling tensor are as follows.
[0045] S41. Based on the Mel spectrum diagram, calculate the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term to construct the initial spectral stress excitation tensor. The mathematical model of the spectral stress excitation tensor is as follows: ; in, For the logarithmic Mel spectrum, the first... The value at each position, For the principal stress terms of the local spectrum, For time-oriented variables, For spectral interaction coupling terms, For time-oriented local derivatives, For element-wise multiplication, , , The spectral stress tensor adjustment coefficient is set to in this embodiment. =1, =0.6, =0.4, For spectral stress excitation tensor; In this embodiment, The calculation formula is: ; in, Let be the spectral energy value corresponding to the m-th Mel filter channel in the t-th frame; In this embodiment, The calculation formula is: ; in, The complex result of the STFT at the f-th frequency point in the t-th frame. The power spectrum of the t-th frame at the f-th frequency point; In this embodiment, the complex result of the STFT at the f-th frequency point of the t-th frame. The calculation formula is: ; In this embodiment, the power spectrum of the t-th frame at the f-th frequency point The calculation formula is: ; in, for The real part value, for The imaginary part of the value; In this embodiment, according to It can be seen that, The calculation formula is: ; The calculation formula is: ; In this embodiment, the local derivative in time orientation The calculation formula is: ; S42. Based on the frequency-phase residual traction factor matrix, the spectral stress excitation tensor is modulated and enhanced to obtain the modulated spectral response tensor. The mathematical model is as follows: ; in, , In this embodiment, the intensity adjustment parameter is set as follows: =0.5、 =0.3, This represents the energy value of the phase residual block. The modulated spectral response tensor; In this embodiment, the energy value of the phase residual block The calculation formula is: .
[0046] S43. Input the modulated spectral response tensor into the channel mapping function, project it from the frequency dimension to the structural domain dimension, and construct a multi-channel spectral coupling tensor. In this embodiment, the spectral response tensor is projected from the frequency dimension f to the structural domain dimension m to construct a multi-channel spectral coupling tensor. The calculation formula is: ; in, For tensor It belongs to a real matrix of dimension T×M. It is the field of real numbers.
[0047] S5. The multi-channel spectrum is coupled with a tensor and input into the feature extraction module. Nonlinear transformation and dynamic response convolution are performed to obtain a multi-dimensional deep feature map. The multi-dimensional deep feature map is weighted and aggregated based on a perturbation guidance mechanism to obtain emotion tags corresponding to traditional cultural language audio materials.
[0048] Furthermore, in step S5, a weighted aggregation of the multidimensional deep feature map is performed based on a perturbation-guided mechanism, the process of which is as follows: Figure 5 As shown, the specific steps are as follows.
[0049] S51. Input the multi-channel spectrum coupling tensor into the feature extraction module, perform a nonlinear transformation on it, and generate the first intermediate response tensor. ; Where c is the graph channel index, In this embodiment, C is set to 64; Let be the first intermediate response tensor of the c-th channel at position (t, m); Let be the value of the c-th channel in the input spectral coupling tensor at position (t, m+k); The offset index of the sliding window is used to define the local receptive field of the convolution kernel on the Mel channel m, K=1; For channel c, the weight parameter is the kth kernel weight parameter in the convolution kernel. This is the bias term for channel c; For activation functions; In this embodiment, the kernel weight parameter With bias term It is obtained through iterative updates using the backpropagation algorithm and optimizer during the training process. Its mathematical model is as follows: ; ; Where η is the learning rate, which is set to 0.0005 in this embodiment. The cross-entropy loss function; Cross-entropy loss function The mathematical model is as follows: ; in, According to the emotion category, =6; This is a one-hot encoded vector of the true emotion label. The value is 1 at the index position of the one-hot encoded vector corresponding to the emotion label, and 0 at the other positions. In this embodiment, the emotion categories are happy, calm, solemn, dignified, sorrowful, and excited, and the corresponding positions are... ; This is the i-th normalized output probability vector; In this implementation, the activation function The calculation formula is: ; Where x is the value inside the input activation function.
[0050] S52. Perform dynamic response convolution operation on the first intermediate response tensor to obtain the structural perturbation sensing feature map; In this embodiment, the structural disturbance sensing map By each position Structural perturbation sensing feature values The structural disturbance sensing feature value is obtained based on the weighted calculation of local directional differences. The calculation formula is: ; in, The spectral coupling tensor at the c-th channel position The value, Let be the value of the spectral coupling tensor at the c-th channel position (t+i, m+j). In this embodiment, the directional harmonic coefficient is used. ; In this embodiment, the specific method for constructing the structural disturbance sensing feature map is as follows: ;
[0051] S53. Stack the structural perturbation-aware feature maps according to the channel dimension to construct a multi-dimensional deep feature map, and perform perturbation-guided weighted aggregation operation on the multi-dimensional deep feature map to obtain the channel aggregation output vector. The mathematical model is as follows: ; in, for The characteristic response compression function; c is the graph channel index. In this embodiment, C is set to 64; For the c-th channel, the first... Location-based structural perturbation sensing features; The disturbance sensing factor for channel c; H and W are the local structural fuzzy response adjustment terms; H and W are the height and width of the feature map, and in this embodiment, H=32 and W=32 are set. In this embodiment, the feature response compression function The calculation formula is: ; in, The adjustment coefficient is set to 3 in this embodiment; In this embodiment, the disturbance sensing factor The calculation formula is: ; in, For A 3×3 local neighborhood window centered on the target; For the current point In this embodiment, the structural perturbation sensing feature values at other locations within the same neighborhood are... , .
[0052] Furthermore, in step S53, a local structural fuzzy response adjustment term is constructed, specifically as follows: S531. Based on the multi-dimensional deep feature map, perform directional gradient magnitude extraction operation within its local spatial neighborhood to obtain the spatial location. Disturbance structure strength value ; In this embodiment, the strength value of the disturbed structure The calculation formula is: ; in, The is the direction weighting coefficient. In this embodiment, the principal axis direction is set to 1, and the diagonal direction is set to 0.5.
[0053] S532, The strength value of the disturbed structure The input is fed into the response smoothing adjustment function to obtain the local structural fuzzy response adjustment term at the corresponding location. The mathematical model of the response regulation function is as follows: ; in, Indicates position The strength value of the disturbed structure; In this embodiment, the slope factor of the adjustment function is set as follows: =6, The threshold for structural fuzzy discrimination is set to . =0.35.
[0054] S54. Perform a normalized exponential mapping operation on each element of the channel aggregation output vector to form a normalized output probability vector. and to Take the maximum index to obtain the emotion tag corresponding to the audio material of traditional cultural language. ; In this embodiment, a normalized exponential mapping operation is performed on each element of the channel aggregation output vector to form a normalized output probability vector. The calculation formula is: ; in, For the i-th normalized output probability vector, ; In this embodiment, emotion labels are obtained. The calculation formula is: .
[0055] S6. Construct an AI-based model for extracting features from traditional cultural language audio. Input the traditional cultural language audio dataset, use cross-entropy as the loss function, and execute steps S2 to S5 sequentially. Use emotion labels for supervised training until the training process converges, thus completing the feature extraction of traditional cultural language audio.
[0056] In step S6, an AI-based immersive traditional cultural language audio feature extraction model is constructed, and training stages such as perturbation structure encoding modeling, spectrum-phase co-fusion, and response aggregation classification mapping are executed sequentially. First, a traditional cultural language audio dataset containing emotion labels is input. In the perturbation structure encoding modeling stage, multimodal feature representations are extracted through frequency-phase residual modeling and spectral coupling tensor construction, forming structurally consistent feature inputs. Subsequently, based on the structural coupling relationship between spectral amplitude and phase perturbation, a co-fusion operation is performed to enhance the structural expressive ability of key rhythmic and emotional transition segments. In the classification response stage, channel aggregation of deep feature maps is completed using a perturbation-guided weighting mechanism, and emotion labels corresponding to the traditional cultural language audio materials are obtained through a category mapping function. Finally, using cross-entropy as the loss function, supervised training is conducted using emotion labels. Under an end-to-end training framework, combined with iterative optimization strategies and perturbation structure guidance mechanisms, the model achieves robust extraction of traditional cultural language audio features and high-precision recognition of emotional states, improving the model's adaptability and discriminative performance in multi-style traditional language expression environments.
[0057] Furthermore, in step S6, the immersive language audio feature extraction model based on artificial intelligence proposed in this invention is developed using the Python programming language and implemented using the PyTorch framework. The model input is audio feature data in the form of a time-frequency tensor with a size of 1×128×128. During training, the RMSprop optimizer is used, with a learning rate of 0.0005, a batch size of 32, and 200 training epochs. The loss function is optimized using the cross-entropy function. By jointly training the spectral stress modulation tensor and the structural perturbation perception feature response map, the channel aggregation weights of the multi-channel deep features are gradually optimized during training. When the number of training epochs reaches about 120, the loss function converges to a stable range, indicating that the model has a strong frequency-phase structure co-expression ability and can accurately extract multi-dimensional emotion recognition labels from traditional cultural language audio.
[0058] Furthermore, in step S6, the traditional cultural language audio dataset is input into the constructed AI-based immersive language audio feature extraction model for processing. The experimental results are as follows: Figure 6 and Figure 7 As shown; Figure 6 This is a confusion matrix for emotion classification based on AI-driven immersive traditional cultural language audio feature extraction. The horizontal axis represents the model's predicted label, and the vertical axis represents the actual emotion label. Grayscale blocks represent the distribution of the number of recognized samples for each category. Figure 6 It can be seen that the recognition frequency of "happy", "solemn", and "exhilarated" emotions is significantly higher in the diagonal area than in other areas, indicating that the model has a strong ability to distinguish these typical categories. At the same time, there is some confusion between the labels "calm" and "sorrowful" and "solemn" and "dignified", which is consistent with the characteristics of natural transition of intonation boundaries and interweaving of subjective expressions in traditional cultural language. Figure 7 This image shows the label prediction accuracy results for AI-based immersive traditional cultural language audio feature extraction. It illustrates the trend of the total number of samples for each label versus the number of correctly predicted labels. Figure 7 As can be seen, the model accurately predicted a number of samples close to the total number of samples in the categories of "calm" and "happy", and the overall sample distribution was relatively balanced with a stable recognition trend, further verifying the effectiveness and robustness of the present invention in the task of extracting audio features of traditional cultural languages.
[0059] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A method for extracting audio features of immersive traditional cultural languages based on artificial intelligence, characterized in that, Specifically, the following steps are included: S1. Obtain audio materials of traditional cultural languages and corresponding emotion tags, preprocess the audio materials, extract logarithmic Mel spectrograms and frequency-phase spectrograms, and construct a structured audio dataset of traditional cultural languages; S2. Calculate the phase difference between adjacent frames based on the frequency phase spectrum to obtain the frequency phase residual information; use a nonlinear reconstruction function combined with the frequency band amplitude variability to modulate the frequency phase residual information, and convert it into a phase residual matrix with the same dimension as the logarithmic Mel spectrum through a structural coupling mapping mechanism. S3. Define a sliding window region on the phase residual matrix, extract local variation features, construct a frequency phase residual traction factor matrix, and normalize the matrix to obtain a normalized traction factor matrix. S4. Calculate the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term based on the logarithmic Mel spectrum, and construct the spectral stress excitation tensor. The spectral stress excitation tensor is modulated and enhanced based on the normalized traction factor matrix, and a multi-channel spectral coupling tensor is constructed by combining the channel mapping function. S5. Couple the multi-channel spectrum tensor into the feature extraction module, perform nonlinear transformation and dynamic response convolution to obtain a multi-dimensional deep feature map; based on the perturbation guidance mechanism, perform weighted aggregation on the multi-dimensional deep feature map to obtain the emotion tags corresponding to the audio materials of traditional cultural language. S6. Construct an AI-based model for extracting features from traditional cultural language audio. Input the traditional cultural language audio dataset, use cross-entropy as the loss function, and execute steps S2 to S5 sequentially. Use emotion labels for supervised training until the training process converges, thus completing the feature extraction of traditional cultural language audio.
2. The method for extracting audio features of immersive traditional cultural languages based on artificial intelligence according to claim 1, characterized in that, The specific method for constructing a structured audio dataset of traditional cultural languages is as follows: Acquire audio materials of traditional cultural language, including recitations, chanting, and storytelling with cultural characteristics; Collect the emotion tag information corresponding to the audio material, and the emotion tags include solemn, solemn, mournful and passionate emotions; The audio material is standardized and sampled to obtain the original speech signal data; the speech signal data is then subjected to frame segmentation, window function processing, and short-time Fourier transform to extract spectral amplitude information and frequency phase information. Based on the spectral amplitude information and the predefined Mel filter bank processing method, a logarithmic Mel spectrum is generated; the phase angle component in the short-time Fourier transform is extracted to obtain the frequency phase spectrum. The logarithmic Mel spectrogram and frequency-phase spectrogram are respectively bound to the emotion label information to form structured traditional cultural language audio samples; all structured sample data are summarized and arranged and indexed in a unified format to construct a traditional cultural language audio dataset for speech emotion recognition tasks.
3. The method for extracting audio features of immersive traditional cultural languages based on artificial intelligence according to claim 2, characterized in that, The specific method for establishing the phase residual matrix is as follows: S21. Calculate the phase difference between adjacent frames based on the frequency phase spectrum to obtain the inter-frame frequency phase residual information; S22. Based on the frequency phase residual information, a nonlinear reconstruction function combined with the frequency band amplitude variability is used to modulate the frequency phase residual information. The mathematical model of the reconstruction function is as follows: ; in, Indicates the index of the time frame. Indicates the frequency channel index. The average phase value within the frequency band. The standard deviation of the residuals in the frequency direction. For the first Frame, First Phase residual information at the frequency, This refers to the frequency phase residual information after modulation. S23. The frequency-phase residual information reconstructed nonlinearly is mapped to the frequency dimension based on a structural coupling mapping mechanism. A weighted coupling model is used to construct a phase residual matrix with the same dimension as the logarithmic Mel spectrum. The mathematical model is as follows: ; in, This represents the Mel spectrum dimension channel number corresponding to the structure mapping. For frequency aisle Mapping weights between them This is the phase residual matrix.
4. The method for extracting audio features of immersive traditional cultural languages based on artificial intelligence according to claim 3, characterized in that, The specific method for constructing the frequency-phase residual traction factor matrix is as follows: S31. Define a sliding window local region on the phase residual matrix, and calculate the variation coefficient for the frequency phase residual signal within each local region. The mathematical model is as follows: ; in, Position of the sliding window The phase residual block inside, For variance, The mean, A small, positive constant. S32. Based on the coefficient of variation and the frequency sensitivity rate adjustment mechanism, construct the frequency phase residual traction factor matrix. The mathematical model for the frequency-phase residual traction factor is: ; in, For the response sensitivity coefficient, The average intensity of the frequency gradient, This is a frequency disturbance suppression term; S33, Regarding the traction factor matrix After normalization, the normalized frequency-phase residual traction factor matrix is obtained. .
5. The method for extracting audio features of immersive traditional cultural languages based on artificial intelligence according to claim 4, characterized in that, The specific method for constructing the spectral coupling tensor is as follows: S41. Based on the Mel spectrum diagram, calculate the local spectral principal stress term, time-directed variation term, and spectral interaction coupling term to construct the initial spectral stress excitation tensor. The mathematical model of the spectral stress excitation tensor is as follows: ; in, For the logarithmic Mel spectrum, the first... The value at each position, For the principal stress terms of the local spectrum, For time-oriented variables, For spectral interaction coupling terms, For time-oriented local derivatives, For element-wise multiplication, , , This is the adjustment coefficient for the spectral stress tensor. For the spectral stress excitation tensor; S42. Based on the frequency-phase residual traction factor matrix, the spectral stress excitation tensor is modulated and enhanced to obtain the modulated spectral response tensor. The mathematical model is as follows: ; in, , For intensity adjustment parameters, This represents the energy value of the phase residual block. The modulated spectral response tensor; S43. Input the modulated spectral response tensor into the channel mapping function, project it from the frequency dimension to the structural domain dimension, and construct a multi-channel spectral coupling tensor.
6. The method for extracting audio features of immersive traditional cultural languages based on artificial intelligence according to claim 5, characterized in that, The specific method for weighted aggregation of multidimensional deep feature maps based on the perturbation guidance mechanism is as follows: S51. Input the multi-channel spectrum coupling tensor into the feature extraction module, perform a nonlinear transformation on it, and generate the first intermediate response tensor. S52. Perform dynamic response convolution operation on the first intermediate response tensor to obtain the structural perturbation sensing feature map; S53. Stack the structural perturbation-aware feature maps according to the channel dimension to construct a multi-dimensional deep feature map, and perform perturbation-guided weighted aggregation operation on the multi-dimensional deep feature map to obtain the channel aggregation output vector. The mathematical model is as follows: ; in, for The characteristic response compression function, where c is the graph channel index. For the c-th channel, the first... Location-based structural perturbation sensing features; The disturbance sensing factor for channel c; H represents the local structural fuzzy response adjustment term; H and W represent the height and width of the feature map. S54. Perform a normalized exponential mapping operation on each element of the channel aggregation output vector to form a normalized output probability vector. and to Take the maximum index to obtain the emotion tag corresponding to the audio material of traditional cultural language. .
7. The method for extracting audio features of immersive traditional cultural language based on artificial intelligence according to claim 6, characterized in that, The method for constructing a local structural fuzzy response modifier is as follows: S531. Based on the multi-dimensional deep feature map, perform directional gradient magnitude extraction operation within its local spatial neighborhood to obtain the spatial location. Disturbance structure strength value ; S532, The strength value of the disturbed structure The input is fed into the response smoothing adjustment function to obtain the local structural fuzzy response adjustment term at the corresponding location. The mathematical model of the response regulation function is as follows: ; in, Indicates position The strength value of the disturbed structure; This is the slope factor of the adjustment function. The threshold for structural fuzziness discrimination.
Citation Information
Patent Citations
Helping equipment for vulnerable groups to go out
CN109875832A
Automatic detection method of digital audio tampering based on grid frequency fluctuation super vector
CN108766464A
Speech emotion recognition method based on spectral features and ELM
CN110827857A
Patient voice anger emotion recognition method and system
CN112002348A
Speech emotion recognition method based on speech spectrum
CN112581979A