Method and device for carrying out sound effect control on played audio

Through self-attention audio analysis network, the music structure and emotional characteristics are identified, and combined with the breathing points of the music phrase to generate and optimize the sound effect parameters, the problem of the existing sound effect control methods being disconnected from the music content is solved, and more natural sound effect adjustment and emotional transmission effects are achieved.

CN120279952AInactive Publication Date: 2025-07-08SHENZHEN ZUNTE DIGITAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510764794.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing audio sound control methods cannot identify the key points of the music structure and the breathing points of the music phrase, resulting in the disconnection of the sound effect processing from the music content, and the inability to dynamically adjust the sound effect parameters according to the changes in the music emotions, affecting the auditory experience.

Method used

Through the audio analysis network of the self-attention algorithm, the key points of the audio structure and breathing points of the phrase are identified, the emotional characteristics of the audio are extracted and mapped into emotionally driven sound effects parameters, combined with the sound effects parameters of the phrase structure, the optimized sound effects parameters are generated using a multi-objective optimization algorithm, and the audio signal characteristics are adjusted in real time.

Benefits of technology

It enhances the expressiveness and emotional expression of music, provides an auditory experience close to live performance, avoids mechanical and stiffness, is suitable for a variety of musical styles and performs well in real-time and computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279952A_ABST
    Figure CN120279952A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio playing, and discloses a sound effect control method and device for played audios, and the sound effect control method comprises the following steps: obtaining to-be-played audios, carrying out the preprocessing of the audios to obtain standardized audio signals, identifying structure key points and phrase breathing points in the audio by using an audio analysis network, and generating music structure marks and phrase breathing point marks; extracting emotional features of the audio based on the music structure marks, and converting the emotional features into emotion-driven sound effect parameters through an emotional sound effect parameter mapping network; according to the invention, by identifying and strengthening the internal structure of the music and the breathing characteristics of the phrases, the recorded music playback is more vitality, the emotion expression is richer, the auditory experience is close to the on-site playing effect, the expressive force of the music recording is improved, and meanwhile, the music playing has a more natural flow sense due to the identification and strengthening of the breathing characteristics of the phrases; and the mechanical feeling and the stiff feeling caused by a traditional sound effect processing method are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio playback, and more specifically, to a method and device for controlling the sound effects of played audio. Background Art

[0002] Audio sound effect control technology is an important research direction in the field of audio processing. Its purpose is to enhance the listening experience of audio playback by processing and adjusting audio signals. With the development of digital audio technology, sound effect control has evolved from basic functions such as simple volume adjustment and pitch adjustment to a complex processing system that includes equalizers, dynamic processing, reverb effects, stereo processing, and more.

[0003] The existing audio sound effect control methods mainly have the following technical problems: Traditional sound effect processing methods ignore the structural features and semantic information in audio content. Specifically, traditional methods cannot identify the structural key points (such as peaks, transitions, introductions, etc.) and phrase breathing points (such as phrase starts, peaks, and ends) in music works, resulting in a disconnect between sound effect processing and music content and an inability to provide corresponding sound effect enhancements for different structural parts. Moreover, the existing sound effect control systems lack the ability to understand the emotional expression needs of music. As a carrier of emotional expression, different paragraphs and parts of music often carry different emotional colors and expression intentions. However, the existing technologies cannot identify and analyze these emotional features, so they cannot dynamically adjust sound effect parameters according to changes in music emotions, resulting in poor emotional transmission effects and an inability to effectively enhance the emotional resonance experience of listeners. Summary of the Invention

[0004] The present invention provides a method and device for controlling the sound effects of played audio, which solves the technical problems of audio playback sound effect control in related technologies.

[0005] The present invention provides a method for controlling the sound effects of played audio, including the following steps: Obtain the audio to be played, preprocess the audio to obtain a standardized audio signal, and use an audio analysis network to identify the structural key points and phrase breathing points in the audio, generating music structure marks and phrase breathing point marks; Based on the music structure marks, extract the emotional features of the audio, and convert the emotional features into emotion-driven sound effect parameters through an emotion sound effect parameter mapping network; Based on the phrase breathing point marks, generate phrase structure sound effect parameters according to the phrase breathing point marks; Combine the emotion-driven sound effect parameters and the phrase structure sound effect parameters, and generate optimized sound effect parameters through a multi-objective optimization algorithm; Apply the optimized audio effect parameters to the audio playback process, adjust the audio signal characteristics in real time, and output the processed audio signal.

[0006] As a further optimization scheme of the present invention, the audio analysis network is a network based on the self-attention algorithm, including: An input layer for receiving an audio feature vector sequence; A self-attention layer for calculating the attention weights between each time point in the feature sequence; A multi-head attention layer for enhancing the model's expressive ability; A bidirectional long short-term memory network layer for capturing temporal information; An output layer for outputting the probability distributions of the structure tag and the breathing point tag.

[0007] As a further optimization scheme of the present invention, the steps of extracting the emotional features of the audio include: Extract the rhythm features, harmony features, timbre features, dynamic features, and structural association features of the audio; Use an emotion classification model to map the extracted features into a two-dimensional emotion space to form an emotion state sequence, and the two-dimensional emotion space includes a valence dimension and an arousal dimension.

[0008] As a further optimization scheme of the present invention, the emotion audio effect parameter mapping network includes: An input layer for receiving an emotion state vector; Multiple hidden layers for performing non-linear transformations; An output layer for generating an audio effect parameter vector including equalizer parameters, dynamic range parameters, reverb parameters, and stereo parameters.

[0009] As a further optimization scheme of the present invention, the steps of generating the musical phrase structure audio effect parameters according to the musical phrase breathing point tag include: Construct a musical phrase breathing model, and abstract the musical phrase breathing process into three stages: start, peak, and end; For the start stage of the musical phrase, apply progressive clarity enhancement and directional focusing parameters; For the peak stage of the musical phrase, apply sound field expansion and dynamic range enhancement parameters; For the end stage of the musical phrase, apply natural decay and acoustic space blanking parameters; Process the transition between adjacent musical phrases to ensure the continuity and naturalness of the audio effect changes.

[0010] As a further optimization scheme of the present invention, the multi-objective optimization algorithm includes the following steps: Detect potential conflicts between the emotion-driven audio effect parameters and the musical phrase structure audio effect parameters; Build a multi-objective optimization model, including an emotional parameter distance function, a phrase parameter distance function, and a smoothness penalty function; Dynamically adjust the weights of the optimization objectives according to the characteristics and conflict situations of the music content; Adopt an iterative optimization algorithm to solve the multi-objective optimization problem and obtain the final sound effect parameter sequence.

[0011] As a further optimization scheme of the present invention, the step of applying the optimized sound effect parameters to the audio playback process includes: Build a processing link including an equalizer unit, a dynamic processing unit, a reverb unit, a stereo processing unit, and a harmonic processing unit; For each audio sampling point, calculate the corresponding parameter interpolation according to its time position; Apply the interpolated sound effect parameters to the sound effect processing link to realize real-time processing of the audio signal; Adaptive adjust the processing strategy according to the system resource status and real-time requirements.

[0012] As a further optimization scheme of the present invention, the adaptive adjustment of the processing strategy includes: Use a sliding window to calculate the processing delay, and trigger performance optimization when the average delay exceeds the threshold; According to the delay situation, divide the processing flow into different precision levels, and reduce the processing precision in the case of high delay; For parameters that change slowly, adopt a frame interval update strategy to reduce the calculation frequency; For computationally intensive operations, use an acceleration algorithm and cache intermediate results to reduce the computational complexity.

[0013] A device for sound effect control of the played audio, used to implement the method for sound effect control of the played audio described above, includes: An audio analysis module, used to identify the structural key points and phrase breathing points in the audio, and generate music structure marks and phrase breathing point marks; An emotional feature extraction module, used to extract the emotional features of the audio; An emotional sound effect parameter mapping module, used to convert emotional features into emotion-driven sound effect parameters; A phrase structure sound effect generation module, used to generate phrase structure sound effect parameters according to the phrase breathing point marks; A parameter collaborative optimization module, used to combine emotion-driven sound effect parameters and phrase structure sound effect parameters, and generate optimized sound effect parameters through a multi-objective optimization algorithm; A sound effect processing module, used to apply the optimized sound effect parameters to the audio playback process, adjust the audio signal characteristics in real time, and output the processed audio signal.

[0014] The beneficial effects of the present invention are as follows: By identifying and enhancing the internal structure of music and the breathing characteristics of musical phrases, the present invention makes the playback of recorded music more vivid, with richer emotional expression, and the auditory experience approaching the live performance effect, improving the expressiveness of music recording. At the same time, the identification and enhancement of the breathing characteristics of musical phrases make the music playback have a more natural sense of flow, avoiding the mechanical and rigid feeling caused by traditional sound effect processing methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a flowchart of the method for controlling the sound effect of the played audio according to the present invention; Figure 2 is a detailed flowchart of generating music structure marks and musical phrase breathing point marks according to the present invention; Figure 3 is a detailed flowchart of converting emotion-driven sound effect parameters according to the present invention; Figure 4 is a detailed flowchart of generating dynamic sound effects for musical phrase structures according to the present invention; Figure 5 is a detailed flowchart of collaborative optimization of multi-dimensional sound effect parameters according to the present invention; Figure 6 is a detailed flowchart of real-time sound effect processing and application according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0017] In at least one embodiment of the present invention, a method for controlling the sound effect of the played audio is disclosed, as Figures 1 to 6 shown, including the following steps: Step 1: Obtain the audio to be played, preprocess the audio to obtain a standardized audio signal, and use an audio analysis network to identify the structural key points and musical phrase breathing points in the audio, generating music structure marks and musical phrase breathing point marks; In this step, the input audio is processed by an audio analysis network based on the self-attention algorithm to extract frequency domain and time domain features, identify the chapter boundaries, structural key points, and musical phrase breathing points in the music, and generate music structure marks and musical phrase breathing point marks. Specifically, it includes the following steps: Step 1.1: Obtain a standardized audio signal; Obtain the audio file to be played and preprocess the audio to obtain a standardized audio signal, specifically including: Audio file reading: Support reading of multiple common audio formats (such as WAV, MP3, FLAC, etc.), and load the audio data into memory.

[0018] Audio signal standardization: Resample the audio to unify the sampling rate (such as 44.1kHz or 48kHz); convert multi-channel audio to mono; perform volume normalization to unify the amplitude range of the audio signal; remove the DC component; apply a pre-emphasis filter to enhance the high-frequency components.

[0019] Noise processing: Use methods such as median filtering to remove impulse noise; apply a band-pass filter to filter out irrelevant frequency bands; perform dynamic range compression to reduce signal fluctuations.

[0020] Signal segmentation: Segment the audio into fixed-length segments (such as 10 seconds); retain a certain overlap (such as 1 second) between adjacent segments; apply a window function (such as Hamming window) to each segment of the signal.

[0021] After the above preprocessing, a standardized audio signal is obtained, providing a basis for subsequent feature extraction and analysis.

[0022] Step 1.2, Audio feature extraction; Preprocess and extract features from the input audio signal, which may specifically include: Process the audio signal in frames, with each frame having a length of 20 milliseconds and a frame shift of 10 milliseconds; Perform a short-time Fourier transform (STFT) on each frame of the signal to obtain a time-frequency spectrogram representation; Extract frequency-domain features, including Mel-frequency cepstral coefficients (MFCC), spectral centroid, spectral flux, zero-crossing rate, etc.; Extract time-domain features, including energy envelope, amplitude envelope, transient features, etc.; Combine the extracted various features to form a feature vector sequence as the input for subsequent analysis.

[0023] The feature extraction process can be expressed as: ; Among them, represents the feature sequence; , , respectively represent the feature vectors of the , , frame; represents the total number of frames of the audio.

[0024] Step 1.3, Construction and training of the self-attention audio analysis network; According to an embodiment of the present application, an audio analysis network based on the self-attention algorithm is constructed. This network can capture long-range dependencies in the audio feature sequence and identify music structures and phrase breathing points. The network structure mainly includes: Input layer: Receives the feature vector sequence ; Self-attention layer: Calculates the attention weights between each time point in the feature sequence to capture long-range dependencies. The self-attention calculation is as follows: ; Among them, represents the attention function; represents the query matrix; represents the key matrix; represents the value matrix; represents the dimension of the key; represents the normalization function; represents the multiplication of the query matrix and the transpose of the key matrix; Feed-forward neural network layer: Performs a non-linear transformation on the output of the self-attention layer; Multi-head attention algorithm: Uses multiple groups of different linear projections to enhance the expressive power of the model; Residual connection and layer normalization: Improves the stability of model training; Bidirectional long short-term memory network (BiLSTM) layer: Further captures temporal information; Output layer: Outputs the probability distributions of the structure label and the breathing point label through a fully connected layer and an activation function.

[0025] This network is trained using a labeled music dataset, and the network parameters are optimized through the cross-entropy loss function and the backpropagation algorithm.

[0026] In addition, it should be understood that the specific structure of the above audio analysis network can be adjusted according to actual needs. For example, the number of network layers can be increased or decreased, the type of activation function can be changed, etc., as long as it can effectively identify music structures and phrase breathing points.

[0027] In a specific embodiment of the present application, the self-attention audio analysis network can be implemented using a variant of the Transformer architecture. Specifically, this network includes an input embedding layer that maps 256-dimensional audio feature vectors to a 512-dimensional embedding space; 4 layers of Transformer encoders, each layer containing 8 heads of self-attention mechanisms (each head with a dimension of 64) and a 1024-dimensional feed-forward neural network; finally, a bidirectional LSTM layer (with 256 hidden units) and two parallel fully connected layers that respectively output the probability distributions of the structure type and the breathing point type.

[0028] In this embodiment, the calculation of multi-head attention can be expressed as: ; ; It should be noted that , , ; are the parameter matrices of different heads, is the output parameter matrix; Among them, represents the multi-head attention function; represents the query matrix; represents the key matrix; represents the value matrix; represents the concatenation operation; represents the output of attention heads; represents the th attention head output; represents the attention calculation function; represents the th head's query weight matrix, with dimension ; represents the th head's key weight matrix, with dimension ; represents the th head's value weight matrix, with dimension ; represents the output weight matrix, with dimension ; represents the set of real numbers.

[0029] In the application example of the classical music scenario, the network can accurately identify the movement structure and phrase breathing points, providing a reliable structural information basis for subsequent sound effect enhancement. For example, when processing Beethoven's piano sonatas, the network can identify the large structures such as the exposition, development, and recapitulation of the sonata form, and at the same time identify the starting, peak, and ending points of the phrases at a finer granularity.

[0030] In some alternative embodiments, the self-attention audio analysis network may adopt a hybrid architecture that combines a convolutional neural network and a self-attention mechanism. For example, a 1D convolutional layer can be used first to extract local time-domain features, and then the extracted features are input into the self-attention layer to capture long-range dependencies. This hybrid architecture performs better when processing popular music with obvious rhythm features because the convolutional layer can effectively capture rhythm patterns and repetitive structures.

[0031] Optionally, to improve the adaptability of the network to different music styles, a domain adaptation training method can be adopted. That is, first pre-train the network on a large-scale general music dataset, and then fine-tune it on a music dataset of a specific style. This method can enable the network to better adapt to the structural and breathing characteristics of different music styles.

[0032] In addition, in some embodiments, music theory knowledge constraints can be introduced into the network. For example, the rules regarding phrase structure in music theory are added as regularization terms to the loss function to guide the network to learn structural representations that conform to music theory. This knowledge-guided learning method can improve the generalization ability of the network on music styles lacking labeled data.

[0033] Step 1.4, Music structure and phrase breathing point recognition; Use the trained audio analysis network to perform structure and breathing point recognition on the input audio: Input the preprocessed feature sequence into the audio analysis network; The network outputs the probability distributions of different structural types (such as peaks, transitions, introductions, etc.) and breathing point types (such as starts, peaks, ends) at each time point; Determine the final structure label and breathing point label sequences through probability thresholds and smoothing processing.

[0034] The structure label sequence can be expressed as: ; where represents the structure label sequence of the audio; , , respectively represent the structural types of the , , th frames; represents the total number of frames of the audio; Specifically, represents the structural type of the th frame.

[0035] The breathing point label sequence can be expressed as: ; where represents the breathing point label sequence of the audio; , , respectively represent the breathing point types of the , , th frames; represents the total number of frames of the audio; Specifically, Indicates the breathing point type of the frame.

[0036] In addition, it should be understood that the specific structure of the above audio analysis network can be adjusted according to actual needs. For example, the number of network layers can be increased or decreased, the type of activation function can be changed, etc., as long as it can effectively identify the music structure and the breathing points of musical phrases.

[0037] Step 2: Based on the music structure markers, extract the emotional features of the audio, and convert the emotional features into emotion-driven sound effect parameters through an emotion sound effect parameter mapping network; In this step, analyze the correlations between features such as music structure, rhythm, and harmony with emotions, establish a music emotion feature library, and use the emotion sound effect mapping network to convert the recognized emotional states into corresponding sets of sound effect parameters, forming the basis of emotion-driven sound effect parameters.

[0038] Step 2.1: Music emotion feature extraction; According to an embodiment of the present application, based on the original audio features and the structure information recognized in Step 1, extract high-level features related to emotional expression: Rhythm features: including rhythm density, rhythm regularity, beat strength, etc., which are extracted from the audio signal through a beat tracking algorithm; Harmony features: extract features such as chord progressions, key information, and harmony complexity in the audio through a harmony analysis algorithm; Timbre features: extract features such as timbre brightness, roughness, and spectral entropy through spectral analysis; Dynamic features: extract features such as the volume change range, volume change speed, and volume contrast; Structure correlation features: based on the structure markers recognized in Step 1, extract features such as the structure change rate and structure contrast.

[0039] The emotion feature vector can be expressed as: ; where represents the emotion feature sequence; , , respectively represent the emotion feature vectors of the , , frame; represents the total number of frames of the audio; Step 2.2: Emotion state recognition; According to an embodiment of the present application, use an emotion classification model to map the extracted emotion features into a predefined emotion space to recognize the emotion states of each paragraph of the audio: Construct a two-dimensional emotion space, where the horizontal axis represents valence and the vertical axis represents arousal; Use the trained emotion classification model to map the emotion feature vectors to the coordinates in the two-dimensional emotion space; Smooth the emotion trajectory to avoid frequent fluctuations in the emotional state; Identify the emotion change points and emotion stable intervals to provide a basis for subsequent sound effect parameter mapping.

[0040] The emotion state sequence can be expressed as: ; Among them, represents the emotion state sequence; , , respectively represent the emotion states of the , , th frames; represents the total number of frames of the audio; Specifically, represents the emotion state of the th frame, where, and respectively represent the valence and arousal values.

[0041] It should be noted that the above two-dimensional emotion space is only one implementation manner, and the technical solution of this application can also adopt other-dimensional emotion spaces, such as a three-dimensional emotion space (adding a dominance dimension) or a higher-dimensional emotion representation.

[0042] Step 2.3, Emotional sound effect parameter mapping; According to the embodiments of this application, construct an emotional sound effect parameter mapping network to convert the identified emotion state into corresponding sound effect parameters: Establish a mapping relationship database between emotion states and sound effect parameters, including the optimal sound effect parameter combinations under typical emotion states; Construct a mapping network structure, including a multi-layer perceptron and an attention algorithm, to achieve a non-linear mapping from emotion states to sound effect parameters; Use the labeled data set to train the mapping network to optimize the mapping relationship from emotion states to sound effect parameters; For the input emotion state sequence, generate the corresponding emotion-driven sound effect parameter sequence through the trained mapping network.

[0043] The emotion-driven sound effect parameters can be expressed as: ; Among them, represents the emotion-driven sound effect parameter sequence; , , respectively represent the sound effect parameter vectors of the , , th frames; represents the total number of frames of the audio; Specifically, represents the sound effect parameter vector mapped from the emotional state , including equalizer parameters, dynamic range parameters, reverb parameters, etc.

[0044] The emotional sound effect parameter mapping function can be expressed as: ; Among them, represents the sound effect parameter vector of the th frame; represents the functional relationship of the mapping network; represents the emotional state of the th frame; represents the valence value of the th frame; represents the arousal value of the th frame.

[0045] In addition, it should be understood that the specific implementation manner of the above-mentioned emotional sound effect parameter mapping network can be adjusted according to actual application requirements. For example, different network structures or training methods can be adopted as long as an effective mapping from the emotional state to the sound effect parameters can be achieved.

[0046] In a specific embodiment of the present application, the emotional sound effect parameter mapping network is implemented by a structure combining a multi-layer perceptron and a residual connection. Specifically, the network includes: an input layer that receives a 2D emotional state vector (valence and arousal); three hidden layers (with dimensions of 64, 128, and 64 respectively), each followed by batch normalization and a ReLU activation function; an output layer that generates a vector containing 24 sound effect parameters (including 10 equalizer band parameters, 4 dynamic range parameters, 6 reverb parameters, and 4 stereo parameters). In addition, a residual connection is added between the first hidden layer and the last hidden layer to improve training stability and model expression ability.

[0047] In this embodiment, for the emotional state (where represents valence, represents arousal, both normalized to the [-1, 1] interval), the mapping function can be expressed as: ; ; ; ; Specifically, , , , , ; Among them, , , respectively represent the output vectors of the first hidden layer, the second hidden layer, and the third hidden layer; represents the rectified linear unit activation function; represents the batch normalization operation; represents the weight matrix from the input layer to the first hidden layer; represents the weight matrix from the first hidden layer to the second hidden layer; represents the weight matrix from the second hidden layer to the third hidden layer; represents the weight matrix from the third hidden layer to the output layer; represents the weight matrix of the residual connection; , , , respectively represent the bias vectors of the first hidden layer, the second hidden layer, the third hidden layer, and the output layer; represents the valence value; represents the arousal value; represents the output sound effect parameter vector, that is, the output layer; represents the set of real numbers.

[0048] In the application example in the pop music scene, when the emotional state is high valence and high arousal (such as a lively and exciting mood), the sound effect parameters generated by this network will enhance the gain of high and middle frequencies, expand the dynamic range, reduce the reverberation time, and increase the stereo width, thereby strengthening the brightness and vitality of the music; when the emotional state is low valence and low arousal (such as a sad and calm mood), the generated sound effect parameters will enhance the warmth of the low frequencies, compress the dynamic range, increase the reverberation time, and reduce the stereo width, thereby creating a more introverted and immersive sound effect experience. In the subjective listening evaluation, the score of the enhanced emotional expression effect of the music processed with the sound effect parameters generated by this mapping network is higher than that of the traditional fixed parameter sound effect processing.

[0049] In some alternative embodiments, the emotional sound effect parameter mapping network may adopt a memory-based neural network architecture, such as a Memory-Augmented Neural Network (MANN). This architecture introduces an external memory module to store the correspondence between typical emotional states and sound effect parameters. When the network performs mapping, it queries the memory module based on the input emotional state, extracts relevant sound effect parameter templates, and then makes adaptive adjustments. This memory-based method is particularly suitable for processing sparse emotional space regions and can better handle emotional states that are not fully covered in the training data.

[0050] Optionally, in order to handle the differences in emotional expressions of different music styles, a Conditional Variational Autoencoder (CVAE) structure can be used to implement the emotional sound effect parameter mapping. In this implementation, in addition to the emotional state, the music style is also used as a conditional input, enabling the network to generate adapted sound effect parameters according to the characteristics of different styles. For example, classical music and jazz expressing the same sad emotion may require different parameter settings in sound effect processing.

[0051] In addition, in some embodiments, a reinforcement learning method can be used to optimize the emotional sound effect parameter mapping. By defining a reward function for enhancing the emotional effect and using algorithms such as policy gradient or deep Q-network, the mapping strategy is iteratively optimized. The advantage of this method is that it can directly optimize the final listening effect, rather than just fitting the parameters annotated by experts. In experiments, this reinforcement learning-based method performs better when dealing with music works with complex emotional changes.

[0052] Step 3: Based on the phrase breathing point markers, generate phrase structure sound effect parameters according to the phrase breathing point markers; According to an embodiment of the present application, in this step, based on the identified phrase structure and breathing points, a phrase structure sound effect processing algorithm is constructed to enhance clarity and directivity at the beginning of the phrase, expand the sound field and dynamic range at the peak, and naturally decay and leave an acoustic space at the end, generating a set of breathing-style dynamic sound effect parameters.

[0053] Step 3.1: Construction of the phrase breathing model; Based on music performance theory and the breathing patterns of performers, a mathematical model is constructed to describe the phrase breathing characteristics: The phrase breathing process is abstracted into three stages: inhalation (start), sustain (peak), and exhalation (end); Characteristic functions are defined for each stage to describe the variation law of sound effect parameters over time; Construct a breathing curve function to control the smooth change of sound effect parameters.

[0054] The breathing curve function of a musical phrase can be expressed as: ; Where represents the breathing curve function of a musical phrase; represents the time variable; , , and represent the times of the start point, peak point, end point of the musical phrase and the start point of the next musical phrase respectively; , and represent the characteristic functions of the start stage, peak stage and end stage respectively.

[0055] It should be noted that the specific form of the above breathing curve function of the musical phrase can be adjusted according to different music types and performance requirements. For example, different mathematical functions can be used to describe the change rules of each stage.

[0056] In a specific embodiment of the present application, the breathing curve of the musical phrase can be implemented in the following function form: ; ; ; Where , and represent the characteristic functions of the start stage, peak stage and end stage respectively; represents the time variable; , and represent the amplitude parameters of the start stage, peak stage and end stage respectively; and represent the parameters controlling the curve shapes of the start stage and the end stage respectively; and represent the times of the start point and end point of the musical phrase respectively; and represent the characteristic functions of the peak stage and the end stage respectively; represents the base of the natural logarithm.

[0057] In the start stage, an exponential growth function is used to simulate the gradual increase of music intensity; in the peak stage, a constant function is used to maintain a stable intensity; in the end stage, an exponential decay function is used to simulate the natural weakening of music intensity.

[0058] In some alternative embodiments, the phrase breathing curve can adopt a parametric representation based on Bezier curves. Bezier curves can provide more flexible shape control and are particularly suitable for simulating the phrase breathing characteristics of different music styles. For example, for long phrases in classical music, a third-order Bezier curve can be used to achieve a smoother start and end transition; while for short phrases in jazz, a second-order Bezier curve can be used to achieve a more direct intensity change.

[0059] Optionally, the phrase breathing model can be adaptively adjusted according to the beat and tempo information of the music. For example, in fast movements, the time parameters of the start and end phases can be correspondingly shortened, while in slow movements, these parameters can be extended to match the natural flow of the music. This adaptive phrase breathing model can be expressed as: ; ; where, parameters representing the curve shape of the start phase; representing the reference parameter value of the start phase; representing the tempo of the music (expressed in beats per minute); parameters representing the curve shape of the end phase; representing the reference parameter value of the end phase.

[0060] In addition, in some embodiments, different breathing model templates can be defined for different types of phrases. For example, interrogative phrases can use a curve with an upward ending, while conclusive phrases can use a curve with a downward ending. These templates can be automatically selected according to the phrase types identified in step 1, so as to achieve a breathing effect that better conforms to musical grammar. Experiments show that this adaptive breathing model based on phrase types can make the sound effect processing better fit the grammatical structure of the music, improving the understanding and acceptance of the listeners.

[0061] Step 3.2, generating phrase structure sound effect parameters; According to the embodiments of the present application, based on the phrase breathing model and the phrase structure and breathing points identified in step 1, sound effect parameters adapted to the phrase structure are generated: For the start phase of the phrase, progressive clarity enhancement and directional focusing are applied to enhance the clarity and directivity of the starting note: ; where, representing the sound effect parameters of the start phase of the phrase; representing the basic sound effect parameters; representing the weight coefficient of the start phase; representing the characteristic function of the start phase; represents a time variable; represents the parameter adjustment amount in the starting stage.

[0062] For the peak stage of a musical phrase, apply sound field expansion and dynamic range enhancement to highlight the expressiveness of the peak part: ; Among them, represents the sound effect parameters in the peak stage of a musical phrase; represents the basic sound effect parameters; represents the weight coefficient in the peak stage; represents the characteristic function in the peak stage; represents a time variable; represents the parameter adjustment amount in the peak stage.

[0063] For the ending stage of a musical phrase, apply natural attenuation and acoustic space blanking to create an auditory effect of natural ending: ; Among them, represents the sound effect parameters in the ending stage of a musical phrase; represents the basic sound effect parameters; represents the weight coefficient in the ending stage; represents the characteristic function in the ending stage; represents a time variable; represents the parameter adjustment amount in the ending stage.

[0064] Integrate to form a sequence of sound effect parameters for the musical phrase structure: ; Among them, represents the sequence of sound effect parameters for the musical phrase breathing; 、 、 respectively represent the 、 、 th sound effect parameters in the sequence; represents the length of the parameter sequence.

[0065] Step 3.3, Multi-phrase transition processing; According to an embodiment of the present application, process the transition between adjacent musical phrases to ensure the continuity and naturalness of sound effect changes: Detect the connection point of adjacent musical phrases and analyze its musical coherence characteristics; Select a suitable transition strategy according to the connection characteristics, such as smooth transition, contrast transition or interruption transition; Apply a transition function to process the connection of sound effect parameters of adjacent musical phrases to avoid sudden changes or unnatural transitions.

[0066] The transition processing function can be expressed as: ; wherein, represents the sound effect parameter of the transition region; represents the transition weight function, satisfying and ; represents the sound effect parameter at the end stage of the previous musical phrase; represents the sound effect parameter at the starting stage of the next musical phrase; represents the time variable.

[0067] In addition, it should be understood that the specific algorithm for generating the sound effect parameters of the above musical phrase structure can be adjusted according to different music styles and performance requirements. For example, different parameter adjustment strategies can be designed for different types of music.

[0068] Step 4: Combine the emotion-driven sound effect parameters and the sound effect parameters of the musical phrase structure, and generate optimized sound effect parameters through a multi-objective optimization algorithm; According to an embodiment of the present application, in this step, the emotion-driven sound effect parameters and the sound effect parameters of the musical phrase structure are combined, and the final sound effect parameter combination is determined through a multi-objective optimization algorithm to achieve the collaborative enhancement of emotion expression and musical phrase breathing, while ensuring the smooth transition of the music emotion change points.

[0069] Step 4.1: Sound effect parameter conflict detection; First, detect the potential conflict between the emotion-driven sound effect parameters and the sound effect parameters of the musical phrase structure: For each time frame, calculate the difference vector of the two sets of sound effect parameters: ; wherein, represents the parameter difference vector of the th frame, and respectively represent the emotion-driven and musical phrase structure sound effect parameters.

[0070] Set the difference threshold , and mark it as a conflict point when the difference exceeds the threshold: ; wherein, represents whether there is a conflict in the th frame, represents the norm of the difference vector.

[0071] Analyze the distribution characteristics of the conflict points to determine the conflict region and the conflict parameter dimension.

[0072] Step 4.2: Multi-objective optimization model construction; According to the embodiments of the present application, a multi-objective optimization model is constructed to find the optimal parameter combination that simultaneously meets the requirements of emotional expression and phrase breathing: Define the objective function:

[0073] Among them, represents minimizing the objective function; represents the sequence of sound effect parameters to be optimized; represents the weight coefficient of the emotion parameter; represents the distance function from the emotion parameter; represents the emotion-driven sound effect parameter; represents the weight coefficient of the phrase parameter; represents the distance function from the phrase parameter; represents the phrase structure sound effect parameter; represents the weight coefficient of the smoothness penalty; represents the smoothness penalty function.

[0074] Definition of the distance function: ; Among them, represents the distance function from the emotion parameter; represents the sequence of sound effect parameters to be optimized; represents the emotion-driven sound effect parameter; represents from the th frame to the th frame, the sum of the squares of the Euclidean distances between the parameter to be optimized and the emotion-driven parameter for each frame; represents the frame number; represents the total number of frames of the audio; represents the parameter to be optimized at the th frame; represents the emotion-driven parameter at the th frame; represents the square of the Euclidean distance between the parameter to be optimized and the emotion-driven parameter at the th frame.

[0075] ; Among them, represents the distance function from the phrase parameter; represents the sequence of sound effect parameters to be optimized; represents the phrase structure sound effect parameter; represents from the th frame to the th frame, the sum of the squares of the Euclidean distances between the parameter to be optimized and the phrase structure parameter for each frame; represents the frame number; Indicates the total number of audio frames; Indicates the parameter to be optimized for the Indicates the phrase structure parameter for the Indicates the square of the Euclidean distance between the parameter to be optimized for the

[0076] Definition of the smoothness penalty function: ; Wherein, Indicates the smoothness penalty function; Indicates the sequence of sound effect parameters to be optimized; Indicates from the frame to the sum of the squared norms of the differences between adjacent frame parameters; Indicates the parameter to be optimized for the Indicates the parameter to be optimized for the Indicates the frame number; Indicates the total number of audio frames; Indicates the squared norm of the difference between adjacent frame parameters.

[0077] Setting of the constraint conditions: ; Wherein, Indicates the minimum value of the parameter; Indicates the parameter vector for the Indicates the maximum value of the parameter; Indicates the frame number; Indicates the total number of audio frames.

[0078]

[0079] Wherein, Indicates the parameter vector for the i-th frame; Indicates the parameter vector for the Indicates the maximum allowable value of the change in adjacent frame parameters; Indicates the frame number; Indicates the total number of audio frames.

[0080] Specifically, and indicate the value range of the parameter, Indicates the maximum allowable value of the change in adjacent frame parameters.

[0081] It should be understood that the above multi-objective optimization model can be adjusted according to actual application requirements. For example, the terms in the objective function can be increased or decreased, or different weight allocation strategies can be adopted.

[0082] Step 4.3, Dynamic weight adjustment; According to an embodiment of the present application, based on the music content characteristics and conflict situations, dynamically adjust the weights of the optimization objectives: Near the emotional change point, increase the weight of the emotional parameter: ; Among them, represents the weight of the emotional parameter of the th frame, represents the basic weight, represents the adjustment coefficient, represents the norm of the emotional change gradient.

[0083] Near the phrase breathing point, increase the weight of the phrase structure parameter: ; Among them, represents the weight of the phrase structure parameter of the th frame, represents the basic weight, represents the adjustment coefficient, represents the th frame's indicating function of whether it is a breathing point; represents the frame number; B represents the set of breathing points.

[0084] In the conflict area, adjust the weight ratio according to the dominant factors of the music content.

[0085] In addition, it should be noted that the above dynamic weight adjustment strategy can be customized according to different music types and application scenarios. For example, for music types where emotional expression is more important, the basic weight of the emotional parameter can be increased.

[0086] Step 4.4, Collaborative optimization solution; According to the embodiment of the present application, use an iterative optimization algorithm to solve the multi-objective optimization problem and obtain the final sound effect parameter sequence: Initialize the parameter sequence, and the weighted average of the emotional parameter and the phrase parameter can be used as the initial value; Adopt gradient descent or other optimization algorithms to iteratively solve and update the parameter sequence; Check the convergence condition and stop when the change of the objective function is less than the preset threshold or the maximum number of iterations is reached; Output the optimized sound effect parameter sequence 。

[0087] The optimized sound effect parameter sequence can be expressed as: ; wherein, represents the optimized sound effect parameter sequence; , , respectively represent the optimized sound effect parameter vectors of the , , th frames; represents the total number of frames of the audio.

[0088] Therefore, through the above multi-dimensional sound effect parameter collaborative optimization process, this application realizes the collaborative enhancement of emotional expression and phrase breathing, avoids parameter conflicts in sound effect processing, and ensures smooth transitions at music emotional change points.

[0089] In a specific embodiment of this application, the multi-objective optimization problem adopts an iterative solution method based on the Adam optimizer. Specifically, the objective function is defined as: ; ; wherein, represents the objective function; represents the sound effect parameter sequence; represents the sum of the squares of the Euclidean distances between the parameter to be optimized and the emotion-driven parameter for each frame from the th frame to the th frame; represents the sum of the squares of the Euclidean distances between the parameter to be optimized and the phrase structure parameter for each frame from the th frame to the th frame; represents the sum of the squared norms of the differences between adjacent frame parameters from the th frame to the th frame; represents the parameter to be optimized for the th frame; represents the emotion-driven parameter for the th frame; represents the phrase structure parameter for the th frame; represents the parameter to be optimized for the th frame; represents the weight coefficient of the emotion parameter; represents the weight coefficient of the phrase structure parameter; represents the weight coefficient of the smoothness constraint; represents the frame number; represents the total number of frames of the audio.

[0090] Among them, the weight coefficient is dynamically adjusted according to the characteristics of the music content, and the basic weight is set to , , .

[0091] For each frame , according to the emotional change gradient and the breathing point indication function , the weight is dynamically adjusted: ; ; ; Among them, represents the weight of the emotional parameter of the -th frame; represents the basic weight of the emotional parameter; represents the adjustment coefficient of the emotional parameter; represents the norm of the emotional change gradient of the -th frame; represents the weight of the phrase structure parameter of the -th frame; represents the basic weight of the phrase structure parameter; represents the adjustment coefficient of the phrase structure parameter; represents the indication function of whether the -th frame is a breathing point; represents the smoothness constraint weight of the -th frame; The Adam optimizer is used in the optimization process, the learning rate is set to 0.01, the batch size is 64 frames, the maximum number of iterations is 100, and the convergence threshold is . The constraint on the parameter value range is realized through the projection operation, and the constraint on the parameter change between adjacent frames is realized by adding a penalty term.

[0092] In the application example of the classical and pop music hybrid scenario, when processing a music work with obvious emotional changes and structural transitions, the optimization algorithm can preferentially retain the emotional-driven sound effect parameter characteristics at the emotional change points (such as the transition from calm to excited), and preferentially retain the sound effect parameter characteristics of the phrase structure at the phrase breathing points (such as the transition between the end of a phrase and the start of a new phrase), while maintaining the smoothness of the parameter changes throughout the audio. In the test, the music processed by this collaborative optimization algorithm has a 32% increase in the listener satisfaction score compared to using only emotional parameters or phrase parameters, especially in the processing of music turning points and emotional change points.

[0093] In some alternative embodiments, the multi-objective optimization problem can be solved using a method based on Particle Swarm Optimization (PSO). Particle Swarm Optimization is a global optimization algorithm inspired by the behavior of biological groups and is particularly suitable for solving complex optimization problems with multiple local extrema. In this implementation, each particle represents a possible sequence of sound effect parameters, and the particle swarm searches for the optimal solution in the parameter space. Compared with local optimization methods such as gradient descent, PSO performs better in avoiding getting trapped in local optimal solutions and is particularly suitable for dealing with the optimization problem of sound effect parameters with complex objective functions.

[0094] Optionally, multi-objective optimization can adopt a method based on Genetic Algorithm (GA). The Genetic Algorithm iteratively optimizes the parameter sequence by simulating natural selection and genetic mechanisms. In this implementation, each individual represents a possible sequence of sound effect parameters, and a new generation of individuals is generated through selection, crossover, and mutation operations, gradually approaching the optimal solution. The advantage of this method is that it can effectively explore different regions of the parameter space, discover diverse high-quality solutions, and provide more possibilities for sound effect processing.

[0095] In addition, in some embodiments, a multi-objective Pareto optimization method can be adopted, considering three objectives of emotional expression, phrase breathing, and smoothness simultaneously, rather than combining them with weights into a single objective. This method can generate a set of Pareto optimal solutions (Pareto front), and then select the most suitable solution according to specific music content characteristics. For example, for music passages with rich emotions, a solution with priority given to emotional expression can be selected; for passages with obvious structures, a solution with priority given to phrase breathing can be selected. Experiments show that this content-adaptive solution selection strategy can better balance the sound effect processing requirements of different music passages.

[0096] Step 5: Apply the optimized sound effect parameters to the audio playback process, adjust the audio signal characteristics in real time, and output the processed audio signal; According to an embodiment of the present application, in this step, the optimized sound effect parameters are applied to the audio playback process to adjust characteristics such as the frequency response, dynamic range, spatial sense, and timbre of the audio signal in real time, output the enhanced audio signal, and provide an auditory experience close to a live performance.

[0097] Step 5.1: Construction of the sound effect processing module; According to the embodiment of the present application, a processing link including multiple sound effect processing units is constructed to achieve comprehensive control of the audio signal: Equalizer unit: Control the gain of different frequency bands and adjust the frequency response characteristics of the audio; Dynamic processing unit: Includes compressors, limiters, expanders, etc., to control the dynamic range of the audio; Reverberation unit: Simulates the reflection characteristics in different acoustic environments to enhance the sense of space; Stereo processing unit: Controls the sound image position and width to enhance the spatial positioning effect; Harmonic processing unit: Adds appropriate harmonic components to enrich the timbre characteristics.

[0098] The processing chain can be represented as a cascaded function combination: ; Where: represents the processed output signal; represents the input audio signal; represents the function of the represents the function of the represents the function of the represents the function of the represents the parameter component in the frame optimization parameter vector that controls the -th processing unit; represents the parameter component in the frame optimization parameter vector that controls the -th processing unit; represents the parameter component in the frame optimization parameter vector that controls the -th processing unit; represents the total number of processing units.

[0099] It should be understood that the specific implementation of the above sound effect processing module can be adjusted according to application requirements. For example, the number of processing units can be increased or decreased, or the connection method of the processing units can be changed.

[0100] In a specific embodiment of the present application, the sound effect processing module adopts the following specific implementation: Equalizer unit: Implemented as a 10 - band parametric equalizer with center frequencies of 31.5 Hz, 63 Hz, 125 Hz, 250 Hz, 500 Hz, 1 kHz, 2 kHz, 4 kHz, 8 kHz, and 16 kHz respectively. Each band has parameters of gain (in the range of ±12 dB), Q value (in the range of 0.1 - 10), and filter type (peak / valley, high / low shelf, high / low pass). The digital implementation uses a cascaded structure of bi - quadratic IIR filters; Dynamic processing unit: It includes a multi-band compressor (3 bands: low, mid, high), and the threshold (-60dB to 0dB), ratio (1:1 to 20:1), attack time (0.1ms to 100ms), and release time (10ms to 1000ms) parameters can be set independently for each band; Reverb unit: It is implemented based on convolution reverb technology and includes parameters such as room size (1 - 100 cubic meters), initial reflection intensity (-30dB to 0dB), reverb time (0.1 second to 10 seconds), reverb density (50% to 100%), high-frequency attenuation (0 to 12dB / Oct), and dry-wet ratio (0% to 100%); Stereo processing unit: It realizes panning and stereo width control and includes parameters such as left-right balance (-100% to 100%), width (0% to 200%), stereo correlation (-100% to 100%), and phase correlation (-180° to 180°); Harmonic processing unit: It is implemented based on waveform shaping technology and generates harmonics through non-linear functions, and includes parameters such as harmonic type (even, odd, or mixed), harmonic intensity (0% to 100%), harmonic frequency range (20Hz to 20kHz), and harmonic spectrum tilt (-12dB / Oct to 12dB / Oct).

[0101] The connection order of the processing units is: input signal → equalizer → dynamic processing → harmonic processing → stereo processing → reverb → output signal. Each processing unit can be adjusted in real time according to the optimized audio effect parameters.

[0102] In an application example in the podcast scenario, when processing a podcast content containing dialogue and background music, this audio effect processing module can automatically adjust the parameters according to the content characteristics: in the dialogue paragraph, enhance the clarity of the vocal frequency range (250Hz - 4kHz), moderately compress the dynamic range to improve intelligibility, and reduce reverb to avoid blurring; in the music transition paragraph, expand the frequency response range, relax the dynamic range, increase the stereo width and reverb effect to create a more open sound field. Tests show that for the podcast content processed by this module, the speech clarity score has increased significantly, the music expressiveness score has also increased, and the overall listening coherence and professional texture have been maintained.

[0103] In some alternative embodiments, the audio effect processing module may adopt a parallel processing structure instead of the serial processing link described above. In this implementation, the input signal is simultaneously sent to each processing unit. After each unit independently processes it, the outputs of each unit are mixed according to a certain ratio through a mixing network. The advantage of this parallel structure is that it avoids the influence of the processing order on the final effect and can more finely control the contribution degree of each processing unit. For example, for some solo instrument segments, the mixing ratio of the reverb unit can be reduced to maintain the clarity of the original timbre; while for ensemble segments, the mixing ratio of the stereo processing unit can be increased to enhance the sense of space.

[0104] Optionally, the audio effect processing module may adopt an adaptive processing architecture to dynamically adjust the parameters and processing flow of the processing unit according to the characteristics of the input signal. For example, a side-chain processing mechanism can be introduced. When detecting vocal content, the volume of other frequency bands is automatically reduced to highlight the clarity of the vocals; when detecting transient signals (such as drum beats), the compression ratio is temporarily reduced to retain the dynamic impact. This adaptive processing can finely adjust according to the characteristics of different audio contents while maintaining the overall audio effect style.

[0105] In addition, in some embodiments, a neural network audio effect processing unit can be introduced into the processing link, such as a neural network for timbre transfer, audio denoising, or spectral enhancement based on deep learning. These neural network units can learn complex audio effect processing mapping relationships and achieve effects that are difficult to achieve with traditional digital signal processing. For example, a neural network can be trained to convert audio recorded in a recording studio into the sound effect of a live concert in a concert hall, adding natural acoustic characteristics and a sense of space while retaining the original content. Experiments show that this method of combining traditional digital signal processing and neural network processing can achieve a good balance between computational efficiency and audio effect quality.

[0106] Step 5.2: Parameter interpolation processing; According to an embodiment of the present application, since the frame rate of the audio effect parameter sequence is usually lower than the sampling rate of the audio signal, parameter interpolation is required to achieve smooth parameter changes: For each audio sampling point, calculate the corresponding frame index (which may be a non-integer) according to its time position; Use linear interpolation or a higher-order interpolation method to calculate the interpolation parameter at this time point: ; Where: represents the interpolation parameter at time ; represents the parameter value of the th frame; represents the parameter value of the th frame; represents the interpolation coefficient, satisfying ; represents the time variable; represents the frame index.

[0107] In addition, it should be noted that the choice of interpolation method can be adjusted according to the computing resources and real-time requirements. For example, in the case of limited computing resources, a linear interpolation method with lower computational complexity can be selected.

[0108] Step 5.3, Real-time sound effect processing; According to the embodiments of the present application, the interpolated sound effect parameters are applied to the sound effect processing chain to achieve real-time processing of audio signals: The audio signal is processed in segments, and the length of each segment is L sampling points; The current sound effect processing parameters are applied to each segment of the signal; The processed signal segments are connected using the Overlap-Add method to avoid discontinuities between segments; The processed audio signal is output.

[0109] Step 5.4, Adaptive performance optimization; According to an embodiment of the present application, based on the system resource status and real-time requirements, the processing strategy is adaptively adjusted: Monitor the processing delay. When the delay exceeds the threshold, simplify the processing flow to ensure real-time performance; Dynamically adjust the processing accuracy and interpolation method according to the device performance; Perform cache optimization on the processing algorithm to reduce repeated calculations.

[0110] Therefore, through the above real-time sound effect processing and application steps, the present application realizes intelligent sound effect control of the played audio and provides an auditory experience close to a live performance.

[0111] In a specific embodiment of the present application, the adaptive performance optimization can be achieved through the following methods: Delay monitoring: Use a sliding window to calculate the average processing delay of the most recent N frames. When the average delay exceeds 80% of the target delay (such as 10 milliseconds), trigger performance optimization; Processing level adjustment: According to the delay situation, divide the processing flow into three precision levels: high, medium, and low. The high-precision level enables all processing units and the complete parameter set; the medium-precision level simplifies some computationally intensive units (such as simplifying a 10-band equalizer to a 5-band equalizer); the low-precision level only retains the most critical processing units (such as the equalizer and dynamic processing); Cache Optimization: For parameters that change slowly (such as reverb parameters), adopt a frame-interval update strategy to reduce the calculation frequency; for computationally intensive convolution operations, use FFT for acceleration and cache the frequency-domain conversion results.

[0112] Tests show that this adaptive performance optimization method can maintain the real-time performance of audio processing on various hardware platforms (from low-end mobile devices to high-performance workstations), while maximizing the audio quality.

[0113] In some alternative embodiments, an importance-based processing resource allocation strategy can be adopted. Different audio processing units have different degrees of influence on the final listening experience. Computational resources can be dynamically allocated according to the characteristics of the audio content and the listening importance of the processing units. For example, for vocal content, the equalizer and dynamic processing have a greater impact on clarity, and the processing precision of these units should be ensured first; while for environmental sound effects, the reverb unit has a more significant impact, and its priority should be appropriately increased.

[0114] Optionally, pre-computation and look-up table techniques can be used to optimize performance. For some processing units that are computationally intensive but have a limited input parameter space (such as certain types of filters or non-linear effects), the response results for different parameter combinations can be pre-computed and stored in a look-up table. During runtime, the results can be directly retrieved from the table or obtained through simple interpolation, significantly reducing the real-time computational burden.

[0115] In addition, in some embodiments, a parallel computing architecture can be utilized to improve the processing efficiency. Modern computing devices (such as multi-core CPUs, GPUs, or dedicated DSPs) generally support parallel computing. The audio processing tasks can be decomposed into sub-tasks suitable for parallel computing. For example, the processing tasks for different frequency bands can be assigned to different processing cores, or the SIMD (Single Instruction Multiple Data) instruction set can be used to process multiple samples simultaneously. For devices that support GPU acceleration, some processing units suitable for parallel computing (such as convolutional reverb) can be migrated to the GPU for execution to further improve the processing efficiency.

[0116] Finally, for cloud application scenarios, a distributed processing architecture can be adopted. The computationally intensive audio analysis and parameter generation tasks can be deployed on cloud servers, while the real-time audio processing tasks can be deployed on terminal devices. This architecture can make full use of cloud computing resources for complex analysis, while ensuring the real-time responsiveness of terminal devices, which is particularly suitable for mobile devices or embedded systems with limited computing resources.

[0117] The method for audio sound control of the played audio proposed in this application realizes intelligent sound control based on content understanding by integrating music structure analysis, emotion calculation, and phrase breathing characteristic recognition technologies, and has the following technical effects: Enhance music expressiveness: By identifying and enhancing the internal structure of music and the breathing characteristics of musical phrases, the playback of recorded music becomes more vivid, with richer emotional expression, and the auditory experience is closer to that of a live performance, thus improving the expressiveness of music recording.

[0118] Improve the effect of emotional transmission: The sound effect control based on emotional computing can accurately capture and enhance the emotional changes in music, significantly improving the effect of emotional transmission. The emotional resonance experience of the audience is deeper. It not only improves the accuracy of emotion recognition but also enhances the effect of sound effects in emotional transmission.

[0119] Create a natural and flowing auditory experience: The recognition and enhancement of the breathing characteristics of musical phrases make the music playback have a more natural sense of flow, avoiding the mechanical and rigid feeling caused by traditional sound effect processing methods.

[0120] Strong adaptability: This method is applicable to various music styles and genres, and can adaptively adjust the processing strategy according to the characteristics of different music contents, with strong universality. According to the embodiments of the present application, significant sound effect enhancement effects have been achieved in tests covering 10 music styles such as classical, jazz, and pop.

[0121] High computational efficiency: Through an optimized algorithm structure and adaptive performance adjustment, this method can achieve real-time processing on ordinary consumer devices, with the processing delay controlled within 10 milliseconds, meeting the requirements of real-time playback, without the support of professional audio equipment, and facilitating wide application.

[0122] It should be understood that the above embodiments are merely exemplary, and the protection scope of the present application is not limited to the specific embodiments described. After understanding the basic idea of the present application, those skilled in the art can make various changes or modifications to these embodiments without departing from the spirit and scope of the present application.

[0123] Application examples of this embodiment 1 Application scenario: In this embodiment, the method for sound effect control of the played audio proposed in the present application is applied to the classical music playback scenario. The specific test object is a recording of a piano concerto. Although the recording has clear sound quality, there are obvious gaps compared with the live performance in terms of emotional expression and the sense of music flow, lacking a natural sense of musical phrase breathing and emotional level changes.

[0124] 2 Example of the implementation process: Step 1: Recognition of music structure and musical phrase breathing points; The target piano concerto audio is input into the self-attention audio analysis network for processing. The network first extracts frequency-domain features such as MFCC, spectral centroid, and spectral flux of the audio, as well as time-domain features such as energy envelope and amplitude envelope, to form a sequence of feature vectors. Then, the feature sequence is analyzed through 4 layers of Transformer encoders and a bidirectional LSTM layer to output structure markers and breathing point markers.

[0125] In this example, the network successfully identifies multiple structural parts of the music, including large structures such as the exposition (0 - 185 seconds), development section (186 - 412 seconds), and recapitulation (413 - 603 seconds), as well as the theme presentation, transitional passages, and peak parts within each section. At the same time, the network identifies approximately 127 phrase units, with an average phrase length of 4.7 seconds, and each phrase contains three types of breathing point markers: starting point, peak point, and ending point.

[0126] A detailed analysis of the first theme section (15 - 42 seconds) shows that the network successfully captures the natural fluctuations of the phrase, accurately marks the starting, peak, and ending points of the phrase, and the recognition accuracy reaches 91% (compared with the reference data annotated by professional music scholars).

[0127] Step 2: Emotional feature extraction and emotional mapping; Based on the previously recognized structural information, the system extracts high-level emotional features of the audio, including features such as rhythm density, rhythm regularity (average value 0.72, range 0.45 - 0.96), harmonic complexity (average value 0.61, range 0.33 - 0.89), timbre brightness (average value 0.56, range 0.21 - 0.85), and dynamic contrast (average value 0.68, range 0.32 - 0.94).

[0128] Through the emotion classification model, the system maps these features into a two-dimensional emotional space, generating an emotional state sequence covering the entire music. The analysis shows that the emotional trajectory of this piano concerto presents an obvious change pattern in the valence-arousal space: the theme part of the exposition is located in the high valence - medium arousal region (valence 0.75, arousal 0.55), showing bright and elegant emotions; the middle part of the development section is located in the medium valence - high arousal region (valence 0.45, arousal 0.82), presenting tense and intense emotions; the ending part of the recapitulation is located in the high valence - low arousal region (valence 0.85, arousal 0.25), showing peaceful and soothing emotions.

[0129] The emotional sound effect parameter mapping network converts the recognized emotional state into sound effect parameters. For example, in the high-arousal section of the development section (about 240 - 280 seconds), the system generated the following sound effect parameters: mid-high frequency gain (+3.5 dB in the 2 - 4 kHz band, +2.8 dB in the 4 - 8 kHz band), dynamic range expansion (compression ratio 1:1.2), reverberation time shortening (1.2 seconds), and stereo width increase (135%). These parameters are designed to enhance the tension and dynamic impact of the music, which is consistent with the results of the emotional analysis.

[0130] Step 3: Generate dynamic sound effects for the phrase structure; Based on the phrase breathing points recognized in Step 1, the system constructs a phrase breathing model and generates corresponding sound effect parameters for different phrase stages. In a typical example of phrase processing (the third phrase of the exposition, about 23 - 29 seconds), the changes in the sound effect parameters generated by the system are as follows: At the beginning stage of the phrase (23 - 24.5 seconds): Progressive clarity enhancement parameters are applied, and the mid-frequency (1 - 2 kHz) gain gradually increases from +1.2 dB to +2.8 dB, and the directional focusing parameter increases from 0.2 to 0.6, making the notes at the beginning of the phrase clearer and more directional.

[0131] At the peak stage of the phrase (24.5 - 27 seconds): Sound field expansion parameters are applied, the stereo width expands from 110% to 145%, the dynamic range expands by 20% from the original dynamic, and the reverberation humidity increases by 15%, making the expressiveness of the peak part more abundant.

[0132] At the end stage of the phrase (27 - 29 seconds): Natural decay parameters are applied, the high-frequency (4 - 8 kHz) gain gradually decreases from +2.0 dB to +0.5 dB, the reverberation time extends from 1.8 seconds to 2.3 seconds, and the stereo width gradually narrows from 145% to 125%, creating an auditory effect of natural ending.

[0133] For the transition between adjacent phrases (29 - 30 seconds), the system applies a smooth transition function to ensure that the sound effect parameters change smoothly between phrases, avoiding sudden changes or incoherence in the listening experience.

[0134] Step 4: Coordinate and optimize multi-dimensional sound effect parameters; During the implementation process, the system detected some potential conflicts between the emotion-driven sound effect parameters and the phrase structure sound effect parameters. For example, in a section of the development section (290 - 310 seconds), the emotional analysis indicates that the low frequency needs to be enhanced (to express a sense of heaviness), while the phrase structure analysis suggests enhancing the high frequency (to highlight the clarity at the beginning of the phrase).

[0135] The system solves these conflicts through a multi-objective optimization algorithm. Specifically, a dynamic weight adjustment strategy is used in this paragraph: Since 290 - 295 seconds is the emotional change point (from excitement to heaviness), the system increases the weight of the emotional parameter (from 0.4 to 0.65); while at the phrase breathing point (298 seconds, a new phrase starts), the system increases the weight of the phrase structure parameter (from 0.4 to 0.62).

[0136] Through iterative solution using the Adam optimizer (convergence is achieved after 53 iterations), the system generates optimized sound effect parameters that take into account both emotional expression and phrase breathing. While retaining the low-frequency enhancement (meeting the emotional needs), the optimized parameters moderately enhance the mid-high frequencies at the phrase start point (meeting the phrase clarity needs) and achieve a smooth transition between the two.

[0137] Step Five: Real-time Sound Effect Processing and Application; The system applies the optimized sound effect parameters to the audio playback process. In the processing chain, each processing unit is connected in the order of equalizer → dynamic processing → harmonic processing → stereo processing → reverb. For real-time processing, the system uses linear interpolation method to calculate the sound effect parameters of each sampling point and applies them to the corresponding processing unit.

[0138] In the test scenario of resource-constrained mobile devices, the system monitors that the processing delay occasionally exceeds the target value (10 milliseconds), and at this time, adaptive performance optimization is triggered: simplify the equalizer from 10 bands to 5 bands, reduce the reverberation calculation accuracy by 20%, and increase the processing frame length from 512 samples to 1024 samples. These optimization measures reduce the processing delay by about 65% and ensure the real-time performance on various hardware platforms.

[0139] 3. Verification of Technical Effects: To verify the effectiveness of this method, we conducted objective measurements and subjective evaluation tests on the audio before and after processing.

[0140] Objective Measurements: Through acoustic parameter analysis, the processed audio has obvious improvements in aspects such as dynamic range, spectral balance, and sound field width. The dynamic range is extended from 12 dB of the original recording to 17 dB, with an increase of about 42%; the spectral energy distribution is more balanced, especially the energy distribution in the mid-frequency band (500 Hz - 4 kHz) has a similarity improvement of 38% with the reference pattern of professional on-site recording; the stereo width parameter increases from the original 0.68 to 0.85, with an increase of about 25%.

[0141] Subjective Evaluation Test: Invite 50 listeners (including 25 music professionals and 25 ordinary listeners) to conduct a double-blind comparison test on the audio before and after processing, and focus on evaluating two core technical effects: music expressiveness and emotional transmission effect.

[0142] Improved musical expressiveness: Listeners evaluate the vitality, expressiveness of the music and its proximity to a live performance. The results show that the processed audio has an average 45% increase in expressiveness scores (+42% for professional listeners, +48% for ordinary listeners). Especially in terms of the sense of breathing in musical phrases, 95% of professional listeners believe that the processed audio has a more natural musical flow and is closer to the breathing rhythm of a live performance.

[0143] Improved emotional transmission effect: Listeners evaluate the effect of the music in transmitting emotions and the emotional resonance experience. The results show that the processed audio has an average 53% increase in emotional transmission effect scores (+47% for professional listeners, +59% for ordinary listeners). At emotional change points (such as the transition from excitement to calm), 88% of listeners believe that the processed audio can express emotional changes more clearly and enhance the emotional resonance experience.

[0144] In addition, in the adaptability test of different music styles, when this method is applied to 10 different styles of music such as classical, jazz, and pop, significant sound effect enhancement results are achieved, demonstrating the wide applicability of this method. In terms of computational efficiency, through adaptive performance optimization, real-time processing can be achieved even on mid-range mobile devices, with the processing delay controlled within 10 milliseconds, meeting the real-time requirements of various application scenarios.

[0145] The above describes the embodiments of the present invention, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A method for sound effect control of played audio, characterized in that, It includes the following steps: Obtain the audio to be played, preprocess the audio to obtain a standardized audio signal, and use an audio analysis network to identify the structural key points and phrase breathing points in the audio, and generate music structure markers and phrase breathing point markers; Based on the music structure markers, extract the emotional features of the audio, and convert the emotional features into emotion-driven sound effect parameters through an emotion sound effect parameter mapping network; Based on the phrase breathing point markers, generate phrase structure sound effect parameters according to the phrase breathing point markers; Combine the emotion-driven sound effect parameters and the phrase structure sound effect parameters, and generate optimized sound effect parameters through a multi-objective optimization algorithm; Apply the optimized sound effect parameters to the audio playback process, adjust the audio signal characteristics in real time, and output the processed audio signal.

2. The method for sound effect control of played audio according to claim 1, wherein, The audio analysis network is a network based on the self-attention algorithm, including: An input layer for receiving an audio feature vector sequence; A self-attention layer for calculating the attention weights between each time point in the feature sequence; A multi-head attention layer for enhancing the model's expressive ability; A bidirectional long short-term memory network layer for capturing temporal information; An output layer for outputting the probability distributions of the structure markers and the breathing point markers.

3. The method for sound effect control of played audio according to claim 1, wherein The steps of extracting the emotional features of the audio include: Extract the rhythm features, harmony features, timbre features, dynamic features, and structural association features of the audio; Use an emotion classification model to map the extracted features into a two-dimensional emotion space to form an emotion state sequence, and the two-dimensional emotion space includes a valence dimension and an arousal dimension.

4. The method for sound effect control of played audio according to claim 1, wherein, The emotion sound effect parameter mapping network includes: An input layer for receiving an emotion state vector; Multiple hidden layers for performing non-linear transformations; An output layer for generating a sound effect parameter vector including equalizer parameters, dynamic range parameters, reverb parameters, and stereo parameters.

5. The method for sound effect control of played audio according to claim 1, characterized in that, The steps of generating phrase structure sound effect parameters according to the phrase breathing point markers include: Construct a phrase breathing model, and abstract the phrase breathing process into three stages: start, peak, and end; For the phrase start stage, apply progressive clarity enhancement and directional focusing parameters; For the phrase peak stage, apply sound field expansion and dynamic range enhancement parameters; For the phrase end stage, apply natural decay and acoustic space blanking parameters; Process the transition between adjacent phrases to ensure the continuity and naturalness of the sound effect changes.

6. The method for sound effect control of played audio according to claim 1, characterized in that The multi-objective optimization algorithm includes the following steps: Detect potential conflicts between the emotion-driven sound effect parameters and the phrase structure sound effect parameters; Construct a multi-objective optimization model, including an emotion parameter distance function, a phrase parameter distance function, and a smoothness penalty function; Dynamically adjust the weights of the optimization objectives according to the music content characteristics and conflict situations; Use an iterative optimization algorithm to solve the multi-objective optimization problem to obtain the final sound effect parameter sequence.

7. The method for sound effect control of played audio according to claim 1, characterized in that, The steps of applying the optimized sound effect parameters to the audio playback process include: Construct a processing link including an equalizer unit, a dynamic processing unit, a reverb unit, a stereo processing unit, and a harmonic processing unit; For each audio sampling point, calculate the corresponding parameter interpolation according to its time position; Apply the interpolated sound effect parameters to the sound effect processing link to achieve real-time processing of the audio signal. Adaptive adjustment of the processing strategy according to the system resource status and real-time requirements.

8. The method for controlling the sound effect of the played audio according to claim 7, wherein The said adaptive adjustment of the processing strategy includes: Using a sliding window to calculate the processing delay and triggering performance optimization when the average delay exceeds the threshold; Dividing the processing flow into different precision levels according to the delay situation and reducing the processing precision in the case of high delay; For parameters that change slowly, adopting a frame interval update strategy to reduce the calculation frequency; For computationally intensive operations, using an acceleration algorithm and caching intermediate results to reduce the computational complexity.

9. A device for sound effect control of played audio, which is used to implement the method for sound effect control of played audio according to any one of claims 1-8, characterized in that, Including: An audio analysis module for identifying the structural key points and phrase breathing points in the audio and generating music structure markers and phrase breathing point markers; An emotional feature extraction module for extracting the emotional features of the audio; An emotional sound effect parameter mapping module for converting the emotional features into emotion-driven sound effect parameters; A phrase structure sound effect generation module for generating phrase structure sound effect parameters according to the phrase breathing point markers; A parameter collaborative optimization module for combining the emotion-driven sound effect parameters and the phrase structure sound effect parameters and generating optimized sound effect parameters through a multi-objective optimization algorithm; A sound effect processing module for applying the optimized sound effect parameters to the audio playback process, adjusting the audio signal characteristics in real time, and outputting the processed audio signal.