Audio processing method and system, computing device and storage medium
By constructing a feature compensation model and adjusting audio features in conjunction with the frequency response relationship of human hearing, the problem of flattened auditory perception caused by existing audio processing methods is solved, and the spatial sense and layering are preserved and the auditory experience is improved during the standardization of audio loudness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHUHAI KINGSOFT ONLINE GAME TECH CO LTD
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-17
AI Technical Summary
Existing audio processing methods ignore the non-linear characteristics of human hearing during loudness standardization, resulting in a flattened sound when the processed audio is integrated into the final product, destroying the sense of space and layering.
By acquiring various types of initial and reference audio features, the feature relationships between audios are determined, a feature compensation model is constructed, and audio features are adjusted to obtain the target audio under the constraint of auditory frequency response. This ensures that the difference in human ear sensitivity to different frequency bands is considered during loudness standardization, and maintains the spatial sense and layering of the audio design.
During the loudness standardization process, the differences in human ear sensitivity to loudness of different frequency bands are fully considered to avoid destroying the spatial sense and layering of the sound design, thereby improving the overall listening experience and spatial layering of the audio and adapting it to the overall auditory experience of the target scene.
Smart Images

Figure CN121884835A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of digital cultural products, and in particular to an audio processing method, system, computing device, and storage medium. Background Technology
[0002] In industrial audio production scenarios, it is necessary to standardize the loudness of a large number of different types of audio files (such as human dialogue, skill sound effects, ambient sounds, background music, narration, etc.), and ensure that these audios, after being integrated into the final product (games, movies, audio programs, etc.), have a unified loudness experience that conforms to the mainstream industry standards.
[0003] Most existing audio processing methods are based on the overall loudness target value of the audio to be processed. They boost or attenuate the loudness of the audio to be processed, and then standardize the loudness of each audio element using these methods before integrating the standardized audio into the final product (such as games, movies, audio programs, etc.). However, while the above solutions can achieve the target loudness value, the actual listening experience still shows significant differences, destroying the spatial sense and layering of the audio design, and making the processed audio sound "flat".
[0004] In conclusion, standardizing the loudness of the audio to be processed using existing audio processing methods will result in a poor listening experience when the processed audio is integrated into the final product. Summary of the Invention
[0005] In view of this, embodiments of this specification provide an audio processing method. One or more embodiments of this specification also relate to an audio processing system, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0006] According to a first aspect of the embodiments of this specification, an audio processing method is provided, comprising: Acquire various types of initial audio and reference audio features; Based on the initial audio features of the initial audio, determine the audio feature relationships between different types of initial audio. A feature compensation model is constructed based on the reference audio features and the relationship between audio features, constrained by the auditory frequency response relationship. Based on the feature compensation model, the feature compensation values of the initial audio are obtained; Based on feature compensation values, the initial audio features are adjusted to obtain various types of target audio.
[0007] According to a second aspect of the embodiments of this specification, an audio processing system is provided, comprising: The file import and export interface is configured to obtain various types of initial audio; The target parameter configuration interface is configured to obtain reference audio features; The multidimensional audio feature analysis engine is configured to determine the audio feature relationships between different types of initial audio based on the initial audio features of the initial audio. The human ear hearing model integration unit is configured to construct a feature compensation model based on reference audio features and audio feature relationships, constrained by the auditory frequency response relationship. The processing strategy decision unit is configured as follows: Based on the feature compensation model, the feature compensation values of the initial audio are obtained; Based on feature compensation values, the initial audio features are adjusted to obtain various types of target audio.
[0008] According to a third aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which implement the steps of the above method when executed by the processor.
[0009] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0010] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0011] This specification provides an audio processing method in one embodiment that acquires multiple types of initial audio and reference audio features; determines the audio feature relationships between different types of initial audio based on the initial audio features; constructs a feature compensation model based on the reference audio features and the audio feature relationships, constrained by auditory frequency response relationships; obtains feature compensation values for the initial audio based on the feature compensation model; and adjusts the initial audio features based on the feature compensation values to obtain multiple types of target audio. This method can fully consider the differences in human ear sensitivity to loudness at different frequency bands when standardizing the loudness of the initial audio, allowing for targeted gain or loss of different frequency bands, avoiding damage to the spatial sense and layering of the sound design, and improving the listening experience of the processed audio. Furthermore, when adjusting the loudness of each type of audio, it strictly adheres to the relative loudness relationships and spectral balance relationships that each type of audio should maintain in the final mixing context. This ensures that each type of audio meets loudness standards while adapting to the target scene, improving the overall listening experience and spatial layering of each type of audio after integration into the final product, and avoiding unreasonable masking effects between different types of audio. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating an audio processing method provided in one embodiment of this specification; Figure 2 This is a schematic diagram of an audio feature analysis result provided in one embodiment of this specification; Figure 3 This is a schematic diagram of the execution flow of an audio processing method provided in an embodiment of this specification in a software application; Figure 4 This is a schematic diagram of the architecture of an audio processing system provided in one embodiment of this specification; Figure 5 This is a schematic diagram of the structure of an audio processing system provided in one embodiment of this specification; Figure 6 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0013] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0014] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all of one or more of the associated listed items that may be combined.
[0015] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0016] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0017] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0018] Equal loudness curves: Equal loudness curves are a set of curves describing how the perceived loudness of sounds at different frequencies changes with the actual sound pressure level. They reflect the physiological characteristics of the human auditory system, which is sensitive to mid-frequency sounds and relatively insensitive to low and high frequencies (especially at low sound pressure levels). For example, on a 40-square equal loudness curve, 40 dB SPL at 1 kHz sounds the same as approximately 52 dB SPL at 100 Hz.
[0019] Loudness: Loudness is the subjective perception of volume produced by a sound signal in the human ear. It is a comprehensive psychoacoustic quantity that combines time, frequency, and intensity. For example, short-term loudness describes the perceived loudness that changes over time, while the overall loudness (such as the Loudness Units Full Scale (LUFS) value) characterizes the average perceived loudness of the sound signal as a whole.
[0020] Audio: Audio refers to vibrational signals that can cause the sensation of sound in the auditory system, and their digital or analog representation. For example, a .wav format file encoded with PCM at a sampling rate of 48kHz and a bit depth of 24bit.
[0021] Computing devices: Computer devices are computer devices designed to perform one or more dedicated tasks. Compared to personal mainframes, they are weaker in performance but have significant advantages in terms of size and power consumption. They are commonly used in various electronic and mechanical control devices.
[0022] When standardizing audio loudness, existing techniques primarily rely on a single overall loudness target value for simple gain boosting or attenuation, ignoring the non-linear characteristics of human hearing. Isoloudness curves clearly show that the human ear's sensitivity to loudness varies across different frequency bands. This means that even if the audio loudness value meets the standard, the actual listening experience (such as volume perception, sound distance perception, and spatial layering) can still show significant differences, disrupting the spatial and layered feel of the sound design and resulting in a "flattened" audio experience.
[0023] Furthermore, existing technologies typically process individual audio files in isolation, lacking the ability to deeply analyze the inherent characteristics of the source audio material, including the audio's dynamic envelope features (Attack-Decay-Sustain-Release, ADSR), spectral energy distribution, spatial information (depth of field and sound field width), and rendering features (such as equalizer (EQ), distortion, reverberation, etc.). At the same time, they fail to consider the relative loudness relationships and spectral balance requirements of these characteristics in the final mixing context, resulting in processed audio that is difficult to adapt to the target scene and deviates from the overall auditory context.
[0024] Based on this, an audio processing method is provided in this specification. This specification also relates to an audio processing system, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0025] See Figure 1 , Figure 1 A flowchart of an audio processing method according to an embodiment of this specification is shown, including the following specific steps: Step 102: Obtain various types of initial audio and reference audio features.
[0026] The embodiments in this specification apply to scenarios where multiple types of audio need to be integrated into the final product, such as digital cultural products, movies, and audio programs. It is essential to ensure that these audio files, once integrated into the final product, provide a unified loudness experience that conforms to mainstream industry standards (such as EBU R128, ITU-R BS.1770, ATSC A / 85, etc.).
[0027] The initial audio is the raw audio without any audio feature processing. For example, a user selects a dialogue audio clip from an audio library as the initial audio.
[0028] The various types of initial audio are audio that play different roles and have different acoustic characteristics in the final mix product. For example, in a role-playing game, the initial audio includes "human dialogue" for character conversations, "skill sound effects" for spell casting, "ambient sounds" for creating scene atmosphere, "background music" for enhancing emotions, and "narration" for story narration.
[0029] Reference audio features are the desired audio feature metrics achieved after adjusting (feature compensation) the initial audio features. They define the ideal state of each type of audio in the target mixing context. For example, reference audio features can be a set of preset numerical targets, including the overall target loudness, the target spectral balance curve (such as a broadcast speech enhancement curve), the target dynamic range, and the target loudness ratio between different audio types (such as dialogue being 6dB higher than ambient sound).
[0030] Optionally, audio features may include: Dynamic features (i.e. envelope features): Accurately analyze the attack, decay, sustain, and release envelope patterns of audio, calculate dynamic range values, and quantify the dynamic variation patterns of audio intensity.
[0031] Spectral characteristics: By constructing an audio spectral energy distribution map, we focus on analyzing the energy distribution of the frequency bands sensitive to the human ear (such as 2kHz-5kHz) and their relative relationship with other frequency bands, and identify the locations of spectral peaks and valleys.
[0032] Spatial characteristics: For stereo / surround sound audio, assess perceived depth of field (distance perception), sound field width (spatial extension range), and identify spatial parameters such as sound image positioning coordinates and reverberation decay time.
[0033] Rendering features: Identify whether the audio contains strong EQ, distortion, reverb, delay and other rendering effects, quantify the feature parameters of these effects (such as reverb wet-dry ratio, distortion, EQ gain), and define the sound "color" attribute.
[0034] Indicatively, Figure 2 This diagram illustrates an audio feature analysis result provided by one embodiment of this specification.
[0035] like Figure 2 As shown, audio features may include: Envelope characteristics: Attack time: 0.05ms; Decay time: 200ms; Hold level: -15dBFS; Release time: 1.2s; Dynamic range: 11.5LU.
[0036] 1 / 3 octave band spectrum: Spectral centroid: 108 kHz; Sensitive frequency band energy percentage: 32%; Low frequency / high frequency energy ratio: 1.2:1; Anomaly marker: 60 kHz resonant peak.
[0037] Sound field distribution diagram: Sound field width: 85%; Perceived depth of field: 4.2M; Channel correlation coefficient: 0.72; Sound image localization: Horizontal 30° / Vertical 0°.
[0038] Rendering parameter radar chart: EQ: Low frequency +2dB, High frequency +dB; Reverb: RT=0.8s / wet-to-dry ratio 30%; Total harmonic distortion: 1.2%; Rendering style code: Warm (0110).
[0039] After obtaining the audio feature analysis results, each audio feature can be written into the feature data master table (XML / Excel file), and a feature difference matrix can be further generated (generated based on the audio features of the initial audio and the reference audio features).
[0040] Optionally, one way to obtain multiple types of initial audio is to batch read audio files from a preset audio resource library or file system under a specified path. For example, the system automatically scans and loads all .wav and .mp3 format files under the / assets / audio / directory as initial audio. Another way to obtain multiple types of initial audio is to receive audio data streams from an upstream content management system through an Application Programming Interface (API). For example, a game engine dynamically submits sound effects and dialogue clips that need to be processed through an audio API at runtime. This specification does not limit the methods for obtaining multiple types of initial audio in the embodiments.
[0041] Optionally, one way to obtain reference audio features is to select a matching template from a preset configuration template library based on the type label or metadata of the initial audio. For example, when "voice" type audio is detected, the "podcast voice optimization" template is automatically loaded as the reference audio feature. Another way to obtain reference audio features is to dynamically generate or modify the reference features by receiving interactive input or parameter adjustments from the user. For example, a mixing engineer drags the "warmth" and "clarity" sliders on the graphical interface, and the system calculates and generates the corresponding target spectral curve in real time. The embodiments in this specification do not limit the method of obtaining reference audio features.
[0042] For example, in a film post-production scenario, an audio engineer needs to process five initial audio tracks: one lead actor's dialogue (type: human voice), two ambient sounds (type: ambient sound), one explosion sound effect (type: sound effect), and one end credits theme (type: music). The system loads the corresponding reference audio features based on the project's preset "Film Mixing - Action Scene" template. This template specifies: the overall loudness target is -27 LKFS; the dialogue's spectral energy must dominate in the 1kHz-4kHz range; the transient peaks of the sound effects can briefly exceed the dialogue's by 3dB; and the average loudness of the music must be 9dB lower than the dialogue's.
[0043] Step 102 provides a data foundation for determining the audio feature relationships between different types of initial audio and constructing a feature compensation model by acquiring various types of initial audio and reference audio features.
[0044] Step 104: Based on the initial audio features of the initial audio, determine the audio feature relationships between different types of initial audio.
[0045] Initial audio features are a set of multidimensional parameters extracted from an initial audio signal to quantify its acoustic properties. For example, initial audio features may include dynamic envelope features describing the change of signal amplitude over time (such as root mean square loudness, true peak value, dynamic range, and transient density), and spectral energy distribution features describing the distribution of signal energy along the frequency axis (such as spectral centroid, 1 / 3 octave band energy, and harmonic structure).
[0046] Audio characteristic relationships refer to the relative loudness and spectral balance relationships of various audio types with other audio in the final mixing context. They define how different audio elements should coordinate and coexist in a shared auditory space. For example, a relative loudness relationship could be "the perceived loudness of human dialogue should be 4-6 dB higher than that of background music," while a spectral balance relationship could be "the low frequencies (60-120 Hz) of explosion sound effects should avoid the bass drum sound effects to prevent masking."
[0047] Optionally, one way to determine the audio feature relationships between different types of initial audio based on the initial audio features is to first calculate the perceived loudness of each audio (using frequency weighting combined with equal loudness curves), and then calculate the target loudness ratio between them according to the preset priority rules for audio types (e.g., dialogue > sound effects > music > ambient sound). Another way to determine the audio feature relationships between different types of initial audio based on the initial audio features is to analyze the spectral energy distribution of each audio, use an auditory masking model to calculate potential spectral conflict regions between them, and generate coordination rules for spectral avoidance or complementarity. Yet another way to determine the audio feature relationships between different types of initial audio based on the initial audio features is to combine a machine learning model to learn common feature relationship patterns between different audio types from a large number of excellent mixing works and apply them to the current audio combination. This specification does not limit the scope of this embodiment.
[0048] For example, after extracting features from each audio segment of a movie, the system found that: explosion sound effects have high peak values but low average loudness; dialogue has a prominent spectrum at 2kHz; and ambient sound has a wide and flat spectrum. The determined audio feature relationships include: loudness relationship: average loudness of dialogue > average loudness of explosion sound effects > average loudness of ambient sound; spectrum relationship: explosion sound effects need to retain sufficient energy at 80Hz to convey impact, but dialogue is in the core frequency band of 300Hz-3kHz, and explosion sound effects and ambient sound need to be appropriately avoided (e.g., attenuated by 3dB) to ensure dialogue clarity.
[0049] Step 104 determines the audio feature relationships between different types of initial audio based on the initial audio features, providing a data foundation for the subsequent construction of the feature compensation model.
[0050] Step 106: Using the auditory frequency response relationship as a constraint, construct a feature compensation model based on the reference audio features and the audio feature relationship.
[0051] Auditory frequency response is a quantitative model or processing technique based on the non-uniform response characteristics of the human ear to sound frequencies. It characterizes the differences in sensitivity of the auditory system to sounds in different frequency bands. As a constraint, it ensures that any audio adjustment result conforms to the perceptual characteristics of the human ear at different playback volumes, avoiding the loss of low-frequency power at low volumes or the harshness of high frequencies at high volumes. For example, auditory frequency response can be equal-loudness curves, auditory masking curves, and "psychoacoustic models" of perceptual audio coding.
[0052] Constraints are conditions that limit the solution space in the mathematical feature compensation model, ensuring that the calculated compensation value is not only mathematically optimal, but also sounds natural and reasonable.
[0053] The feature compensation model is a mathematical feature compensation model. Its inputs are the initial audio features, the reference audio features, and the audio feature relationships. The output is the optimal feature compensation value. Its goal is to make the adjusted audio features as close as possible to the reference audio features while satisfying the constraints of the equal loudness curve and the audio feature relationships.
[0054] Optionally, one implementation of the feature compensation model, constrained by the auditory frequency response relationship and based on the reference audio features and the audio feature relationship, can be as follows: Establish a constrained optimization problem whose objective function is to minimize the Euclidean distance between the adjusted features and the reference features. The constraints include the equal-loudness curve selected based on the currently estimated playback sound pressure level (limiting the perceptual equalization of gain adjustment across frequency bands) and the audio feature relationship determined in step 104 (transformed into linear inequality constraints between the feature values of each audio element). Another implementation of the feature compensation model, constrained by the auditory frequency response relationship and based on the reference audio features and the audio feature relationship, can be as follows: Construct a framework based on weighted least squares, where the auditory frequency response relationship is incorporated into the objective function through a frequency-related weight matrix, and the audio feature relationship is added as an equality or inequality constraint using the Lagrange multiplier method. This specification does not limit the implementation of this method in the embodiments.
[0055] Continuing with the previous example of movie audio, assuming the target playback level is cinematic, the system selects the corresponding equal-loudness curve. The constructed feature compensation model is formalized as follows: Decision variables: gain compensation value for each audio track, gain compensation value for each frequency band of the multi-band equalizer.
[0056] Objective function: Minimize the difference between the dialogue-adjusted features and the reference features + the difference between the explosion sound effect-adjusted features and the reference features + the difference between the ambient sound-adjusted features and the reference features.
[0057] Constraint 1 (Equal Loudness Curve): After adjustment, the perceived loudness change of each audio frequency band (e.g., 125Hz) must conform to the sensitivity ratio specified by the equal loudness curve at 85 dBSPL.
[0058] Constraint 2 (Audio Feature Relationship): The energy difference between dialogue and explosion sound effects at 2kHz is ≥4dB (to ensure clear dialogue); the energy of explosion sound effects at 80Hz is ≤120% of the bass drum reference energy (to avoid low-frequency overload).
[0059] Step 106 uses the auditory frequency response relationship as a constraint and constructs a feature compensation model based on the reference audio features and the audio feature relationship, providing a prerequisite for obtaining the feature compensation value of the initial audio in the future.
[0060] Step 108: Obtain the feature compensation values of the initial audio based on the feature compensation model.
[0061] Feature compensation values are digital representations used to gain or attenuate the audio features of the initial audio in order to achieve loudness normalization, spectral balance, and mixing context harmony.
[0062] Optionally, based on the feature compensation model, one way to obtain the feature compensation value of the initial audio is to directly solve the constrained convex optimization problem using numerical optimization algorithms such as the interior-point method or the effective set method. Another way to obtain the feature compensation value of the initial audio is to transform the constrained optimization problem into an unconstrained optimization problem (e.g., using the augmented Lagrange method) and then solve it using a quasi-Newton method. This specification does not limit the embodiments in this way.
[0063] For example, after solving the model constructed in step 106, the following feature compensation values are obtained: For the "lead actor's dialogue" audio, an overall gain of +2.1dB needs to be applied, with a boost of +1.5dB at 300Hz and a boost of +2.0dB at 3kHz; For the "explosion sound effect" audio, a compression of -18dB at a ratio of 3:1 needs to be applied, with a decay of -3.0dB at 2kHz; For the "ambient sound" audio, a decay of -2.0dB at 500Hz needs to be applied, with a stereo width extension of +20%.
[0064] Step 108, based on the feature compensation model, obtains the feature compensation values of the initial audio, which can provide a data foundation for subsequent adjustments to the initial audio features.
[0065] Step 110: Adjust the initial audio features based on the feature compensation value to obtain various types of target audio.
[0066] Adjusting the initial audio features involves converting the feature compensation values into specific digital signal processing parameters and applying them to the corresponding initial audio signal.
[0067] The target audio is the final audio after feature compensation value adjustment. Its audio features not only meet the overall loudness standard, but also satisfy the requirements of relative loudness relationship and spectral balance in the target mixing context.
[0068] Optionally, one way to adjust the initial audio features based on feature compensation values to obtain multiple types of target audio is to call the corresponding audio processing library to perform batch processing on the audio files according to the gain and EQ parameters in the compensation values. Another way to adjust the initial audio features based on feature compensation values to obtain multiple types of target audio is to inject the compensation values as dynamic parameters into the audio playback pipeline in real time in a real-time audio engine (such as a game engine or digital audio workstation) to achieve non-destructive real-time processing. This specification does not limit the embodiments in this way.
[0069] For example, the system applies the compensation value obtained in step 108 to the original audio file. The processed "lead dialogue" sounds clearer and more prominent; the "explosion sound effects" no longer mask the dialogue while maintaining impact; and the "ambient sounds" provide appropriate background atmosphere without being overpowering. When these three target audio tracks are integrated into the final film mix, a professional and harmonious auditory effect is achieved without manual rebalancing.
[0070] Step 110 adjusts the initial audio features based on the feature compensation value to obtain various types of target audio. This allows for the output of audio materials that meet a unified loudness standard while maintaining a balanced, clear, and layered artistic listening experience, greatly improving the efficiency and quality of industrial audio production.
[0071] In this embodiment, multiple types of initial audio and reference audio features are acquired; based on the initial audio features, the audio feature relationships between different types of initial audio are determined; a feature compensation model is constructed based on the reference audio features and the audio feature relationships, constrained by auditory frequency response relationships; feature compensation values for the initial audio are obtained based on the feature compensation model; and the initial audio features are adjusted based on the feature compensation values to obtain multiple types of target audio. This allows for targeted gain or loss of different frequency bands when standardizing the loudness of the initial audio, taking into full account the differences in human ear sensitivity to loudness at different frequency bands. This avoids damaging the spatial sense and layering of the sound design, improving the listening experience of the processed audio. Furthermore, when adjusting the loudness of each type of audio, the relative loudness relationship and spectral balance relationship that each type of audio should maintain in the final mixing context are strictly adhered to. This ensures that each type of audio meets the loudness standard requirements while adapting to the target scenario, improving the overall listening experience and spatial layering of each type of audio after integration into the final product, and avoiding unreasonable masking effects of different types of audio.
[0072] In one optional embodiment of this specification, the auditory frequency response relationship includes equal loudness curves, and step 106 includes: Under the condition of satisfying the audio feature relationship, a feature compensation model is constructed with equal loudness curves as constraints and minimizing the feature difference between the reference audio features and the initial audio features as the objective. Step 108 includes: Solve the feature compensation model to obtain the target feature differences; Based on the differences in target features, the feature compensation value of the initial audio is determined.
[0073] Satisfying the audio feature relationship requires that the adjustment scheme (i.e., the feature compensation value and the resulting adjusted audio features) solved by the feature compensation model must conform to the constraint rules defined by the audio feature relationship. For example, if the relative loudness relationship stipulates that "the perceived loudness of human voice dialogue must be at least 4dB higher than that of background music", then any model solution must satisfy "adjusted human voice loudness - adjusted background music loudness ≥ 4dB".
[0074] Feature difference is the quantized distance between reference audio features and initial audio features, which can be measured using the norm of a vector (such as Euclidean distance) or a weighted sum. For example, loudness feature difference can be the difference (3 LUFS) between a target loudness value (such as -23 LUFS) and the current actual loudness value of the audio (such as -20 LUFS).
[0075] The target feature difference is the optimal feature difference value obtained by solving the feature compensation model that satisfies all constraints. It represents the minimum difference between the initial audio features and the reference audio features that can be achieved under the constraints of equal loudness curves and audio feature relationships.
[0076] Optionally, one way to determine the initial audio feature compensation value based on the target feature differences is to directly map the target feature differences to corresponding signal processing parameters. For example, loudness feature differences are directly converted into linear gain values, and spectral energy distribution feature differences are converted into gain values for each frequency band of a multi-band equalizer (EQ). For example, if the target loudness difference is +2 LUFS, then the gain compensation value is determined to be +2 dB. Another way to determine the initial audio feature compensation value based on the target feature differences is to use the target feature differences as input to an iterative optimization algorithm, and output more refined dynamic processing parameters (such as compression threshold and release time) by looking up a preset parameter mapping table or using a lightweight regression model. For example, a large envelope feature difference (indicating an excessively large dynamic range) can be mapped to a lower compressor threshold (such as -30 dB) and a higher compression ratio (such as 4:1). Yet another way to determine the initial audio feature compensation value based on the target feature differences is to use the target feature differences to initialize a fast, deterministic compensation parameter calculation process that simultaneously considers the interactions between multiple feature differences. For example, the target differences in loudness, spectrum, and envelope can be simultaneously input into a pre-trained neural network, which directly outputs a set of coordinated gain, EQ, and compression parameters. This specification does not limit the scope of the embodiments described herein.
[0077] For example, for an initial audio recording of a human voice with excessive dynamic range and insufficient mid-frequency response, the reference characteristics are: target loudness -18 LUFS, target dynamic range 10 dB, and a target spectrum boost of +3 dB at 2 kHz. After solving the model, the target feature differences are obtained: loudness needs to be increased by +1.5 LUFS, dynamic range needs to be reduced by 8 dB, and energy in the 2 kHz band needs to be increased by +4 dB. Based on this, the determined feature compensation values are: overall gain +1.5 dB, compressor parameters (threshold -20 dB, ratio 3:1), and EQ gain at 2 kHz +4 dB.
[0078] In this embodiment, a feature compensation model is constructed by using equal loudness curves as constraints and minimizing the feature difference between the reference audio features and the initial audio features, while satisfying the audio feature relationships. The feature compensation model is solved to obtain the target feature difference. Based on the target feature difference, the feature compensation value of the initial audio is determined. The equal loudness curve constraint ensures that any adjustment conforms to the basic perceptual characteristics of the human ear, avoiding auditory imbalance at different playback volumes. Furthermore, the audio feature relationship constraint forces the optimization result to meet the balance rules between audio elements in mixed contexts, preventing the problem of "loudness meets the standard but auditory confusion" caused by isolated optimization. This not only ensures accurate matching of the overall target loudness value but also achieves accurate adaptation of multi-dimensional features.
[0079] In one optional embodiment of this specification, the initial audio features include initial loudness features, initial envelope features, and initial spectral energy distribution features, and the reference audio features include reference loudness features, reference envelope features, and reference spectral energy distribution features; Step 104 includes: Based on the initial loudness characteristics, the relative loudness relationships between different types of initial audio are determined; Based on the initial envelope features, the relative envelope relationships between different types of initial audio are determined; Based on the initial spectral energy distribution characteristics, the relative spectral balance relationship between different types of initial audio is determined; Under the condition of satisfying the audio feature relationship, and with equal loudness curves as constraints, a feature compensation model is constructed with the objective of minimizing the feature difference between the reference audio features and the initial audio features. This model includes: Under the conditions of satisfying the relative loudness relationship, relative envelope relationship and relative spectral balance relationship, a feature compensation model is constructed with equal loudness curves as constraints, aiming to minimize the difference in loudness features between the initial loudness features and the reference loudness features, the difference in envelope features between the initial envelope features and the reference envelope features, and the difference in spectral energy distribution features between the initial spectral energy distribution features and the reference spectral energy distribution features.
[0080] Initial loudness features are quantized parameters that describe the perceived volume of an audio signal, either as a whole or in a localized area. For example, they can be overall loudness, short-term loudness, or instantaneous loudness.
[0081] The initial envelope features are a set of parameters describing the change of the audio signal amplitude over time, reflecting the dynamic characteristics of the signal. For example, they may include the signal's root mean square level, true peak value, dynamic range, and the attack time and release time, which describe the rise and fall of the amplitude.
[0082] The initial spectral energy distribution characteristics are a set of parameters describing how the energy of an audio signal is distributed across different frequency ranges. For example, they may include the spectral centroid, the energy value of each 1 / 3 octave, or the energy proportion of a specific frequency band (such as low frequency 20-250Hz, mid frequency 250-4kHz, and high frequency 4-20kHz).
[0083] Similarly, the reference loudness feature is the target loudness value or range to be achieved after the initial audio adjustment. For example, the reference overall loudness is -14 LUFS. The reference envelope feature is the target dynamic range to be achieved after the initial audio adjustment. For example, the reference dynamic range is 8-12 dB. The reference spectral energy distribution feature is the target spectral shape to be achieved after the initial audio adjustment. For example, the reference spectral energy distribution in the 2-4 kHz range needs to be 2 dB higher than the overall average level.
[0084] Relative loudness relationship refers to the perceived volume ratio that different initial audios should maintain in a mixed audio context. For example, in a game scenario, it is stipulated that "the loudness of key character voice lines should be 6-8 dB higher than the background ambient sound, but 2-3 dB lower than major combat sound effects."
[0085] The relative envelope relationship refers to the coordination rules that different initial audios should follow in the dynamic characteristics of a mixed audio context to prevent mutual interference. For example, it is stipulated that "the dynamic range of continuous background music should be kept large to provide a sense of breathing, while the dynamic range of transient sound effects (such as gunshots) can be appropriately compressed to enhance the impact. However, when the two overlap transiently, the background music needs to be compressed through sidechains to make room for the sound effects."
[0086] The relative spectral balance relationship refers to the rules (such as avoidance, complementarity, etc.) that should be followed in the spectral energy distribution of different initial audio in a mixed audio context to reduce the masking effect. For example, it is stipulated that "human voice dialogue should dominate in the core clarity frequency band of 1kHz-3kHz, and background music and sound effects need to be appropriately attenuated in this frequency band."
[0087] Optionally, one way to determine the relative loudness relationship between different types of initial audio based on initial loudness characteristics is to search for or calculate the corresponding loudness ratio relationship according to the audio type label and a preset priority rule library. For example, the system has built-in rules: "For audio of type dialogue, the loudness baseline value is set to 0dB; for audio of type music, the baseline value is -6dB; for audio of type sound effects, the baseline value is -3dB." Another way to determine the relative loudness relationship between different types of initial audio based on initial loudness characteristics is to analyze the current loudness feature distribution of each audio and, combined with a clustering algorithm, automatically infer a more reasonable relative loudness relationship under the current material combination. For example, if there is a prominent element in a group of audio with a significantly higher loudness than all other audio, the system can automatically set it as the baseline and use it to calculate the relative target of other elements. This specification does not limit this aspect in the embodiments.
[0088] Similarly, the implementation methods for determining the relative envelope relationship between different types of initial audio based on initial envelope features, and the implementation methods for determining the relative spectral balance relationship between different types of initial audio based on initial spectral energy distribution features, can refer to the implementation methods for determining the relative loudness relationship between different types of initial audio based on initial loudness features, which will not be elaborated here.
[0089] Loudness feature difference is a distance metric between a reference loudness feature vector and an initial loudness feature vector. For example, loudness feature difference can be the absolute difference between the target LUFS value and the current LUFS value in the logarithmic domain.
[0090] Envelope feature difference is a distance metric between the reference envelope feature vector and the initial envelope feature vector. For example, envelope feature difference can be the difference between the target dynamic range and the current dynamic range, and the weighted sum of the differences between the target root mean square (RMS) level and the current RMS level.
[0091] The difference in spectral energy distribution characteristics is a distance metric between the reference spectral energy distribution characteristic vector and the initial spectral energy distribution characteristic vector. For example, the difference in spectral energy distribution characteristics could be the root mean square error between the target energy value and the current energy value in each 1 / 3 octave band.
[0092] For example, a game audio clip is processed, containing three types: player voice (initial loudness -22 LUFS, dynamic range 15 dB, prominent mid-frequency spectrum), battle background music (initial loudness -18 LUFS, dynamic range 20 dB, uniform spectrum), and explosion sound effects (initial loudness -12 LUFS but high peak value, large dynamic range, strong low-frequency spectrum). The feature relationships determined in step 104 are: relative loudness relationship: voice > sound effects ≈ music; relative envelope relationship: when sound effects appear, music needs sidechain compression to avoid them; relative spectral balance relationship: the 1-3 kHz frequency band of voice should remain clear, sound effects should emphasize low frequencies, and music should emphasize mid-high frequencies. The model constructed in step 106 aims to: minimize the difference between voice loudness and the target -20 LUFS, the difference between music dynamic range and the target 12 dB, and the difference between the sound effect spectrum and the target low-frequency enhancement curve, while satisfying the above three relationships and conforming to the equal loudness curve constraint. After solving the model, the compensation values are: overall speech gain +2dB, music compression applied (threshold -24dB, ratio 2.5:1, sidechain triggered), sound effects boosted by +5dB at 80Hz and attenuated by -2dB at 2kHz.
[0093] In the embodiments of this specification, the relative loudness relationship between different types of initial audio is determined based on initial loudness characteristics; the relative envelope relationship between different types of initial audio is determined based on initial envelope characteristics; and the relative spectral balance relationship between different types of initial audio is determined based on initial spectral energy distribution characteristics. Under the condition of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, a feature compensation model is constructed with equal loudness curves as constraints, aiming to minimize the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the initial envelope feature and the reference envelope feature, and the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature. This model can decouple the loudness feature, envelope feature, and spectral feature of audio processing and establish collaborative rules separately, thereby constructing a more refined multi-dimensional collaborative feature compensation model that better conforms to professional mixing practices. This ensures that the final target audio not only meets the overall loudness standard but also achieves balance in dynamic contrast and spectral coordination.
[0094] In one optional embodiment of this specification, determining the feature compensation value of the initial audio based on the target feature difference includes: Based on the difference in target loudness features, the loudness feature compensation value of the initial audio is determined through a pre-constructed first mapping relationship; Based on the differences in target envelope features, the envelope feature compensation value of the initial audio is determined through a pre-constructed second mapping relationship; Based on the differences in the target spectral energy distribution characteristics, the compensation value of the spectral energy distribution characteristics of the initial audio is determined through a pre-constructed third mapping relationship.
[0095] The target loudness feature difference is the minimum quantization difference that can be achieved between the initial loudness feature and the reference loudness feature under all constraints. For example, the target loudness feature difference can be ΔL = L_ref - L_initial, where L_ref is the target overall loudness - 23 LUFS, and L_initial is the overall loudness of the initial audio - 20 LUFS, then ΔL = -3 LUFS.
[0096] The first mapping relationship is a mapping relationship (such as a function, lookup table, or rule set) between the difference in target loudness characteristics and the desired applied gain value. For example, the first mapping relationship can be a linear function Gain_dB=k×ΔL (where k is a scaling factor, such as 0.9), or a nonlinear mapping table that takes into account the equal loudness effect.
[0097] The loudness feature compensation value is a gain parameter calculated based on the first mapping relationship, used to adjust the initial audio to achieve the target loudness. For example, based on the above linear function and ΔL=-3LUFS, the overall gain with a loudness feature compensation value of -2.7dB is calculated.
[0098] The target envelope feature difference is the minimum quantization difference that can be achieved between the initial envelope characteristics (such as dynamic range, RMS level, etc.) and the reference envelope characteristics, while satisfying all constraints. For example, the target envelope feature difference may include a reduction in dynamic range of ΔDR = 5dB and an increase in average RMS level of ΔRMS = 2dB.
[0099] The second mapping relationship is a mapping relationship (such as a function, lookup table, or rule set) that characterizes the difference in target envelope features and the compression / limiting parameters. For example, the second mapping relationship could be: for cases where the dynamic range needs to be reduced by ΔDR, it can be mapped to a set of compressor parameter suggestions, such as Threshold = -20 - 0.5 × ΔDR (dBFS) and Ratio = 1 + 0.2 × ΔDR.
[0100] Envelope feature compensation values are a set of dynamic processor parameters calculated based on the second mapping relationship, used to adjust the initial audio dynamic characteristics. For example, based on the above rules and ΔDR=5dB, the calculated envelope feature compensation values include: compressor threshold = -22.5dBFS, compression ratio = 2:1, and may include suggested start / release times.
[0101] The target spectral energy distribution feature difference is the optimal result obtained after solving the feature compensation model. It represents the minimum quantization difference vector that can be achieved between the energy distribution of the initial audio in each frequency band and the energy distribution feature of the reference spectrum under all constraints. For example, the target spectral energy distribution feature difference can be a vector ΔE=[ΔE_1, ΔE_2, ..., ΔE_n], where ΔE_i represents the difference (in dB) between the target energy and the initial energy in the i-th frequency band (e.g., 1 / 3 octave band).
[0102] The third mapping relationship is a mapping relationship (such as a function, lookup table, or rule set) between the differences in the target spectral energy distribution characteristics and the required gain values of the multi-band equalizer in each corresponding frequency band. For example, the third mapping relationship can be a gain allocation matrix G=M×ΔE, where M is a matrix that considers the smooth transition between adjacent frequency bands and perceptual weighting; the third mapping relationship can also be a set of rules adjusted based on an auditory masking model.
[0103] The spectral energy distribution characteristic compensation value is a set of equalizer parameters used to adjust the initial audio spectrum shape, calculated based on the third mapping relationship. It is typically represented by the center frequency, bandwidth, and gain value of each frequency band. For example, based on the aforementioned vector and mapping relationship, the calculated spectral energy distribution characteristic compensation value could include: a +2dB boost at a center frequency of 100Hz, a -1dB attenuation at a center frequency of 1kHz, and a +3dB boost at a center frequency of 8kHz.
[0104] Optionally, one way to construct the first mapping relationship is to establish an empirical correspondence table of required gains for different loudness differences based on statistical analysis of a large number of audio samples. For example, through experimental measurement, the actual gain value applied when adjusting audio at different loudnesses to the same target loudness is recorded, forming data pairs (ΔL, Gain), and a mapping curve is fitted or a lookup table is created. Another way to construct the first mapping relationship is to derive a theoretical calculation model of the relationship between perceived loudness differences and physical gain, considering equal loudness curves, based on a psychoacoustic model. For example, by combining equal loudness curves at a specific playback level, the gain required for each frequency band to achieve a change of perceived loudness ΔL is calculated, and then synthesized into an overall gain mapping. The embodiments in this specification do not limit this approach.
[0105] Optionally, one way to construct the second mapping relationship is to collect cases of different dynamic processing requirements (such as "tightening dynamics" and "boosting average level") in professional mixing projects and their corresponding compressor / limiter parameter settings, and train a predictive model from feature differences to processing parameters through machine learning (such as regression analysis). For example, a deep neural network can be used to learn a nonlinear mapping from (ΔDR, ΔRMS) to (threshold, ratio, start time, release time). Another way to construct the second mapping relationship is to establish a parameterized rule engine based on a theoretical model of dynamic processing (such as compression curve equation) and auditory dynamic perception characteristics. For example, several preset sets of "light compression", "medium compression", and "heavy compression" parameter templates can be selected according to the magnitude of ΔDR, and then the threshold can be fine-tuned according to ΔRMS. The embodiments in this specification do not limit this approach.
[0106] Optionally, one way to construct the third mapping relationship is to train a neural network model using a large amount of professionally processed audio spectrum comparison data before and after equalization, enabling it to directly predict the optimal equalizer parameters based on the spectral energy difference vector. For example, a convolutional neural network can be used to learn the mapping from the spectral difference map to multi-band EQ gain parameters. Another way to construct the third mapping relationship is to design a regularized mapping algorithm based on perceptual audio coding principles and auditory filters. For example, the target spectral energy distribution difference can be input into a filter bank that simulates the critical frequency band analysis of human hearing, and then the gain of the corresponding EQ band can be determined according to the difference value of each critical frequency band through a preset gain calculation formula. This specification does not limit the embodiments in this way.
[0107] For example, consider an initial human voice audio clip with a target loudness feature difference of +2.5 LUFS, a target envelope feature difference of a 4 dB reduction in dynamic range and a 1 dB increase in average RMS, and a target spectral energy distribution feature difference of -2 dB attenuation at 250 Hz, +3 dB increase at 2 kHz, and +1 dB increase at 8 kHz. Applying the first mapping relationship (linear coefficient k=0.95), the overall gain of the loudness feature compensation is determined to be +2.375 dB. Applying the second mapping relationship (lookup table), the envelope feature compensation is determined to be the compressor parameters: threshold -22 dBFS, ratio 2.8:1, start-up time 10 ms, and release time 150 ms. By applying the third mapping relationship (gain allocation matrix), the spectral energy distribution characteristic compensation values are determined to be the parameters of the three equalizer bands: attenuation of -2dB at 250Hz (Q=2.0), boost of +2.8dB at 2kHz (Q=1.4), and boost of +1dB at 8kHz (Q=1.0). The system applies these three sets of compensation values to the initial audio, thereby synergistically achieving loudness normalization, dynamic control, and spectral shaping.
[0108] In the embodiments of this specification, the differences in target loudness characteristics, target envelope characteristics, and target spectral energy distribution characteristics are synchronously converted into gain, dynamic processing, and equalization compensation parameters that can directly drive digital signal processing through pre-constructed first, second, and third mapping relationships. This not only normalizes the loudness to the target value but also controls dynamic fluctuations through envelope compensation to improve clarity and impact. Furthermore, it optimizes timbre balance through spectral compensation to reduce masking and enhance details. As a result, the processed audio not only meets the loudness standard but also improves the listening experience, thereby solving the problem of "flattening" of the listening experience caused by simply adjusting the loudness gain.
[0109] In one optional embodiment of this specification, after determining the spectral energy distribution feature compensation value of the initial audio based on the difference in target spectral energy distribution features through a pre-constructed third mapping relationship, the method further includes: Receive adjustment instructions from the front end for compensation values based on spectral energy distribution characteristics; In response to the adjustment command, the spectral energy distribution characteristic compensation value is adjusted to obtain the adjusted spectral energy distribution characteristic compensation value.
[0110] Adjustment commands are user-issued commands on the front-end interface to modify the compensation values for spectral energy distribution characteristics. For example, an adjustment command could be a user dragging the slider representing the "2kHz" frequency point down from the system-suggested "+3dB" to "+1dB" on a graphical equalizer interface.
[0111] For example, in a movie sound mixing scene, the system automatically calculates a spectral energy distribution characteristic compensation value for an ambient sound audio of "howling wind," suggesting a -4dB attenuation in the "300-600Hz" frequency band to reduce the masking of dialogue. However, if the user wants to preserve the low frequencies of the wind sound to enhance the atmosphere, they can drag the attenuation slider for that frequency band back from "-4dB" to "-1dB" on the front-end EQ interface. This updates the original compensation value to "attenuate by -1dB in the 300-600Hz frequency band," and applies this adjusted compensation value to process the audio. This ensures that the wind sound neither completely masks the dialogue nor fails to retain the low-frequency presence required for the artistic intent.
[0112] In this embodiment, by receiving an adjustment command sent by the front end for the spectral energy distribution characteristic compensation value, and responding to the adjustment command, the spectral energy distribution characteristic compensation value is adjusted to obtain the adjusted spectral energy distribution characteristic compensation value. This avoids over-clarification of certain sounds that would otherwise require artistic masking, thus preserving the efficiency and scientific nature of automated processing while giving the user ultimate artistic control. This allows the user to flexibly choose to automatically process such sounds or manually adjust the relevant parameters to achieve a satisfactory effect.
[0113] In one optional embodiment of this specification, the initial audio features further include at least one of initial spatial features and initial rendering features, and the reference audio features further include at least one of reference spatial features and reference rendering features; Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objective of minimizing the difference in loudness features between the initial loudness feature and the reference loudness feature, the difference in envelope features between the initial envelope feature and the reference envelope feature, and the difference in spectral energy distribution features between the initial spectral energy distribution feature and the reference spectral energy distribution feature. The model includes at least one of the following: Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the difference in loudness features between the initial loudness feature and the reference loudness feature, the difference in envelope features between the initial envelope feature and the reference envelope feature, the difference in spectral energy distribution features between the initial spectral energy distribution feature and the reference spectral energy distribution feature, and the difference in spatial features between the initial spatial feature and the reference spatial feature. Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the difference in loudness features between the initial loudness features and the reference loudness features, the difference in envelope features between the initial envelope features and the reference envelope features, the difference in spectral energy distribution features between the initial spectral energy distribution features and the reference spectral energy distribution features, and the difference in rendering features between the initial rendering features and the reference rendering features. Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the envelope feature and the reference envelope feature, the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature, the spatial feature difference between the initial spatial feature and the reference spatial feature, and the rendering feature difference between the initial rendering feature and the reference rendering feature.
[0114] The initial spatial features are a set of parameters describing the spatial properties of audio signals (such as sound source localization, sound field width, and environmental perception). For example, they may include inter-channel correlation, inter-channel intensity difference, inter-channel time difference of stereo or multi-channel signals, as well as apparent sound source width, reverberation time, and the ratio of early reflections to late reverberation energy obtained based on reverberation analysis.
[0115] Initial rendering features are a set of parameters describing the "coloring" characteristics of an audio signal due to specific encoding, processing history, or artistic effects. They reflect the nonlinear distortion, frequency response modification, or signature of dynamic processing of the signal. For example, they may include harmonic distortion features (such as total harmonic distortion, odd-even harmonic ratio), residual features of specific equalization curves (such as the bandpass feature of a "telephone tone"), modulation characteristics introduced by compression / limiting processing, etc.
[0116] The reference spatial characteristics are the target spatial attributes that are expected to be achieved after the initial audio adjustment. For example, they may include the target sound image position (azimuth, elevation), the target stereo width value, the target room reverberation type (such as "small room", "hall"), etc.
[0117] Reference rendering features are the target coloring characteristics that are achieved or retained after the initial audio adjustments. For example, these may include the target harmonic warmth level, the characteristics of the target frequency response curve (such as "retro analog feel" or "modern clarity"), or markers indicating whether the original specific encoding / processing effects should be retained.
[0118] For example, examples of building feature compensation models based on the above three items are as follows: Example 1 (Adding Initial Spatial Features): When processing a 5.1 channel movie clip containing dialogue (centered) and ambient sound (surround), the initial spatial features show insufficient surround sound. The reference spatial features require a stronger sense of immersion in the ambient sound. During model building, in addition to the basic loudness, dynamics, and spectral objectives, an objective term is added to minimize the difference between the initial spatial features of the ambient sound (e.g., the energy proportion of the surround channels) and the reference spatial features. After solving, the model can output a compensation value that increases the ambient sound level in the surround channels.
[0119] Example 2 (Adding Initial Rendering Features): When processing a vocal audio track from an old recording studio with noticeable "tape saturation" coloration, the initial rendering features detect abundant even harmonics. The reference rendering features require "moderate preservation of retro warmth." During model building, in addition to the basic objective, an objective term is added to minimize the difference between the initial rendering features (current harmonic structure) and the reference rendering features (target harmonic structure, such as reducing the proportion of odd harmonics). After solving, the model can output a set of equalization and dynamic compensation parameters that improve clarity while preserving specific harmonic colorations.
[0120] Example 3 (simultaneously adding initial spatial features and initial rendering features): When processing skill sound effects in a game that need to be adapted to different output devices (such as TVs, mobile phones, VR headsets), the initial spatial feature is wide stereo, and the initial rendering feature is high-fidelity uncompressed. The reference spatial feature can be mono (mobile phone) or 3D spatial audio (VR) depending on the target device, and the reference rendering feature can be slightly lossy encoded coloration (streaming media). During model construction, both spatial feature differences and rendering feature differences need to be minimized simultaneously. After solving, the model can output different channel downmixing parameters and preprocessing filtering parameters for different target devices to achieve optimal adaptation.
[0121] In the embodiments of this specification, by incorporating the initial spatial features and / or initial rendering features, and their corresponding reference spatial features and / or reference rendering features, into the construction of the feature compensation model, not only can basic technical loudness standardization and balance be achieved, but also the unified shaping of spatial sense (such as adapting to different playback systems) and the matching or protection of artistic style can be further achieved, thereby further improving the audio listening experience.
[0122] In one optional embodiment of this specification, determining the feature compensation value of the initial audio based on the target feature difference includes: Based on the difference in target loudness features, the loudness feature compensation value of the initial audio is determined through a pre-constructed first mapping relationship; Based on the differences in target envelope features, the envelope feature compensation value of the initial audio is determined through a pre-constructed second mapping relationship; Based on the differences in the target spectral energy distribution characteristics, the compensation value of the spectral energy distribution characteristics of the initial audio is determined through a pre-constructed third mapping relationship; Based on the differences in target spatial features, the spatial feature compensation value of the initial audio is determined through a pre-constructed fourth mapping relationship, and / or, based on the differences in target rendering features, the rendering feature compensation value of the initial audio is determined through a pre-constructed fifth mapping relationship.
[0123] The meanings of the first, second, and third mapping relationships have been introduced above and will not be repeated here.
[0124] The target spatial feature difference is the minimum quantization difference that can be achieved between the spatial properties of the initial audio and the reference spatial features, under all constraints. For example, the target spatial feature difference may include the amount of sound image position offset (e.g., the azimuth angle needs to be adjusted 15 degrees to the left), the amount of increase in sound field width (e.g., the stereo width index needs to be increased from 0.3 to 0.6), or the difference in reverberation characteristics (e.g., the reverberation time RT60 needs to be reduced by 0.2 seconds).
[0125] The fourth mapping relationship is a mapping relationship (such as a function, lookup table, or rule set) between the differences in target spatial features and spatial audio processing parameters (such as the pan controller, upmix / downmix matrix, and reverb parameters). For example, the fourth mapping relationship could be: for differences in the azimuth angle of the sound image position, mapping to the adjustment coefficient of the stereo or multichannel gain matrix; for differences in sound field width, mapping to the gain adjustment amount of the lateral signal in the correlation processing.
[0126] Spatial feature compensation values are a set of spatial audio processing parameters calculated based on the fourth mapping relationship, used to adjust the initial audio spatial properties. For example, based on the above differences and mapping relationship, the calculated spatial feature compensation values include: left channel gain adjustment +1.5dB, right channel gain adjustment -1.5dB.
[0127] The target rendering feature difference is the minimum quantization difference that can be achieved between the initial audio coloration characteristics and the reference rendering features, while satisfying all constraints. For example, the target rendering feature difference may include the amount that the harmonic distortion structure needs to be adjusted (e.g., the third harmonic energy needs to be reduced by 2dB), the amount that the specific frequency response "coloration" needs to be changed (e.g., the mid-frequency bulge of the "telephone effect" needs to be attenuated by 4dB), or the amount that the dynamic "coloration" needs to be adapted (e.g., the "sucking" effect caused by compression needs to be reduced).
[0128] The fifth mapping relationship is a mapping (such as a function, lookup table, or rule set) between the differences in target rendering features and the audio processing parameters (such as harmonic exciter parameters, specific equalization curves, and dynamic processor modulation parameters) used to achieve a specific coloring. For example, the fifth mapping relationship could be: for differences that need to increase the "warmth of the tape" (characterized by even harmonics), it could be mapped to the even harmonic generation intensity parameters of the harmonic exciter in the low and mid frequencies; for differences that need to reduce "coding artifacts," it could be mapped to the coefficients of a specific artifact removal linear or nonlinear filter.
[0129] The rendering feature compensation value is a set of audio parameters used to adjust the initial audio coloring characteristics, calculated according to the fifth mapping relationship. For example, the rendering feature compensation value calculated based on the above differences and mapping relationships may include: applying +2% even harmonic distortion to the harmonic exciter at 100Hz; and a notch equalizer parameter with -4dB attenuation at 3.5kHz and a Q value of 1.5.
[0130] Optionally, one way to construct the fourth mapping relationship is to establish an analytical calculation model based on the acoustic and psychoacoustic principles of spatial audio (such as the Head Related Transfer Function (HRTF)) to map spatial attribute differences to multi-channel gain or delay parameters. For example, based on the differences in the target sound image position (azimuth, elevation), the gain coefficients of each channel are calculated using an HRTF database or a simplified model to form a mapping relationship. Another way to construct the fourth mapping relationship is to collect a large number of spatial processing settings used in professional mixing projects to achieve specific spatial effects (such as "widening", "backing up", "immersion"), and use data mining or machine learning methods to summarize the statistical laws or predictive models from spatial feature differences to processing parameters. For example, a regression tree model can be used to learn the mapping from (ΔWidth, ΔDistance) to "reverberation delivery" and "early reflection ratio". The embodiments in this specification do not limit this approach.
[0131] Optionally, one way to construct the fifth mapping relationship is to perform parameter scanning on representative coloring effect processors (such as tape simulators, tube simulators, and specific equalizers) and extract the rendering feature changes of their input and output audio, thereby establishing a corresponding database from the processor parameter space to the rendering feature difference space, and then fitting or interpolating the inverse mapping relationship (from difference to parameter). For example, systematically adjusting the driving parameters of a "saturator", measuring the changes in the harmonic structure of the output audio, and establishing a mapping table between the driving parameter values and the harmonic difference vector for inverse lookup. Another way to construct the fifth mapping relationship is to use signal processing and auditory perception models to mathematically model the generation mechanism of a specific coloring effect, thereby deriving the signal processing transformation parameters required to achieve the target rendering feature difference. For example, approximating the distortion characteristics of a specific amplifier based on a memory polynomial nonlinear model, and solving for the input signal preprocessing parameters required to obtain the target harmonic difference through model inverse deduction. The embodiments in this specification do not limit this approach.
[0132] Optionally, one way to determine the initial audio spatial feature compensation value based on the target spatial feature difference through a pre-constructed fourth mapping relationship is to input the target spatial feature difference vector into a spatial rendering engine based on the Vector Base Amplitude Panning (VBAP) algorithm. This engine calculates the updated gain coefficients of each channel in real time according to the mapping relationship, which are then used as compensation values. For example, for a difference that requires moving the sound source from the center to the right front by 30 degrees, the mapping relationship calls the VBAP algorithm to output the gain adjustment coefficients of the right front, right surround, and other channels. Another way to determine the initial audio spatial feature compensation value based on the target spatial feature difference through a pre-constructed fourth mapping relationship is to decompose the target spatial feature difference into two parts: reverberation attributes and direct sound attributes. Specific reverberator parameters (such as room size and decay time) and stereo processing parameters (such as center-side balance adjustment) are obtained by querying a pre-set "reverberation parameter library" and "sound image / width parameter mapping table," respectively, and then combined to form the spatial feature compensation value. For example, to address the difference in "increasing the sense of distance and space," the mapping relationship provides a set of longer reverberation pre-delay and decay time parameters, as well as a set of gain coefficients that slightly reduce the proportion of direct sound. This specification does not limit the embodiments described herein.
[0133] For example, for an initial stereo music clip, the target spatial feature difference is that the sound field needs to be significantly widened (width exponent difference ΔWidth = +0.4) and a slight increase in "hall" reverberation is needed (reverberation time difference ΔRT60 = +0.5s). Applying the fourth mapping relationship (by looking up the width control table based on mid-side processing), the width compensation value is determined to be: a +6dB increase in side signal gain. Simultaneously, by querying the reverberation parameter mapping table, the reverberation compensation value is determined to be: loading a "medium-sized hall" reverberation algorithm and setting the decay time to 1.8 seconds. Another initial narration audio clip, with slight low-quality MP3 encoding artifacts (manifested as harsh high-frequency quantization noise), has a target rendering feature difference of eliminating these artifacts. Applying the fifth mapping relationship (based on the artifact removal filter model), the rendering feature compensation value is determined to be: applying a low-pass filter with a smooth roll-off (-6dB / oct) above 12kHz, and superimposing a slight noise shaping process.
[0134] In the embodiments of this specification, by introducing a fourth mapping relationship and a fifth mapping relationship, the target spatial feature differences and / or target rendering feature differences solved by the feature compensation model are mapped into executable spatial audio processing parameters and sound effect processing parameters. This enables automated and refined adjustment of audio spatial attributes (such as positioning, width, and environment) and artistic coloring characteristics (such as style, texture, and device signature). This not only expands the processing capabilities for advanced mixing tasks but also ensures a high degree of consistency between the adjustment results in terms of artistry and technicality. Furthermore, it improves the coordination of various types of initial audio after adjustment when integrated into the final multi-audio context.
[0135] In one optional embodiment of this specification, after adjusting the initial audio features based on feature compensation values to obtain multiple types of target audio, the method further includes: Feed the target audio back to the front end; Receive optimization instructions sent from the front end for the target audio; In response to the optimization command, the reference audio features are adjusted to obtain the adjusted reference audio features; Return to the execution steps that use equal loudness curves as constraints and construct feature compensation models based on reference audio features and audio feature relationships, until optimized target audio of various types is obtained.
[0136] The front end is the client application or user interface that interacts with the user, displays processing results, and receives user input. For example, the front end could be a plugin interface for a digital audio workstation or the interactive screen of a mobile application.
[0137] Optimization commands are actions issued by users after listening to or evaluating target audio to adjust the final effect of the initial audio. These commands specify how the characteristics of the reference audio should be adjusted. For example, an optimization command might be for the user to drag the "warmth" slider to increase it by 10%.
[0138] Optionally, one way to feed the target audio back to the front end is to send the processed target audio data to the front-end application for real-time preview or download via an audio streaming protocol or file transfer method. Another way to feed the target audio back to the front end is to directly embed an audio player control in the front-end interface and render and display the waveform, spectrum, and other visual information of the target audio along with the audio data. This specification does not limit this approach in the embodiments.
[0139] Optionally, one implementation of receiving optimization instructions for the target audio sent by the front end can be: capturing user operations on interface controls (such as sliders, buttons, and drawing areas) through an event listener mechanism in the graphical user interface, and converting them into numerical adjustments to specific parameters in the reference audio features. Another implementation of receiving optimization instructions for the target audio sent by the front end can be: receiving descriptive feedback input by the user in natural language or structured text, parsing the user's intent through a natural language processing model, and mapping it to adjustments to the reference audio features. This specification does not limit the scope of this implementation.
[0140] For example, after a user processes a podcast audio clip (containing vocals and background music), the front-end plays the resulting target audio. The user feels the background music is too loud, so they drag the "Music Loudness" slider to the left (decrease it) on the front-end interface. The platform receives this optimization instruction and adjusts the reference loudness feature value of "Background Music" in the reference audio features from -16 LUFS to -18 LUFS. Subsequently, based on the adjusted reference audio features and previously determined audio feature relationships (such as the relative loudness relationship between vocals and music), the platform reconstructs and solves the feature compensation model to obtain new feature compensation values and generate new target audio. The user can repeat this process until a satisfactory "optimized target audio" is obtained.
[0141] This embodiment introduces iterative optimization based on user feedback, allowing users to make minor adjustments to the overall satisfactory automated initial version, thereby improving user satisfaction while ensuring processing efficiency.
[0142] In one optional embodiment of this specification, there are multiple reference audio features, which are extracted from multiple reference audios. The target audio of any type consists of multiple features. After adjusting the initial audio features based on feature compensation values to obtain various types of target audio, the process also includes: Feedback of various types of target audio to the front end; Receive evaluation instructions for the target audio sent from the front end; In response to the evaluation command, multiple reference audio features are filtered to obtain the filtered reference audio features.
[0143] Any type of target audio is a set of multiple audio files obtained by processing any type of initial audio file based on the features of each reference audio file using the audio processing method described above. In other words, for any type of audio file, the number of reference audio files is the same as the number of target audio files. For example, for an initial human voice audio file, if the system provides three reference audio files of different styles, the system will process the human voice based on the features of these three sets of reference audio files respectively, generating three target audio files of different styles for the user to choose from.
[0144] Evaluation instructions are commands issued by users to rate multiple sets of target audio generated based on different reference audio after listening to them. For example, an evaluation instruction could be a user clicking to select "I like style A the most".
[0145] The process of filtering multiple reference audio features involves selecting one or more features with the highest user preference from the currently used reference audio features, based on the evaluation instructions, as the reference standard for subsequent iterations or the final determination.
[0146] Optionally, one way to filter multiple reference audio features to obtain the filtered reference audio features is to directly determine the selected reference audio features based on the user's explicit single-selection, multiple-selection, or sorting operation. For example, if the user checks the checkboxes corresponding to "Style B" and "Style C" on the front-end interface, the system will filter out these two reference audio features. Another way to filter multiple reference audio features to obtain the filtered reference audio features is to perform statistical analysis based on user feedback (such as the playback duration and number of repetitions of each target audio), infer user preferences, and then filter out the reference audio features that best meet the user's expectations. This specification does not limit this approach in the embodiments.
[0147] For example, a user has initial audio of a game character's voice acting. The system has built-in reference audio features for five different character types and generates five target audio tracks with different styles based on these features. The front end displays these five audio tracks and their style tags in a list format. After listening to them, the user selects one of the target audio tracks by clicking the "Favorite" button. Upon receiving this evaluation instruction, the system filters out the reference audio features corresponding to that target audio track, which serve as the final basis for subsequent processing.
[0148] In this embodiment, various target audio samples of different types are fed back to the front end; evaluation instructions for the target audio samples are received from the front end; in response to the evaluation instructions, multiple reference audio features are filtered to obtain filtered reference audio features. This allows users to listen to and select from multiple preset or exemplary reference audio samples and features when they are unsure how to configure the specific audio features of the reference audio. Through user evaluation feedback, the reference standard that best matches the user's current project or subjective preferences is selected, thereby lowering the user's learning curve and improving the match between the automated processing results and user expectations.
[0149] In one optional embodiment of this specification, obtaining reference audio features includes at least one of the following: Receive configuration instructions for reference audio features sent by the front end, wherein the configuration instructions carry the reference audio features; Obtain multiple reference audio files; extract the audio features from the multiple reference audio files as multi-reference audio features.
[0150] Configuration instructions are reference audio feature data containing specific parameter values specified by the user or upstream system.
[0151] For example, obtaining reference audio features can be achieved in two modes: 1. Reference Audio Benchmark Mode: The system performs the multi-dimensional feature analysis described above on the user-specified reference audio, and uses the analysis results as the target feature benchmark for the audio to be processed. For example, if the user does not know the features of the reference audio, they can upload a reference audio, and the system will automatically extract its loudness features, spectral energy distribution features, envelope features, etc., to form a set of reference audio features to guide the processing of the initial audio.
[0152] 2. Custom Parameter Mode: Supports users to directly set complex combinations of target parameters such as target loudness standards, spectral balance curves, dynamic range thresholds, and spatial width parameters to form a personalized processing benchmark. For example, if the user knows the specific reference audio characteristics, they can directly set the target loudness to -24 LUFS in the reference audio characteristic configuration interface, draw a target equalization curve, and set the vocal compression ratio to 2.5:1. These parameters are sent to the system via configuration commands as the reference audio characteristics for this processing.
[0153] In the embodiments of this specification, two methods are provided for users to obtain reference audio features. For professional users, a configuration interface for reference audio features is provided to meet their needs for precise control. For non-professional users or when the target is difficult to describe with parameters, a method for extracting reference audio features based on input reference audio is provided, which reduces the difficulty of use and adapts to the needs of users with different knowledge backgrounds and operating habits.
[0154] For example, Figure 3 This is a schematic diagram illustrating the execution flow of an audio processing method provided in one embodiment of this specification within a software application. (Reference) Figure 3 The above audio processing method can be implemented as a software application, and its specific implementation process is as follows: Step 1, File Input and Reference Settings: 1) Users can select the audio files or folders to be processed through the file import interface of the software (batch import is supported). The system will automatically recognize the file format and channel configuration and display the import list (including file name, format, duration, number of channels, etc.).
[0155] 2) Optional operation: The user imports one or more reference audio files through the reference audio setting interface. The system prompts "Do you want to use the reference audio as the processing benchmark?" After confirmation, the reference audio is included in the analysis queue. The multidimensional audio features of the audio to be processed are compared with the multidimensional audio features of the reference audio to generate a feature difference matrix. This allows the attention allocation and resource scheduling for difference perception to be performed based on the feature difference matrix during the automatic analysis of multidimensional features in step 2. If no reference audio is needed, the user can directly proceed to the target parameter configuration (step 3) step.
[0156] Step 2: Automatic analysis of multidimensional features: 1) The system starts a background analysis task to perform four-dimensional feature extraction and quantization on all imported audio files to be processed and reference audio files (if any): Dynamic feature analysis: ADSR parameters are extracted using the envelope detection algorithm, the dynamic range (the difference between the peak value and the effective value) is calculated, and a dynamic feature curve is generated.
[0157] Spectral feature analysis: A spectrum diagram was constructed using Fast Fourier Transform, and the energy proportion of each frequency band was statistically analyzed. The energy distribution data of the frequency band sensitive to human ears (2kHz-5kHz) was marked in particular.
[0158] Spatial feature analysis: For stereo / surround sound audio, the sound field width is calculated through inter-channel correlation analysis, the perceived depth of field is estimated based on reverberation decay time, and a spatial parameter matrix is output.
[0159] Rendering Feature Analysis: Through effects feature recognition algorithms, the presence of effects such as EQ, distortion, and reverb is detected, effect parameter values are quantified, and rendering feature differences are formed.
[0160] 2) If a reference audio exists, the system compares the multidimensional audio features of the audio to be processed with the multidimensional audio features of the reference audio, generates a feature difference matrix, and clarifies the deviation values of each dimension (such as loudness feature difference, spectral energy deviation, spatial width deviation, etc.).
[0161] Step 3, Target Parameter Configuration: 1) Loudness standard selection: Users can select the target standard from the system's preset mainstream standards (EBU R128, ATSC A / 85, ITU-RBS.1770, etc.), and the system will automatically load the loudness threshold parameters of the corresponding standard; users can also input the target loudness value (such as -16 LUFS) through custom options.
[0162] 2) Feature parameter adjustment: The system defaults to using reference audio features (if available) or industry-optimal parameters as the target feature benchmark. Users can fine-tune parameters such as target spectrum balance (e.g., high-frequency gain ±2dB), dynamic range (e.g., 8-12dB), and spatial width (e.g., 0-100%) through the parameter configuration interface, and preview the adjustment results in real time.
[0163] 3) Batch processing rule settings: Users set the output folder path, choose whether to retain the original file directory structure, enable "resource optimization mode" (automatically adjust resource usage), and confirm the batch processing start conditions.
[0164] Step 4: Generate adaptive processing scheme: 1) The system receives the audio feature data to be processed, target parameters (including loudness standards and feature parameters), and feature difference matrix (if any), and inputs them into the adaptive processing engine. The feature difference matrix acts as a structured difference navigation map and intelligent optimization scheduler in the feature compensation model. Guided by the feature difference matrix, the feature compensation model allocates attention and schedules resources based on difference perception—assigning higher optimization weights and more lenient adjustment boundaries to feature dimensions and frequency bands with significant differences, while applying protective constraints to areas with minor differences to prevent overprocessing. Simultaneously, the inter-feature difference correlations contained in the feature difference matrix (such as the synchronous manifestation of loudness and spectral differences in specific frequency bands) enable the feature compensation model to identify the root causes, thereby formulating a collaborative rather than isolated compensation strategy. This ensures that while accurately achieving the overall loudness standard, it simultaneously optimizes the audio's spectral balance, dynamic appropriateness, and mixing context harmony, significantly improving the clarity, layering, and artistic integration of the output audio in subjective listening perception.
[0165] 2) The engine combines a human hearing model (incorporating equal loudness curves and frequency sensitivity data) for comprehensive calculations to generate personalized processing solutions: The overall gain adjustment is calculated based on the differences in loudness characteristics to ensure that the final loudness meets the target standard.
[0166] Based on the differences in spectral energy distribution characteristics, frequency band gain compensation parameters are generated (e.g., a 1dB deviation in the frequency band sensitive to human hearing corresponds to a compensation of 0.8dB).
[0167] Based on the results of dynamic feature analysis, compression thresholds, ratios, and attack / release times are set to avoid over-compression.
[0168] Based on spatial characteristic differences, level adjustment coefficients and reverberation parameter compensation values are generated for each channel.
[0169] Based on the differences in rendering features, a color protection strategy is formulated (such as preserving the original distortion and adjusting only loudness-related parameters).
[0170] 3) After the processing solution is generated, the system allows users to preview the processing effect of a single file (compare the original audio with the processed audio). Users can confirm the solution or return to the parameter adjustment stage to regenerate it.
[0171] Step 5: Processing and Monitoring 1) After the user confirms the processing plan, the batch processing task is started. The system allocates computing resources through the resource scheduling module, executes audio processing in sequence, and displays the processing progress bar and resource usage status (Central Processing Unit (CPU) and memory usage) in real time.
[0172] 2) During the process, users can perform operations such as pause, continue, and cancel through the monitoring interface and feed the monitoring data back to the batch management module; if the resource usage is close to the threshold (such as CPU utilization reaching 90%), the system will automatically start the load balancing mechanism to reduce the single task processing rate and avoid device overload.
[0173] Step 6: Results Output and Report Generation 1) After a single file is processed, the processed audio is automatically output to a preset folder, retaining the original file name and extension; after all files are processed, the system generates a batch processing report (Excel / PDF format report), which includes processing details of each file and overall statistics (such as pass rate and average processing time).
[0174] Optionally, users can preview the processed audio using the system's preview function. If the listening experience does not meet expectations (e.g., insufficient spatiality, spectral imbalance), they can select the corresponding file, return to the parameter configuration section to adjust the target parameters, regenerate the processing plan, and execute the processing. Secondary optimization does not require re-importing files; the system retains the original audio data and initial analysis results, only updating the processing strategy and output files, thus improving optimization efficiency.
[0175] For example, Figure 4 This is a schematic diagram of the architecture of an audio processing system provided in one embodiment of this specification. (Reference) Figure 4 The system includes an audio processing platform, a multi-dimensional audio feature analysis engine, an adaptive processing engine, and a batch and engineering management module, the details of which are as follows: 1) Audio Processing Platform: This platform serves as the sole core platform for audio file input, analysis, processing, and output, achieving fully integrated operation. It does not require specific DAW software, enhancing its versatility and flexibility. The platform's core interfaces include: (1) File import / export interface: Supports batch import of single files, multi-track files, complete scene recordings / video clips and folders, and is compatible with mainstream audio formats such as WAV, MP3 and FLAC; output files automatically retain the original format and naming conventions.
[0176] (2) Reference audio setting interface: allows users to specify one or more "reference audio files / segments" as the source of the processing target benchmark.
[0177] (3) Target parameter configuration interface: It supports users to select preset mainstream loudness standards or customize loudness target values, and can finely adjust characteristic parameters such as target spectrum balance, dynamic range, and spatial width.
[0178] (4) Processing monitoring and preview interface: Real-time display of processing progress and computer resource usage status (such as CPU and memory usage), supports real-time preview of processing effect, and facilitates users to adjust parameters in a timely manner.
[0179] (5) Cross-platform underlying support layer: enables the core processing program to run independently without relying on any specific DAW software or operating system, with strong cross-platform adaptability; supports mainstream audio formats and multi-channel configurations, and is compatible with the input and output needs of various industrial audio production scenarios, improving the versatility and application scope of the solution.
[0180] 2) Multi-dimensional audio feature analysis engine, including: Automated audio analysis and scheduling unit: Performs fully automated background analysis on imported audio files (including single files, multi-track files, and batch files) and reference audio, without manual intervention.
[0181] Envelope feature extraction unit: accurately analyzes the envelope shape of audio during attack, decay, sustain, and release, calculates dynamic range values, and quantifies the dynamic variation patterns of audio intensity.
[0182] Spectral feature extraction unit: Constructs an audio spectrum energy distribution map, focuses on analyzing the energy distribution of the human ear sensitive frequency band (such as 2kHz-5kHz) and its relative relationship with other frequency bands, and identifies the location of spectral peaks and valleys.
[0183] Spatial feature extraction unit: For stereo / surround sound audio, it evaluates perceived depth of field (distance perception), sound field width (spatial extension range), and identifies spatial parameters such as sound image positioning coordinates and reverberation decay time.
[0184] Rendering Feature Extraction Unit: Identifies whether the audio contains strong EQ, distortion, reverb, delay and other rendering effects, quantifies the feature parameters of these effects (such as reverb wet-dry ratio, distortion, EQ gain), and defines the sound "color" attribute.
[0185] Feature quantization and encoding unit: used to quantize and encode rendering features, spatial features, spectral features, and envelope features.
[0186] Reference feature comparison unit: If a reference audio exists, the system compares the multidimensional features of the audio to be processed with the features of the reference audio, generates a feature difference matrix, and clarifies the deviation values of each dimension (such as loudness deviation, spectral energy deviation, spatial width deviation, etc.).
[0187] 3) Adaptive processing engine, including: The Human Hearing Model Integration Unit combines multi-dimensional feature analysis results with the human hearing model (emphasizing frequency response sensitivity characteristics and equal-loudness curve patterns) to generate an adaptive processing scheme that goes beyond simple gain adjustment. This not only ensures accurate matching of the overall target loudness value but also achieves precise adaptation of multi-dimensional features through dynamic adjustment of the processing strategy.
[0188] The processing strategy decision-making unit includes: Dynamic range control unit: Based on ADSR envelope analysis results, set differentiated compression thresholds and ratios to avoid excessive compression that could damage the original dynamic characteristics (such as changes in music intensity and sound effect impact), while ensuring that the dynamic range meets the requirements of the target context.
[0189] Frequency gain compensation unit: Based on the spectral feature analysis results and the human hearing sensitivity model, it performs fine-grained gain adjustment (non-global linear adjustment) on different frequency bands, focusing on optimizing the relative energy relationship of the frequency bands that are sensitive to human hearing, and ensuring spectral balance.
[0190] Spatial feature adaptation unit: Based on depth of field and sound field analysis results, fine-tune the level or spatial parameters of each channel (such as reverberation and sound image position) to maintain or optimize the spatial sense required by the target context and avoid spatial imbalance after multi-channel audio processing.
[0191] Rendering Feature Protection Unit: Identifies the coloring characteristics of the original audio (such as specific EQ curves and distortion textures) and preserves its features through parameter compensation during loudness adjustment; if the target context has special requirements, coloring adaptation optimization can be performed based on reference audio features.
[0192] Multi-channel synchronous processing unit: performs synchronous analysis and processing on 5.1, 7.1 and other surround sound multi-channel audio files, maintains the relative level relationship, phase relationship and spatial positioning between each channel, and ensures the integrity of the overall spatial characteristics of multi-channel audio.
[0193] 4) Batch and engineering management module, including: Batch task scheduling unit: Supports importing dozens to hundreds of audio files of different types or an entire folder at once. The system automatically assigns processing tasks according to preset rules to achieve fully automated batch processing.
[0194] File management optimization unit: The processed output files automatically retain the original filename, extension and storage path structure, avoiding users' later renaming and reorganization operations and reducing management costs.
[0195] Dynamic resource allocation unit: The system's underlying layer adopts task sharding and dynamic resource allocation algorithms to reasonably control CPU and memory usage, avoid overloading computer resources during large-scale batch processing, and ensure stable and efficient processing.
[0196] Processing report generation unit and historical record traceability unit: After batch processing is completed, a detailed processing report is automatically generated, which includes the initial loudness value, multi-dimensional feature parameters, processing strategy parameters, final indicators (the degree to which the target standard is met) and effect comparison summary for each file, which facilitates project auditing and traceability.
[0197] The aforementioned audio processing platform replaces fragmented, multi-tool workflows, integrating audio feature analysis, reference audio feature configuration, initial audio processing, and target audio naming within a single platform. This significantly reduces operational steps and time costs, making it particularly suitable for industrial batch processing tasks. It also eliminates frequent switching between different DAWs, editors, and plugins. By automatically maintaining source file names, it saves the tedious task of renaming. Based on an independent operating environment and optimized underlying processing logic, it mitigates the risk of computer resource overload caused by DAW reliance during large-scale batch processing. Furthermore, it supports multiple mainstream loudness standards and target platform requirements, enabling the processing of multi-channel audio.
[0198] It should be noted that the audio processing methods provided in this manual can be applied to various industries or scenarios, such as virtual reality processing software, home entertainment product software, digital cultural product production software, digital cultural creative software, digital cultural creative design, education, news, cultural content industry software, digital publishing software, digital music development and production, and digital mobile multimedia development and production. In some cases, they can also be applied to fields such as animation and game production engine software and development systems, game and animation software, animation and game digital content services, digital film and television development and production, and digital performance development and production.
[0199] Corresponding to the above method embodiments, this application also provides an audio processing system. Figure 5 A schematic diagram of an audio processing system according to an embodiment of this application is shown. Figure 5 As shown, the system includes: The file import and export interface 502 is configured to acquire various types of initial audio and reference audio features. The target parameter configuration interface 504 is configured to determine the audio feature relationships between different types of initial audio based on the initial audio features. The multidimensional audio feature analysis engine 506 is configured to construct a feature compensation model based on reference audio features and audio feature relationships, constrained by auditory frequency response relationships. The human ear auditory model integration unit 508 is configured to obtain the feature compensation value of the initial audio based on the feature compensation model. The processing strategy decision unit 510 is configured to adjust the initial audio features based on feature compensation values to obtain various types of target audio.
[0200] Optionally, the auditory frequency response relationship, including equal loudness curves, and the multi-dimensional audio feature analysis engine 506, are further configured as follows: Under the condition of satisfying the audio feature relationship, a feature compensation model is constructed with equal loudness curves as constraints and minimizing the feature difference between the reference audio features and the initial audio features as the objective. The processing strategy decision unit 510 is further configured as follows: Solve the feature compensation model to obtain the target feature differences; Based on the differences in target features, the feature compensation value of the initial audio is determined.
[0201] Optionally, the initial audio features include initial loudness features, initial envelope features, and initial spectral energy distribution features, and the reference audio features include reference loudness features, reference envelope features, and reference spectral energy distribution features; The multidimensional audio feature analysis engine 506 is further configured as follows: Based on the initial loudness characteristics, the relative loudness relationships between different types of initial audio are determined; Based on the initial envelope features, the relative envelope relationships between different types of initial audio are determined; Based on the initial spectral energy distribution characteristics, the relative spectral balance relationship between different types of initial audio is determined; The human ear auditory model integration unit 508 is further configured as follows: Under the conditions of satisfying the relative loudness relationship, relative envelope relationship and relative spectral balance relationship, a feature compensation model is constructed with equal loudness curves as constraints, aiming to minimize the difference in loudness features between the initial loudness features and the reference loudness features, the difference in envelope features between the initial envelope features and the reference envelope features, and the difference in spectral energy distribution features between the initial spectral energy distribution features and the reference spectral energy distribution features.
[0202] Optionally, the processing strategy decision unit 510 is further configured as follows: Based on the difference in target loudness features, the loudness feature compensation value of the initial audio is determined through a pre-constructed first mapping relationship; Based on the differences in target envelope features, the envelope feature compensation value of the initial audio is determined through a pre-constructed second mapping relationship; Based on the differences in the target spectral energy distribution characteristics, the compensation value of the spectral energy distribution characteristics of the initial audio is determined through a pre-constructed third mapping relationship.
[0203] Optionally, the initial audio features further include at least one of initial spatial features and initial rendering features, and the reference audio features further include at least one of reference spatial features and reference rendering features; The human ear auditory model integration unit 508 is further configured as follows: Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the difference in loudness features between the initial loudness features and the reference loudness features, the difference in envelope features between the initial envelope features and the reference envelope features, the difference in spectral energy distribution features between the initial spectral energy distribution features and the reference spectral energy distribution features, and the difference in spatial features between the initial spatial features and the reference spatial features. Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the difference in loudness features between the initial loudness features and the reference loudness features, the difference in envelope features between the initial envelope features and the reference envelope features, the difference in spectral energy distribution features between the initial spectral energy distribution features and the reference spectral energy distribution features, and the difference in rendering features between the initial rendering features and the reference rendering features. Under the conditions of satisfying the relative loudness relationship, relative envelope relationship, and relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the initial envelope feature and the reference envelope feature, the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature, the spatial feature difference between the initial spatial feature and the reference spatial feature, and the rendering feature difference between the initial rendering feature and the reference rendering feature.
[0204] Optionally, the processing strategy decision unit 510 is further configured as follows: Based on the difference in target loudness features, the loudness feature compensation value of the initial audio is determined through a pre-constructed first mapping relationship; Based on the differences in target envelope features, the envelope feature compensation value of the initial audio is determined through a pre-constructed second mapping relationship; Based on the differences in the target spectral energy distribution characteristics, the compensation value of the spectral energy distribution characteristics of the initial audio is determined through a pre-constructed third mapping relationship; Based on the differences in target spatial features, the spatial feature compensation value of the initial audio is determined through a pre-constructed fourth mapping relationship, and / or, based on the differences in target rendering features, the rendering feature compensation value of the initial audio is determined through a pre-constructed fifth mapping relationship.
[0205] Optionally, the processing strategy decision unit 510 is further configured as follows: Feed the target audio back to the front end; Receive optimization instructions sent from the front end for the target audio; In response to the optimization command, the reference audio features are adjusted to obtain the adjusted reference audio features; Return to the execution steps that use equal loudness curves as constraints and construct feature compensation models based on reference audio features and audio feature relationships, until optimized target audio of various types is obtained.
[0206] Optionally, there are multiple reference audio features, which are extracted from multiple reference audio sources. The target audio of any type has multiple features. The processing strategy decision unit 510 is further configured as follows: Feedback of various types of target audio to the front end; Receive evaluation instructions for the target audio sent from the front end; In response to the evaluation command, multiple reference audio features are filtered to obtain the filtered reference audio features.
[0207] Optionally, the file import and export interface 502 is further configured as follows: Receive configuration instructions for reference audio features sent by the front end, wherein the configuration instructions carry the reference audio features; Obtain multiple reference audio files; extract the audio features from the multiple reference audio files as multi-reference audio features.
[0208] The audio processing system provided in this application acquires various types of initial audio and reference audio features; determines the audio feature relationships between different types of initial audio based on the initial audio features; constructs a feature compensation model based on the reference audio features and the audio feature relationships, constrained by auditory frequency response relationships; obtains feature compensation values for the initial audio based on the feature compensation model; and adjusts the initial audio features based on the feature compensation values to obtain various types of target audio. When standardizing the loudness of the initial audio, it fully considers the differences in human ear sensitivity to loudness at different frequency bands, performing targeted gain or loss for different frequency bands, avoiding damage to the spatial sense and layering of the sound design, and improving the listening experience of the processed audio. Furthermore, when adjusting the loudness of each type of audio, it strictly adheres to the relative loudness relationships and spectral balance relationships that each type of audio should maintain in the final mixing context. This ensures that each type of audio meets loudness standards while adapting to the target scene, improving the overall listening experience and spatial layering of each type of audio after integration into the final product, and avoiding unreasonable masking effects between different types of audio.
[0209] The above is an illustrative scheme of an audio processing system according to this embodiment. It should be noted that the technical solution of this audio processing system and the technical solution of the aforementioned audio processing method belong to the same concept. Details not described in detail in the technical solution of the audio processing system can be found in the description of the technical solution of the aforementioned audio processing method. Furthermore, the components in the system embodiment should be understood as functional modules necessary to implement each step of the program flow or each step of the method; these functional modules are not actual functional divisions or separations. The system claims defined by such a set of functional modules should be understood as a functional module architecture that primarily implements the solution through the computer program described in the specification, and not as a physical system that primarily implements the solution through hardware.
[0210] Figure 6 A structural block diagram of a computing device according to one embodiment of this specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected to the memory 610 via a bus 630, and a database 650 is used to store data.
[0211] The computing device 600 also includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 640 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0212] In one embodiment of this specification, the above-described components of the computing device 600 and Figure 6 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 6 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0213] The computing device 600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 600 can also be a mobile or stationary server.
[0214] The memory 610 is used to store computer programs / instructions, and the processor 620 is used to execute the following computer programs / instructions, which, when executed by the processor, implement the steps of the above-described audio processing method.
[0215] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the audio processing method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the audio processing method described above.
[0216] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described audio processing method.
[0217] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the audio processing method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the audio processing method described above.
[0218] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described audio processing method.
[0219] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above-described audio processing method belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the above-described audio processing method.
[0220] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or advantageous.
[0221] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0222] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0223] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0224] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An audio processing method, characterized in that, include: Acquire various types of initial audio and reference audio features; Based on the initial audio features of the initial audio, determine the audio feature relationships between different types of initial audio; Using auditory frequency response as a constraint, a feature compensation model is constructed based on the reference audio features and the relationship between the audio features; Based on the feature compensation model, the feature compensation value of the initial audio is obtained; Based on the feature compensation value, the initial audio features are adjusted to obtain various types of target audio.
2. The method according to claim 1, characterized in that, The auditory frequency response relationship includes equal loudness curves. The feature compensation model, constrained by the auditory frequency response relationship and based on the reference audio features and the audio feature relationship, includes: Under the condition of satisfying the audio feature relationship, a feature compensation model is constructed with equal loudness curves as constraints and minimizing the feature difference between the reference audio feature and the initial audio feature as the objective. The step of obtaining the feature compensation value of the initial audio based on the feature compensation model includes: Solve the feature compensation model to obtain the target feature differences; Based on the differences in the target features, the feature compensation value of the initial audio is determined.
3. The method according to claim 2, characterized in that, The initial audio features include initial loudness features, initial envelope features, and initial spectral energy distribution features; the reference audio features include reference loudness features, reference envelope features, and reference spectral energy distribution features. The determination of audio feature relationships between different types of initial audio based on the initial audio features includes: Based on the initial loudness characteristics, the relative loudness relationships between the initial audio types are determined; Based on the initial envelope features, the relative envelope relationships between the initial audio types are determined; Based on the initial spectral energy distribution characteristics, the relative spectral balance relationship between each type of initial audio is determined; Under the condition of satisfying the audio feature relationship, and with equal loudness curves as constraints, a feature compensation model is constructed with the objective of minimizing the feature difference between the reference audio features and the initial audio features, including: Under the conditions of satisfying the relative loudness relationship, the relative envelope relationship, and the relative spectral balance relationship, a feature compensation model is constructed with equal loudness curves as constraints, aiming to minimize the difference in loudness features between the initial loudness feature and the reference loudness feature, the difference in envelope features between the initial envelope feature and the reference envelope feature, and the difference in spectral energy distribution features between the initial spectral energy distribution feature and the reference spectral energy distribution feature.
4. The method according to claim 2 or 3, characterized in that, Determining the feature compensation value of the initial audio based on the target feature differences includes: Based on the target loudness feature differences, the loudness feature compensation value of the initial audio is determined through a pre-constructed first mapping relationship; Based on the differences in target envelope features, the envelope feature compensation value of the initial audio is determined through a pre-constructed second mapping relationship; Based on the differences in the target spectral energy distribution characteristics, the compensation value of the spectral energy distribution characteristics of the initial audio is determined through a pre-constructed third mapping relationship.
5. The method according to claim 4, characterized in that, After determining the spectral energy distribution feature compensation value of the initial audio based on the target spectral energy distribution feature difference through a pre-constructed third mapping relationship, the method further includes: Receive adjustment instructions from the front end for the compensation values of the spectral energy distribution characteristics; In response to the adjustment command, the spectral energy distribution characteristic compensation value is adjusted to obtain the adjusted spectral energy distribution characteristic compensation value.
6. The method according to claim 3, characterized in that, The initial audio features further include at least one of initial spatial features and initial rendering features, and the reference audio features further include at least one of reference spatial features and reference rendering features; Under the condition of satisfying the relative loudness relationship, the relative envelope relationship, and the relative spectral balance relationship, and with equal loudness curves as constraints, a feature compensation model is constructed with the objective of minimizing the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the initial envelope feature and the reference envelope feature, and the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature. This model includes at least one of the following: Under the conditions of satisfying the relative loudness relationship, the relative envelope relationship, and the relative spectral balance relationship, a feature compensation model is constructed with equal loudness curves as constraints, aiming to minimize the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the initial envelope feature and the reference envelope feature, the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature, and the spatial feature difference between the initial spatial feature and the reference spatial feature. Under the conditions of satisfying the relative loudness relationship, the relative envelope relationship, and the relative spectral balance relationship, a feature compensation model is constructed with equal loudness curves as constraints, aiming to minimize the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the initial envelope feature and the reference envelope feature, the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature, and the rendering feature difference between the initial rendering feature and the reference rendering feature. Under the conditions of satisfying the relative loudness relationship, the relative envelope relationship, and the relative spectral balance relationship, and with the equal loudness curve as a constraint, a feature compensation model is constructed with the objectives of minimizing the loudness feature difference between the initial loudness feature and the reference loudness feature, the envelope feature difference between the initial envelope feature and the reference envelope feature, the spectral energy distribution feature difference between the initial spectral energy distribution feature and the reference spectral energy distribution feature, the spatial feature difference between the initial spatial feature and the reference spatial feature, and the rendering feature difference between the initial rendering feature and the reference rendering feature.
7. The method according to claim 6, characterized in that, Determining the feature compensation value of the initial audio based on the target feature differences includes: Based on the target loudness feature differences, the loudness feature compensation value of the initial audio is determined through a pre-constructed first mapping relationship; Based on the differences in target envelope features, the envelope feature compensation value of the initial audio is determined through a pre-constructed second mapping relationship; Based on the differences in the target spectral energy distribution characteristics, the compensation value for the spectral energy distribution characteristics of the initial audio is determined through a pre-constructed third mapping relationship; Based on the differences in target spatial features, the spatial feature compensation value of the initial audio is determined through a pre-constructed fourth mapping relationship, and / or, based on the differences in target rendering features, the rendering feature compensation value of the initial audio is determined through a pre-constructed fifth mapping relationship.
8. The method according to claim 1, characterized in that, After adjusting the initial audio features based on the feature compensation value to obtain multiple types of target audio, the method further includes: The target audio is fed back to the front end; Receive optimization instructions sent by the front end for the target audio; In response to the optimization instruction, the reference audio features are adjusted to obtain the adjusted reference audio features; Return to the step of constructing a feature compensation model based on the reference audio features and the relationship between the audio features, constrained by the equal loudness curve, until optimized target audio of various types is obtained.
9. The method according to claim 1, characterized in that, The reference audio features are multiple, and the multiple reference audio features are extracted from the multiple reference audios. The target audio of any type has multiple features. After adjusting the initial audio features based on the feature compensation value to obtain multiple types of target audio, the method further includes: Feedback of various types of target audio to the front end; Receive the evaluation instruction sent by the front end for the target audio; In response to the evaluation instruction, the multiple reference audio features are filtered to obtain the filtered reference audio features.
10. The method according to claim 1, characterized in that, Obtain reference audio features, including at least one of the following: Receive a configuration instruction for reference audio features sent by the front end, wherein the configuration instruction carries the reference audio features; Obtain multiple reference audios; extract the audio features of the multiple reference audios as multi-reference audio features.
11. An audio processing system, characterized in that, include: The file import and export interface is configured to obtain various types of initial audio; The target parameter configuration interface is configured to obtain reference audio features; A multidimensional audio feature analysis engine is configured to determine the audio feature relationships between different types of initial audio based on the initial audio features of the initial audio. The human ear hearing model integration unit is configured to construct a feature compensation model based on the reference audio features and the audio feature relationship, constrained by the auditory frequency response relationship. The processing strategy decision unit is configured as follows: Based on the feature compensation model, the feature compensation value of the initial audio is obtained; Based on the feature compensation value, the initial audio features are adjusted to obtain various types of target audio.
12. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, It stores a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
14. A computer program product, characterized in that, Includes a computer program / instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.