Sound balance degree real-time evaluation method and system based on frequency spectrum characteristics
By acquiring the standard voice frequency band and energy value, segmenting and analyzing the audio signal, calculating and correcting the voice deviation, the problem of inaccurate identification of multi-voice audio energy distribution in existing technologies is solved, and dynamic evaluation and optimization guidance of voice balance state are realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN WANBO ZHIHUI CLOUD EDUCATION TECHNOLOGY CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-01
AI Technical Summary
Existing audio quality assessment methods cannot accurately identify and quantify the independent energy distribution in multi-part mixed audio. In particular, they cannot dynamically adjust the assessment sensitivity when the number of parts changes, which makes it impossible to locate the imbalance problem of a specific part and to distinguish whether the balance state is a unidirectional drift or a dynamic oscillation.
By acquiring the frequency band range and energy value of the standard voice part, the audio signal is analyzed in segments, the voice part deviation is calculated, and correction is made based on the voice part number threshold. Combined with the deviation difference analysis of adjacent time windows, the evaluation results of dynamic trend classification are generated.
It achieves accurate quantitative evaluation of multi-voice audio, can identify changes in the energy state of specific voices, enhances sensitivity to minor imbalances, distinguishes between technical imbalances and artistic design, and provides targeted optimization guidance.
Smart Images

Figure CN121963776A_ABST
Abstract
Description
A Real-Time Evaluation Method and System for Acoustic Part Balance Based on Spectral Characteristics Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a method and system for real-time evaluation of acoustic balance based on spectral characteristics. Background Technology
[0002] In fields such as music production, live sound reinforcement, and audio post-processing, it is crucial to accurately assess the dynamic balance of multi-part mixed audio signals, but existing audio quality evaluation methods have limitations.
[0003] Specifically, existing spectral analysis typically focuses only on macroscopic indicators such as overall frequency response or total harmonic distortion, failing to specifically identify and quantify the independent energy distribution of specific standard voices such as vocals, strings, and percussion in mixed audio. This results in the inability to locate imbalances in specific voices. Secondly, when processing time-varying audio, existing methods often employ static energy statistics (such as overall loudness or spectral flatness) over a fixed time window, lacking fine-grained tracking of the evolution of energy ratios between voices over time. Especially when the number of voices changes (such as the overlapping or removal of instruments in an orchestral arrangement), existing algorithms cannot dynamically adjust their evaluation sensitivity. For example, in symphonic passages or choral works, when the number of simultaneously existing voices increases, small energy deviations in individual voices can be masked by the voice overlap effect. However, in actual auditory perception, the balance tolerance in such multi-voice coexistence scenarios is actually lower, and current technology lacks a dynamic evaluation threshold adjustment mechanism based on the number of voices. Furthermore, mainstream evaluation models typically only output a single overall score, failing to distinguish whether the balance state is a unidirectional drift (such as a continuous enhancement of the bass) or a dynamic oscillation (such as alternating prominence of the lead vocals and harmonies), making it difficult for audio engineers to optimize mixing strategies in a targeted manner. Summary of the Invention
[0004] To address the technical problem that existing technologies struggle to automatically generate quantitative assessment results reflecting the dynamic balance defects and evolution patterns of acoustic parts, thus limiting the accuracy and practicality of intelligent audio quality diagnosis, this invention provides a real-time acoustic part balance assessment method and system based on spectral characteristics.
[0005] A real-time evaluation method for acoustic part balance based on spectral characteristics includes: acquiring multiple different standard acoustic parts, acquiring the standard frequency band range and standard energy value corresponding to each standard acoustic part, acquiring the audio signal to be evaluated, segmenting the audio signal to be evaluated into multiple segments by a preset time window, and acquiring multiple frequency points and corresponding amplitude values for each segment; acquiring the standard acoustic parts contained in the i-th segment based on the multiple frequency points and the standard frequency band range corresponding to each standard acoustic part in the i-th segment, and evaluating the standard acoustic parts contained in the i-th segment based on the standard acoustic parts. The actual energy values of each standard voice part contained in the i-th audio segment are obtained from the frequency points and the amplitude values corresponding to each frequency point. The initial voice part deviation of the i-th audio segment is obtained based on the actual energy values and standard energy values of each standard voice part contained in the i-th audio segment. The initial voice part deviation of the i-th audio segment is corrected according to the number of standard voice parts contained in the i-th audio segment to obtain the target voice part deviation of the i-th audio segment. The deviation difference between the target voice part deviations of adjacent audio segments is obtained, and the voice part balance evaluation result of the audio signal is obtained based on the deviation difference.
[0006] Optionally, obtaining the standard voices contained in the i-th audio segment based on multiple frequency points of the i-th audio segment and the standard frequency band range corresponding to each standard voice part includes: matching the standard frequency band range corresponding to each standard voice part with all frequency points of the i-th audio segment; when there is a frequency point belonging to the i-th audio segment within the standard frequency band range of a certain standard voice part, it is determined that the frequency point belongs to the standard voice part, and the standard voice part containing the frequency point of the i-th audio segment is marked as the standard voice part contained in the i-th audio segment.
[0007] Optionally, obtaining the actual energy value of each standard voice part in the i-th audio segment based on the frequency points of each standard voice part and the amplitude value corresponding to each frequency point includes: adding the amplitude values of all frequency points in the j-th standard voice part contained in the i-th audio segment to obtain the actual energy value of the j-th standard voice part contained in the i-th audio segment.
[0008] Optionally, obtaining the initial voice deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment includes: taking the absolute value of the difference between the actual energy value and the standard energy value of each standard voice part contained in the i-th audio segment, and obtaining the absolute energy difference of each standard voice part; dividing the absolute energy difference of each standard voice part contained in the i-th audio segment by the standard energy value, and obtaining the energy difference ratio of each standard voice part; dividing the sum of the energy difference ratios of all standard voice parts contained in the i-th audio segment by the total number of all standard voice parts contained in the i-th audio segment, and obtaining the initial voice deviation of the i-th audio segment.
[0009] Optionally, correcting the initial voice deviation of the i-th audio segment based on the number of standard voices contained in the i-th audio segment and obtaining the target voice deviation of the i-th audio segment includes: obtaining a preset quantity threshold and a standard correction ratio; subtracting the preset quantity threshold from the number of standard voices contained in the i-th audio segment and dividing by the preset quantity threshold to obtain a quantity ratio; if the quantity ratio is greater than zero, incrementing the quantity ratio by 1 to obtain an adjustment ratio; multiplying the adjustment ratio by the standard correction ratio to obtain a target correction ratio; incrementing the target correction ratio by 1 to obtain an increase ratio; multiplying the initial voice deviation of the i-th audio segment by the increase ratio to obtain the target voice deviation of the i-th audio segment; if the quantity ratio is not greater than zero, using the initial voice deviation of the i-th audio segment as the target voice deviation of the i-th audio segment.
[0010] Optionally, obtaining the audio signal's acoustic balance assessment result based on the deviation difference includes: analyzing the continuous change direction of the deviation difference between adjacent audio target acoustic parts; when the deviation difference continuously changes in the same direction, determining that the acoustic balance is in a unidirectional evolution state; when the deviation difference alternately changes in opposite directions, determining that the acoustic balance is in a dynamic fluctuation state; and generating a balance assessment result including a balance trend classification based on the duration of the unidirectional evolution state or the dynamic fluctuation state.
[0011] A real-time sound balance evaluation system based on spectral characteristics is also provided. The system includes: an acquisition module, used to acquire multiple different standard sound parts, acquire the standard frequency band range and standard energy value corresponding to each standard sound part, acquire the audio signal to be evaluated, segment the audio signal to be evaluated into multiple segments by a preset time window, and acquire multiple frequency points and the amplitude value corresponding to each frequency point of each segment audio; a first data processing module, used to acquire the standard sound parts contained in the i-th segment audio based on the multiple frequency points of the i-th segment audio and the standard frequency band range corresponding to each standard sound part, and process the standard sound parts contained in the i-th segment audio according to the standard frequency band range corresponding to each standard sound part. The first data processing module obtains the actual energy value of each standard voice part contained in the i-th audio segment based on the frequency point and the amplitude value corresponding to each frequency point; the second data processing module obtains the initial voice part deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment, corrects the initial voice part deviation of the i-th audio segment based on the number of standard voice parts contained in the i-th audio segment, and obtains the target voice part deviation of the i-th audio segment; the third data processing module obtains the deviation difference between the target voice part deviations of adjacent audio segments, and obtains the voice part balance evaluation result of the audio signal based on the deviation difference.
[0012] Optionally, the first data processing module is further configured to: match the standard frequency band range corresponding to each standard voice part with all frequency points of the i-th segment audio based on each standard voice part; when there is a frequency point belonging to the i-th segment audio within the standard frequency band range of a certain standard voice part, determine that the frequency point belongs to the standard voice part, and mark the standard voice part containing the frequency point of the i-th segment audio as the standard voice part contained in the i-th segment audio.
[0013] Optionally, the first data processing module is further configured to: sum the amplitude values of all frequency points present in the j-th standard voice part contained in the i-th audio segment, and obtain the actual energy value of the j-th standard voice part contained in the i-th audio segment.
[0014] Optionally, the second data processing module is further configured to: take the absolute value of the difference between the actual energy value and the standard energy value of each standard voice part contained in the i-th audio segment, and obtain the absolute energy difference of each standard voice part; divide the absolute energy difference of each standard voice part contained in the i-th audio segment by the standard energy value, and obtain the energy difference ratio of each standard voice part; divide the sum of the energy difference ratios of all standard voice parts contained in the i-th audio segment by the total number of all standard voice parts contained in the i-th audio segment, and obtain the initial voice part deviation of the i-th audio segment.
[0015] The beneficial effects of this invention are as follows: In the entire real-time evaluation method of voice balance based on spectral characteristics, firstly, by predefining the spectral characteristics and energy benchmark of standard voices, and combining the frequency point energy collection of segmented audio, it is possible to accurately extract and quantify the energy state of each independent voice at any time from the mixed spectrum (such as identifying the absence or abnormal prominence of a specific instrument), overcoming the defect of existing overall frequency response analysis that cannot locate the imbalance of specific voices; secondly, it innovatively introduces a deviation correction mechanism based on a threshold for the number of voices. When the number of active voices exceeds a preset threshold (such as a full symphony), the deviation correction mechanism is applied. The system automatically amplifies the initial deviation value to make the evaluation criteria more stringent, thus aligning with the heightened sensitivity of the human ear to minor imbalances in complex mixes. This effectively eliminates the sensitivity inaccuracy caused by changes in the number of voices. Finally, through the difference sequence analysis of the target deviation between adjacent time windows, the system analyzes the continuous direction of the balance state's evolution over time (such as continuous degradation or oscillation) and its duration, generating evaluation results that include dynamic trend classification. This allows audio engineers to specifically distinguish between technical imbalances (continuous offsets that need to be prioritized for repair) and artistic design (reasonable fluctuations that can be retained). Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0017] Figure 1 is a partial flowchart of steps S1 to S2 in the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 2 is a partial flowchart of steps S2 to S3 in the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 3 is a partial flowchart of steps S3 to S4 in the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 4 is a step diagram of the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 5 is a partial step diagram of step S2 in the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 6 is a partial step diagram of step S3 in the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 7 is a partial step diagram of step S3 in the real-time sound balance evaluation method based on spectral characteristics of the present invention; Figure 8 is a partial step diagram of step S5 in the real-time sound balance evaluation method based on spectral characteristics of the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0019] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0020] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, the terms "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0021] As shown in Figures 1 to 4, a real-time evaluation method for acoustic part balance based on spectral characteristics is provided. In one embodiment, the method includes: S1, acquiring multiple different standard acoustic parts, acquiring the standard frequency band range and standard energy value corresponding to each standard acoustic part, acquiring the audio signal to be evaluated, segmenting the audio signal to be evaluated into segments according to a preset time window and acquiring multiple segmented audios, and acquiring multiple frequency points and the amplitude value corresponding to each frequency point of each segmented audio; S2, acquiring the standard acoustic parts contained in the i-th segmented audio based on the multiple frequency points of the i-th segmented audio and the standard frequency band range corresponding to each standard acoustic part, and according to the i-th segment... S3. Obtain the actual energy value of each standard voice part in the i-th audio segment by taking the frequency points of each standard voice part and the amplitude value corresponding to each frequency point; S4. Obtain the initial voice part deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part in the i-th audio segment, correct the initial voice part deviation of the i-th audio segment based on the number of standard voice parts in the i-th audio segment, and obtain the target voice part deviation of the i-th audio segment; S5. Obtain the deviation difference between the target voice part deviations of adjacent audio segments, and obtain the voice part balance evaluation result of the audio signal based on the deviation difference.
[0022] In this embodiment, it should be noted that in S1, an evaluation benchmark is constructed and a preliminary time-frequency conversion is performed on the input audio signal to be analyzed, laying the foundation for subsequent refined part balance calculations. First, multiple "standard parts" representing different instruments or vocal types need to be predefined and obtained. These parts are the target objects of evaluation. For example, in a pop music production scenario, standard parts may include lead vocals, electric bass, rhythm guitar, bass drum, snare drum, etc.; in a symphony orchestra scenario, they may include violin, cello, woodwind, brass, etc. Each standard part is associated with the most representative frequency range of its acoustic characteristics, i.e., the "standard frequency band range" (e.g., electric bass may mainly occupy 80Hz-250Hz, and lead vocals may be concentrated in the core part of 200Hz-2kHz). At the same time, it is necessary to obtain or preset the energy level that each standard part should have under ideal balanced mixing conditions, i.e., the standard energy value. These values are set as empirical values (e.g., target values for the mixing ratio of each part under standard studio monitoring conditions) or are average values obtained through statistical analysis of a limited number of historical reference tracks or reference tracks in a database.
[0023] Subsequently, the actual audio signal to be evaluated (e.g., a song that has just completed its initial mixing or a recording of a live performance) is received. To capture the dynamic changes in the balance of the vocal parts in the audio signal over time (corresponding to the time-varying problem mentioned in the background art), the signal is divided into multiple consecutive audio segments according to a preset time window (e.g., 50 milliseconds or 100 milliseconds). This segmentation ensures that subsequent analysis can be tracked in the time dimension. Finally, for each obtained audio segment, its frequency components need to be precisely located. This is usually achieved by performing a short-time Fourier transform on the audio segment, which is a relatively mature technique; the core output of the Fourier transform is a complex number, where each frequency point corresponds to a specific frequency component (e.g., discrete frequencies spaced at 1 Hz), and the "magnitude" (i.e., the absolute value of the complex number) of the transform result at that frequency point quantifies the "amplitude" of that frequency component in the current audio segment (intuitively, it can approximate the strength of the sound vibration at that frequency, and is a direct reflection of its energy contribution). This step completes the conversion from a continuous time-domain signal to a discrete frequency-domain representation. The extracted frequency points and their amplitude values will be used in subsequent steps (S2) to identify which actual active standard voices and their specific energy distributions are contained in the current time segment.
[0024] In S2, based on time window segmentation, refined voice identification and energy aggregation are performed on each segmented audio. This technical implementation relies on the standard voice spectral feature library established in S1 and the discrete spectral data of the segmented audio (i.e., all frequency points and their amplitude values). Specifically, the first step is to identify the actual valid standard voices present in the i-th segmented audio. This is achieved by iterating through all preset standard voices (such as electric bass, lead vocals, snare drum, etc. defined in S1) and matching their standard frequency band range (e.g., 80Hz-250Hz for electric bass) one by one with all frequency points (discrete frequency components decomposed by STFT) of the current segmented audio.
[0025] When at least one frequency point belonging to a segment of audio exists within the frequency band of a standard voice (e.g., 1kHz-3.5kHz for a violin voice) (e.g., a significant amplitude value is detected at 1.2kHz), that frequency point is determined to belong to that standard voice, and that voice is simultaneously marked as a valid voice in the current segment of audio. This is an acoustic object separation operation based on physical frequency domain characteristics. In this way, the problem of being unable to locate the imbalance of a specific voice can be solved to some extent. Existing overall spectrum analysis cannot distinguish voices that are superimposed, but by matching a predefined frequency band range, the energy points in the composite audio can be decomposed into independent voice entities.
[0026] For example, in a pop music segment, if there are no significant frequency points in the vocal frequency band (200Hz-2kHz) within a certain time window, it can be determined that the lead vocal part is missing during that period; if there are high amplitude points in the bass drum frequency band (60Hz-100Hz), it can be confirmed that the bass drum is active.
[0027] Furthermore, after identifying the effective voice parts, S2 calculates the actual energy contribution of each effective voice part within the current time segment. This requires aggregating the amplitude data of all discrete frequency points assigned to a specific voice part. Specifically, for the j-th standard voice part (e.g., the identified violin voice part) marked in the i-th audio segment, all frequency points belonging to it within its standard frequency band (e.g., multiple frequency components falling within the 1kHz-3.5kHz range) are extracted, and the amplitude values corresponding to these frequency points are accumulated. This accumulation result is the actual energy value of the j-th voice part within the current time window. Its physical meaning is: by dividing the frequency band, the energy points in the composite spectrum are decomposed into independent voice parts, and then the relative energy intensity of the voice part in the time domain segment is obtained through the linear superposition of the amplitude values of each frequency point within the same voice part (rather than the sum of squares or logarithmic operations).
[0028] For example, if a violin part in a certain audio segment is matched to 10 frequency points with amplitude values from A1 to A10, its actual energy value is equal to the algebraic sum of these 10 values. This design specifically solves the problem of independent energy quantization for each part—existing methods can only calculate the total energy across the entire frequency band or the energy of a wideband because they cannot distinguish between parts, while here the energy value precisely corresponds to a specific acoustic object (such as an electric bass or lead vocal), thus supporting subsequent part-level deviation analysis. It should be noted that the direct summation of amplitude magnitudes (rather than the sum of squares) is used to avoid excessive amplification of low-frequency energy due to squaring operations (such as the energy of an electric bass, which is usually concentrated in the low frequencies), and to maintain comparability with the standard energy value under the same dimensions (because the standard energy value is also based on the amplitude summation logic).
[0029] In S3, based on the actual energy values of each voice part provided in S2, the deviation of the overall balance state of the multi-voice components within the current time window from the preset standard is quantified. First, for each valid standard voice part (such as the violin, electric bass, etc., determined in the previous steps) identified in the i-th audio segment, the absolute difference between its actual energy value and the preset standard energy value is calculated. Taking the absolute value ensures that the deviation of all voice parts is considered as a loss (whether too strong or too weak), avoiding the overall balance problem being underestimated due to positive and negative cancellation.
[0030] Subsequently, to eliminate the dimensional differences caused by different standard energy benchmarks for different vocal parts (such as the bass drum naturally having higher energy than the hi-hat), the absolute difference is divided by its corresponding standard energy value to obtain a normalized energy difference ratio. This ratio represents the degree of relative imbalance of that vocal part independent of other vocal parts (for example, a difference ratio of 0.15 indicates that the actual energy of this vocal part deviates from the ideal value by 15%).
[0031] Next, the sum of the differences in proportions for all effective voices (only those existing in the current time window) is divided by the total number of effective voices to obtain the initial voice deviation. This average value reflects the overall coordinated deviation level of the energy states of all active voices within the current time window. For example, in a rock chorus containing four voices—electric bass, lead vocals, kick drum, and snare drum—if the electric bass is 10% stronger, the lead vocals 5% weaker, the kick drum 8% weaker, and the snare drum 2% weaker, then the initial deviation is the average of the sum of all proportions, characterizing the degree of deterioration in the overall balance of the mix. In this way, voices are quantified independently to a certain extent. The initial deviation is essentially an average statistical measure of the balance state of all effective voices themselves; the higher the value, the further the overall deviation from the ideal balance state.
[0032] Furthermore, after obtaining the initial deviation of each voice part, the deviation value is dynamically adjusted based on the current number of effective voice parts. The logic of the adjustment stems from an acoustic psychophysical phenomenon: when the number of simultaneously existing voice parts increases (such as in a full symphony passage or a dense arrangement in pop music), the auditory tolerance for imbalance in a single voice part decreases (even if each voice part has only a small deviation, the sense of chaos is significantly enhanced after superposition), and existing technologies lack a response mechanism to this. In implementation, a key threshold is preset (e.g., 2 or 3 voice parts). If the current number of effective voice parts exceeds this threshold, the excess ratio is calculated as a base, and combined with a preset standard correction coefficient (engineering experience value) to generate the final target correction ratio, thereby amplifying the initial deviation value. Conversely, if the number of voice parts does not exceed the threshold, the initial value is directly used.
[0033] For example, when the quantity threshold is set to 2 and the standard correction coefficient is 0.5, there are currently 4 voices (over-limit ratio = (4-2) / 2 = 1), the adjustment ratio = 1+1 = 2, the target correction ratio = 2 × 0.5 = 1.0, the increase ratio = 1.0 + 1 = 2.0, and the final target deviation = initial value × 2.0. This correction essentially simulates the change in human ear sensitivity in multi-voice scenarios—when the number of voices increases, even the same energy deviation ratio needs to be evaluated as a more severe imbalance (for example, a 20% deviation in 4 voices is actually perceived as equivalent to a 40% degradation in 2 voices), thus solving to some extent the defect of "mismatch in sensitivity evaluation when the number of voices changes" in the background technology. The corrected target voice deviation becomes a standardized indicator for subsequent time dynamic analysis. Its value can objectively quantify the absolute balance state of the current window and perform perceptual alignment calibration based on the voice coexistence relationship, providing a time series input of consistent order of magnitude for the evolution trend analysis of S4.
[0034] In S4, the time series of target part deviations output from S3 is analyzed (each time window corresponds to a corrected equilibrium quantification value). The dynamic evolution of the part's equilibrium state is explored through the deviation differences between adjacent segments. Specifically, the difference in target part deviation between two adjacent segments (e.g., segment i and segment i+1) is calculated. The sign of this difference (positive / negative) indicates the direction of change in the equilibrium state: if the deviation of the current segment is higher than that of the previous segment, the difference is positive, indicating that the equilibrium state is deteriorating (e.g., the continuous enhancement of the electric bass part leads to increased imbalance); conversely, if the difference is consistently negative, it indicates that the equilibrium state is continuously improving (e.g., the mixing engineer gradually weakens an overly strong snare drum). The continuity analysis of adjacent differences aims to identify the persistent characteristics of deviation changes, to some extent solving the deficiency of being unable to distinguish between unidirectional drift and dynamic fluctuations. Specifically, existing methods only output static comprehensive scores and cannot reveal how the equilibrium state evolves over time. For example, in the development section of a symphony, the string section may continuously intensify (resulting in positive deviation differences for multiple consecutive windows), while in the transitional section of pop music, the lead vocals and backing vocals may alternately emphasize (resulting in alternating positive and negative deviation differences). This dynamic difference chain based on discrete time windows transforms the continuous audio timeline into a discrete event sequence with computable evolutionary patterns, providing a quantitative basis for subsequent trend classification.
[0035] Furthermore, after constructing the difference sequence, S4 determines the state of the continuous occurrence pattern of the difference symbols. When multiple consecutive difference symbols move in the same direction (e.g., the difference sequence is [+0.1,+0.2,+0.15]), the vocal balance is determined to be in a unidirectional evolution state. This state usually corresponds to an imbalance in vocal energy (e.g., an unnoticed continuous boost in the bass part during mixing), requiring audio engineers to check the automated control parameters of specific channels. When the difference sequence frequently shows sign reversals (e.g., [+0.3,-0.2,+0.4,-0.3]), it is determined to be a dynamic fluctuation state, which often reflects intentional changes driven by creative intent (e.g., the dynamic design of vocals between verses and choruses).
[0036] Furthermore, the final evaluation result is generated by combining the duration of the two states: long-term unidirectional evolution (such as continuous deterioration exceeding 10 seconds) indicates a serious risk of imbalance requiring intervention; short-term high-frequency fluctuations (such as multiple oscillations within 2 seconds) may be a reasonable artistic expression. The trend classification results output in this step (unidirectional / fluctuation + duration) provide targeted guidance for mixing optimization—engineers can prioritize correcting long-term unidirectional deterioration segments while preserving artistic dynamic fluctuations. This mechanism not only overcomes the limitation of existing evaluation methods in not being able to distinguish trend types, but also achieves automatic quantitative diagnosis of the evolution pattern of dynamic energy balance defects in vocal parts through fine-grained tracking at the time window level, upgrading audio quality analysis from static scoring to process monitoring.
[0037] In summary, the real-time assessment method for voice balance based on spectral characteristics achieves several key improvements. First, by defining the spectral characteristics and energy benchmark of a predefined standard voice and combining this with the frequency energy aggregation of segmented audio, it accurately isolates and quantifies the energy state of each independent voice at any given moment from the mixed spectrum (e.g., identifying the absence or abnormal prominence of a specific instrument), overcoming the limitation of existing overall frequency response analysis in locating specific voice imbalances. Second, it innovatively introduces a deviation correction mechanism based on a voice quantity threshold, which corrects for deviations when the number of active voices exceeds a preset threshold (e.g., in a full symphony passage). The system automatically amplifies the initial deviation value, making the evaluation criteria more stringent to match the heightened sensitivity of the human ear to minor imbalances in complex mixes, effectively eliminating sensitivity inaccuracies caused by changes in the number of voices. Finally, through the difference sequence analysis of the target deviation between adjacent time windows, the system analyzes the continuous direction of the balance state's evolution over time (such as continuous degradation or oscillation) and its duration, generating evaluation results that include dynamic trend classification. This allows audio engineers to specifically distinguish between technical imbalances (continuous offsets that need to be prioritized for repair) and artistic design (reasonable fluctuations that can be retained).
[0038] As shown in Figure 5, in one embodiment, obtaining the standard sound part contained in the i-th segment audio based on multiple frequency points of the i-th segment audio and the standard frequency band range corresponding to each standard sound part in S2 includes: S21, matching the standard frequency band range corresponding to each standard sound part with all frequency points of the i-th segment audio; S22, when there is a frequency point belonging to the i-th segment audio within the standard frequency band range of a certain standard sound part, determining that the frequency point belongs to the standard sound part, and marking the standard sound part containing the frequency point of the i-th segment audio as the standard sound part contained in the i-th segment audio.
[0039] In this embodiment, it should be noted that in S21, when performing the matching operation between the standard voice and the segmented audio frequency points, the preset standard frequency band range of each standard voice is used as the comparison benchmark (e.g., the violin voice is defined as 1kHz-3.5kHz). All discrete frequency points (e.g., 100Hz, 1.2kHz, 2.5kHz, etc.) extracted by S1 in the i-th segmented audio are scanned one by one. The matching principle is based on the frequency domain separability of physical acoustics—the energy of a specific instrument or human voice is mainly concentrated within its characteristic frequency band (e.g., the 80Hz main frequency energy of the bass drum is much higher than the frequency band above 1kHz). When a frequency point is detected to fall within the frequency band range of a standard voice (e.g., the 1.2kHz point falls within the preset 1kHz-3.5kHz range for the violin voice), an attribution relationship is established. This operation does not rely on amplitude thresholds (even weak signals are included), ensuring sensitive identification of the voice's presence. For example, in a symphonic passage, if a 1.5kHz frequency point appears in the oboe part (preset frequency band 800Hz-2kHz) (even if the amplitude is small), the oboe is determined to be present, so as to avoid missing weak playing parts.
[0040] In S22, marking is performed based on the matching results of S21. When at least one assigned frequency point (regardless of amplitude) exists within the frequency band of a standard voice part, it is marked as a valid voice part of the i-th audio segment (the standard voice part contained in the i-th audio segment). Simultaneously, dynamic resolution is supported for multi-voice coexistence scenarios: the same frequency point is only assigned to the voice part that was matched first (the preset voice part library is sorted by priority), ensuring the uniqueness of energy assignment. For example, in a rock music section, if a 1.5kHz frequency point is detected in the overlapping frequency band area of the electric guitar (preset 100Hz-5kHz) and the lead vocals (200Hz-2kHz), it is preferentially assigned to the lead vocals voice part according to the preset order.
[0041] In one embodiment, S2 obtaining the actual energy value of each standard sound part contained in the i-th segment audio based on the frequency points of each standard sound part contained in the i-th segment audio and the amplitude value corresponding to each frequency point includes: adding the amplitude values of all frequency points present in the j-th standard sound part contained in the i-th segment audio to obtain the actual energy value of the j-th standard sound part contained in the i-th segment audio.
[0042] In this embodiment, it should be noted that for each valid voice part marked in S22 (such as the j-th violin voice part), the process of calculating the actual energy value is to perform linear accumulation of the amplitude magnitude values of all frequency points belonging to that voice part. Specifically, firstly, the sum of squares (the conventional energy calculation method) is discarded to avoid excessive amplification of the energy proportion of low-frequency voice parts (such as electric bass) (square operations would widen the gap between high and low frequencies), maintaining balanced comparability with high-frequency voice parts (such as hi-hats); secondly, the magnitude value (the absolute value of complex amplitude) is used instead of the real / imaginary part to preserve the original frequency domain intensity information; finally, linear accumulation ensures alignment with the dimensions of the standard energy value in S1 (the standard value is also set based on amplitude accumulation).
[0043] For example, when a violin part is assigned to three frequency points (2kHz, 2.5kHz, 3kHz) in a certain time window, its actual energy value is the algebraic sum of the amplitude moduli of the three frequencies. This value directly represents the relative energy contribution intensity of the part in the current segment, providing input data of the same scale for the deviation calculation of S3.
[0044] As shown in Figure 6, in one embodiment, obtaining the initial voice deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment in step S3 includes: S31, taking the absolute value of the difference between the actual energy value and the standard energy value of each standard voice part contained in the i-th audio segment, and obtaining the absolute energy difference of each standard voice part; S32, dividing the absolute energy difference of each standard voice part contained in the i-th audio segment by the standard energy value, and obtaining the energy difference ratio of each standard voice part; S33, dividing the sum of the energy difference ratios of all standard voice parts contained in the i-th audio segment by the total number of all standard voice parts contained in the i-th audio segment, and obtaining the initial voice deviation of the i-th audio segment.
[0045] In this embodiment, it should be noted that in S31, the actual energy value of each effective sound part (such as the marked electric bass sound part) provided in S2 is subtracted from the preset standard energy value, and the absolute value of the result is taken. The core of this operation is to eliminate the directional bias (both excessively strong and excessively weak are considered unbalanced), ensuring that positive and negative biases do not cancel each other out during subsequent aggregation. For example, in a certain segment, if the snare drum's actual energy exceeds the standard value by 20 units and the violin's is below the standard value by 15 units, the absolute difference between the two is recorded as 20 and 15 respectively, avoiding the risk of underestimating the overall balance problem due to opposite directions in existing methods.
[0046] In S32, to address the dimensional mismatch caused by differences in standard energy benchmarks for different vocal parts (e.g., a benchmark value of 500 for the bass drum and 80 for the hi-hat), the absolute difference obtained in S31 is divided by its corresponding standard energy value, transforming it into a dimensionless relative proportion. This proportion represents the relative imbalance intensity of that vocal part independent of other vocal parts (if the standard energy of the electric bass is 100, but the actual value is 120, then a proportion of 0.2 indicates a deviation of 20%). After normalization, the deviations of high-frequency, low-energy vocal parts (e.g., hi-hat) and low-frequency, high-energy vocal parts (e.g., bass drum) become comparable, preventing a 20-unit deviation in the bass drum from being misjudged as more severe than a 10-unit deviation in the hi-hat (the latter proportion may actually be higher), thus making subsequent overall statistics more fair.
[0047] In S33, the overall coordination imbalance of multiple voices within the current time window is quantified. The algebraic sum of the energy difference ratios of all effective voices calculated in S32 is taken (e.g., if the ratios of the four voices are 0.2 / 0.15 / 0.1 / 0.05, the sum is 0.5), and then divided by the total number of effective voices (4 in this example) to obtain the initial voice deviation value of 0.125. This average value reflects the average imbalance state of all active voices: if a single voice is severely imbalanced (e.g., a certain voice ratio is 0.5) but other voices are balanced (ratios close to 0), the average value will be significantly higher than 0; if all voices have small ratio deviations (e.g., all are 0.1), the average value will still be 0.1.
[0048] As shown in Figure 7, in one embodiment, step S3, correcting the initial voice deviation of the i-th audio segment based on the number of standard voice parts contained in the i-th audio segment and obtaining the target voice deviation of the i-th audio segment, includes: S34, obtaining a preset quantity threshold and a standard correction ratio, subtracting the preset quantity threshold from the number of standard voice parts contained in the i-th audio segment and dividing by the preset quantity threshold to obtain a quantity ratio; S35, if the quantity ratio is greater than zero, incrementing the quantity ratio by 1 to obtain an adjustment ratio, multiplying the adjustment ratio by the standard correction ratio to obtain a target correction ratio, incrementing the target correction ratio by 1 to obtain an increase ratio, multiplying the initial voice deviation of the i-th audio segment by the increase ratio to obtain the target voice deviation of the i-th audio segment; S36, if the quantity ratio is not greater than zero, using the initial voice deviation of the i-th audio segment as the target voice deviation of the i-th audio segment.
[0049] In this embodiment, it should be noted that in S34, first, a preset quantity threshold and a standard correction ratio are obtained. Among them, when taking the value of the preset quantity threshold, through a limited number of listening tests, the inflection point of the number of voices with a significantly enhanced sense of imbalance perception by the subjects is statistically counted, and the sudden increase point of the sensitivity of 80% of the subjects is taken as the threshold (usually 2.5 - 3.5, rounded to 3). Among them, when taking the value of the standard correction ratio, first, multiple groups of audio segments with an increasing number of voices are generated, and the same proportion of voice deviation (such as all voices +5%) is implanted in each group. Then, professional sound mixers are required to score the intensity of the imbalance perception of each group of segments. Finally, a growth curve is fitted with the number of voices as the horizontal axis and the perception intensity as the vertical axis. The standard correction ratio K is taken as the slope of the curve at the threshold point. When the number of voices increases from 4 to 8, the perception intensity increases from 3 points to 5 points, and the growth slope is (5 - 3) / (8 - 4)=0.5, so K≈0.5.
[0050] Furthermore, the evaluation sensitivity is dynamically adjusted according to the number of effective voices. The preset quantity threshold N (such as N = 3) is used as the turning point of the auditory tolerance. Calculate the proportion of the current number of effective voices M exceeding the threshold: (M - N) / N. For example, in the full orchestra passage, M = 7 and N = 3, then the proportion=(7 - 3) / 3≈1.33. This proportion is quantified as the amplification requirement of the voice density on the evaluation strictness. The larger the value, the more stringent the evaluation required due to the multi-voice superposition effect (even small deviations need to be amplified).
[0051] In S35, when the proportion calculated in S34 is greater than 0 (that is, the number of voices exceeds the threshold), sensitivity enhancement correction is performed. First, add 1 to the proportion value as the basic adjustment factor (in the above example, 1.33→2.33), then multiply it by the preset standard correction coefficient K (such as K = 0.6, set according to engineering experience) to obtain the target correction ratio (2.33×0.6≈1.4). Finally, add 1 to the target correction ratio to convert it into a multiplication factor (1.4 + 1 = 2.4). This design achieves a non-linear amplification effect: the more the number of voices exceeds the threshold, the faster the multiplication factor grows (such as when there are 5 voices, the factor is about 1.8, and when there are 7 voices, it reaches 2.4), simulating the exponential increase characteristic of the human ear's sensitivity to imbalance in complex sound mixing.
[0052] In S36, based on the multiplication factor generated in S35 (such as 2.4), a multiplication operation is performed on the initial deviation in S33 (the initial value 0.125→0.3 after calibration) to generate the target voice deviation. This value has dual characteristics: it not only retains the objective quantification of the absolute deviation by the initial deviation but also reflects the perception weight calibration through the voice number adaptive factor. If the number of voices does not exceed the threshold (such as M = 2 < N = 3), the initial value is directly output to avoid overcorrection. The final target deviation forms a time window sequence (such as [0.3, 0.28, 0.35…]), and its numerical magnitude is unified and perception-aligned, providing time-series data input for the trend analysis in S4.
[0053] As shown in Figure 8, in one embodiment, obtaining the audio signal's acoustic balance evaluation result based on the deviation difference in S4 includes: S41, analyzing the continuous change direction of the deviation difference between adjacent audio target acoustic parts; S42, when the deviation difference continuously shows the same direction of change, determining that the acoustic balance is in a unidirectional evolution state; S43, when the deviation difference alternately shows the opposite direction of change, determining that the acoustic balance is in a dynamic fluctuation state; S44, generating a balance evaluation result including a balance trend classification based on the duration of the unidirectional evolution state or the dynamic fluctuation state.
[0054] In this embodiment, it should be noted that in S41, a difference operation is performed on the deviation of the target voice in adjacent time windows to extract the sign sequence (positive / negative / zero). A positive difference indicates that the balance state of the current window has deteriorated compared to the previous window (e.g., an unexpected increase in the bass voice leads to an increase in the overall deviation by 0.1); a negative difference indicates an improvement in the state (e.g., turning down an overly loud snare drum reduces the deviation by 0.05). By analyzing the consistency of the signs of multiple consecutive differences (e.g., three consecutive positive differences [+,+,+]), the continuous evolution characteristics of the balance state are initially captured. This operation transforms the time series into a directional event chain, providing raw input for trend classification. For example, in the crescendo section of a symphony, the continuous increase in the string section may lead to 10 consecutive windows with positive differences, suggesting a risk of imbalance.
[0055] In S42, when multiple consecutive (e.g., ≥3) deviation difference values are in the same direction (all positive or all negative), the vocal balance is considered to have entered a unidirectional evolution state. A sustained positive difference value indicates a continued worsening of the imbalance (such as an undetected linear enhancement of the bass), while a sustained negative difference value reflects restorative adjustments (such as a mixing engineer gradually correcting a weak vocal). This state usually indicates an unintentional technical imbalance requiring manual intervention. For example, in a rock track, if the difference values are all positive for 15 consecutive time windows, an inspection reveals an abnormal gain of 2dB in the automatic control of the kick drum channel, which meets the criteria for "unidirectional deterioration."
[0056] In S43, when the deviation difference sequence frequently exhibits sign reversals (e.g., alternating between [+,-,+,-]), and the reversal frequency exceeds a certain preset threshold (e.g., ≥4 reversals within a 5-window period), it is determined to be a dynamic fluctuation state. This state characterizes the artistic oscillation of vocal energy (e.g., vocals weaken in the verse to highlight the guitar solo, vocals strengthen and return in the chorus), reflecting creative intent rather than technical defects. For example, the alternating prominence of the saxophone and piano in a jazz improvisation corresponds to the difference sequence [+0.3,-0.4,+0.2,-0.3].
[0057] In S44, based on the preliminary classification in S42 / S43, the final evaluation result is further output by combining the duration. For example, in the unidirectional evolution state, if the duration of continuous deterioration exceeds a certain threshold (e.g., 10 seconds), it is marked as a high-risk imbalance (e.g., a 12-second sustained bass enhancement may mask the main melody); if the continuous improvement exceeds the threshold, it is marked as an effective correction, but the rate of improvement still needs to be monitored. In the dynamic fluctuation state, high-frequency oscillations (e.g., 3 fluctuations within 2 seconds) are marked as artistic fluctuations and do not require repair; low-frequency long-period fluctuations (e.g., 30-second periodic changes) indicate structural design issues, and it is recommended to review whether they conform to the compositional intent.
[0058] A real-time sound balance evaluation system based on spectral characteristics is also provided. The system includes: an acquisition module, used to acquire multiple different standard sound parts, acquire the standard frequency band range and standard energy value corresponding to each standard sound part, acquire the audio signal to be evaluated, segment the audio signal to be evaluated into multiple segments by a preset time window, and acquire multiple frequency points and the amplitude value corresponding to each frequency point of each segment audio; a first data processing module, used to acquire the standard sound parts contained in the i-th segment audio based on the multiple frequency points of the i-th segment audio and the standard frequency band range corresponding to each standard sound part, and process the standard sound parts contained in the i-th segment audio according to the standard frequency band range corresponding to each standard sound part. The first data processing module obtains the actual energy value of each standard voice part contained in the i-th audio segment based on the frequency point and the amplitude value corresponding to each frequency point; the second data processing module obtains the initial voice part deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment, corrects the initial voice part deviation of the i-th audio segment based on the number of standard voice parts contained in the i-th audio segment, and obtains the target voice part deviation of the i-th audio segment; the third data processing module obtains the deviation difference between the target voice part deviations of adjacent audio segments, and obtains the voice part balance evaluation result of the audio signal based on the deviation difference.
[0059] In one embodiment, the first data processing module is further configured to: match the standard frequency band range corresponding to each standard voice part with all frequency points of the i-th segment audio based on each standard voice part; when there is a frequency point belonging to the i-th segment audio within the standard frequency band range of a certain standard voice part, determine that the frequency point belongs to the standard voice part, and mark the standard voice part containing the frequency point of the i-th segment audio as the standard voice part contained in the i-th segment audio.
[0060] In one embodiment, the first data processing module is further configured to: sum the amplitude values of all frequency points present in the j-th standard voice part contained in the i-th segment audio, and obtain the actual energy value of the j-th standard voice part contained in the i-th segment audio.
[0061] In one embodiment, the second data processing module is further configured to: take the absolute value of the difference between the actual energy value and the standard energy value of each standard voice part contained in the i-th audio segment, and obtain the absolute energy difference of each standard voice part; divide the absolute energy difference of each standard voice part contained in the i-th audio segment by the standard energy value, and obtain the energy difference ratio of each standard voice part; divide the sum of the energy difference ratios of all standard voice parts contained in the i-th audio segment by the total number of all standard voice parts contained in the i-th audio segment, and obtain the initial voice part deviation of the i-th audio segment.
[0062] In this embodiment, it should be noted that the specific method of performing the above-mentioned real-time sound balance evaluation system based on spectral characteristics has been described in detail in the embodiments of the real-time sound balance evaluation method based on spectral characteristics, and will not be elaborated here.
[0063] The preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings. However, the present disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of the present disclosure, various simple modifications can be made to the technical solutions of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0064] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0065] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. A real-time evaluation method for acoustic part balance based on spectral characteristics, characterized in that, include: Acquire multiple different standard voice parts, and obtain the standard frequency band range and standard energy value corresponding to each standard voice part. Acquire the audio signal to be evaluated, divide the audio signal to be evaluated into segments with a preset time window and acquire multiple segmented audio, and acquire multiple frequency points and the amplitude value corresponding to each frequency point of each segmented audio. Based on multiple frequency points of the i-th audio segment and the standard frequency band range corresponding to each standard voice part, the standard voice parts contained in the i-th audio segment are obtained. The actual energy value of each standard voice part contained in the i-th audio segment is obtained according to the frequency points of each standard voice part contained in the i-th audio segment and the amplitude value corresponding to each frequency point. The initial voice part deviation of the i-th audio segment is obtained according to the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment. The initial voice part deviation of the i-th audio segment is corrected according to the number of standard voice parts contained in the i-th audio segment to obtain the target voice part deviation of the i-th audio segment. The deviation difference between the target voice part deviations of adjacent audio segments is obtained, and the voice part balance evaluation result of the audio signal is obtained according to the deviation difference.
2. The real-time evaluation method for acoustic part balance based on spectral characteristics according to claim 1, characterized in that, The step of obtaining the standard sound parts contained in the i-th segment of audio based on multiple frequency points of the i-th segment and the standard frequency band range corresponding to each standard sound part includes: matching the standard frequency band range corresponding to each standard sound part with all frequency points of the i-th segment of audio; when there is a frequency point belonging to the i-th segment of audio within the standard frequency band range of a certain standard sound part, it is determined that the frequency point belongs to the standard sound part, and the standard sound part containing the frequency point of the i-th segment of audio is marked as the standard sound part contained in the i-th segment of audio.
3. The real-time evaluation method for acoustic part balance based on spectral characteristics according to claim 1, characterized in that, The step of obtaining the actual energy value of each standard sound part contained in the i-th audio segment based on the frequency points of each standard sound part contained in the i-th audio segment and the amplitude value corresponding to each frequency point includes: adding the amplitude values of all frequency points present in the j-th standard sound part contained in the i-th audio segment, and obtaining the actual energy value of the j-th standard sound part contained in the i-th audio segment.
4. The real-time evaluation method for acoustic part balance based on spectral characteristics according to claim 1, characterized in that, The step of obtaining the initial voice deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment includes: taking the absolute value of the difference between the actual energy value and the standard energy value of each standard voice part contained in the i-th audio segment to obtain the absolute energy difference of each standard voice part; dividing the absolute energy difference of each standard voice part contained in the i-th audio segment by the standard energy value to obtain the energy difference ratio of each standard voice part; and dividing the sum of the energy difference ratios of all standard voice parts contained in the i-th audio segment by the total number of all standard voice parts contained in the i-th audio segment to obtain the initial voice deviation of the i-th audio segment.
5. The real-time evaluation method for acoustic part balance based on spectral characteristics according to claim 1, characterized in that, The step of correcting the initial voice deviation of the i-th audio segment based on the number of standard voice parts contained in the i-th audio segment and obtaining the target voice deviation of the i-th audio segment includes: obtaining a preset quantity threshold and a standard correction ratio; subtracting the preset quantity threshold from the number of standard voice parts contained in the i-th audio segment and dividing by the preset quantity threshold to obtain a quantity ratio; if the quantity ratio is greater than zero, incrementing the quantity ratio by 1 to obtain an adjustment ratio; multiplying the adjustment ratio by the standard correction ratio to obtain a target correction ratio; incrementing the target correction ratio by 1 to obtain an increase ratio; multiplying the initial voice deviation of the i-th audio segment by the increase ratio to obtain the target voice deviation of the i-th audio segment; if the quantity ratio is not greater than zero, using the initial voice deviation of the i-th audio segment as the target voice deviation of the i-th audio segment.
6. The real-time evaluation method for acoustic part balance based on spectral characteristics according to claim 1, characterized in that, The method of obtaining the audio signal's acoustic balance assessment result based on the deviation difference includes: analyzing the continuous change direction of the deviation difference between adjacent audio target acoustic parts; when the deviation difference continuously changes in the same direction, it is determined that the acoustic balance is in a unidirectional evolution state; when the deviation difference alternately changes in opposite directions, it is determined that the acoustic balance is in a dynamic fluctuation state; and generating a balance assessment result including a balance trend classification based on the duration of the unidirectional evolution state or the dynamic fluctuation state.
7. A real-time sound part balance evaluation system based on spectral characteristics, characterized in that, The system includes: an acquisition module, used to acquire multiple different standard acoustic parts, acquire the standard frequency band range and standard energy value corresponding to each standard acoustic part, acquire the audio signal to be evaluated, segment the audio signal to be evaluated into multiple segments by a preset time window, and acquire multiple frequency points and the amplitude value corresponding to each frequency point of each segment audio; and a first data processing module, used to acquire the standard acoustic parts contained in the i-th segment audio based on the multiple frequency points of the i-th segment audio and the standard frequency band range corresponding to each standard acoustic part, and to acquire the standard acoustic parts contained in the i-th segment audio according to the frequency points and the amplitude value corresponding to each frequency point of each standard acoustic part. The corresponding amplitude value is used to obtain the actual energy value of each standard voice part contained in the i-th audio segment; the second data processing module is used to obtain the initial voice part deviation of the i-th audio segment based on the actual energy value and standard energy value of each standard voice part contained in the i-th audio segment, correct the initial voice part deviation of the i-th audio segment based on the number of standard voice parts contained in the i-th audio segment, and obtain the target voice part deviation of the i-th audio segment; the evaluation data processing module is used to obtain the deviation difference between the target voice part deviations of adjacent audio segments, and obtain the voice part balance evaluation result of the audio signal based on the deviation difference.
8. The real-time acoustic balance evaluation system based on spectral characteristics according to claim 7, characterized in that, The first data processing module is further configured to: match the standard frequency band range corresponding to each standard voice part with all frequency points of the i-th segment audio based on each standard voice part; when there is a frequency point belonging to the i-th segment audio within the standard frequency band range of a certain standard voice part, determine that the frequency point belongs to the standard voice part, and mark the standard voice part containing the frequency point of the i-th segment audio as the standard voice part contained in the i-th segment audio.
9. The real-time acoustic balance evaluation system based on spectral characteristics according to claim 7, characterized in that, The first data processing module is further configured to: sum the amplitude values of all frequency points in the j-th standard voice part contained in the i-th audio segment, and obtain the actual energy value of the j-th standard voice part contained in the i-th audio segment.
10. The real-time acoustic balance evaluation system based on spectral characteristics according to claim 7, characterized in that, The second data processing module is further configured to: take the absolute value of the difference between the actual energy value and the standard energy value of each standard voice part contained in the i-th segment audio, and obtain the absolute energy difference of each standard voice part; divide the absolute energy difference of each standard voice part contained in the i-th segment audio by the standard energy value, and obtain the energy difference ratio of each standard voice part. Divide the sum of the energy difference ratios of all standard voice parts contained in the i-th audio segment by the total number of standard voice parts contained in the i-th audio segment to obtain the initial voice part deviation of the i-th audio segment.