Sound dynamic gain optimization method and system based on audio characteristics

By constructing an audio feature temporal evolution chain and a gain response mode library, the problem that existing audio gain adjustment methods cannot be dynamically adjusted is solved, enabling precise audio output optimization of audio equipment and improving the audio quality of audio equipment.

CN122054045APending Publication Date: 2026-05-15SOUTHWEST UNIVERSITY FOR NATIONALITIES +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHWEST UNIVERSITY FOR NATIONALITIES
Filing Date
2026-02-09
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing audio gain adjustment methods cannot make real-time, dynamic adjustments based on changes in audio characteristics over continuous time, making it difficult to achieve accurate and efficient gain optimization when faced with complex and ever-changing audio signals, thus affecting the audio output quality of audio equipment.

Method used

By acquiring a set of audio signals to be optimized, performing feature temporal evolution analysis, constructing an audio feature temporal evolution chain, establishing a gain response pattern library, and refining gain response patterns based on historical audio gain optimization cases, a dynamic adaptation relationship is formed, a dynamic gain scheme is generated, and gain adjustment of audio equipment is realized.

Benefits of technology

It enables precise gain optimization based on real-time changes in audio signals, improving the audio output quality of audio equipment, and enhances the gain response mode library through a feedback mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122054045A_ABST
    Figure CN122054045A_ABST
Patent Text Reader

Abstract

The invention provides a sound dynamic gain optimization method and system based on audio characteristics, and relates to the technical field of audio processing, and the method comprises the steps: firstly obtaining a to-be-optimized audio signal set collected in different scenes, then carrying out the characteristic time sequence evolution analysis of the to-be-optimized audio signal set, and constructing an audio characteristic time sequence evolution chain; and extracting a gain response mode based on the historical audio gain optimization case set to form a gain response mode library. And establishing a dynamic adaptation relationship between the audio characteristic time sequence evolution chain and a gain adjustment sequence in the gain response mode library, wherein the relationship is updated along with the real-time change of the audio characteristic time sequence evolution chain. And generating a dynamic gain scheme according to the dynamic adaptation relationship, applying the dynamic gain scheme to gain adjustment of the sound equipment, executing gain optimization processing on an audio signal set, obtaining an audio output signal set after gain optimization, and feeding back characteristic time sequence evolution information to a gain response mode library to perfect contents, thereby realizing accurate optimization of the dynamic gain of the sound equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and more specifically, to a method and system for dynamic gain optimization of audio based on audio features. Background Technology

[0002] In the current field of audio processing, the methods for adjusting audio gain are mostly traditional and fixed. The common practice is to amplify or reduce the audio signal as a whole based on a preset fixed gain value. For example, in some simple audio playback systems, the audio gain is adjusted uniformly according to a preset volume level. This adjustment method does not take into account the dynamic changes in audio signals under different scenarios.

[0003] Audio signals acquired in different scenarios exhibit significant differences. For example, audio captured in a noisy outdoor environment differs drastically from audio captured in a quiet indoor environment. However, existing gain adjustment methods cannot adjust in real-time and dynamically based on changes in audio characteristics over continuous time. Furthermore, traditional gain optimization methods lack effective utilization of historical audio gain optimization cases, failing to extract targeted gain adjustment strategies from past experience. This makes it difficult to achieve accurate and efficient gain optimization when faced with complex and varied audio signals, thus affecting the audio output quality of audio equipment. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for optimizing dynamic gain of audio based on audio features, the method comprising: Obtain a set of audio signals to be optimized, which includes continuous audio signal units collected under different scenarios; The set of audio signals to be optimized is subjected to feature temporal evolution analysis, and an audio feature temporal evolution chain corresponding to the set of audio signals to be optimized is constructed. The audio feature temporal evolution chain includes the change sequence of audio features in the continuous time dimension and the correlation between features. Gain response patterns are extracted from a set of historical audio gain optimization cases to form a gain response pattern library, which contains gain adjustment sequences corresponding to the temporal evolution chains of different audio features. A dynamic adaptation relationship is established between the audio feature temporal evolution chain and the gain adjustment sequence in the gain response pattern library. The dynamic adaptation relationship is updated in real time with the changes in the audio feature temporal evolution chain. Based on the dynamic adaptation relationship, a dynamic gain scheme corresponding to the set of audio signals to be optimized is generated. The dynamic gain scheme is applied to the gain adjustment of the audio equipment. Gain optimization processing is performed on the set of audio signals to be optimized to obtain a set of audio output signals with optimized gain. The characteristic timing evolution information corresponding to the set of audio output signals is fed back to the gain response pattern library to improve the content of the pattern library.

[0005] Furthermore, embodiments of the present invention also provide an audio dynamic gain optimization system based on audio features, characterized in that it includes: A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to perform the above-described audio feature-based dynamic gain optimization method by executing the machine-executable instructions.

[0006] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, a processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described audio feature-based dynamic gain optimization method.

[0007] Based on the above, by acquiring sets of audio signals to be optimized under different scenarios, a comprehensive range of audio information in various complex environments is covered. Feature temporal evolution analysis is performed on the audio signal sets, and an audio feature temporal evolution chain is constructed. This deeply captures the changing sequences of audio features and the correlations between features in the continuous time dimension, enabling precise understanding of the dynamic characteristics of audio signals. Gain response patterns are extracted from a set of historical audio gain optimization cases, forming a gain response pattern library. This fully utilizes past experience and provides diverse reference strategies for audio gain adjustment. A dynamic adaptation relationship is established between the audio feature temporal evolution chain and the gain adjustment sequences in the gain response pattern library, and updated in real time as the audio feature temporal evolution chain changes, achieving dynamic optimization of gain adjustment strategies. Dynamic gain schemes are generated based on the dynamic adaptation relationship and applied to the gain adjustment of audio equipment. This enables precise gain optimization processing based on real-time changes in the audio signal, effectively improving the audio output quality of the audio equipment. Simultaneously, the feature temporal evolution information corresponding to the gain-optimized audio output signal set is fed back to the gain response pattern library, further enriching the content of the pattern library. Attached Figure Description

[0008] Figure 1 This is a schematic diagram of the execution flow of the audio dynamic gain optimization method based on audio features provided in the embodiments of the present invention.

[0009] Figure 2 This is a schematic diagram of exemplary hardware and software components of the audio dynamic gain optimization system based on audio features provided in an embodiment of the present invention. Detailed Implementation

[0010] The present invention will now be described in detail with reference to the accompanying drawings. Figure 1 This is a flowchart illustrating an audio dynamic gain optimization method based on audio features according to an embodiment of the present invention. The following is a detailed description of the audio dynamic gain optimization method based on audio features.

[0011] Step S110: Obtain the set of audio signals to be optimized, which includes continuous audio signal units collected under different scenarios.

[0012] Professional audio acquisition equipment, such as multi-channel microphones and high-fidelity recording equipment, is used to collect audio signals in various audio-generating scenarios. For example, in music performances, meetings, and natural environmental sound recording scenarios, audio signals are continuously collected at set time intervals. The collected audio signals are then organized chronologically to form a set of audio signals to be optimized. During the acquisition process, if audio signals containing personal privacy information, such as audio containing personal voice, are involved, privacy protection technologies are employed. For instance, homomorphic encryption algorithms are used to encrypt the audio data, ensuring that only authorized parties with the decryption key can access the original audio information during subsequent processing. Alternatively, audio desensitization technology is used to replace or obfuscate privacy-related features in the audio while retaining the main characteristic information of the audio to ensure that privacy data is not leaked.

[0013] Step S120: Perform feature temporal evolution analysis on the set of audio signals to be optimized, and construct the audio feature temporal evolution chain corresponding to the set of audio signals to be optimized. The audio feature temporal evolution chain includes the change sequence of audio features in the continuous time dimension and the correlation between features.

[0014] Step S121: Extract time-domain features for each audio signal unit in the set of audio signals to be optimized, and obtain the time-domain features of each audio signal unit. The time-domain features include the amplitude variation statistics of the audio signal unit in the time dimension. The amplitude variation statistics cover the average amplitude, amplitude fluctuation range and amplitude variation frequency.

[0015] Step S1211: Perform uniform sampling in the time dimension on each audio signal unit to obtain the amplitude value of each audio signal unit at multiple equally spaced sampling moments. The number of sampling moments is determined according to the duration of the audio signal unit and the preset sampling frequency.

[0016] For each audio signal unit, the number of sampling moments is determined based on the duration of the audio signal unit and the preset sampling frequency. For example, if the duration of the audio signal unit is t and the preset sampling frequency is f, then the number of sampling moments is t multiplied by f. Then, each sampling moment is determined at equal time intervals, and a high-precision analog-to-digital converter is used to collect the amplitude of the audio signal unit at these sampling moments to obtain the corresponding amplitude values. During the sampling process, the accuracy and stability of the sampling equipment must be ensured to guarantee that the collected amplitude values ​​accurately reflect the amplitude of the audio signal unit at that moment.

[0017] Step S1212: Calculate the arithmetic mean of the amplitude values ​​at multiple sampling times, and use the arithmetic mean as the average amplitude of the audio signal unit. The average amplitude reflects the overall amplitude level of the audio signal unit in the time dimension.

[0018] The amplitude values ​​obtained from multiple sampling moments in step S1211 are summed, and then the sum is divided by the number of sampling moments to obtain the average amplitude of the audio signal unit. By calculating the average amplitude, the overall amplitude of the audio signal unit in the time dimension can be understood.

[0019] Step S1213: Find the maximum and minimum values ​​of amplitude values ​​among multiple sampling times, calculate the difference between the maximum and minimum values, and use this difference as the amplitude fluctuation range of the audio signal unit. The amplitude fluctuation range reflects the amplitude change of the audio signal unit in the time dimension.

[0020] From the amplitude values ​​obtained at multiple sampling times in step S1211, the maximum and minimum values ​​are selected. Then, the minimum value is subtracted from the maximum value, and the difference is the amplitude fluctuation range of the audio signal unit. The amplitude fluctuation range reflects the amplitude change of the audio signal unit in the time dimension. By analyzing the amplitude fluctuation range, the dynamic changes of the audio signal can be understood, such as whether the audio signal is stable or has large fluctuations.

[0021] Step S1214: Analyze the trend of amplitude values ​​at multiple sampling times, identify the turning points where the amplitude values ​​change from rising to falling or from falling to rising, count the number of turning points per unit time, and use this number as the amplitude change frequency of the audio signal unit. The amplitude change frequency reflects how fast the amplitude of the audio signal unit changes in the time dimension.

[0022] The amplitude values ​​obtained from multiple sampling moments in step S1211 are arranged in chronological order, and then the changing trends of adjacent amplitude values ​​are analyzed sequentially. When the changing trend of two adjacent amplitude values ​​changes from rising to falling, or from falling to rising, this sampling moment is recorded as a turning point. Then, the number of turning points per unit time is counted; this number is the amplitude change frequency of the audio signal unit. The amplitude change frequency reflects how fast the amplitude of the audio signal unit changes in the time dimension. By analyzing the amplitude change frequency, the rhythmic change characteristics of the audio signal can be understood, such as whether the audio signal changes rhythmically quickly or slowly.

[0023] Step S1215: Construct a temporal feature association network, establish association edges between the temporal features of each audio signal unit and the temporal features of adjacent units, and determine the weight of the association edges based on feature similarity. The higher the feature similarity, the greater the weight.

[0024] The temporal features of each audio signal unit (including average amplitude, amplitude fluctuation range, and amplitude variation frequency) are used as node attributes in the network. Then, association edges are established between the temporal feature nodes of adjacent audio signal units. The weight of the association edge is determined by calculating the temporal feature similarity between two adjacent audio signal units. Various methods can be used to calculate feature similarity, such as calculating the cosine similarity between two temporal feature vectors or the reciprocal of their Euclidean distance. Higher feature similarity indicates greater similarity in the temporal features of the two audio signal units, resulting in a larger weight for the corresponding association edge. By constructing a temporal feature association network, the temporal features of audio signal units can be visualized and quantified along the time dimension.

[0025] Step S1216: Based on the temporal feature association network, extract the common patterns of temporal features across audio signal units, integrate the common patterns into the temporal feature representation of each audio signal unit, and optimize the representation accuracy of temporal features.

[0026] By analyzing temporal feature association networks, such as using graph mining algorithms like community detection and centrality analysis, common patterns in temporal features across audio signal units are extracted. These common patterns may include the trend of the average amplitude of audio signal units within a certain time range, the distribution pattern of amplitude fluctuation range, and the concentration interval of amplitude variation frequency. Then, these common patterns are integrated into the temporal feature representation of each audio signal unit. For example, by adjusting the temporal features of each audio signal unit to better conform to the extracted common patterns, the accuracy of the temporal feature representation is optimized. This allows the temporal features to more accurately reflect the essential characteristics of the audio signal, improving the accuracy of subsequent analysis.

[0027] Step S1217: Integrate the optimized average amplitude, amplitude fluctuation range, and amplitude change frequency into the time domain characteristics of the audio signal unit.

[0028] The optimized average amplitude, amplitude fluctuation range, and amplitude change frequency from step S1216 are integrated to form the time-domain characteristics of the audio signal unit. The integrated time-domain characteristics contain statistical information on the amplitude changes of the audio signal unit in the time dimension, and have been optimized for common patterns, thus more accurately describing the time-domain characteristics of the audio signal unit.

[0029] Step S1218: Based on the association relationship of the temporal feature association network, the temporal features of different audio signal units are collaboratively optimized so that the temporal features of adjacent audio signal units maintain consistency in the trend of change.

[0030] Based on the weights and relationships of the edges in the temporal feature association network, the temporal features of different audio signal units are collaboratively optimized. For example, for adjacent audio signal units with large edge weights, their temporal features are adjusted to ensure that the trends of their average amplitude, amplitude fluctuation range, and amplitude variation frequency are consistent. Specific adjustment methods can be determined based on the edge weights and feature similarity. For instance, when the temporal feature trends of two adjacent audio signal units are inconsistent, the temporal features of one or two audio signal units are fine-tuned based on their feature similarity and edge weights to gradually bring their trends closer together. Through collaborative optimization, the temporal features of the entire audio signal set can change more smoothly and reasonably over time, improving the reliability and consistency of the temporal features.

[0031] Step S122: Extract frequency domain features for each audio signal unit in the set of audio signals to be optimized, and obtain the frequency domain features of each audio signal unit. The frequency domain features include the energy distribution information of the audio signal unit in the frequency dimension. The energy distribution information covers the total energy in different frequency ranges and the proportion of energy in each frequency range to the total energy.

[0032] Signal processing algorithms such as Fourier transform are used to transform each audio signal unit from the time domain to the frequency domain. In the frequency domain, the frequency range is divided into multiple frequency intervals, for example, according to equal frequency intervals or the frequency characteristics of the audio signal. Then, the total energy within each frequency interval and the proportion of energy in each frequency interval to the total energy of the audio signal unit are calculated. The total energy can be calculated by summing the squares of the amplitudes of the spectral components within that frequency interval, or by using other energy calculation methods. Through frequency domain feature extraction, the energy distribution of the audio signal unit along the frequency dimension can be understood.

[0033] Step S123: Sort the time-domain features of all audio signal units in chronological order to form a time-domain feature sequence. At the same time, sort the frequency-domain features of all audio signal units in chronological order to form a frequency-domain feature sequence.

[0034] Based on the acquisition time sequence of the audio signal units, the time-domain features of all audio signal units are arranged chronologically to form a time-domain feature sequence. Similarly, the frequency-domain features of all audio signal units are also arranged chronologically to form a frequency-domain feature sequence. During the sorting process, the accuracy of the time order must be ensured to guarantee that both the time-domain and frequency-domain feature sequences accurately reflect the temporal changes of the audio signal.

[0035] Step S124: Perform adjacent feature correlation analysis on the time-domain feature time series, calculate the amplitude change range and trend between the time-domain features of two adjacent audio signal units, wherein the amplitude change range is the sum of the absolute values ​​of the relative difference of the average amplitude, the relative difference of the amplitude fluctuation range, and the relative difference of the amplitude change frequency in the adjacent time-domain features, wherein the relative difference is the ratio of the difference to the corresponding feature value of the previous audio signal unit, and the trend is the direction of increase or decrease of each statistical information in the adjacent time-domain features.

[0036] For the temporal characteristics of two adjacent audio signal units in a time-domain feature sequence, the relative differences in their average amplitude, amplitude fluctuation range, and amplitude change frequency are calculated. The relative difference is calculated by subtracting the characteristic value of the preceding audio signal unit from the characteristic value of the following audio signal unit, and then dividing the resulting difference by the characteristic value of the preceding audio signal unit. The absolute values ​​of these three relative differences are then added together to obtain the amplitude change magnitude. Simultaneously, the direction of increase or decrease of various statistical information (average amplitude, amplitude fluctuation range, and amplitude change frequency) in the temporal characteristics of two adjacent audio signal units is analyzed. For example, is the average amplitude increasing or decreasing, is the amplitude fluctuation range expanding or shrinking, and is the amplitude change frequency accelerating or slowing down? These directions of increase or decrease are used to describe the trend of change. Through the correlation analysis of adjacent features, the changes in temporal characteristics along the time dimension can be understood.

[0037] Step S125: Perform adjacent feature correlation analysis on the frequency domain feature time sequence, calculate the energy transfer ratio and transfer direction between the frequency domain features of two adjacent audio signal units, wherein the energy transfer ratio is the ratio of the difference of the total energy in the same frequency interval in adjacent frequency domain features to the total energy in the same frequency interval of the previous audio signal unit, and the transfer direction is the direction of increase or decrease of energy in each frequency interval in adjacent frequency domain features.

[0038] Step S1251: Select two adjacent audio signal units in the frequency domain feature time sequence, denoted as the preceding audio signal unit and the following audio signal unit, respectively, and extract the frequency domain features of the preceding audio signal unit and the following audio signal unit.

[0039] In the frequency domain characteristic time sequence, two adjacent audio signal units are selected, and the former is determined as the preceding audio signal unit and the latter as the following audio signal unit. Then, the frequency domain characteristics of the preceding and following audio signal units are extracted, including information such as the total energy and energy ratio of each frequency interval.

[0040] Step S1252: Determine the common frequency interval in the frequency domain features of the two audio signal units. Compare the frequency intervals in the frequency domain features of the two audio signal units and select the frequency intervals that exist in both frequency domain features as the common frequency interval. If any frequency interval appears only in the frequency domain features of one of the audio signal units, then the frequency interval is regarded as a frequency interval with zero energy and included in the analysis.

[0041] By comparing the frequency ranges in the frequency domain characteristics of the preceding and following audio signal units, frequency ranges that exist simultaneously in both frequency domain characteristics are identified; these frequency ranges are the common frequency ranges. Frequency ranges that appear only in the frequency domain characteristics of one audio signal unit are considered to have zero energy and are also included in the subsequent analysis. This ensures that the frequency ranges of the two audio signal units correspond when analyzing the energy transfer ratio and direction, improving the accuracy of the analysis.

[0042] Step S1253: For each common frequency interval, calculate the difference between the total energy of the subsequent audio signal unit in that frequency interval and the total energy of the preceding audio signal unit in that frequency interval, and obtain the energy difference of that frequency interval.

[0043] For each common frequency range, the energy difference is calculated by subtracting the total energy of the preceding audio signal units in that frequency range from the total energy of the subsequent audio signal units. By calculating this energy difference, we can understand how the energy changes within that frequency range.

[0044] Step S1254: Calculate the ratio of the energy difference of each frequency interval to the total energy of the preceding audio signal unit in that frequency interval, and use this ratio as the energy transfer ratio of that frequency interval. If the total energy of the preceding audio signal unit in that frequency interval is zero, then set the energy transfer ratio of that frequency interval according to the preset rules.

[0045] The energy transfer ratio for a given frequency range is calculated by dividing the energy difference of each frequency range by the total energy of the preceding audio signal units within that range. If the total energy of the preceding audio signal units in that frequency range is zero, the energy transfer ratio can be set according to a preset rule. For example, it can be set to a fixed value or estimated based on the energy transfer ratios of other frequency ranges. The energy transfer ratio reflects the degree of energy transfer within that frequency range.

[0046] Step S1255: Determine whether the energy difference value of each frequency interval is positive or negative. If the energy difference value is positive, the energy transfer direction of the frequency interval is increasing from front to back. If the energy difference value is negative, the energy transfer direction of the frequency interval is decreasing from front to back. If the energy difference value is zero, the energy transfer direction of the frequency interval remains unchanged.

[0047] By determining the sign of the energy difference in each frequency interval, the direction of energy transfer within that interval can be identified. A positive energy difference indicates that the subsequent audio signal unit has more energy than the preceding audio signal unit in that frequency interval, and the energy transfer direction is from front to back increasing. A negative energy difference indicates that the subsequent audio signal unit has less energy than the preceding audio signal unit in that frequency interval, and the energy transfer direction is from front to back decreasing. A zero energy difference indicates that the energy in that frequency interval remains unchanged between the two audio signal units, and the energy transfer direction is unchanged. By determining the direction of energy transfer, the changing trend of frequency domain characteristics across frequency intervals can be understood.

[0048] Step S1256: Construct a frequency domain feature transfer map, using the energy transfer ratio and transfer direction of each frequency interval as map node attributes. The connection relationship between nodes is determined based on the adjacency of frequency intervals, and connection edges are established between nodes in adjacent frequency intervals.

[0049] Each frequency range is treated as a node in the spectrum, with node attributes including the energy transfer ratio and direction within that frequency range. Then, based on the adjacency of frequency ranges, connections are established between nodes in adjacent frequency ranges. For example, nodes in adjacent frequency ranges on the frequency axis are connected. By constructing a frequency domain feature transfer spectrum, the changes in frequency domain features can be visualized and networked, facilitating subsequent analysis and processing.

[0050] Step S1257: Based on the frequency domain feature transfer spectrum, identify the dominant frequency range of energy transfer. The dominant frequency range is the frequency range with the largest absolute value of energy transfer ratio. Use the energy transfer characteristics of the dominant frequency range as a reference for the correlation of adjacent frequency domain features.

[0051] By comparing the absolute values ​​of the energy transfer ratios at each node in the frequency domain feature transfer map, the frequency interval with the largest absolute value is identified as the dominant frequency interval for energy transfer. Then, the energy transfer characteristics of the dominant frequency interval, including the energy transfer ratio and direction, are used as a reference for correlating adjacent frequency domain features. By identifying the dominant frequency interval, we can focus on the most significant parts of the frequency domain feature changes, improving the efficiency and accuracy of the analysis.

[0052] Step S1258: Integrate the energy transfer ratio, transfer direction and dominant frequency range identifier of all frequency ranges to form the correlation information between the frequency domain characteristics of two adjacent audio signal units. Each pair of adjacent audio signal units corresponds to a set of the above correlation information.

[0053] The energy transfer ratio, transfer direction, and dominant frequency range identifier for each frequency range are integrated to form a set of correlation information. This set of correlation information corresponds to the relationship between the frequency domain characteristics of two adjacent audio signal units. By integrating this information, the changes between adjacent frequency domain characteristics can be comprehensively described.

[0054] Step S1259: Repeat the above steps for all adjacent audio signal units in the frequency domain feature time sequence to obtain a set of adjacent frequency domain feature association information covering the entire frequency domain feature time sequence.

[0055] For all adjacent audio signal units in the frequency domain feature time sequence, steps S1251 to S1258 are performed to obtain the frequency domain feature association information of each pair of adjacent audio signal units. Then, the above association information is collected to form a set of adjacent frequency domain feature association information covering the entire frequency domain feature time sequence, thereby providing a comprehensive understanding of the changes in frequency domain features over time.

[0056] Step S126: Extract the feature mutation points in the time-domain feature time series. The feature mutation points are the time-domain features corresponding to the time points in the time-domain features where the amplitude change exceeds a preset amplitude threshold. At the same time, extract the feature mutation points in the frequency-domain feature time series. The feature mutation points are the frequency-domain features corresponding to the time points in the frequency-domain features where the energy transfer ratio exceeds a preset ratio threshold.

[0057] For time-domain feature time series, the amplitude change of the audio signal unit at each time point is calculated (according to the method in step S124), and then compared with a preset amplitude threshold. If the amplitude change exceeds the preset amplitude threshold, then the time-domain feature corresponding to that time point is a feature mutation point. Similarly, for frequency-domain feature time series, the energy transfer ratio of the audio signal unit at each time point is calculated (according to the method in step S125), and then compared with a preset ratio threshold. If the energy transfer ratio exceeds the preset ratio threshold, then the frequency-domain feature corresponding to that time point is a feature mutation point. By extracting feature mutation points, significant changes in audio features in the time dimension can be identified.

[0058] Step S127: Construct a feature association weight matrix, using each feature parameter in the time-domain feature time series and the frequency-domain feature time series as matrix row vectors, and the correlation strength between features as matrix element values, thereby strengthening the intrinsic correlation between features through matrix operations.

[0059] Each feature parameter (such as the average amplitude, amplitude fluctuation range, and amplitude change frequency in the time-domain feature time series, and the total energy and energy percentage of each frequency interval in the frequency-domain feature) in the time-domain feature time series is used as a row vector of a matrix. Then, the correlation strength between every two feature parameters is calculated, using methods such as correlation coefficient calculation and mutual information calculation. The calculated correlation strength is used as the element values ​​of the matrix to construct a feature correlation weight matrix. Finally, matrix operations, such as matrix multiplication and matrix decomposition, are used to process the feature correlation weight matrix to strengthen the intrinsic correlation between features.

[0060] Step S128: Based on the feature association weight matrix, the adjacent time-domain feature association information, the adjacent frequency-domain feature association information, and the two types of feature mutation points are integrated into the matrix association relationship and connected in series along the time dimension to form a continuous audio feature time-series evolution chain. Each time node in the audio feature time-series evolution chain contains the time-domain features, frequency-domain features, and association information with adjacent nodes of the corresponding audio signal unit.

[0061] Based on the correlation relationships in the feature correlation weight matrix, the correlation information of adjacent time-domain features (the information obtained in step S124), the correlation information of adjacent frequency-domain features (the information obtained in step S125), and the two types of feature mutation points (the information obtained in step S126) are integrated into the correlation relationships of the matrix. Then, in the order of the time dimension, the time-domain features, frequency-domain features, and correlation information with adjacent nodes of the audio signal unit corresponding to each time node are concatenated to form a continuous audio feature temporal evolution chain. In this evolution chain, each time node contains rich feature information and correlation information, which can comprehensively reflect the evolution process of audio features in the time dimension.

[0062] Step S130: Extract gain response patterns based on the historical audio gain optimization case set to form a gain response pattern library, which contains gain adjustment sequences corresponding to different audio feature temporal evolution chains.

[0063] Step S131: Obtain a set of historical audio gain optimization cases. Each historical audio gain optimization case includes a set of historical audio signals to be optimized, a temporal evolution chain of audio features corresponding to the set of historical audio signals to be optimized, a gain adjustment sequence applied to the historical audio gain optimization case, and a set of historical audio output signals after gain optimization.

[0064] We collected past cases of audio gain optimization processing. Each case includes a set of historical audio signals to be optimized, the corresponding audio feature time-series evolution chain, the gain adjustment sequence used, and a set of optimized historical audio output signals. These cases can come from different audio processing scenarios, such as music production, audio broadcasting, and audio playback. During the collection process, we ensured the completeness and accuracy of the cases to facilitate subsequent analysis and refinement.

[0065] Step S132: Perform feature analysis on the audio feature temporal evolution chain in each historical audio gain optimization case, extract the temporal subsequence of temporal features, the temporal subsequence of frequency features, the correlation parameters of adjacent temporal features, the correlation parameters of adjacent frequency features, and the distribution information of temporal and frequency feature mutation points in each historical audio feature temporal evolution chain, and integrate them into historical feature evolution elements.

[0066] For each historical audio gain optimization case, the audio feature temporal evolution chain is processed according to steps S121 to S128. This process extracts the temporal subsequence of temporal features, the temporal subsequence of frequency features, the correlation parameters between adjacent temporal features (such as amplitude variation and trend), the correlation parameters between adjacent frequency features (such as energy transfer ratio and direction), and the distribution information of temporal and frequency feature abrupt change points. Then, the extracted information is integrated to form the historical feature evolution elements of that historical audio gain optimization case.

[0067] Step S133: Perform similarity clustering on the historical feature evolution elements of all historical audio gain optimization cases, calculate the similarity between any two historical feature evolution elements. The similarity calculation covers the similarity of time-domain feature time-series subsequences, the similarity of frequency-domain feature time-series subsequences, the similarity of correlation parameters, and the similarity of mutation point distribution. The similarity of each part is weighted and summed according to preset weights to obtain the comprehensive similarity.

[0068] For all historical audio gain optimization cases, the similarity between any two historical feature evolution elements is calculated. The similarity calculation includes several aspects: the similarity of time-domain feature time-series subsequences can be determined by calculating the dynamic time warping distance and the longest common subsequence length between two subsequences; the similarity of frequency-domain feature time-series subsequences can be determined by calculating the energy distribution similarity of two subsequences in the frequency range; the similarity of correlation parameters can be determined by calculating the cosine similarity between two correlation parameter vectors; and the similarity of abrupt change point distributions can be determined by calculating the distance or overlap between two abrupt change point distributions. Then, the similarities from these different aspects are weighted and summed according to pre-defined weights to obtain the comprehensive similarity. Through similarity clustering, historical cases with similar evolutionary patterns can be grouped into one category, facilitating subsequent pattern extraction.

[0069] Step S134: Group historical feature evolution elements with a comprehensive similarity greater than the similarity threshold into the same feature evolution category. Each feature evolution category contains multiple historical feature evolution elements with similar feature change patterns.

[0070] Based on a pre-defined similarity threshold, historical feature evolution elements with a comprehensive similarity greater than the threshold are grouped into the same feature evolution category. Historical feature evolution elements within each category exhibit similar feature change patterns, such as similar temporal feature change trends, frequency domain feature energy distribution change patterns, correlation parameter change patterns, and mutation point distribution. Through clustering, historical cases can be classified according to their feature evolution patterns.

[0071] Step S135: Extract patterns from the historical gain adjustment sequences in each feature evolution category, analyze the correlation between the historical gain adjustment sequences and the corresponding historical feature evolution elements under the feature evolution category, and determine the correspondence rules between audio feature changes and gain adjustments in the feature evolution category. The correspondence rules include the gain adjustment amount corresponding to the feature change amplitude, the gain adjustment direction corresponding to the feature change trend, and the gain adjustment timing corresponding to the feature mutation point.

[0072] For example, step S1351: Select a feature evolution category and extract the historical feature evolution elements and corresponding historical gain adjustment sequences from all historical audio gain optimization cases under that feature evolution category.

[0073] From each feature evolution category, select all historical audio gain optimization cases within that category, and extract the historical feature evolution elements and corresponding historical gain adjustment sequences for each case. The historical gain adjustment sequence contains the specific parameters for gain adjustment of the audio signal in that case, such as the gain adjustment amount, adjustment direction, and adjustment time.

[0074] Step S1352: Extract feature change parameters from the historical feature evolution elements of each historical audio gain optimization case. The feature change parameters include the temporal feature change magnitude, temporal feature change trend, frequency domain feature energy transfer ratio, frequency domain feature energy transfer direction, and the time of occurrence of the time domain and frequency domain feature abrupt change points.

[0075] For each historical audio gain optimization case, the historical feature evolution elements are extracted, including feature change parameters. The temporal feature change amplitude is calculated according to the method in step S124, and the temporal feature change trend is determined based on the increase / decrease direction analyzed in step S124. The frequency domain feature energy transfer ratio is calculated according to the method in step S125, and the frequency domain feature energy transfer direction is determined based on the increase / decrease direction determined in step S125. The occurrence time of the abrupt change points in the temporal and frequency domain features is determined based on the time points of the abrupt change points extracted in step S126. By extracting the feature change parameters, key change information in the historical feature evolution elements can be extracted.

[0076] Step S1353: Extract adjustment parameters from the historical gain adjustment sequence for each historical audio gain optimization case. The adjustment parameters include the gain adjustment amount, gain adjustment direction, gain adjustment start time, and gain adjustment duration.

[0077] For each historical audio gain optimization case, extract the adjustment parameters from the historical gain adjustment sequence. The gain adjustment amount refers to the specific value of amplifying or attenuating the audio signal; the gain adjustment direction refers to whether the gain adjustment is amplification or attenuation; the gain adjustment start time refers to the time point at which the gain adjustment begins; and the gain adjustment duration refers to the duration of the gain adjustment.

[0078] Step S1354: Establish a correspondence table between feature change parameters and adjustment parameters under the feature evolution category. Each historical audio gain optimization case corresponds to a row in the correspondence table, which includes the feature change parameter combination and the corresponding adjustment parameter combination for that historical audio gain optimization case.

[0079] The characteristic change parameter combinations and corresponding adjustment parameter combinations for each historical audio gain optimization case are compiled into a correspondence table. In the correspondence table, each row corresponds to a historical case, which includes the characteristic change parameter combination of the case (such as the combination of time domain characteristic change amplitude, time domain characteristic change trend, frequency domain characteristic energy transfer ratio, frequency domain characteristic energy transfer direction, and the time of occurrence of the change point) and the corresponding adjustment parameter combination (such as the combination of gain adjustment amount, gain adjustment direction, gain adjustment start time, and gain adjustment duration).

[0080] Step S1355: Construct a parameter correlation model, using the combination of feature change parameters as model input and the combination of adjustment parameters as model output, and use the parameter correlation model to train and explore the nonlinear correlation between feature change parameters and adjustment parameters.

[0081] Choose a suitable machine learning model, such as a neural network model, decision tree model, or support vector machine model, and construct a parameter correlation model. Use the combination of feature variation parameters as the model input and the combination of adjustment parameters as the model output. Train the model using data from a historical case mapping table. During training, the model learns the non-linear relationship between feature variation parameters and adjustment parameters; for example, different combinations of time-domain feature variation magnitudes and frequency-domain feature energy transfer ratios will correspond to different gain adjustment amounts and directions.

[0082] Step S1356: Based on the parameter association model, perform statistical analysis on the data in the corresponding relationship table, calculate the frequency of occurrence of the adjustment parameter combination corresponding to the same or similar feature change parameter combination, and select the adjustment parameter combination with the highest frequency of occurrence as the optimal adjustment parameter combination corresponding to the feature change parameter combination.

[0083] Using a trained parameter association model, statistical analysis is performed on the data in the correspondence table. For parameter combinations with the same or similar feature changes, the frequency of their corresponding adjustment parameter combinations is calculated. Then, the adjustment parameter combination with the highest frequency is selected as the optimal adjustment parameter combination for that feature change parameter combination. This determines the most frequently used gain adjustment method under that feature change condition.

[0084] Step S1357: Analyze the correlation between each parameter in the characteristic change parameters and each parameter in the adjustment parameters, and determine the correlation between the amplitude of time-domain characteristic change and the amount of gain adjustment, the correlation between the trend of time-domain characteristic change and the direction of gain adjustment, the correlation between the energy transfer ratio of frequency-domain characteristics and the amount of gain adjustment, the correlation between the energy transfer direction of frequency-domain characteristics and the direction of gain adjustment, and the correlation between the time of occurrence of characteristic mutation point and the time of start of gain adjustment.

[0085] Correlation analysis is performed on each parameter in the characteristic variation parameters and each parameter in the adjustment parameters, such as calculating the Pearson correlation coefficient and Spearman correlation coefficient. Through correlation analysis, the relationship between the magnitude of time-domain characteristic variation and the amount of gain adjustment is determined; for example, does a larger magnitude of time-domain characteristic variation necessarily lead to a larger gain adjustment? The relationship between the trend of time-domain characteristic variation and the direction of gain adjustment is also determined; for example, if the trend of time-domain characteristic variation is increasing, does the direction of gain adjustment indicate amplification? Furthermore, the relationship between the proportion of frequency-domain characteristic energy transfer and the amount of gain adjustment, the direction of frequency-domain characteristic energy transfer and the direction of gain adjustment, and the relationship between the time of occurrence of the characteristic abrupt change point and the time of start of gain adjustment are also determined.

[0086] Step S1358: Based on the statistical analysis results, the correlation and correlation analysis results mined by the parameter association model, a gain adjustment rule for the feature evolution category is generated. The gain adjustment rule includes: the range of gain adjustment values ​​when the time-domain feature change amplitude is any range; the gain adjustment direction when the time-domain feature change trend is increasing or decreasing; the gain adjustment correction coefficient when the frequency domain feature energy transfer ratio is any value; the basis for correcting the gain adjustment direction when the frequency domain feature energy transfer direction is increasing or decreasing; and the rule for setting the start time of gain adjustment when a feature mutation point occurs.

[0087] Based on the combined results of statistical analysis (results from step S1356), correlation mining from parameter association models (results from step S1355), and correlation analysis (results from step S1357), gain adjustment rules for this feature evolution category are generated. For example, based on different ranges of temporal feature variation amplitude, the corresponding gain adjustment range is determined; based on whether the temporal feature variation trend is increasing or decreasing, the corresponding gain adjustment direction (amplification or attenuation) is determined; based on different values ​​of the frequency domain feature energy transfer ratio, the corresponding gain adjustment correction coefficient is determined to correct the basic gain adjustment; based on whether the frequency domain feature energy transfer direction is increasing or decreasing, the corresponding gain adjustment direction correction basis is determined to correct the basic gain adjustment direction; and based on the occurrence time of the feature mutation point, the rule for setting the corresponding gain adjustment start time is determined, for example, starting gain adjustment at a certain time point after the mutation point occurs.

[0088] Step S1359: Construct a rule adaptability model, dynamically associate the gain adjustment rule with the historical feature change parameters under the feature evolution category, and calculate the adaptability coverage of the gain adjustment rule for different feature change scenarios.

[0089] A rule-based adaptability model is constructed to dynamically correlate the generated gain adjustment rule with historical feature change parameters under the feature evolution category. The model takes historical feature change parameters as input and outputs the adaptability of the gain adjustment rule to that feature change scenario. Then, the adaptability coverage of the gain adjustment rule for different feature change scenarios is calculated. Adaptability coverage refers to the proportion of feature change scenarios that the gain adjustment rule can adapt to out of the total number of feature change scenarios. By calculating the adaptability coverage, the applicability and effectiveness of the gain adjustment rule can be understood.

[0090] Step S13510: Based on the rule adaptability model, dynamically optimize the gain adjustment rule to expand the adaptability of the gain adjustment rule to scenarios with changing edge features.

[0091] Based on the calculation results of the rule-based adaptability model, the gain adjustment rules are dynamically optimized. For example, for edge feature variation scenarios with low adaptability coverage, the feature variation parameters and corresponding adjustment parameters of these scenarios are analyzed, and relevant parameters in the gain adjustment rules, such as the value range of the gain adjustment amount and the correction coefficient, are adjusted to expand the adaptability of the gain adjustment rules to the aforementioned edge scenarios. Through dynamic optimization, the applicability and accuracy of the gain adjustment rules can be improved.

[0092] Step S13511: Perform the above steps for all feature evolution categories so that each feature evolution category has a corresponding adaptive gain adjustment rule.

[0093] For each feature evolution category, steps S1351 to S13510 are performed to ensure that each feature evolution category has a corresponding gain adjustment rule with strong adaptability. This provides targeted gain adjustment guidance for audio signals with different feature evolution patterns, improving the effect of gain optimization.

[0094] Step S136: Combine each feature evolution category and its corresponding gain adjustment rule into a gain response pattern. Each gain response pattern includes a feature evolution category identifier, a description of the feature evolution law, and a corresponding set of gain adjustment rules.

[0095] Each feature evolution category and its corresponding gain adjustment rule are combined to form a gain response pattern. The gain response pattern includes a feature evolution category identifier to uniquely identify the feature evolution category; a feature evolution law description to describe the changes in audio features within that category; and a corresponding set of gain adjustment rules to guide gain adjustment for audio signals of that category. By combining gain response patterns, feature evolution laws and gain adjustment rules can be associated, facilitating subsequent lookup and application.

[0096] Step S137: Based on the correlation strength between feature evolution category and gain adjustment rule, prioritize all gain response modes to form an ordered queue of gain response modes.

[0097] The correlation strength between each feature evolution category and its corresponding gain adjustment rule is calculated. This correlation strength can be determined by analyzing the effects of using the gain adjustment rule on audio signals of that category in historical cases, such as the degree of audio quality improvement after gain optimization and the accuracy of the adjustment. Then, all gain response modes are prioritized according to their correlation strength, forming an ordered queue of gain response modes. In subsequent applications, higher-priority gain response modes will be considered first to improve the efficiency and effectiveness of gain optimization.

[0098] Step S138: Enter the ordered gain response pattern queue into the initial pattern library, establish a feature matching index for the gain response pattern, and the index keywords include key feature parameters in the feature evolution category identifier and adjustment parameters in the gain adjustment rule set.

[0099] An ordered queue of gain response patterns is entered into an initial gain response pattern library. Then, a feature matching index for the gain response patterns is established. The index keywords include key feature parameters from the feature evolution category identifier, such as the range of temporal feature variation amplitude and the range of frequency domain feature energy transfer ratio; and adjustment parameters from the gain adjustment rule set, such as the range of gain adjustment values ​​and the direction of gain adjustment. By establishing the feature matching index, gain response patterns that match the feature evolution patterns of the audio signal to be optimized can be quickly retrieved, improving the efficiency of gain optimization.

[0100] Step S139: Dynamically optimize the merged pattern library. When a new historical audio gain optimization case is added, process the new historical audio gain optimization case according to the above steps, extract the new gain response pattern, and if the new gain response pattern is different from the existing gain response patterns in the gain response pattern library, add the new gain response pattern to the gain response pattern library and update the feature matching index to achieve dynamic improvement of the gain response pattern library.

[0101] When a new historical audio gain optimization case is available, it is processed according to steps S131 to S138 to extract a new gain response pattern. Then, the new gain response pattern is compared with existing gain response patterns in the gain response pattern library. If they differ, the new gain response pattern is added to the library, and the feature matching index is updated. By dynamically optimizing the gain response pattern library, its content can be continuously enriched and improved, enhancing the applicability and accuracy of gain response patterns to adapt to different audio feature evolution patterns.

[0102] Step S140: Establish a dynamic adaptation relationship between the audio feature temporal evolution chain and the gain adjustment sequence in the gain response pattern library. The dynamic adaptation relationship is updated in real time with the changes in the audio feature temporal evolution chain.

[0103] Step S141: Extract key evolution nodes in the audio feature temporal evolution chain. The key evolution nodes are the feature information and related information corresponding to the feature mutation points and time nodes in the audio feature temporal evolution chain where the feature change amplitude exceeds the preset feature change amplitude threshold. Each key evolution node contains time domain features, frequency domain features, related parameters with the previous node, and feature change type identifier.

[0104] The temporal evolution chain of audio features is analyzed to extract key evolution nodes. Key evolution nodes include feature mutation points (determined according to the method in step S126) and time nodes where the feature change amplitude exceeds a preset feature change amplitude threshold. For each key evolution node, its corresponding temporal features, frequency domain features, correlation parameters with the previous node (such as correlation parameters between adjacent temporal features and adjacent frequency domain features), and feature change type identifiers (such as the type of feature mutation and the type of feature change amplitude) are extracted.

[0105] Step S142: Extract typical evolution nodes from the feature evolution category corresponding to each gain response pattern from the gain response pattern library. The typical evolution nodes are representative feature mutation points and feature change amplitudes exceeding the preset typical feature change amplitude thresholds in the historical feature evolution elements under the feature evolution category. Each typical evolution node includes a time domain feature template, a frequency domain feature template, an association parameter template, and a feature change type template.

[0106] From the feature evolution category corresponding to each gain response pattern in the gain response pattern library, typical evolution nodes are extracted. Typical evolution nodes are representative feature mutation points and nodes whose feature change amplitude exceeds a preset typical feature change amplitude threshold among the historical feature evolution elements under that feature evolution category. For each typical evolution node, its time-domain feature template, frequency-domain feature template, correlation parameter template, and feature change type template are extracted. The time-domain feature template is a standard description of the time-domain features of the typical evolution node; the frequency-domain feature template is a standard description of the frequency-domain features of the typical evolution node; the correlation parameter template is a standard description of the correlation parameters between the typical evolution node and the previous node; and the feature change type template is a standard description of the feature change type of the typical evolution node.

[0107] Step S143: Based on the time-domain features, frequency-domain features, correlation parameters, and feature change type identifiers of the key evolutionary nodes, and the corresponding templates of typical evolutionary nodes as model references, calculate the feature matching degree between the key evolutionary nodes and typical evolutionary nodes; the calculation of the feature matching degree covers the matching degree of time-domain features, the matching degree of frequency-domain features, the matching degree of correlation parameters, and the matching degree of feature change types, and weights the matching degree of each part according to a preset ratio to obtain the comprehensive matching degree.

[0108] Based on the temporal and frequency domain features, associated parameters, and feature change type identifiers of key evolutionary nodes, as well as the corresponding templates of typical evolutionary nodes, the feature matching degree between key and typical evolutionary nodes is calculated. The matching degree of temporal features can be determined by calculating the similarity between the temporal features of key evolutionary nodes and the temporal feature templates of typical evolutionary nodes, such as using cosine similarity or Euclidean distance. The matching degree of frequency domain features can be determined by calculating the similarity between the frequency domain features of key evolutionary nodes and the frequency domain feature templates of typical evolutionary nodes. The matching degree of associated parameters can be determined by calculating the similarity between the associated parameters of key evolutionary nodes and the associated parameter templates of typical evolutionary nodes. The matching degree of feature change types can be determined by judging whether the feature change type identifier of key evolutionary nodes is consistent with the feature change type template of typical evolutionary nodes. Then, the matching degrees of the above different parts are weighted and summed according to a pre-set ratio to obtain the comprehensive matching degree. By calculating the comprehensive matching degree, the similarity between key and typical evolutionary nodes can be quantified.

[0109] Step S144: Sort all gain response modes according to the overall matching degree, and select the gain response mode with the highest overall matching degree as the candidate matching mode. If the highest overall matching degree is lower than the preset matching threshold, select the top preset number of gain response modes with the highest overall matching degree as the candidate matching mode set.

[0110] Based on the calculated overall matching degree, all gain response patterns in the gain response pattern library are sorted, with the gain response pattern ranking higher for those with higher overall matching degrees. Then, the gain response pattern with the highest overall matching degree is selected as a candidate matching pattern. If the highest overall matching degree is lower than a preset matching threshold, it means that no gain response pattern can match the current key evolution node well. In this case, a preset number of gain response patterns with the highest overall matching degree are selected to form a candidate matching pattern set. By selecting candidate matching patterns, gain response patterns that may be applicable to the current audio feature temporal evolution chain can be determined.

[0111] Step S145: When there is only one candidate matching pattern, analyze the difference between the audio feature temporal evolution chain and the feature evolution category corresponding to the candidate matching pattern, and calculate the difference parameters. The difference parameters include temporal feature difference values, frequency domain feature difference values, correlation parameter difference values, and key evolution node distribution difference values.

[0112] When there is only one candidate matching pattern, the current audio feature temporal evolution chain is compared with the feature evolution category corresponding to that candidate matching pattern, and the differences between them are analyzed. Difference parameters are calculated, including temporal feature difference values ​​(the difference between the temporal features of the current audio feature temporal evolution chain and the temporal feature templates in the feature evolution category), frequency domain feature difference values ​​(the difference between the frequency domain features of the current audio feature temporal evolution chain and the frequency domain feature templates in the feature evolution category), association parameter difference values ​​(the difference between the association parameters of the current audio feature temporal evolution chain and the association parameter templates in the feature evolution category), and key evolution node distribution difference values ​​(the difference between the key evolution node distribution of the current audio feature temporal evolution chain and the typical evolution node distribution in the feature evolution category). By calculating these difference parameters, the differences between the current audio feature temporal evolution chain and the feature evolution category corresponding to the candidate matching pattern can be understood.

[0113] Step S146: Generate a difference compensation rule based on the difference parameter. The difference compensation rule is used to correct the gain adjustment sequence corresponding to the candidate matching mode, so that the corrected gain adjustment sequence is more suitable for the audio feature temporal evolution chain. The difference compensation rule includes the gain adjustment amount compensation coefficient and the adjustment direction compensation basis corresponding to each difference value.

[0114] For example, step S1461: standardize the time-domain feature difference value, frequency-domain feature difference value, correlation parameter difference value and key evolution node distribution difference value in the difference parameters, and set a corresponding compensation weight for each standardized difference value. The compensation weight is determined by analyzing the degree of influence of different types of differences on the gain adjustment effect. The greater the degree of influence, the higher the compensation weight.

[0115] Each difference value in the difference parameter is standardized, for example, using a min-max standardization method to convert the difference values ​​to values ​​within the range [0,1]. Then, the impact of different types of differences on the gain adjustment effect is analyzed. For example, through experiments or historical data statistics, the impact of time-domain feature differences, frequency-domain feature differences, correlation parameter differences, and key evolution node distribution differences on the gain adjustment effect is determined. Based on the degree of impact, a corresponding compensation weight is assigned to each standardized difference value; the greater the impact, the higher the compensation weight. Through standardization and setting compensation weights, different types of difference values ​​can be uniformly quantified, facilitating subsequent compensation coefficient calculations.

[0116] Step S1462: Calculate the product of each standardized difference value and the corresponding compensation weight to obtain the preliminary compensation coefficient corresponding to each difference value. The preliminary compensation coefficient reflects the initial influence of the difference value on the gain adjustment sequence correction.

[0117] Each standardized difference value is multiplied by its corresponding compensation weight to obtain the preliminary compensation coefficient for each difference value. The preliminary compensation coefficient reflects the initial impact of the difference value on the correction of the gain adjustment sequence. For example, the larger the standardized difference value and the higher the compensation weight, the larger the preliminary compensation coefficient, indicating that the difference has a greater impact on the correction of the gain adjustment sequence.

[0118] Step S1463: Analyze the structure of the gain adjustment sequence corresponding to the candidate matching mode, decompose the constituent parameters of the gain adjustment sequence, and determine the adjustable parameter types in the gain adjustment sequence. The adjustable parameter types include gain adjustment amount, gain adjustment direction, and gain adjustment time.

[0119] The structure of the gain adjustment sequence corresponding to the candidate matching pattern is analyzed, and the constituent parameters of the gain adjustment sequence are decomposed, such as gain adjustment amount, gain adjustment direction, gain adjustment time, and gain adjustment duration. Then, the types of adjustable parameters are determined, which typically include gain adjustment amount, gain adjustment direction, and gain adjustment time. These adjustable parameter types will be used as the objects of subsequent difference compensation.

[0120] Step S1464: Establish the mapping relationship between standardized variance values ​​and adjustable parameter types, and determine the adjustable parameter type corresponding to each standardized variance value through parameter correlation analysis.

[0121] By performing parameter correlation analysis, such as calculating the correlation coefficient between standardized variance values ​​and adjustable parameter types, a mapping relationship between standardized variance values ​​and adjustable parameter types can be established. For example, time-domain characteristic variance values ​​may be mainly related to the gain adjustment amount, while frequency-domain characteristic variance values ​​may be mainly related to the gain adjustment direction. By establishing this mapping relationship, the adjustable parameter type corresponding to each standardized variance value can be determined.

[0122] Step S1465: Based on the mapping relationship and the preliminary compensation coefficient, generate a compensation sub-rule for each adjustable parameter. The compensation sub-rule includes the adjustment basis, adjustment range calculation method and adjusted value constraint for the adjustable parameter.

[0123] Based on the mapping relationship and the initial compensation coefficient, compensation sub-rules are generated for each adjustable parameter. For example, for the adjustable parameter of gain adjustment, the adjustment is based on the time-domain characteristic difference value, the adjustment magnitude is calculated by multiplying the initial compensation coefficient by the base gain adjustment, and the constraint on the adjusted value is that the gain adjustment cannot exceed the preset maximum and minimum values.

[0124] Step S1466: Construct a compensation sub-rule association network, with compensation sub-rules with different adjustable parameters as network nodes. The connection relationship between nodes is determined based on the collaborative correction relationship between rules. The stronger the collaboration, the greater the connection weight.

[0125] A network of compensation sub-rules with different adjustable parameters is constructed, acting as network nodes. The connections between nodes are determined based on the collaborative correction relationships between rules. For example, if a compensation sub-rule adjusting the gain adjustment amount and a compensation sub-rule adjusting the gain adjustment direction have a collaborative effect—that is, adjusting the gain adjustment amount affects the correction of the gain adjustment direction—then a connection edge is established between them. The connection weight is determined by the strength of the collaboration; the stronger the collaboration, the larger the connection weight. By constructing a network of compensation sub-rules, the collaborative effects between different compensation sub-rules can be considered, improving the accuracy of the correction.

[0126] Step S1467: Based on the compensation sub-rule association network, integrate all compensation sub-rules, optimize the synergy between rules, generate compensation order rules, and determine the correction order of each adjustable parameter in the gain adjustment sequence through parameter adjustment priority analysis.

[0127] Based on the connection relationships and weights in the compensation sub-rules' association network, all compensation sub-rules are integrated to optimize their synergy. For example, the order of highly synergistic compensation sub-rules is adjusted so that their correction effects can mutually reinforce each other. Then, through parameter adjustment priority analysis, the correction order of each adjustable parameter in the gain adjustment sequence is determined; for example, first correcting the gain adjustment amount, then the gain adjustment direction, and finally the gain adjustment timing. By generating compensation order rules, the execution order of the compensation sub-rules can be ensured to be reasonable, thereby improving the correction effect.

[0128] Step S1468: Based on the difference parameters and difference compensation rules, predict the adaptation effect of the corrected gain adjustment sequence to the audio feature temporal evolution chain; The gain adjustment sequence of the candidate matching pattern is corrected using difference parameters and generated difference compensation rules to obtain the corrected gain adjustment sequence. Then, the adaptation effect of the corrected gain adjustment sequence to the temporal evolution chain of audio features is predicted, for example, by simulating the gain optimization process of the audio signal, and analyzing the optimized audio quality and the rationality of feature changes.

[0129] Step S1469: Based on the compensation effect prediction model, iteratively optimize the difference compensation rules to improve the correction accuracy of the difference compensation rules for different difference scenarios, and combine the optimized compensation sub-rules and compensation order rules into complete difference compensation rules.

[0130] A compensation effect prediction model is constructed, taking the corrected gain adjustment sequence and the temporal evolution chain of audio features as input, and outputting an evaluation index of the adaptation effect. Then, based on the output of the prediction model, the difference compensation rules are iteratively optimized. For example, if the predicted adaptation effect is unsatisfactory, the calculation method of the adjustment magnitude and value constraints in the compensation sub-rules are adjusted, or the correction order in the compensation order rule is adjusted. Through iterative optimization, the correction accuracy of the difference compensation rules for different difference scenarios can be improved. Finally, the optimized compensation sub-rules and compensation order rules are combined into a complete difference compensation rule.

[0131] Step S147: Correct the gain adjustment sequence of the candidate matching mode according to the difference compensation rule to obtain the target gain adjustment sequence that is adapted to the audio feature temporal evolution chain, and determine the correspondence between the audio feature temporal evolution chain and the target gain adjustment sequence as the initial dynamic adaptation relationship.

[0132] Based on the generated difference compensation rules, the gain adjustment sequence of the candidate matching pattern is corrected. For example, following the order of the compensation order rules, each compensation sub-rule is applied sequentially to correct adjustable parameters such as gain adjustment amount, gain adjustment direction, and gain adjustment timing. The corrected gain adjustment sequence is the target gain adjustment sequence that adapts to the current audio feature temporal evolution chain. Then, the correspondence between the audio feature temporal evolution chain and the target gain adjustment sequence is determined as the initial dynamic adaptation relationship.

[0133] Step S148: When the candidate matching patterns are a set, perform difference analysis, difference compensation rule generation and gain adjustment sequence correction operations on each candidate matching pattern in the candidate matching pattern set to obtain the corrected gain adjustment sequence corresponding to each candidate matching pattern. When the candidate matching patterns are a set, steps S145 to S147 are performed for each candidate matching pattern in the set, namely, performing difference analysis, generating difference compensation rules, and correcting the gain adjustment sequence to obtain the corrected gain adjustment sequence corresponding to each candidate matching pattern. By processing each candidate matching pattern, multiple corrected gain adjustment sequences can be obtained.

[0134] Step S149: Based on the key evolution nodes of the audio feature temporal evolution chain and the adjustment parameters of the modified gain adjustment sequence, calculate the adaptation score of each modified gain adjustment sequence and the audio feature temporal evolution chain, select the modified gain adjustment sequence with the highest adaptation score as the target gain adjustment sequence, and establish an initial dynamic adaptation relationship. Based on the key evolution nodes of the audio feature temporal evolution chain and the adjustment parameters of each modified gain adjustment sequence, a fit score is calculated for each modified gain adjustment sequence and the audio feature temporal evolution chain. The calculation method for the fit score can include multiple aspects, such as the degree of matching between the gain adjustment amount and the feature change amplitude, the degree of matching between the gain adjustment direction and the feature change trend, and the degree of matching between the gain adjustment time and the feature abrupt change point. The fit scores are obtained by weighted summation of the matching degrees in these aspects. Then, the modified gain adjustment sequence with the highest fit score is selected as the target gain adjustment sequence, establishing an initial dynamic fit relationship. By calculating the fit score and selecting the target gain adjustment sequence, the best fit between the final gain adjustment sequence and the audio feature temporal evolution chain can be ensured.

[0135] Step S1410: Monitor the real-time changes of the audio feature temporal evolution chain. When a new key evolution node appears in the audio feature temporal evolution chain or the features of an existing key evolution node change, re-execute the above steps of key evolution node extraction, candidate matching mode selection, difference analysis, correction and adaptation relationship establishment, update the dynamic adaptation relationship, and make the dynamic adaptation relationship consistent with the current state of the audio feature temporal evolution chain.

[0136] The system monitors changes in the temporal evolution chain of audio features in real time. When a new key evolution node appears or the features of an existing key evolution node change, steps S141 to S149 are re-executed: extracting key evolution nodes, selecting candidate matching patterns, performing difference analysis, correcting the gain adjustment sequence, establishing adaptation relationships, and updating the dynamic adaptation relationships. Real-time monitoring and updating ensure that the dynamic adaptation relationships always remain consistent with the current state of the audio feature temporal evolution chain, improving the real-time performance and accuracy of gain optimization.

[0137] Step S150: Generate a dynamic gain scheme corresponding to the set of audio signals to be optimized based on the dynamic adaptation relationship, apply the dynamic gain scheme to the gain adjustment of the audio equipment, perform gain optimization processing on the set of audio signals to be optimized, obtain the set of audio output signals after gain optimization, and feed back the characteristic timing evolution information corresponding to the set of audio output signals to the gain response pattern library to improve the content of the pattern library.

[0138] Step S151: Extract the target gain adjustment sequence and the corresponding audio feature temporal evolution chain segment from the dynamic adaptation relationship, and determine the time node and feature conditions corresponding to each gain adjustment parameter in the target gain adjustment sequence; The target gain adjustment sequence and the corresponding audio feature temporal evolution chain segment are extracted from the dynamic adaptation relationship. Then, the target gain adjustment sequence is analyzed to determine the time node and feature conditions corresponding to each gain adjustment parameter (such as gain adjustment amount, gain adjustment direction, gain adjustment time, etc.). For example, the time node corresponding to a certain gain adjustment amount is from t1 to t2, and the feature conditions are that the change amplitude of the time domain feature is within a certain range, and the energy transfer ratio of the frequency domain feature is within a certain range, etc.

[0139] Step S152: Divide the time axis of the audio signal set to be optimized into multiple consecutive time intervals according to a preset time interval. The duration of each time interval is determined according to the sampling frequency and gain adjustment accuracy requirements of the audio signal set to be optimized. Through a preset adjustment accuracy verification process, determine whether the changes in audio characteristics within each time interval meet the minimum unit requirement for gain adjustment. If not, adjust the duration of the time interval until it meets the requirement. Based on the sampling frequency and gain adjustment accuracy requirements of the audio signal set to be optimized, a preset time interval is determined, dividing the time axis of the audio signal set into multiple consecutive time intervals. For example, if the sampling frequency is f and the gain adjustment accuracy requirement is that the audio feature changes within each time interval can be accurately adjusted, then the duration of the time interval can be determined based on f and the accuracy requirement. Then, through a preset adjustment accuracy verification process, it is checked whether the audio feature changes within each time interval meet the minimum unit requirement for gain adjustment. If not, the duration of the time interval is adjusted, for example, shortening or lengthening the time interval until the requirement is met. By dividing the time intervals and verifying the adjustment accuracy, the accuracy and effectiveness of the gain adjustment can be ensured.

[0140] Step S153: For each time interval, extract the corresponding audio feature temporal evolution chain segment within that time interval. The audio feature temporal evolution chain segment contains the temporal domain features, frequency domain features, and adjacent association information of all audio signal units within that time interval. For each defined time interval, the corresponding audio feature temporal evolution chain segment is extracted from the audio feature temporal evolution chain. This segment contains the temporal features (such as average amplitude, amplitude fluctuation range, amplitude change frequency, etc.), frequency domain features (such as the total energy of each frequency interval, energy percentage, etc.), and adjacent correlation information (such as adjacent temporal feature correlation parameters, adjacent frequency domain feature correlation parameters, etc.) of all audio signal units within the time interval.

[0141] Step S154: Associate the audio feature temporal evolution chain segment of each time interval with the corresponding segment of the target gain adjustment sequence, and output the recommended value of the gain adjustment parameter corresponding to the time interval; The audio feature temporal evolution chain segment for each time interval is associated with the corresponding segment of the target gain adjustment sequence. For example, based on the feature changes in the audio feature temporal evolution chain segment, the corresponding gain adjustment parameter in the target gain adjustment sequence is found, and the recommended value of the gain adjustment parameter for that time interval is output. The recommended value of the gain adjustment parameter includes the gain adjustment amount, gain adjustment direction, and gain adjustment timing.

[0142] Step S155: Based on the dynamic adaptation relationship and the recommended value of the gain adjustment parameter, find the target gain adjustment sequence segment that matches the audio feature temporal evolution chain segment of the time interval, and determine the gain adjustment parameter corresponding to the time interval. The gain adjustment parameter includes the gain adjustment amount, adjustment direction and adjustment start time of the time interval. Based on the dynamic adaptation relationship and recommended values ​​for gain adjustment parameters, a segment matching the temporal evolution chain of audio features within the target gain adjustment sequence is searched. Then, the gain adjustment parameters corresponding to that time interval are determined, including the gain adjustment amount, adjustment direction, and adjustment start time. For example, based on the parameter information in the matched target gain adjustment sequence segment, the gain adjustment amount for that time interval, whether the adjustment direction is amplification or attenuation, and the adjustment start time within that time interval are determined.

[0143] Step S156: Based on the gain adjustment parameters of the time interval and combined with the changing trend of audio features within the time interval, generate a local gain adjustment sub-scheme for the time interval. The local gain adjustment sub-scheme includes the gain adjustment value and corresponding adjustment basis for each sampling moment within the time interval. Based on the gain adjustment parameters for that time interval and the changing trends of audio features within that time interval, a local gain adjustment sub-scheme is generated for that time interval. For example, based on the gain adjustment amount and direction, as well as the changing trends of audio features (such as a gradual increase or decrease in amplitude, a gradual change in frequency energy distribution, etc.), the gain adjustment value for each sampling moment within that time interval is determined. Simultaneously, the adjustment basis corresponding to the gain adjustment value at each sampling moment is recorded, such as whether it is based on the amplitude of changes in time-domain features, the proportion of energy transfer in the frequency domain, or a point of abrupt change in features.

[0144] Step S157: Calculate the smoothness of the transition between local gain adjustment sub-schemes in adjacent time intervals, output transition optimization suggestions, and perform transition processing on local gain adjustment sub-schemes in adjacent time intervals according to the transition optimization suggestions. Integrate all the transition-processed local gain adjustment sub-schemes in chronological order to form a preliminary dynamic gain scheme covering the entire time axis of the audio signal set to be optimized. The preliminary dynamic gain scheme includes the gain adjustment value and the corresponding timestamp at each sampling moment. The smoothness of transition between local gain adjustment sub-schemes in adjacent time intervals is calculated, such as the rate of change and consistency of trend of gain adjustment values ​​in adjacent time intervals. Based on the calculation results, transition optimization suggestions are output, such as adjusting the gain adjustment value in a certain time interval to make the gain adjustment transition between adjacent time intervals smoother. Then, based on the transition optimization suggestions, the local gain adjustment sub-schemes in adjacent time intervals are processed for transition, such as interpolating the gain adjustment values ​​in adjacent time intervals to make their changes smoother. Finally, all the processed local gain adjustment sub-schemes are integrated in chronological order to form a preliminary dynamic gain scheme, which includes the gain adjustment value and corresponding timestamp for each sampling moment. Through transition processing and integration, a smooth transition of the dynamic gain scheme in the time dimension can be ensured, improving the effect of gain optimization.

[0145] Step S158: Establish a correlation between the local gain adjustment sub-scheme of each time interval and the corresponding audio feature temporal evolution chain segment to obtain the dynamic gain scheme correlation network; A dynamic gain scheme association network is constructed by associating local gain adjustment sub-schemes for each time interval with corresponding audio feature temporal evolution chain segments. In this network, each node contains a local gain adjustment sub-scheme for a time interval and a corresponding audio feature temporal evolution chain segment, and the connection relationships between nodes are determined based on temporal order and feature association relationships. By establishing a dynamic gain scheme association network, gain adjustment schemes and audio feature evolution chains can be linked, facilitating subsequent analysis and optimization.

[0146] Step S159: Based on the dynamic gain scheme association network, extract the gain adjustment rules across time intervals, optimize the overall consistency of the dynamic gain scheme, and determine the optimized dynamic gain scheme as the final dynamic gain scheme. The final dynamic gain scheme includes the gain adjustment value, timestamp, adjustment basis and device adaptation instructions for each sampling time.

[0147] The dynamic gain scheme's correlation network is analyzed to extract gain adjustment patterns across time intervals, such as the variation patterns of gain adjustment amount, direction, and usage patterns of the adjustment criteria. Based on these extracted patterns, the dynamic gain scheme is optimized to improve its overall consistency. For example, gain adjustment parameters in certain time intervals are adjusted to better conform to the overall gain adjustment patterns. The optimized dynamic gain scheme is the final dynamic gain scheme, which includes the gain adjustment value, timestamp, adjustment criteria, and device compatibility specifications for each sampling time (e.g., which types of audio equipment the gain adjustment scheme is suitable for and what equipment parameters are required). By optimizing and determining the final dynamic gain scheme, the rationality and applicability of the gain adjustment scheme can be ensured.

[0148] Step S1510: Analyze the dynamic gain scheme, extract the gain adjustment value and the corresponding timestamp at each sampling time, and generate a table showing the correspondence between the gain adjustment value and the timestamp; The final dynamic gain scheme is analyzed to extract the gain adjustment value and corresponding timestamp for each sampling moment. Then, the above gain adjustment values ​​and timestamps are organized into a correspondence table, where each row corresponds to a sampling moment and contains the gain adjustment value and timestamp for that moment.

[0149] Step S1511: Convert the gain adjustment values ​​in the correspondence table into a control signal format that can be recognized by the audio equipment gain adjustment. Each gain adjustment value corresponds to a control signal, forming a control signal sequence. The control signals in the control signal sequence are arranged in order of timestamp. Based on the gain adjustment interface requirements of the audio equipment, the gain adjustment values ​​in the correspondence table are converted into a control signal format that the audio equipment can recognize. For example, if the audio equipment's gain adjustment interface uses digital signal input, the gain adjustment values ​​are converted into the corresponding digital signals; if it uses analog signal input, the gain adjustment values ​​are converted into the corresponding voltage or current signals. Each gain adjustment value corresponds to a control signal, and these control signals are arranged in the order of their timestamps to form a control signal sequence.

[0150] Step S1512: The control signal sequence is transmitted to the gain adjustment component of the audio equipment in chronological order. After receiving the control signal, the gain adjustment component adjusts its own gain parameter according to the gain adjustment value corresponding to the control signal. At each sampling time, the current gain parameter of the gain adjustment component is compared with the gain adjustment value at the corresponding sampling time in the dynamic gain scheme. If there is a difference, the gain parameter of the gain adjustment component is adjusted to be consistent with the gain adjustment value. The control signal sequence is transmitted sequentially to the gain adjustment unit of the audio equipment. Upon receiving the control signal, the gain adjustment unit adjusts its own gain parameter according to the corresponding gain adjustment value in the control signal. At each sampling moment, the gain adjustment unit compares the current gain parameter with the gain adjustment value for the corresponding sampling moment in the dynamic gain scheme. If a difference exists, such as the current gain parameter being greater than or less than the gain adjustment value, the gain adjustment unit adjusts its own gain parameter to match the gain adjustment value. By transmitting control signals and adjusting the gain parameter, dynamic adjustment of the audio equipment's gain can be achieved.

[0151] Step S1513: The set of audio signals to be optimized is transmitted to the signal processing unit of the audio equipment in chronological order. After receiving the audio signal units, the signal processing unit transmits the audio signal units to the gain adjustment unit, which performs gain amplification or attenuation processing on the audio signal units according to the gain parameters at the current time. The set of audio signals to be optimized is transmitted sequentially to the signal processing unit of the audio equipment. Upon receiving the audio signal units, the signal processing unit transmits them to the gain adjustment unit. The gain adjustment unit amplifies or attenuates the audio signal units based on the current gain parameter. For example, if the gain parameter is positive, gain amplification is performed by multiplying the amplitude of the audio signal unit by the gain parameter; if the gain parameter is negative, gain attenuation is performed by dividing the amplitude of the audio signal unit by the absolute value of the gain parameter. By transmitting the audio signal and performing gain processing, gain optimization of the audio signal can be achieved.

[0152] Step S1514: Collect the audio signal units processed by the gain adjustment component, sort the processed audio signal units according to the timestamp order, and form a preliminary audio output signal set; The audio signal units, processed by the gain adjustment unit, are collected. These audio signal units contain timestamp information. Then, the processed audio signal units are sorted according to the timestamp order to form a preliminary audio output signal set. By sorting, it can be ensured that the time order of the audio output signal set is consistent with the time order of the original audio signal set to be optimized, thus guaranteeing the continuity of the audio.

[0153] Step S1515: Analyze the transition characteristics of adjacent audio signal units based on the preliminary audio output signal set, and output transition optimization parameters; The initial set of audio output signals is analyzed, focusing on the transition characteristics of adjacent audio signal units, such as whether the amplitude changes of adjacent audio signal units are smooth and whether the changes in frequency energy distribution are natural. Based on the analysis results, transition optimization parameters are output, such as the amplitude transition coefficient and frequency energy distribution transition coefficient between adjacent audio signal units.

[0154] Step S1516: Based on the transition optimization parameters, perform signal continuity processing on the initial audio output signal set to make the audio output signal set continuous, and use the audio output signal set after continuity processing as the audio output signal set after gain optimization. Based on the output transition optimization parameters, signal continuity processing is performed on the initial audio output signal set. For example, interpolation algorithms are used to process the transition between adjacent audio signal units, making amplitude changes smoother and frequency energy distribution changes more natural. The resulting audio output signal set has good continuity and is used as the gain-optimized audio output signal set. Through signal continuity processing, the quality of audio output can be improved, avoiding problems such as stuttering and noise.

[0155] Step S1517: Extract features from the gain-optimized audio output signal set and construct the audio feature temporal evolution chain corresponding to the audio output signal set; For the set of audio output signals after gain optimization, feature extraction is performed according to the methods in steps S121 to S128, including time-domain feature extraction, frequency-domain feature extraction, feature time sequence construction, adjacent feature correlation analysis, feature mutation point extraction, feature correlation weight matrix construction, and audio feature time sequence evolution chain construction, etc. Finally, the audio feature time sequence evolution chain corresponding to the set of audio output signals is constructed.

[0156] Step S1518: Extract the audio feature temporal evolution chain corresponding to the set of audio signals to be optimized, the applied dynamic gain scheme, and the audio feature temporal evolution chain corresponding to the set of audio output signals, and integrate them into gain optimization effect correlation data; The audio feature time-series evolution chain corresponding to the set of audio signals to be optimized, the applied dynamic gain scheme, and the audio feature time-series evolution chain corresponding to the set of audio output signals are extracted. Then, the above data are integrated to form gain optimization effect correlation data. This data includes the correlation information between the original audio features, the gain adjustment scheme, and the optimized audio features.

[0157] Step S1519: Construct an optimization effect correlation graph, taking the audio feature temporal evolution chain of the audio signal set to be optimized, the dynamic gain scheme, and the audio feature temporal evolution chain of the audio output signal set as graph nodes, and the edges between nodes are determined based on the causal correlation strength; An optimization effect correlation graph is constructed, with the temporal evolution chain of audio features of the audio signal set to be optimized, the dynamic gain scheme, and the temporal evolution chain of audio features of the audio output signal set as nodes in the graph. The edges between nodes are determined based on the strength of their causal relationships. For example, the causal relationship between the temporal evolution chain of audio features of the audio signal set to be optimized and the dynamic gain scheme is based on their matching degree, while the causal relationship between the temporal evolution chain of audio features of the dynamic gain scheme and the audio output signal set is based on the effect of gain optimization. By constructing the optimization effect correlation graph, the relationships between the data in different parts of the gain optimization process can be visually displayed.

[0158] Step S1520: Based on the correlation graph of the optimization effect, extract the key influencing factors of gain optimization, and transmit the correlation data of gain optimization effect and key influencing factors to the update interface of the gain response pattern library. Analyze the corresponding feature evolution law and gain adjustment rules. If there is no corresponding gain response pattern in the gain response pattern library, generate a new gain response pattern and add it to the gain response pattern library; if there is already a corresponding pattern for the feature evolution law, modify the gain adjustment rules of the existing pattern. The correlation graph of the optimization effect is analyzed to extract key influencing factors of gain optimization, such as which feature parameters have the greatest impact on the gain optimization effect and which adjustment parameters of gain adjustment have the most significant effect improvement. Then, the correlation data of gain optimization effect and key influencing factors are transferred to the update interface of the gain response pattern library. In the update interface, the corresponding feature evolution laws and gain adjustment rules are analyzed. If there is no corresponding gain response pattern for a feature evolution law in the gain response pattern library, a new gain response pattern is generated and added to the library; if there is already a corresponding pattern for the feature evolution law, the gain adjustment rules of the existing pattern are revised to make it more complete and accurate. By updating the gain response pattern library, the applicability of the pattern library and the effect of gain optimization can be continuously improved.

[0159] Step S1521: Record the updated content of the gain response pattern library, including information on newly added gain response patterns or changes before and after correcting gain response patterns, form a pattern library update log, and store the pattern library update log in a preset log storage area.

[0160] Record updates to the gain response pattern library, such as feature evolution category identifiers, feature evolution law descriptions, and gain adjustment rule sets for newly added gain response patterns, or changes in the gain response patterns before and after modification, such as changes in the range of gain adjustment values ​​or changes in the basis for correcting the gain adjustment direction. Compile these updates into a pattern library update log and store it in a pre-defined log storage area, such as a database or file system. Recording and storing the pattern library update log facilitates subsequent querying and analysis of the pattern library's update history.

[0161] Based on the same inventive concept, please refer to Figure 2 This paper shows a schematic block diagram of an audio feature-based audio dynamic gain optimization system 100 for performing the above-described audio feature-based dynamic gain optimization method, provided in an embodiment of this application. The audio feature-based dynamic gain optimization system 100 may include a communication unit 110, a machine-readable storage medium 120, and a processor 130.

[0162] The machine-readable storage medium 120 is used to store machine-executable instructions for executing the scheme of this application, and the processor 130 is used to execute the machine-executable instructions stored in the machine-readable storage medium 120 to implement the audio dynamic gain optimization method based on audio features provided in the aforementioned method embodiments.

[0163] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. A method for optimizing dynamic gain of audio based on audio features, characterized in that, The method includes: Obtain a set of audio signals to be optimized, which includes continuous audio signal units collected under different scenarios; The set of audio signals to be optimized is subjected to feature temporal evolution analysis, and an audio feature temporal evolution chain corresponding to the set of audio signals to be optimized is constructed. The audio feature temporal evolution chain includes the change sequence of audio features in the continuous time dimension and the correlation between features. Gain response patterns are extracted from a set of historical audio gain optimization cases to form a gain response pattern library, which contains gain adjustment sequences corresponding to the temporal evolution chains of different audio features. A dynamic adaptation relationship is established between the audio feature temporal evolution chain and the gain adjustment sequence in the gain response pattern library, and the dynamic adaptation relationship is updated in real time with the changes of the audio feature temporal evolution chain. Based on the dynamic adaptation relationship, a dynamic gain scheme corresponding to the set of audio signals to be optimized is generated. The dynamic gain scheme is applied to the gain adjustment of the audio equipment. Gain optimization processing is performed on the set of audio signals to be optimized to obtain a set of audio output signals with optimized gain. The characteristic timing evolution information corresponding to the set of audio output signals is fed back to the gain response pattern library to improve the content of the pattern library.

2. The audio dynamic gain optimization method based on audio features according to claim 1, characterized in that, The step of performing feature temporal evolution analysis on the set of audio signals to be optimized, and constructing the audio feature temporal evolution chain corresponding to the set of audio signals to be optimized, includes: Temporal features are extracted from each audio signal unit in the set of audio signals to be optimized to obtain the temporal features of each audio signal unit. The temporal features include the amplitude variation statistics of the audio signal unit in the time dimension. The amplitude variation statistics cover the average amplitude, amplitude fluctuation range and amplitude variation frequency. Frequency domain features are extracted for each audio signal unit in the set of audio signals to be optimized, and the frequency domain features of each audio signal unit are obtained. The frequency domain features include the energy distribution information of the audio signal unit in the frequency dimension. The energy distribution information covers the total energy in different frequency ranges and the proportion of energy in each frequency range to the total energy. The time-domain features of all audio signal units are sorted in chronological order to form a time-domain feature sequence. At the same time, the frequency-domain features of all audio signal units are sorted in chronological order to form a frequency-domain feature sequence. Adjacent feature correlation analysis is performed on the time-domain feature time series to calculate the amplitude variation range and trend between the time-domain features of two adjacent audio signal units. The amplitude variation range is the sum of the absolute values ​​of the relative difference of the average amplitude, the relative difference of the amplitude fluctuation range, and the relative difference of the amplitude variation frequency in adjacent time-domain features. The relative difference is the ratio of the difference to the corresponding feature value of the previous audio signal unit. The trend is the direction of increase or decrease of each statistical information in adjacent time-domain features. Adjacent feature correlation analysis is performed on the frequency domain feature time series to calculate the energy transfer ratio and transfer direction between the frequency domain features of two adjacent audio signal units. The energy transfer ratio is the ratio of the difference in the total energy of the same frequency interval in adjacent frequency domain features to the total energy of the previous audio signal unit in the same frequency interval. The transfer direction is the direction of increase or decrease of energy in each frequency interval in adjacent frequency domain features. Extract feature mutation points from the time-domain feature time series. The feature mutation points are the time-domain features corresponding to the time points in the time-domain features where the amplitude change exceeds a preset amplitude threshold. At the same time, extract feature mutation points from the frequency-domain feature time series. The feature mutation points are the frequency-domain features corresponding to the time points in the frequency-domain features where the energy transfer ratio exceeds a preset ratio threshold. A feature association weight matrix is ​​constructed, with each feature parameter in the time-domain feature time series and the frequency-domain feature time series as a matrix row vector and the association strength between features as a matrix element value. The intrinsic association between features is strengthened through matrix operations. Based on the aforementioned feature association weight matrix, the association information of adjacent time-domain features, the association information of adjacent frequency-domain features, and the two types of feature mutation points are integrated into the matrix association relationship and connected in series along the time dimension to form a continuous audio feature time-series evolution chain. Each time node in the audio feature time-series evolution chain contains the time-domain features, frequency-domain features, and association information with adjacent nodes of the corresponding audio signal unit.

3. The audio dynamic gain optimization method based on audio features according to claim 2, characterized in that, The step of extracting time-domain features from each audio signal unit in the set of audio signals to be optimized, to obtain the time-domain features of each audio signal unit, includes: Each audio signal unit is uniformly sampled in the time dimension to obtain the amplitude value of each audio signal unit at multiple equally spaced sampling moments. The number of sampling moments is determined according to the duration of the audio signal unit and the preset sampling frequency. Calculate the arithmetic mean of the amplitude values ​​at multiple sampling times, and use this arithmetic mean as the average amplitude of the audio signal unit. The average amplitude reflects the overall amplitude level of the audio signal unit in the time dimension. Find the maximum and minimum values ​​of amplitude values ​​at multiple sampling times, calculate the difference between the maximum and minimum values, and use this difference as the amplitude fluctuation range of the audio signal unit. The amplitude fluctuation range reflects the amplitude change of the audio signal unit in the time dimension. The amplitude values ​​at multiple sampling times are analyzed for trend changes to identify the turning points where the amplitude values ​​change from rising to falling or from falling to rising. The number of turning points per unit time is counted and used as the amplitude change frequency of the audio signal unit. The amplitude change frequency reflects how fast the amplitude of the audio signal unit changes in the time dimension. A temporal feature association network is constructed, and association edges are established between the temporal features of each audio signal unit and the temporal features of adjacent units. The weight of the association edge is determined based on the feature similarity. The higher the feature similarity, the greater the weight. Based on the temporal feature association network, common patterns of temporal features across audio signal units are extracted, and these common patterns are integrated into the temporal feature representation of each audio signal unit to optimize the representation accuracy of temporal features. The optimized average amplitude, amplitude fluctuation range, and amplitude change frequency are integrated into the time-domain characteristics of the audio signal unit. Based on the correlation relationship of the temporal feature association network, the temporal features of different audio signal units are collaboratively optimized so that the temporal features of adjacent audio signal units maintain consistency in their changing trends.

4. The audio dynamic gain optimization method based on audio features according to claim 2, characterized in that, The step of performing adjacent feature correlation analysis on the frequency domain feature time sequence, and calculating the energy transfer ratio and transfer direction between the frequency domain features of two adjacent audio signal units, includes: Two adjacent audio signal units in the frequency domain feature time sequence are selected and denoted as the preceding audio signal unit and the following audio signal unit, respectively. The frequency domain features of the preceding audio signal unit and the frequency domain features of the following audio signal unit are extracted. Determine the common frequency interval in the frequency domain features of two audio signal units, compare the frequency intervals in the frequency domain features of the two audio signal units, and select the frequency interval that exists in both frequency domain features as the common frequency interval. If any frequency interval appears only in the frequency domain features of one audio signal unit, then the frequency interval is regarded as a frequency interval with zero energy and included in the analysis. For each common frequency interval, calculate the difference between the total energy of the subsequent audio signal unit in that frequency interval and the total energy of the preceding audio signal unit in that frequency interval, and obtain the energy difference of that frequency interval; Calculate the ratio of the energy difference in each frequency range to the total energy of the preceding audio signal unit in that frequency range, and use this ratio as the energy transfer ratio for that frequency range. If the total energy of the preceding audio signal unit in that frequency range is zero, then set the energy transfer ratio for that frequency range according to a preset rule. Determine the sign of the energy difference in each frequency range. If the energy difference is positive, the energy transfer direction of that frequency range is increasing from front to back. If the energy difference is negative, the energy transfer direction of that frequency range is decreasing from front to back. If the energy difference is zero, the energy transfer direction of that frequency range remains unchanged. Construct a frequency domain feature transfer map, and use the energy transfer ratio and transfer direction of each frequency interval as the map node attributes. The connection relationship between nodes is determined based on the adjacency of frequency intervals, and nodes in adjacent frequency intervals are connected by edges. Based on the frequency domain feature transfer spectrum, the dominant frequency range of energy transfer is identified. The dominant frequency range is the frequency range with the largest absolute value of energy transfer ratio. The energy transfer characteristics of the dominant frequency range are used as a reference for the correlation of adjacent frequency domain features. The energy transfer ratio, transfer direction and dominant frequency range identifier of all frequency ranges are integrated to form the correlation information between the frequency domain characteristics of two adjacent audio signal units. Each pair of adjacent audio signal units corresponds to a set of the above correlation information. For all adjacent audio signal units in the frequency domain feature time sequence, return to the execution and denote them as the preceding audio signal unit and the following audio signal unit, respectively. Extract the frequency domain features of the preceding audio signal unit and the following audio signal unit to obtain a set of adjacent frequency domain feature association information covering the entire frequency domain feature time sequence.

5. The audio dynamic gain optimization method based on audio features according to claim 1, characterized in that, The gain response patterns extracted from the historical audio gain optimization case set form a gain response pattern library, including: Obtain a set of historical audio gain optimization cases. Each historical audio gain optimization case includes a set of historical audio signals to be optimized, a temporal evolution chain of audio features corresponding to the set of historical audio signals to be optimized, a gain adjustment sequence applied to the historical audio gain optimization case, and a set of historical audio output signals after gain optimization. For each historical audio gain optimization case, the audio feature temporal evolution chain is analyzed to extract the temporal subsequence of temporal features, the temporal subsequence of frequency features, the correlation parameters of adjacent temporal features, the correlation parameters of adjacent frequency features, and the distribution information of temporal and frequency feature abrupt change points in each historical audio feature temporal evolution chain, and integrate them into historical feature evolution elements. Similarity clustering is performed on the historical feature evolution elements of all historical audio gain optimization cases. The similarity between any two historical feature evolution elements is calculated. The similarity calculation covers the similarity of time-domain feature time-series subsequences, the similarity of correlation parameters, and the similarity of mutation point distribution. The similarity of each part is weighted and summed according to preset weights to obtain the comprehensive similarity. Historical feature evolution elements with a comprehensive similarity greater than the similarity threshold are grouped into the same feature evolution category. Each feature evolution category contains multiple historical feature evolution elements with similar feature change patterns. The historical gain adjustment sequence in each feature evolution category is extracted to extract the regularity, and the correlation between the historical gain adjustment sequence and the corresponding historical feature evolution element under the feature evolution category is analyzed. The correspondence rules between audio feature changes and gain adjustment in the feature evolution category are determined. The correspondence rules include the gain adjustment amount corresponding to the feature change amplitude, the gain adjustment direction corresponding to the feature change trend, and the gain adjustment timing corresponding to the feature mutation point. Each feature evolution category and its corresponding gain adjustment rule are combined into a gain response pattern. Each gain response pattern contains a feature evolution category identifier, a description of the feature evolution law, and a corresponding set of gain adjustment rules. Based on the correlation strength between feature evolution categories and gain adjustment rules, all gain response modes are prioritized and sorted to form an ordered queue of gain response modes. An ordered queue of gain response patterns is entered into the initial pattern library, and a feature matching index for the gain response patterns is established. The index keywords include key feature parameters in the feature evolution category identifier and adjustment parameters in the gain adjustment rule set. The merged pattern library is dynamically optimized. When a new historical audio gain optimization case is added, the step of performing feature parsing on the audio feature temporal evolution chain in each historical audio gain optimization case is re-executed. The new historical audio gain optimization case is processed to extract a new gain response pattern. If the new gain response pattern is different from the existing gain response patterns in the gain response pattern library, the new gain response pattern is added to the gain response pattern library and the feature matching index is updated.

6. The audio dynamic gain optimization method based on audio features according to claim 1, characterized in that, The process of establishing a dynamic adaptation relationship between the audio feature temporal evolution chain and the gain adjustment sequences in the gain response pattern library includes: Extract key evolution nodes in the audio feature temporal evolution chain. The key evolution nodes are the feature information and related information corresponding to the feature mutation points and time nodes in the audio feature temporal evolution chain where the feature change amplitude exceeds the preset feature change amplitude threshold. Each key evolution node contains time domain features, frequency domain features, related parameters with the previous node, and feature change type identifier. Extract typical evolution nodes from the feature evolution category corresponding to each gain response pattern from the gain response pattern library. The typical evolution nodes are representative feature mutation points and feature change amplitudes that exceed the preset typical feature change amplitude thresholds in the historical feature evolution elements under the feature evolution category. Each typical evolution node includes a time domain feature template, a frequency domain feature template, an association parameter template, and a feature change type template. Based on the time-domain features, frequency-domain features, correlation parameters, and feature change type identifiers of the key evolutionary nodes, and the corresponding templates of typical evolutionary nodes as model references, the feature matching degree between the key evolutionary nodes and typical evolutionary nodes is calculated. The calculation of the feature matching degree covers the matching degree of time-domain features, the matching degree of frequency-domain features, the matching degree of correlation parameters, and the matching degree of feature change types. The matching degree of each part is weighted according to a preset ratio to obtain the comprehensive matching degree. All gain response modes are sorted according to the comprehensive matching degree, and the gain response mode with the highest comprehensive matching degree is selected as the candidate matching mode. If the highest comprehensive matching degree is lower than the preset matching threshold, the top preset number of gain response modes with the highest comprehensive matching degree are selected as the candidate matching mode set. When there is only one candidate matching pattern, the difference between the audio feature temporal evolution chain and the feature evolution category corresponding to the candidate matching pattern is analyzed, and the difference parameters are calculated. The difference parameters include temporal feature difference values, frequency domain feature difference values, correlation parameter difference values, and key evolution node distribution difference values. Based on the difference parameters, a difference compensation rule is generated. The difference compensation rule is used to correct the gain adjustment sequence corresponding to the candidate matching mode. The difference compensation rule includes the gain adjustment amount compensation coefficient and the adjustment direction compensation basis corresponding to each difference value. The gain adjustment sequence of the candidate matching mode is corrected according to the difference compensation rule to obtain the target gain adjustment sequence that is adapted to the audio feature temporal evolution chain. The correspondence between the audio feature temporal evolution chain and the target gain adjustment sequence is determined as the initial dynamic adaptation relationship. When the candidate matching patterns are a set, for each candidate matching pattern in the set, differential analysis, differential compensation rule generation and gain adjustment sequence correction operations are performed to obtain the corrected gain adjustment sequence for each candidate matching pattern. Based on the key evolution nodes of the audio feature temporal evolution chain and the adjustment parameters of the modified gain adjustment sequence, calculate the fit score between each modified gain adjustment sequence and the audio feature temporal evolution chain, select the modified gain adjustment sequence with the highest fit score as the target gain adjustment sequence, and establish an initial dynamic fit relationship. The system monitors the real-time changes of the audio feature temporal evolution chain. When a new key evolution node appears in the audio feature temporal evolution chain or the features of an existing key evolution node change, the system re-executes the above steps of key evolution node extraction, candidate matching mode selection, difference analysis, correction, and adaptation relationship establishment, and updates the dynamic adaptation relationship to ensure that the dynamic adaptation relationship is consistent with the current state of the audio feature temporal evolution chain.

7. The audio dynamic gain optimization method based on audio features according to claim 1, characterized in that, The step of generating a dynamic gain scheme corresponding to the set of audio signals to be optimized based on the dynamic adaptation relationship includes: Extract the target gain adjustment sequence and the corresponding audio feature temporal evolution chain segment from the dynamic adaptation relationship, and determine the time node and feature conditions corresponding to each gain adjustment parameter in the target gain adjustment sequence; The time axis of the audio signal set to be optimized is divided into multiple continuous time intervals according to a preset time interval. The duration of each time interval is determined according to the sampling frequency and gain adjustment accuracy requirements of the audio signal set to be optimized. Through a preset adjustment accuracy verification process, it is determined whether the changes in audio characteristics in each time interval meet the minimum unit requirement of gain adjustment. If not, the duration of the time interval is adjusted until it meets the requirement. For each time interval, the corresponding audio feature time-series evolution chain segment contains the time-domain features, frequency-domain features, and adjacent association information of all audio signal units in that time interval; The audio feature temporal evolution chain segment of each time interval is associated with the corresponding segment of the target gain adjustment sequence, and the recommended value of the gain adjustment parameter corresponding to that time interval is output. Based on the dynamic adaptation relationship and recommended values ​​of gain adjustment parameters, a target gain adjustment sequence segment that matches the audio feature temporal evolution chain segment of the time interval is found, and the gain adjustment parameters corresponding to the time interval are determined. The gain adjustment parameters include the gain adjustment amount, adjustment direction and adjustment start time of the time interval. Based on the gain adjustment parameters of the time interval, and combined with the changing trend of audio features within the time interval, a local gain adjustment sub-scheme for the time interval is generated. The local gain adjustment sub-scheme includes the gain adjustment value and the corresponding adjustment basis for each sampling moment within the time interval. Calculate the smoothness of the transition between local gain adjustment sub-schemes in adjacent time intervals, output transition optimization suggestions, and perform transition processing on the local gain adjustment sub-schemes in adjacent time intervals according to the transition optimization suggestions. Integrate all the transition-processed local gain adjustment sub-schemes in chronological order to form a preliminary dynamic gain scheme covering the entire time axis of the audio signal set to be optimized. The preliminary dynamic gain scheme includes the gain adjustment value and the corresponding timestamp at each sampling moment. By associating the local gain adjustment sub-schemes of each time interval with the corresponding audio feature temporal evolution chain segments, a dynamic gain scheme association network is obtained. Based on the dynamic gain scheme association network, the gain adjustment rules across time intervals are extracted, the overall consistency of the dynamic gain scheme is optimized, and the optimized dynamic gain scheme is determined as the final dynamic gain scheme. The final dynamic gain scheme includes the gain adjustment value, timestamp, adjustment basis and device adaptation instructions for each sampling time.

8. The audio dynamic gain optimization method based on audio features according to claim 1, characterized in that, The process of applying the dynamic gain scheme to the gain adjustment of audio equipment, performing gain optimization processing on the set of audio signals to be optimized, obtaining a set of audio output signals after gain optimization, and feeding back the characteristic temporal evolution information corresponding to the set of audio output signals to the gain response pattern library to improve the content of the pattern library includes: The dynamic gain scheme is analyzed, and the gain adjustment value and corresponding timestamp at each sampling time are extracted to generate a table showing the correspondence between the gain adjustment value and the timestamp. The gain adjustment values ​​in the corresponding table are converted into a control signal format that can be recognized by the audio equipment gain adjustment. Each gain adjustment value corresponds to a control signal, forming a control signal sequence. The control signals in the control signal sequence are arranged in order of timestamp. The control signal sequence is transmitted to the gain adjustment component of the audio equipment in chronological order, so that after receiving the control signal, the gain adjustment component adjusts its own gain parameter according to the gain adjustment value corresponding to the control signal. At each sampling time, the current gain parameter of the gain adjustment component is compared with the gain adjustment value at the corresponding sampling time in the dynamic gain scheme. If there is a difference, the gain parameter of the gain adjustment component is adjusted to be consistent with the gain adjustment value. The set of audio signals to be optimized is transmitted to the signal processing unit of the audio equipment in chronological order, so that after receiving the audio signal units, the signal processing unit transmits the audio signal units to the gain adjustment unit, and the gain adjustment unit performs gain amplification or attenuation processing on the audio signal units according to the gain parameters at the current time. The audio signal units processed by the gain adjustment unit are collected and sorted according to the timestamp order to form a preliminary audio output signal set. Based on the preliminary audio output signal set analysis, the transition characteristics of adjacent audio signal units are analyzed, and transition optimization parameters are output; Based on the transition optimization parameters, the initial audio output signal set is subjected to signal continuity processing, and the audio output signal set after continuity processing is used as the audio output signal set after gain optimization. Feature extraction is performed on the gain-optimized audio output signal set to construct the audio feature temporal evolution chain corresponding to the audio output signal set; Extract the audio feature time-series evolution chain corresponding to the set of audio signals to be optimized, the applied dynamic gain scheme, and the audio feature time-series evolution chain corresponding to the set of audio output signals, and integrate them into gain optimization effect correlation data; An optimization effect correlation graph is constructed, with the audio feature temporal evolution chain of the audio signal set to be optimized, the dynamic gain scheme, and the audio feature temporal evolution chain of the audio output signal set as graph nodes, and the edges between nodes are determined based on the causal correlation strength. Based on the correlation graph of the optimization effect, the key influencing factors of gain optimization are extracted, and the correlation data of gain optimization effect and key influencing factors are transmitted to the update interface of the gain response pattern library. The corresponding feature evolution law and gain adjustment rules are analyzed. If there is no corresponding gain response pattern in the gain response pattern library, a new gain response pattern is generated and added to the gain response pattern library; if there is already a corresponding pattern for the feature evolution law, the gain adjustment rules of the existing pattern are corrected. Record the updates to the gain response pattern library, including information on newly added gain response patterns or changes before and after correcting gain response patterns, and form a pattern library update log. Store the pattern library update log in a preset log storage area.

9. A dynamic gain optimization system for audio based on audio features, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the audio feature-based dynamic gain optimization method according to any one of claims 1 to 8 by executing the machine-executable instructions.

10. A computer program product, characterized in that, The computer program product includes machine-executable instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the machine-executable instructions from the computer-readable storage medium and executes the machine-executable instructions to cause the computer device to perform the audio feature-based dynamic gain optimization method as described in any one of claims 1 to 8.