An intelligent regulation method, system and storage medium for speaker audio signals
By conducting trial operation of speakers and ambient noise acquisition, segmenting the audio signal into short-time frames and performing feature analysis, the problem of failure to effectively consider the impact of environmental noise in the existing technology is solved, and more efficient and accurate audio signal regulation is achieved to ensure clear sound quality.
Patent Information
- Application Number
- CN202510312349.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The prior art usually focuses on processing and analysis of the audio signal itself in audio signal regulation, and fails to effectively consider the impact of environmental noise, resulting in the possibility of distortion of the audio signal or inaccurate regulation parameters.
By conducting trial operation of the speaker to be regulated at startup, the real-time performance of the voice coil and diaphragm is monitored, and the audio signal and ambient noise are collected. After initially enhancing the audio signal, it is divided into short-time frames, the frequency domain characteristics and formant peak parameters of each frame are calculated, and the clarity and sound quality evaluation values are comprehensively analyzed to generate targeted regulation.
It improves the processing efficiency and accuracy of audio signals, can more accurately capture and analyze the instantaneous characteristics of the signal, realize targeted regulation, ensure clearer sound quality, and avoid audio signal distortion caused by speaker performance defects.
Smart Images

Figure CN119835590B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio signal analysis, and particularly to an intelligent regulation method, system and storage medium for speaker audio signals. Background Art
[0002] With the popularization of digital audio technology, the processing and transmission of audio signals have become more flexible and complex, and more advanced regulation methods are required to ensure sound quality. The popularization of devices such as smart speakers and intelligent vehicle systems has led to higher and higher user requirements for audio quality, which requires more flexible regulation methods.
[0003] The prior art, such as the invention patent with publication number: CN114724580B, is an audio signal regulation method. The method includes: accessing an original audio signal, inputting the original audio signal into a post-processing module for adjustment to obtain a first audio signal; adjusting the gain of the original audio signal to obtain a second audio signal; filtering out medium and high frequency signals in the second audio signal to obtain a third audio signal; performing equalization calibration on the third audio signal to obtain a fourth audio signal; performing dynamic range control on the fourth audio signal to obtain a fifth audio signal; and performing clipping on the fifth audio signal to obtain a sixth audio signal. A speaker system includes a sound source input, a digital signal processor with multi-channel independent control, a power amplifier, a box body, and a first speaker unit and a second speaker unit respectively receiving two audio signals obtained by amplifying the first audio signal and the sixth audio signal respectively.
[0004] The prior art, such as the invention patent with publication number: CN101453532B, is a sound processing device for a speaker switch, including a microphone input signal envelope estimator, a microphone input signal background noise envelope estimator, a microphone input signal useful audio envelope estimator, a microphone input signal attenuator, a speaker receiving signal envelope estimator, a speaker receiving signal background noise envelope estimator, a speaker receiving signal useful audio envelope estimator, a speaker receiving signal attenuator, and a state machine for controlling the conversion between the above two states.
[0005] Based on the above solutions, it can be seen that the prior art in the field of audio signal regulation usually focuses on processing and analyzing the audio signal itself. However, in practical applications, the input and output of audio signals are affected by various factors. For example, due to environmental noise, if only the audio signal is analyzed and processed, audio signal distortion or inaccurate regulation parameters usually occur. Summary of the Invention
[0006] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention provides an intelligent regulation method, system and storage medium for a speaker audio signal, which can split the audio signal into short-time frames, improve the processing efficiency and accuracy, and achieve targeted regulation.
[0007] On the one hand, an embodiment of the present invention provides an intelligent regulation method for a speaker audio signal, including:
[0008] After receiving the start signal, the speaker to be regulated conducts a trial operation, monitors the real-time performance of the voice coil and the real-time performance of the speaker diaphragm of the speaker to be regulated during the trial operation, and obtains the hardware trial operation monitoring result of the speaker to be regulated;
[0009] When the hardware trial operation monitoring result of the speaker to be regulated is normal in performance, a permission formal start instruction is sent. At the same time, the audio signal to be played and the ambient noise of the speaker to be regulated are collected, the audio signal is initially enhanced in sound quality, and the enhanced audio signal is split into short-time frames according to a preset time period, which are recorded as each frame of audio signal;
[0010] Calculate the frequency-domain characteristics of each frame of audio signal, comprehensively analyze and process to generate a clarity evaluation value of the audio signal to be played. At the same time, extract the linear prediction coding coefficients, obtain the formant parameters of each frame of audio signal, and comprehensively analyze and process to obtain an audio quality evaluation value of the audio signal to be played;
[0011] Based on the clarity evaluation value and the audio quality evaluation value of the audio signal to be played, comprehensively analyze and process to obtain an audio signal regulation value set of the audio signal to be played and perform regulation and output.
[0012] According to some embodiments of the present invention, the specific process of monitoring the real-time performance of the voice coil and the real-time performance of the speaker diaphragm of the speaker to be regulated during the trial operation is as follows:
[0013] Monitor the real-time performance of the voice coil of the speaker to be regulated during the trial operation, and obtain the voice coil temperature performance parameters, including the temperature change rate, the highest temperature, the lowest temperature and the average temperature of the voice coil during the trial operation period, and analyze and process therefrom to obtain the thermal stability evaluation value of the voice coil of the speaker to be regulated;
[0014] Monitor the real-time performance of the speaker diaphragm of the speaker to be regulated during the trial operation, and obtain the speaker diaphragm performance parameters, including the amplitude change rate, the maximum amplitude, the minimum amplitude, the average amplitude and the average frequency response time of the speaker diaphragm during the trial operation period, and analyze and process therefrom to obtain the amplitude stability evaluation value of the speaker diaphragm of the speaker to be regulated.
[0015] According to some embodiments of the present invention, the specific process of obtaining the hardware trial operation monitoring result of the speaker to be regulated is as follows:
[0016] Based on the thermal stability evaluation value of the voice coil of the loudspeaker to be regulated and the amplitude stability evaluation value of the loudspeaker diaphragm, they are respectively compared and analyzed with the thermal stability evaluation threshold of the voice coil and the amplitude stability evaluation threshold of the loudspeaker diaphragm pre-stored in the database. When the thermal stability evaluation value of the voice coil of the loudspeaker to be regulated is greater than the thermal stability evaluation threshold and the amplitude stability evaluation value of the loudspeaker diaphragm of the loudspeaker to be regulated is greater than the amplitude stability evaluation threshold, it is determined that the hardware trial operation monitoring result of the loudspeaker to be regulated is normal in performance;
[0017] If the thermal stability evaluation value of the voice coil of the loudspeaker to be regulated is less than or equal to the thermal stability evaluation threshold, or the amplitude stability evaluation value of the loudspeaker diaphragm of the loudspeaker to be regulated is less than or equal to the amplitude stability evaluation threshold, it is determined that the hardware trial operation monitoring result of the loudspeaker to be regulated is abnormal in performance.
[0018] According to some embodiments of the present invention, collecting the audio signal to be played and the ambient noise of the loudspeaker to be regulated, performing preliminary sound quality enhancement on the audio signal, and splitting the enhanced audio signal into short-time frames according to a preset time period, specifically including:
[0019] After sending a permission formal start instruction to the control terminal of the loudspeaker to be regulated, connecting the audio source to be played to the audio processing unit of the loudspeaker to be regulated, and collecting the audio signal;
[0020] Based on the external configuration ambient noise sensor of the loudspeaker to be regulated, extracting the ambient noise data of the loudspeaker to be regulated within a preset period, including the average noise decibel value, equivalent continuous noise level, peak noise level, and noise spectrum, and analyzing and processing to obtain the ambient noise evaluation value of the loudspeaker to be regulated;
[0021] The ambient noise evaluation value is used to characterize the influence degree of the ambient noise on the audio signal to be played;
[0022] Based on the ambient noise evaluation value of the loudspeaker to be regulated, performing mapping matching with the set of sound quality enhancement parameters corresponding to each ambient noise evaluation value interval pre-stored in the database, obtaining the sound quality enhancement parameters of the audio signal to be played, and performing preliminary sound quality enhancement on the audio signal based on the sound quality enhancement parameters to obtain the enhanced audio signal;
[0023] Extracting the preset time period size for splitting from the database, splitting the enhanced audio signal into short-time frames, and recording them as each frame of audio signal.
[0024] According to some embodiments of the present invention, calculating the frequency domain characteristics of each frame of audio signal, and comprehensively analyzing and processing to generate the clarity evaluation value of the audio signal to be played, specifically including:
[0025] Calculate the frequency domain features of each frame of audio signal based on the fast Fourier transform, including the frequency center, bandwidth, and spectral flatness of each frame of audio signal. Analyze and process the frequency domain features of each frame of audio signal comprehensively to obtain the average bandwidth, frequency center volatility, and average spectral flatness of the audio signal to be played;
[0026] Extract the clear index parameter set from the database, including the bandwidth index, frequency center volatility index, and spectral flatness index. Compare and analyze the average bandwidth, frequency center volatility, and average spectral flatness of the audio signal to be played with the clear index parameter set to obtain the clarity evaluation value of the audio signal to be played.
[0027] According to some embodiments of the present invention, the extracting linear prediction coding coefficients to obtain the formant parameters of each frame of audio signal, and comprehensively analyzing and processing to obtain the audio quality evaluation value of the audio signal to be played, specifically including:
[0028] Calculate the autocorrelation coefficient of each frame of audio signal, and calculate again based on the autocorrelation coefficient of each frame of audio signal to obtain the linear prediction coding coefficient of each frame of audio signal;
[0029] Obtain the formant parameters of each frame of audio signal based on the linear prediction coding coefficient of each frame of audio signal, including formant position, formant bandwidth, and formant intensity. Analyze and process the formant parameters of each frame of audio signal comprehensively to obtain the formant position volatility, average formant bandwidth, and formant intensity volatility of the audio signal to be played;
[0030] Extract the audio quality index parameters from the database, including the formant position volatility index, average formant bandwidth index, and formant intensity volatility index. Compare and analyze the formant position volatility, average formant bandwidth, and formant intensity volatility of the audio signal to be played with the audio quality index parameters to obtain the audio quality evaluation value of the audio signal to be played.
[0031] According to some embodiments of the present invention, the comprehensively analyzing and processing to obtain the audio signal regulation value set of the audio signal to be played and performing regulation and output, the specific processing conditions are:
[0032] Based on the clarity evaluation value and audio quality evaluation value of the audio signal to be played, comprehensively analyze and process to obtain the comprehensive feature value of the audio signal to be played. Map and match the comprehensive feature value of the audio signal to be played with the audio signal regulation parameter set corresponding to each comprehensive feature value interval pre-stored in the database to obtain the regulation parameter set of the audio signal to be played, including high-frequency regulation value, low-frequency regulation value, and audio compression value;
[0033] Based on the regulation parameter set of the audio signal, perform advanced regulation on the audio signal to be played.
[0034] According to some embodiments of the present invention, the comprehensive analysis process obtains the comprehensive eigenvalue of the audio signal to be played, and the specific processing conditions are as follows:
[0035] ;
[0036] Among them, is the comprehensive eigenvalue of the audio signal to be played, is the clarity evaluation value of the audio signal to be played, is the audio quality evaluation value of the audio signal to be played, is the clarity weight factor, is the audio quality weight factor, is the activation function, .
[0037] In a second aspect, an intelligent regulation system for a speaker audio signal provided by an embodiment of the present invention includes:
[0038] A trial operation module, which is used to perform a trial operation after the speaker to be regulated receives a start signal, monitor the real-time performance of the voice coil and the real-time performance of the speaker diaphragm of the speaker to be regulated during the trial operation, and obtain the hardware trial operation monitoring result of the speaker to be regulated;
[0039] A preliminary sound quality enhancement module, which is used to send a permission formal start instruction when the hardware trial operation monitoring result of the speaker to be regulated is normal in performance. At the same time, it collects the audio signal to be played and the ambient noise of the speaker to be regulated, performs preliminary sound quality enhancement on the audio signal, and divides the enhanced audio signal into short-time frames according to a preset time period, denoted as each frame of audio signal;
[0040] An audio feature analysis module, which is used to calculate the frequency domain features of each frame of audio signal, comprehensively analyze and process to generate the clarity evaluation value of the audio signal to be played. At the same time, it extracts the linear predictive coding coefficients, obtains the formant parameters of each frame of audio signal, and comprehensively analyzes and processes to obtain the audio quality evaluation value of the audio signal to be played;
[0041] An audio regulation output module, which is used to comprehensively analyze and process based on the clarity evaluation value and the audio quality evaluation value of the audio signal to be played, obtain the audio signal regulation value set of the audio signal to be played, and perform regulation and output.
[0042] In a third aspect, an embodiment of the present invention provides a storage medium, including: the storage medium has one or more programs, and the one or more programs are executed by one or more processors.
[0043] Compared with the prior art, the embodiments of the present invention have at least the following beneficial effects:
[0044] (1)The present invention provides an intelligent regulation method for the audio signal of a loudspeaker. By conducting a trial operation on the loudspeaker to be regulated at startup and monitoring the real-time performance of the voice coil and the loudspeaker diaphragm, it can prevent the damage of the loudspeaker caused by overheating and avoid mechanical damage caused by excessive amplitude. At the same time, potential fault signs can be detected in advance, such as abnormal temperature rise or abnormal amplitude fluctuation, and to a large extent, it can avoid the audio signal distortion problem caused by the performance defects of the loudspeaker itself.
[0045] (2)The present invention divides the audio signal into short-time frames and then analyzes and processes each frame of the audio signal, which can more accurately capture and analyze the instantaneous characteristics of the signal, contribute to fine processing, improve the analysis efficiency and accuracy. By calculating the clarity evaluation value and audio quality evaluation value of the audio signal to be played through synthesizing the analysis results of each frame of the audio signal, the efficiency of audio analysis and processing is improved, and corresponding regulation processing can be carried out according to specific audio characteristics, which is beneficial to accelerating the analysis speed. At the same time, due to focusing on local characteristics, the analysis accuracy is also improved.
[0046] (3)The present invention collects the ambient noise of the loudspeaker to be regulated, analyzes and processes it to obtain the ambient noise evaluation value of the loudspeaker to be regulated. According to the ambient noise evaluation value, the audio signal can be adjusted specifically for sound quality enhancement to adapt to different noise environments and ensure that the sound quality of the actually output audio signal is clearer.
[0047] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic flow chart of the method of the present invention.
[0049] Figure 2 It is a schematic diagram of the system module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0051] In the description of the present invention, it should be understood that the meaning of "several" is one or more, the meaning of "multiple" is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, and "above", "below", "within", etc. are understood as including the present number. If there is a description of "first", "second", etc., it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features. The terms "opening", "upper", "lower", "thickness", "top", "middle", "length", "inner", "periphery", etc. indicate the orientation or position relationship, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the components or elements referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.
[0052] Please refer to Figure 1 As shown, an intelligent regulation method for the audio signal of a loudspeaker provided by an embodiment of the present invention includes:
[0053] After the loudspeaker to be regulated receives a start signal, it conducts a trial operation, monitors the real-time performance of the voice coil and the real-time performance of the loudspeaker diaphragm during the trial operation, and obtains the hardware trial operation monitoring result of the loudspeaker to be regulated.
[0054] The specific process of monitoring the real-time performance of the voice coil and the real-time performance of the loudspeaker diaphragm of the loudspeaker to be regulated during the trial operation is as follows:
[0055] Monitor the real-time performance of the voice coil of the loudspeaker to be regulated during the trial operation, and obtain the voice coil temperature performance parameters, including the temperature change rate, the highest temperature, the lowest temperature, and the average temperature of the voice coil during the trial operation cycle.
[0056] The temperature change rate refers to the change rate of the voice coil temperature during the trial operation cycle. It reflects the sensitivity and response speed of the voice coil to temperature changes. A higher temperature change rate usually means that the voice coil has experienced severe temperature fluctuations in a short period of time.
[0057] The highest temperature refers to the highest temperature value reached by the voice coil during the trial operation cycle. It indicates the thermal performance of the voice coil under extreme conditions. An excessively high highest temperature usually causes the aging, performance degradation, or damage of the voice coil material.
[0058] The lowest temperature refers to the lowest temperature value reached by the voice coil during the trial operation cycle. If the temperature is too low, it will cause the material to become brittle or the performance to change.
[0059] The average temperature refers to the average temperature value of the voice coil during the trial operation period, which provides the overall thermal performance of the voice coil during the trial operation period. An excessively high average temperature usually indicates insufficient heat dissipation or excessive power, while an excessively low average temperature usually indicates that the speaker is not being effectively utilized. When evaluating thermal stability, the average temperature should be maintained within a reasonable range to ensure the long-term stable operation of the voice coil.
[0060] Extract the maximum allowable temperature change rate of the voice coil from the database, and analyze and process it to obtain the thermal stability evaluation value of the voice coil of the speaker to be regulated, specifically including:
[0061] ;
[0062] Among them, is the thermal stability evaluation value of the voice coil of the speaker to be regulated, is the temperature change rate of the voice coil, is the maximum temperature of the voice coil, is the minimum temperature of the voice coil, is the average temperature of the voice coil, is the maximum allowable temperature change rate of the voice coil.
[0063] It should also be noted that there is a certain correlation among the temperature change rate, maximum temperature, minimum temperature, and average temperature of the voice coil during the trial operation period, and these correlations can reflect the thermal stability of the voice coil. The temperature change rate refers to the rate of temperature change within a certain period of time. If the temperature change rate is high, it usually means that the voice coil has experienced a rapid temperature rise or fall, which usually results in a higher maximum temperature or a lower minimum temperature. A large temperature change rate usually indicates that the speaker has experienced a large power change or thermal load during operation, which usually affects the thermal stability of the voice coil. The maximum temperature is the highest temperature value reached during the trial operation period, and the average temperature is the average value of the temperature throughout the period. If the maximum temperature is much higher than the average temperature, this usually indicates a transient overheating phenomenon, usually caused by instantaneous high power input or poor heat dissipation. The minimum temperature is the lowest temperature value reached during the trial operation period. If the minimum temperature is much lower than the average temperature, this usually indicates that the speaker has hardly been working during a certain period or the ambient temperature is low. The temperature change rate also affects the average temperature. If the temperature change rate is large, but the difference between the maximum temperature and the minimum temperature is small, then the average temperature is usually relatively stable. On the contrary, if the temperature change rate is small, but the difference between the maximum temperature and the minimum temperature is large, then the average temperature is usually greatly affected. These parameters together reflect the thermal performance of the voice coil.
[0064] Monitor the real-time performance of the speaker diaphragm of the speaker to be regulated during the trial operation, obtain the performance parameters of the speaker diaphragm, including the amplitude change rate, maximum amplitude, minimum amplitude, average amplitude, and average frequency response time of the speaker diaphragm within the trial operation period, and extract the maximum allowable amplitude change rate and the target average frequency response time from the database.
[0065] The amplitude change rate refers to the change in amplitude within the trial operation period. In an oscillatory phenomenon, the amplitude refers to the displacement from the equilibrium position to other positions. Subtract the amplitude at the last moment of the trial operation period from the amplitude at the beginning of the trial operation period to obtain the change in amplitude within the trial operation period. The amplitude change rate reflects the dynamic response ability and stability of the speaker diaphragm. When evaluating stability, a lower and stable amplitude change rate is an ideal choice.
[0066] The maximum amplitude refers to the maximum amplitude value recorded by the sensor within the trial operation period. It is used to indicate the maximum vibration amplitude that the speaker diaphragm can withstand.
[0067] The minimum amplitude refers to the minimum amplitude value of the speaker diaphragm that the sensor can record within the trial operation period. During the trial operation period, the test amplitude range covers the rated minimum amplitude value and the rated maximum amplitude value of the speaker diaphragm.
[0068] The average amplitude refers to the average value of the amplitude within the trial operation period. It provides the overall level of the amplitude. A stable average amplitude indicates that the speaker diaphragm performs stably under normal operating conditions.
[0069] The average frequency response time refers to the average time for the speaker diaphragm to respond to frequency changes. It reflects the frequency response characteristics of the speaker diaphragm, that is, the adaptability of the speaker diaphragm to different frequency signals. A shorter average frequency response time means that the speaker diaphragm can adapt to frequency changes faster, improving stability.
[0070] From this, analyze and process to obtain the amplitude stability evaluation value of the speaker diaphragm of the speaker to be regulated, specifically including:
[0071] ;
[0072] Among them, is the amplitude stability evaluation value of the speaker diaphragm of the speaker to be regulated, is the amplitude change rate of the speaker diaphragm, is the maximum amplitude of the speaker diaphragm, is the minimum amplitude of the speaker diaphragm, is the average amplitude of the speaker diaphragm, is the average frequency response time of the speaker diaphragm, is the maximum allowable amplitude change rate, is the target average frequency response time.
[0073] It should also be noted that there is a certain correlation among several parameters of the speaker diaphragm during the trial operation period, such as the amplitude change rate, the maximum amplitude, the minimum amplitude, the average amplitude, and the average frequency response time. A higher amplitude change rate means that the speaker diaphragm has experienced a rapid change from the minimum amplitude to the maximum amplitude. A stable amplitude change rate is usually related to a smaller maximum amplitude and a larger minimum amplitude, indicating that the speaker diaphragm is operating within the normal working range. The difference between the maximum amplitude and the minimum amplitude reflects the amplitude range that the speaker diaphragm can respond to. The average amplitude is the intermediate value between the maximum amplitude and the minimum amplitude, reflecting the overall amplitude level of the speaker diaphragm during the trial operation period. A stable average amplitude indicates that the speaker diaphragm performs stably under normal working conditions. A short average frequency response time means that the speaker diaphragm can quickly adapt to frequency changes, which is usually related to a lower amplitude change rate, indicating that the speaker diaphragm has a good adaptability to rapidly changing audio signals. A longer average frequency response time usually means that the speaker diaphragm responds slowly to frequency changes, and the average frequency response time is also usually related to the maximum amplitude and the minimum amplitude, especially when the amplitude responses of the speaker diaphragm at different frequencies are different. A fast response time usually means that the speaker diaphragm can maintain good performance at different amplitudes.
[0074] The specific process of obtaining the hardware trial operation monitoring result of the speaker to be regulated is as follows:
[0075] Based on the thermal stability evaluation value of the voice coil of the speaker to be regulated and the amplitude stability evaluation value of the speaker diaphragm, they are respectively compared and analyzed with the pre-stored thermal stability evaluation threshold of the voice coil and the amplitude stability evaluation threshold of the speaker diaphragm in the database. When the thermal stability evaluation value of the voice coil of the speaker to be regulated is greater than the thermal stability evaluation threshold and the amplitude stability evaluation value of the speaker diaphragm of the speaker to be regulated is greater than the amplitude stability evaluation threshold, it is determined that the hardware trial operation monitoring result of the speaker to be regulated is normal in performance.
[0076] If the thermal stability evaluation value of the voice coil of the speaker to be regulated is less than or equal to the thermal stability evaluation threshold, or the amplitude stability evaluation value of the speaker diaphragm of the speaker to be regulated is less than or equal to the amplitude stability evaluation threshold, it is determined that the hardware trial operation monitoring result of the speaker to be regulated is abnormal in performance.
[0077] The thermal stability evaluation threshold is processed as follows: In a specific embodiment, the thermal stability evaluation threshold is directly set in the database during the development of the intelligent control system for the speaker audio signal involved in the embodiments of the present invention. There are various methods for setting the thermal stability evaluation threshold. For example, it is obtained through statistical analysis, including statistically analyzing the thermal stability evaluation values under normal historical thermal stability evaluations (for example, the average temperature of the voice coil is less than 40 degrees Celsius), and performing mean processing on the thermal stability evaluation values obtained multiple times to obtain the thermal stability evaluation threshold.
[0078] The amplitude stability evaluation threshold is processed as follows: In a specific embodiment, the amplitude stability evaluation threshold is directly set in the database during the development of the intelligent control system for the speaker audio signal involved in the embodiments of the present invention. There are various methods for setting the amplitude stability evaluation threshold. For example, it is obtained through statistical analysis, including statistically analyzing the amplitude stability evaluation values under normal historical amplitude stability evaluations (for example, the average amplitude of the speaker diaphragm is less than 1 mm), and performing mean processing on the amplitude stability evaluation values obtained multiple times to obtain the amplitude stability evaluation threshold.
[0079] When the hardware trial operation monitoring result of the speaker to be controlled is normal in performance, a permission to officially start the instruction is sent. At the same time, the audio signal to be played and the ambient noise of the speaker to be controlled are collected, the audio signal is preliminarily denoised, and the enhanced audio signal is segmented into short-time frames according to a preset time period, which are recorded as each frame of the audio signal.
[0080] The collection of the audio signal to be played and the ambient noise of the speaker to be controlled, the preliminary denoising of the audio signal, and the segmentation of the enhanced audio signal into short-time frames according to a preset time period specifically include:
[0081] After sending the permission to officially start the instruction to the control terminal of the speaker to be controlled, connect the audio source to be played to the audio processing unit of the speaker to be controlled and collect the audio signal.
[0082] Based on the external configuration of the speaker to be controlled with an ambient noise sensor, extract the ambient noise data of the speaker to be controlled within a preset period, including the average noise decibel value, equivalent continuous noise level, peak noise level, and noise spectrum.
[0083] The average noise decibel value refers to the average value of the noise intensity within a preset period, usually in decibels (dB). It is used to evaluate the average noise level around the speaker to be controlled within a preset period.
[0084] The equivalent continuous noise level refers to converting each measured noise within a preset period into the corresponding sound power, calculating the average value of the sound power within the preset period, also in decibels (dB). It is usually used to represent the average noise level within the preset period. The equivalent continuous noise level is a commonly used indicator for evaluating the impact of noise on people and the environment.
[0085] It should be noted that the conversion of each measured noise within the preset period into the corresponding sound power specifically includes:
[0086] Converting the sound pressure level to sound pressure. The relationship between the sound pressure level Lp and the sound pressure p is: ;
[0087] where p is the sound pressure, is the reference sound pressure, usually 20 micropascals, and Lp is the sound pressure level.
[0088] The sound power is proportional to the square of the sound pressure. The sound power P can be calculated by the following formula:
[0089] ;
[0090] where, is the sound power, p is the sound pressure, is the acoustic impedance of the medium. For air, the acoustic impedance is 420 pascal·seconds / meter (at standard temperature and pressure).
[0091] For each measured noise level, calculate the corresponding sound power according to the above steps, and then average these sound power values within the preset period to obtain the equivalent sound power.
[0092] The peak noise level refers to the maximum sound pressure level reached by the noise at a certain moment within the preset period, in decibels (dB). The peak noise level is used to represent the instantaneous maximum intensity of the noise and is very important for evaluating the sudden and impactful effects of the noise.
[0093] The noise spectrum refers to the distribution of the noise at different frequencies, usually represented by a graph with frequency as the abscissa and sound pressure level as the ordinate. The noise spectrum is used to analyze the frequency components of the noise.
[0094] The analysis and processing to obtain the environmental noise evaluation value of the loudspeaker to be regulated specifically includes:
[0095] Based on the noise spectrum, analyze and process to obtain the frequency band width and spectrum slope of the environmental noise.
[0096] The frequency band width refers to the frequency range of the noise energy distribution, that is, the interval between the lowest frequency and the highest frequency covered by the noise. It is usually expressed in hertz (Hz).
[0097] The spectral slope refers to the rate of change of the sound pressure level with frequency in the noise spectrogram, usually expressed in decibels per octave (dB / octave) or decibels per decade (dB / decade). The spectral slope describes the attenuation or growth trend of the noise at different frequencies. A positive slope indicates that the sound pressure level increases as the frequency increases. A negative slope indicates that the sound pressure level decreases as the frequency increases.
[0098] Extract the average noise decibel indication value, equivalent continuous noise level indication value, peak noise level indication value, frequency band indication width, and spectral indication slope from the database, and jointly analyze and process them to obtain the environmental noise evaluation value of the loudspeaker to be regulated. The specific processing conditions are as follows:
[0099] ;
[0100] Among them, is the environmental noise evaluation value of the loudspeaker to be regulated, is the average noise decibel value of the environmental noise, is the equivalent continuous noise level of the environmental noise, is the peak noise level of the environmental noise, is the frequency band width of the environmental noise, is the spectral slope of the environmental noise, is the average noise decibel indication value, is the equivalent continuous noise level indication value, is the peak noise level indication value, is the frequency band indication width, is the spectral indication slope, is the weight factor of the average noise decibel value, is the weight factor of the equivalent continuous noise level, is the weight factor of the peak noise level, is the weight factor of the frequency band width, is the weight factor of the spectral slope. The softplus function is a built-in function in Python, .
[0101] It should be noted that the value ranges of the weight factors of the average noise decibel value, equivalent continuous noise level, peak noise level, frequency band width, and spectral slope are all between 0 and 1, and satisfy , the weight factor of the average noise decibel value is the influence factor corresponding to the environmental noise evaluation value of the loudspeaker to be regulated preset in the database, which represents the numerical value of the influence degree of the average noise decibel value on the environmental noise evaluation value of the loudspeaker to be regulated. The weight factor of the equivalent continuous noise level is the influence factor corresponding to the environmental noise evaluation value of the loudspeaker to be regulated preset in the database, which represents the numerical value of the influence degree of the equivalent continuous noise level on the environmental noise evaluation value of the loudspeaker to be regulated. The weight factor of the peak noise level is the influence factor corresponding to the environmental noise evaluation value of the loudspeaker to be regulated preset in the database, which represents the numerical value of the influence degree of the peak noise level on the environmental noise evaluation value of the loudspeaker to be regulated. The weight factor of the frequency band width is the influence factor corresponding to the environmental noise evaluation value of the loudspeaker to be regulated preset in the database, which represents the numerical value of the influence degree of the frequency band width on the environmental noise evaluation value of the loudspeaker to be regulated. The weight factor of the spectral slope is the influence factor corresponding to the environmental noise evaluation value of the loudspeaker to be regulated preset in the database, which represents the numerical value of the influence degree of the spectral slope on the environmental noise evaluation value of the loudspeaker to be regulated. When in use, the weight factor of the average noise decibel value, the weight factor of the equivalent continuous noise level, the weight factor of the peak noise level, the weight factor of the frequency band width and the weight factor of the spectral slope can be directly obtained from the database, and the obtaining method is the preset mapping relationship. For example: the average noise decibel value, the equivalent continuous noise level, the peak noise level, the frequency band width and the spectral slope involved in the embodiments of the present invention are input into the mapping relationship for mapping and matching to obtain the weight factor of the average noise decibel value, the weight factor of the equivalent continuous noise level, the weight factor of the peak noise level, the weight factor of the frequency band width and the weight factor of the spectral slope involved in the embodiments of the present invention, and the mapping relationship therein is one-to-one correspondence.
[0102] It should also be noted that there is a certain correlation among several parameters such as the average noise decibel value, equivalent continuous A-weighted sound pressure level (Leq), peak noise level, frequency band width, and spectral slope. The average noise decibel value reflects the overall intensity level of the noise. The equivalent continuous A-weighted sound pressure level (Leq) is also an indicator reflecting the overall intensity of the noise, but it takes into account the time distribution characteristics of the noise. The peak noise level reflects the instantaneous maximum intensity of the noise. The frequency band width refers to the frequency range covered by the noise. The wider the frequency band width, the more frequency components the noise contains. The spectral slope describes the inclination degree of the noise spectrum. Both the average noise decibel value and the equivalent continuous A-weighted sound pressure level are indicators for measuring the overall intensity of the noise and have a certain positive correlation. A high peak noise level does not necessarily mean a high average decibel value or equivalent continuous A-weighted sound pressure level because the peak reflects the instantaneous maximum value, while the average or equivalent value reflects the overall level. The wider the frequency band width, the more high-intensity frequency components the noise usually contains, resulting in an increase in the average decibel value or equivalent continuous A-weighted sound pressure level. The spectral slope affects the frequency distribution of the noise. For example, a steeper slope usually means stronger low-frequency components, which usually affects the average decibel value or equivalent continuous A-weighted sound pressure level of the noise. Noises with a wide frequency band width usually contain more high-intensity frequency components, thus usually having a higher peak noise level. The spectral slope affects the distribution of the frequency components of the noise, and thus affects the peak noise level. For example, if the high-frequency components are stronger, the peak noise level will be higher.
[0103] The environmental noise evaluation value is used to characterize the influence degree of environmental noise on the audio signal to be played.
[0104] Based on the environmental noise evaluation value of the loudspeaker to be regulated, map and match with the set of sound quality enhancement parameters corresponding to each environmental noise evaluation value interval stored in the database to obtain the sound quality enhancement parameters of the audio signal to be played, and perform preliminary sound quality enhancement on the audio signal based on the sound quality enhancement parameters to obtain the enhanced audio signal.
[0105] The sound quality enhancement parameters include volume enhancement parameters, dynamic range compression parameters, and reverberation enhancement parameters.
[0106] The volume enhancement parameter controls the amplification factor of the input signal. Increasing the gain can increase the volume.
[0107] The dynamic range compression parameter controls the proportional relationship between the input signal and the output signal. A higher compression ratio will result in a smaller dynamic range and a more balanced sound.
[0108] The reverberation enhancement parameter controls the duration of the reverberation effect. A longer reverberation time will make the sound fuller, and a shorter reverberation time will make the sound clearer.
[0109] For example, the set of sound quality enhancement parameters corresponding to the environmental noise evaluation value of a certain loudspeaker to be regulated is:
[0110] Volume enhancement parameter: +3dB;
[0111] Dynamic range compression parameter: compression ratio is 2:1, threshold is -80dB;
[0112] Reverberation enhancement parameter: reverberation time is 0.5 seconds;
[0113] Based on the sound quality enhancement parameters for the sound quality enhancement parameters, the specific process is as follows:
[0114] Overall increase the volume of the audio signal by 3dB to offset the influence of ambient noise. Apply a compression ratio of 2:1 and a threshold of -80dB to reduce the dynamic range of the audio signal and ensure that weak sounds in the ambient noise can also be heard. Add a reverberation time of 0.5 seconds to make the audio signal sound more full and natural.
[0115] The sound quality enhancement tool involved in the embodiments of the present invention is Audacity.
[0116] Extract the segmentation preset period size from the database, segment the enhanced audio signal into short-time frames, and record them as each frame of audio signal.
[0117] Calculate the frequency domain characteristics of each frame of audio signal, comprehensively analyze and process to generate the clarity evaluation value of the audio signal to be played. At the same time, extract the linear prediction coding coefficients, obtain the formant parameters of each frame of audio signal, and comprehensively analyze and process to obtain the audio quality evaluation value of the audio signal to be played.
[0118] The calculation of the frequency domain characteristics of each frame of audio signal, comprehensively analyze and process to generate the clarity evaluation value of the audio signal to be played, specifically includes:
[0119] Based on the fast Fourier transform, calculate the frequency domain characteristics of each frame of audio signal, including the frequency center, bandwidth, and spectral flatness of each frame of audio signal. Comprehensively analyze and process the frequency domain characteristics of each frame of audio signal to obtain the average bandwidth, frequency center volatility, and average spectral flatness of the audio signal to be played.
[0120] The fast Fourier transform (Fast Fourier Transform, FFT) is an efficient algorithm for calculating the discrete Fourier transform (Discrete Fourier Transform, DFT). DFT is a mathematical transform that converts a discrete time-domain signal into a discrete frequency-domain signal, while FFT provides a fast method for calculating DFT. FFT utilizes the symmetry and periodicity of DFT, and reduces the computational amount by decomposing the DFT of a long sequence into the DFT of short sequences. Common FFT algorithms include the Cooley-Tukey algorithm, which calculates the DFT by recursively dividing the sequence into two halves.
[0121] The frequency center is the energy center of the audio signal spectrum, representing the average frequency of the signal energy. It reflects the pitch of the audio signal, and a higher frequency center usually corresponds to a sharper pitch.
[0122] The bandwidth refers to the main frequency range in which the energy of the audio signal is distributed. The larger the bandwidth, the richer the frequency components contained in the signal, and the more complex the sound quality is usually.
[0123] Spectral flatness is an index that measures the uniformity of the energy distribution in the audio signal spectrum. A higher spectral flatness indicates that the energy is more evenly distributed across various frequencies and is usually related to noise. A lower spectral flatness indicates that the energy is concentrated on certain frequencies and is usually related to pure tones or specific pitches.
[0124] The average bandwidth refers to the average value of the bandwidths of each frame of the audio signal. It reflects the average frequency range of the entire audio signal.
[0125] The frequency center volatility refers to the degree of fluctuation of the frequency center in the time series. A large volatility usually indicates a large change in the pitch of the audio signal, which usually affects the stability of auditory perception. Calculate the power spectral density (PSD) to determine the frequency center of the signal at each time point. The frequency center is the frequency corresponding to the maximum value of the power spectral density. Calculate the standard deviation of the time-varying sequence of the frequency center. The larger the standard deviation, the greater the degree of fluctuation of the frequency center. For example, assume a time series data that records the frequency center once per second for 100 seconds. Calculate the standard deviation of the 100 frequency center values, and this standard deviation is the frequency center volatility.
[0126] The average spectral flatness refers to the average value of the spectral flatness of each frame of the audio signal. It reflects the average distribution of the spectral energy of the entire audio signal.
[0127] Extract a set of clear index parameters from the database, including bandwidth index, frequency center volatility index, and spectral flatness index. Compare and analyze the average bandwidth, frequency center volatility, and average spectral flatness of the audio signal to be played with the set of clear index parameters to obtain the clarity evaluation value of the audio signal to be played, specifically including:
[0128] ;
[0129] Among them, is the clarity evaluation value of the audio signal to be played, is the average bandwidth of the audio signal to be played, is the frequency center volatility of the audio signal to be played, is the average spectral flatness of the audio signal to be played, is the bandwidth index, is the frequency center volatility index, is the spectral flatness index, is the average bandwidth weight factor, is the frequency center volatility weight factor, is the average spectral flatness weight factor.
[0130] It should be noted that the value ranges of the average bandwidth weight factor, the frequency center volatility weight factor, and the average spectral flatness weight factor are all between 0 and 1, and satisfy , where the average bandwidth weight factor is the influence factor corresponding to the clarity evaluation value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the average bandwidth on the clarity evaluation value of the audio signal to be played. The frequency center volatility weight factor is the influence factor corresponding to the clarity evaluation value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the frequency center volatility on the clarity evaluation value of the audio signal to be played. The average spectral flatness weight factor is the influence factor corresponding to the clarity evaluation value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the average spectral flatness on the clarity evaluation value of the audio signal to be played. When in use, the average bandwidth weight factor, the frequency center volatility weight factor, and the average spectral flatness weight factor can be directly obtained from the database, and the acquisition method is a preset mapping relationship. For example: input the frequency center volatility, average spectral flatness, and average bandwidth involved in the embodiments of the present invention into the mapping relationship for mapping and matching to obtain the average bandwidth weight factor, the frequency center volatility weight factor, and the average spectral flatness weight factor involved in the embodiments of the present invention, and the mapping relationship therein is one-to-one correspondence.
[0131] It also should be noted that there is a certain correlation among the parameters of the average bandwidth, frequency center volatility, and average spectral flatness of the audio signal to be played. An audio signal with a larger average bandwidth usually has more frequency components, which usually leads to a larger fluctuation of the frequency center in time, thereby increasing the frequency center volatility. Conversely, if the frequency center volatility is larger, it usually also means that the audio signal contains different ranges of frequency components in different time frames, thereby affecting the calculation of the average bandwidth. An audio signal with a larger average bandwidth usually contains a wider range of frequency components, which usually leads to a higher spectral flatness because the energy is distributed over a wider frequency range. If the average spectral flatness is higher, it means that the energy distribution is more uniform, which usually means that the signal contains a wider frequency range, thereby affecting the average bandwidth. An audio signal with a large frequency center volatility usually indicates that the pitch or the main frequency component changes greatly in time, which usually affects the calculation of the spectral flatness because the uniformity of the energy distribution usually changes with the change of the frequency center. These three parameters affect each other and jointly determine the frequency domain characteristics of the audio signal.
[0132] Extracting the linear prediction coding coefficients, obtaining the formant parameters of each frame of audio signal, and comprehensively analyzing and processing to obtain the audio quality evaluation value of the audio signal to be played, specifically including:
[0133] Calculating the autocorrelation coefficients of each frame of audio signal, and based on the autocorrelation coefficients of each frame of audio signal, calculating again to obtain the linear prediction coding (LPC) coefficients of each frame of audio signal.
[0134] The autocorrelation coefficients of each frame of audio signal are a statistical measure of the similarity between each frame of audio signal and itself at different time delays. The autocorrelation coefficient can reveal the periodicity, stationarity, and noise characteristics of the signal. In audio signal processing, the autocorrelation coefficient is used to identify the fundamental frequency (i.e., pitch) of the audio. At the same time, the autocorrelation coefficient is the basis for calculating the linear prediction coding (LPC) coefficients.
[0135] The linear prediction coding (LPC) coefficients of each frame of audio signal are parameters used to simulate the audio generation mechanism. The LPC coefficients can effectively represent the main characteristics of the audio signal, and are used for audio compression and coding. They are also used to analyze the spectral structure and resonance characteristics of the audio signal.
[0136] Based on the linear prediction coding (LPC) coefficients of each frame of audio signal, using an LPC filter, obtaining the formant parameters of each frame of audio signal, including the formant position sequence, formant bandwidth, and formant intensity. Comprehensively analyzing and processing the formant parameters of each frame of audio signal to obtain the formant position volatility, average formant bandwidth, and formant intensity volatility of the audio signal to be played.
[0137] The formant position sequence refers to the frequency points where the energy is concentrated in the audio signal spectrum, usually corresponding to the resonance frequencies of the vocal tract. The formant position determines the timbre and tone color of the audio, and the formant position sequences of different vowels are different.
[0138] The formant bandwidth refers to the frequency range of the formant, that is, the width of the formant. The formant bandwidth is related to the damping characteristics of the vocal tract and affects the clarity and naturalness of the sound.
[0139] The formant intensity refers to the amplitude of the formant, which reflects the energy size of this frequency component. The formant intensity affects the loudness and perceived volume of the sound.
[0140] The formant position volatility refers to the degree of change of the formant position sequence in the audio signal over time. A small volatility indicates a stable sound, and a large volatility usually indicates an unstable sound or the presence of interference. Calculate the standard deviation of each formant position sequence as the formant position volatility.
[0141] The average formant bandwidth is the average of all formant bandwidths in the entire audio signal. The average formant bandwidth can be used to evaluate the overall clarity and naturalness of the sound.
[0142] The formant intensity volatility refers to the degree to which the formant intensity in the audio signal changes over time. A small intensity volatility indicates a stable sound, while a large volatility usually indicates a large dynamic range of the sound or the presence of noise. The standard deviation of each formant intensity is calculated as the formant intensity volatility.
[0143] The formant position and bandwidth directly affect the clarity of the audio. Stable formant positions and moderate bandwidths generally mean higher clarity.
[0144] The volatility of the formant position and intensity can reflect the stability of the sound. A small volatility indicates a stable sound, while a large volatility usually indicates an unstable sound or the presence of interference.
[0145] Extract the audio quality index parameters from the database, including the formant position volatility index, the average formant bandwidth index, and the formant intensity volatility index. Compare and analyze the formant position volatility, the average formant bandwidth, and the formant intensity volatility of the audio signal to be played with the audio quality index parameters to obtain the audio quality evaluation value of the audio signal to be played, specifically including:
[0146] ;
[0147] Among them, is the audio quality evaluation value of the audio signal to be played, is the formant position volatility of the audio signal to be played, is the average formant bandwidth of the audio signal to be played, is the formant intensity volatility of the audio signal to be played, is the formant position volatility index, is the average formant bandwidth index, is the formant intensity volatility index, is the formant position volatility weight factor, is the average formant bandwidth weight factor, is the formant intensity volatility weight factor.
[0148] It should be noted that the formant position volatility weight factor, the average formant bandwidth weight factor, and the formant intensity volatility weight factor all have a value range between 0 and 1, and satisfy , the formant position volatility weight factor is an influence factor corresponding to the audio quality evaluation value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the formant position volatility on the audio quality evaluation value of the audio signal to be played. The average formant bandwidth weight factor is an influence factor corresponding to the audio quality evaluation value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the average formant bandwidth on the audio quality evaluation value of the audio signal to be played. The formant intensity volatility weight factor is an influence factor corresponding to the audio quality evaluation value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the formant intensity volatility on the audio quality evaluation value of the audio signal to be played. When in use, the formant position volatility weight factor, the average formant bandwidth weight factor and the formant intensity volatility weight factor can be directly obtained from the database, and the obtaining method is a preset mapping relationship. For example: the formant position volatility, the average formant bandwidth and the formant intensity volatility involved in the embodiments of the present invention are input into the mapping relationship for mapping and matching to obtain the formant position volatility weight factor, the average formant bandwidth weight factor and the formant intensity volatility weight factor involved in the embodiments of the present invention, and the mapping relationship therein is one-to-one correspondence.
[0149] It should also be noted that there is a certain correlation among the formant position volatility, the average formant bandwidth and the formant intensity volatility of the audio signal to be played. A large formant position volatility usually means that the vocal tract shape or the pronunciation method is changing rapidly, which usually leads to a change in the formant bandwidth. A large formant position volatility is usually accompanied by a change in bandwidth, especially in fast speech or noisy environments. A large formant position volatility usually means that the dynamic range of the audio is large, which usually leads to fluctuations in the formant intensity. To a certain extent, both reflect the stability and consistency of the audio. Rapid audio changes or noise interference usually lead to fluctuations in both the formant position and intensity. A large average formant bandwidth usually means that the vocal tract attenuates frequency components relatively quickly, which usually leads to fluctuations in the formant intensity over time. A large average formant bandwidth usually makes the formant intensity more susceptible to environmental noise or vocal tract changes, resulting in intensity fluctuations. These three parameters together reflect the stability and consistency of the audio. A stable audio usually has a small formant position volatility, a moderate average formant bandwidth and a small formant intensity volatility.
[0150] Based on the clarity evaluation value and the audio quality evaluation value of the audio signal to be played, a set of audio signal regulation values of the audio signal to be played is obtained through comprehensive analysis and processing and output after regulation.
[0151] The specific processing conditions for obtaining a set of audio signal regulation values of the audio signal to be played through the above comprehensive analysis and processing and outputting after regulation are as follows:
[0152] Based on the clarity evaluation value and audio quality evaluation value of the audio signal to be played, a comprehensive feature value of the audio signal to be played is obtained through comprehensive analysis and processing, specifically including:
[0153] ;
[0154] Among them, is the comprehensive feature value of the audio signal to be played, is the clarity evaluation value of the audio signal to be played, is the audio quality evaluation value of the audio signal to be played, is the clarity weight factor, is the audio quality weight factor, is the activation function, .
[0155] It should be noted that the clarity weight factor and the audio quality weight factor both have a value range between 0 and 1 and satisfy , the clarity weight factor is an influencing factor of the comprehensive feature value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the clarity evaluation value on the comprehensive feature value of the audio signal to be played, and the audio quality weight factor is an influencing factor of the comprehensive feature value of the audio signal to be played preset in the database, representing the numerical value of the influence degree of the audio quality evaluation value on the comprehensive feature value of the audio signal to be played. When in use, the clarity weight factor and the audio quality weight factor can be directly obtained from the database, and the acquisition method is a preset mapping relationship. For example: the clarity evaluation value and audio quality evaluation value calculated in the embodiments of the present invention are input into the mapping relationship for mapping and matching to obtain the clarity weight factor and audio quality weight factor involved in the embodiments of the present invention, and the mapping relationship therein is one-to-one correspondence.
[0156] The comprehensive feature value of the audio signal to be played is mapped and matched with the set of audio signal regulation parameters corresponding to each comprehensive feature value interval pre-stored in the database to obtain the set of regulation parameters of the audio signal to be played, including the high-frequency regulation value, the low-frequency regulation value, and the audio compression value.
[0157] The high-frequency regulation value refers to the parameter for increasing or decreasing the high-frequency components (usually frequencies above 2 kHz) in the audio signal. Increasing the high-frequency gain can improve the clarity and detail performance of the audio and make the sound brighter.
[0158] The low-frequency regulation value refers to the parameter for increasing or decreasing the low-frequency components (usually frequencies below 200 Hz) in the audio signal. Increasing the low-frequency gain can make the audio sound more full and powerful and enhance the bass effect.
[0159] The audio compression value refers to the parameter for compressing the dynamic range of an audio signal. Compressing the dynamic range of the audio can make the volume more stable and avoid excessive volume fluctuations. Through compression, the overall volume of the audio can be increased, making weaker sounds more prominent.
[0160] For example, collect an audio signal to be played. The set of control parameters corresponding to the interval where the comprehensive characteristic value of this audio signal is located is: high-frequency control value +3dB, low-frequency control value +2dB, audio compression value -1dB.
[0161] Apply the obtained control parameters to the audio signal for corresponding sound quality adjustment, which specifically includes:
[0162] Increase the high frequency by 3dB to improve the clarity of the audio, increase the low frequency by 2dB to enhance the fullness of the audio, and reduce the audio compression value by 1dB to reduce distortion.
[0163] Based on the set of audio signal control parameters, perform advanced control on the audio signal to be played.
[0164] In this embodiment, the present invention provides an intelligent control system for a loudspeaker audio signal, which specifically includes:
[0165] A trial operation module, which is used to perform a trial operation after the loudspeaker to be controlled receives a start signal, monitor the real-time performance of the voice coil and the real-time performance of the loudspeaker diaphragm during the trial operation, and obtain the hardware trial operation monitoring result of the loudspeaker to be controlled.
[0166] A preliminary sound quality enhancement module, which is used to send a permission to officially start the instruction when the hardware trial operation monitoring result of the loudspeaker to be controlled is normal. At the same time, collect the audio signal to be played and the ambient noise of the loudspeaker to be controlled, perform preliminary sound quality enhancement on the audio signal, and divide the enhanced audio signal into short-time frames according to a preset time period, which are recorded as each frame of audio signal.
[0167] An audio feature analysis module, which is used to calculate the frequency domain features of each frame of audio signal, comprehensively analyze and process to generate a clarity evaluation value of the audio signal to be played. At the same time, extract the linear predictive coding coefficients, obtain the formant parameters of each frame of audio signal, and comprehensively analyze and process to obtain the audio quality evaluation value of the audio signal to be played.
[0168] An audio control output module, which is used to comprehensively analyze and process based on the clarity evaluation value and the audio quality evaluation value of the audio signal to be played to obtain a set of audio signal control values of the audio signal to be played and perform control and output.
[0169] A storage medium, including: the storage medium has one or more programs, and the one or more programs are executed by one or more processors to implement the method described in the above content.
[0170] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus.
[0171] The preferred embodiments of the present invention disclosed above are only used to assist in explaining the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention. As long as it does not deviate from the structure of the present invention or exceed the scope defined by the present invention, it should fall within the protection scope of the present invention.
Claims
1. An intelligent control method for a loudspeaker audio signal, characterized in that: include: The speaker to be controlled performs a trial run after receiving the start signal, monitors the real-time performance of the voice coil and the real-time performance of the speaker diaphragm of the speaker to be controlled during the trial run, and obtains the hardware trial run monitoring result of the speaker to be controlled; When the hardware trial operation monitoring result of the speaker to be controlled shows that the performance is normal, a permission formal start instruction is sent, and the audio signal to be played and the environmental noise of the speaker to be controlled are collected at the same time, and the audio signal is preliminarily enhanced in sound quality, and the enhanced audio signal is divided into short time frames according to a preset time period, which are recorded as each frame audio signal; Calculate the frequency domain features of each frame of the audio signal, comprehensively analyze and process to generate a clarity evaluation value of the audio signal to be played, and extract the linear prediction coding coefficient to obtain the formant parameters of each frame of the audio signal, and comprehensively analyze and process to obtain the audio quality evaluation value of the audio signal to be played; Based on the clarity evaluation value and the audio quality evaluation value of the audio signal to be played, a comprehensive analysis and processing is performed to obtain an audio signal control value set of the audio signal to be played, which is then controlled and output.
2. The intelligent control method of a loudspeaker audio signal according to claim 1, characterized in that: The real-time performance of the voice coil and the real-time performance of the speaker diaphragm to be regulated during the monitoring trial operation is specifically as follows: Monitor the real-time performance of the voice coil of the speaker to be regulated during the trial operation, and obtain the temperature performance parameters of the voice coil, including the temperature change rate, maximum temperature, minimum temperature and average temperature of the voice coil during the trial operation period, thereby analyzing and processing to obtain the thermal stability evaluation value of the voice coil of the speaker to be regulated; Monitor the real-time performance of the speaker diaphragm of the speaker to be controlled during the trial operation, and obtain the performance parameters of the speaker diaphragm, including the amplitude change rate, maximum amplitude, minimum amplitude, average amplitude and frequency response average time of the speaker diaphragm during the trial operation period, thereby analyzing and processing to obtain the amplitude stability evaluation value of the speaker diaphragm of the speaker to be controlled.
3. The intelligent control method of a loudspeaker audio signal according to claim 2, characterized in that: The specific process of obtaining the hardware trial operation monitoring result of the speaker to be controlled is as follows: Based on the thermal stability evaluation value of the voice coil of the speaker to be controlled and the amplitude stability evaluation value of the speaker diaphragm, a comparison analysis is performed with the thermal stability evaluation threshold of the voice coil and the amplitude stability evaluation threshold of the speaker diaphragm pre-stored in the database respectively; when the thermal stability evaluation value of the voice coil of the speaker to be controlled is greater than the thermal stability evaluation threshold and the amplitude stability evaluation value of the speaker diaphragm of the speaker to be controlled is greater than the amplitude stability evaluation threshold, it is determined that the hardware trial operation monitoring result of the speaker to be controlled is normal in performance; If the thermal stability evaluation value of the voice coil of the speaker to be controlled is less than or equal to the thermal stability evaluation threshold, or the amplitude stability evaluation value of the speaker diaphragm of the speaker to be controlled is less than or equal to the amplitude stability evaluation threshold, then the hardware trial operation monitoring result of the speaker to be controlled is determined to be performance abnormality.
4. The intelligent control method of a loudspeaker audio signal according to claim 1, characterized in that: The collecting of the audio signal to be played and the ambient noise of the speaker to be adjusted, performing preliminary sound quality enhancement on the audio signal, and dividing the enhanced audio signal into short time frames according to a preset time period specifically includes: After sending a permission formal start instruction to the control terminal of the speaker to be controlled, the audio source to be played is connected to the audio processing unit of the speaker to be controlled to collect audio signals; Based on the externally configured environmental noise sensor of the speaker to be controlled, the environmental noise data of the speaker to be controlled within a preset period is extracted, including the average decibel value of the noise, the equivalent continuous noise level, the peak noise level and the noise spectrum, and the environmental noise evaluation value of the speaker to be controlled is obtained by analysis and processing; The environmental noise evaluation value is used to characterize the degree of influence of the environmental noise on the audio signal to be played; Based on the environmental noise evaluation value of the speaker to be controlled, a mapping and matching is performed with a sound quality enhancement parameter set corresponding to each environmental noise evaluation value interval pre-stored in the database to obtain a sound quality enhancement parameter of the audio signal to be played, and the audio signal is preliminarily enhanced in sound quality based on the sound quality enhancement parameter to obtain an enhanced audio signal; The preset time period size is extracted from the database, and the enhanced audio signal is divided into short time frames, which are recorded as each frame audio signal.
5. The intelligent control method of a loudspeaker audio signal according to claim 1, characterized in that: The step of calculating the frequency domain features of each frame of the audio signal and comprehensively analyzing and processing to generate a clarity evaluation value of the audio signal to be played specifically includes: The frequency domain characteristics of each frame of audio signal are calculated based on fast Fourier transform, including the frequency center, bandwidth and spectral flatness of each frame of audio signal, and the frequency domain characteristics of each frame of audio signal are analyzed and processed to obtain the average bandwidth, frequency center fluctuation rate and average spectral flatness of the audio signal to be played; A set of clarity index parameters is extracted from the database, including bandwidth index, frequency center fluctuation index and spectral flatness index. The average bandwidth, frequency center fluctuation rate and average spectral flatness of the audio signal to be played are compared and analyzed with the clarity index parameter set to obtain a clarity evaluation value of the audio signal to be played.
6. The intelligent control method of a loudspeaker audio signal according to claim 1, characterized in that: The extracting of linear prediction coding coefficients to obtain the formant parameters of each frame of the audio signal and the comprehensive analysis and processing to obtain the audio quality evaluation value of the audio signal to be played specifically include: Calculating the autocorrelation coefficient of each frame of audio signal, and calculating again based on the autocorrelation coefficient of each frame of audio signal to obtain the linear prediction coding coefficient of each frame of audio signal; Based on the linear prediction coding coefficients of each frame of audio signal, the formant parameters of each frame of audio signal are obtained, including the formant position, the formant bandwidth and the formant intensity, and the formant parameters of each frame of audio signal are analyzed and processed to obtain the formant position fluctuation rate, the average formant bandwidth and the formant intensity fluctuation rate of the audio signal to be played; Audio quality index parameters are extracted from a database, including a resonance peak position fluctuation rate index, an average resonance peak bandwidth index and a resonance peak intensity fluctuation rate index. The resonance peak position fluctuation rate, the average resonance peak bandwidth and the resonance peak intensity fluctuation rate of the audio signal to be played are compared and analyzed with the audio quality index parameters to obtain an audio quality evaluation value of the audio signal to be played.
7. The intelligent control method of a loudspeaker audio signal according to claim 1, characterized in that: The comprehensive analysis and processing obtains the audio signal control value set of the audio signal to be played and outputs it after control. The specific processing conditions are: Based on the clarity evaluation value and the audio quality evaluation value of the audio signal to be played, a comprehensive characteristic value of the audio signal to be played is obtained by comprehensive analysis and processing, and the comprehensive characteristic value of the audio signal to be played is mapped and matched with a set of audio signal control parameters corresponding to each comprehensive characteristic value interval pre-stored in the database to obtain a set of control parameters of the audio signal to be played, including a high-frequency control value, a low-frequency control value and an audio compression value; Based on the control parameter set of the audio signal, the audio signal to be played is advanced controlled.
8. The intelligent control method of a loudspeaker audio signal according to claim 7, characterized in that: The comprehensive analysis process obtains the comprehensive characteristic value of the audio signal to be played, and the specific processing conditions are: ; in, is the comprehensive characteristic value of the audio signal to be played, is the clarity evaluation value of the audio signal to be played, is the audio quality evaluation value of the audio signal to be played, is the clarity weight factor, is the audio quality weight factor, is the activation function, .
9. An intelligent control system for loudspeaker audio signals, characterized in that: A trial operation module is used to perform a trial operation on the speaker to be controlled after receiving a start signal, monitor the real-time performance of the voice coil and the real-time performance of the speaker diaphragm of the speaker to be controlled during the trial operation, and obtain the hardware trial operation monitoring result of the speaker to be controlled; The sound quality preliminary enhancement module is used to send a permission formal start instruction when the hardware trial operation monitoring result of the speaker to be controlled is normal, collect the audio signal to be played and the environmental noise of the speaker to be controlled, perform preliminary sound quality enhancement on the audio signal, and divide the enhanced audio signal into short time frames according to a preset time period, which are recorded as each frame audio signal; The audio feature analysis module is used to calculate the frequency domain features of each frame of the audio signal, comprehensively analyze and process to generate the clarity evaluation value of the audio signal to be played, and extract the linear prediction coding coefficient to obtain the formant parameters of each frame of the audio signal, and comprehensively analyze and process to obtain the audio quality evaluation value of the audio signal to be played; The audio control output module is used to comprehensively analyze and process the clarity evaluation value and audio quality evaluation value of the audio signal to be played, obtain the audio signal control value set of the audio signal to be played, and output it after control.
10. A storage medium, characterized in that: include: The storage medium has one or more programs, and the one or more programs are executed by one or more processors to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Sound processing equipment used in loudspeaker switch
CN101453532B
Audio signal control method and speaker system
CN114724580B
Audio power management system
CN102196336A
Improved speech intelligibility
CN106257584A