Human voice dynamic calibration method and device for voice test and storage medium

By dynamically determining the external amplifier gain and software gain on the vehicle, and combining audio correction coefficients and dynamic gain coefficients, the problem of low accuracy in voice testing under the fixed sound pressure level calibration method is solved, realizing automated calibration and improving testing accuracy and efficiency.

CN121306098APending Publication Date: 2026-01-09DONGFENG MOTOR CO LTD DONGFENG NISSAN PASSENGER VEHICLE CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511572401.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-30
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In existing technologies, the calibration method with fixed sound pressure levels cannot fully cover different user scenarios, which affects the accuracy of voice tests under different ambient sound conditions and fails to truly reflect the voice recognition effect in actual use scenarios.

Method used

By acquiring the recorded audio from the recording equipment on the vehicle, the external power amplifier gain and software gain are dynamically determined. Combined with the audio correction coefficient and dynamic gain coefficient, the human voice is dynamically calibrated, and the target software gain and external power amplifier gain are automatically adjusted to adapt to different ambient sound conditions.

Benefits of technology

It achieves accurate speech recognition under various background music and volume changes, improves test accuracy and consistency, reduces manual intervention, and increases calibration efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306098A_ABST
    Figure CN121306098A_ABST
Patent Text Reader

Abstract

The invention discloses a human voice dynamic calibration method and device for voice testing and a storage medium, and relates to the technical field of voice testing, and the method comprises the steps: in a current human voice audio playing process, obtaining a recording audio collected by a recording device on a vehicle, and determining an external power amplifier gain according to the recording audio; determining an initial software gain based on the current human voice audio and an external power amplifier gain; determining an audio correction coefficient according to the original voice audio to be played and the initial software gain; determining an audio dynamic gain coefficient based on the multimedia audio to be played; determining a target software gain according to the audio correction coefficient and the audio dynamic gain coefficient; human voice dynamic calibration is performed according to the external power amplifier gain and the target software gain, the target software gain can be accurately determined by comprehensively considering the external power amplifier gain, the initial software gain, the audio correction coefficient and the audio dynamic gain coefficient, and automatic measurement and calibration of audio are realized according to the target software gain and the external power amplifier gain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech testing technology, and in particular to a method, apparatus and storage medium for dynamic calibration of human voice for speech testing. Background Technology

[0002] The development of voice recognition and other functions in vehicles requires testing by playing back human voices using dummies or audio playback devices such as speakers. However, current calibration processes typically rely on manual calibration, using fixed sound pressure levels. According to the Lombard effect, as ambient noise intensifies, people increase their vocal volume to adapt. Since in-vehicle media background noise dynamically changes with different music, volumes, and other factors, the current fixed sound pressure testing method cannot comprehensively cover diverse user scenarios. This results in the accuracy of voice testing being affected under varying ambient noise conditions, failing to accurately reflect the voice recognition performance in real-world usage scenarios. Summary of the Invention

[0003] The main purpose of this application is to provide a method, device and storage medium for dynamic calibration of human voice for speech testing, which aims to solve the technical problem that the calibration method with fixed sound pressure value is difficult to meet the needs of dynamic testing, resulting in low accuracy of speech testing.

[0004] To achieve the above objectives, this application proposes a method for dynamic voice calibration in speech testing, the method comprising: During the current playback of human voice audio, the recorded audio collected by the recording device on the vehicle is acquired, and the gain of the external power amplifier is determined based on the recorded audio. The initial software gain is determined based on the current human voice audio and the external power amplifier gain. The audio correction coefficients are determined based on the original audio of the human voice to be played and the initial software gain. Determine the audio dynamic gain coefficient based on the multimedia audio to be played; The target software gain is determined based on the audio correction coefficient and the audio dynamic gain coefficient. The human voice dynamic calibration is performed based on the external power amplifier gain and the target software gain.

[0005] In one embodiment, the step of determining the initial software gain based on the current human voice audio and the external power amplifier gain includes: Acquire standard audio, wherein the standard audio is obtained by recording after calibrating the current human voice audio to reach a calibration state; Calculate the first root mean square value of the current human voice audio and the second root mean square value of the standard audio within a preset time window; The initial software gain is determined based on the first root mean square value, the second root mean square value, and the external power amplifier gain.

[0006] In one embodiment, the step of determining the initial software gain based on the first root mean square value, the second root mean square value, and the external power amplifier gain includes: The first root mean square value and the second root mean square value are obtained respectively based on the first root mean square value and the second root mean square value; Obtain the root mean square maximum value of the current human voice audio, the root mean square maximum value of the standard audio, and the first relationship between the external power amplifier gain and the initial software gain; The initial software gain is calculated based on the first relationship, the first root mean square maximum value, the second root mean square maximum value, and the external power amplifier gain.

[0007] In one embodiment, the step of determining the audio correction coefficient based on the original human voice audio to be played and the initial software gain includes: Obtain the third root mean square value data of the original human voice audio to be played in multiple preset time windows; The maximum value of the third root mean square is obtained based on the third root mean square value data; Get the first root mean square maximum value of the current human voice audio; The audio correction coefficients are calculated based on the first root mean square maximum value, the initial software gain, and the third root mean square maximum value.

[0008] In one embodiment, the step of determining the audio dynamic gain coefficient based on the multimedia audio to be played includes: During the playback of the multimedia audio to be played, multimedia audio data is collected; Calculate the sound pressure level at the human ear based on the multimedia audio data; The audio dynamic gain coefficient is calculated based on the sound pressure amplitude.

[0009] In one embodiment, the step of calculating the audio dynamic gain coefficient based on the sound pressure amplitude includes: The maximum sound pressure amplitude value is obtained based on the sound pressure amplitude value; Obtain the sound pressure threshold of the human ear affected by ambient noise; The audio dynamic gain coefficient is determined based on the maximum sound pressure amplitude and the sound pressure threshold.

[0010] In one embodiment, the step of performing dynamic calibration of human voice based on the external power amplifier gain and the target software gain includes: The gain of the external playback device is adjusted according to the gain of the external power amplifier. The gain of the in-vehicle audio player is adjusted according to the target software gain; Dynamic calibration of human voices is performed based on the adjusted external playback device and the in-vehicle audio player.

[0011] In one embodiment, the step of determining the external power amplifier gain based on the recorded audio includes: Obtain the reference root mean square value for human voice reconstruction; Calculate the fourth root mean square value of the recorded audio within a preset time window; The fourth root mean square value is obtained based on the fourth root mean square value; The external power amplifier gain is calculated based on the reference root mean square value and the fourth root mean square maximum value.

[0012] Furthermore, to achieve the above objectives, this application also proposes a voice dynamic calibration device for speech testing, the voice dynamic calibration device for speech testing comprising: The steps for determining the external power amplifier gain based on the recorded audio include: Obtain the reference root mean square value for human voice reconstruction; Calculate the fourth root mean square value of the recorded audio within a preset time window; The fourth root mean square value is obtained based on the fourth root mean square value; The external power amplifier gain is calculated based on the reference root mean square value and the fourth root mean square maximum value.

[0013] In addition, to achieve the above objectives, this application also proposes a voice dynamic calibration device for speech testing, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice dynamic calibration method for speech testing as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the human voice dynamic calibration method for speech testing as described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the human voice dynamic calibration method for speech testing as described above.

[0016] This application proposes one or more technical solutions that, during the current playback of human voice audio, acquire recorded audio from a recording device in the vehicle and determine the external power amplifier gain based on the recorded audio; determine an initial software gain based on the current human voice audio and the external power amplifier gain; determine an audio correction coefficient based on the original human voice audio to be played and the initial software gain; determine an audio dynamic gain coefficient based on the multimedia audio to be played; determine a target software gain based on the audio correction coefficient and the audio dynamic gain coefficient; and perform dynamic calibration of the human voice based on the external power amplifier gain and the target software gain. By comprehensively considering the external power amplifier gain, the initial software gain, the audio correction coefficient, and the audio dynamic gain coefficient, the target software gain can be accurately determined, effectively adapting to the voice testing needs under different ambient sound conditions. By dynamically adjusting the target software gain and the external power amplifier gain, accurate voice recognition results are ensured under various background music and volume changes. This achieves dynamic calibration of the test human voice according to the in-vehicle audio environment, which is more in line with user scenarios. Furthermore, it automates the calibration process from manual calibration, improving testing accuracy and reducing manual intervention, thus increasing calibration efficiency and consistency. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the method for dynamic voice calibration used in speech testing according to this application. Figure 2 A structural diagram of a voice dynamic calibration system for speech testing provided in an embodiment of the voice dynamic calibration method for speech testing according to this application; Figure 3 This is a schematic diagram of the overall architecture of a voice dynamic calibration system for speech testing, provided in an embodiment of the voice dynamic calibration method for speech testing according to this application. Figure 4 This is a flowchart illustrating Embodiment 2 of the method for dynamic voice calibration for speech testing in this application. Figure 5 This is a flowchart illustrating Embodiment 3 of the method for dynamic voice calibration for speech testing in this application; Figure 6 This is a flowchart illustrating Embodiment 4 of the method for dynamic voice calibration for speech testing in this application; Figure 7 A simplified flowchart is provided for an embodiment of the human voice dynamic calibration method for speech testing in this application; Figure 8 This is a schematic diagram of the module structure of the voice dynamic calibration device for voice testing according to an embodiment of this application; Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the voice dynamic calibration method for voice testing in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] The main solution of this application embodiment is as follows: during the current human voice audio playback, acquire the recorded audio collected by the recording device on the vehicle, and determine the external power amplifier gain based on the recorded audio; determine the initial software gain based on the current human voice audio and the external power amplifier gain; determine the audio correction coefficient based on the original human voice audio to be played and the initial software gain; determine the audio dynamic gain coefficient based on the multimedia audio to be played; determine the target software gain based on the audio correction coefficient and the audio dynamic gain coefficient; and perform human voice dynamic calibration based on the external power amplifier gain and the target software gain.

[0024] Current technologies primarily rely on external, independent testing equipment for dynamic voice calibration during speech testing. This involves playing voice data using a separate computer, audio playback software, and audio playback device. During each test, the gain (K) of the audio playback software and the gain (S) of the audio playback device are manually adjusted to bring the test audio to the calibration state. However, due to the influence of microphone, recording environment, and speaker factors during recording, consistency cannot be guaranteed. A fixed sound pressure level calibration method is insufficient for dynamic testing requirements, thus impacting the performance optimization and user experience of the speech recognition system. Therefore, current methods require manual calibration of K and S before playing each individual audio segment. Specifically, this involves manually measuring and adjusting K and S multiple times at the audio playback device output until calibration is complete. This manual calibration method is inefficient and prone to errors due to human factors, making it difficult to guarantee the consistency and reliability of the calibration results.

[0025] This application provides a solution that, through the design of the in-vehicle infotainment (IVI) system and audio playback device, deploys a recording device in the vehicle to collect human voice audio during actual playback, and then uses this audio data to dynamically determine the external power amplifier gain and software gain parameters, achieving a technological breakthrough from fixed calibration to dynamic calibration. This enables the test human voice to be dynamically calibrated according to the in-vehicle audio environment, which is more in line with user scenarios and improves test accuracy.

[0026] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone; or an electronic device capable of performing the above functions, or a voice dynamic calibration device for voice testing, such as a controller in a vehicle. The following description uses a controller in a vehicle as an example to illustrate this embodiment and the subsequent embodiments.

[0027] Based on this, embodiments of this application provide a method for dynamic calibration of human voices in speech testing, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the human voice dynamic calibration method for speech testing according to this application.

[0028] In this embodiment, the method for dynamic voice calibration for speech testing includes steps S10 to S60: Step S10: During the current human voice audio playback, acquire the recorded audio collected by the recording device on the vehicle, and determine the external power amplifier gain based on the recorded audio.

[0029] It should be noted that currently, when performing voice dynamic calibration for speech testing, external, independent testing equipment is primarily used. This involves playing voices using a separate computer, audio playback software, and audio playback device. This necessitates manually adjusting the gain of both the audio playback software and the audio playback device for each test, resulting in low efficiency and a high risk of errors. However, in the embodiments of this application, as shown... Figure 2 As shown, Figure 2 This is a structural diagram of the human voice dynamic calibration system used for voice testing in this embodiment. It mainly includes a playback device, a recording device, a computing device, a storage device, an external device driver, and a control device. Audio playback and acquisition are performed through the vehicle's own microphone and the audio playback device. An audio output path is designed within the controller to drive the external playback device to produce sound, achieving a closed-loop test. An audio reference signal acquisition path is designed to serve as a prerequisite for adjusting the input signal of the external playback device within the SOC chip for calculation. The audio playback device controls the playback of human voice audio. The recording device is a functional module that performs recording through the vehicle's microphone, recording the audio being played for calculation. The computing device calculates the gain of the audio playback device and the external audio playback device, as well as audio feature calculations, based on standard audio, the recording, and the reference signal. The storage device stores software default parameters, latency for different vehicle models, volume curves, and other inherent information, as well as data storage during software operation. The driver device drives the external audio playback device to produce sound, and the control device schedules the operation of the aforementioned devices. Based on this audio data, the external power amplifier gain and software gain are dynamically calculated without manual intervention, greatly improving calibration efficiency and accuracy. Figure 3 As shown, Figure 3 This is a schematic diagram of the overall architecture of a voice dynamic calibration system used for voice testing. The in-vehicle infotainment system (IVI) includes multimedia apps and VR devices for voice recognition and other applications. The media playback and dummy voice driving are controlled through the in-vehicle app. The system performs audio gain, playback, and recording. The microphone in the vehicle can collect the dummy's voice and media sound, which are sent to the application through the system's hardware abstraction layer (HAL) and framework to achieve dynamic audio calibration. At the same time, external audio playback devices can also communicate with the dummy via USB.

[0030] It should be noted that the in-vehicle infotainment system contains audio data from various human voices, which are then aggregated to form the original audio to be played. Any segment of this original audio can be selected for playback, representing the current audio clip. The audio clip is the audio data emitted by the human voice system.

[0031] In practice, the original audio of the voice to be played is first stored in a storage device. When playback is needed, the controller selects the corresponding audio from the storage device according to the instruction. During playback, the recording equipment on the vehicle collects the audio data in real time to obtain a recorded audio.

[0032] It should be noted that the current human voice audio h[n] is an arbitrarily selected human voice audio from the original audio Ui to be played. Audio h[n] is used as the audio to be played, and is played through the audio player and playback device. Simultaneously, the vehicle-mounted microphone records the played sound again, resulting in the recorded audio file P[n]. It should also be noted that in this process, the experimenter only needs to select audio h[n] and execute the start command; other audio playback and recording operations are automatically executed by the internal control and recording equipment. The recorded audio includes not only the currently playing human voice audio but may also be affected by factors such as ambient noise and background music within the vehicle. To calibrate the human voice audio, it is necessary to determine the software gain K and the external playback device gain S. During system startup, to ensure the human voice system can produce sound correctly for subsequent calibration, default values ​​K0 and S0 need to be set for K and S. The default value setting requirement is: 0. <K0<K m ;0 <S0<S m , where K m S m These are the system default parameters, stored in the storage device, ensuring that the amplitude of the audio system is within a reasonable range for human voices and that no audio distortion occurs. During the playback of the current human voice audio, the default gain is applied, and the recording device synchronously begins recording until the human voice audio playback is complete, at which point the recording ends, resulting in the recorded audio P[n]. During audio playback and synchronous recording, the software considers the delay Δt in sound propagation in space, i.e., the time difference between when the host starts emitting sound and when the host again picks up the audio through the microphone. This depends on the spatial distance and the software processing time; the time Δt is determined by different vehicle models and the computing power of the vehicle's infotainment system, and is a fixed value for each vehicle model. The external amplifier gain is the gain S of the external audio playback device.

[0033] In one feasible implementation, the step of determining the external power amplifier gain in step S10 may include steps A11 to A14: Step A11: Obtain the reference root mean square value for voice reconstruction; It should be noted that the reference root mean square value R for human voice reconstruction s The reference RMS for human voice reconstruction can be set in advance and serves as a benchmark value for subsequent calculations to measure the energy level of human voice audio.

[0034] Step A12: Calculate the fourth root mean square value of the recorded audio within a preset time window; The preset time window can be set according to actual needs, such as a time period of a few seconds or tens of seconds. In this embodiment, the preset time window is set to t, and m sampling points can be collected within time t. The RMS value is calculated at each time interval t. x i This represents audio data, specifically recorded audio data, from which the RMS value of the recorded audio is obtained, denoted as... .

[0035] Step A13: Obtain the maximum value of the fourth root mean square based on the fourth root mean square value; In practice, all the calculated fourth root mean square values ​​can be compared to obtain the maximum fourth root mean square value. .

[0036] Step A14: Calculate the external power amplifier gain based on the reference root mean square value and the fourth root mean square maximum value.

[0037] In practical implementation, the external power amplifier gain can be calculated based on the reference root mean square value and the fourth root mean square maximum value, S= / .

[0038] Step S20: Determine the initial software gain based on the current human voice audio and the external power amplifier gain.

[0039] In practice, the initial software gain can be obtained by calculating the audio characteristics based on the current human voice audio h[n] and the external power amplifier gain S.

[0040] This includes calculating the root mean square (RMS) value between the current human voice audio and the standard audio, and then calculating the initial software gain based on the RMS value and the external power amplifier gain to ensure that the audio can achieve the best results in subsequent processing.

[0041] Step S30: Determine the audio correction coefficient based on the original audio of the human voice to be played and the initial software gain.

[0042] It should be noted that the original audio to be played consists of audio data of different human voices. The audio correction coefficient is the correction coefficient of different human voices in a quiet scene. By calculating the audio correction coefficient, the calibration of different audio files in a quiet scene can be achieved.

[0043] Step S40: Determine the audio dynamic gain coefficient based on the multimedia audio to be played.

[0044] It is understandable that the multimedia audio to be played is the audio played by the vehicle's audio playback devices, such as speakers. For example, playing different music through the vehicle's audio system results in multimedia audio.

[0045] The audio dynamic gain coefficient is a coefficient dynamically calibrated in a music scene. It can be obtained by acquiring the digital signal of the audio currently being played by the speaker through the control device, thereby obtaining the audio data to be played, and then calculating the audio dynamic gain coefficient based on the audio data to be played. The corresponding audio dynamic gain coefficient is different in different music scenes.

[0046] Step S50: Determine the target software gain based on the audio correction coefficient and the audio dynamic gain coefficient.

[0047] In practical implementation, the target software gain can be calculated based on the audio correction coefficient and the audio dynamic gain coefficient. The target software gain is the final gain that needs to be adjusted by the software, and the target software gain K = K i *K r , where K i K is the audio correction factor. r This represents the audio dynamic gain coefficient.

[0048] Step S60: Perform dynamic calibration of human voice based on the external power amplifier gain and the target software gain.

[0049] It should be noted that after obtaining the external amplifier gain S and the target software gain K, the human voice output process passes through gains S and K respectively, thus achieving automated calibration of the tested human voice audio and enabling dynamic calibration based on changes in the in-vehicle ambient sound. During the test, the audio needs to undergo two gain processes: ① software gain from the audio playback module inside the controller, with a gain value of K; ② built-in gain from the external audio playback device, with a gain value of S.

[0050] In one feasible implementation, step S60 may include steps A21 to A23: Step A21: Adjust the gain of the external playback device according to the gain of the external power amplifier; Step A22: Adjust the gain of the in-vehicle audio player according to the target software gain; It should be noted that during the test, the human voice needs to be amplified twice. The gain of the external playback device is adjusted by the gain S of the external power amplifier, and at the same time, the target software gain K is applied to the audio player inside the controller.

[0051] Step A23: Perform dynamic calibration of human voices based on the adjusted external playback device and the in-vehicle audio player.

[0052] Therefore, the human voice can be dynamically calibrated based on the adjusted external playback device and audio player. For example, if the human voice audio to be played is h[n], the audio h[n] is played through the control module and the gain K of the audio player and the gain S of the external power amplifier device are adjusted until the audio reaches the calibration state. At this time, the audio in the calibration state is recorded through the recording device to obtain the recording file H[n]. The RMS and other audio characteristics of H[n] are obtained through the calculation module and used as the standard for human voice audio calibration.

[0053] This embodiment provides a method for dynamic calibration of human voices for speech testing. During the playback of a current human voice audio, the method acquires recorded audio from a recording device in a vehicle and determines the external power amplifier gain based on the recorded audio. An initial software gain is determined based on the current human voice audio and the external power amplifier gain. An audio correction coefficient is determined based on the original human voice audio to be played and the initial software gain. An audio dynamic gain coefficient is determined based on the multimedia audio to be played. A target software gain is determined based on the audio correction coefficient and the audio dynamic gain coefficient. Finally, dynamic calibration of the human voice is performed based on the external power amplifier gain and the target software gain. By comprehensively considering the external amplifier gain, initial software gain, audio correction coefficient, and audio dynamic gain coefficient, the target software gain can be accurately determined, effectively adapting to the voice testing needs under different ambient sound conditions. By dynamically adjusting the target software gain and external amplifier gain, accurate voice recognition results can be obtained under various background music and volume changes. The test voice is dynamically calibrated according to the in-vehicle audio environment, which is more in line with user scenarios. Furthermore, the current manual calibration has been automated, improving test accuracy. The automated calibration process reduces manual intervention and improves calibration efficiency and consistency.

[0054] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 4 Step S20 includes steps S201 to S203: Step S201: Obtain standard audio, wherein the standard audio is obtained by recording after calibrating the current human voice audio to reach the calibration state.

[0055] It should be noted that during the development process, in order to quickly calibrate different human voices, it is necessary to first set a standard for human voice calibration. After the standard is set, audio gain can be calculated in other vehicles or environments to achieve calibration in other vehicles or environments. Therefore, the standard audio is obtained by playing audio h[n] through the control module and adjusting the gain of the audio playback module and the gain of external devices until the audio reaches the calibration state. At this time, the audio in the calibration state is recorded through the recording device to obtain the recording file H[n], which is the standard audio.

[0056] It should be noted that, under ideal conditions, the audio obtained after gaining the recorded audio is the standard audio.

[0057] Step S202: Calculate the first root mean square value of the current human voice audio and the second root mean square value of the standard audio within a preset time window.

[0058] In practical implementation, the preset time window is set to t, and there are m sampling points within time t. Using the formula for calculating the root mean square value mentioned above, the first root mean square value of the current human voice audio and the second root mean square value of the standard audio can be calculated at time intervals of t. The first root mean square value is represented as { The second root mean square value is expressed as .

[0059] Step S203: Determine the initial software gain based on the first root mean square value, the second root mean square value, and the external power amplifier gain.

[0060] Understandably, it can be achieved through { }、 The initial software gain is calculated together with the external power amplifier gain S.

[0061] In one feasible implementation, step S203 may include steps B11-B13: Step B11: Obtain the first root mean square maximum value and the second root mean square maximum value based on the first root mean square value and the second root mean square value, respectively; It should be noted that the first root mean square values ​​can be compared, and the largest root mean square value can be selected, which is the maximum value of the first root mean square. Similarly, we obtain the second root mean square maximum value. .

[0062] Step B12: Obtain the root mean square maximum value of the current human voice audio, the root mean square maximum value of the standard audio, and the first relationship between the external power amplifier gain and the initial software gain; It should be noted that there is a primary relationship between the current root mean square maximum value of the human voice audio, the root mean square maximum value of the standard audio, the external power amplifier gain, and the initial software gain, expressed as: .

[0063] Step B13: Calculate the initial software gain based on the first relationship, the first root mean square maximum value, the second root mean square maximum value, and the external power amplifier gain.

[0064] In practical implementation, due to the first root mean square maximum value The second root mean square maximum value Since the external power amplifier gain S is known, the initial software gain can be calculated. The calculation is as follows:

[0065] This embodiment acquires standard audio, which is obtained by recording the current human voice audio after calibration to a calibration state. Within a preset time window, the first root mean square (RMS) value of the current human voice audio and the second RMS value of the standard audio are calculated. An initial software gain is determined based on the first RMS value, the second RMS value, and the external power amplifier gain. By introducing the standard audio as a reference, and combining the RMS value of the current human voice audio with the external power amplifier gain, the initial software gain is accurately calculated. Specifically, the standard audio setting provides a stable audio energy reference for the system, ensuring consistency in calibration under different vehicles or environments. By calculating the RMS values ​​of the current human voice audio and the standard audio within a preset time window, the dynamic characteristics of the audio signal can be captured, and then, combined with the external power amplifier gain, the initial software gain can be accurately calculated. This process not only improves the accuracy of calibration but also lays a solid foundation for subsequent dynamic audio calibration, ensuring the stability and reliability of the entire voice testing system under different environments.

[0066] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 Step S30 includes steps S301 to S304: Step S301: Obtain the third root mean square value data of the original human voice audio to be played in multiple preset time windows.

[0067] It should be noted that in a quiet environment, initial calibration is required based on the audio's amplitude and gain coefficient to meet the testing requirements. The calibration method involves obtaining the maximum audio value and ensuring that this maximum value matches the expected maximum audio value for calibration.

[0068] Therefore, for the original audio to be played, the RMS value of each time window of the original audio Ui can be obtained by preset time window t, and the third root mean square value data can be obtained by summing them up.

[0069] Step S302: Obtain the maximum value of the third root mean square based on the third root mean square value data.

[0070] In practical implementation, the maximum RMS value of each human voice audio segment can be selected from the third root mean square (RMS) value data, i.e., the third root mean square maximum value. .

[0071] Step S303: Obtain the first root mean square maximum value of the current human voice audio.

[0072] Understandably, calculating the audio correction coefficient requires not only obtaining the third root mean square maximum value of the original audio to be played, but also obtaining the first root mean square maximum value of the current audio within a preset time window. This value reflects the peak energy of the current human voice audio.

[0073] Step S304: Calculate the audio correction coefficient based on the first root mean square maximum value, the initial software gain, and the third root mean square maximum value.

[0074] It is understandable that the first root mean square maximum value is obtained. Initial software gain and the third root mean square maximum value The process of calculating the audio correction coefficient is as follows:

[0075] In the above formula, Audio correction coefficients, specifically the correction coefficients for different human voices in a quiet field, are used to characterize the difference in energy level between the original human voice audio to be played and the current human voice audio. By calculating these audio correction coefficients, accurate calibration of different human voices in quiet scenarios can be achieved, ensuring the accuracy and consistency of the speech testing system under various human voice conditions. This process not only improves the flexibility of calibration but also provides an important reference for subsequent dynamic audio calibration.

[0076] This embodiment acquires the third root mean square (RMS) values ​​of the original audio to be played over multiple preset time windows; obtains the third RMS maximum value based on the RMS data; acquires the first RMS maximum value of the current audio; and calculates the audio correction coefficient based on the first RMS maximum value, the initial software gain, and the third RMS maximum value. By introducing the third RMS values ​​of the original audio to be played and the first RMS maximum value of the current audio, combined with the initial software gain, the accurate calculation of the audio correction coefficient is achieved. The audio correction coefficient is calculated using a formula based on the initial software gain, and this coefficient accurately characterizes the difference in energy level between the original audio to be played and the current audio. This process not only improves the accuracy of calibration but also provides an important reference for subsequent dynamic audio calibration, ensuring the stability and reliability of the voice testing system under different voice conditions.

[0077] Based on the first embodiment of this application, in the fourth embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 6Step S40 includes steps S401 to S403: Step S401: During the playback of the multimedia audio to be played, collect multimedia audio data.

[0078] It should be noted that the multimedia audio to be played is the music data played by the speakers in the vehicle. During the playback of the multimedia audio r[n], multimedia audio data is collected at time intervals of T.

[0079] Step S402: Calculate the sound pressure amplitude at the human ear based on the multimedia audio data.

[0080] It should be noted that the sound pressure level at the human ear within time t can be calculated using control and computing devices. Specifically, the sound pressure level yi at the human ear can be calculated based on the volume curve yi=f(r[n]).

[0081] Step S403: Calculate the audio dynamic gain coefficient based on the sound pressure amplitude.

[0082] It should be noted that after obtaining the sound pressure level (SPL) amplitude at the human ear, the audio dynamic gain coefficient Kr can be calculated, for example, based on a preset correspondence between SPL amplitude and gain. Specifically, the system internally presets gain adjustment values ​​corresponding to different SPL amplitude ranges. By searching or interpolating, the audio dynamic gain coefficient that matches the current SPL amplitude is determined. This process ensures that the audio dynamic gain coefficient accurately reflects the impact of ambient sound on speech testing, providing a key parameter for subsequent dynamic speech calibration. With the introduction of the audio dynamic gain coefficient, the system can automatically adapt to different background music and volume changes, maintaining the accuracy and stability of speech recognition.

[0083] In one feasible implementation, step S403 may include steps C11 to C13: Step C11: Obtain the maximum sound pressure amplitude value based on the sound pressure amplitude value; It should be noted that after calculating the sound pressure amplitude at the human ear, the maximum sound pressure amplitude Y at the human ear within time T can be obtained, Y=max{yi}.

[0084] Step C12: Obtain the sound pressure threshold of the human ear affected by ambient noise; It should be noted that the sound pressure thresholds that affect the human ear from ambient noise can be preset, including a minimum sound pressure threshold and a maximum sound pressure threshold. The minimum sound pressure threshold is... The highest sound pressure threshold is .

[0085] Step C13: Determine the audio dynamic gain coefficient based on the maximum sound pressure amplitude and the sound pressure threshold.

[0086] It should be noted that the maximum sound pressure amplitude Y and the sound pressure threshold can be used as a reference. and The relationship between the maximum sound pressure level Y and the sound pressure threshold determines the audio dynamic gain coefficient. and When the magnitude relationship between the two is different, the resulting audio dynamic gain coefficient will be different.

[0087] Specifically, based on the sound pressure level at the human ear and the Lombard effect, the audio dynamic gain coefficient is determined in a music scene as follows:

[0088] It is understandable that if the maximum sound pressure amplitude Y is greater than or equal to the highest sound pressure threshold... Then the audio dynamic gain coefficient = If the maximum sound pressure amplitude Y is within the sound pressure threshold and Between, the audio dynamic gain coefficient = If the maximum sound pressure amplitude Y is less than or equal to the minimum sound pressure threshold Then the audio dynamic gain coefficient =1. Therefore, the gain coefficient of the human voice Ui can be dynamically adjusted via the multimedia audio r[n] to be played. .

[0089] This embodiment collects multimedia audio data during playback; calculates the sound pressure level (SPL) at the human ear based on the multimedia audio data; and calculates the dynamic gain coefficient of the audio based on the SPL. This achieves dynamic perception and quantification of the impact of ambient noise. By accurately calculating the dynamic gain coefficient of the audio, the system can effectively compensate for the impact of different background music and volume changes on speech testing. Specifically, when the ambient noise is strong, the system automatically reduces the gain to avoid speech signal distortion; when the ambient noise is weak, the system maintains or appropriately increases the gain to ensure speech signal clarity. This dynamic adjustment mechanism significantly improves the adaptability and reliability of the speech testing system in complex acoustic environments, providing more accurate testing assurance for application scenarios such as in-vehicle voice interaction.

[0090] For example, to help understand the implementation process of the voice dynamic calibration method for speech testing obtained by combining this embodiment with the above embodiment one, please refer to... Figure 7 , Figure 7A simplified flowchart of a method for dynamic voice calibration in speech testing is provided. Specifically: standard audio settings are performed at the vehicle end, followed by audio playback and recording. Audio features are calculated based on the collected data to obtain the audio playback device gain S, and an initial software gain K is calculated based on the audio playback device gain S. h After obtaining the initial software gain K h Then, calculate other voice correction coefficients K. i After the music signal is acquired, the dynamic gain coefficient K is calculated. r Finally, the target software gain K is obtained, and the target software gain K and the external power amplifier gain S are output.

[0091] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the human voice dynamic calibration method used in this application for speech testing. Any simple modifications based on this technical concept are within the protection scope of this application.

[0092] This application also provides a human voice dynamic calibration device for speech testing; please refer to... Figure 8 The voice dynamic calibration device for voice testing includes: The acquisition module 10 is used to acquire the recorded audio collected by the recording device on the vehicle during the current human voice audio playback, and to determine the external power amplifier gain based on the recorded audio.

[0093] The determination module 20 is used to determine the initial software gain based on the current human voice audio and the external power amplifier gain.

[0094] The determining module 20 is further configured to determine the audio correction coefficient based on the original human voice audio to be played and the initial software gain.

[0095] The determining module 20 is also used to determine the audio dynamic gain coefficient based on the multimedia audio to be played.

[0096] The determining module 20 is further configured to determine the target software gain based on the audio correction coefficient and the audio dynamic gain coefficient.

[0097] The calibration module 30 is used to perform dynamic calibration of human voice based on the external power amplifier gain and the target software gain.

[0098] The voice dynamic calibration device for speech testing provided in this application employs the voice dynamic calibration method for speech testing described in the above embodiments, which can solve the technical problem that the calibration method with fixed sound pressure levels is difficult to meet the requirements of dynamic testing, resulting in low accuracy in speech testing. Compared with the prior art, the beneficial effects of the voice dynamic calibration device for speech testing provided in this application are the same as those of the voice dynamic calibration method for speech testing provided in the above embodiments, and other technical features in the voice dynamic calibration device for speech testing are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0099] In one embodiment, the determining module 20 is further configured to acquire standard audio, wherein the standard audio is obtained by recording after calibrating the current human voice audio to reach a calibration state; Calculate the first root mean square value of the current human voice audio and the second root mean square value of the standard audio within a preset time window; The initial software gain is determined based on the first root mean square value, the second root mean square value, and the external power amplifier gain.

[0100] In one embodiment, the determining module 20 is further configured to obtain a first root mean square maximum value and a second root mean square maximum value based on the first root mean square value and the second root mean square value, respectively. Obtain the root mean square maximum value of the current human voice audio, the root mean square maximum value of the standard audio, and the first relationship between the external power amplifier gain and the initial software gain; The initial software gain is calculated based on the first relationship, the first root mean square maximum value, the second root mean square maximum value, and the external power amplifier gain.

[0101] In one embodiment, the determining module 20 is further configured to acquire the third root mean square value data of the original audio to be played over multiple preset time windows; The maximum value of the third root mean square is obtained based on the third root mean square value data; Get the first root mean square maximum value of the current human voice audio; The audio correction coefficients are calculated based on the first root mean square maximum value, the initial software gain, and the third root mean square maximum value.

[0102] In one embodiment, the determining module 20 is further configured to collect multimedia audio data during the playback of the multimedia audio to be played; Calculate the sound pressure level at the human ear based on the multimedia audio data; The audio dynamic gain coefficient is calculated based on the sound pressure amplitude.

[0103] In one embodiment, the determining module 20 is further configured to obtain the maximum sound pressure amplitude value based on the sound pressure amplitude value; Obtain the sound pressure threshold of the human ear affected by ambient noise; The audio dynamic gain coefficient is determined based on the maximum sound pressure amplitude and the sound pressure threshold.

[0104] In one embodiment, the calibration module 30 is further configured to adjust the gain of the external playback device according to the gain of the external power amplifier; The gain of the in-vehicle audio player is adjusted according to the target software gain; Dynamic calibration of human voices is performed based on the adjusted external playback device and the in-vehicle audio player.

[0105] In one embodiment, the acquisition module 10 is further configured to acquire a reference root mean square value for human voice reconstruction; Calculate the fourth root mean square value of the recorded audio within a preset time window; The fourth root mean square value is obtained based on the fourth root mean square value; The external power amplifier gain is calculated based on the reference root mean square value and the fourth root mean square maximum value.

[0106] This application provides a voice dynamic calibration device for voice testing. The voice dynamic calibration device for voice testing includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the voice dynamic calibration method for voice testing in the above embodiment 1.

[0107] The following is for reference. Figure 9 This document illustrates a structural schematic diagram of a voice dynamic calibration device suitable for implementing embodiments of this application. The voice dynamic calibration device for voice testing in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 9 The illustrated voice dynamic calibration device for voice testing is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0108] like Figure 9As shown, a voice dynamic calibration device for voice testing may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the voice dynamic calibration device for voice testing. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the voice dynamic calibration device used for voice testing to exchange data wirelessly or via wired communication with other devices. Although the figure shows a voice dynamic calibration device for voice testing with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0109] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0110] The voice dynamic calibration device for speech testing provided in this application employs the voice dynamic calibration method for speech testing described in the above embodiments, which solves the technical problem that the calibration method with fixed sound pressure levels is difficult to meet the requirements of dynamic testing, resulting in low accuracy in speech testing. Compared with the prior art, the beneficial effects of the voice dynamic calibration device for speech testing provided in this application are the same as those of the voice dynamic calibration method for speech testing provided in the above embodiments, and other technical features of the voice dynamic calibration device for speech testing are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0111] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0112] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0113] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the human voice dynamic calibration method for speech testing in the above embodiments.

[0114] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or Flash Memory), optical fibers, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0115] The aforementioned computer-readable storage medium may be included in a voice dynamic calibration device for speech testing; or it may exist independently and not assembled into a voice dynamic calibration device for speech testing.

[0116] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a voice dynamic calibration device for voice testing, cause the voice dynamic calibration device for voice testing to: acquire recorded audio from a recording device on a vehicle during the current playback of voice audio, and determine the external power amplifier gain based on the recorded audio; determine an initial software gain based on the current voice audio and the external power amplifier gain; determine an audio correction coefficient based on the original voice audio to be played and the initial software gain; determine an audio dynamic gain coefficient based on the multimedia audio to be played; determine a target software gain based on the audio correction coefficient and the audio dynamic gain coefficient; and perform voice dynamic calibration based on the external power amplifier gain and the target software gain.

[0117] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0119] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0120] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for dynamic voice calibration in speech testing. This solves the technical problem that calibration methods using fixed sound pressure levels are insufficient to meet the demands of dynamic testing, leading to low accuracy in speech testing. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the dynamic voice calibration method for speech testing provided in the above embodiments, and will not be repeated here.

[0121] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method for dynamic voice calibration for speech testing.

[0122] The computer program product provided in this application can solve the technical problem that the calibration method with fixed sound pressure levels is difficult to meet the needs of dynamic testing, resulting in low accuracy of speech testing. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the human voice dynamic calibration method for speech testing provided in the above embodiments, and will not be repeated here.

[0123] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A method for dynamic calibration of human voice for speech testing, characterized in that, The method for dynamic voice calibration used in speech testing includes: During the current playback of human voice audio, the recorded audio collected by the recording device on the vehicle is acquired, and the gain of the external power amplifier is determined based on the recorded audio. The initial software gain is determined based on the current human voice audio and the external power amplifier gain. The audio correction coefficients are determined based on the original audio of the human voice to be played and the initial software gain. Determine the audio dynamic gain coefficient based on the multimedia audio to be played; The target software gain is determined based on the audio correction coefficient and the audio dynamic gain coefficient. The human voice dynamic calibration is performed based on the external power amplifier gain and the target software gain.

2. The method as described in claim 1, characterized in that, The step of determining the initial software gain based on the current human voice audio and the external power amplifier gain includes: Acquire standard audio, wherein the standard audio is obtained by recording after calibrating the current human voice audio to reach a calibration state; Calculate the first root mean square value of the current human voice audio and the second root mean square value of the standard audio within a preset time window; The initial software gain is determined based on the first root mean square value, the second root mean square value, and the external power amplifier gain.

3. The method as described in claim 2, characterized in that, The step of determining the initial software gain based on the first root mean square value, the second root mean square value, and the external power amplifier gain includes: The first root mean square value and the second root mean square value are obtained respectively based on the first root mean square value and the second root mean square value; Obtain the root mean square maximum value of the current human voice audio, the root mean square maximum value of the standard audio, and the first relationship between the external power amplifier gain and the initial software gain; The initial software gain is calculated based on the first relationship, the first root mean square maximum value, the second root mean square maximum value, and the external power amplifier gain.

4. The method as described in claim 1, characterized in that, The step of determining the audio correction coefficient based on the original human voice audio to be played and the initial software gain includes: Obtain the third root mean square value data of the original human voice audio to be played in multiple preset time windows; The maximum value of the third root mean square is obtained based on the third root mean square value data; Get the first root mean square maximum value of the current human voice audio; The audio correction coefficients are calculated based on the first root mean square maximum value, the initial software gain, and the third root mean square maximum value.

5. The method as described in claim 1, characterized in that, The step of determining the audio dynamic gain coefficient based on the multimedia audio to be played includes: During the playback of the multimedia audio to be played, multimedia audio data is collected; Calculate the sound pressure level at the human ear based on the multimedia audio data; The audio dynamic gain coefficient is calculated based on the sound pressure amplitude.

6. The method as described in claim 5, characterized in that, The step of calculating the audio dynamic gain coefficient based on the sound pressure amplitude includes: The maximum sound pressure amplitude value is obtained based on the sound pressure amplitude value; Obtain the sound pressure threshold of the human ear affected by ambient noise; The audio dynamic gain coefficient is determined based on the maximum sound pressure amplitude and the sound pressure threshold.

7. The method as described in claim 1, characterized in that, The step of performing dynamic calibration of human voice based on the external power amplifier gain and the target software gain includes: The gain of the external playback device is adjusted according to the gain of the external power amplifier. The gain of the in-vehicle audio player is adjusted according to the target software gain; Dynamic calibration of human voices is performed based on the adjusted external playback device and the in-vehicle audio player.

8. The method as described in claim 1, characterized in that, The steps for determining the external power amplifier gain based on the recorded audio include: Obtain the reference root mean square value for human voice reconstruction; Calculate the fourth root mean square value of the recorded audio within a preset time window; The fourth root mean square value is obtained based on the fourth root mean square value; The external power amplifier gain is calculated based on the reference root mean square value and the fourth root mean square maximum value.

9. A human voice dynamic calibration device for speech testing, characterized in that, The device includes: The acquisition module is used to acquire the recorded audio collected by the recording device on the vehicle during the current human voice audio playback, and determine the external power amplifier gain based on the recorded audio. The determination module is used to determine the initial software gain based on the current human voice audio and the external power amplifier gain; The determining module is further configured to determine the audio correction coefficient based on the original human voice audio to be played and the initial software gain; The determining module is also used to determine the audio dynamic gain coefficient based on the multimedia audio to be played. The determining module is further configured to determine the target software gain based on the audio correction coefficient and the audio dynamic gain coefficient; The calibration module is used to perform dynamic calibration of human voice based on the external power amplifier gain and the target software gain.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the method for dynamic calibration of human voice for speech testing as described in any one of claims 1 to 8.