Playing control method and device, electronic equipment, storage medium and program product
By obtaining the frequency response deviation information of the audio playback device and using the preset frequency response information for equalization control, the problem of poor audio playback effect is solved, and the audio playback effect and test accuracy are improved efficiently and at a low cost.
Patent Information
- Application Number
- CN202511039167.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-25
- Publication Date
- 2025-09-26
AI Technical Summary
The audio playback effect of the existing audio playback devices is poor, and the use of audio analyzers for equalization control is costly, the equipment is redundant and inflexible, and it is difficult to meet the automated testing requirements of rich voice interaction scenarios.
By obtaining the deviation information between the actual frequency response of the target playback device and the target frequency response, the target audio signal is determined using the preset frequency response information, and the playback device is controlled to play the target audio signal, achieving balanced control and reducing costs.
It improves audio playback effects, reduces equalization control costs, simplifies equipment configuration, adapts to multi-channel equalization control needs, and improves test accuracy and flexibility.
Smart Images

Figure CN120708643A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the application of acoustic technology in the field of automotive intelligent cockpits, and in particular to a playback control method, device, electronic device, storage medium, and program product. Background Art
[0002] With the development of smart devices, various audio playback devices have also developed accordingly. In related application scenarios, it is usually necessary to perform equalization control on the audio playback device to improve the audio playback effect. Summary of the Invention
[0003] In order to overcome the problem of poor audio playback effect in the related art, the present disclosure provides a playback control method, device, electronic device, storage medium and program product.
[0004] According to a first aspect of an embodiment of the present disclosure, a playback control method is provided, comprising: obtaining preset frequency response information and an audio signal to be played, the preset frequency response information being used to represent a deviation between an actual frequency response of a target playback device and a target frequency response; determining a target audio signal based on the preset frequency response information and the audio signal to be played; and controlling the target playback device to play the target audio signal.
[0005] The target audio signal is determined by combining the audio signal to be played and preset frequency response information representing the deviation between the target playback device's actual frequency response and the target frequency response, thereby controlling the target playback device to play the target audio signal. Based on the audio signal to be played, the target audio signal is determined in combination with the deviation between the target playback device's actual frequency response and the target frequency response. By processing the audio signal, the frequency response deviation of the target playback device can be compensated in advance, ensuring that the playback effect of the target audio signal meets the desired playback effect of the audio signal to be played. Compared to directly playing the audio signal to be played by the target playback device, this can solve the problem of poor playback effect of the audio signal to be played due to performance defects of the target playback device, achieve balanced control of the target playback device, and thus improve the audio playback effect. Furthermore, this technical solution can achieve balanced control of the target playback device by simply utilizing the preset frequency response information, without the need for additional audio analysis equipment, thereby reducing the cost of balanced control of the target playback device.
[0006] In some possible implementations, determining the target audio signal based on the preset frequency response information and the audio signal to be played includes: determining a target filter coefficient based on the preset frequency response information; and filtering the audio signal to be played based on the target filter coefficient to obtain the target audio signal.
[0007] The filter coefficient is determined by presetting the frequency response information, and then the audio signal to be played is filtered with high precision and efficiency according to the filter coefficient.
[0008] In some possible implementations, obtaining preset frequency response information includes: obtaining a real frequency response curve, where the real frequency response curve is used to represent the real frequency response of the target playback device; obtaining a target frequency response curve, where the target frequency response curve is used to represent the required frequency response of the target playback device; and determining the preset frequency response information based on the real frequency response curve and the target frequency response curve.
[0009] In this way, the preset frequency response information can be effectively determined with low determination cost.
[0010] In some possible implementations, the real frequency response curve is used to represent a plurality of first frequencies and first response values respectively corresponding to the plurality of first frequencies, and the target frequency response curve is used to represent a plurality of second frequencies and second response values respectively corresponding to the plurality of second frequencies. Determining the preset frequency response information based on the real frequency response curve and the target frequency response curve includes: for each first frequency among the plurality of first frequencies, determining a target frequency identical to the first frequency from the plurality of second frequencies; determining a target response value corresponding to the first frequency based on the second response value corresponding to the target frequency identical to the first frequency and the first response value corresponding to the first frequency; and determining the preset frequency response information based on the plurality of first frequencies and the target response values respectively corresponding to the plurality of first frequencies.
[0011] In this way, the deviation between the actual frequency response curve and the target frequency response curve can be analyzed, and the preset frequency response information can be determined based on the deviation, so that the preset frequency response information can compensate for the deviation and achieve balanced control of the target playback device.
[0012] In some possible implementations, determining the target audio signal based on the preset frequency response information and the audio signal to be played includes: preprocessing the audio signal to be played to obtain a preprocessed audio signal; and determining the target audio signal based on the preset frequency response information and the preprocessed audio signal.
[0013] By preprocessing the audio signal, invalid signals can be removed and the audio signal can be normalized, thereby improving the quality of the target audio signal.
[0014] In some possible implementations, preprocessing the audio signal to be played to obtain a preprocessed audio signal includes: determining a low-frequency audio signal having a frequency lower than a preset frequency from the audio signal to be played; and removing the low-frequency audio signal from the audio signal to be played to obtain the preprocessed audio signal.
[0015] Through this preprocessing method, invalid low-frequency audio signals can be removed and the quality of the audio signal can be improved.
[0016] In some possible implementations, preprocessing the audio signal to be played to obtain a preprocessed audio signal includes: determining a silent audio signal from the audio signal to be played; determining a root mean square value of the audio signal to be played; and adjusting the root mean square value of the silent audio signal based on the root mean square value of the audio signal to be played to obtain the preprocessed audio signal.
[0017] Through this preprocessing method, the signals of each frequency band of the audio signal to be played can be normalized to a calibrated level, thereby realizing normalized processing of the audio signal and improving the accuracy of subsequent audio signal processing.
[0018] In some possible implementations, preprocessing the audio signal to be played to obtain a preprocessed audio signal includes: determining a non-silent audio signal from the audio signal to be played; determining a proportion of the non-silent audio signal in the audio signal to be played; and processing the amplitude of the audio signal to be played based on the proportion to obtain the preprocessed audio signal.
[0019] Through this amplitude processing method, the signals of each frequency band of the audio signal to be played can be normalized to a relatively uniform level, thereby achieving normalized processing of the audio signal and improving the accuracy of subsequent audio signal processing.
[0020] It should be noted that the above-mentioned low-frequency audio signal removal processing and audio signal normalization processing can be reasonably combined according to actual needs. For example, the audio signal to be played is pre-processed, including at least one of the following processing methods: removing the low-frequency audio signal of the audio signal to be played, normalizing the audio signal by determining a silent audio signal, and normalizing the audio signal by determining a non-silent audio signal.
[0021] According to a second aspect of an embodiment of the present disclosure, a playback control device is provided, which is configured to execute the playback control method described in the first aspect of the present disclosure.
[0022] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the playback control method described in the first aspect of the present disclosure is implemented.
[0023] According to a fourth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to: execute the executable instructions to implement the playback control method as described in the first aspect of the present disclosure.
[0024] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the playback control method as described in the first aspect of the present disclosure.
[0025] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0027] Figure 1 The figure is a flowchart of a playback control method according to an exemplary embodiment.
[0028] Figure 2 The figure is a schematic diagram of a recording scenario of an audio signal according to an exemplary embodiment.
[0029] Figure 3 FIG. 4 is a diagram showing an example of a frequency response curve according to an exemplary embodiment.
[0030] Figure 4 The figure is a flow chart of playback control of an artificial mouth according to an exemplary embodiment.
[0031] Figure 5 The figure is a schematic diagram of a test scenario of a smart cockpit according to an exemplary embodiment.
[0032] Figure 6 The figure is a block diagram of a playback control device according to an exemplary embodiment.
[0033] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0034] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0035] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0036] As mentioned in the background art, in relevant application scenarios, it is usually necessary to perform equalization control on an audio playback device to improve the audio playback effect.
[0037] As an example application scenario, in acoustic testing, an artificial mouth can be used to play a large amount of test material. Artificial mouths are commonly used in acoustic testing. Their purpose is to simulate the human vocal apparatus and accurately reproduce the human voice. The frequency response and distortion of the artificial mouth are important considerations. The flatter the frequency response and the lower the distortion, the closer the sound produced by the artificial mouth is to a real person.
[0038] Frequency response can be used to describe the differences in the playback device's ability to process audio signals. The closer the frequency response is to the ideal frequency response, the smaller the distortion of the audio signal and the better the playback effect.
[0039] However, in reality, due to the performance limitations of the hardware itself, it is difficult to achieve the requirement of smaller distortion. Therefore, it is usually necessary to perform equalization control on the artificial mouth so that the playback effect of the artificial mouth on the test corpus meets the expected requirements.
[0040] In some related technologies, audio analyzers are used to perform equalization control on audio playback devices to reduce playback distortion. However, this equalization control solution has the following problems: The cost is high. The price of an audio analyzer is usually around hundreds of thousands, which leads to high costs for acoustic testing.
[0041] Redundant functions and complicated configuration: The function of an audio analyzer is to analyze the indicators of acoustic components, such as frequency response, distortion, signal-to-noise ratio, etc. The equalization of the artificial mouth is only a preparatory work in the audio analyzer test.
[0042] Dependence on equipment manufacturers or third-party software. Audio analyzers require specialized software for control and operation, resulting in inflexible usage and limited scalability, making it difficult to meet the deployment requirements of automated testing solutions for increasingly diverse voice interaction scenarios.
[0043] Their bulk limits their application scenarios. For example, in-vehicle voice testing typically requires placing an artificial mouth on each seat, each requiring calibration and equalization. However, most audio analyzers do not support multi-channel equalization. Furthermore, deploying multiple audio analyzers in a vehicle is inconvenient for real-world road testing.
[0044] Based on this, an embodiment of the present disclosure provides a technical solution for determining a target audio signal based on an audio signal to be played and preset frequency response information representing the deviation between the actual frequency response of a target playback device and a target frequency response, and controlling the target playback device to play the target audio signal.
[0045] Based on the audio signal to be played, the target audio signal is determined in combination with the deviation between the actual frequency response of the target playback device and the target frequency response. Through audio signal processing, the frequency response deviation of the target playback device can be compensated in advance, so that the playback effect of the target audio signal can meet the required playback effect of the audio signal to be played.
[0046] Compared with directly playing the audio signal to be played by the target playback device, it can solve the problem of poor playback effect of the audio signal to be played due to performance defects of the target playback device, realize balanced control of the target playback device, and thus improve the audio playback effect.
[0047] Furthermore, this technical solution can achieve balanced control of the target playback device by simply utilizing preset frequency response information, without the need for additional audio analysis equipment, thereby reducing the balanced control cost of the target playback device.
[0048] Therefore, this technical solution can be applied to various acoustic test scenarios, such as: acoustic test scenarios of smart cockpits, acoustic test scenarios of smart devices such as speakers and mobile phones, etc.
[0049] It can be understood that in addition to acoustic test scenarios, this technical solution can also be applied to other scenarios, in which there is a need for equalization control of audio playback devices to improve the audio playback effect of the audio playback devices.
[0050] In the embodiment of the present disclosure, the audio playback device may be an artificial mouth, earphones, speakers, etc.
[0051] The hardware execution entity of this technical solution depends on the specific application scenario. For example, in an acoustic testing scenario, the hardware execution entity can be a host computer, a main control computer, etc.
[0052] Figure 1 is a flow chart of a playback control method according to an exemplary embodiment. Figure 1 As shown, the playback control method includes: Step S11 : obtaining preset frequency response information and the audio signal to be played. The preset frequency response information is used to represent the deviation between the actual frequency response of the target playback device and the target frequency response.
[0053] Step S12: determining a target audio signal according to the preset frequency response information and the audio signal to be played.
[0054] Step S13: Control the target playback device to play the target audio signal.
[0055] In some embodiments, the target playback device may be an artificial mouth, or other device that simulates the sound production of a human mouth.
[0056] In some embodiments, the actual frequency response of the target playback device can represent the target playback device's audio signal processing capabilities, which depends on the hardware performance of the target playback device. The target frequency response of the target device can represent the target playback device's audio signal processing capabilities.
[0057] Taking the acoustic test scenario as an example, the requirement for the target playback device may be that the frequency response curve of the target playback device is a flat frequency response curve, indicating that the sound performance of the target playback device is balanced throughout the entire frequency range, and there is no obvious tendency in the bass, mid-range and treble in the sound output, and the audio content can be presented in a relatively realistic manner.
[0058] However, affected by the hardware performance of the target playback device, the actual frequency response curve of the target playback device is not a flat frequency response curve, and there is a deviation between the actual frequency response curve and the flat frequency response curve (ie, the target frequency response).
[0059] Taking audio playback as an example, the requirement for the target playback device may be that the frequency response curve of the target playback device is a frequency response curve that emphasizes low frequencies, so that the target playback device can enhance the processing of the low-frequency part and achieve a "low-frequency boost" effect, such as a rich, heavy bass playback effect.
[0060] However, due to the hardware performance of the target playback device, the actual frequency response curve of the target playback device is not a frequency response curve that emphasizes low frequencies, and there is a deviation between the actual frequency response curve and the frequency response curve that emphasizes low frequencies (i.e., the target frequency response).
[0061] When there is a deviation between the actual frequency response of the target playback device and the target frequency response, the preset frequency response information can be used to compensate (equalize) the deviation so that the final audio playback effect of the target playback device meets the required audio playback effect.
[0062] As an optional implementation, obtaining preset frequency response information includes: obtaining a real frequency response curve, the real frequency response curve is used to represent the real frequency response of the target playback device; obtaining a target frequency response curve, the target frequency response curve is used to represent the required frequency response of the target playback device; and determining the preset frequency response information based on the real frequency response curve and the target frequency response curve.
[0063] In this embodiment, the preset frequency response information may be determined by a real frequency response curve representing the real frequency response of the target playback device and a target frequency response curve representing the required frequency response of the target playback device.
[0064] The preset frequency response information may be a preset frequency response curve, which may also be understood as an EQ (Equalizer Curve) curve. The EQ curve is a response curve that describes the gain or attenuation adjustment of an equalizer to different frequencies of an audio signal.
[0065] In this way, the preset frequency response information can be effectively determined with low determination cost.
[0066] In some embodiments, the actual frequency response curve can be determined based on actual measurements of the target playback device.
[0067] As an optional implementation, obtaining a true frequency response curve includes: obtaining a test audio signal; obtaining a recorded audio signal, where the recorded audio signal is an audio signal obtained by recording the test audio signal played by a target playback device; and determining the true frequency response curve based on the test audio signal and the recorded audio signal.
[0068] With this implementation, it is only necessary to record the test audio signal played by the target playback device and combine the test audio signal with the recorded audio signal to determine the true frequency response curve. No additional audio analysis equipment is required, thus achieving a simple and effective determination of the true frequency response curve.
[0069] In some embodiments, obtaining a recorded audio signal may include: controlling a target playback device to play a test audio signal; recording the test audio signal played by the target playback device through an audio acquisition device to obtain an original recorded audio signal; and converting the original recorded audio signal through an audio conversion device to obtain a target recorded audio signal.
[0070] By means of this embodiment, a simple and low-cost acquisition of recorded audio signals can be achieved.
[0071] In some embodiments, the test audio signal may be in the form of white noise.
[0072] In some embodiments, the audio acquisition device may be a microphone, and the audio conversion device may be a sound card. Accordingly, the original recorded audio signal may be an analog signal, and the target recorded audio signal may be a digital signal.
[0073] In some embodiments, the target playback device, audio acquisition device, and audio conversion device may be deployed in a semi-anechoic indoor environment to reduce interference from other noises.
[0074] Figure 2 FIG. 1 is a schematic diagram of a recording scenario of an audio signal according to an exemplary embodiment. Figure 2 As shown in FIG, in this recording scene, there are an artificial mouth, a microphone, a sound card, and an audio analysis device. The microphone is connected to the sound card, and the sound card is connected to the audio analysis device.
[0075] Among them, audio processing software can be installed on the audio analysis device to process the target recorded audio signal to obtain a real frequency response curve.
[0076] Furthermore, the distance between the microphone and the artificial mouth may be a preset distance, for example, 10 cm. The preset distance may be reasonably set according to actual conditions.
[0077] In this recording scenario, the artificial mouth plays white noise, the microphone records the white noise played by the artificial mouth, and transmits it to the sound card. The sound card converts the signal and then transmits it to the audio analysis device. The audio analysis device processes the audio to obtain the real frequency response curve.
[0078] In some embodiments, the target frequency response curve can be determined based on the desired audio playback effect. For example, if the desired audio playback effect is to present audio content realistically with minimal distortion, the target frequency response curve is a flat frequency response curve. This desired audio playback effect is commonly found in acoustic testing scenarios.
[0079] For example, the required audio playback effect is to emphasize the low-frequency part of the audio content to achieve low-frequency enhancement, and the target frequency response curve is a frequency response curve that emphasizes low frequencies.
[0080] For example, if the desired audio playback effect is strong, with prominent bass and treble, and relatively weak mid-range, to provide users with a more impactful and entertaining audio experience, the target frequency response curve could be a V-shaped frequency response curve. This type of audio playback effect is commonly seen in audio playback devices such as music headphones.
[0081] In some embodiments, a time-domain adaptive filtering algorithm or a least squares method may be used to solve a true frequency response curve based on the recorded audio signal and the test audio signal.
[0082] As an example, the relationship between the recorded audio signal, the test audio signal, and the actual frequency response curve can be: target=conv(input,h), where target represents the recorded audio signal, input represents the test audio signal, h represents the actual frequency response curve, and conv represents the convolution algorithm.
[0083] In some embodiments, the actual frequency response curve can also be understood as the EQ curve of a filter equivalent to the EQ curve of the target playback device. It is understood that the target playback device has an internal equalizer (filter), and the EQ curve of this equalizer represents the audio processing capability of the target playback device.
[0084] In some embodiments, the real frequency response curve is used to represent multiple first frequencies and the first response values corresponding to the multiple first frequencies, and the target frequency response curve is used to represent multiple second frequencies and the second response values corresponding to the multiple second frequencies.
[0085] The first response value and the second response value may be gain values or attenuation values. If they are gain values, the audio effect may be amplified; if they are attenuation values, the audio effect may be weakened.
[0086] Therefore, as an optional implementation, determining the preset frequency response information based on the actual frequency response curve and the target frequency response curve includes: for each first frequency among the multiple first frequencies, determining a target frequency that is the same as the first frequency from the multiple second frequencies; determining a target response value corresponding to the first frequency based on a second response value corresponding to the target frequency that is the same as the first frequency and a first response value corresponding to the first frequency; and determining the preset frequency response information based on the multiple first frequencies and the target response values respectively corresponding to the multiple first frequencies.
[0087] In this embodiment, since the frequencies involved in the actual frequency response curve may be different from the frequencies involved in the target frequency response curve, the same frequencies involved in both the target frequency response curve and the actual frequency response curve can be determined, and the response value of the same frequencies in the actual frequency response curve can be adjusted using the response value of the same frequencies in the target frequency response curve.
[0088] In this way, the deviation between the actual frequency response curve and the target frequency response curve can be analyzed, and the preset frequency response information can be determined based on the deviation, so that the preset frequency response information can compensate for the deviation and achieve balanced control of the target playback device.
[0089] In some embodiments, determining the target response value corresponding to the first frequency based on the second response value corresponding to the target frequency the same as the first frequency and the first response value corresponding to the first frequency may include: using a least squares algorithm to solve the target response value by which the first response value corresponding to the first frequency can be compensated to the second response value.
[0090] The preset frequency response information may be a preset frequency response curve, which may represent a plurality of target frequencies and target response values corresponding to the plurality of target frequencies.
[0091] As an example, the real frequency response curve is understood as the EQ curve of the filter equivalent to the EQ curve of the target playback device. Then, the preset frequency response curve can be understood as the EQ curve of the inverse filter, and the target frequency response curve can be understood as the EQ curve required by the filter of the target playback device.
[0092] Then, the relationship between the actual frequency response curve, the preset frequency response curve and the target frequency response curve can be expressed as: impulse=conv(hi,hi), where impulse represents the target frequency response curve, h represents the actual frequency response curve, hi represents the preset frequency response curve, and hi needs to be solved.
[0093] Figure 3 is an example diagram of a frequency response curve according to an exemplary embodiment. Figure 3 As shown, there are three frequency response curves involved, namely: real frequency response curve, target frequency response curve and preset frequency response curve.
[0094] By superimposing the actual frequency response curve with the preset frequency response curve, a target frequency response curve can be obtained. The target frequency response curve is a flat frequency response curve.
[0095] Therefore, the audio signal to be played is first processed according to the preset frequency response curve to obtain the target audio signal, and then the target audio signal is played by the target playback device. The target playback device will adjust the target audio signal in a gain or attenuation adjustment manner represented by the real frequency response curve before playing it. This is equivalent to processing the audio signal to be played according to the target frequency response curve before playing it, so that the playback effect of the target audio signal is the playback effect required by the audio signal to be played.
[0096] For example, frequency A has a response value of 40 in the preset frequency response curve, a response value of 0 in the target frequency response curve, and a response value of -40 in the actual frequency response curve. Therefore, using the preset frequency response curve, we first apply a gain of +40 to audio signal B corresponding to frequency A in the audio signal being played. This changes the audio signal corresponding to frequency A to B+40, resulting in the target audio signal.
[0097] Then, based on the actual frequency response curve of the target playback device, the target playback device will perform a -40 attenuation process on the audio signal B+40 corresponding to frequency A. The audio signal corresponding to frequency A becomes B, which is consistent with the audio signal corresponding to frequency A in the original audio signal to be played, so that the playback effect of the target audio signal is the playback effect required by the audio signal to be played.
[0098] In some embodiments, step S12 may include: determining a target filter coefficient according to preset frequency response information; and filtering the audio signal to be played according to the target filter coefficient to obtain a target audio signal.
[0099] In this embodiment, filtering processing can be achieved through a filter.
[0100] In some embodiments, the filter can be called an equalizer, and the preset frequency response information can be understood as the EQ curve of the filter. Therefore, the filter coefficient of the filter can be determined through the EQ curve, so that the filter can filter the audio signal to be played according to the filter coefficient to obtain the target audio signal.
[0101] Furthermore, in step S13, the target playback device may be controlled to play the target audio signal so that the playback effect of the target audio signal is consistent with the required playback effect of the audio signal to be played.
[0102] In some embodiments, in an acoustic test scenario, the target playback device is required to continuously play the audio signal.
[0103] Therefore, as an optional implementation, obtaining the audio signal to be played includes: obtaining the audio signal to be played from an audio signal source according to a preset sampling frequency.
[0104] The preset sampling frequency may be 48K or other sampling frequencies configured according to the scenario.
[0105] In some embodiments, the sampling frequency can be determined in real time. If the preset sampling frequency is met, the audio signal processing can continue. If the preset sampling frequency is not met, resampling is performed according to the preset sampling frequency.
[0106] By controlling the sampling frequency, it is possible to achieve effective and stable acquisition of the audio signal to be played, thereby ensuring the processing efficiency and accuracy of the audio signal to be played.
[0107] In some embodiments, determining the target audio signal based on preset frequency response information and the audio signal to be played includes: preprocessing the audio signal to be played to obtain a preprocessed audio signal; and determining the target audio signal based on the preset frequency response information and the preprocessed audio signal.
[0108] In this embodiment, the audio signal to be played may be pre-processed before being processed according to the preset frequency response information. The pre-processing of the audio signal may remove invalid signals, normalize the audio signal, and so on, thereby improving the quality of the target audio signal.
[0109] As an optional preprocessing method, a low-frequency audio signal with a frequency lower than a preset frequency is determined from the audio signal to be played; the low-frequency audio signal is removed from the audio signal to be played to obtain a preprocessed audio signal.
[0110] The preset frequency can be configured according to different application scenarios. For example, in an acoustic test scenario, the preset frequency can be 50Hz. Audio signals below 50Hz can correspond to low-frequency background noise, vibration, and other invalid audio signals.
[0111] Through this preprocessing method, invalid low-frequency audio signals can be removed and the quality of the audio signal can be improved.
[0112] In some embodiments, the low-frequency audio signal removal process may be implemented directly through a high-pass filter.
[0113] As an optional preprocessing method, a silent audio signal is determined from the audio signal to be played; the root mean square value of the audio signal to be played is determined; and the root mean square value of the silent audio signal is adjusted according to the root mean square value of the audio signal to be played to obtain a preprocessed audio signal.
[0114] The silent audio signal in the audio signal to be played may be detected based on energy or short-time zero-crossing rate. These two detection algorithms are mature technologies in this field and will not be introduced in detail here.
[0115] In some embodiments, the root mean square value of the audio signal to be played may be used as the target root mean square value, and the root mean square value of the silent audio signal may be adjusted to the target root mean square value.
[0116] Through this preprocessing method, the signals of each frequency band of the audio signal to be played can be normalized to a calibrated level, thereby realizing normalized processing of the audio signal and improving the accuracy of subsequent audio signal processing.
[0117] As an optional preprocessing method, a non-silent audio signal is determined from the audio signal to be played; the proportion of the non-silent audio signal in the audio signal to be played is determined; and based on the proportion, the amplitude of the audio signal to be played is processed to obtain a preprocessed audio signal.
[0118] In some embodiments, audio signals other than silent audio signals may be determined as non-silent audio signals by detecting silent audio signals.
[0119] In some embodiments, the overall amplitude or average amplitude of the audio signal to be played can be adaptively reduced or increased according to the occupancy ratio.
[0120] Through this amplitude processing method, the signals of each frequency band of the audio signal to be played can be normalized to a relatively uniform level, thereby achieving normalized processing of the audio signal and improving the accuracy of subsequent audio signal processing.
[0121] It is understood that the various preprocessing methods described above can be combined. In such a scenario, the order of the preprocessing methods can be: first perform low-frequency signal removal, then perform normalization. Normalization includes processing of silent audio signals and processing of audio signal amplitude.
[0122] Figure 4 FIG. 1 is a flow chart of playback control of an artificial mouth according to an exemplary embodiment. Figure 4 As shown, the control process includes: Read the audio file and determine whether the sampling rate of the audio file is 48k. If so, proceed to the next step. If not, resample it to 48k.
[0123] High-pass filtering is used to remove low-frequency noise, vibration and other non-valid data.
[0124] Calculate the effective point ratio, that is, calculate the ratio of the non-silent audio signal in the audio signal to be played.
[0125] The calculated effective point ratio is used to perform normalization processing, including the root mean square value adjustment and amplitude adjustment introduced in the above embodiment.
[0126] Applying an EQ filter, that is, applying the preset frequency response information based on the aforementioned EQ filter, to filter the audio signal.
[0127] Finally, the final processed audio signal is played through the artificial mouth.
[0128] In an acoustic test scenario, the playback control method may further include: obtaining identification information of a target audio signal played by the device under test to a target playback device; and determining test information based on the identification information and the target audio signal, the test information being used to characterize the audio recognition capability of the device under test.
[0129] In some embodiments, the device to be tested may be a vehicle, a mobile phone, a speaker, or other smart device with voice recognition capabilities.
[0130] In some embodiments, the identification information may be analysis data and response data of the target audio signal. The analysis data may be the content obtained by analyzing the target audio signal, and the response data may be the result of responding to the target audio signal, including response content, response time, etc.
[0131] For example, if the device under test is a vehicle, the accuracy of the vehicle's in-car assistant's voice algorithm can be tested. The target audio signal could represent a voice control command, and the response information could indicate the assistant's recognition of the command. Successful recognition and correct command content indicate the assistant's audio recognition capabilities are relatively accurate. Failed recognition and / or incorrect command content indicate poor audio recognition capabilities.
[0132] Therefore, by combining the recognition information and the target audio information, the audio recognition capability of the device under test can be determined.
[0133] It is understandable that the specific method of determining the test information may be implemented differently in different application scenarios, and will not be introduced one by one here.
[0134] In some embodiments, taking the test of the voice algorithm of the smart cockpit as an example, the target playback device can be an artificial mouth, which can be deployed at various locations in the vehicle's cockpit.
[0135] As one of the means of human-machine interaction within the cabin, voice is connected to the vast majority of the cabin's application functions. Speech recognition systems need to convert speech into text. Speech algorithms can provide purer speech input, thereby improving speech recognition accuracy.
[0136] During the speech algorithm testing process, an artificial mouth can be used to play a large amount of corpus to test its actual speech recognition performance. The artificial mouth must play the corpus in a way that highly restores the user's original voice without distortion or distortion.
[0137] Figure 5 FIG. 1 is a schematic diagram of a test scenario of a smart cockpit according to an exemplary embodiment. Figure 5 As shown, the smart cockpit includes a vehicle computer. Artificial mouths can be deployed in the driver's seat, front passenger seat, left rear seat, rear center seat, and right rear seat of the smart cockpit to emit human-like audio signals.
[0138] Furthermore, the test corpus that the artificial mouth needs to play can be sent to the artificial mouth through the control device and the sound card, so that the artificial mouth plays the specific test corpus.
[0139] Among them, the control device is used to generate the test corpus that the artificial mouth needs to play, and the sound card is used to convert the test corpus sent by the control device into the form of audio signals that can be played by the artificial mouth, so that the artificial mouth can directly play the test corpus.
[0140] In this test scenario, the control device can apply the technical solutions introduced in the above embodiments to obtain the test corpus that the artificial mouth needs to play.
[0141] Based on the technical solutions of the embodiments of the present disclosure, in the voice algorithm test scenario of the smart cockpit, high-precision restoration of real human voices can be achieved, thereby improving the test accuracy of the smart cockpit voice algorithm.
[0142] Also, it can simplify the test equipment and dependent hardware devices.
[0143] In addition, it can save testing costs. Compared with the solution using an audio analyzer, the cost of this test solution is 1 / 30 of that of using an audio analyzer, and the test effect is the same.
[0144] Figure 6 FIG. 6 is a block diagram of a playback control device 600 according to an exemplary embodiment. Figure 6 , the device comprises: The acquisition module 601 is configured to acquire preset frequency response information and an audio signal to be played, wherein the preset frequency response information is used to represent the deviation between the actual frequency response of the target playback device and the target frequency response.
[0145] The determination module 602 is configured to determine a target audio signal according to the preset frequency response information and the audio signal to be played.
[0146] The control module 603 is configured to control the target playback device to play the target audio signal.
[0147] Optionally, the determination module 602 is further configured to: determine a target filter coefficient according to the preset frequency response information; and perform filtering processing on the audio signal to be played according to the target filter coefficient to obtain a target audio signal.
[0148] Optionally, the acquisition module 601 is further configured to: acquire a real frequency response curve, where the real frequency response curve is used to characterize the real frequency response of the target playback device; acquire a target frequency response curve, where the target frequency response curve is used to characterize the required frequency response of the target playback device; and determine the preset frequency response information based on the real frequency response curve and the target frequency response curve.
[0149] Optionally, the determination module 602 is further configured to: determine, for each first frequency among the multiple first frequencies, a target frequency that is the same as the first frequency from the multiple second frequencies; determine a target response value corresponding to the first frequency based on a second response value corresponding to the target frequency that is the same as the first frequency and a first response value corresponding to the first frequency; and determine the preset frequency response information based on the multiple first frequencies and the target response values respectively corresponding to the multiple first frequencies.
[0150] Optionally, the determination module 602 is further configured to: pre-process the audio signal to be played to obtain a pre-processed audio signal; and determine the target audio signal according to the preset frequency response information and the pre-processed audio signal.
[0151] Optionally, the determination module 602 is further configured to: determine a low-frequency audio signal having a frequency lower than a preset frequency from the audio signal to be played; and remove the low-frequency audio signal from the audio signal to be played to obtain a preprocessed audio signal.
[0152] Optionally, the determination module 602 is further configured to: determine a silent audio signal from the audio signal to be played; determine the root mean square value of the audio signal to be played; and adjust the root mean square value of the silent audio signal according to the root mean square value of the audio signal to be played to obtain a preprocessed audio signal.
[0153] Optionally, the determination module 602 is further configured to: determine a non-silent audio signal from the audio signal to be played; determine the proportion of the non-silent audio signal in the audio signal to be played; and process the amplitude of the audio signal to be played according to the proportion to obtain a preprocessed audio signal.
[0154] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0155] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the playback control method provided by the present disclosure are implemented.
[0156] Figure 7 7 is a block diagram of an electronic device 700 according to an exemplary embodiment. For example, the electronic device 700 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0157] Reference Figure 7 The electronic device 700 may include one or more of the following components: a processing component 702 , a memory 704 , a power component 706 , a multimedia component 708 , an audio component 710 , an input / output interface 712 , a sensor component 714 , and a communication component 716 .
[0158] The processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 702 may include one or more processors 720 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 702 may include one or more modules to facilitate interaction between the processing component 702 and other components. For example, the processing component 702 may include a multimedia module to facilitate interaction between the multimedia component 708 and the processing component 702.
[0159] The memory 704 is configured to store various types of data to support operations on the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, etc. The memory 704 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0160] The power supply component 706 provides power to the various components of the electronic device 700. The power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 700.
[0161] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the electronic device 700 is in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.
[0162] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, or a voice recognition mode. The received audio signals may be further stored in the memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 also includes a speaker for outputting audio signals.
[0163] The input / output interface 712 provides an interface between the processing component 702 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0164] The sensor assembly 714 includes one or more sensors for providing various aspects of status assessment for the electronic device 700. For example, the sensor assembly 714 can detect the open / closed state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor assembly 714 can also detect changes in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and temperature changes of the electronic device 700. The sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 714 may also include a light sensor and an image sensor for use in imaging applications. In some embodiments, the sensor assembly 714 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0165] The communication component 716 is configured to facilitate wired or wireless communication between the electronic device 700 and other devices. The electronic device 700 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0166] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the above-mentioned playback control method.
[0167] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions. The instructions can be executed by the processor 720 of the electronic device 700 to implement the above-described playback control method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0168] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned playback control method when executed by the programmable device.
[0169] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.
[0170] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. With particular regard to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. In addition, although particular features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "include," "have," "have," "have," or variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0171] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
[0172] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
[0173] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.
[0174] Although terms such as "first", "second" and "third" may be used herein to describe various components, parts, regions, layers or sections, these components, parts, regions, layers or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer or section from another component, part, region, layer or section. Therefore, without departing from the teachings of each example, the first component, part, region, layer or section mentioned in the examples described herein may also be referred to as the second component, part, region, layer or section. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" can explicitly or implicitly include at least one such feature. In the description herein, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise clearly and specifically defined.
Claims
1. A playback control method, characterized in that: include: Obtaining preset frequency response information and an audio signal to be played, wherein the preset frequency response information is used to represent a deviation between an actual frequency response of a target playback device and a target frequency response; determining a target audio signal according to the preset frequency response information and the audio signal to be played; Control the target playback device to play the target audio signal.
2. The playback control method according to claim 1, characterized in that: The determining of the target audio signal according to the preset frequency response information and the audio signal to be played includes: Determining a target filter coefficient according to the preset frequency response information; The audio signal to be played is filtered according to the target filter coefficient to obtain the target audio signal.
3. The playback control method according to claim 1, wherein: Get preset frequency response information, including: Obtaining a real frequency response curve, where the real frequency response curve is used to represent the real frequency response of the target playback device; Obtaining a target frequency response curve, where the target frequency response curve is used to characterize the required frequency response of the target playback device; The preset frequency response information is determined according to the actual frequency response curve and the target frequency response curve.
4. The playback control method according to claim 3, characterized in that: The real frequency response curve is used to represent a plurality of first frequencies and first response values respectively corresponding to the plurality of first frequencies, and the target frequency response curve is used to represent a plurality of second frequencies and second response values respectively corresponding to the plurality of second frequencies. Determining the preset frequency response information based on the real frequency response curve and the target frequency response curve includes: For each first frequency of the plurality of first frequencies, determining a target frequency that is the same as the first frequency from the plurality of second frequencies; determining a target response value corresponding to the first frequency according to a second response value corresponding to a target frequency that is the same as the first frequency and a first response value corresponding to the first frequency; The preset frequency response information is determined according to the multiple first frequencies and the target response values respectively corresponding to the multiple first frequencies.
5. The playback control method according to claim 1, characterized in that: The determining of the target audio signal according to the preset frequency response information and the audio signal to be played includes: Preprocessing the audio signal to be played to obtain a preprocessed audio signal; The target audio signal is determined according to the preset frequency response information and the preprocessed audio signal.
6. The playback control method according to claim 5, characterized in that: The preprocessing of the audio signal to be played to obtain a preprocessed audio signal includes: Determining a low-frequency audio signal having a frequency lower than a preset frequency from the audio signal to be played; The low-frequency audio signal is removed from the audio signal to be played to obtain a preprocessed audio signal.
7. The playback control method according to claim 5 or 6, characterized in that: The preprocessing of the audio signal to be played to obtain a preprocessed audio signal includes: Determining a silent audio signal from the audio signals to be played; Determining a root mean square value of the audio signal to be played; The root mean square value of the silent audio signal is adjusted according to the root mean square value of the audio signal to be played to obtain a preprocessed audio signal.
8. The playback control method according to claim 5 or 6, characterized in that: The preprocessing of the audio signal to be played to obtain a preprocessed audio signal includes: Determining a non-silent audio signal from the audio signal to be played; Determining a proportion of the non-silent audio signal in the audio signal to be played; The amplitude of the audio signal to be played is processed according to the occupancy ratio to obtain a preprocessed audio signal.
9. A playback control device, characterized in that: The playback control device is configured to execute the playback control method according to any one of claims 1 to 8.
10. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions to implement the playback control method according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the playback control method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that The present invention comprises a computer program, which implements the playback control method according to any one of claims 1 to 8 when executed by a processor.