Method, apparatus, storage medium and electronic device for enhancing audio recording quality
Patent Information
- Application Number
- CN202510895916.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2045-06-30
AI Technical Summary
[0004]本申请实施例提供了一种音频录制质量的增强方法、装置、存储介质、电子装置,以至少解决无法动态调整算法中使用的音频增益参数的问题
[0010]通过本申请,使用包含预设调整参数集合的音频算法组件对智能终端获取的第一音频进行处理,得到第二音频。随后,将第二音频同步至第一服务器进行识别,获取第一服务器识别所述第二音频生成的第一文本数据。在知晓智能终端对应的性能类型的基础上,根据第一文本数据与智能终端上预先设置的关键词集合,计算出关键词对应的第一相似值;再确定第一相似值与预设关键词阈值之间的第一大小关系,依据大小关系更新预设调整参数集合中的参数值,生成第一调整参数集合,后续音频算法组件通过使用第一调整参数集合对音频录制质量进行处理时的质量增强,采用上述方法,解决了相关技术中存在的无法动态调整算法中使用的音频增益参数的问题,达到查找适用于音频算法组件的调整参数集合,从而提升智能终端采集与用户在交互中产生的录制音频的音频质量,确保智能终端执行高质量语音交互的目的。
Smart Images

Figure CN120690216B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio data processing technology, and more specifically, to a method, apparatus, storage medium, and electronic device for enhancing audio recording quality. Background Technology
[0002] With the maturity of artificial intelligence technology, smart terminals, such as smart speakers, smartphones, and tablets, have become an indispensable part of daily life. Especially in the field of digital humans, these devices offer unprecedented human-computer interaction experiences. However, digital human applications rely heavily on the audio quality of the smart terminal; higher audio quality leads to higher accuracy in digital human recognition. Therefore, high-quality speech recognition is the foundation of a series of interactions, and the quality of audio recording directly affects the accuracy of the digital human's response and the user experience. In related technologies, most smart terminals have a set of fixed gain parameters preset locally. When the location or volume of the caller changes, the local preset gain parameters are obtained based on the recorded audio characteristics at that time, and the gain is adjusted and compared. That is, when the smart terminal interacts with the user, the audio collected during the interaction is processed locally on the smart device based on preset gain parameters. However, local adjustment requires the smart terminal to have high computing power and storage space. Therefore, the local parameter adjustment process is simplistic and has application limitations. Furthermore, it cannot guarantee that the recorded audio quality will remain unchanged when the smart terminal's environment changes subsequently, leading to a deterioration in the smart terminal's speech recognition performance.
[0003] Regarding the issue of the inability to dynamically adjust the audio gain parameters used in the algorithm, no effective solution has yet been proposed. Therefore, it is necessary to improve the relevant techniques to overcome the aforementioned shortcomings. Summary of the Invention
[0004] This application provides a method, apparatus, storage medium, and electronic device for enhancing audio recording quality, to at least address the problem of the inability to dynamically adjust the audio gain parameters used in the algorithm.
[0005] According to one aspect of the embodiments of this application, a method for enhancing audio recording quality is provided, comprising: processing a first audio obtained by a smart terminal using an audio algorithm component containing a preset set of adjustment parameters to obtain a second audio, wherein the first audio is audio generated by a target object by reading audio verification text displayed on the display interface of the smart terminal or a preset verification audio in the smart terminal; obtaining first text data generated by a first server recognizing the second audio, and performing similarity calculation based on the first text data and a preset set of keywords set on the smart terminal to obtain a first similarity value corresponding to the keywords; and, upon confirming the performance type corresponding to the smart terminal, updating the parameter values in the preset set of adjustment parameters according to a first size relationship between the first similarity value and a preset keyword threshold to obtain a first set of adjustment parameters for enhancing the audio recording quality processed by the audio algorithm component.
[0006] According to another aspect of the embodiments of this application, an audio recording quality enhancement device is also provided, comprising: a processing module, configured to process a first audio obtained by a smart terminal using an audio algorithm component containing a preset set of adjustment parameters to obtain a second audio, wherein the first audio is audio generated by a target object reading audio verification text displayed on the display interface of the smart terminal or a preset verification audio in the smart terminal; a calculation module, configured to obtain first text data generated by the first server recognizing the second audio, and perform similarity calculation based on the first text data and a preset set of keywords set on the smart terminal to obtain a first similarity value corresponding to the keyword; and an update module, configured to, upon confirming the performance type corresponding to the smart terminal, update the parameter values in the preset set of adjustment parameters according to a first size relationship between the first similarity value and a preset keyword threshold to obtain a first set of adjustment parameters for enhancing the audio recording quality processed by the audio algorithm component.
[0007] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, which is configured to execute the above-described method for enhancing audio recording quality when running.
[0008] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described method for enhancing audio recording quality through the computer program.
[0009] According to another aspect of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0010] This application utilizes an audio algorithm component containing a preset set of adjustment parameters to process a first audio file acquired by a smart terminal, resulting in a second audio file. The second audio file is then synchronized to a first server for recognition, generating first text data from the recognized second audio file. Based on knowledge of the smart terminal's performance type, a first similarity value is calculated between the first text data and a pre-set set of keywords on the smart terminal. A first size relationship is then determined between the first similarity value and a preset keyword threshold. The parameter values in the preset set of adjustment parameters are updated according to this relationship to generate a first set of adjustment parameters. Subsequent audio algorithm components enhance the audio recording quality by using this first set of adjustment parameters. This method solves the problem in related technologies where the audio gain parameters used in the algorithm cannot be dynamically adjusted. It achieves the goal of finding a suitable set of adjustment parameters for the audio algorithm component, thereby improving the audio quality of the recorded audio acquired by the smart terminal during user interaction and ensuring high-quality voice interaction on the smart terminal. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and, together with the description thereof, serve to explain this application and do not constitute an undue limitation thereof. In the drawings:
[0012] Figure 1 This is a hardware structure block diagram of a computer terminal for an audio recording quality enhancement method according to an embodiment of this application;
[0013] Figure 2 This is a flowchart of a method for enhancing audio recording quality according to an embodiment of this application;
[0014] Figure 3 This is a schematic diagram of audio recording interaction according to an embodiment of this application;
[0015] Figure 4 This is a flowchart illustrating an audio recording quality enhancement method according to an embodiment of this application;
[0016] Figure 5 This is a structural block diagram of an audio recording quality enhancement device according to an embodiment of this application. Detailed Implementation
[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] The methods and embodiments provided in this application can be executed on a computer terminal or similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for an audio recording quality enhancement method according to an embodiment of this application. For example... Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the image. A processor 102 (which may include, but is not limited to, a microprocessor unit (MPU) or a programmable logic device (PLD)) and a memory 104 for storing data.
[0020] In one exemplary embodiment, the computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 Equivalent functions or ratios shown Figure 1 The functions shown have more different configurations.
[0021] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the audio recording quality enhancement method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0022] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0023] This embodiment provides a method for enhancing audio recording quality. Figure 2 This is a flowchart of an audio recording quality enhancement method according to an embodiment of this application, such as... Figure 2 As shown, the steps of this method include:
[0024] Step S202: The first audio obtained by the smart terminal is processed using an audio algorithm component containing a preset set of adjustment parameters to obtain a second audio. The first audio is either the audio generated by the target object by reading the audio verification text displayed on the display interface of the smart terminal or the verification audio preset in the smart terminal.
[0025] The smart terminal displays a specific text and asks the user to read it aloud. The audio of the user's reading is captured by the device and used as a baseline signal for audio quality adjustment. In addition to the verification audio generated by the user on-site, the smart terminal may also have some carefully designed audio clips pre-stored to test the performance of the audio processing algorithm under different conditions. In either case, the "first audio" is processed by the smart terminal's audio algorithm component, which has a built-in set of preset adjustment parameters to control functions such as Automatic Gain Control (AGC), Acoustic Echo Cancellation (AEC), and Active Noise Reduction (ANR). After processing by the algorithm component, the enhanced "second audio" is obtained.
[0026] Step S204: Obtain the first text data generated by the first server recognizing the second audio, and perform similarity calculation based on the first text data and the preset keyword set set on the smart terminal to obtain the first similarity value corresponding to the keyword;
[0027] Optionally, the smart terminal synchronizes the "second audio" to an external first server. This server is typically an Automatic Speech Recognition (ASR) server, responsible for converting audio into text data. The device obtains the recognition result of the "second audio" from the first server, which is the "first text data." Then, the smart terminal calculates the similarity between the "first text data" and a locally preset set of keywords (preset speech recognition results) to obtain a "first similarity value." This "first similarity value" reflects the degree of match between the recognition result and the expected keywords. Furthermore, if the smart terminal is determined to be a high-performance device, the similarity between the "first text data" and a set of onomatopoeic words (preset noise words that should not be recognized) can also be calculated to obtain a "second similarity value." The "first similarity value" reflects the degree of match between the recognition result and the expected keywords, while the "second similarity value" is used to assess whether unnecessary onomatopoeic words are misrecognized, which helps to avoid the adverse effects of background noise on the recognition results.
[0028] Step S206: After confirming the performance type corresponding to the smart terminal, update the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold, and obtain the first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component.
[0029] Understandably, the smart terminal determines whether to update the preset set of adjustment parameters based on the first relationship between the "first similarity value" and the "preset keyword threshold," and the second relationship between the "second similarity value" and the "preset onomatopoeia threshold." If the accuracy and noise filtering capability of the recognition result do not meet the expected standards—that is, if the "first similarity value" is less than or equal to the keyword threshold, or the "second similarity value" is greater than or equal to the onomatopoeia threshold—the smart terminal will fine-tune the parameters of the 3A algorithm until the accuracy and noise filtering capability of the recognition result meet the requirements. This dynamic parameter tuning mechanism based on recognition feedback allows the smart terminal to self-adjust to its optimal state without affecting the user experience, adapting to different ASR servers and environmental conditions, and ultimately obtaining the "first set of adjustment parameters," which is a set of optimized parameters used to enhance audio recording quality.
[0030] Understandably, for smart terminal devices identified as "low-performance," their computing power is limited, making resource management particularly important. Therefore, in this embodiment, the smart terminal device will focus only on calculating the keyword set during the similarity calculation process; that is, it will only calculate the similarity between the first text data and the keyword set. Simply put, accurate keyword recognition is the most critical part of voice interaction, and the computing resources of low-performance devices should be prioritized to ensure the quality of keyword recognition. Simultaneously, for the similarity calculation of the onomatopoeic word set, the smart terminal will automatically adjust the result to a preset value, which is less than a preset onomatopoeic word threshold. This is to avoid unnecessary calculations on low-performance devices, as onomatopoeic word calculations can consume significant resources, and ensuring the accuracy of keyword recognition is of higher priority when resources are limited.
[0031] Conversely, if a smart terminal is categorized as "high-performance," it means the device has sufficient computing power to handle more complex tasks. In this case, the similarity calculation process can simultaneously include calculations on both the keyword set and the onomatopoeia set to evaluate the accuracy of the audio algorithm components in keyword recognition and their filtering effect on background noise and non-speech signals. Because high-performance devices can withstand the additional computational burden, they can more comprehensively optimize audio recording quality, ensuring accurate keyword recognition while avoiding misrecognition of unnecessary onomatopoeia, thereby improving overall speech recognition performance and user experience.
[0032] Optionally, the entity performing the above steps can be a smart terminal, computer, cloud platform, etc., but is not limited to these.
[0033] This application embodiment uses an audio algorithm component (mainly including AGC, AEC, and ANR) containing a preset set of adjustment parameters to process the first audio acquired by the smart terminal to obtain the second audio. The first audio can be: audio generated by the target object reading audio verification text displayed on the smart terminal's screen, or preset verification audio in the smart terminal. Subsequently, the second audio is synchronized to a first server (i.e., the ASR server) for recognition, obtaining the first text data generated by the first server recognizing the second audio. Next, based on knowing the performance type of the smart terminal, a first similarity value corresponding to the keywords is calculated according to the first text data and a preset keyword set on the smart terminal; then, a first size relationship is determined between the first similarity value and a preset keyword threshold, and the parameter values in the preset set of adjustment parameters are updated according to this relationship to generate a first set of adjustment parameters. The audio algorithm component then enhances the audio recording quality by using the first set of adjustment parameters. This implementation solves the problem in related technologies where the audio gain parameters used in the algorithm cannot be dynamically adjusted, achieving the goal of finding a suitable set of adjustment parameters for the audio algorithm component, thereby improving the audio quality of the recorded audio acquired by the smart terminal during user interaction, and ensuring that the smart terminal performs high-quality voice interaction.
[0034] Furthermore, by employing the above method, even with limited resources, the smart terminal can automatically find the optimal 3A algorithm parameters (equivalent to the first set of adjustment parameters in the above embodiments) through efficient interaction with an external ASR server to improve audio recording quality. This, in turn, reduces reliance on the smart terminal's hardware resources while maintaining audio quality, and reduces the recognition error rate caused by mismatched algorithm parameters. This allows users to obtain high-quality speech recognition and response without intervention, ensuring smooth voice interaction between the smart terminal and the user in both quiet and noisy environments.
[0035] In one example embodiment, after obtaining the first text data generated by the first server recognizing the second audio, the method further includes: obtaining the hardware configuration information of the smart terminal; determining the performance type of the smart terminal based on the hardware configuration, wherein the performance type includes at least one of the following: a low-performance device, a high-performance device. Simply put, if the hardware configuration of the smart terminal is lower than a predetermined performance benchmark, it is classified as a low-performance device. In this case, the device may be more inclined to simplify some computationally intensive tasks to ensure the smooth operation of critical functions. Conversely, if the hardware configuration of the device exceeds the performance benchmark, it is considered a high-performance device. This type of device has more powerful computing and processing resources and can handle more complex algorithms and tasks.
[0036] In one example embodiment, updating the parameter values in the preset adjustment parameter set according to a first size relationship between the first similarity value and a preset keyword threshold includes: when the smart terminal is a low-performance device, allowing parameter value updates only based on the first similarity value; wherein, updating parameter values based on the first similarity value includes: when the first size relationship is that the first similarity value is less than or equal to the preset keyword threshold, updating the default values of the preset automatic gain parameter and / or the preset echo cancellation parameter in the preset adjustment parameter set to obtain a target gain parameter value and a target echo cancellation parameter value; and updating the parameter values of the preset adjustment parameter set according to the target gain parameter value and the target echo cancellation parameter value.
[0037] Optionally, if the first similarity value is greater than the preset keyword threshold, the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set will not be updated; if the second similarity value is less than the preset onomatopoeia threshold, the preset noise reduction parameter in the preset adjustment parameter set will not be updated.
[0038] Optionally, after the smart terminal performs preliminary processing on the first audio (verification audio) and obtains the second audio, the device will synchronize the second audio to the first server for speech recognition. The first text data of the recognition result is then compared with the smart terminal's preset keyword set to calculate a first similarity value. If the first similarity value is greater than the preset keyword threshold, it means that the ASR server has correctly identified the keywords in the audio, and the preset AGC parameters are sufficient to optimize the audio signal, requiring no further adjustment. If the first similarity value is less than or equal to the preset keyword threshold, the matching degree between the recognition result and the preset keyword set is insufficient, indicating that the current AGC parameters or AEC parameters have failed to effectively improve the audio quality, and parameter updates are required to obtain better speech recognition results. During the update process, different gain values or echo cancellation parameter values may be tried until the target gain parameter value and target echo cancellation parameter value are found. The target gain parameter value and target echo cancellation parameter value are then used to update the parameter values in the audio algorithm component, thereby improving the accuracy of the audio algorithm component in keyword recognition during audio recording.
[0039] In one example embodiment, updating the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold includes: when the smart terminal is a high-performance type device, determining the similarity between the first text data and the preset onomatopoeia set to obtain a second similarity value corresponding to the onomatopoeia; determining the second similarity value and the second size relationship corresponding to the preset onomatopoeia threshold; and updating the parameter values in the preset adjustment parameter set simultaneously through the second size relationship and the first size relationship.
[0040] Optionally, updating the parameter values in the preset adjustment parameter set simultaneously through the second size relationship and the first size relationship includes: when the first size relationship has a first similarity value less than or equal to a preset keyword threshold, updating the parameter values of the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set to obtain a target gain parameter value and a target echo cancellation parameter value; when the second size relationship has a second similarity value greater than or equal to a preset onomatopoeia threshold, updating the parameter values of the preset noise reduction parameter in the preset adjustment parameter set to obtain a target noise reduction parameter value; and updating the parameter values in the preset adjustment parameter set based on the target gain parameter value, the target echo cancellation parameter value, and the target noise reduction parameter value.
[0041] Based on the established first similarity, if the smart terminal is a high-performance device, it will further evaluate whether the recognition result contains too many onomatopoeic words. This is achieved by calculating the similarity between the first text data and a preset onomatopoeic word set to obtain a second similarity value. If the second similarity value is less than the preset onomatopoeic word threshold, it indicates that no onomatopoeic words that should not be recognized have appeared in the recognition result, and the current ANR parameters are valid and do not need modification. If the second similarity value is greater than or equal to the preset onomatopoeic word threshold, it indicates that the recognition result is adversely affected by background noise, and too many non-speech signals (onomatopoeic words) have been recognized. In this case, the preset noise reduction parameters need to be optimized to reduce the second similarity value until the target noise reduction value is found, effectively filtering out background noise.
[0042] Once the type of parameter to be optimized (AGC, AEC, or ANR) is determined, the smart terminal will try different parameter values one by one until it obtains the target gain value, target echo cancellation parameter value, and target noise reduction parameter value that meet the requirements of keyword matching and onomatopoeia filtering. These target parameter values will be summarized to form the first set of adjustment parameters, which will be used for subsequent audio signal processing to enhance the quality of the recorded audio.
[0043] Through a dynamic optimization mechanism based on recognition result feedback, the smart terminal can intelligently adjust the 3A algorithm parameters to adapt to different recording conditions and the recognition preferences of the ASR server. This method not only reduces the hardware requirements caused by resource-intensive local parameter tuning but also significantly improves audio recording quality, ensuring a better user experience when using digital human applications or other voice interaction services. Furthermore, through interaction with the ASR server, the smart terminal can evaluate the effectiveness of parameters in real time, avoiding the problem of ineffective preset parameters in complex or changing environments. This allows the device to maintain good performance and compatibility even in scenarios with extremely high speech recognition requirements.
[0044] In an exemplary embodiment, after updating the parameter values in a preset adjustment parameter set based on a first size relationship and a second size relationship to obtain a first adjustment parameter set for enhancing the audio recording quality processed by the audio algorithm component, the method further includes: performing verification processing on the first audio using an audio algorithm component containing the first adjustment parameter set to obtain a third audio; synchronizing the third audio to a first server and receiving second text data fed back by the first server; performing similarity calculations based on the performance type corresponding to the smart terminal, the second text data, a preset keyword set, and a preset onomatopoeia set to obtain a third similarity value corresponding to the keyword and a fourth similarity value corresponding to the onomatopoeia; and determining the first adjustment parameter set as the final result of parameter value update when the third size relationship between the third similarity value and the preset keyword threshold and the fourth size relationship between the fourth similarity value and the preset onomatopoeia threshold meet preset relationship conditions, wherein the preset relationship conditions include at least: the similarity value at the keyword level is greater than the preset keyword threshold, and the similarity value at the onomatopoeia level is less than the preset onomatopoeia threshold.
[0045] It should be noted that, in order to facilitate cyclic processing, when the smart terminal is a low-performance device, the second similarity can be automatically set to any value less than the preset onomatopoeia threshold to avoid entering the second similarity calculation process in case of anomalies.
[0046] Optionally, the smart terminal uses preset audio algorithm components (such as AGC, AEC, ANR) again, but this time it uses an updated set of first adjustment parameters to process the previous first audio, generating a new "third audio." This verifies the validity of the determined first adjustment parameter set to determine whether further parameter updates are necessary. As with the initial process, the device synchronizes the third audio to the first server (ASR server) and receives the second text data generated by the server based on this audio. Subsequently, the smart terminal performs a second round of similarity calculation, calculating the similarity between the second text data and the preset keyword set and onomatopoeia set. If the smart terminal is a low-performance device, only the "third similarity value" corresponding to the keywords is obtained. However, if the smart terminal is a high-performance device, both the "third similarity value" and the "fourth similarity value" corresponding to the onomatopoeia need to be determined simultaneously. The calculation of the third and fourth similarity values is then used to deeply verify whether the actual effect of the first adjustment parameter set has improved the accuracy of keyword recognition as expected, while reducing the misidentification of onomatopoeia.
[0047] Ultimately, the smart terminal will determine whether the currently generated first set of adjustment parameters can enhance the audio algorithm component's processing of audio recording quality by checking the third similarity value and the preset keyword threshold, and the fourth similarity value and the preset onomatopoeia threshold. For example, if the third similarity value is greater than the preset keyword threshold, it means that the accuracy of keyword recognition has been improved; if the fourth similarity value is less than the preset onomatopoeia threshold, it means that the filtering effect of background noise and non-speech signals has been improved. If all the above conditions are met, the smart terminal device determines that the first set of adjustment parameters is the final result of parameter value update, that is, the parameter values of this set can effectively enhance the audio recording quality and become the standard parameter set for subsequent audio processing.
[0048] Through the above verification process, smart terminal devices can confirm whether the adjusted 3A algorithm parameters truly help improve audio quality, especially the accuracy of keyword recognition and the efficiency of onomatopoeia filtering.
[0049] In an exemplary embodiment, after collecting numerical information of different parameters in the first set of adjustment parameters and updating the preset set of adjustment parameters using the numerical information, the method further includes: when the audio recording module on the smart terminal generates target audio data, processing the target audio using an audio algorithm component that has completed parameter value updates to obtain target enhanced audio, wherein the target enhanced audio contains a normal interaction identifier; acquiring target text data generated by the first server processing the target enhanced audio, and receiving skill instructions issued by the Natural Language Processing (NLP) server based on the target text data; and providing voice interaction feedback to the target object operating the current smart terminal based on the target text data and the skill instructions.
[0050] After the smart terminal device updates the parameter values of the preset adjustment parameter set to obtain the first adjustment parameter set used to enhance audio recording quality, the device enters a normal user interaction state. The audio recording module continues to collect real-time audio data of the target object (user) in the environment, generating "target audio data". Then, the device's audio algorithm component applies this updated set of parameters to process the target audio, generating "enhanced audio data". In addition to containing the optimized audio signal, i.e., "target enhanced audio", the enhanced audio data also includes a "normal interaction identifier" to inform the first server that this is data from normal user interaction, rather than previous test or calibration audio. The enhanced audio data is uploaded to the first server (ASR server) for processing. The server uses its internal automatic speech recognition technology to convert the audio signal into target text data. The generated target text data is not only returned to the smart terminal device but also synchronized to the natural language processing server. The task of the natural language processing server is to deeply analyze the meaning of the text data and transform it into specific semantic instructions. Based on the parsed semantics, the natural language processing server issues corresponding "skill instructions" to the smart terminal device, guiding the device to perform specific operations or services. After receiving a skill command, the smart terminal device will provide precise voice interaction feedback to the target object (i.e., the user) based on the target text data and the content of the skill command. This feedback can be an answer to a question, a broadcast of the execution result of a command, or the provision of further interactive options, aiming to quickly respond to user needs and provide a smooth dialogue experience.
[0051] In an exemplary embodiment, after providing voice interaction feedback to the target object operating the current smart terminal based on target text data and skill instructions, the method further includes: determining the upgrade type of the active upgrade when the target object actively upgrades the voice interaction function of the smart terminal; synchronizing the second audio to the second server when the upgrade type is to replace the first server; acquiring the third text data generated by the second server recognizing the second audio, performing similarity calculation based on the performance type corresponding to the smart terminal, the third text data, the preset keyword set, and the preset onomatopoeia set to obtain a fifth similarity value corresponding to the keyword and a sixth similarity value corresponding to the onomatopoeia; determining the fifth size relationship between the fifth similarity value and the preset keyword threshold and the sixth size relationship between the sixth similarity value and the preset onomatopoeia threshold, updating the parameter values of the preset adjustment parameter set based on the fifth size relationship and the sixth size relationship to obtain a second adjustment parameter set, wherein the second adjustment parameter set is a set of parameters determined after the smart terminal connects to the second server for enhancing the audio recording quality of the audio algorithm component.
[0052] Optionally, when the target (user) consciously chooses to upgrade the voice interaction function of the smart terminal, such as upgrading to a higher-level digital human application or replacing the ASR server, the smart terminal will detect this operation and determine the specific upgrade type. Replacing the ASR server means the device needs to be adapted to the new server to ensure that the audio recording quality remains ideal in the new environment. After determining that the upgrade type is replacing the first server (i.e., the currently used ASR server), the smart terminal will synchronize the previously used "second audio" (containing the device's built-in test voice data) to the new "second server." A unified test audio is used to evaluate the recognition capabilities of different ASR servers, thereby adjusting the parameters of the audio algorithm components accordingly. After processing the second audio, the second server generates "third text data," which is the recognition result of the new server. The smart terminal then calculates the similarity between the third text data and the local preset keyword set and preset onomatopoeia set. When the smart terminal is a low-performance device, only the "fifth similarity value" corresponding to the keywords is obtained; while when the smart terminal is a high-performance device, both the "third similarity value" and the "sixth similarity value" corresponding to the onomatopoeia need to be determined. Next, the smart terminal will dynamically update its preset adjustment parameter set based on the fifth similarity value and the preset keyword threshold, and the sixth similarity value and the preset onomatopoeia threshold, forming a "second adjustment parameter set".
[0053] Specifically: If the fifth similarity value is greater than the preset keyword threshold, it indicates that the keyword recognition accuracy meets the requirements, and there is no need to adjust the corresponding 3A algorithm parameters; if the fifth similarity value is less than or equal to the preset keyword threshold, it indicates that there is a problem with keyword recognition, and it is necessary to adjust the automatic gain control (AGC) or echo cancellation (AEC) parameters to improve the accuracy of keyword recognition; if the sixth similarity value is less than the preset onomatopoeia threshold, it indicates that the onomatopoeia recognition control is good, and there is no need to adjust the noise reduction (ANR) parameters; if the sixth similarity value is greater than or equal to the preset onomatopoeia threshold, it is necessary to enhance the noise reduction capability and reduce the misrecognition of onomatopoeia by adjusting the ANR parameters.
[0054] It should be noted that the second set of adjustment parameters includes all 3A algorithm parameters that the smart terminal needs to adjust to optimize audio recording quality after switching to the second server. The determination of this parameter set allows the smart terminal to seamlessly switch between different servers without affecting the end-user's interactive experience with the digital human application. More importantly, it ensures that regardless of server changes, the smart terminal can maintain or improve audio recording quality by adjusting its own audio processing strategy, meeting users' expectations for speech recognition accuracy.
[0055] In one example embodiment, after updating the parameter values of a preset adjustment parameter set based on a first size relationship and a second size relationship to obtain a first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component, the method further includes: collecting configuration information of a first server; associating and saving the configuration information with the preset adjustment parameter set updated using the first adjustment parameter set to generate recorded data for a smart terminal configuration server; and, when the smart terminal is connected to a third server and the configuration information of the third server is the same as that of the first server, setting up audio recording quality enhancement for the smart terminal based on the recorded data.
[0056] Optionally, after enhancing audio recording quality using the optimized first set of adjustment parameters, and after this set has proven effective, the smart terminal then collects the configuration information of the current "first server" (i.e., the verified ASR server). This configuration information may include the server type, version, and technology stack used, reflecting the server's characteristics and identification capabilities. After obtaining the first server's configuration information, the smart terminal associates this information with the updated preset set of adjustment parameters and saves it. This data recording essentially creates a mapping between the smart terminal and a specific server configuration, meaning the smart terminal can quickly identify and call the optimized parameter set that matches the server in the future. When the smart terminal needs to establish a connection with a new "third server," if the configuration information of this third server is the same as the previously saved first server configuration information, the smart terminal can immediately call the associated recorded data. Using this data, the device can automatically apply the audio recording quality enhancement settings previously optimized for the first server—the second set of adjustment parameters—without needing to perform a lengthy parameter calibration and testing process again. In summary, by storing a set of optimized parameters related to specific server configurations, smart terminals enhance their personalized service capabilities, enabling them to provide the most suitable audio recording and processing solutions based on each user's currently selected ASR server. This allows them to automatically make the best parameter selection in complex multi-server environments without manual user intervention.
[0057] Furthermore, the following examples will provide further details.
[0058] As an optional implementation, this application proposes an audio recording quality improvement method, which includes: a smart terminal device (equivalent to the smart terminal in the above embodiment) locally presets audio data processed by a 3A algorithm with preset parameters, and the audio data has known text information. A corresponding keyword set and onomatopoeia set are also preset locally. When adapting to a digital human platform, the audio is sent to an ASR server for processing, and the target text data is returned to the terminal device. The device calculates the similarity between the target text data and the locally preset keyword set. If the similarity meets a preset threshold, the preset parameters of the 3A algorithm are considered effective, and the 3A algorithm module records and loads the parameters. If the preset threshold is not met, the preset parameters of the 3A algorithm are considered ineffective, and the algorithm parameters on the 3A algorithm module are adaptively adjusted. The reprocessed target audio data is then uploaded to the ASR server for recognition until the similarity meets the requirements. The 3A algorithm module then records and loads the optimized parameters. The device's 3A algorithm module processes the continuously collected audio data and uploads it to the ASR server. By using the above technical solution and dynamically adjusting the 3A algorithm parameters of the terminal device, it is possible not only to effectively reduce the use of device resources by the 3A algorithm, but also to improve the efficiency of adapting to different ASR servers and enhance the user experience.
[0059] Furthermore, the digital human application on the smart terminal device interacts and calculates with the ASR server via the internet. The device's 3A algorithm obtains optimal parameters, the audio collected by the device is processed by the 3A algorithm, and then the processed audio data is sent to the ASR server. The ASR server correctly recognizes the text information and then sends the corresponding text information to the NLP server. After the NLP server recognizes the correct speech, it notifies the skill server, which then sends the corresponding skill message to the smart terminal device. After receiving the skill instruction from the server, the digital human application correctly executes the specified interaction with the user. This method improves the recorded audio quality with limited device resources, ensuring a good user experience for the digital human application. It also allows smart terminal devices to quickly adapt to different ASR servers, reduces the resource usage of the 3A algorithm, and allows users to seamlessly switch ASR servers, effectively improving the user experience.
[0060] Optionally, the above-mentioned methods for improving audio recording quality can be applied to... Figure 3 In the system architecture shown, Figure 3This is a schematic diagram of audio recording interaction according to an embodiment of this application; the system architecture includes: a smart terminal device 32, an ASR server 34, an NLP server 36, and a skill server 38. The smart terminal device includes a 3A algorithm module, an audio acquisition module, a similarity calculation module, and a digital human application; wherein, the 3A algorithm module includes AGC, AEC, and ANR, wherein AGC has multiple parameters to adjust the audio gain, with an adjustment range of 0-100, AEC has two parameters, AEC and AIAEC, and ANR has two parameters, NR and AINR; the ASR server includes an ASR module and a network transmission and reception device; the NLP server includes an NLP module and a network transmission and reception device; the skill server includes a skill distribution module, a skill execution module, and a network transmission and reception device.
[0061] Optionally, the interaction steps are as follows: (1) The smart terminal device uploads the target audio processed by the 3A algorithm to the ASR server; (2) The ASR server processes the audio and identifies the text information, and returns the identified text information to the smart terminal device; (3) The ASR server sends the identified text information to the NLP server; (4) The NLP server processes the text information to obtain accurate semantics and passes it to the skill server; (5) The skill server issues skill instructions to the smart terminal device according to the semantics.
[0062] As an optional implementation method, Figure 4 This is a flowchart illustrating an audio recording quality enhancement method according to an embodiment of this application. The specific process includes:
[0063] Step S1: System initialization. The smart terminal device uses the preset audio data for 3A algorithm processing. Specifically, the smart terminal device (hereinafter referred to as the device) sends the pre-collected audio data to the 3A algorithm module (equivalent to the audio algorithm component in the above embodiment). The AGC algorithm is configured with empirical values by default, the AEC algorithm is configured with ordinary AEC parameters by default, and the ANR algorithm is configured with ordinary NR parameters by default. The target audio is obtained after processing.
[0064] Step S2: The intelligent underfloor heating device sends the processed data to the ASR server. Specifically, the digital human application on the device uploads the target audio to the ASR server and identifies the request as audio quality verification mode.
[0065] Step S3: After processing the audio data, the ASR server returns the target text data to the smart terminal device. That is, the ASR server processes the uploaded audio, generates the target text data, and returns the target text data to the device.
[0066] Step S4: The smart terminal device calculates the similarity between the target text information and a locally preset keyword data set. Specifically, the device receives the target text data and calculates the similarity between it and the locally preset keyword data set. If the calculated similarity S1 (equivalent to the first similarity value in the above embodiment) is greater than the preset keyword threshold T1 and the similarity S2 (equivalent to the second similarity value in the above embodiment) is less than the preset onomatopoeia threshold T2, then the current 3A algorithm parameters are considered to meet the ASR server's processing capabilities and are used for audio processing during subsequent interaction with the ASR server, proceeding to S8. If the similarity S1 is less than or equal to the preset keyword threshold, or the similarity S2 is greater than or equal to the preset onomatopoeia threshold, proceed to S5-S6-S7. Similarity S2 primarily serves to enhance the verification of audio quality processing in the 3A algorithm. It ensures that the ASR server does not recognize unnecessary onomatopoeia when processing audio, thus avoiding impacting subsequent NLP and skills. For example, the audio might be expected to recognize the text data "I want to watch a movie," but due to onomatopoeia interference, it might recognize the text data "I want to watch a movie," affecting subsequent NLP processing. However, this situation generally only occurs in complex scenarios with high external noise. Therefore, on low-performance devices, similarity S2 can be omitted, and only similarity S1 can be calculated. Then, similarity comparisons can be performed, proceeding through S6-S7, while ignoring S5, reducing the device's overhead in similarity calculation and the 3A algorithm.
[0067] It should be noted that the target text data acquired by the device is DT, which can be considered as the set of all characters in the text data. The keyword set corresponding to the locally recorded audio is {PA, PB, PC, ..., PN}, which depends on the preset locally recorded audio. Each keyword can be considered as the set of characters in the keyword. The onomatopoeic word set is OT. The similarity S1 between each keyword set and DT is calculated using the formula: S1=(DT∩PA) / (DT∪PA)+(DT∩PB) / (DT∪PB)+(DT∩PC) / (DT∪PC)+……+(DT∩PN) / (DT∪PN); the similarity S2 between the onomatopoeic word set is calculated using the formula: S2=(DT∩OT) / (DT∪OT).
[0068] Furthermore, since different keywords appear with varying frequencies during user interactions, to improve the accurate identification of high-frequency keywords, keyword weights can be added to the above calculation formula. For example, the keyword weight set could be {WA, WB, WC, ..., WN}, and the keyword calculation formula could then be adjusted to: S1 = WA(DT∩PA) / (DT∪PA) + WB(DT∩PB) / (DT∪PB) + WC(DT∩PC) / (DT∪PC) + ... + WN(DT∩PN) / (DT∪PN). Thus, by adjusting the weight ratio of different keywords in the similarity calculation formula, high-frequency words can be effectively identified.
[0069] If the similarity S1 is less than or equal to the preset keyword threshold T1, it indicates that the target text data returned by ASR contains missing characters (missing characters will reduce the intersection of the keyword set and DT, while the union remains unchanged; in the calculation formula, the numerator decreases while the denominator remains unchanged, leading to a decrease in similarity) or redundant text (redundant text will increase the length of DT; in the similarity calculation formula, the union of DT and the keyword set increases while the intersection remains unchanged; the denominator increases while the numerator remains unchanged, leading to a decrease in similarity). It can be assumed that the target audio data processed by the 3A algorithm has insufficient energy or that the AEC processing capability is insufficient. Since setting the AEC algorithm to AIAEC parameters places high demands on device performance, the AGC parameters should be adjusted first. If adjusting the AGC algorithm parameters still cannot make the similarity S1 greater than the preset keyword threshold T1, the AEC algorithm needs to be set to AIAEC to improve the AEC algorithm processing capability. If the similarity S1 is greater than the preset keyword threshold T1, there is no need to adjust the AGC and AEC parameters of the 3A algorithm module. Using ordinary AEC and AGC parameters can meet the current ASR server processing capability and reduce the device resource usage of the 3A algorithm module.
[0070] Step S5: The smart terminal device calculates the similarity between the target text information and the locally preset onomatopoeic word data set. Specifically, if the similarity S2 is less than the onomatopoeic word threshold T2, there is no need to adjust the ANR parameters of the 3A algorithm module; ordinary NR parameters are sufficient, which can meet the current ASR server's processing capacity and reduce the 3A algorithm module's resource usage on the device. If the similarity S2 is greater than or equal to the preset onomatopoeic word threshold T2, it is considered that the NR parameters of the ANR algorithm cannot match the current ASR server's processing capacity. In this case, the ANR algorithm cannot handle environmental noise well. Some ASR servers may identify incompletely eliminated environmental noise in the target audio data, such as keyboard sounds or wind sounds, as onomatopoeia. In this case, it is necessary to improve the ANR capability of the device's 3A algorithm module by setting the ANR algorithm to AINR parameters to enhance the ANR algorithm capability. After the 3A algorithm parameters are optimized, proceed to step S6.
[0071] Step S6: Save the optimized 3A algorithm parameters. The device's 3A algorithm module loads the optimized parameters and reprocesses the preset audio data. The device sends the reprocessed target audio data to the ASR server. The ASR server returns the processed target text data to the device. Then, the device recalculates the similarity. If the similarity still does not meet the preset threshold (similarity S1 is less than or equal to the preset keyword threshold, or similarity S2 is greater than or equal to the preset onomatopoeia threshold), proceed to S4-S5. On low-performance devices, proceed to S4 and ignore S5. If similarity S1 is greater than the preset keyword threshold T1 and similarity S2 is less than the preset onomatopoeia threshold T2, the audio quality verification stage ends. The device saves the final adjusted 3A algorithm parameters and proceeds to S7.
[0072] Step S7: The user uses the AI voice application on the smart terminal device and performs voice interaction. The audio recording module of the smart terminal device collects audio; the 3A algorithm module loads the optimized parameters; that is, the device collects audio data in the environment through the audio recording module and sends the audio data to the 3A algorithm module for processing. The 3A algorithm module uses the optimal 3A algorithm parameters.
[0073] Step S8: The smart terminal device sends the collected raw audio data to the 3A algorithm module for processing. The AI voice application on the smart terminal device obtains the processed audio data from the 3A algorithm module and uploads it to the ASR server. Furthermore, when sending the processed audio data to the ASR server, it also identifies the request as a normal interaction mode.
[0074] Step S9: The ASR server processes the audio data and returns the recognized text information (i.e., the target text data) to the AI voice application on the smart terminal device, while simultaneously sending the data to the NLP server. Optionally, the device can also display the processed target text data for the user to view in real time.
[0075] Step S10: The NLP server receives the text information sent by the ASR server, performs semantic processing on the text information, identifies the correct semantics, obtains the skills that the user wants the device to provide, and notifies the skills server.
[0076] Step S11: The skill server sends the skill command corresponding to the skill specified by the user to the AI voice application on the smart terminal device. The AI voice application on the smart terminal device executes the skill command and interacts with the user to form a complete voice interaction.
[0077] It should be noted that if you need to adapt to a new ASR server, you can repeat steps S1-S11.
[0078] As an optional implementation, the aforementioned 3A algorithm parameter optimization can be applied to a digital human application on a smart speaker. The specific process includes: the smart speaker initiates 3A algorithm parameter optimization; the digital human application processes locally preset audio data using the 3A algorithm with preset parameters and sends the target audio data to the ASR server; the ASR server processes the audio data and returns the text data identified from the audio data to the digital human application on the smart speaker; after receiving the text information, the digital human application compares it with a preset keyword set and an onomatopoeia set; if the similarity does not meet the preset similarity threshold, the 3A algorithm parameters are adjusted, and the preset audio data is reprocessed; the reprocessed audio data is then sent to the ASR server until the similarity between the target text data and the keyword set is greater than the preset keyword set threshold, and the similarity with the onomatopoeia set is less than the preset onomatopoeia set threshold. Then, the 3A algorithm parameter optimization step ends, and the parameters are saved for subsequent audio data processing interactions with the ASR server.
[0079] Optionally, users can interact with the digital human in real time. For example, a user can speak into the smart speaker's microphone, such as "What's the weather like today?" The smart speaker's audio acquisition module collects the raw audio and sends it to the 3A algorithm module for processing. The digital human application obtains the processed audio data from the 3A algorithm module and sends it to the ASR server. The ASR server processes the audio data, recognizes the text "What's the weather like today?", and sends it to the NLP server. The NLP server identifies the corresponding semantics, knows the user wants to check the weather, and then sends the specific semantics to the skill server. The skill server then issues the specified skill to the smart speaker. The smart speaker's digital human application receives the specified skill from the skill server, such as weather, and displays the day's weather image and provides voice announcements.
[0080] Optionally, when faced with ASR servers provided by different manufacturers, the locally preset algorithm parameters in related technologies are often difficult to meet the requirements of all servers at once, resulting in low efficiency in multi-device debugging scenarios. Furthermore, when the environment of the smart terminal changes, the fixed parameters may no longer be applicable, leading to a decrease in the quality of recorded audio and affecting the performance of speech recognition.
[0081] To ensure the quality of subsequent audio, when users pay to upgrade to a more powerful AI digital human, and need to adapt to a new ASR server, the smart terminal device needs to recalibrate the 3A algorithm parameters and retain a new set of 3A algorithm parameters. When users switch to different versions of AI digital human, different 3A algorithm parameters are used to process the audio data. This not only dynamically adapts to different ASR servers, but also effectively reduces the use of device resources by the 3A algorithm, improving the user's AI application experience on the smart terminal device.
[0082] In summary, through the above solution, the smart terminal device sends the collected audio data to the audio 3A algorithm module on the device for processing to obtain the target audio; then, the target audio is sent to the ASR server for automatic speech recognition processing; the ASR server returns the processed target text data to the device, and the device calculates the similarity between the target text data and the keyword set and onomatopoeia set. If the similarity does not meet the preset threshold, the device modifies the audio 3A algorithm parameters until the similarity meets the preset threshold; the smart terminal device saves the optimized 3A algorithm parameters, and the 3A algorithm module uses these parameters to process the recorded audio, improving audio quality, and then sends the processed target audio to the ASR server, which can effectively improve the accuracy of text recognition. By applying the technical solution of this application, the application software on the smart terminal device can efficiently adapt when accessing different ASR servers, improving the user experience.
[0083] This embodiment also provides an audio recording quality enhancement device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.
[0084] Figure 5 This is a structural block diagram of an audio recording quality enhancement device according to an embodiment of this application. Figure 5 As shown, the audio recording quality enhancement device includes:
[0085] Processing module 52 is used to process the first audio obtained by the smart terminal using an audio algorithm component containing a preset set of adjustment parameters to obtain a second audio, wherein the first audio is the audio generated by the target object by reading the audio verification text displayed on the display interface of the smart terminal or the preset verification audio in the smart terminal.
[0086] The calculation module 54 is used to obtain the first text data generated by the first server recognizing the second audio, and to perform similarity calculation based on the first text data and the preset keyword set set on the smart terminal to obtain the first similarity value corresponding to the keyword;
[0087] The update module 56 is used to update the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold when the performance type corresponding to the smart terminal is confirmed, so as to obtain a first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component.
[0088] Using the aforementioned device, an audio algorithm component containing a preset set of adjustment parameters processes the first audio acquired by the smart terminal to obtain the second audio. Subsequently, the second audio is synchronized to a first server for recognition, obtaining the first text data generated by the first server recognizing the second audio. Based on knowledge of the smart terminal's performance type, a first similarity value is calculated according to the first text data and a preset set of keywords on the smart terminal. Then, a first size relationship is determined between the first similarity value and a preset keyword threshold. The parameter values in the preset set of adjustment parameters are updated according to this relationship to generate a first set of adjustment parameters. The audio algorithm component then enhances the audio recording quality by using this first set of adjustment parameters. This method solves the problem in related technologies where the audio gain parameters used in the algorithm cannot be dynamically adjusted. It achieves the goal of finding a suitable set of adjustment parameters for the audio algorithm component, thereby improving the audio quality of the recorded audio acquired by the smart terminal during user interaction and ensuring high-quality voice interaction by the smart terminal.
[0089] Furthermore, by employing the above method, even with limited resources, the smart terminal can automatically find the optimal 3A algorithm parameters (equivalent to the first set of adjustment parameters in the above embodiments) through efficient interaction with an external ASR server to improve audio recording quality. This, in turn, reduces reliance on the smart terminal's hardware resources while maintaining audio quality, and reduces the recognition error rate caused by mismatched algorithm parameters. This allows users to obtain high-quality speech recognition and response without intervention, ensuring smooth voice interaction between the smart terminal and the user in both quiet and noisy environments.
[0090] In one exemplary embodiment, the above apparatus further includes: a determining module, configured to, after obtaining first text data generated by the first server recognizing the second audio when the second audio is synchronized to the first server, obtain hardware configuration information of the smart terminal; and determine the performance type of the smart terminal according to the hardware configuration, wherein the performance type includes at least one of the following: a low-performance device, a high-performance device.
[0091] In an exemplary embodiment, the above-described updating module is further configured to allow parameter value updates based solely on the first similarity value when the smart terminal is a low-performance device; wherein, updating the parameter value based on the first similarity value includes: updating the default values of the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set when the first size relationship is such that the first similarity value is less than or equal to a preset keyword threshold, to obtain a target gain parameter value and a target echo cancellation parameter value; and updating the preset adjustment parameter set based on the target gain parameter value and the target echo cancellation parameter value.
[0092] In an exemplary embodiment, the above-mentioned updating module is further configured to, when the smart terminal is a high-performance type device, determine the similarity between the first text data and the preset onomatopoeia set to obtain a second similarity value corresponding to the onomatopoeia; determine a second size relationship between the second similarity value and the preset onomatopoeia threshold; and update the parameter values in the preset adjustment parameter set simultaneously through the second size relationship and the first size relationship.
[0093] In an exemplary embodiment, the above-described updating module is further configured to: update the parameter values of the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set when the first size relationship is a first similarity value less than or equal to a preset keyword threshold, to obtain a target gain parameter value and a target echo cancellation parameter value; update the parameter values of the preset noise reduction parameter in the preset adjustment parameter set when the second size relationship is a second similarity value greater than or equal to a preset onomatopoeia threshold, to obtain a target noise reduction parameter value; and update the parameter values in the preset adjustment parameter set based on the target gain parameter value, the target echo cancellation parameter value, and the target noise reduction parameter value.
[0094] In an exemplary embodiment, the above-described apparatus further includes: a verification module, configured to update the parameter values of a preset adjustment parameter set based on a first size relationship and a second size relationship to obtain a first adjustment parameter set for enhancing the audio recording quality processed by the audio algorithm component; then, using an audio algorithm component containing the first adjustment parameter set to perform verification processing on the first audio to obtain a third audio; synchronize the third audio to the first server and receive second text data fed back by the first server; perform similarity calculation based on the performance type corresponding to the smart terminal, the second text data, the preset keyword set, and the preset onomatopoeia set to obtain a third similarity value corresponding to the keyword and a fourth similarity value corresponding to the onomatopoeia; and, if the third size relationship between the third similarity value and a preset keyword threshold and the fourth size relationship between the fourth similarity value and a preset onomatopoeia threshold meet preset relationship conditions, determine the first adjustment parameter set as the final result of the parameter value update, wherein the preset relationship conditions include at least: the similarity value at the keyword level is greater than the preset keyword threshold, and the similarity value at the onomatopoeia level is less than the preset onomatopoeia threshold.
[0095] In an exemplary embodiment, the above-described apparatus further includes: an interaction module, configured to update the parameter values of a preset adjustment parameter set based on a first size relationship and a second size relationship to obtain a first adjustment parameter set for enhancing the audio recording quality processed by the audio algorithm component; and, when the audio recording module on the smart terminal generates target audio data, process the target audio using the audio algorithm component that has completed the parameter value update to obtain enhanced audio data, wherein the enhanced audio data includes at least: target enhanced audio and a normal interaction identifier for indicating the interaction state between the smart terminal and the first server; acquire target text data generated by the first server processing the enhanced audio data, and receive skill instructions issued by the natural language processing server based on the target text data; and provide voice interaction feedback to the target object operating the current smart terminal based on the target text data and the skill instructions.
[0096] In an exemplary embodiment, the above-mentioned interaction module further includes: an upgrade unit, configured to determine the upgrade type of the active upgrade when the target object actively upgrades the voice interaction function of the smart terminal; when the upgrade type is to replace the first server, synchronize the second audio to the second server; obtain third text data generated by the second server recognizing the second audio, and perform similarity calculation based on the performance type corresponding to the smart terminal, the third text data, the preset keyword set, and the preset onomatopoeia set to obtain a fifth similarity value corresponding to the keyword and a sixth similarity value corresponding to the onomatopoeia; determine a fifth size relationship between the fifth similarity value and the preset keyword threshold and a sixth size relationship between the sixth similarity value and the preset onomatopoeia threshold, and update the parameter values of the preset adjustment parameter set based on the fifth size relationship and the sixth size relationship to obtain a second adjustment parameter set, wherein the second adjustment parameter set is a parameter set determined after the smart terminal connects to the second server for enhancing the audio recording quality of the audio algorithm component.
[0097] In an exemplary embodiment, the above-described apparatus further includes: a recording module, configured to update the parameter values of a preset adjustment parameter set based on a first size relationship and a second size relationship to obtain a first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component; collect configuration information of the first server; associate and save the configuration information with the preset adjustment parameter set updated using the first adjustment parameter set to generate recorded data of the smart terminal configuration server; and, when the smart terminal is connected to a third server and the configuration information of the third server is the same as that of the first server, perform audio recording quality enhancement settings on the smart terminal based on the recorded data.
[0098] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, smart speaker, or network device, etc.) to execute the methods of the various embodiments of this application.
[0099] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0100] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0101] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0102] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0103] S1, using an audio algorithm component containing a preset set of adjustment parameters to process the first audio obtained by the smart terminal to obtain the second audio, wherein the first audio is the audio generated by the target object by reading the audio verification text displayed on the display interface of the smart terminal or the preset verification audio in the smart terminal.
[0104] S2, obtain the first text data generated by the first server recognizing the second audio, and perform similarity calculation based on the first text data and the preset keyword set set on the smart terminal to obtain the first similarity value corresponding to the keyword;
[0105] S3, upon confirming the performance type corresponding to the smart terminal, update the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold, to obtain a first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component.
[0106] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0107] Optionally, in this embodiment, the electronic device may also be configured to perform step S1 via a computer program.
[0108] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0109] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0110] Embodiments of this application also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0111] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0112] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0113] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for enhancing audio recording quality, characterized in that, include: The first audio obtained by the smart terminal is processed using an audio algorithm component containing a set of preset adjustment parameters to obtain a second audio, wherein the first audio is the audio generated by the target object by reading the audio verification text displayed on the display interface of the smart terminal or the verification audio preset in the smart terminal. The system obtains the first text data generated by the first server recognizing the second audio, and performs similarity calculation based on the first text data and the preset keyword set set on the smart terminal to obtain the first similarity value corresponding to the keyword. Upon confirming the performance type of the smart terminal, the parameter values in the preset adjustment parameter set are updated according to the first size relationship between the first similarity value and the preset keyword threshold, thereby obtaining a first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component. The step of updating the parameter values in the preset adjustment parameter set based on the first size relationship between the first similarity value and the preset keyword threshold includes: In the case where the smart terminal is a low-performance device, parameter value updates are allowed only based on the first similarity value; wherein, updating parameter values based on the first similarity value includes: when the first similarity value is less than or equal to a preset keyword threshold, updating the default values of the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set to obtain the target gain parameter value and the target echo cancellation parameter value; updating the preset adjustment parameter set based on the target gain parameter value and the target echo cancellation parameter value; When the smart terminal is a high-performance device, the similarity between the first text data and the preset onomatopoeia set is determined to obtain the second similarity value corresponding to the onomatopoeia; the second similarity value and the second size relationship corresponding to the preset onomatopoeia threshold are determined; the parameter values in the preset adjustment parameter set are updated simultaneously through the second size relationship and the first size relationship, wherein the preset onomatopoeia set contains preset noise words that should not be recognized.
2. The method for enhancing audio recording quality according to claim 1, characterized in that, After obtaining the first text data generated by the first server recognizing the second audio, the method further includes: Obtain the hardware configuration information of the smart terminal; The performance type of the smart terminal is determined based on the hardware configuration, wherein the performance type includes at least one of the following: low-performance device, high-performance device.
3. The method for enhancing audio recording quality according to claim 1, characterized in that, Updating the parameter values in the preset adjustment parameter set simultaneously through the second size relationship and the first size relationship includes: When the first similarity value is less than or equal to the preset keyword threshold, update the parameter values of the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set to obtain the target gain parameter value and the target echo cancellation parameter value. When the second similarity value is greater than or equal to the preset onomatopoeia threshold, the parameter values of the preset noise reduction parameters in the preset adjustment parameter set are updated to obtain the target noise reduction parameter value. The parameter values in the preset adjustment parameter set are updated based on the target gain parameter value, the target echo cancellation parameter value, and the target noise reduction parameter value.
4. The method for enhancing audio recording quality according to claim 1, characterized in that, After updating the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold, and obtaining the first adjustment parameter set for enhancing the audio recording quality processed by the audio algorithm component, the method further includes: The first audio is verified using an audio algorithm component containing the first set of adjustment parameters to obtain the third audio. The third audio is synchronized to the first server, and the second text data fed back by the first server is received. Based on the performance type of the smart terminal, the second text data, the preset keyword set, and the preset onomatopoeia set, similarity calculation is performed to obtain the third similarity value corresponding to the keyword and the fourth similarity value corresponding to the onomatopoeia. If the third similarity value and the preset keyword threshold and the fourth similarity value and the preset onomatopoeia threshold meet the preset relationship conditions, the first set of adjustment parameters is determined as the final result of parameter value update. The preset relationship conditions include at least: the similarity value is greater than the preset keyword threshold at the keyword level, and the similarity value is less than the preset onomatopoeia threshold at the onomatopoeia level.
5. The method for enhancing audio recording quality according to claim 1, characterized in that, After updating the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold, and obtaining the first adjustment parameter set for enhancing the audio recording quality processed by the audio algorithm component, the method further includes: When the audio recording module on the smart terminal generates target audio data, the target audio is processed by the audio algorithm component that has completed parameter value updates to obtain enhanced audio data. The enhanced audio data includes at least: target enhanced audio and a normal interaction identifier for indicating the interaction status between the smart terminal and the first server. The system acquires target text data generated by the first server processing the enhanced audio data, and receives skill instructions issued by the natural language processing server based on the target text data. Based on the target text data and the skill instructions, voice interaction feedback is provided to the target object operating the current smart terminal.
6. The method for enhancing audio recording quality according to claim 5, characterized in that, After providing voice interaction feedback to the target object operating the current smart terminal based on the target text data and the skill command, the method further includes: When the target object actively upgrades the voice interaction function of the smart terminal, determine the upgrade type of the active upgrade; If the upgrade type is to replace the first server, the second audio will be synchronized to the second server; The third text data generated by the second server recognizing the second audio is obtained. Based on the performance type corresponding to the smart terminal, the third text data, the preset keyword set, and the preset onomatopoeia set, similarity calculation is performed to obtain the fifth similarity value corresponding to the keyword and the sixth similarity value corresponding to the onomatopoeia. The fifth similarity value and the preset keyword threshold are determined, and the sixth similarity value and the preset onomatopoeia threshold are determined. Based on the fifth similarity value and the sixth similarity value, the preset adjustment parameter set is updated to obtain the second adjustment parameter set. The second adjustment parameter set is a set of parameters determined by the smart terminal after connecting to the second server to enhance the audio recording quality of the audio algorithm component.
7. The method for enhancing audio recording quality according to claim 1, characterized in that, After updating the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold, and obtaining the first adjustment parameter set for enhancing the audio recording quality processed by the audio algorithm component, the method further includes: Collect the configuration information of the first server; The configuration information is associated with and saved with the preset adjustment parameter set updated using the first adjustment parameter set, thereby generating the record data of the smart terminal configuration server; When the smart terminal is connected to a third server and the configuration information of the third server is the same as that of the first server, the audio recording quality of the smart terminal is enhanced based on the recorded data.
8. An audio recording quality enhancement device, characterized in that, include: The processing module is used to process the first audio obtained by the smart terminal using an audio algorithm component containing a preset set of adjustment parameters to obtain the second audio, wherein the first audio is the audio generated by the target object by reading the audio verification text displayed on the display interface of the smart terminal or the preset verification audio in the smart terminal. The calculation module is used to obtain the first text data generated by the first server recognizing the second audio, and to perform similarity calculation based on the first text data and the preset keyword set set on the smart terminal to obtain the first similarity value corresponding to the keyword; The update module is used to update the parameter values in the preset adjustment parameter set according to the first size relationship between the first similarity value and the preset keyword threshold when the performance type corresponding to the smart terminal is confirmed, so as to obtain a first adjustment parameter set for enhancing the audio recording quality of the audio algorithm component. The update module is further configured to, when the smart terminal is a low-performance device, allow parameter value updates solely based on the first similarity value; wherein, updating parameter values based on the first similarity value includes: updating the default values of the preset automatic gain parameter and / or preset echo cancellation parameter in the preset adjustment parameter set when the first similarity value is less than or equal to a preset keyword threshold, to obtain a target gain parameter value and a target echo cancellation parameter value; updating the preset adjustment parameter set based on the target gain parameter value and the target echo cancellation parameter value; when the smart terminal is a high-performance device, determining the similarity between the first text data and a preset onomatopoeia set, to obtain a second similarity value corresponding to the onomatopoeia; determining a second similarity relationship between the second similarity value and a preset onomatopoeia threshold; and simultaneously updating the parameter values in the preset adjustment parameter set using both the second similarity relationship and the first similarity relationship, wherein the preset onomatopoeia set includes preset noise words that should not be recognized.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method described in any one of claims 1 to 7 when it is run.
10. An electronic device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program and the processor is configured to perform the method described in any one of claims 1 to 7 via the computer program.
11. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Mobile terminal audio calibrating method and automatic testing system
CN101917735A
Method and device for improving audio processing performance
CN104980337A