Terminal device and method for quality enhancement of an audio signal

CN122821996APending Publication Date: 2026-09-25HISENSE VISUAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610905002.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供一种终端设备及音频信号的质量增强方法,以解决现有音频采集技术中采集到的音频信号质量低的问题

Benefits of technology

[0007]上述技术方案具有以下有益效果或优点:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821996A_ABST
    Figure CN122821996A_ABST
Patent Text Reader

Abstract

The application provides a terminal device and a quality enhancement method of an audio signal. The method can receive PDM audio data collected by a sound collector and convert the PDM audio data into PCM audio data, and perform audio quality detection on the PCM audio data. When it is detected that the audio quality detection results of a continuous preset number of times indicate abnormalities, the sampling point position of the PDM audio data is gradually adjusted in a preset adjustment range by a preset step unit, audio quality detection is performed on the PCM audio data after each adjustment of the sampling point position, and the audio quality detection results corresponding to the sampling point position are recorded. The optimal sampling point position is determined in the sampling point position indicated by the audio quality detection result as normal, and subsequent sampling is performed according to the optimal sampling point position. The method dynamically and adaptively calibrates the sampling point position, so that the sampling time always adapts to the current hardware timing state of the device, and the audio signal quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal equipment technology, and in particular to a terminal equipment and a method for enhancing the quality of audio signals. Background Technology

[0002] Terminal devices refer to devices capable of performing specific functions, such as smart TVs, communication terminals, smart advertising screens, projectors, printers, scanners, etc. Taking smart TVs as an example, a smart TV is a television product based on Internet application technology, equipped with an open operating system and chip, possessing an open application platform, enabling two-way human-computer interaction, and integrating multiple functions such as audio-visual, entertainment, and data to meet diverse and personalized user needs.

[0003] Terminal devices can be configured with a sound acquisition unit to enable long-distance voice wake-up and voice command recognition. Specifically, the controller in the terminal device (such as a system-on-chip (SOC)) can communicate with the sound acquisition unit via the I²S interface protocol. The sound acquisition unit converts the acquired audio signal into Pulse Density Modulation (PDM) audio data, which is transmitted to the controller via clock and data signal lines. The controller samples the level of the data signal lines according to preset sampling point positions and converts the sampled PDM audio data into Pulse Code Modulation (PCM) audio data for subsequent speech recognition processing.

[0004] Ideally, the controller should sample PDM audio data at specific edges of the clock signal (such as rising or falling edges), at which point the level on the data signal line is stable and accurately reflects the audio signal acquired by the sound acquisition unit. However, in actual operation, due to inherent delays in hardware circuits, clock signal jitter, electromagnetic interference, component aging, and ambient temperature drift, the timing relationship between the clock signal and the data signal will shift. This timing shift causes the controller to read data at the preset sampling time that is not the stable level actually output by the sound acquisition unit, but rather a random value in the data transition region. This introduces sampling errors and high-frequency noise, reducing the quality of the far-field audio signal and consequently degrading far-field speech recognition performance, thus affecting the user's voice interaction experience. Summary of the Invention

[0005] This application provides a terminal device and a method for enhancing the quality of audio signals to solve the problem of low audio signal quality acquired in existing audio acquisition technologies.

[0006] In a first aspect, this application provides a terminal device, including: Sound acquisition device; The controller is configured as follows: Receive pulse density modulation (PDM) audio data output by the sound acquisition unit, and convert the PDM audio data into pulse code modulation (PCM) audio data; The PCM audio data is subjected to audio quality detection to obtain audio quality detection results; When an abnormal audio quality detection result is detected for a preset number of consecutive times, the sampling point position of the PDM audio data is gradually adjusted within a preset adjustment range in a preset step unit, and audio quality detection is performed on the PCM audio data after each adjustment of the sampling point position, and the audio quality detection result corresponding to the sampling point position is recorded. The optimal sampling point position is determined from the sampling point positions indicated by the audio quality detection results as normal. The PDM audio data output by the sound acquisition device is sampled according to the optimal sampling point position.

[0007] The above technical solution has the following beneficial effects or advantages: This solution samples PDM audio data output from a sound acquisition unit and converts it to PCM audio data, while simultaneously performing audio quality checks. When a preset number of consecutive audio quality anomalies are detected, the sampling point positions are automatically adjusted step-by-step within a preset adjustment range. Audio quality checks are performed on the PCM audio data after each adjustment, and the results are recorded. Finally, the optimal sampling point position is determined from all normally detected sampling point positions, and subsequent sampling is performed accordingly. This technical solution triggers the adjustment mechanism only after multiple consecutive anomaly detections, avoiding false triggers caused by occasional noise or transient interference. By recording the audio quality check results at each sampling point position, the audio quality of each candidate sampling point position can be comprehensively evaluated throughout the preset adjustment range, thereby accurately locating the optimal sampling point position and ensuring that the audio signal quality acquired in the far field is always at its best. Dynamic adjustment of the sampling point positions ensures that the sampling time always adapts to the current hardware timing state of the device, improving audio signal quality and far-field voice interaction performance.

[0008] In some embodiments of this application, the controller performs audio quality detection on the PCM audio data to obtain audio quality detection results, specifically configured as follows: Perform a Fast Fourier Transform on the PCM audio data to obtain the spectrum of the PCM audio data; When the energy of a preset high-frequency band in the spectrum exceeds a preset energy threshold, the audio quality detection result is determined to be abnormal. When the energy of a preset high-frequency band in the spectrum is less than or equal to a preset energy threshold, the audio quality detection result is determined to be normal.

[0009] The above technical solution has the following beneficial effects or advantages: The audio data is converted from the time domain to the frequency domain for analysis using the Fast Fourier Transform (FFT). The audio quality is determined by comparing the high-frequency energy with a preset energy threshold, and the signal characteristics of high-frequency noise being amplified after sampling errors caused by timing offset are accurately matched.

[0010] In some embodiments of this application, the controller determines the optimal sampling point position from the sampling point positions indicated by the audio quality detection results, specifically configured as follows: Within the sampling point locations where the audio quality detection results indicate normal operation, determine the range of locations with the largest consecutive sampling points. Calculate the midpoint position within the maximum continuous sampling point position interval to obtain the optimal sampling point position.

[0011] The above technical solution has the following beneficial effects or advantages: By selecting the range of the largest continuous sampling point locations instead of individual sampling point locations, the randomness that may exist in the locations of isolated sampling points can be eliminated, making the selected optimal sampling point location more robust. Furthermore, taking the midpoint of the range as the optimal sampling point location ensures that the sampling point has sufficient fault tolerance margin on both sides, so that even if slight timing drift occurs later, it can remain within the normal range, thereby extending the stable operating time after adjustment.

[0012] In some embodiments of this application, the controller calculates the midpoint position within the maximum continuous sampling point position interval to obtain the optimal sampling point position, specifically configured as follows: Determine the position of the maximum sampling point and the position of the minimum sampling point within the maximum continuous sampling point position interval; The average value between the value at the maximum sampling point position and the value at the minimum sampling point position is calculated to obtain the value at the midpoint position.

[0013] The above technical solution has the following beneficial effects or advantages: The optimal sampling point location can be determined by averaging, which has low computational complexity and reduces the consumption of computing resources.

[0014] In some embodiments of this application, the controller is further configured to: The preset step unit is a fraction of the clock cycle corresponding to the PDM audio data. The preset adjustment range is defined as at least half of the clock cycle.

[0015] The above technical solution has the following beneficial effects or advantages: Using fractions of a clock cycle as the step unit enables fine-tuning of the sampling point position. Setting the adjustment range to at least half a clock cycle ensures that even with large timing offsets, the effective sampling point position can still be fully scanned.

[0016] In some embodiments of this application, the controller performs audio quality detection on the PCM audio data to obtain an audio quality detection result, and is configured as follows: The PCM audio data is subjected to audio quality testing according to a preset time period.

[0017] The above technical solution has the following beneficial effects or advantages: Using a timed polling method for periodic detection can ensure timely detection of timing deviations while avoiding the problem of continuous detection consuming computing resources.

[0018] In some embodiments of this application, the controller gradually adjusts the sampling point position of the PDM audio data within a preset adjustment range in preset step units, specifically configured as follows: Within a preset adjustment range, the register value of the sampling phase register is gradually adjusted in preset step units; the sampling phase register is used to control the sampling point position of the PDM audio data; The controller samples the PDM audio data acquired by the sound collector according to the optimal sampling point position, specifically configured as follows: Write the register value corresponding to the optimal sampling point position into the sampling phase register.

[0019] The above technical solution has the following beneficial effects or advantages: The sampling point position is adjusted and locked by reading and writing registers. The sampling timing is controlled directly by the built-in hardware registers without modifying the hardware circuit, which is low-cost and easy to deploy.

[0020] In some embodiments of this application, the controller converts the PDM audio data into Pulse Code Modulation (PCM) audio data, specifically configured as follows: The PDM audio data is low-pass filtered and sampled and quantized to obtain the PCM audio data, and the PCM audio data is stored in a preset buffer. The controller performs audio quality detection on the PCM audio data and obtains the audio quality detection result, specifically configured as follows: The PCM audio data is obtained from the buffer, and the audio quality of the PCM audio data is detected.

[0021] The above technical solution has the following beneficial effects or advantages: The conversion from PDM audio data to PCM audio data is completed by low-pass filtering and sampling quantization. The buffer is used to temporarily store and asynchronously detect the audio data, decoupling the audio quality detection process from the main audio sampling path. The detection operation will not interfere with the normal transmission of the real-time audio stream and subsequent speech recognition processing, ensuring the continuity of far-field speech function operation.

[0022] In some embodiments of this application, the controller is further configured to: Create an exception counter; After each audio quality detection result is obtained, if the audio quality detection result indicates an abnormality, the abnormality counter is incremented; if the audio quality detection result indicates a normality, the abnormality counter is cleared. When the count value of the abnormality counter reaches the preset number of times, the sampling point position of the PDM audio data is gradually adjusted within the preset adjustment range in preset step units.

[0023] The above technical solution has the following beneficial effects or advantages: An anomaly counter is used to count the number of consecutive anomalies. The sampling point position adjustment is only triggered when there are a preset number of consecutive anomalies, thus avoiding false triggering caused by instantaneous interference or sudden noise.

[0024] Secondly, this application provides a method for enhancing the quality of an audio signal, comprising: Receive pulse density modulation (PDM) audio data output from the sound acquisition unit, and convert the PDM audio data into pulse code modulation (PCM) audio data; The PCM audio data is subjected to audio quality detection to obtain audio quality detection results; When an abnormal audio quality detection result is detected for a preset number of consecutive times, the sampling point position of the PDM audio data is gradually adjusted within a preset adjustment range in a preset step unit, and audio quality detection is performed on the PCM audio data after each adjustment of the sampling point position, and the audio quality detection result corresponding to the sampling point position is recorded. The optimal sampling point position is determined from the sampling point positions indicated by the audio quality detection results as normal. The PDM audio data output by the sound acquisition device is sampled according to the optimal sampling point position.

[0025] The above technical solution has the following beneficial effects or advantages: This solution samples PDM audio data output from a sound acquisition unit and converts it to PCM audio data, while simultaneously performing audio quality checks. When a preset number of consecutive audio quality anomalies are detected, the sampling point positions are automatically adjusted step-by-step within a preset adjustment range. Audio quality checks are performed on the PCM audio data after each adjustment, and the results are recorded. Finally, the optimal sampling point position is determined from all normally detected sampling point positions, and subsequent sampling is performed accordingly. This technical solution triggers the adjustment mechanism only after multiple consecutive anomaly detections, avoiding false triggers caused by occasional noise or transient interference. By recording the audio quality check results at each sampling point position, the audio quality of each candidate sampling point position can be comprehensively evaluated throughout the preset adjustment range, thereby accurately locating the optimal sampling point position and ensuring that the audio signal quality acquired in the far field is always at its best. Dynamic adjustment of the sampling point positions ensures that the sampling time always adapts to the current hardware timing state of the device, improving audio signal quality and far-field voice interaction performance. Attached Figure Description

[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 This is a schematic diagram illustrating the operational scenarios between a terminal device and a control device provided in some embodiments of this application; Figure 2 This is a schematic diagram of the hardware configuration of a terminal device provided in some embodiments of this application; Figure 3 This is a schematic diagram of the software configuration of a terminal device provided in some embodiments of this application; Figure 4 A flowchart illustrating an audio signal quality enhancement method provided in some embodiments of this application; Figure 5 A schematic diagram of the audio quality detection process provided in some embodiments of this application; Figure 6 A flowchart illustrating the process of determining the optimal sampling point location provided in some embodiments of this application; Figure 7 A system architecture diagram of an audio signal quality enhancement method provided in some embodiments of this application; Figure 8 Timing diagrams for audio signal quality enhancement provided in some embodiments of this application; Figure 9A schematic diagram illustrating the sampling point locations provided in some embodiments of this application; Figure 10 A schematic diagram of spectrum analysis of PCM audio data when the sampling point position is at an abnormal offset position, provided in some embodiments of this application; Figure 11 This is a schematic diagram of the spectrum analysis of PCM audio data when the sampling point position is at the optimal sampling point position, provided in some embodiments of this application. Detailed Implementation

[0028] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0029] In this embodiment, the terminal device 200 generally refers to a device with data processing capabilities. For example, the terminal device 200 includes, but is not limited to, smart TVs, mobile terminals, computers, monitors, advertising screens, wearable devices, virtual reality devices, augmented reality devices, etc.

[0030] Figure 1 This is a schematic diagram illustrating an operational scenario between a terminal device and a control device provided in some embodiments of this application. For example... Figure 1 As shown, a user can operate a terminal device 200 via touch operation, a mobile terminal 300, and a control device 100. The control device 100 receives user input commands and converts them into control commands that the terminal device 200 can recognize and respond to. For example, the control device 100 can be a remote control, a stylus, a gamepad, etc.

[0031] The mobile terminal 300 can function as a control device for human-computer interaction between the user and the terminal device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the terminal device 200 and exchanging data. In some embodiments, the mobile terminal 300 can install software applications with the terminal device 200 to establish a connection and communication via network communication protocols, achieving one-to-one control operations and data communication. It can also transmit audio and video content displayed on the mobile terminal 300 to the terminal device 200 for synchronized display.

[0032] In some embodiments, the mobile terminal 300 or other electronic devices may also simulate the functions of the control device 100 by running the application of the control terminal device 200.

[0033] like Figure 1The document also shows that the terminal device 200 communicates with the server 400 via various communication methods. The terminal device 200 can communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.

[0034] Terminal device 200 can provide broadcast television reception function, and can also be equipped with intelligent network television function with computer support function, including but not limited to network television, smart television, Internet Protocol television (IPTV), etc.

[0035] Figure 2 Provided for some embodiments of this application Figure 1 Hardware configuration block diagram of terminal device 200.

[0036] In some embodiments, the terminal device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface 280.

[0037] In some embodiments, detector 230 is used to acquire signals from the external environment or to interact with the outside world. For example, detector 230 includes a light receiver, a sensor for acquiring ambient light intensity; or, detector 230 includes an image acquisition device, such as a camera, which can be used to acquire external environmental scenes, user attributes, or user interaction gestures; or, detector 230 includes a sound acquisition device, such as a microphone, for receiving external sounds.

[0038] In some embodiments, the display 260 includes display function components for presenting images and driving components for driving image display. The display 260 is used to receive and display image signals output from the controller 250. For example, the display 260 can be used to display video content, image content, menu control interface components, and user control UI interfaces, etc.

[0039] In some embodiments, the communication device 220 is a component used to communicate with external devices or the server 400 according to various communication protocol types. The terminal device 200 may have multiple communication devices 220 depending on the supported communication methods. For example, when the terminal device 200 supports wireless network communication, it may have a communication device 220 with WiFi functionality. When the terminal device 200 supports Bluetooth connection communication, it needs to have a communication device 220 with Bluetooth functionality.

[0040] The communication device 220 enables the terminal device 200 to communicate with external devices or the server 400 via wireless or wired connections. Wired connections utilize data cables, interfaces, or other components to connect the terminal device 200 to external devices. Wireless connections utilize wireless signals or wireless networks. The terminal device 200 can establish a direct connection with external devices or indirectly through gateways, routers, or other connection devices.

[0041] In some embodiments, the controller 250 may include at least one of a central processing unit, a video processor, an audio processor, a graphics processor, and a power processor, and a first to an nth interface for input / output. The controller 250 controls the operation of the terminal device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the terminal device 200.

[0042] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0043] In some embodiments, a user can input user commands through a graphical user interface (GUI) displayed on a display 260, and the user input interface 280 receives the user input commands through the graphical user interface (GUI).

[0044] In some embodiments, the audio output device 270 can be the built-in speaker of the terminal device 200 or an external audio output device connected to the terminal device 200. For the external audio output device connected to the terminal device 200, the terminal device 200 may also be provided with an external audio output terminal, through which the audio output device can be connected to the terminal device 200 to output sound from the terminal device 200.

[0045] In some embodiments, the user input interface 280 can be used to receive instructions from user input.

[0046] In order to perform user interaction, in some embodiments, the terminal device 200 may run an operating system. The operating system is a computer program used to manage and control the hardware and software resources in the terminal device 200. The operating system can control the terminal device to provide a user interface; for example, the operating system can directly control the terminal device to provide a user interface, or it can provide a user interface by running applications. The operating system also allows users to interact with the terminal device 200.

[0047] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system that is deeply customized based on a specific operating platform, or an independent operating system specifically developed for terminal devices.

[0048] An operating system can be divided into different modules or levels based on the functions it implements, for example... Figure 3 As shown, in some embodiments, the system is divided into four layers, from top to bottom: the Applications layer (referred to as the "Application Layer"), the Application Framework layer (referred to as the "Framework Layer"), the System Runtime Library layer, and the Kernel Layer.

[0049] In some embodiments, the application layer provides services and interfaces for applications, enabling the terminal device 200 to run applications and interact with the user based on the applications. The application layer may contain at least one application, which may be a built-in Windows program, system settings program, or clock program of the operating system; or it may be an application developed by a third-party developer. In specific implementations, the application packages in the application layer are not limited to the examples above.

[0050] The framework layer provides application programming interfaces (APIs) and a programming framework for applications. The application framework layer includes predefined functions. It acts as a central processing unit, determining the actions taken by applications within the application layer. Through the API, applications can access system resources and obtain system services during execution.

[0051] like Figure 3As shown, the application framework layer in this embodiment includes a view system, managers, and content providers. The view system designs and implements the application's interface and interactions, and includes lists, grids, text boxes, and buttons. The managers include at least one of the following modules: an activity manager for interacting with all running activities in the system; a location manager for providing system services or applications with access to system location services; a package manager for retrieving various information related to application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.

[0052] In some embodiments, the Activity Manager manages the lifecycle of individual applications and common navigation and back functions, such as controlling application exit, opening, and back actions. The Window Manager manages all window programs, such as obtaining the screen size, determining if a status bar is present, locking the screen, capturing the screen, and controlling changes to the display window, such as shrinking the display window, shaking the display, or distorting the display.

[0053] In some embodiments, the system runtime library layer can provide support for the framework layer. When the framework layer is used, the operating system runs the instruction library contained in the system runtime library layer, such as the C / C++ instruction library, to implement the functions to be performed by the framework layer.

[0054] In some embodiments, the kernel layer is a functional layer situated between the hardware and software of the terminal device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. For example, ... Figure 3 As shown, hardware drivers can be configured in the kernel layer. The kernel layer can contain at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.

[0055] It should be noted that the above examples are merely a simple division of operating system functions and do not limit the specific form of the operating system of the terminal device 200 in this application embodiment. Depending on the functions of the terminal device, the type of operating system, and other factors, the number of layers and the specific type of the operating system may take other forms.

[0056] Terminal device 200 can be configured with a sound acquisition device to achieve long-distance voice wake-up and voice command recognition functions. Specifically, the controller 250 in terminal device 200 (such as a system-on-chip (SOC)) can be connected via I... 2 The S (Inter-IC Sound) interface protocol communicates with the sound acquisition unit, which includes an analog-to-digital converter module. This module converts the acquired audio signal into pulse density modulation (PDM) audio data, which is then transmitted to the controller 250 via the clock (CLK) and data (DATA) lines. 2 The S-interface includes a clock signal (CLK) line and a data signal (DATA) line. The sound acquisition unit transmits a clock signal to the controller 250 via the clock signal line and transmits PDM audio data to the controller 250 via the data signal line. Driven by the clock signal, the controller 250 samples the PDM audio data on the data signal line according to preset sampling point positions. The sampled PDM audio data stream is then low-pass filtered, sampled, and quantized before being converted into Pulse Code Modulation (PCM) audio data for subsequent speech recognition processing.

[0057] The sampling point position refers to the specific moment (phase position) at which the controller 250 samples the PDM audio data on the data signal line within a single clock cycle.

[0058] Ideally, the controller 250 should sample the PDM audio data at specific edges of the clock signal (such as rising or falling edges), at which point the level on the data signal line is stable and accurately reflects the audio signal acquired by the sound acquisition unit. However, in actual operation, due to inherent delays in the hardware circuit, clock signal jitter, electromagnetic interference, component aging, and ambient temperature drift, the timing relationship between the clock signal and the data signal will shift. This timing shift causes the controller 250 to read data at the preset sampling time that is not the stable level actually output by the sound acquisition unit, but rather a random value in the data transition region, thus introducing sampling errors. After filtering and conversion, this manifests as significant high-frequency noise components in the PCM audio data, reducing the quality of the far-field audio signal and consequently degrading far-field speech recognition performance, affecting the user's voice interaction experience.

[0059] To address the aforementioned issues, the terminal device 200 provided in this embodiment can execute an audio signal quality enhancement method. By dynamically and adaptively calibrating the sampling point position of the PDM audio data, the sampling time is always adapted to the current hardware timing state of the device, thereby improving the quality of far-field acquired audio signals and enhancing far-field voice interaction performance.

[0060] The terminal device 200 capable of applying the audio signal quality enhancement method includes at least a sound acquisition unit and a controller 250. The sound acquisition unit is configured to acquire audio signals. The controller 250 is configured to execute specific steps of the audio signal quality enhancement method. Figure 4 The diagram shown is a flowchart illustrating the audio signal quality enhancement method provided in this application embodiment, which specifically includes the following steps: S401 receives PDM audio data output from the sound acquisition unit and converts the PDM audio data into PCM audio data.

[0061] During the operation of terminal device 200, the sound acquisition unit collects audio signals and transmits PDM audio data to controller 250 via the I²S interface protocol. After receiving the PDM audio data, controller 250 converts the PDM audio data into PCM audio data.

[0062] PDM audio data is an audio data format that uses a high sampling rate (typically in the MHz range) and 1 bit quantization depth, and can characterize the amplitude of an analog audio signal by the level of pulse density. PCM audio data is an audio data format that uses a fixed sampling rate (such as 16kHz or 48kHz) and multiple bits of quantization depth (such as 16 bits or 24 bits).

[0063] In some embodiments, the controller 250 converts PDM audio data into PCM audio data through low-pass filtering and sampling quantization. Specifically, after receiving the PDM audio data stream sent by the sound acquisition unit, the controller 250 first filters out high-frequency quantization noise in the PDM audio data using a low-pass filter, then samples the filtered signal according to a target sampling rate (e.g., 16kHz), and performs multi-bit quantization (e.g., 16-bit quantization) on the sampled signal to obtain PCM audio data. The converted PCM audio data can be stored in a preset buffer for subsequent audio quality detection and speech recognition processing.

[0064] The buffer can be a circular buffer, a first-in, first-out (FIFO) data structure that efficiently supports continuous data write and read operations. The audio path (such as a speech recognition engine) can read PCM audio data from the front of the buffer for speech recognition processing. Simultaneously, the audio quality detection module can also obtain PCM audio data from the buffer and independently detect and analyze the audio quality.

[0065] The conversion from PDM audio data to PCM audio data is completed by low-pass filtering and sampling quantization. The buffer is used to temporarily store and asynchronously detect the audio data, so that the audio quality detection process is completely decoupled from the main audio sampling path. The detection operation will not interfere with the normal transmission of real-time audio stream and subsequent speech recognition processing, ensuring the continuity of far-field speech function.

[0066] S402 performs audio quality detection on the PCM audio data and obtains the audio quality detection results.

[0067] After the controller 250 converts the PCM audio data, it can perform audio quality detection on the PCM audio data to determine whether the current audio signal quality is normal.

[0068] In some embodiments, the controller 250 can acquire PCM audio data from a preset buffer and then perform audio quality detection on the acquired PCM audio data. By temporarily storing the PCM audio data in the buffer, the audio quality detection process can be decoupled from the main audio sampling path, so that the detection operation will not interfere with the normal transmission of the real-time audio stream and subsequent speech recognition processing, ensuring the continuity of far-field speech function operation.

[0069] In some embodiments, the controller 250 can perform audio quality detection on the PCM audio data according to a preset time period. That is, it uses a timed polling method to periodically detect the audio signal quality, for example, triggering an audio quality detection process every 60 seconds. By using timed polling, it is possible to ensure timely detection of timing deviations while avoiding excessive computational resource consumption from continuous detection, thus balancing detection timeliness and system power consumption.

[0070] In some embodiments, such as Figure 5 The diagram shown is a flowchart of the audio quality detection process provided in an embodiment of this application, which specifically includes the following steps: S501 performs a Fast Fourier Transform (FFT) on the PCM audio data to obtain the spectrogram of the PCM audio data.

[0071] The controller 250 can perform an FFT on the PCM audio data, converting the audio signal from the time domain to the frequency domain to obtain the spectrum of the PCM audio data frame. In the spectrum, the horizontal axis represents frequency, and the vertical axis represents the energy amplitude corresponding to each frequency component.

[0072] The controller 250 selects a preset high-frequency band in the spectrum, calculates the total energy or average energy within the preset high-frequency band, and compares it with a preset energy threshold to determine whether the audio quality is abnormal.

[0073] The preset high-frequency band can be a higher frequency range outside the human voice frequency band (usually 300Hz to 3400Hz), such as the frequency band above 6kHz. Since sampling errors caused by timing offsets mainly manifest as high-frequency noise in PCM audio data, analyzing the energy level of the high-frequency band can effectively determine whether sampling errors exist. The preset energy threshold can be set based on experimental data or empirical values, for example, set as the reference energy value of a normal audio signal in the high-frequency band plus a certain tolerance margin.

[0074] Under normal operating conditions, when the sampling point is correctly positioned, there is no abnormal noise in the PCM audio data. However, due to factors such as hardware delay, clock jitter, signal interference, hardware aging, and temperature drift, the timing difference between the clock signal and the data signal may shift. When the timing shift exceeds a certain limit, it will affect the normal sampling of the controller 250, causing the bit values ​​sampled by the controller 250 to no longer be the actual output values ​​of the sound acquisition unit, but rather random 0s or 1s. This bit error caused by the sampling point shift is unpredictable and is equivalent to injecting random noise into the original bitstream. At the same time, because the sampling rate of PDM audio data is high (MHz level), the bit errors are distributed over a very short time scale, and its spectral energy is naturally concentrated in the high-frequency band. The low-pass filter used in the conversion process from PDM audio data to PCM audio data will further amplify and expose this high-frequency noise. At this time, the controller 250 will detect the PCM audio data through FFT and find that the high-frequency energy far exceeds the normal energy threshold, thereby determining the audio abnormality and triggering the sampling point position adjustment process.

[0075] S502, when the energy of a preset high-frequency band in the spectrum exceeds a preset energy threshold, the audio quality detection result is determined to be abnormal.

[0076] If the energy of the preset high-frequency band is greater than the preset energy threshold, it indicates that there is an abnormal high-frequency noise component in the current PCM audio data, and the controller 250 determines that the audio quality detection result is abnormal.

[0077] S503: When the energy of a preset high-frequency band in the spectrum is less than or equal to a preset energy threshold, the audio quality detection result is determined to be normal.

[0078] If the energy of the preset high-frequency band is less than or equal to the preset energy threshold, it indicates that the high-frequency noise level in the current PCM audio data is within the normal range, and the controller 250 determines that the audio quality detection result is normal.

[0079] In this embodiment, the audio data is converted from the time domain to the frequency domain for analysis using Fast Fourier Transform. The audio quality is determined by comparing the high-frequency energy with a preset energy threshold. This method can accurately match the signal characteristics of high-frequency noise being amplified after sampling errors caused by timing offset, thus achieving efficient and accurate detection of audio quality anomalies.

[0080] S403, when an abnormal audio quality test result is detected for a preset number of consecutive times, the sampling point position of the PDM audio data is gradually adjusted within a preset adjustment range in a preset step unit, and audio quality test is performed on the PCM audio data after each adjustment of the sampling point position, and the audio quality test result corresponding to the sampling point position is recorded.

[0081] After each audio quality detection, the controller 250 obtains an audio quality detection result, indicating whether the current audio signal quality is normal or abnormal. To avoid false triggering caused by occasional noise or transient interference, the controller 250 does not immediately trigger the sampling point position adjustment when an abnormality is detected once. Instead, it needs to detect a preset number of audio quality detection results indicating an abnormality (e.g., 3 or 5 consecutive times) before considering that there is indeed a persistent audio quality problem caused by timing offset, and then triggering the sampling point position adjustment process.

[0082] Once the continuous triggering condition is met, the controller 250 enters the sampling point position adjustment phase. Within a preset adjustment range, the controller 250 gradually adjusts the sampling point position of the PDM audio data in preset step units. After adjusting to a new sampling point position, the controller 250 samples the subsequently received PDM audio data based on that sampling point position, converts the sampled PDM audio data into PCM audio data, performs audio quality detection on the PCM audio data, and records the audio quality detection result corresponding to that sampling point position. This process is repeated, gradually traversing each candidate sampling point position within the adjustment range, to complete a comprehensive evaluation of the audio quality at each sampling point position.

[0083] In some embodiments, the preset step unit is a fraction of the clock cycle (PDM_CLK cycle) corresponding to the PDM audio data. For example, one-eighth, one-sixteenth, or one-thirty-second of the clock cycle can be used as the step unit to ensure adjustment accuracy. Using a fraction of the clock cycle as the step unit enables fine adjustment of the sampling point position, scanning each candidate sampling position within the clock cycle with sufficient resolution within a finite time.

[0084] In some embodiments, the preset adjustment range is at least half a clock cycle. Setting the adjustment range to at least half a clock cycle ensures that even if a large timing offset causes the original sampling point to fall completely into the data transition region, an effective stable sampling interval can still be found within half a clock cycle.

[0085] In some embodiments, the controller 250 can create an error counter. After each audio quality detection result is obtained, if the audio quality detection result indicates an error, the error counter is incremented; if the audio quality detection result indicates a normal result, the error counter is reset to zero. When the error counter reaches a preset number of counts, the sampling point position adjustment process is triggered.

[0086] Specifically, the controller 250 can maintain an anomaly counter locally, initially set to 0. After each audio quality detection, the anomaly counter is updated based on the detection result. When the audio quality detection result is abnormal, the anomaly counter is incremented (e.g., by 1); when the result is normal, the anomaly counter is reset to 0. This anomaly counter reset mechanism ensures that only consecutive anomalies, rather than cumulative anomalies, trigger sampling point position adjustments, effectively preventing false triggering during intermittent fluctuations in audio signal quality.

[0087] The controller 250 continuously monitors the count value of the anomaly counter. When the count value reaches a preset number (e.g., 3 or 5), indicating that an abnormal audio quality detection result has been detected for the preset number of consecutive times, the controller 250 determines that there is a persistent audio quality problem caused by timing offset, and then triggers the sampling point position adjustment process. The anomaly counter can be reset to zero for subsequent detection. By statistically confirming the number of consecutive anomalies through the anomaly counter, the sampling position adjustment is only triggered when there are a preset number of consecutive anomalies. This effectively avoids false triggering problems caused by single sporadic noise, transient electromagnetic interference, or sudden environmental noise, and improves the accuracy and stability of the adjustment and calibration mechanism.

[0088] S404 determines the optimal sampling point position from the sampling point positions indicated by the audio quality test results.

[0089] After completing the traversal of all candidate sampling point positions and audio quality detection within the entire preset adjustment range, the controller 250 can obtain the audio quality detection results corresponding to each sampling point position. The controller 250 then filters out sampling point positions where the audio quality detection results indicate normal positions, and further determines the optimal sampling point position from these normal sampling point positions.

[0090] S405 samples the PDM audio data output by the sound acquisition unit according to the optimal sampling point position.

[0091] After determining the optimal sampling point position, the controller 250 sets this optimal sampling point position as the working position for subsequent PDM audio data sampling, and samples the PDM audio data subsequently acquired by the sound acquisition unit according to this optimal sampling point position. Thereafter, the controller 250 continues to sample audio data at the optimal sampling point position and continues to periodically perform audio quality detection to continuously monitor the audio signal quality. If an abnormal audio quality detection result is detected again for a preset number of consecutive times during subsequent operation, a new round of sampling point position adjustment is triggered, thereby achieving continuous tracking and adaptive adjustment of hardware timing changes.

[0092] In some embodiments, the controller 250 is configured with a sampling phase register, which is used to control the sampling point position of the PDM audio data. Therefore, in the sampling point position adjustment process, the controller 250 can gradually adjust the register value of the sampling phase register within a preset adjustment range in preset step units, thereby adjusting the sampling point position. Correspondingly, the controller 250 samples the PDM audio data acquired by the sound acquisition device according to the optimal sampling point position, specifically by writing the register value corresponding to the optimal sampling point position into the sampling phase register.

[0093] In other words, the sampling phase register, as a hardware register inside the controller 250, controls the specific timing (i.e., sampling phase) at which the controller 250 samples the data signal lines during PDM audio data transmission. By modifying the register value of the sampling phase register, the relative position of the sampling point within the clock cycle can be adjusted, thereby changing the alignment between the sampling time and the stable data interval.

[0094] Specifically, after triggering the sampling point position adjustment process, the controller 250 first determines the preset adjustment range and preset step unit. Using the preset step unit as the step increment, starting from the initial register value (corresponding to the original sampling point position), the controller gradually modifies the register value of the sampling phase register according to the step unit. Each modification of the register value corresponds to an adjustment to a new sampling point position. After each new register value is written, the controller 250 obtains the PCM audio data down-converted at that sampling point position and performs audio quality detection, recording the audio quality detection result corresponding to that register value (i.e., the sampling point position), until all candidate sampling point positions within the entire preset adjustment range have been traversed.

[0095] After determining the optimal sampling point position, the controller 250 writes the register value corresponding to the optimal sampling point position into the sampling phase register, thereby locking the optimal sampling configuration. Subsequently, under the control of the sampling phase register, the controller 250 samples the PDM audio data on the data signal line according to the optimal sampling point position. This method of adjusting and locking the sampling point position through register read / write allows for precise control of the sampling timing directly through the controller 250's built-in hardware registers, eliminating the need for external circuit modifications or additional hardware investment. This results in low cost and ease of deployment and application in actual products.

[0096] In some embodiments, such as Figure 6 The diagram shown is a flowchart illustrating the process of determining the optimal sampling point location according to an embodiment of this application, specifically including the following steps: S601 determines the range of the maximum continuous sampling point positions within the sampling point positions indicated by the audio quality test results.

[0097] After completing the traversal detection of all candidate sampling point positions, the controller 250 can filter out all sampling point positions indicating normal audio quality detection results. Since timing offsets typically cause sampling points to fall within a certain range of data transition regions, and sampling points far from these transition regions usually yield normal audio quality detection results, there are usually one or more intervals within the adjustment range consisting of consecutive normal sampling point positions. The controller 250 identifies the largest consecutive sampling point position interval among all sampling point positions indicating normal audio quality detection results—that is, the interval containing the most consecutive normal sampling point positions.

[0098] S602, calculate the midpoint position in the maximum continuous sampling point position interval to obtain the optimal sampling point position.

[0099] After determining the maximum continuous sampling point position range, the controller 250 can determine the maximum sampling point position and the minimum sampling point position within the maximum continuous sampling point position range. Then, it calculates the average value between the maximum sampling point position and the minimum sampling point position to obtain the value of the midpoint position, which is taken as the optimal sampling point position.

[0100] For example, if the register value corresponding to the minimum sampling point position in the maximum continuous sampling point position interval is B, and the register value corresponding to the maximum sampling point position is C, then the register value corresponding to the optimal sampling point position is M = (B + C) / 2.

[0101] By selecting the range of locations with the largest continuous sampling points instead of individual sampling points, the randomness that may exist in the locations of isolated normal sampling points can be eliminated, making the selected optimal sampling point location more robust. Furthermore, taking the midpoint of the range of locations with the largest continuous sampling points as the optimal sampling point location ensures that this sampling point location has sufficient fault tolerance margin on both sides of the range. Even if slight timing drift occurs subsequently, the sampling point location can still remain within the normal range, thereby extending the stable operating time after adjustment and reducing the adjustment frequency.

[0102] In some embodiments, the sound acquisition unit configured in the terminal device 200 integrates a micro-electro-mechanical system (MEMS) sensor and a PDM modulator. The MEMS sensor is responsible for picking up audio signals from the environment and converting them into analog electrical signals. The PDM modulator is used to perform pulse density modulation on the analog electrical signals to generate single-bit PDM audio data. Then, the sound acquisition unit can transmit the PDM audio data to the controller 250 according to the I²S interface protocol via clock signal lines and data signal lines. The clock signal is used to synchronize the data transmission timing, and the data signal is used to carry the actual audio data. The controller 250 receives the PDM data through its built-in PDM interface controller. Based on the sampling point position configured in the current sampling phase register, it samples the level of the data signal line at a specific phase of the clock signal (such as a fixed offset position after the rising edge), and sequentially reads the level values ​​(0 or 1) to form a complete bit stream of PDM audio data. Subsequently, the controller 250 performs low-pass filtering and sampling quantization on the PDM audio data stream, down-converting the MHz-level high-frequency PDM audio data into PCM audio data, and storing the converted PCM audio data in an internal preset buffer for asynchronous use by subsequent speech recognition algorithms and audio quality detection modules.

[0103] For example, such as Figure 7The diagram shown is a system architecture diagram of the audio signal quality enhancement method provided in this application embodiment. In this application embodiment, the sound acquisition device is a digital microphone (MIC), which establishes communication with the SOC through an I²S interface. The MIC internally includes a clock interface, a MEMS sensor, and a PDM modulator. The clock interface is used to receive the clock (CLK) signal output by the SOC, providing a unified synchronous clock reference for signal modulation and data transmission within the MIC. The PDM modulator, in conjunction with the MEMS sensor, converts the acquired audio signal into a single-bit digital audio signal, i.e., PDM data, and outputs it to the SOC terminal through a data signal line.

[0104] The System-on-a-Chip (SoC) is integrated into the controller 250 of the terminal device 200. Internally, it includes at least a clock control unit, a PDM interface controller, a sampling phase controller, and an audio quality detection module. The clock control unit generates and outputs the CLK signal required for PDM audio data transmission. One CLK signal is output through a chip pin to the clock interface at the MIC (Microphone Interface) to provide the operating clock for the digital microphone. The other signal serves as an internal clock reference, supplying the various audio processing modules within the SoC to ensure end-to-end timing synchronization. The PDM interface controller receives PDM audio data output from the MIC, samples the levels on the data signal lines according to the currently configured sampling point positions, and transmits the sampled PDM audio data stream to subsequent processing units. The sampling phase controller receives control commands and adjusts the phase offset of the sampling clock to change the sampling timing of the PDM interface controller, achieving fine-tuning of the sampling point positions. The audio quality detection module performs low-pass filtering and sampling quantization on the sampled PDM audio data, converting it into PCM (Physical Conversion Model) audio data. It also performs spectral analysis and audio quality detection on the PCM audio data. When multiple consecutive audio quality test results indicating abnormality are detected, the module outputs adjustment control commands to the sampling phase controller, driving the sampling phase controller to perform sampling point scanning and optimization, and finally locks the optimal sampling point position and fixes the configuration.

[0105] like Figure 8 The diagram shown is a timing diagram for audio signal quality enhancement provided in an embodiment of this application. The MEMS sensor at the MIC end is responsible for picking up audio signals from the environment, converting them into analog electrical signals, and outputting them to the PDM modulator. The PDM modulator performs pulse density modulation on the analog electrical signals, converting them into single-bit PDM audio data. Simultaneously, the clock interface at the MIC end receives the CLK signal output from the clock control unit at the SOC end, providing a unified synchronous clock reference for signal modulation and data transmission within the MIC, ensuring that the output timing of the PDM audio data remains synchronized with the sampling timing at the SOC end.

[0106] The clock control unit on the SOC side generates the CLK signal required for PDM audio data transmission. One of this signal is output to the clock interface on the MIC side through the chip pin to provide the working clock for the digital microphone. The other signal serves as an internal clock reference, supplying various audio processing modules within the SOC to ensure timing consistency across the entire chain.

[0107] The PDM interface controller receives PDM audio data output from the PDM modulator at the MIC end. According to the sampling point position currently configured by the sampling phase controller, it samples the level of the data signal line at a specified phase of the CLK signal, sequentially reading the level values ​​to form a complete PDM audio data stream. Subsequently, the PDM interface controller transmits the PDM audio data to the audio quality detection module. This module performs low-pass filtering and sampling quantization processing on the PDM audio data, converting the high-frequency single-bit PDM audio data into standard PCM audio data, which is then stored in an internal buffer for subsequent quality detection and speech recognition.

[0108] The audio quality detection module uses a timer polling mechanism to read one frame of PCM audio data from the buffer at a preset time period (e.g., every 60 seconds) and perform audio quality detection. First, an FFT is performed on the PCM audio data to convert the time-domain signal into a frequency-domain spectrogram. Then, the average energy of a preset high-frequency band (e.g., above 6kHz) in the spectrogram is detected and compared with a preset energy threshold (e.g., -50 dB). If the energy of the high-frequency band is lower than the preset energy threshold, the current audio quality is considered normal, and the anomaly counter is reset to zero. If the energy of the high-frequency band is higher than the preset energy threshold, the current audio quality is considered abnormal, and the anomaly counter is incremented. Only when the anomaly counter reaches a preset number of counts is the audio anomaly confirmed as being caused by a persistent timing offset rather than transient interference, and the sampling point position adjustment process is then triggered.

[0109] During the sampling point position adjustment process, the audio quality detection module outputs an adjustment control command for the sampling point position to the sampling phase controller. Upon receiving the command, the sampling phase controller modifies the register value of the sampling phase register, adjusting the position in steps of fractions of a clock cycle (e.g., 1 / 8 of a clock cycle), at least... Within the adjustment range of one clock cycle, the phase offset (sampling point position) of the sampling clock is gradually fine-tuned, thereby driving the sampling time of the PDM interface controller to move sequentially in two directions: phase leading and phase lagging.

[0110] After each adjustment to a candidate sampling point position, the audio quality detection module performs an audio quality detection on the currently output PCM audio data and records the audio detection result corresponding to that sampling point position.

[0111] After completing the scan within the adjustment range, the audio quality detection module iterates through the audio detection results of all sampling point positions. Among all sampling point positions where the audio detection results indicate normal, it selects the interval with the largest continuous sampling point position, and calculates the midpoint position of the interval based on the maximum and minimum sampling point positions, which is then used as the optimal sampling point position.

[0112] Finally, the audio quality detection module sends the optimal sampling point position to the sampling phase controller, which then writes it into the sampling phase register to solidify the parameters. Afterward, the PDM interface controller samples the PDM audio data input from the MIC according to this optimal sampling point position.

[0113] After a single sampling point position adjustment, the audio quality detection module resets the anomaly counter and continues to perform audio quality detection of PCM audio data according to a preset cycle, forming a closed-loop working mechanism. When the device experiences timing deviations again due to factors such as hardware aging, ambient temperature drift, and clock jitter during long-term operation, resulting in continuous audio quality anomalies, the system automatically repeats the above scanning and optimization adjustment process to achieve dynamic adaptive compensation throughout the device's entire lifecycle, ensuring the audio acquisition quality of the far-field microphone.

[0114] like Figure 9 The diagram shown is a schematic representation of the sampling point locations provided in an embodiment of this application. It illustrates the timing correspondence between the CLK and DATA signals. The upper waveform represents the CLK signal, and the lower waveform represents the DATA signal. The horizontal axis represents the time axis, and the vertical axis represents the signal voltage amplitude. The DATA signal contains a level transition region and a level stabilization region within a single CLK cycle. Figure 9 Position D is located in the level transition region of the DATA signal, which is an abnormal sampling point. At this time, the SOC will read unstable random level values, generating sampling errors and introducing high-frequency noise. Positions B and C are the left and right boundaries of the stable range of the DATA signal level. The continuous area enclosed by them is an effective sampling range with sufficient timing margin. Position M is the midpoint of this effective sampling range, i.e., the optimal sampling point position. The timing margin at this point is equal to that at both boundaries of the range, and it has the strongest tolerance to timing disturbances such as clock jitter and temperature drift. Position Q corresponds to the stable level region of the DATA signal and is the effective data reading point when the sampling point is in the optimal position.

[0115] like Figure 10 The figure shown is a schematic diagram of the spectrum analysis of PCM audio data when the sampling point position is at an abnormal offset position according to the embodiment of this application. This figure illustrates the situation when the sampling point position shifts to the level transition region of the DATA signal (…). Figure 9The image shows the spectral analysis results of the PCM audio data at position D. The left side is the time-domain analysis plot, with the upper part being the spectrogram and the lower part being the time-domain energy distribution plot. The vertical axis represents frequency, and the horizontal axis represents time. The right side is the frequency-domain analysis plot obtained through FFT, where the horizontal axis represents frequency, and the vertical axis represents the signal energy amplitude in decibels. At this point, the PDM audio data obtained from SOC sampling contains a large number of random bit errors, resulting in high-frequency noise in the PCM audio data. Figure 10 As shown, the high-frequency energy above 6kHz in PCM audio data is approximately -20 dB, far exceeding the preset energy threshold of -50 dB. This results in significant audio noise, leading to a decrease in the accuracy of far-field speech recognition.

[0116] like Figure 11 The figure shown is a schematic diagram of the spectrum analysis of PCM audio data when the sampling point position is at the optimal sampling point position according to the embodiment of this application. This figure illustrates the adjustment of the sampling point position to the optimal sampling point position (…). Figure 9 The spectral analysis results of PCM audio data at position M in the image are shown. When the sampling point position is adjusted to the stable level region of the DATA signal, the SOC can accurately read the stable level value of the microphone output, the random bit errors in the PDM audio data are eliminated, and the high-frequency noise introduced by sampling is effectively suppressed. Compared with the spectrum of the abnormal sampling state, the high-frequency noise of the PCM audio data is reduced at this time, and the energy of the high-frequency band above 6kHz drops to below -75dB, which is far below the energy threshold of -50dB, so that the audio quality is restored to the normal level, thereby ensuring the accuracy of far-field speech recognition.

[0117] Based on the aforementioned terminal device 200, this application embodiment also provides an audio signal quality enhancement method, which may include the following steps: It receives pulse density modulation (PDM) audio data output from the sound acquisition unit and converts the PDM audio data into pulse code modulation (PCM) audio data.

[0118] The audio quality of the PCM audio data is tested to obtain the audio quality test results.

[0119] When an abnormal audio quality test result is detected for a preset number of consecutive times, the sampling point position of the PDM audio data is gradually adjusted within a preset adjustment range in preset step units. Audio quality is then tested on the PCM audio data after each adjustment of the sampling point position, and the audio quality test result corresponding to the sampling point position is recorded.

[0120] The optimal sampling point location is determined from the sampling point locations indicated by the audio quality test results.

[0121] The PDM audio data output by the sound acquisition unit is sampled according to the optimal sampling point location.

[0122] The same or similar parts among the various embodiments in this specification can be referred to mutually, and will not be repeated here.

[0123] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

[0124] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be obtained based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the principles and practical applications, thereby enabling those skilled in the art to better utilize the described embodiments and various different variations of embodiments suitable for specific use considerations.

Claims

1. A terminal device, characterized in that, include: Sound acquisition device; The controller is configured as follows: Receive pulse density modulation (PDM) audio data output by the sound acquisition unit, and convert the PDM audio data into pulse code modulation (PCM) audio data; The PCM audio data is subjected to audio quality detection to obtain audio quality detection results; When an abnormal audio quality detection result is detected for a preset number of consecutive times, the sampling point position of the PDM audio data is gradually adjusted within a preset adjustment range in a preset step unit, and audio quality detection is performed on the PCM audio data after each adjustment of the sampling point position, and the audio quality detection result corresponding to the sampling point position is recorded. The optimal sampling point position is determined from the sampling point positions indicated by the audio quality detection results as normal. The PDM audio data output by the sound acquisition device is sampled according to the optimal sampling point position.

2. The terminal device according to claim 1, characterized in that, The controller performs audio quality detection on the PCM audio data and obtains the audio quality detection result, specifically configured as follows: Perform a Fast Fourier Transform on the PCM audio data to obtain the spectrum of the PCM audio data; When the energy of a preset high-frequency band in the spectrum exceeds a preset energy threshold, the audio quality detection result is determined to be abnormal. When the energy of a preset high-frequency band in the spectrum is less than or equal to a preset energy threshold, the audio quality detection result is determined to be normal.

3. The terminal device according to claim 1, characterized in that, The controller determines the optimal sampling point position from the sampling point positions indicated by the audio quality detection results as normal, and is specifically configured as follows: Within the sampling point locations where the audio quality detection results indicate normal operation, determine the range of locations with the largest consecutive sampling points. Calculate the midpoint position within the maximum continuous sampling point position interval to obtain the optimal sampling point position.

4. The terminal device according to claim 3, characterized in that, The controller calculates the midpoint position within the maximum continuous sampling point position interval to obtain the optimal sampling point position, specifically configured as follows: Determine the position of the maximum sampling point and the position of the minimum sampling point within the maximum continuous sampling point position interval; The average value between the value at the maximum sampling point position and the value at the minimum sampling point position is calculated to obtain the value at the midpoint position.

5. The terminal device according to claim 1, characterized in that, The controller is also configured to: The preset step unit is a fraction of the clock cycle corresponding to the PDM audio data. The preset adjustment range is defined as at least half of the clock cycle.

6. The terminal device according to claim 1, characterized in that, The controller performs audio quality detection on the PCM audio data and obtains the audio quality detection result, which is configured as follows: The PCM audio data is subjected to audio quality testing according to a preset time period.

7. The terminal device according to claim 1, characterized in that, Within a preset adjustment range, the controller gradually adjusts the sampling point position of the PDM audio data in preset step units, specifically configured as follows: Within a preset adjustment range, the register value of the sampling phase register is gradually adjusted in preset step units; the sampling phase register is used to control the sampling point position of the PDM audio data; The controller samples the PDM audio data acquired by the sound collector according to the optimal sampling point position, specifically configured as follows: Write the register value corresponding to the optimal sampling point position into the sampling phase register.

8. The terminal device according to claim 1, characterized in that, The controller converts the PDM audio data into Pulse Code Modulation (PCM) audio data, specifically configured as follows: The PDM audio data is low-pass filtered and sampled and quantized to obtain the PCM audio data, and the PCM audio data is stored in a preset buffer. The controller performs audio quality detection on the PCM audio data and obtains the audio quality detection result, specifically configured as follows: The PCM audio data is obtained from the buffer, and the audio quality of the PCM audio data is detected.

9. The terminal device according to claim 1, characterized in that, The controller is also configured to: Create an exception counter; After each audio quality detection result is obtained, if the audio quality detection result indicates an abnormality, the abnormality counter is incremented; if the audio quality detection result indicates a normality, the abnormality counter is cleared. When the count value of the abnormality counter reaches the preset number of times, the sampling point position of the PDM audio data is gradually adjusted within the preset adjustment range in preset step units.

10. A method for enhancing the quality of an audio signal, characterized in that, include: Receive pulse density modulation (PDM) audio data output from the sound acquisition unit, and convert the PDM audio data into pulse code modulation (PCM) audio data; The PCM audio data is subjected to audio quality detection to obtain audio quality detection results; When an abnormal audio quality detection result is detected for a preset number of consecutive times, the sampling point position of the PDM audio data is gradually adjusted within a preset adjustment range in a preset step unit, and audio quality detection is performed on the PCM audio data after each adjustment of the sampling point position, and the audio quality detection result corresponding to the sampling point position is recorded. The optimal sampling point position is determined from the sampling point positions indicated by the audio quality detection results as normal. The PDM audio data output by the sound acquisition device is sampled according to the optimal sampling point position.