Voice playing method and device
By dynamically adjusting the VAD output confidence and adaptive window of the voice acquisition channel between the mobile terminal and the translation headset, combined with a flexible switching mechanism, the problem of inconvenient audio path switching is solved, thus improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN TIMEKETTLE TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-01
AI Technical Summary
The inconvenience of switching audio paths in existing technologies leads to a poor user experience, especially since they cannot intelligently identify and switch audio input and output devices in different scenarios, resulting in interruptions in information transmission.
By dynamically adjusting the VAD output confidence of the voice acquisition channel between the mobile terminal and the translation headset, and combining short-time energy and zero crossover rate, the window of the audio frame is adaptively adjusted to achieve intelligent switching and smooth transition of the audio channel. Flexible switching mechanisms such as gain adjustment and buffer overlap, priority arbitration, etc. are used to ensure the continuity of the audio path.
It enables dynamic adaptation and automatic management of audio paths, improving user experience and solving the problems of inconvenient audio path switching and discontinuous interaction.
Smart Images

Figure CN121967963A_ABST
Abstract
Description
Voice playback method and device Technical Field
[0001] Embodiments of this application relate to the field of data processing, and more particularly to voice playback methods, apparatus, devices, and computer-readable storage media. Background Technology
[0002] With the widespread use of wireless headphones and smartphones in voice interaction, more and more application scenarios are beginning to adopt real-time speech-to-text (TTS) and text-to-speech (TTS) methods to provide users with auxiliary communication or language translation services. To adapt to the needs of different scenarios, some systems will distinguish between "TTS mode" and "speaker mode," that is, switch the audio input and output paths between the headphones and the mobile phone as needed.
[0003] Currently, the common method is to manually switch audio input / output devices. For example, users need to manually set headphones or mobile phones as the audio output device, or select the input source through system settings. These methods are not only cumbersome but also fail to intelligently switch according to changes in the actual scenario. For instance, when a user brings their phone close to a speaker, the preferred mode should be phone-based audio pickup and headphone-based audio playback; conversely, when the user is wearing headphones and away from the phone, the preferred mode should be headphone-based audio pickup and phone-based audio playback. However, existing systems often fail to recognize these scenario differences, resulting in a poor user experience and even interrupted information transmission. Summary of the Invention
[0004] According to embodiments of this application, a voice playback solution is provided that can solve the problems of inconvenient audio path switching and discontinuous interaction in the prior art, and greatly improve the user experience.
[0005] In a first aspect of this application, a voice playback method is provided. Applicable to a mobile terminal connected to a translation headset, the mobile terminal being configured with a first voice acquisition channel and a first voice playback channel, and the translation headset being configured with a second voice acquisition channel and a second voice playback channel, the method comprising: monitoring a first original voice signal input from the first voice acquisition channel and a second original voice signal input from the second voice acquisition channel; if the first original voice signal is detected, then operating a first voice playback mode and outputting a first translated voice signal corresponding to the first original voice signal from the second voice playback channel; if the second original voice signal is detected, then operating a second voice playback mode and outputting a second translated voice signal corresponding to the second original voice signal from the first voice playback channel.
[0006] Furthermore, the monitoring of the first raw voice signal input from the first voice acquisition channel and the second raw voice signal input from the second voice acquisition channel includes: monitoring the first voice acquisition channel and the second voice acquisition channel respectively through VAD detection logic to determine the VAD output confidence of the first voice acquisition channel and the second voice acquisition channel; and based on the VAD output confidence, determining whether the first voice playback channel and / or the second voice playback channel have a corresponding raw voice signal.
[0007] Further, determining whether the first and / or second voice playback channels have corresponding original voice signals based on the VAD output confidence includes: acquiring the short-time energy and zero crossover rate of the current audio frame; dynamically adjusting the adaptive window of the VAD output confidence based on the short-time energy and zero crossover rate of the audio frame; acquiring the VAD output confidence of a preset number of consecutive audio frames within the range of the adaptive window of the VAD output confidence; if the VAD output confidence of the preset number of consecutive audio frames is greater than a preset threshold and the zero crossover rate is lower than the current noise background threshold, then it is determined that the current second voice playback channel has corresponding original voice signals; otherwise, it is determined that the current first voice playback channel has corresponding original voice signals.
[0008] Furthermore, the step of outputting the second translated speech signal corresponding to the second original speech signal from the first speech playback channel includes: sequentially recognizing, translating, and synthesizing the second original speech signal, and then playing it through the first speech playback channel.
[0009] Furthermore, the step of outputting a second translated voice signal corresponding to the second original voice signal from the first voice playback channel further includes: after playback is completed, switching to the original monitoring mode through a preset flexible switching mechanism.
[0010] Furthermore, the preset flexible switching mechanism includes: gain adjustment, buffer overlap, and / or priority arbitration switching methods.
[0011] In a second aspect of this application, a voice playback device is provided. The device includes: a monitoring module for monitoring a first original voice signal input from a first voice acquisition channel and a second original voice signal input from a second voice acquisition channel; and a switching module for, if the first original voice signal is detected, operating a first voice playback mode and outputting a first translated voice signal corresponding to the first original voice signal from the second voice playback channel; and if the second original voice signal is detected, operating a second voice playback mode and outputting a second translated voice signal corresponding to the second original voice signal from the first voice playback channel.
[0012] In a third aspect of this application, an electronic device is provided. The electronic device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described above.
[0013] In a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the method according to the first aspect of this application.
[0014] The voice playback method provided in this application listens to a first original voice signal input from the first voice acquisition channel and / or a second original voice signal input from the second voice acquisition channel. If the first original voice signal is detected, a first voice playback mode is run, and a first translated voice signal corresponding to the first original voice signal is output from the second voice playback channel. If the second original voice signal is detected, a second voice playback mode is run, and a second translated voice signal corresponding to the second original voice signal is output from the first voice playback channel. This achieves automatic management of dynamic audio channel adaptation and smooth transition, solves the problems of inconvenient audio path switching and discontinuous interaction in the prior art, and greatly improves the user experience.
[0015] It should be understood that the description in the Summary Section is not intended to limit the key or essential features of the embodiments of this application, nor is it intended to restrict the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0016] The above and other features, advantages, and aspects of the embodiments of this application will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein: FIG1 is a flowchart of a voice playback method according to an embodiment of this application; FIG2 is a block diagram of a voice playback device according to an embodiment of this application; FIG3 is a structural schematic diagram of a terminal device or server suitable for implementing the embodiments of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0018] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0019] Figure 1 shows a flowchart of a voice playback method according to an embodiment of the present disclosure. The method implemented in this disclosure can be applied to a mobile terminal connected to a translation headset. The mobile terminal is configured with a first voice acquisition channel and a first voice playback channel, and the translation headset is configured with a second voice acquisition channel and a second voice playback channel. The method includes: S110, monitoring a first raw voice signal input from the first voice acquisition channel and a second raw voice signal input from the second voice acquisition channel.
[0020] In some embodiments, after a user launches the translation application on a terminal device (such as a mobile phone), the system initializes audio input / output control and begins executing the audio channel management logic of this disclosure. This is applicable to scenarios where users use headphones (classic Bluetooth module BT or Bluetooth Low Energy module BLE) for voice translation.
[0021] In some embodiments, two audio acquisition channels are activated simultaneously: a mobile phone microphone acquisition channel, i.e., the first voice acquisition channel, used to acquire voice signals from the user's surrounding environment or the target user; and an earphone microphone acquisition channel, i.e., the second voice acquisition channel, typically a single earphone microphone, used to detect whether the person wearing the earphone actively speaks.
[0022] Furthermore, by executing VAD (Voice Activation Detection) logic separately, it is determined whether there is a voice signal in each microphone acquisition channel.
[0023] Specifically, the VAD detection logic is used to monitor the first voice acquisition channel and the second voice acquisition channel respectively to determine the VAD output confidence of the first voice acquisition channel and the second voice acquisition channel.
[0024] The system obtains the short-time energy and zero-crossing rate of the current audio frame calculated in real time, and dynamically adjusts the adaptive window of the VAD output confidence (range 200ms–600ms, instead of a fixed 300ms): within the adaptive window of the VAD output confidence, the system obtains the VAD output confidence of a consecutive preset number of audio frames (which can be preset according to the application scenario); if the VAD output confidence of the consecutive preset number of audio frames is greater than a preset threshold and the zero-crossing rate is lower than the current noise background threshold, it is determined that the current second voice playback channel has a corresponding original voice signal; otherwise, it is determined that the current first voice playback channel has a corresponding original voice signal.
[0025] For example, when the confidence level is higher than the threshold for N consecutive frames (e.g., 3–5 frames), it is determined that the user has spoken voluntarily; the value of N can be dynamically adjusted according to the scene modality; the value is larger in the meeting environment (5) and smaller in the quiet environment (3).
[0026] Unlike existing fixed time windows, this disclosure adopts an adaptive window adjustment method based on acoustic features and VAD output confidence: if the energy rises continuously and the zero crossover rate is lower than the noise range, the window can be shortened to 200–300ms; if the ambient noise increases, the window can be extended to 400–600ms; if the VAD confidence rises rapidly, N (number of consecutive frames) automatically decreases; if the environment is complex, N automatically increases to reduce misjudgments.
[0027] Here are some examples: In a conference room setting (with stable but continuous background noise): Window = 550ms; N = 5; effectively distinguishes between speech and keyboard sounds; In a quiet office: Window = 250ms; N = 3; more sensitive switching and reduced latency.
[0028] Furthermore, the consecutive N-frame determination logic is as follows: "Active speaking from the headset" is determined when the following conditions are met: within the adaptive window, the confidence score of the VAD output for N consecutive frames is >0.7, and the E_weight shows an upward trend; the zero crossover rate (ZCR) is lower than the noise background threshold; and the VAD state is inconsistent with the mobile phone microphone (no obvious sound from the mobile phone channel). After meeting the above conditions, the system considers the user to be speaking through the headset and enters speaking mode. That is, it determines that the current second voice playback channel contains a corresponding original voice signal.
[0029] Furthermore, to address the issue of significant fluctuations in the traditional zero-crossing rate under noisy conditions, the present disclosure provides the following stabilization measures: only counting the number of crosses where the amplitude significantly crosses the noise baseline; and dynamically adjusting the baseline threshold through background noise self-learning to better distinguish between "voice initiation" and "slight friction noise".
[0030] In some embodiments, unlike the traditional "sum of squares of energy" formula, this disclosure incorporates an "energy trend compensation factor" when calculating energy, taking into account the close proximity of the headphone microphone to the mouth (sound emission point): that is, calculating the average energy E_base for each frame; calculating the energy difference E_diff between adjacent frames; and calculating the weighted energy: E_weight = E_base + k * E_diff; where k is an adaptive coefficient; the transient changes at the moment the user begins to speak can be captured more quickly through the above method. S120, if the first original speech signal is detected, a first speech playback mode is run, and a first translated speech signal corresponding to the first original speech signal is output from the second speech playback channel; if the second original speech signal is detected, a second speech playback mode is run, and a second translated speech signal corresponding to the second original speech signal is output from the first speech playback channel.
[0031] In some embodiments, if the second original voice signal is detected, a second voice playback mode is activated, and a second translated voice signal corresponding to the second original voice signal is output from the first voice playback channel. That is, the second original voice signal is sequentially recognized, translated, and synthesized, and then played through the first voice playback channel (phone speaker).
[0032] Furthermore, it also includes switching to the original monitoring mode through a preset flexible switching mechanism after the first voice playback channel is completed; wherein, the flexible switching mechanism includes switching methods such as gain adjustment, buffer overlap, and / or priority arbitration.
[0033] For example: the mobile phone microphone is muted briefly (e.g., 100ms) before resuming recording to eliminate tail noise interference; when the headphone resumes broadcasting, a fade-in method is used to start playback; preferably, a buffer overlap of 50–100ms can be maintained to avoid voice interruption; a dynamic cooling timer is maintained to adjust the cooling time in real time according to the user's speaking interval and the ambient noise level to avoid misjudgment or stuttering caused by fixed cooling parameters.
[0034] The cooling-off time is constructed as follows: a cooling timer records the time interval starting from "exiting speaking mode". The default cooling interval includes: T_base: base cooling-off time; T_noise: dynamically increased based on ambient noise estimation; T_gap: dynamically increased or decreased based on the user's historical speaking interval; final cooling-off time: T_cool = T_base + T_noise + T_gap. Examples are provided below: Quiet environment: T_noise ≈ 0, therefore T_cool can be as low as 300ms. Noisy scenario: T_noise increases → T_cool ≈ 800–1200ms. Furthermore, the cooling logic and switching jitter suppression include: during the cooling-off period, the system still executes dual-path VAD while disabling switching, only used for "pre-judging trends".
[0035] If the energy trend at the headphone end is detected to continue rising beyond a threshold during the cooling process, the cooling will be terminated early to improve the response speed.
[0036] According to the embodiments of this disclosure, the following technical effects are achieved: the problems of inconvenient audio path switching and discontinuous interaction in the prior art are solved, and the user experience is greatly improved.
[0037] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0038] The above is an introduction to the method embodiments. The following describes the solution described in this application through device embodiments.
[0039] Figure 2 shows a block diagram 200 of a voice playback device according to an embodiment of this application. As shown in Figure 2, it includes: a monitoring module 210, used to monitor a first original voice signal input from the first voice acquisition channel and a second original voice signal input from the second voice acquisition channel; and a switching module 220, used to, if the first original voice signal is detected, run a first voice playback mode and output a first translated voice signal corresponding to the first original voice signal from the second voice playback channel; and if the second original voice signal is detected, run a second voice playback mode and output a second translated voice signal corresponding to the second original voice signal from the first voice playback channel. Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the described modules can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0040] Figure 3 shows a schematic diagram of a terminal device or server suitable for implementing the embodiments of this application.
[0041] As shown in Figure 3, the terminal device or server includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 302 or programs loaded from storage section 308 into random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the terminal device or server. The CPU 301, ROM 302, and RAM 303 are interconnected via bus 304. An input / output (I / O) interface 305 is also connected to bus 304.
[0042] The following components are connected to I / O interface 305: an input section 306 including a keyboard, mouse, etc.; an output section 307 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN card, modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to I / O interface 305 as needed. A removable medium 311, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 310 as needed so that computer programs read from it can be installed into storage section 308 as needed.
[0043] Specifically, according to embodiments of this application, the above method flow steps can be implemented as a computer software program. For example, embodiments of this application include a computer program product comprising a computer program carried on a machine-readable medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 309, and / or installed from removable medium 311. When the computer program is executed by central processing unit (CPU) 301, it performs the functions defined in the system of this application.
[0044] It should be noted that the computer-readable medium shown in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0045] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0046] The units or modules described in the embodiments of this application can be implemented in software or hardware. The described units or modules can also be located in a processor. The names of these units or modules do not, in certain circumstances, constitute a limitation on the unit or module itself.
[0047] In another aspect, this application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable storage medium stores one or more programs that, when used by one or more processors, execute the methods described in this application.
[0048] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the foregoing application concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions claimed in this application.
Claims
1. A voice playback method, characterized in that, An application is made in a mobile terminal connected to a translation headset. The mobile terminal is configured with a first voice acquisition channel and a first voice playback channel, and the translation headset is configured with a second voice acquisition channel and a second voice playback channel. The method includes: monitoring a first original voice signal input from the first voice acquisition channel and a second original voice signal input from the second voice acquisition channel; if the first original voice signal is detected, then a first voice playback mode is activated, and a first translated voice signal corresponding to the first original voice signal is output from the second voice playback channel; if the second original voice signal is detected, then a second voice playback mode is activated, and a second translated voice signal corresponding to the second original voice signal is output from the first voice playback channel.
2. The method according to claim 1, characterized in that, The monitoring of the first raw voice signal input from the first voice acquisition channel and the second raw voice signal input from the second voice acquisition channel includes: monitoring the first voice acquisition channel and the second voice acquisition channel respectively through VAD detection logic to determine the VAD output confidence of the first voice acquisition channel and the second voice acquisition channel; and determining whether the first voice playback channel and the second voice playback channel have corresponding raw voice signals based on the VAD output confidence.
3. The method according to claim 2, characterized in that, The step of determining whether the first voice playback channel and / or the second voice playback channel have a corresponding original voice signal based on the VAD output confidence score includes: acquiring the short-time energy and zero crossover rate of the current audio frame; dynamically adjusting the adaptive window of the VAD output confidence score based on the short-time energy and zero crossover rate of the audio frame; acquiring the VAD output confidence scores of a preset number of consecutive audio frames within the adaptive window of the VAD output confidence score; if the VAD output confidence scores of the preset number of consecutive audio frames are greater than a preset threshold and the zero crossover rate is lower than the current noise background threshold, then it is determined that the current second voice playback channel has a corresponding original voice signal; otherwise, it is determined that the current first voice playback channel has a corresponding original voice signal.
4. The method according to claim 1, characterized in that, The step of outputting the second translated speech signal corresponding to the second original speech signal from the first speech playback channel includes: sequentially recognizing, translating, and synthesizing the second original speech signal, and then playing it through the first speech playback channel.
5. The method according to claim 1, characterized in that, The process of outputting a second translated voice signal corresponding to the second original voice signal from the first voice playback channel further includes: after playback is completed, switching back to the original monitoring mode through a preset flexible switching mechanism.
6. The method according to claim 5, characterized in that, The preset flexible switching mechanism includes: gain adjustment, buffer overlap, and / or priority arbitration switching methods.
7. A voice playback device, characterized in that, include: The monitoring module is used to monitor the first original voice signal input from the first voice acquisition channel, and / or the second original voice signal input from the second voice acquisition channel; the switching module is used to run the first voice playback mode and output the first translated voice signal corresponding to the first original voice signal from the second voice playback channel if the first original voice signal is detected. If the second original voice signal is detected, the second voice playback mode is activated, and the second translated voice signal corresponding to the second original voice signal is output from the first voice playback channel.
8. The apparatus according to claim 7, characterized in that, The monitoring of the first raw voice signal input from the first voice acquisition channel and the second raw voice signal input from the second voice acquisition channel includes: monitoring the first voice acquisition channel and the second voice acquisition channel respectively through VAD detection logic to determine the VAD output confidence of the first voice acquisition channel and the second voice acquisition channel; and determining whether the first voice playback channel and the second voice playback channel have corresponding raw voice signals based on the VAD output confidence.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.