Voice processing method, electronic device, and computer program product

By acquiring and recognizing the voices of multiple Bluetooth audio devices, translating and playing them, the problem of headphones being unable to uniformly process multiple voice translations is solved, enabling fast translation and playback and improving the user experience.

WO2026098705A1PCT designated stage Publication Date: 2026-05-15ZTE CORP
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ZTE CORP
Filing Date
2025-11-11
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing headphones only display the translated content on the interface after recording audio, without supporting the playback of the translated audio, and cannot uniformly process multiple voice translations, resulting in an inconvenient user experience.

Method used

By acquiring voice data from multiple Bluetooth audio devices, identifying language types, and translating them, the system enables rapid many-to-one or many-to-many translation. It also plays the target voice data according to a preset method, supporting unified translation and playback of multiple voice data.

Benefits of technology

It enables rapid translation and playback of translated audio in multiple audio scenarios, improving user convenience, especially in scenarios such as meetings, promotions, and training, where the speaker can quickly understand the content of multiple speakers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025134073_15052026_PF_FP_ABST
    Figure CN2025134073_15052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a voice processing method, an electronic device, a computer program product, and a computer-readable storage medium. The voice processing method comprises: acquiring voices collected from multiple Bluetooth audio devices, to obtain multiple different voices, the voice of each Bluetooth audio device corresponding to one language type; translating the multiple different voices on the basis of the language type of each voice and a target language type, to obtain corresponding target voices for each; and playing back the target voices on the basis of a preset playback mode.
Need to check novelty before this filing date? Find Prior Art

Description

Speech processing methods, electronic devices and computer program products

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese patent application CN 202411603685.9, filed on November 11, 2024, entitled “A speech processing method, electronic device and program product”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of wireless communication, and in particular to a voice processing method, electronic device, and computer program product. Background Technology

[0004] Current headphones typically translate audio via a corresponding app after recording it, displaying the translated content only on the interface and not playing the translated audio through the headphones. Furthermore, the current headphone's recording mechanism is not ideal for collecting multiple voice recordings for unified translation. Additionally, current voice translation applications only provide the translated content and do not support playback of the translated audio, causing inconvenience for users. Summary of the Invention

[0005] This disclosure provides a voice processing method, an electronic device, and a program product.

[0006] This disclosure provides a voice processing method, including: acquiring voice data collected from multiple Bluetooth audio devices to obtain multiple different voice data, wherein the voice data from each Bluetooth audio device corresponds to a language type; translating the multiple different voice data based on the language type of each voice data and a target language type to obtain corresponding target voice data; and playing the target voice data according to a preset playback method.

[0007] This disclosure also provides an electronic device, including: one or more processors; and a memory storing one or more programs thereon, which, when executed by the one or more processors, cause the one or more processors to implement a voice processing method according to an embodiment of this disclosure.

[0008] This disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, causes the processor to implement the speech processing method according to embodiments of this disclosure. Attached Figure Description

[0009] In the accompanying drawings of the embodiments disclosed herein:

[0010] Figure 1 is a schematic flowchart of the speech processing method provided in an embodiment of this disclosure;

[0011] Figure 2 is a schematic diagram of audio communication between the speaker and the listener via Bluetooth in a meeting scenario provided by an embodiment of this disclosure;

[0012] Figure 3 is a schematic diagram of the parallel translation process provided in the embodiments of this disclosure;

[0013] Figure 4 is a schematic diagram of the serial translation process provided in the embodiments of this disclosure;

[0014] Figure 5 is a schematic diagram of the first target information display method provided in the embodiments of this disclosure;

[0015] Figure 6 is a schematic diagram of the second target information display method provided in the embodiments of this disclosure;

[0016] Figure 7 is a schematic block diagram of the electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0017] To enable those skilled in the art to better understand the technical solutions of this disclosure, the communication-sensing data processing method and computer-readable storage medium provided in the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.

[0018] The present disclosure will be described more fully below with reference to the accompanying drawings; however, the embodiments shown may be embodied in different forms, and the present disclosure should not be construed as limited to the embodiments set forth below. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will enable those skilled in the art to fully understand the scope of the disclosure.

[0019] The accompanying drawings of the embodiments disclosed herein are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the detailed embodiments to explain this disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the description of the detailed embodiments with reference to the accompanying drawings.

[0020] This disclosure may be described with reference to plan and / or cross-sectional views using the ideal schematic diagrams of this disclosure. Therefore, the example illustrations may be modified according to manufacturing techniques and / or tolerances.

[0021] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0022] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. The term "and / or" as used in this disclosure includes any and all combinations of one or more of the associated enumerated entries. The singular forms "a" and "the" as used in this disclosure are also intended to include the plural forms, unless the context clearly indicates otherwise. The terms "comprising," "made of," etc., as used in this disclosure specify the presence of the stated feature, integral, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof.

[0023] Unless otherwise specified, all terms used in this disclosure (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so specified in this disclosure.

[0024] Current headphones typically translate audio via a corresponding app after recording it, displaying the translated content only on the interface and not playing the translated audio through the headphones. Furthermore, the current headphone's recording mechanism is not well-suited for collecting multiple voice recordings for simultaneous translation. Additionally, current voice translation applications (e.g., including but not limited to meetings, promotional events, training sessions, and educational settings) only provide the translated content and do not support playback of the translated audio, hindering users' ability to quickly understand what the other person is saying.

[0025] According to the scheme of this disclosure embodiment, voice data collected from multiple Bluetooth audio devices is acquired, resulting in multiple different voices, wherein the voice data from each Bluetooth audio device corresponds to a language type. Based on the language type of each voice and the target language type, the multiple different voices are translated to obtain corresponding target voices, enabling the same terminal to uniformly translate multiple voices to be translated. This achieves fast many-to-one translation in scenarios with multiple voice senders and one voice receiver, and is also applicable to many-to-many translation scenarios. The target voice is played according to a preset playback method, realizing the playback of the translated voice, avoiding the need for the user to view the translated content after the voice translation is completed, allowing the user to quickly hear the translated content, achieving fast and unified translation of different languages, and providing convenience for the user.

[0026] The solutions disclosed herein can be applied to any terminal with voice translation functionality and to any voice translation application scenario, including but not limited to conference scenarios (e.g., international conferences, ethnic conferences, etc.), publicity scenarios, training scenarios, and teaching scenarios. The terminal can be used by the speaker in the aforementioned scenarios, enabling the speaker to conveniently and quickly listen to the voice information of multiple speakers. The terminal may have a Bluetooth module, a translation module, a display module, and a playback module, and may include, but is not limited to, mobile phones, smart wearable devices (e.g., headphones, glasses, wristbands), in-vehicle devices, and teaching equipment.

[0027] The embodiments of this disclosure will be described in detail below.

[0028] This disclosure provides a speech processing method, as shown in FIG1, which includes steps S11 to S13.

[0029] In step S11, voice data collected from multiple Bluetooth audio devices is acquired, resulting in multiple different voice data, where the voice data from each Bluetooth audio device corresponds to a language type.

[0030] In this embodiment of the disclosure, before acquiring the voice collected by multiple Bluetooth audio devices (i.e., step S11), the method may further include: connecting different Bluetooth audio devices to the terminal's Bluetooth Isochronous Streams (BIS) group via the Enhanced Attribute Protocol (ATT).

[0031] In this embodiment of the disclosure, a meeting scenario is used as an example to illustrate the solution of this embodiment. A meeting scenario may include a speaker and numerous listeners. During the meeting, at least one listener will ask the speaker a question. If the speaker and at least one listener speak different languages ​​(e.g., languages ​​of different countries, languages ​​of different ethnic groups, etc.), the solution of this embodiment can be applied to translate the listener's speech on the speaker's terminal. The solution of this embodiment can be applied to the speaker's terminal, which may include, but is not limited to, headphones.

[0032] In this embodiment of the disclosure, as shown in Figure 2, the speaker and listeners in a conference scenario can communicate via Bluetooth. The speaker's terminal (e.g., headphones) can be equipped with a Bluetooth module (e.g., a Bluetooth Low Energy (BLE) audio device), and the listeners can be equipped with Bluetooth audio devices (e.g., BLE audio devices). Any listener's BLE audio device can be added to the speaker's terminal device's Bluetooth BIS group at any time via the Enhanced Attribute Protocol. Different listeners can correspond to different Bluetooth audio devices, and different voices are voices generated by different listeners.

[0033] In this embodiment of the disclosure, each listener's Bluetooth audio device can correspond to a unique identifier, which can be a device identifier or any other suitable identifier, such as, but not limited to, a serial number, a number, a seat number, or a listener's nickname.

[0034] In this embodiment, the speaker's terminal includes a Bluetooth module that can connect to the BLE audio devices of multiple listeners. When multiple listeners ask questions together, the speaker's Bluetooth module can individually acquire the speaker's voice from each speaker's Bluetooth audio device. Each speaker's voice can be collected via the microphone (e.g., a microphone) of their respective BLE audio device. The duration of each speaker's question can be limited to a certain time; therefore, the acquisition duration of the BLE audio devices can be set, for example, 30-40 seconds per person. The speaker's voice can be output to the speaker's terminal's translation module based on the acquisition duration. In other exemplary embodiments, the question duration may be unrestricted.

[0035] In this embodiment of the disclosure, after acquiring the voice collected from multiple Bluetooth audio devices (i.e., step S11), the method may further include: segmenting each voice with a duration greater than or equal to a preset duration threshold according to the preset duration threshold, so that the duration of each voice obtained after segmentation is less than the duration threshold.

[0036] In this embodiment of the disclosure, the acquisition duration may not be set in the listener's Bluetooth audio device. For long audio recordings, they can be segmented in the terminal before translation.

[0037] In this embodiment of the disclosure, each speaker's voice stream corresponds to a corresponding language type, and each speaker's voice stream is associated with the identifier of that speaker's Bluetooth audio device, in order to distinguish the voices of different speakers. For the voice obtained after segmentation, the same identifier is associated with the corresponding long voice (i.e., the voice before segmentation whose duration is greater than or equal to the duration threshold).

[0038] In this embodiment of the disclosure, when the listener's Bluetooth audio device is added to the Bluetooth BIS group, the terminal device can record the identifier corresponding to each Bluetooth audio device. After the Bluetooth module in the speaker's terminal receives the voice stream sent by the Bluetooth audio device, the Bluetooth module can notify the terminal of the correspondence between each identifier and the Bluetooth audio device and the voice stream. For example, it can notify the translation module in the terminal.

[0039] In this embodiment, the translation module may include, but is not limited to, a preset translation application, a translation mini-program, or a translation chip. When the translation module is a translation application, the application can add each identifier and the language type of the corresponding speech to the application interface.

[0040] In step S12, multiple different speech samples are translated based on the language type of each speech sample and the target language type to obtain the corresponding target speech samples.

[0041] In this embodiment of the disclosure, based on the language type of each speech, multiple different speech can be translated in parallel or sequentially by a translation module to obtain the target speech, and the language type of the target speech is the language type used by the terminal.

[0042] In this embodiment of the disclosure, the parallel translation scheme can be as shown in Figure 3, which translates multiple different speech based on the language type of each speech and the target language type to obtain the corresponding target speech (i.e., step S12) including the following steps S21 to S23.

[0043] In step S21, a target number of threads are started based on the number of different voices.

[0044] In step S22, a separate thread is used for timbre recognition and speech translation for each speech stream.

[0045] In step S23, the speech obtained after speech translation is combined with the corresponding timbre to obtain the target speech corresponding to each speech.

[0046] In this embodiment, multiple different voices can be transmitted to the translation module via a digital signal processor (DSP). The translation module activates a target number of threads based on the number of different voices, performs voice translation on the received voice streams based on the multiple threads, and performs timbre recognition on each voice stream to obtain the speaker's voice with the corresponding timbre as the target voice. Based on timbre recognition, the target voice can have the original timbre of the speaker.

[0047] In this embodiment of the disclosure, the serial translation scheme can be as shown in FIG4, which translates multiple different speech based on the language type of each speech and the target language type to obtain the corresponding target speech respectively (i.e., step S12) including the following steps S31 to S34.

[0048] In step S31, multiple different voices are mixed to obtain a voice stream.

[0049] In step S32, the speech stream is split according to the waveform of the speech stream to obtain multiple sub-speech streams.

[0050] In step S33, each of the split sub-speech streams is translated sequentially, and the timbre of each sub-speech stream is identified.

[0051] In step S34, the speech obtained after speech translation of each sub-speech stream is combined with the corresponding timbre to obtain the target speech corresponding to each speech.

[0052] In this embodiment, multiple different speech streams can be mixed by a DSP and synthesized into a single speech stream. The DSP then splits this speech stream according to the waveform to obtain multiple sub-speech streams. These sub-speech streams are concatenated into a long audio stream and sent to a translation application for individual speech translation. The translation application also performs timbre recognition to identify the speaker's voice with the corresponding timbre as the target speech.

[0053] In this embodiment of the disclosure, the timbre recognition scheme can be implemented according to actual needs. When no timbre needs to be added, the target speech can be generated according to the preset timbre.

[0054] In this embodiment of the disclosure, if the translation application interface stores the identifiers corresponding to the Bluetooth audio devices of each speaker, the translated target speech can be stored under the corresponding identifier in the translation application interface, for example, under the corresponding serial number.

[0055] In this embodiment of the disclosure, the method may further include: acquiring target information for each target speech and outputting and displaying it, wherein the target information may include, but is not limited to: the identifier and / or translation content corresponding to the target speech.

[0056] In this embodiment of the disclosure, the target information can be displayed on the terminal interface or the interface of the external device connected to the terminal.

[0057] In this embodiment of the disclosure, the text corresponding to the translated content (the text corresponding to the language type used by the speaker) and / or the identifier corresponding to the translated content can be displayed on the terminal interface or the interface of the external device connected to the terminal, so that the speaker can view the questions asked by the listeners at any time.

[0058] In this embodiment of the disclosure, the external devices connected to the terminal may include, but are not limited to, mobile phones, tablets, smart wearable devices, etc.

[0059] In this embodiment of the disclosure, the display method of the target speech information on the interface can be selected according to needs, and no detailed limitation is made here. For example, as shown in Figure 5, the identifiers of each listener and the translation content can be displayed in the same way as the voice chat interface of social media (e.g., including but not limited to WeChat). As shown in Figure 6, the translation information can be displayed on the interface in order of completion time.

[0060] Figure 5 illustrates a voice bar-based interface that displays translated content for each identifier (e.g., seat numbers 105, 211, 415, and 407) by prompting a voice message. For example, the question for seat number 105 is "What is the specific implementation plan?", for seat number 211 it is "What is the specific implementation time?", for seat number 415 it is "Who will be involved?", and for seat number 407 it is "Who will benefit?".

[0061] Figure 6 shows a time-series-based interface that can display the translation content corresponding to each identifier (e.g., seat number 105, 211, 415, 407). For example, the translation completion times for seats 105, 211, 415, and 407 are 14:35:21, 14:35:22, 14:35:37, and 14:36:07 on October 31, 2024, respectively. Based on this time sequence, the translation of the question "What is the specific implementation plan?" completed by seat 105 at 14:35:21, "What is the specific implementation time?" completed by seat 211 at 14:35:22, "What is the specific implementation time?" completed by seat 415 at 14:35:37, and "Who will be involved?" completed by seat 407 at 14:35:37, and 14:36:07, respectively, can be displayed sequentially. The translated question, "Who will benefit?", was completed at 14:36:07. The times 2024 / 10 / 31 14:35:21, 2024 / 10 / 31 14:35:22, 2024 / 10 / 31 14:35:37, and 2024 / 10 / 31 14:36:07 can be selectively displayed or not. Figure 6 shows an example of displaying the times.

[0062] In step S13, the target audio is played according to a preset playback method.

[0063] In this embodiment of the disclosure, the steps of displaying the target speech information on the interface and playing the target speech can be performed together, or the speech can be played directly after obtaining the target speech without displaying the target speech information.

[0064] In this embodiment of the disclosure, the target speech can be played according to a preset playback method.

[0065] In this embodiment of the disclosure, the playback method may include manually selecting the playback method or automatically playing the playback method.

[0066] Manually selecting the playback mode includes playing the corresponding target audio based on the selection result of the displayed icons and / or translation content.

[0067] In this embodiment, the speaker can select an icon or a translation based on the displayed icons and / or translation content, and obtain the corresponding target speech for playback based on the selection result. A selection button can be set after each displayed icon and / or translation content; when any selection button is triggered, it can be determined that the corresponding icon or translation content has been selected.

[0068] Automatic playback methods may include: automatically playing the target speech corresponding to a specified identifier based on a pre-stored specified identifier, or automatically playing the target speech according to the order in which the target speech was translated.

[0069] In this embodiment of the disclosure, the speaker can set a speaker to pay special attention to on the interface, and set the identifier corresponding to the voice of the speaker to be a designated identifier. After the voice corresponding to the designated identifier is translated, the playback module in the terminal can automatically play the corresponding target voice.

[0070] In this embodiment, the speaker can set a mute mode, meaning the DSP will not receive audio from the speaker's BLE audio device. The listener can also turn off or remove BLE audio devices such as headphones to stop receiving the speaker's voice.

[0071] In this embodiment of the disclosure, the playback module can also play the corresponding target audio sequentially according to the order in which the translation was completed.

[0072] In this embodiment of the disclosure, each Bluetooth audio device, the voice collected by the Bluetooth audio device, and the target voice obtained after the voice translation all correspond to the same identifier.

[0073] In this embodiment of the disclosure, the method may further include: after obtaining each target speech, sending the target speech to a Bluetooth audio device corresponding to an identifier other than the target speech's corresponding identifier for playback.

[0074] In this embodiment of the disclosure, each target speech obtained through translation can also be sent to each listener other than its corresponding listener, so as to achieve real-time and rapid unified translation of different languages.

[0075] According to embodiments of this disclosure, a speaker's voice data can be translated into the speaker's language via a terminal and simultaneously sent to the Bluetooth headsets of other listeners, achieving rapid and unified translation across different languages. Furthermore, in scenarios where multiple speakers ask questions simultaneously, the speaker, acting as the receiving end, stores the received translated voice data into corresponding identifiers. This allows the speaker to select which data to play or play it in a preset order, eliminating the noise interference from multiple speakers.

[0076] This disclosure also provides an electronic device 100, as shown in FIG7, including: one or more processors 101; and a memory 102 storing one or more programs thereon, wherein when the one or more programs are executed by the one or more processors 101, the one or more processors 101 implement a voice processing method according to various embodiments of this disclosure.

[0077] Processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU); memory 102 is a device with data storage capabilities, including but not limited to random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH), etc.

[0078] In addition, the wireless communication device according to this disclosure may also include an I / O interface (read / write interface) 103, which is connected between the processor 101 and the memory 102 and configured to realize information interaction between the processor 101 and the memory 102, including but not limited to a data bus.

[0079] In some embodiments, the processor 101, memory 102, and I / O interface 103 are interconnected via a bus, and thus connected to other components of the computing device.

[0080] In the embodiments disclosed herein, any of the embodiments in the foregoing method embodiments are applicable to the electronic device embodiments, and will not be described in detail here.

[0081] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to implement a speech processing method according to various embodiments of this disclosure.

[0082] In the embodiments disclosed herein, any of the embodiments in the foregoing method embodiments are applicable to the computer-readable storage medium embodiments, and will not be described in detail here.

[0083] This disclosure also provides a computer program product, which includes a computer program that, when executed by a processor, causes the processor to implement the speech processing method according to various embodiments of this disclosure.

[0084] In this disclosure, any of the embodiments in the foregoing method embodiments are applicable to the computer program product embodiments, and will not be described in detail here.

[0085] Those skilled in the art will understand that all or some of the functional modules / units disclosed above can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0086] In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components. For example, a physical component may have multiple functions, or a function or step may be executed by several physical components working together.

[0087] Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technique for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; read-only optical disc (CD-ROM), digital versatile disc (DVD) or other optical disc storage; magnetic cartridges, magnetic tapes, disk storage or other magnetic storage; and any other media that can be used to store desired information and can be accessed by a computer. Furthermore, as is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0088] This disclosure has disclosed exemplary embodiments, and although specific terminology has been used, it is for general illustrative purposes only and should not be construed as limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A speech processing method, comprising: Acquire speech from multiple Bluetooth audio devices to obtain multiple different speech, wherein the speech from each Bluetooth audio device corresponds to a language type; The multiple different speech samples are translated based on the language type of each speech sample and the target language type to obtain the corresponding target speech samples respectively; The target audio is played according to the preset playback method.

2. The speech processing method according to claim 1, wherein, Translate multiple different speech samples based on the language type of each speech sample and the target language type to obtain the corresponding target speech samples, including: Start a target number of threads based on the number of different voices; For each speech stream, a separate thread is used for timbre recognition and speech translation. The speech obtained after speech translation is combined with the corresponding timbre to obtain the target speech corresponding to each speech.

3. The speech processing method according to claim 1, wherein, The process of translating multiple different speech sounds based on the language type and target language type of each speech sound to obtain the corresponding target speech sounds includes: The multiple different voices are mixed to obtain a single voice stream; The speech stream is split according to its waveform to obtain multiple sub-speech streams; Each of the split sub-speech streams is translated sequentially, and the timbre of each sub-speech stream is identified. The speech obtained by translating each sub-speech stream is combined with the corresponding timbre to obtain the target speech corresponding to each speech.

4. The speech processing method according to claim 1 further includes: Acquire target information for each target speech and output it for display. The target information includes: the identifier and / or translation content corresponding to the target speech.

5. The speech processing method according to claim 4, wherein, The playback methods include: manually selecting the playback method or automatic playback method; The manual selection of playback mode includes: playing the corresponding target audio based on the selection result of the displayed identifier and / or translation content; The automatic playback method includes: automatically playing the target speech corresponding to the pre-stored specified identifier, or automatically playing the target speech according to the order in which the target speech is translated.

6. The speech processing method according to claim 1, wherein, Each Bluetooth audio device, the voice collected by the Bluetooth audio device, and the target voice obtained after the voice translation correspond to the same identifier.

7. The speech processing method according to claim 6 further includes: After obtaining each target voice, the target voice is sent to the Bluetooth audio device corresponding to an identifier other than the one corresponding to the target voice for playback.

8. The speech processing method according to claim 1, wherein, After acquiring voice data collected from multiple Bluetooth audio devices, the method further includes: Based on a preset duration threshold, each speech segment whose duration is greater than or equal to the duration threshold is segmented, such that the duration of each segmented speech is less than the duration threshold.

9. An electronic device, comprising: One or more processors; A memory having stored one or more programs that, when executed by one or more processors, cause the one or more processors to implement the speech processing method according to any one of claims 1-8.

10. A computer program product comprising a computer program that, when executed by a processor, causes the processor to implement the speech processing method according to any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, wherein when executed by a processor, the computer program causes the processor to implement the speech processing method according to any one of claims 1-8.