Real-time simultaneous interpretation method, system, device, equipment and medium
By establishing a BLE connection between the terminal and the smart wearable device, and utilizing a custom audio transmission protocol and the OPUS algorithm, independent processing and transmission of bidirectional audio streams were achieved, overcoming the limitations of the audio link in mobile terminals and enabling bidirectional, real-time, and smooth simultaneous interpretation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-31
AI Technical Summary
When existing mobile terminals are connected to smart wearable devices, they cannot simultaneously record and play bidirectional audio, and the real-time audio transmission via Bluetooth Low Energy is limited, making it difficult to achieve bidirectional, real-time, and smooth simultaneous interpretation.
By establishing a Bluetooth Low Energy (BLE) connection between the terminal and the smart wearable device, audio is captured and played using the terminal's built-in microphone and speaker. A custom audio transmission protocol and OPUS algorithm are used for audio compression encoding and decoding, enabling independent processing and transmission of bidirectional audio streams.
It solves the limitations of audio links on mobile terminals, enabling bidirectional, real-time, and smooth simultaneous interpretation, ensuring the real-time performance and stability of the audio stream, and reducing the system's dependence on audio devices.
Smart Images

Figure CN121771684A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a real-time simultaneous interpretation method, system, apparatus, device, and medium. Background Technology
[0002] With the rapid development of artificial intelligence technology, speech translation technology based on large models is becoming increasingly efficient and accurate, leading to a growing demand for cross-language communication. Currently, smart wearable devices (such as AI glasses) are becoming more widespread, and many manufacturers have implemented one-way simultaneous interpretation functions, meaning that external speech is received via a mobile phone or glasses, translated, and then output through the glasses.
[0003] However, existing technologies face significant challenges in achieving two-way real-time simultaneous interpretation scenarios where one person holds a mobile phone and the other wears glasses. This is primarily due to the audio link mechanisms of mainstream mobile operating systems (such as Android and iOS): First, the system typically only selects one audio input device as the primary input source at any given time (e.g., either the phone's microphone or a Bluetooth headset microphone), making it difficult to meet the need for simultaneous recording of different speakers' voices on both the phone and glasses. Second, the system typically only selects one audio output device for playback at any given time, making it difficult to simultaneously play the translated audio on both the phone (for the phone holder) and the glasses (for the glasses wearer). Third, if classic Bluetooth audio protocols (such as SCO or A2DP) are used, the phone usually takes over the audio routing, resulting in an inability to flexibly control independent recording and playback on both ends. While Bluetooth Low Energy (BLE) supports data transmission, its data transmission volume is relatively small, making it extremely difficult to directly transmit unprocessed real-time audio streams.
[0004] Therefore, there is an urgent need for a method that can overcome the limitations of existing mobile terminal audio links and achieve bidirectional, real-time, and smooth simultaneous interpretation under low-power Bluetooth connections. Summary of the Invention
[0005] This application provides a real-time simultaneous interpretation method, system, apparatus, and device, aiming to solve the problems that existing mobile terminals cannot simultaneously perform independent two-way audio recording and playback with smart wearable devices when connected to them, as well as the limitations of real-time audio transmission via Bluetooth Low Energy.
[0006] A real-time simultaneous interpretation method is applied to a terminal, wherein the terminal establishes a Bluetooth Low Energy (BLE) connection with a smart wearable device; the method includes: The terminal's built-in microphone is used to capture a first raw audio stream; The system sends a recording command to the smart wearable device via the BLE connection and receives a second raw audio stream returned by the smart wearable device in response to the recording command via the BLE connection, wherein the second raw audio stream is audio data collected and compressed by the smart wearable device; A first translated audio stream is obtained for the first original audio stream, and after decompressing the second original audio stream, a second translated audio stream is obtained for the decompressed second original audio stream. The terminal's built-in speaker is invoked to play the second translated audio stream, and the first translated audio stream is compressed and encoded. The compressed and encoded first translated audio stream is then sent to the smart wearable device via the BLE connection for the smart wearable device to decompress and play.
[0007] Further, the step of sending a recording command to the smart wearable device via the BLE connection and receiving a second raw audio stream returned by the smart wearable device in response to the recording command via the BLE connection includes: The data transmission channel of the BLE connection is encapsulated using a preset custom audio transmission protocol; The recording command is sent through the encapsulated data transmission channel; The data packet returned by the smart wearable device based on the custom audio transmission protocol is received through the encapsulated data transmission channel, and the data packet is parsed to obtain the second original audio stream.
[0008] Furthermore, the second original audio stream is compressed and encoded using the OPUS algorithm; The decompression process for the second original audio stream includes: decoding the second original audio stream using the decoding method corresponding to the OPUS algorithm; The compression encoding of the first translated audio stream includes: encoding the first translated audio stream using the OPUS algorithm.
[0009] Further, the step of obtaining a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtaining a second translated audio stream for the decompressed second original audio stream includes: The first original audio stream is input into a local or cloud-deployed translation model to obtain the first translated audio stream; the decompressed second original audio stream is input into the local or cloud-deployed translation model to obtain the second translated audio stream.
[0010] Furthermore, the terminal is a smartphone or tablet computer, and the smart wearable device is smart glasses, a smart helmet, or smart headphones; During the execution of the method, the system audio input source of the terminal remains configured as the built-in microphone, and the system audio output device of the terminal remains configured as the built-in speaker.
[0011] A real-time simultaneous interpretation system includes a terminal and a smart wearable device, wherein the terminal establishes a Bluetooth Low Energy (BLE) connection with the smart wearable device; The terminal is configured to use its built-in microphone to capture a first raw audio stream, and send a recording command to the smart wearable device via the BLE connection; receive a second raw audio stream returned by the smart wearable device, and decompress the second raw audio stream; obtain a first translated audio stream corresponding to the first raw audio stream and a second translated audio stream corresponding to the decompressed second raw audio stream using a translation model; play the second translated audio stream using the built-in speaker, and compress the first translated audio stream before sending it to the smart wearable device via the BLE connection; The smart wearable device is configured to, in response to the recording command, collect audio data and compress and encode it to generate the second original audio stream, and send the second original audio stream to the terminal via the BLE connection; and to receive the compressed first translated audio stream sent by the terminal, decompress the compressed first translated audio stream and play it.
[0012] Furthermore, the terminal is a smartphone or tablet computer, and the smart wearable device is smart glasses, a smart helmet, or smart headphones; During the execution of the method, the system audio input source of the terminal remains configured as the built-in microphone, and the system audio output device of the terminal remains configured as the built-in speaker.
[0013] A two-way real-time simultaneous interpretation device, configured on a terminal, the device comprising: The audio acquisition and communication module is used to call the built-in microphone of the terminal to acquire a first raw audio stream, send a recording command to the smart wearable device through a Bluetooth Low Energy (BLE) connection established with the smart wearable device, and receive a second raw audio stream transmitted by the smart wearable device through the BLE connection in response to the recording command, wherein the second raw audio stream is audio data acquired by the smart wearable device and compressed and encoded. The processing module is used to obtain a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtain a second translated audio stream for the decompressed second original audio stream. The audio output and transmission module is used to call the built-in speaker of the terminal to play the second translated audio stream, compress and encode the first translated audio stream, and transmit the compressed first translated audio stream to the smart wearable device through the BLE connection so that the smart wearable device can decompress and play it.
[0014] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method as described in any of the preceding claims.
[0015] A computer-readable storage medium having a computer program stored thereon, the computer program, when executed by a processor, implementing the method as described in any of the preceding claims. In one embodiment of this application, the method cleverly circumvents the limitation of Android / iOS systems on the single audio input / output source. By allowing the terminal to utilize its built-in microphone and speaker, the terminal's system audio layer is occupied in local mode, and the system does not automatically switch the audio routing to the smart wearable device. Secondly, for audio interaction on the smart wearable device side, this solution does not rely on the system's Bluetooth audio protocol stack, but instead transmits it as pure data through the BLE channel. This means that the terminal can process two independent audio streams simultaneously: one from the local hardware interface (first raw audio stream) and the other from the BLE data interface (second raw audio stream), thus truly realizing bidirectional, simultaneous voice acquisition and playback. Finally, by compressing and encoding the audio transmitted over the BLE link, the problem of the narrow bandwidth of BLE making it difficult to support real-time, high-quality audio streams is solved, ensuring the real-time performance and fluency of bidirectional simultaneous interpretation. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a real-time simultaneous interpretation method according to an embodiment of this application; Figure 2 This is a system architecture diagram of a real-time simultaneous interpretation system according to one embodiment of this application; Figure 3 This is a schematic diagram of a real-time simultaneous interpretation device according to one embodiment of this application; Figure 4This is a schematic diagram of the structure of a computer device according to one embodiment of this application. Detailed Implementation
[0018] To make the technical problems, technical solutions, and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0019] Example 1 In this embodiment, as Figure 1 As shown, a real-time simultaneous interpretation method is provided, applied to a terminal, wherein the terminal establishes a Bluetooth Low Energy (BLE) connection with a smart wearable device; the method includes: S10. Call the built-in microphone of the terminal to collect the first raw audio stream; S20. Send a recording command to the smart wearable device through the BLE connection, and receive a second raw audio stream returned by the smart wearable device in response to the recording command through the BLE connection, wherein the second raw audio stream is audio data collected and compressed by the smart wearable device; S30. Obtain a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtain a second translated audio stream for the decompressed second original audio stream. S40. The terminal's built-in speaker is invoked to play the second translated audio stream, and the first translated audio stream is compressed and encoded. The compressed and encoded first translated audio stream is then sent to the smart wearable device via the BLE connection for the smart wearable device to decompress and play.
[0020] In this embodiment, the terminal can be an electronic device with computing and communication capabilities, such as a smartphone or tablet. The smart wearable device can be smart glasses, a smart helmet, or smart headphones, etc., and is not specifically limited. Bluetooth Low Energy (BLE) is a low-cost, short-range, interoperable, and robust wireless technology. Unlike classic Bluetooth, this embodiment does not establish a traditional audio link (such as HFP or A2DP), but rather establishes a BLE-based data transmission link.
[0021] Specifically, the first raw audio stream refers to the speech (e.g., native language) emitted by the user holding the terminal, directly captured by the terminal's own microphone; the second raw audio stream refers to the speech (e.g., foreign language) emitted by the user wearing the smart wearable device, captured by the wearable device's microphone. In this method, the terminal acts as the main control and processing center. The terminal is not only responsible for capturing its own audio but also controls the wearable device to simultaneously start recording by sending recording commands. The audio captured by the wearable device is not directly transmitted but is "compressed and encoded" and then sent back to the terminal via a BLE channel. Subsequently, the terminal uses translation capabilities (such as large models) to translate the first raw audio stream into the target language (i.e., the first translated audio stream, e.g., translated into a foreign language) and translates the decompressed second raw audio stream into the native language (i.e., the second translated audio stream, e.g., translated into Chinese). Finally, the terminal plays the translated speech (second translated audio stream) to the user holding the terminal through its own speaker, while simultaneously compressing its own translated speech (first translated audio stream) and sending it back to the wearable device, which then plays it to the wearer.
[0022] The above technical solution achieves the following beneficial effects by establishing a BLE connection and conducting specific data interaction between the terminal and the smart wearable device: First, this method cleverly circumvents the limitation of Android / iOS systems on the single audio input / output source. By allowing the terminal to utilize its built-in microphone and speaker, the terminal's system audio layer is occupied in local mode, and the system does not automatically switch the audio routing to the smart wearable device. Second, for audio interaction on the smart wearable device side, this solution does not rely on the system's Bluetooth audio protocol stack, but instead transmits it as pure data through the BLE channel. This means that the terminal can simultaneously process two independent audio streams: one from the local hardware interface (the first raw audio stream) and the other from the BLE data interface (the second raw audio stream), thus truly realizing bidirectional, simultaneous voice acquisition and playback. Finally, by compressing and encoding the audio transmitted over the BLE link, the problem of the narrow bandwidth of BLE making it difficult to support real-time, high-quality audio streams is solved, ensuring the real-time performance and fluency of bidirectional simultaneous interpretation.
[0023] Example 2 In this embodiment, in conjunction with Embodiment 1 above, the step of sending a recording command to the smart wearable device via the BLE connection and receiving a second raw audio stream returned by the smart wearable device in response to the recording command via the BLE connection includes: The data transmission channel of the BLE connection is encapsulated using a preset custom audio transmission protocol; the recording command is then sent through the encapsulated data transmission channel. The data packet returned by the smart wearable device based on the custom audio transmission protocol is received through the encapsulated data transmission channel, and the data packet is parsed to obtain the second original audio stream.
[0024] In this embodiment, the preset custom audio transmission protocol refers to a set of communication rules specifically defined on top of the application layer of the BLE standard protocol stack (such as GATT Generic Attribute Profile) for the encapsulation, transmission, and decapsulation of audio data. Because Bluetooth Low Energy (BLE) is primarily designed for transmitting low-bandwidth control or sensor data, its Maximum Transmission Unit (MTU) is small, and it cannot provide a guaranteed audio stream channel like Classic Bluetooth. Directly transmitting continuous audio streams is highly susceptible to packet loss, out-of-order delivery, or packet merging.
[0025] To address the aforementioned issues, the custom protocol in this embodiment logically encapsulates the BLE data channel. Specifically, this protocol defines a specific packet structure. As a concrete example, this packet structure can include two parts: a header and a payload. The header can contain a start frame identifier (e.g., 0xAA55, used to identify the start of the data packet), a sequence number (used by the receiver to detect packet loss and reorder out-of-order data packets), a data type identifier (used to distinguish between control commands and audio data), and a data length.
[0026] Payload: Stores the actual audio segment data after compression and encoding.
[0027] Based on this protocol, when a recording command is sent, the terminal constructs a data packet containing a specific command code (e.g., CommandID 0x01 indicates start recording) and sends it to the smart wearable device. When the smart wearable device sends back the second original audio stream, it segments the continuously acquired audio stream, adds a header containing a sequence number, and encapsulates it into a data packet sequence conforming to the BLE MTU size (e.g., 20 bytes or more, depending on the handshake negotiation result) before sending it. After receiving these data packets, the terminal removes the header according to the protocol rules and concatenates the audio data in the payload sequentially according to the sequence number, thereby reconstructing the complete second original audio stream.
[0028] As can be seen, in this embodiment, the custom audio transmission protocol refers to a set of data packetization and depacketization rules defined at the application layer on top of the BLE standard protocol stack (such as GATTprofile). Since BLE is mainly designed for transmitting small amounts of data (such as sensor readings), directly transmitting continuous audio streams is prone to packet loss, out-of-order delivery, or delays. By encapsulating the data transmission channel, the continuous audio stream can be segmented into data packets suitable for the BLE transmission unit (MTU) size, which may include header information (such as sequence number, timestamp, payload length, etc.).
[0029] By employing the aforementioned technical features and encapsulating the BLE channel using a pre-defined custom audio transmission protocol, the transmission stability of audio data in low-bandwidth BLE environments can be significantly improved. The custom protocol standardizes the interaction format between recording commands and audio data, ensuring that the terminal can accurately parse compressed audio data packets from wearable devices, preventing decoding errors caused by packet concatenation or fragmentation, thereby guaranteeing the integrity of the second original audio stream and providing a foundation for accurate subsequent translation.
[0030] Example 3 In this embodiment, in conjunction with Embodiment 1 above, the second original audio stream is compressed and encoded using the OPUS algorithm; the decompression of the second original audio stream includes: decoding the second original audio stream using the decoding method corresponding to the OPUS algorithm; the compression and encoding of the first translated audio stream includes: encoding the first translated audio stream using the OPUS algorithm.
[0031] In this embodiment, the OPUS algorithm refers to the Opus Audio Codec. It is a completely open, free, and feature-rich audio encoding format standardized by the IETF (Internet Engineering Task Force). Compared to traditional formats such as MPEG-1 Audio Layer III (MP3) or Advanced Audio Coding (AAC), Opus offers better sound quality at low bitrates and has extremely low latency, making it ideal for real-time interactive scenarios. In this solution, after the smart wearable device acquires audio, it uses the Opus algorithm to compress it into extremely small data packets (i.e., the second original audio stream) before sending it via BLE. Similarly, after the terminal generates the translated speech (the first translated audio stream), it also uses the Opus algorithm to compress it before sending it to the wearable device.
[0032] Using the Opus algorithm as the compression encoding method brings two main technical benefits: First, Opus's high compression ratio significantly reduces the size of audio data, making it compatible with the limited transmission bandwidth of BLE (Bluetooth Low Energy), avoiding audio transmission stutters or interruptions, and enabling real-time voice transmission over non-traditional Bluetooth audio (non-A2DP / SCO) links. Second, Opus's low latency is crucial for simultaneous interpretation scenarios, shortening the time interval between speaking and hearing the translated speech, enhancing the immersive and real-time experience of two-way communication.
[0033] Example 4 In this embodiment, in conjunction with Embodiment 1 above, the step of obtaining a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtaining a second translated audio stream for the decompressed second original audio stream includes: The first original audio stream is input into a large translation model deployed locally or in the cloud to obtain the first translated audio stream; The decompressed second original audio stream is input into the local or cloud-deployed translation big model to obtain the second translated audio stream.
[0034] In this embodiment, the large-scale translation model refers to a large-scale language model trained based on deep learning techniques (such as the Transformer architecture), possessing powerful natural language understanding and generation capabilities. These models can be deployed on cloud servers to utilize powerful computing power, or they can be deployed locally on the terminal after quantization pruning to protect privacy and reduce latency.
[0035] By employing a large-scale translation model, compared to traditional statistical or rule-based machine translation, this method can more accurately understand the context, slang, and emotions in spoken language, thereby generating more natural and accurate translated audio (first and second translated audio streams). Furthermore, it supports flexible deployment, both locally and in the cloud, enabling the method to adapt to different network environments and user privacy requirements, ensuring the reliability and high quality of the translation service.
[0036] Example 5 In this embodiment, in conjunction with the above embodiments, the terminal is a smartphone or tablet computer, and the smart wearable device is smart glasses, a smart helmet, or smart headphones; during the execution of the method, the system audio input source of the terminal remains configured as the built-in microphone, and the system audio output device of the terminal remains configured as the built-in speaker.
[0037] This embodiment clarifies the hardware configuration and key system settings. Typically, when a mobile phone connects to Bluetooth headphones or glasses, the operating system automatically switches the system audio input source and system audio output device to the smart wearable device. This embodiment specifically emphasizes forcing or maintaining the terminal to use its own microphone and speaker during method execution. This means that the BLE connection between the terminal and the smart wearable device exists only as a data channel, not as an audio channel for the operating system.
[0038] This feature is the core of enabling two-way simultaneous interpretation, and its technical effect is to completely decouple the audio control of the terminal and the wearable device. By keeping the terminal system's audio source as the built-in microphone and speaker, it ensures that the person holding the terminal can record speech without interference and directly hear the translation result; at the same time, the audio of the wearable device is processed independently using the BLE data channel, so that the two audio streams do not compete for system resources. This dual-track processing method (one path goes through the system audio layer, and the other path goes through the application layer BLE data transmission) successfully solves the problem that mainstream mobile phone systems cannot simultaneously schedule two different physical audio devices.
[0039] Example 6 In this embodiment, as Figure 2 As shown, a real-time simultaneous interpretation system is provided, including a terminal and a smart wearable device, wherein the terminal establishes a Bluetooth Low Energy (BLE) connection with the smart wearable device; The terminal is configured to use its built-in microphone to capture a first raw audio stream, and send a recording command to the smart wearable device via the BLE connection; receive a second raw audio stream returned by the smart wearable device, and decompress the second raw audio stream; obtain a first translated audio stream corresponding to the first raw audio stream and a second translated audio stream corresponding to the decompressed second raw audio stream using a translation model; play the second translated audio stream using the built-in speaker, and compress the first translated audio stream before sending it to the smart wearable device via the BLE connection; The smart wearable device is configured to, in response to the recording command, collect audio data and compress and encode it to generate the second original audio stream, and send the second original audio stream to the terminal via the BLE connection; and to receive the compressed first translated audio stream sent by the terminal, decompress the compressed first translated audio stream and play it.
[0040] In this embodiment, the system materializes the process in the above method embodiments into specific device interactions. The terminal acts as the computing and control center, while the smart wearable device acts as a subordinate peripheral with audio acquisition and playback capabilities.
[0041] The system's technological advantage lies in providing a complete hardware solution that enables barrier-free face-to-face cross-language communication. Users holding the terminal and those wearing smart devices can converse naturally, with the system automatically handling bidirectional data acquisition, transmission, translation, and playback. Leveraging the low-power characteristics of BLE connectivity, the system also ensures the long battery life of smart wearable devices during extended operation, enhancing the product's practicality.
[0042] Example 7 In this embodiment, in conjunction with Embodiment Six above, the terminal is a smartphone or tablet computer, and the smart wearable device is smart glasses, a smart helmet, or smart headphones; during the execution of the method, the system audio input source of the terminal remains configured as the built-in microphone, and the system audio output device of the terminal remains configured as the built-in speaker.
[0043] This embodiment corresponds to the hardware limitations and configuration states at the system level. By limiting the specific types of terminals and smart wearable devices and maintaining the configuration state of audio I / O (input / output), the operability of the system on actual physical devices is ensured. Its effect is similar to Embodiment 5, guaranteeing independent parallel processing of dual audio streams from a system architecture perspective, avoiding functional conflicts caused by the operating system's default takeover behavior of Bluetooth audio devices.
[0044] For specific limitations regarding simultaneous interpreting systems, please refer to the limitations on simultaneous interpreting methods mentioned above, which will not be repeated here.
[0045] Example 8 A real-time simultaneous interpretation device is provided, configured on a terminal, and this device corresponds one-to-one with the real-time simultaneous interpretation methods described in the above embodiments. For example... Figure 3 As shown, the real-time simultaneous interpretation device 30 includes an audio acquisition and communication module 301, a processing module 302, and an audio output and transmission module 303. Detailed descriptions of each functional module are as follows: The audio acquisition and communication module 301 is used to call the built-in microphone of the terminal to acquire a first raw audio stream, send a recording command to the smart wearable device through a Bluetooth Low Energy (BLE) connection established with the smart wearable device, and receive a second raw audio stream transmitted by the smart wearable device through the BLE connection in response to the recording command, wherein the second raw audio stream is audio data acquired by the smart wearable device and compressed and encoded. The processing module 302 is used to obtain a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtain a second translated audio stream for the decompressed second original audio stream; The audio output and transmission module 303 is used to call the built-in speaker of the terminal to play the second translated audio stream, compress and encode the first translated audio stream, and transmit the compressed first translated audio stream to the smart wearable device through the BLE connection so that the smart wearable device can decompress and play it.
[0046] In this embodiment, the device is divided into three logical functional modules. The audio acquisition and communication module is responsible for acquiring local audio and receiving remote compressed audio; the processing module is responsible for decompression and translation; and the audio output and transmission module is responsible for playing local audio and sending remote compressed audio.
[0047] The advantages of this modular design are: reduced coupling in software development, allowing each module to be optimized independently. For example, the translation or decompression algorithm in the processing module can be upgraded separately without changing the audio acquisition logic. Simultaneously, this device structure clearly maps the data flow of bidirectional interpretation, facilitating specific code implementation and functional deployment on terminal devices (such as at the APP level).
[0048] For specific limitations regarding the real-time simultaneous interpretation device, please refer to the limitations of the real-time simultaneous interpretation method above, which will not be repeated here. Each module in the aforementioned real-time simultaneous interpretation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0049] Example 9 In this embodiment, as Figure 4 As shown, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the real-time simultaneous interpretation method as described in any of the foregoing embodiments. For example, implementing... Figure 1 S10-S40, as shown, will not be described again here to avoid repetition. Alternatively, in this embodiment, the processor executes a computer program to implement the terminal's function as a real-time simultaneous interpretation device; this will also not be described again to avoid repetition.
[0050] The electronic device includes smartphones or tablets, with no specific limitation. By equipping it with a processor and memory at the hardware level and running specific computer programs, a general-purpose electronic device can be transformed into a dedicated device with two-way real-time simultaneous interpretation capabilities. Its technical advantage lies in utilizing existing, widely available mobile terminal hardware resources; the technical solution of this application can be implemented through software upgrades, eliminating the need for users to purchase expensive dedicated translation devices and reducing usage costs.
[0051] Example 10 In this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the method described in any of the foregoing embodiments. For example, implementing... Figure 1 S10-S40, as shown, will not be described again here to avoid repetition. Alternatively, in this embodiment, the processor executes a computer program to implement the terminal's function as a real-time simultaneous interpretation device; this will also not be described again to avoid repetition.
[0052] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A real-time simultaneous interpretation method, characterized in that, Applied to a terminal, the terminal establishes a Bluetooth Low Energy (BLE) connection with a smart wearable device; the method includes: The terminal's built-in microphone is used to capture a first raw audio stream; The system sends a recording command to the smart wearable device via the BLE connection and receives a second raw audio stream returned by the smart wearable device in response to the recording command via the BLE connection, wherein the second raw audio stream is audio data collected and compressed by the smart wearable device; A first translated audio stream is obtained for the first original audio stream, and after decompressing the second original audio stream, a second translated audio stream is obtained for the decompressed second original audio stream. The terminal's built-in speaker is invoked to play the second translated audio stream, and the first translated audio stream is compressed and encoded. The compressed and encoded first translated audio stream is then sent to the smart wearable device via the BLE connection for the smart wearable device to decompress and play.
2. The method according to claim 1, characterized in that, The step of sending a recording command to the smart wearable device via the BLE connection and receiving a second raw audio stream returned by the smart wearable device in response to the recording command via the BLE connection includes: The data transmission channel of the BLE connection is encapsulated using a preset custom audio transmission protocol; The recording command is sent through the encapsulated data transmission channel; The data packet returned by the smart wearable device based on the custom audio transmission protocol is received through the encapsulated data transmission channel, and the data packet is parsed to obtain the second original audio stream.
3. The method according to claim 1, characterized in that, The second original audio stream is compressed and encoded using the OPUS algorithm; The decompression process for the second original audio stream includes: decoding the second original audio stream using the decoding method corresponding to the OPUS algorithm; The compression encoding of the first translated audio stream includes: encoding the first translated audio stream using the OPUS algorithm.
4. The method according to claim 1, characterized in that, The step of obtaining a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtaining a second translated audio stream for the decompressed second original audio stream includes: The first original audio stream is input into a local or cloud-deployed translation model to obtain the first translated audio stream; the decompressed second original audio stream is input into the local or cloud-deployed translation model to obtain the second translated audio stream.
5. The method according to claim 4, characterized in that, The terminal is a smartphone or tablet computer, and the smart wearable device is a smart glasses, a smart helmet, or a smart headset. During the execution of the method, the system audio input source of the terminal remains configured as the built-in microphone, and the system audio output device of the terminal remains configured as the built-in speaker.
6. A real-time simultaneous interpretation system, characterized in that, Includes a terminal and a smart wearable device, wherein the terminal establishes a Bluetooth Low Energy (BLE) connection with the smart wearable device; The terminal is used to call the built-in microphone to collect a first raw audio stream and send a recording command to the smart wearable device through the BLE connection. Receive the second raw audio stream returned by the smart wearable device, and decompress the second raw audio stream; The first translated audio stream corresponding to the first original audio stream and the second translated audio stream corresponding to the decompressed second original audio stream are obtained by using a large translation model. The built-in speaker is invoked to play the second translated audio stream, and the first translated audio stream is compressed and sent to the smart wearable device via the BLE connection; The smart wearable device is configured to, in response to the recording command, collect audio data and compress and encode it to generate the second original audio stream, and send the second original audio stream to the terminal via the BLE connection; And receive the compressed first translated audio stream sent by the terminal, decompress the compressed first translated audio stream and play it.
7. The system according to claim 6, characterized in that, The terminal is a smartphone or tablet computer, and the smart wearable device is a smart glasses, a smart helmet, or a smart headset. During the execution of the method, the system audio input source of the terminal remains configured as the built-in microphone, and the system audio output device of the terminal remains configured as the built-in speaker.
8. A two-way real-time simultaneous interpretation device, characterized in that, Configured in a terminal, the device includes: The audio acquisition and communication module is used to call the built-in microphone of the terminal to acquire a first raw audio stream, send a recording command to the smart wearable device through a Bluetooth Low Energy (BLE) connection established with the smart wearable device, and receive a second raw audio stream transmitted by the smart wearable device through the BLE connection in response to the recording command, wherein the second raw audio stream is audio data acquired by the smart wearable device and compressed and encoded. The processing module is used to obtain a first translated audio stream for the first original audio stream, and after decompressing the second original audio stream, obtain a second translated audio stream for the decompressed second original audio stream. The audio output and transmission module is used to call the built-in speaker of the terminal to play the second translated audio stream, compress and encode the first translated audio stream, and transmit the compressed first translated audio stream to the smart wearable device through the BLE connection so that the smart wearable device can decompress and play it.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.