Simultaneous interpretation method, device and equipment and storage medium

By deploying a local AI service module on the simultaneous interpretation device for offline conversion and wired transmission of audio data, the problems of data privacy and security and high latency in existing simultaneous interpretation technologies are solved, achieving low-latency and high-security simultaneous interpretation.

CN122369460APending Publication Date: 2026-07-10MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MOORE THREADS TECH CO LTD
Filing Date
2026-04-17
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing simultaneous interpretation technologies suffer from insufficient data privacy and security, strong network dependence, high processing latency, and limited audio output methods, making it difficult to meet the needs of real-time and low-latency simultaneous interpretation.

Method used

Using offline simultaneous interpretation equipment, an AI service module deployed locally converts audio data in the first language into audio data in the second language offline, and sends it to the communication device via wired transmission, avoiding uploading to the cloud server, thus achieving secure conversion and low-latency transmission of audio data.

Benefits of technology

It ensures data security, reduces transmission latency, takes into account the transmission quality of audio data, is suitable for scenarios with strict requirements for data privacy and security, and enhances the robustness and applicability of the simultaneous interpretation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369460A_ABST
    Figure CN122369460A_ABST
Patent Text Reader

Abstract

This application discloses a simultaneous interpretation method, apparatus, device, and storage medium, relating to the field of audio processing technology. The aforementioned simultaneous interpretation method is applied to a simultaneous interpretation device, which includes a locally deployed AI service module. The method includes: acquiring first audio data in a first language; converting the first audio data into second audio data in a second language offline, based on the locally deployed AI service module; and transmitting the second audio data via wired transmission to a first communication device in a call state, wherein the second audio data is provided to a second communication device engaged in a call with the first communication device. This method enables low-latency transmission during simultaneous interpretation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to a simultaneous interpretation method, apparatus, device, and storage medium. Background Technology

[0002] Simultaneous interpreting refers to the process of translating a speaker's speech continuously in the original language into a speech expressed in a specified language without interruption.

[0003] In related technologies, the collected speech content is first uploaded to a cloud server via a terminal device; then, for the speech content expressed in the original language, ASR (Automatic Speech Recognition), MT (Machine Translation) or AST (Automatic Speech Translation), and TTS (Text-to-Speech Synthesis) are sequentially performed on the cloud server to obtain the translated audio content; finally, the text content or audio content processed by the cloud server is sent back to the terminal device for display or playback.

[0004] However, how to avoid privacy leaks during simultaneous interpretation is a problem that urgently needs to be solved. Summary of the Invention

[0005] This application provides a simultaneous interpretation method, apparatus, device, and storage medium. The technical solution provided by this application is as follows: According to one aspect of the embodiments of this application, a simultaneous interpretation method is provided, applied to a simultaneous interpretation device, the simultaneous interpretation device including an AI service module deployed locally, the method comprising: acquiring first audio data in a first language; converting the first audio data into second audio data in a second language in an offline state based on the AI ​​service module deployed locally; and sending the second audio data to a first communication device in a call state via wired transmission, the second audio data being used to provide to a second communication device that is in a call with the first communication device.

[0006] According to one aspect of the embodiments of this application, a simultaneous interpretation method is provided, applied to a first communication device. The method includes: during a conversation between the first communication device and a second communication device, receiving second audio data in a second language sent by a simultaneous interpretation device via wired transmission, wherein the second audio data is obtained by converting first audio data in a first language; and sending the second audio data to the second communication device.

[0007] According to one aspect of the embodiments of this application, a simultaneous interpretation apparatus is provided, the apparatus comprising: an acquisition module for acquiring first audio data in a first language; a processing module for converting the first audio data into second audio data in a second language in an offline state based on an AI service module deployed locally; and a sending module for sending the second audio data to a first communication device in a call state via wired transmission, the second audio data being provided to a second communication device that is in a call with the first communication device.

[0008] According to one aspect of the embodiments of this application, a simultaneous interpretation apparatus is provided, the apparatus comprising: a receiving module, configured to receive second audio data in a second language sent by a simultaneous interpretation device via wired transmission during a conversation between a first communication device and a second communication device, wherein the second audio data is obtained by converting first audio data in a first language; and a sending module, configured to send the second audio data to the second communication device.

[0009] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described simultaneous interpretation method.

[0010] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, in which a computer program is stored, which is loaded and executed by a processor to implement the above-described simultaneous interpretation method.

[0011] According to one aspect of the embodiments of this application, a chip is provided, the chip including programmable logic circuits and / or program instructions, which, when the chip is running, are used to implement the above-described simultaneous interpretation method.

[0012] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program, which is loaded and executed by a processor to implement the above-described simultaneous interpretation method.

[0013] The technical solution provided in this application can bring the following beneficial effects: The offline simultaneous interpretation device converts the acquired first audio data in the first language into second audio data in the second language, and then sends the second audio data to the first communication device in a call state via wired transmission. This allows the first communication device to provide the second audio data in the second language to the second communication device it is talking to. The conversion process from first audio data to second audio data is completed on the AI ​​service module deployed locally on the offline simultaneous interpretation device, without uploading to a cloud server or other devices. This avoids audio content leakage, ensures data security, and reduces latency introduced by data transmission. In addition, the second audio data is sent to the first communication device via wired transmission, which balances the requirements of audio data transmission quality and low latency transmission. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram of a computer system provided in one possible implementation of this application; Figure 2 This is a flowchart of a possible implementation of the simultaneous interpretation method provided in this application; Figure 3 This is a flowchart of the simultaneous interpretation process of spoken Chinese voice provided in one possible implementation of this application; Figure 4 This is a flowchart of a simultaneous interpretation process of incoming English speech provided in one possible implementation of this application; Figure 5 This is a flowchart of a simultaneous interpretation process for calling out Chinese voice and receiving English voice, provided in one possible implementation of this application; Figure 6 This is a structural block diagram of a simultaneous interpretation device provided in one possible implementation of this application; Figure 7 This is a structural block diagram of a simultaneous interpretation device provided in one possible implementation of this application; Figure 8 This is a structural block diagram of a computer device provided in one possible implementation of this application. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0017] Before introducing the technical solution proposed in this application, the relevant technologies will be briefly described below.

[0018] Simultaneous interpreting refers to the continuous translation of a speaker's speech into a specified language while the speaker speaks in the original language. Based on the difference in implementation methods, simultaneous interpreting can be divided into machine simultaneous interpreting and human simultaneous interpreting. Human simultaneous interpreting relies on human interpreters to translate in real time as the original language speech is continuously input; machine simultaneous interpreting uses technologies such as ASR, MT, AST, and TTS to generate the specified language speech simultaneously with the original language speech.

[0019] The following sections will explain ASR, MT / AST, and TTS respectively.

[0020] ASR: Used to convert continuous speech content expressed in the original language into a corresponding text representation.

[0021] MT / AST: Used to convert text in a source language into text in a specified language.

[0022] TTS: Used to convert translated text in a specified language into playable audio content, enabling real-time generation of continuous audio content expressed in a specified language.

[0023] For example, for speech content expressed in the original language, text in the specified language is generated sequentially through ASR and MT / AST, and then continuous speech content expressed in the specified language is generated through TTS.

[0024] Currently, machine simultaneous interpretation in related technologies usually relies on network connections, which not only results in insufficient privacy and security of audio data, but also in high processing latency during the simultaneous interpretation process.

[0025] For example, in some solutions, audio data collected by a communication-enabled terminal device (such as a mobile terminal device or a PC) is uploaded to a cloud server to complete the machine simultaneous interpretation process. However, there is a risk of privacy leakage during the audio data upload process. Furthermore, the terminal device typically needs to wait for the complete recording of a sentence or paragraph before uploading the collected audio data to the cloud server. The audio data transmission process and the cloud server's processing of the audio data take a considerable amount of time, resulting in high latency in the machine simultaneous interpretation process and making it difficult to meet the requirements for real-time, low-latency simultaneous interpretation. In addition, this transmission process relies on the network connection between the terminal device and the cloud server. If the terminal device cannot connect to the network, or if the network connection between the terminal device and the cloud server is unstable, the simultaneous interpretation function will be limited.

[0026] For example, in another solution, the audio data collected by a terminal device with communication capabilities is processed locally using ASR (Automatic Speech Recognition) to obtain the recognized text in the original language. Then, the recognized text is translated over the network to obtain the translated text in the specified language. Finally, the translated text in the specified language is processed locally using TTS (Text-to-Speech) to obtain the translated audio data, which is then played on the terminal device. While this machine simultaneous interpretation process restricts the transmission of the original language audio data, the recognized text in the original language still requires online translation, making completely offline operation impossible. This still presents insufficient privacy and security risks. Furthermore, after acquiring the original language audio data, terminal devices generally use non-streaming processing. Specifically, ASR and online translation are only performed after receiving the complete audio segment. This results in high response latency, making it difficult to meet the low-latency requirements of simultaneous interpretation, especially unsuitable for scenarios requiring immediate feedback. In addition, the translated audio data is mainly played through the terminal device's audio output module (e.g., speaker), limiting the flexibility of playback and making it difficult to meet the needs of teleconferences or cross-device calls.

[0027] Therefore, the relevant technologies suffer from problems such as insufficient data privacy and security, strong network dependence, high processing latency in the simultaneous interpretation process, and limited audio output methods.

[0028] Based on this, this application provides a simultaneous interpretation method. An offline simultaneous interpretation device converts the acquired first audio data in a first language into second audio data in a second language, and sends the second audio data to a first communication device in a call state via wired transmission. This enables the first communication device to provide the second audio data in the second language to the second communication device with which it is in a call. The conversion process from first audio data to second audio data is completed by an AI service module deployed locally on the offline simultaneous interpretation device, without uploading to a cloud server or other devices. This avoids audio content leakage, ensures data security, and reduces latency introduced by data transmission. Furthermore, the second audio data is sent to the first communication device via wired transmission, balancing the requirements for audio data transmission quality and low-latency transmission.

[0029] Please refer to Figure 1 The diagram illustrates a computer system provided in one embodiment of this application. The computer system may include a first communication device 10, a second communication device 20, and an offline simultaneous interpretation device 30.

[0030] Both the first communication device 10 and the second communication device 20 are terminal devices with call functionality. The first communication device 10 and the second communication device 20 include, but are not limited to, networked computers (e.g., via VoIP (Voice over Internet Protocol) software), landline phones, mobile phones, tablets, smart voice interaction devices, PCs (Personal Computers), and other electronic devices. The first communication device 10 may have a client installed for accessing the simultaneous interpretation device 30, or it may be used to access the simultaneous interpretation services provided by the simultaneous interpretation device 30 via a browser or other means.

[0031] Optionally, the first communication device 10 includes at least one wired transmission interface (e.g., a Type-C interface), through which the first communication device 10 can connect with the simultaneous interpretation device 30 and transmit audio signals.

[0032] Optionally, the first communication device 10 may also have audio input or output capabilities. For example, through audio input or output capabilities, the first communication device 10 can receive audio signals (e.g., translated speech content) from the simultaneous interpretation device 30 and transmit them. For example, through audio input or output capabilities, the first communication device 10 can transmit external conversation voice received from the second communication device 20 to the simultaneous interpretation device 30.

[0033] In this embodiment of the application, the offline simultaneous interpretation device 30 is used to perform recognition, translation, speech synthesis and other processing on the acquired audio content during the conversation between the first communication device 10 and the second communication device 20, so as to realize the language conversion of the audio content.

[0034] Optionally, "simultaneous interpretation device 30 offline" means that the device is not connected to the network or has not established a communication connection with the network.

[0035] In some embodiments, the simultaneous interpretation device 30 may include one or more of the following modules: a front-end application module, an AI (Artificial Intelligence) service module, and an audio processing module.

[0036] The following is a description of the various modules included in the 30 simultaneous interpretation devices.

[0037] Front-end application module: This module can run on the operating system of the simultaneous interpretation device 30 and is used to implement functions such as user interface interaction, audio acquisition control, audio playback control, audio routing selection, and text display. Optionally, the front-end application module can be developed based on frameworks such as Electron or React, but this embodiment is not limited thereto.

[0038] AI Service Module: Deployed locally on the simultaneous interpretation device 30, this module provides locally running ASR, AST, and TTS services. Each service provided by the AI ​​Service Module pre-loads offline models (e.g., locally deployed audio processing models), enabling the completion of corresponding audio processing functions without relying on external communication network connections.

[0039] Optionally, the ASR service can be used to recognize audio content as text content; the AST service can be used to translate text content expressed in one language into text content expressed in another language; and the TTS service can be used to synthesize text content into an audio signal that can be played in speech form.

[0040] Optionally, services such as ASR, AST, and TTS can be implemented in the form of audio processing models, algorithm units, modules, or functional services, and the embodiments of this application are not limited thereto.

[0041] This application does not impose any special restrictions on the connection method between the AI ​​service module and various functional services (such as ASR service, AST service and TTS service). For example, the AI ​​service module can connect to each service through the WebSocket interface.

[0042] Audio processing module: may include, but is not limited to, audio acquisition module, audio preprocessing module, audio playback module, audio routing control module, and user interface display module.

[0043] Optionally, the audio acquisition module is used to acquire or receive raw audio data in real time; the audio preprocessing module is used to preprocess the acquired or received raw audio data; the audio playback module is used to play the generated audio signal or the received audio signal through the audio playback module of the simultaneous interpretation device 30 or a specified audio output device; the audio routing control module is used to sequentially list the audio output devices that can be used to play audio, and in response to the selection instruction of the audio output device, to route the audio stream to the selected audio output device; the user interface display module is used to display visual information such as translated text, system status, and device selection.

[0044] This application does not impose special restrictions on the content of audio preprocessing, such as downsampling, format conversion, and encapsulation.

[0045] Optionally, the audio acquisition module, audio playback module, user interface display module, etc., can be external devices that are connected to the simultaneous interpretation device 30 via an external wired connection, or external devices that are connected to the simultaneous interpretation device 30 via a wireless connection (e.g., Bluetooth), or functional modules integrated inside the simultaneous interpretation device 30.

[0046] This application does not impose special limitations on the implementation of the audio acquisition module; for example, it can be a microphone or other sensors used to acquire ambient audio content. Similarly, this application does not impose special limitations on the implementation of the audio playback module; for example, it can be a speaker, headphones, or other audio output devices used to play audio content.

[0047] Optionally, the audio routing control module routes the audio stream to the selected audio output device via the setSinkId API.

[0048] This application does not impose any special restrictions on the way the audio preprocessing module is called. For example, it can be implemented using the AudioWorklet or ScriptProcessorNode of the WebAudio API.

[0049] In this embodiment of the application, there is no special limitation on the number of devices that can communicate with the first communication device 10. The device that establishes a communication connection with the first communication device 10 can be one (e.g., the second communication device 20) or more.

[0050] The first communication device 10 and the second communication device 20 can communicate with each other via a network. This network can be a wired network or a wireless network.

[0051] Optionally, the simultaneous interpretation device 30 can be an independent hardware device connected to the first communication device 10, or it can be a functional module integrated into the first communication device 10.

[0052] Accordingly, the first communication device 10 and the simultaneous interpretation device 30 can communicate with each other via a network. This network can be a wired network or a wireless network.

[0053] In some embodiments, when the connection network between the first communication device 10 and the simultaneous interpretation device 30 is a wired network, the first communication device 10 and the simultaneous interpretation device 30 can be connected using an audio pairing cable. The audio pairing cable can be used for bidirectional transmission of digital audio signals. Using the audio pairing cable can reduce the transmission delay of audio data during a conversation between the first communication device 10 and the second communication device 20, thereby improving the response efficiency during simultaneous interpretation. In this embodiment, there are no special restrictions on the interface type of the audio pairing cable; for example, it can be a cable with dual Type-C interfaces, or a cable containing at least one Type-C interface.

[0054] For example, please refer to Figure 1 User 1 initiates a call through the first communication device 10 and establishes a call connection with the second communication device 20 used by User 2. User 1 and User 2 use different languages ​​during the call. The first communication device 10 is connected to the simultaneous interpretation device 30. The audio content received by the first communication device 10 from the second communication device 20 is processed by the simultaneous interpretation device 30 and presented to User 1 in the form of text content. The audio content input by User 1 into the first communication device 10 is processed by the simultaneous interpretation device 30 and then transmitted to the second communication device 20.

[0055] To address the problems existing in related technologies, this application provides a simultaneous interpretation method. For example, please refer to... Figure 1 , Figure 1 This is a flowchart of a possible implementation of the simultaneous interpretation method provided in this application. The executing entity of this method can be the simultaneous interpretation device 30 described above, which includes a locally deployed AI service module. In the following method embodiments, for ease of description, only the simultaneous interpretation device 30 is described as the executing entity of each step. The above method may include at least one of the following steps (210-230): Step 210: Obtain the first audio data in the first language.

[0056] Optionally, during a call between the first and second communication devices, the first communication device receives second audio data in a second language sent by the simultaneous interpretation device via wired transmission. The second audio data is obtained by converting the first audio data in the first language. The first communication device then sends the second audio data to the second communication device.

[0057] In some embodiments, the first language refers to the language type corresponding to the original speech content that is used as input during simultaneous interpretation.

[0058] This application does not impose any special restrictions on the first language; for example, it can be Chinese, English, Japanese, Korean, or other natural languages.

[0059] In some embodiments, the first audio data refers to the raw audio content acquired by the first communication device or simultaneous interpretation device as input during simultaneous interpretation. Accordingly, the first audio data in the first language refers to the audio data corresponding to the audio content expressed in the first language.

[0060] The method of acquiring the first audio data in this application embodiment is not particularly limited. Optionally, it can be acquired by the audio acquisition module in the simultaneous interpretation device, or it can be transmitted from the first communication device to the simultaneous interpretation device.

[0061] Optionally, before the first audio data is transmitted from the first communication device to the simultaneous interpretation device, the first communication device may preprocess the raw audio data it has acquired (e.g., echo cancellation, noise suppression, or gain adjustment) to improve the accuracy of recognition and translation by the simultaneous interpretation device.

[0062] For example, the simultaneous interpretation device may be connected to an external microphone, through which the user inputs audio content expressed in a first language, and the simultaneous interpretation device uses the microphone to collect the first audio data.

[0063] For example, a user inputs audio content expressed in a first language into a first communication device. The first communication device performs noise reduction processing on the collected raw audio data to obtain first audio data, and sends the first audio data to a simultaneous interpretation device.

[0064] In some embodiments, step 210 includes at least one of the following steps (211-214): Step 211: Obtain the original audio data in the first language.

[0065] Alternatively, raw audio data refers to audio data that has been collected but has not undergone any processing.

[0066] Step 212: Downsample the original audio data to obtain the first intermediate data.

[0067] Optionally, downsampling refers to reducing the sampling rate of the original audio data to reduce the amount of audio data and reduce the computational burden of subsequent processing.

[0068] For example, the original audio data is downsampled to reduce the sampling rate of the original audio data from 48kHz to 16kHz to obtain the first intermediate data.

[0069] Step 213: Perform PCM (Pulse-Code Modulation) format conversion on the first intermediate data to obtain the second intermediate data.

[0070] Optionally, the first intermediate data can be converted to PCM format to ensure that the data format is consistent and improve the efficiency of subsequent processing.

[0071] This application does not impose special restrictions on the quantization bit width of the PCM format. For example, it can be 16 bits, 32 bits, etc.

[0072] For example, the first intermediate data is converted to 16-bit PCM format to obtain the second intermediate data represented in 16-bit PCM format.

[0073] Step 214: Based on the preset time batch length, the second intermediate data is divided into batches and packaged to obtain the first audio data.

[0074] Optionally, batch packaging refers to dividing the continuous second intermediate data according to a preset time batch length, and packaging the data within each time period.

[0075] Optionally, the time batch length refers to the preset time length used when packaging the second intermediate data in batches.

[0076] This application does not impose any special restrictions on the time batch length; for example, it can be 200ms.

[0077] Alternatively, downsampling, format conversion, and batch packaging can be implemented using the AudioWorklet or ScriptProcessorNode of the Web Audio API.

[0078] This application does not impose special restrictions on the audio preprocessing process. For example, noise reduction, filtering, gain control, and other processing can also be performed on the acquired raw audio data.

[0079] In related technologies, translation is usually performed using non-streaming methods after the raw audio data is acquired. This means that audio processing can only be performed after the complete acquisition of the raw audio data, which leads to a delay in translation output and a high latency in simultaneous interpretation.

[0080] In this embodiment of the application, by packaging the continuously received raw audio data in batches, the raw audio data can be processed batch by batch during the acquisition process, thereby reducing the waiting time of the acquisition process and improving the timeliness of simultaneous interpretation.

[0081] Step 220: Based on the AI ​​service module deployed locally, the first audio data is converted into second audio data in the second language in an offline state.

[0082] In some embodiments, the second language refers to the language type in which the audio content of the first language is translated or converted into the corresponding target language during simultaneous interpretation.

[0083] Optionally, the first language and the second language are different language types.

[0084] This application does not impose any special restrictions on the second language; for example, it can be Chinese, English, Japanese, Korean, or other natural languages.

[0085] In some embodiments, the second audio data refers to the target audio content output during simultaneous interpretation. Accordingly, the second audio data in the second language refers to the audio data corresponding to the audio content expressed in the second language.

[0086] In some embodiments, the AI ​​service module includes an audio processing model and a TTS service, and step 220 includes at least one of the following steps (221-222): Step 221: Based on the locally deployed audio processing model, the first audio data is identified and translated offline to obtain the first translation in the second language (i.e., the text information in the second language).

[0087] The locally deployed audio processing model has the ability to recognize ASR services and translate AST services. It can be understood that the audio processing model is a model that integrates speech recognition and translation functions.

[0088] In one aspect of this embodiment, recognizing the first audio data refers to using speech recognition technology to convert the collected audio content in the first language into corresponding text information in the first language. The purpose is to convert the audio data signal into processable text to facilitate translation and / or speech synthesis processing. Specifically, recognizing the first audio data can be achieved through the ASR service provided by the simultaneous interpretation device. This ASR service runs locally on the simultaneous interpretation device, reducing network transmission latency introduced during data transmission compared to the process of speech recognition via a cloud server in related technologies.

[0089] In one aspect of this embodiment, translating the first audio data refers to using machine translation technology to convert the identified text information in the first language into the corresponding text information in the second language. Specifically, translating the first audio data can be achieved through the AST service provided by the simultaneous interpretation device. This AST service runs locally on the simultaneous interpretation device. Compared with online translation methods in related technologies, this embodiment translates the text information in the first language locally on the simultaneous interpretation device, thus balancing the requirements of low latency and high security in simultaneous interpretation.

[0090] This application does not impose any special restrictions on the method by which the first audio data obtained by the simultaneous interpretation device is transmitted to the audio processing model. For example, the first audio data can be transmitted to the audio processing model via a WebSocket connection.

[0091] In some embodiments, after obtaining the first translation, the front-end application module of the simultaneous interpretation device can also display the first translation in the second language on the user interface display module and / or store it in a designated storage module, so that the user on the first communication device side can intuitively read the corresponding translation content.

[0092] Step 222: Based on the locally deployed TTS service, perform speech synthesis on the first translation to obtain the second audio data.

[0093] In one aspect of this embodiment, speech synthesis of the first translation refers to using TTS technology to convert the text content (first translation) in the second language into playable audio content (second audio data). Specifically, speech synthesis of the first translation can be achieved through the TTS service provided by the simultaneous interpretation device. This TTS service runs locally on the simultaneous interpretation device. Compared to speech generation on a cloud server in related technologies, this embodiment performs speech synthesis of the first translation locally on the simultaneous interpretation device, reducing processing latency and improving the fluency of simultaneous interpretation. Furthermore, local speech synthesis ensures stable audio output speed, thereby avoiding playback discontinuity or stuttering caused by network fluctuations or remote processing delays.

[0094] Optionally, the format of the second audio data obtained after speech synthesis is the same as that of the first audio data. For example, if the first audio data is represented in 16-bit PCM format, then the second audio data can also be represented in 16-bit PCM format.

[0095] As can be seen from the above, in this embodiment of the application, the process of converting the first audio data in the first language into the second audio data in the second language (e.g., ASR, AST, TTS, etc.) is implemented locally on the simultaneous interpretation device. There is no need to transmit the intermediate data in the conversion process to the cloud server. Compared with related online conversion technologies, this fundamentally eliminates the risk of data leakage, misuse, or eavesdropping by third parties, ensuring privacy and security. It is suitable for use scenarios with strict requirements for data privacy and security (e.g., business negotiations, confidential meetings, private personal conversations, etc.).

[0096] Meanwhile, compared to related technologies that transmit data to a cloud server and wait for conversion, the conversion process in this embodiment avoids network transmission latency introduced by online transmission. Especially when converting the first audio data obtained through batch encapsulation, the streaming processing mechanism minimizes local processing latency, achieving ultra-low latency simultaneous interpretation. During a conversation between the first and second communication devices, near-real-time interpretation makes cross-language communication more natural and fluent.

[0097] Furthermore, the online conversion methods using cloud servers in related technologies rely entirely on network connectivity. Simultaneous interpretation cannot be performed in environments with unstable or no network connection. Even semi-offline simultaneous interpretation methods (such as local ASR or TTS) are limited by network conditions, affecting translation quality and the number of languages ​​supported. In contrast, in this embodiment, the conversion process is completed completely offline. The AI ​​service module in the simultaneous interpretation device preloads offline models (such as locally deployed audio processing models, TTS services, etc.), ensuring continuous availability of the simultaneous interpretation process. The simultaneous interpretation process is unaffected by network fluctuations. In other words, regardless of the network environment (even a complete network outage), the simultaneous interpretation device can provide stable and high-quality simultaneous interpretation for the first communication device. Compared to related technologies, this enhances robustness, reduces dependence on external infrastructure, and broadens its applicability, making it suitable for outdoor, underground, confidential locations, or other network-restricted areas.

[0098] Step 230: The second audio data is sent to the first communication device in a call state via wired transmission. The second audio data is used to provide the second communication device that is in a call with the first communication device.

[0099] Optionally, step 230 includes: using an audio recording cable to send second audio data to a first communication device in a call state, wherein the simultaneous interpretation device is connected to the first communication device via the audio recording cable.

[0100] By leveraging the bidirectional and efficient transmission characteristics of the audio recording cable, the simultaneous interpretation device can simultaneously interpret the audio content of the first communication device to the second communication device, while also transmitting the audio content of the second communication device to the simultaneous interpretation device in real time. This allows the simultaneous interpretation method provided in this application embodiment to seamlessly integrate into the existing communication ecosystem, ensuring low latency and high stability in bidirectional simultaneous interpretation.

[0101] In some embodiments, the simultaneous interpretation device may also transmit the second audio data to the first communication device in a call state via wireless transmission, such as Bluetooth, WiFi (Wireless Fidelity), NFC (Near Field Communication), etc.

[0102] Optionally, when the audio data (e.g., second audio data) is transmitted wirelessly, the audio data is encrypted. This application does not impose special limitations on the encryption method of the audio data; for example, it can be symmetric encryption, asymmetric encryption, etc.

[0103] Accordingly, when audio data (e.g., second audio data) is transmitted via wired transmission, the audio data can be the raw audio data, i.e., unencrypted data.

[0104] By transmitting encrypted audio data wirelessly, data can be prevented from being tampered with or leaked during wireless transmission, thereby improving the security of data transmission.

[0105] In one aspect of this embodiment, wired transmission refers to a communication method that uses a physical connection cable to transmit second audio data from a simultaneous interpretation device to a first communication device, such as an audio recording cable, a USB (Universal Serial Bus) cable, an HDMI (High-Definition Multimedia Interface) cable, or other physical connection cables.

[0106] This application does not impose special restrictions on the interface of the physical connection cable. For example, it can be a Type-C interface, a USB interface, etc.

[0107] It should be understood that transmitting the second audio data between the simultaneous interpretation device and the first communication device via wired transmission, so that the second audio data is transmitted on a physical medium, not only has high transmission quality, but also avoids the risk of data leakage that may be introduced by transmitting via wireless network or Internet, thereby improving the information security in the simultaneous interpretation process.

[0108] In some embodiments, prior to step 230, the audio routing control module in the simultaneous interpretation device may determine the device to which the second audio data is to be transmitted. Alternatively, the audio routing control module may determine the output device for the audio content corresponding to the second audio data.

[0109] Optionally, in response to a selection instruction provided by the user, second audio data is sent to the output device corresponding to the selection instruction.

[0110] This application does not impose any special restrictions on the way a user selects an output device in a simultaneous interpretation device. Optionally, the user interface in the simultaneous interpretation device displays one or more candidate output devices, and in response to the user selecting one of the one or more candidate output devices, second audio data is sent to the output device corresponding to the selection indication.

[0111] This application does not impose any special restrictions on the type of audio output device. For example, it can be headphones, speakers, or internal speakers of the first communication device.

[0112] For example, the second audio data is routed to the selected audio output device via the setSinkId API.

[0113] In related technologies, the translated audio content is usually played through the device's own speaker, making it difficult to input the translated audio content into the call link of an external communication device. This makes it impossible to achieve synchronous transmission of the translated audio content, affecting the real-time nature of simultaneous interpretation and the smoothness of the call.

[0114] In this embodiment, the desired audio output device is determined by software control (e.g., the setSinkId API). Users can freely switch between audio output devices according to their needs (e.g., local or remote playback). This flexibility is enhanced because the conversion process between the first and second audio data can be flexibly switched without terminating the connection. In particular, the second audio data provided by the simultaneous interpretation device can be directly transmitted to the call link of the second communication device with high transmission quality, making it especially suitable for scenarios such as teleconferences and international calls.

[0115] It is worth noting that since the first audio data is encapsulated and converted based on a streaming processing mechanism, the second audio data can also be continuously generated and sequentially transmitted and routed to the audio output device without waiting for the complete acquisition and processing of the original audio data, thus ensuring the real-time performance and stability of the simultaneous interpretation process.

[0116] It should be understood that the above-described process of transmitting the second audio data via wired transmission is merely an illustrative example. Any transmission method that can characterize the transmission of the second audio data to the user-selected audio output device falls within the protection scope of this application. For example, the second audio data can also be transmitted wirelessly (e.g., via Bluetooth).

[0117] In some embodiments, prior to step 230, the simultaneous interpretation method provided in this application further includes: sending preheating audio data of a preset duration to a first communication device.

[0118] Optionally, the audio playback module in the simultaneous interpretation device preheats the audio pipeline based on the second audio data. The audio pipeline refers to the link that carries audio data from reception, buffering, processing to output playback. By sending preheated audio data to the first communication device in advance, a stable audio transmission channel can be established before sending the second audio data. This allows the audio link, buffer, and audio output device to enter a working state in advance, thereby avoiding interruptions or frame drops when the audio content corresponding to the second audio data begins playing. This improves the stability of audio playback during simultaneous interpretation, ensuring smooth playback.

[0119] In some embodiments, prior to step 230, the simultaneous interpretation method provided in this application further includes: gradually increasing the volume of the second audio data within a preset time period before the audio content corresponding to the second audio data begins to play, until the target volume is reached.

[0120] For example, within the first 50ms before the audio content corresponding to the second audio data begins to play, the volume corresponding to the second audio data is gradually increased to the target volume.

[0121] In this embodiment of the application, by performing a fade-in process in advance, the sudden change in volume when the audio content corresponding to the second audio data begins to play can be avoided, thereby improving the smoothness of audio playback during simultaneous interpretation.

[0122] In some embodiments, using the RxJS reactive programming framework to stream and process audio data (e.g., first audio data, second audio data) can enable parallel execution of audio acquisition, processing, and output processes.

[0123] By streaming and processing audio data, and optimizing the audio data processing pipeline (including batch encapsulation of raw audio data, pre-sending warm-up audio data, and pre-fading processing), the latency introduced by waiting for complete raw audio data can be reduced, thus lowering the processing latency from input to output. This meets the requirements of simultaneous interpretation for low-latency response, enabling the raw audio data input to the simultaneous interpretation device to be translated to the first communication device almost in real time, thereby making cross-language communication between the first and second communication devices more natural.

[0124] The simultaneous interpretation method provided in this application converts first audio data in a first language into second audio data in a second language using a simultaneous interpretation device, and then transmits the second audio data to a first communication device in a call state via wired transmission. This enables the first communication device to provide the second audio data in the second language to the second communication device with which it is in a call. The conversion process from first audio data to second audio data is completed on the simultaneous interpretation device, without the need to upload to a cloud server or other devices, thus reducing the latency introduced by data transmission. At the same time, the second audio data is transmitted to the first communication device via wired transmission, which takes into account both the requirements of audio data transmission quality and low-latency transmission.

[0125] For example, to illustrate more specifically and clearly the detailed process of simultaneously interpreting first audio data in a first language into second audio data in a second language and sending it to a first communication device, please refer to [reference needed]. Figure 3 , Figure 3 This is a flowchart of a simultaneous interpretation process for spoken Chinese, provided in one possible implementation of this application. This simultaneous interpretation process is used to translate Chinese into English. The process may include at least one of steps 310-390. It is worth noting that the specific processes mentioned below are merely illustrative examples.

[0126] Step 310: The user speaks Chinese into the microphone. Specifically, the microphone is an external audio input device connected to the simultaneous interpretation equipment, and the user's mobile phone 1 is currently in a conversation with mobile phone 2.

[0127] Step 320, Audio Acquisition Module. Specifically, when the user speaks Chinese into the microphone, the audio acquisition module acquires raw audio data in Chinese through the microphone.

[0128] Step 330, Audio Preprocessing Module. Specifically, the acquired raw audio data is downsampled, converted to PCM format, and batched and packaged in fixed 200ms batches to obtain the first audio data.

[0129] Step 340, Local End-to-End Speech Translation Model: Chinese speech to English text. Specifically, the first audio data (in PCM format) in Chinese is sent to the local end-to-end speech translation model (i.e., the audio processing model) via a WebSocket connection. This local end-to-end speech translation model integrates speech recognition and translation functions, and can recognize and translate Chinese speech to obtain English text in English.

[0130] Step 350, Front-end application module: Text processing. Specifically, the front-end application module displays English text in the user interface display module.

[0131] Step 360, Local TTS Service: English Speech Synthesis. Specifically, using a locally deployed TTS service, English text is synthesized into audio data corresponding to English speech. The audio data corresponding to the English speech is in PCM format.

[0132] Step 370, Audio Playback Module. Specifically, the audio playback module preheats the audio pipeline based on the audio data corresponding to the synthesized English speech to ensure that the English speech to be played can be played smoothly.

[0133] Step 380, Audio Routing Control Module. Specifically, in response to the user selecting mobile phone 1 from one or more candidate audio output devices, the audio routing control module routes the English voice stream to mobile phone 1.

[0134] Step 390: Output to mobile phone 1 via audio recording cable. Specifically, English speech is sent from the simultaneous interpretation device to mobile phone 1 via audio recording cable, so that mobile phone 1 can provide it to mobile phone 2 for playback.

[0135] The technical solution for simultaneous Chinese-to-English speech translation provided in this application not only achieves the conversion of Chinese to English speech in a completely offline state, but also reduces the latency of the simultaneous interpretation process by batch encapsulating the acquired raw audio data and combining it with the RxJS reactive programming framework. This allows the Chinese speech transmitted by the user through the microphone of the simultaneous interpretation device to be translated to mobile phone 1 almost in real time. Simultaneously, by preheating the audio channel before the English speech begins playing, pre-playing preheated audio data improves the smoothness of the simultaneous interpretation process. Since the simultaneous interpretation process is completely offline, data privacy and security are fundamentally guaranteed. Using an audio recording cable for audio data transmission between the simultaneous interpretation device and mobile phone 1 not only ensures data security but also enables high-quality audio data transmission.

[0136] In addition, based on the audio routing control module, users can select the desired audio output device, thereby accurately sending English speech to the expected audio output device, simplifying the operation process.

[0137] In some embodiments, the simultaneous interpretation method provided in this application further includes at least one of the following steps (240-260). It is worth noting that steps (240-260) and steps (210-230) do not constitute a limitation on the order of execution.

[0138] Step 240: Receive third audio data in a second language sent by the first communication device via wired transmission. The third audio data is sent by the second communication device to the first communication device.

[0139] Optionally, the first communication device receives third audio data in a second language sent by the second communication device.

[0140] Accordingly, the first communication device sends third audio data to the simultaneous interpretation device via wired transmission. The third audio data is used to convert into a second translation in the first language for display, or to convert into a fourth audio data in the first language for playback.

[0141] In some embodiments, the third audio data refers to the audio content acquired by a second communication device during a conversation with the first communication device, which is used as input in the simultaneous interpretation process. Correspondingly, the third audio data in the second language refers to the audio data corresponding to the audio content expressed in the second language.

[0142] Optionally, before the third audio data is transmitted from the second communication device to the first communication device or the simultaneous interpretation device, the first or second communication device may preprocess it (e.g., echo cancellation, noise suppression, or gain adjustment) to improve the accuracy of recognition and translation by the simultaneous interpretation device.

[0143] Optionally, the third audio data is transmitted from the first communication device to the simultaneous interpretation device via wired transmission.

[0144] Optionally, step 240 includes: receiving third audio data sent by the first communication device using an audio recording cable. The bidirectional and efficient transmission characteristics of the audio recording cable enable the simultaneous interpretation device to simultaneously perform: simultaneous interpretation of the audio content (first audio data) of the first communication device to the second communication device, and real-time transmission of the audio content (third audio data) of the second communication device back to the simultaneous interpretation device. This allows the simultaneous interpretation method provided in this embodiment to seamlessly integrate into the existing communication ecosystem, ensuring low latency and high stability in bidirectional simultaneous interpretation.

[0145] In one aspect of this embodiment, wired transmission refers to a communication method that uses a physical connection cable to transmit third audio data from a first communication device to a simultaneous interpretation device, such as an audio recording cable, a USB cable, an HDMI cable, or other physical connection cables.

[0146] This application does not impose special restrictions on the interface of the physical connection cable. For example, it can be a Type-C interface, a USB interface, etc.

[0147] It should be understood that transmitting third audio data between the simultaneous interpretation equipment and the first communication equipment via wired transmission, allowing the third audio data to be transmitted over a physical medium, not only provides higher transmission quality but also avoids the risk of data leakage that may be introduced by transmitting via wireless networks or the Internet, thereby improving information security during the simultaneous interpretation process.

[0148] Step 250: Based on the AI ​​service module deployed locally, the third audio data is converted into a second translation in the first language while offline.

[0149] Optionally, step 250 includes at least one of the following steps (251-252): Step 251: Preprocess the third audio data to obtain the third intermediate data.

[0150] This application does not impose special restrictions on the audio preprocessing process. For example, noise reduction, filtering, gain control, and other processing can also be performed on the acquired raw audio data.

[0151] In some embodiments, step 251 includes at least one of the following steps (2511-2513): Step 2511: Downsample the third audio data to obtain the fourth intermediate data.

[0152] Optionally, downsampling refers to reducing the sampling rate of the third audio data to reduce the amount of audio data and reduce the computational burden of subsequent processing.

[0153] Step 2512: Perform PCM format conversion on the fourth intermediate data to obtain the fifth intermediate data.

[0154] Optionally, the fourth intermediate data can be converted to PCM format to ensure data format consistency and improve subsequent processing efficiency.

[0155] This application does not impose special restrictions on the quantization bit width of the PCM format. For example, it can be 16 bits, 32 bits, etc.

[0156] For example, the fourth intermediate data is converted to 32-bit PCM format to obtain the fifth intermediate data represented in 32-bit PCM format.

[0157] Step 2513: Based on the preset time batch length, the fifth intermediate data is divided into batches and packaged to obtain the third intermediate data.

[0158] Optionally, batch packaging refers to dividing the continuous fifth intermediate data into segments according to a preset time batch length, and packaging the data within each time segment.

[0159] Optionally, the time batch length refers to the preset time length used when packaging the fifth intermediate data in batches.

[0160] This application does not impose any special restrictions on the time batch length; for example, it can be 200ms.

[0161] Alternatively, downsampling, format conversion, and batch packaging can be implemented using the AudioWorklet or ScriptProcessorNode of the Web Audio API.

[0162] Step 252: Based on the locally deployed audio processing model, the third intermediate data is identified and translated to obtain the second translation.

[0163] By recognizing and translating third-party intermediate data using a locally running audio processing model, compared to online translation methods in related technologies, it is possible to prevent data leakage, misuse, or eavesdropping by third parties. At the same time, it can also avoid network latency introduced by online transmission, thus balancing data security and simultaneous interpretation efficiency.

[0164] Step 260: Display the second translation.

[0165] Optionally, the second translation is presented through the user interface display module in the simultaneous interpretation device, so that the user can obtain the corresponding translation while listening to the voice of the second communication device (third audio data).

[0166] In some embodiments, the AI ​​service module in the simultaneous interpretation device further includes a TTS service, and the simultaneous interpretation method provided in this application may further include step 270. It is worth noting that steps (240-260), steps (210-230), and step 270 do not constitute a limitation on the execution order.

[0167] Step 270: Play the third audio data. By playing the third audio data, the audio content of the other end of the call can be avoided due to simultaneous interpretation, thereby improving the reliability of the call.

[0168] Optionally, the audio output device for playing the third audio data can be determined by the audio routing control module. Optionally, in response to a selection instruction provided by the user, the third audio data is sent to the output device corresponding to the selection instruction.

[0169] This application does not impose any special restrictions on how users select output devices in simultaneous interpretation equipment. Optionally, the user interface in the simultaneous interpretation equipment displays one or more candidate output devices, and in response to the user selecting one of the one or more candidate output devices, third audio data is sent to the output device corresponding to the selection indication.

[0170] This application does not impose any special restrictions on the type of audio output device. For example, it can be headphones, speakers, or internal speakers of the first communication device.

[0171] For example, third-party audio data can be routed to the selected audio output device via the setSinkId API.

[0172] In some embodiments, the simultaneous interpretation method provided in this application further includes step 280. It is worth noting that steps (240-260), (210-230), 270, and 280 do not constitute a limitation on the order of execution.

[0173] Step 280: Based on the locally deployed TTS service, the second translation is processed into speech to obtain and play the fourth audio data.

[0174] Optionally, the fourth audio data can be played through the audio playback module.

[0175] Optionally, the format of the fourth audio data obtained after speech synthesis is the same as that of the third audio data. For example, if the third audio data is represented in 16-bit PCM format, then the fourth audio data can also be represented in 16-bit PCM format.

[0176] Simultaneous interpretation equipment converts the third audio data transmitted from the second communication device, which is communicating with the first communication device, into a second translation in the first language. This allows users to view the corresponding translation in real time while listening to the third audio data transmitted from the second communication device. The simultaneous interpretation process is performed locally without the need for an internet connection, which reduces latency while also ensuring data privacy and security.

[0177] For example, to illustrate more specifically and clearly the detailed process of simultaneous interpretation of third audio data in a second language, please refer to [reference needed]. Figure 4 , Figure 4This is a flowchart of a simultaneous interpretation process for incoming English speech provided in one possible implementation of this application. This simultaneous interpretation process is used to convert English to Chinese. The process may include at least one of steps 410-440. It is worth noting that the specific processes mentioned below are merely illustrative examples.

[0178] Step 410: Input English voice from mobile phone 2 via audio recording cable. Specifically, the user's mobile phone 1 is in a call with mobile phone 2, and the English voice is captured by the audio acquisition module of mobile phone 2.

[0179] Step 420, audio receiving module. Specifically, English speech is sent from mobile phone 1 to the simultaneous interpretation device via an audio recording cable. After step 420, steps 430 or 440 can be executed.

[0180] Step 430 involves simultaneous interpretation of the third audio data, which may specifically include at least one of steps 431 to 434.

[0181] Step 431, audio preprocessing module. Specifically, the acquired English speech is downsampled, converted to PCM format, and batched and packaged in fixed 200ms batches to obtain the third intermediate data.

[0182] Step 432, Local End-to-End Speech Translation Model: English speech is converted to Chinese text. Specifically, third-party intermediate data (in PCM format) in English is sent to the local end-to-end speech translation model (i.e., the audio processing model) via a WebSocket connection. This local end-to-end speech translation model integrates speech recognition and translation functions, and can recognize and translate English speech to obtain Chinese text.

[0183] Step 433, Front-end application module: Text processing. Specifically, the front-end application module transmits Chinese text to the user interface display module.

[0184] Step 434, User interface display module: Display Chinese translation.

[0185] Step 440: Play English text through the phone's speaker.

[0186] The technical solution provided in this application for simultaneous interpretation of incoming English speech into Chinese speech uses a simultaneous interpretation device to convert the English speech received from mobile phone 2 into Chinese text, so that users can view the corresponding translation in real time while listening to the English speech received from mobile phone 2. The simultaneous interpretation process is performed locally without the need for a network connection, which reduces latency while also ensuring data privacy and security.

[0187] Furthermore, by batch encapsulating the acquired English speech and combining it with the RxJS reactive programming framework, the latency of the simultaneous interpretation process is reduced, enabling the Chinese speech transmitted by the user through the microphone of the simultaneous interpretation device to be translated to the user interface display module of the simultaneous interpretation device almost in real time.

[0188] For example, to illustrate more specifically and clearly the detailed process of the simultaneous interpretation method provided in this application, please refer to [reference needed]. Figure 5 , Figure 5 This is a flowchart illustrating a simultaneous interpretation process for voicing Chinese speech and receiving English speech, as provided in one possible implementation of this application. This simultaneous interpretation process is used to achieve Chinese-English translation. The process may include at least one step from steps 511 to 517 and at least one step from steps 521 to 526. It is worth noting that the specific processes described below are merely illustrative examples.

[0189] In this example, both Device B and Device C are external communication devices, and Device A is a simultaneous interpretation device. Device B and Device C are in the middle of a conversation. Steps 511-517 relate to the simultaneous interpretation of Chinese speech received from Device B to Device C, while steps 521-526 relate to the simultaneous interpretation of English speech received from Device C to Device B. Steps 511-517 and 521-526 are explained below.

[0190] Step 511: The user speaks Chinese into the microphone. Specifically, the microphone is an external audio input device connected to the simultaneous interpretation equipment.

[0191] Step 512, Local ASR Model: Chinese Recognition. Specifically, using the ASR model deployed locally on the simultaneous interpretation device, the Chinese text corresponding to the Chinese speech is recognized.

[0192] Step 513, Local AST Model: Chinese Text to English Text. Specifically, the AST model deployed locally on the simultaneous interpretation equipment is used to translate Chinese text into English text.

[0193] Step 514, Local TTS Model: English Speech Synthesis. Specifically, using the TTS model deployed locally on the simultaneous interpretation device, the English text is synthesized to obtain English speech.

[0194] Step 515, Software Selection: Audio Pair Output. Specifically, the desired audio playback method for the user's English speech is determined to be audio pair output, and the English speech is output via the audio pair output through the audio routing control module.

[0195] Step 516, audio recording cable. Specifically, the English speech is transmitted from the simultaneous interpretation equipment to device B via the audio recording cable.

[0196] Step 517: Device B receives English audio. Specifically, after receiving the English audio sent by Device B, Device C plays the audio so that the user of Device C can hear the simultaneously interpreted audio.

[0197] Step 521: Device B receives English voice messages from Device C. Specifically, Device B is in a conversation with Device C, and the English voice messages are captured by the audio acquisition module on Device C's side.

[0198] Step 522, Audio recording cable. Specifically, the English speech is transmitted from device B to the simultaneous interpretation device via the audio recording cable.

[0199] Step 523: Device A directly plays the English audio.

[0200] Step 524, Local ASR Model: English Recognition. Specifically, the ASR model deployed locally on the simultaneous interpretation device is used to recognize the English text corresponding to the English speech.

[0201] Step 525, Local AST Model: English Speech to Chinese Text. Specifically, the English text is translated into Chinese text using the AST model deployed locally on the simultaneous interpretation equipment.

[0202] Step 526: The application interface of Device A displays Chinese text.

[0203] The simultaneous interpretation solution provided in this application, which handles both outgoing Chinese and incoming English speech, not only achieves Chinese-to-English and vice versa in a completely offline state, but also batches and encapsulates the acquired Chinese or English speech. Combined with the RxJS reactive programming framework, this reduces latency in the simultaneous interpretation process. Furthermore, by pre-warming the audio pipeline before speech playback, pre-playing pre-warmed audio data improves the smoothness of the simultaneous interpretation process. Since the simultaneous interpretation process is completely offline, both data security and efficiency can be balanced.

[0204] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0205] For example, please refer to Figure 6 , Figure 6 This is a structural block diagram of a simultaneous interpretation apparatus provided in one possible implementation of this application. The apparatus 600 may include: Module 610 is used to acquire the first audio data of the first language; Processing module 620 is used to convert the first audio data into second audio data in a second language in an offline state, based on the AI ​​service module deployed locally; The sending module 630 is used to send the second audio data to the first communication device in a call state via wired transmission. The second audio data is used to provide the second communication device that is in a call with the first communication device.

[0206] In some embodiments, the AI ​​service module includes an audio processing model and a text-to-speech (TTS) service; the processing module 620 is configured to: recognize and translate the first audio data based on the locally deployed audio processing model to obtain a first translation in a second language; and perform speech synthesis on the first translation based on the locally deployed TTS service to obtain second audio data.

[0207] In some embodiments, the acquisition module 610 is configured to: acquire raw audio data in a first language; perform downsampling processing on the raw audio data to obtain first intermediate data; perform PCM format conversion processing on the first intermediate data to obtain second intermediate data; and encapsulate the second intermediate data in batches based on a preset time batch length to obtain the first audio data.

[0208] In some embodiments, the sending module 630 is configured to: send second audio data to a first communication device in a call state using an audio recording cable, wherein the simultaneous interpretation device is connected to the first communication device via the audio recording cable.

[0209] In some embodiments, before sending the second audio data to the first communication device in a call state via wired transmission, the sending module 630 is further configured to: send preheating audio data of a preset duration to the first communication device.

[0210] In some embodiments, the device 600 further includes a receiving module (not shown) and / or a display module (not shown).

[0211] The receiving module is used to: receive third audio data in a second language sent by the first communication device via wired transmission, wherein the third audio data is sent by the second communication device to the first communication device; the processing module 620 is used to: convert the third audio data into a second translation in the first language in an offline state based on the AI ​​service module deployed locally; and the display module is used to: display the second translation.

[0212] In some embodiments, the AI ​​service module includes an audio processing model, and a processing module 620 is configured to: preprocess the third audio data to obtain third intermediate data; and, based on the locally deployed audio processing model, identify and translate the third intermediate data to obtain a second translation.

[0213] In some embodiments, the AI ​​service module further includes a TTS service, and the processing module 620 is further configured to: play third audio data; or, based on the locally deployed TTS service, perform speech synthesis on the second translation to obtain and play fourth audio data.

[0214] For example, please refer to Figure 7 , Figure 7 This is a structural block diagram of a simultaneous interpretation apparatus provided in one possible implementation of this application. The apparatus 700 may include: The receiving module 710 is used to receive second audio data in a second language sent by the simultaneous interpretation device via wired transmission during a call between the first communication device and the second communication device. The second audio data is obtained by converting the first audio data in the first language. The transmitting module 720 is used to transmit second audio data to the second communication device.

[0215] In some embodiments, the receiving module 710 is further configured to: receive third audio data in a second language sent by the second communication device; the sending module 720 is further configured to: send the third audio data to the simultaneous interpretation device via wired transmission, wherein the third audio data is used to convert into a second translation in a first language for display, or to convert into fourth audio data in a first language for playback.

[0216] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0217] The simultaneous interpretation device provided in this application converts acquired first audio data in a first language into second audio data in a second language, and sends the second audio data to a first communication device in a call state via wired transmission. This enables the first communication device to provide the second audio data in the second language to the second communication device with which it is in a call. The conversion process from first audio data to second audio data is completed on the AI ​​service module deployed locally on the simultaneous interpretation device, without uploading to a cloud server or other devices, thus avoiding audio content leakage, ensuring data security, and reducing latency introduced by data transmission. In addition, the second audio data is sent to the first communication device via wired transmission, which takes into account both the requirements of audio data transmission quality and low-latency transmission.

[0218] In addition, the third audio data transmitted by the second communication device is converted into a second translation in the first language by a simultaneous interpretation device, so that users can view the corresponding translation in real time while listening to the third audio data transmitted by the second communication device. The simultaneous interpretation process is performed locally without the need for a network connection, which reduces latency and also ensures data privacy and security.

[0219] For example, please refer to Figure 8 , Figure 8 This is a structural block diagram of a computer device provided in one possible implementation of this application. The computer device 800 can be any electronic device with data computing, processing, and storage functions. The computer device 800 can be used to implement the simultaneous interpretation method provided in the above embodiments.

[0220] Typically, computer device 800 includes a processor 810 and a memory 820.

[0221] Processor 810 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 810 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA, and PLA (Programmable Logic Array). Processor 810 may also include a main processor and a coprocessor. The main processor, also known as the CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 810 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 810 may also include an AI processor, which can be used to handle computational operations related to machine learning.

[0222] The memory 820 may include one or more computer-readable storage media, which may be non-transitory. The memory 820 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 820 are used to store a computer program configured to be executed by one or more processors to implement the simultaneous interpretation method described above.

[0223] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the computer device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0224] In an illustrative embodiment, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor of a computer device, implements the simultaneous interpretation method described above. Optionally, the computer-readable storage medium may be a ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, or optical data storage device, etc.

[0225] In an exemplary embodiment, a chip is also provided, the chip including programmable logic circuitry and / or program instructions stored in a computer-readable storage medium. A processor of a computer device reads the programmable logic circuitry and / or program instructions from the computer-readable storage medium, and executes the programmable logic circuitry and / or program instructions, causing the computer device to perform the simultaneous interpretation method described above.

[0226] In an exemplary embodiment, a computer program product is also provided, which includes a computer program that is loaded and executed by a processor to implement the simultaneous interpretation method described above.

[0227] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.

[0228] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for simultaneous interpretation, characterized in that, The method, applied to a simultaneous interpretation device, wherein the simultaneous interpretation device includes a locally deployed artificial intelligence (AI) service module, comprises: Obtain the first audio data in the first language; Based on the AI ​​service module deployed locally, the first audio data is converted into second audio data in a second language in an offline state; The second audio data is transmitted via wired means to a first communication device that is in a call state. The second audio data is used to provide to a second communication device that is in a call with the first communication device.

2. The method according to claim 1, characterized in that, The AI ​​service module includes an audio processing model and a text-to-speech (TTS) service. The process of converting the first audio data into second audio data in a second language using the locally deployed AI service module in an offline state includes: Based on the locally deployed audio processing model, the first audio data is identified and translated to obtain the first translation in the second language; Based on the locally deployed TTS service, speech synthesis is performed on the first translation to obtain the second audio data.

3. The method according to claim 1, characterized in that, The acquisition of the first audio data in the first language includes: Obtain the raw audio data in the first language; The original audio data is downsampled to obtain the first intermediate data; The first intermediate data is converted to PCM format to obtain the second intermediate data. Based on a preset time batch length, the second intermediate data is divided into batches and packaged to obtain the first audio data.

4. The method according to claim 1, characterized in that, The step of sending the second audio data to a first communication device in a call state via wired transmission includes: Using an audio recording cable, the second audio data is sent to the first communication device which is in a call state, and the simultaneous interpretation device is connected to the first communication device through the audio recording cable.

5. The method according to claim 1, characterized in that, Before sending the second audio data to the first communication device in a call state via wired transmission, the method further includes: Send preheating audio data of a preset duration to the first communication device.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: The third audio data in the second language sent by the first communication device is received via wired transmission. The third audio data is sent by the second communication device to the first communication device. Based on the AI ​​service module deployed locally, the third audio data is converted into a second translation in the first language in an offline state; The second translation is displayed.

7. The method according to claim 6, characterized in that, The AI ​​service module includes an audio processing model; The process of converting the third audio data into a second translation in the first language using the locally deployed AI service module in an offline state includes: The third audio data is preprocessed to obtain the third intermediate data; Based on the locally deployed audio processing model, the third intermediate data is identified and translated to obtain the second translation.

8. The method according to claim 7, characterized in that, The AI ​​service module also includes TTS service; The method further includes: Play the third audio data; or, Based on the locally deployed TTS service, the second translation is synthesized into speech to obtain and play the fourth audio data.

9. A method for simultaneous interpretation, characterized in that, Applied to a first communication device, the method includes: During a conversation between the first and second communication devices, the second audio data in a second language sent by the simultaneous interpretation device is received via wired transmission. The second audio data is obtained by converting the first audio data in the first language. The second audio data is sent to the second communication device.

10. The method according to claim 9, characterized in that, The method further includes: Receive third audio data in the second language sent by the second communication device; The third audio data is transmitted to the simultaneous interpretation device via wired transmission. The third audio data is used to convert into a second translation in the first language for display, or to convert into a fourth audio data in the first language for playback.

11. A simultaneous interpretation device, characterized in that, The device includes: The acquisition module is used to acquire the first audio data in the first language; The processing module is used to convert the first audio data into second audio data in a second language in an offline state, based on the AI ​​service module deployed locally. The sending module is used to send the second audio data to a first communication device in a call state via wired transmission. The second audio data is used to provide to a second communication device that is in a call with the first communication device.

12. A simultaneous interpretation device, characterized in that, The device includes: The receiving module is used to receive second audio data in a second language sent by the simultaneous interpretation device via wired transmission during a call between the first communication device and the second communication device. The second audio data is obtained by converting the first audio data in the first language. The sending module is used to send the second audio data to the second communication device.

13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 8, or to implement the method as claimed in claim 9 or 10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that is executed by a processor to implement the method as claimed in any one of claims 1 to 8, or to implement the method as claimed in claim 9 or 10.

15. A computer program product, characterized in that, The computer program product includes a computer program executed by a processor to implement the method as claimed in any one of claims 1 to 8, or to implement the method as claimed in claim 9 or 10.