Multi-language synchronous broadcasting system and method based on Auracast

By using the Auracast-based multilingual synchronous broadcasting system, real-time and on-demand broadcasting of multilingual content was achieved, solving the multilingual synchronization problem in existing technologies, improving user experience, and optimizing resource utilization.

CN121968023APending Publication Date: 2026-05-01SHENZHEN FENGHEYUAN TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN FENGHEYUAN TECH
Filing Date
2026-01-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing broadcasting systems struggle to achieve synchronous, real-time, and on-demand multilingual broadcasting in cross-border communication and multilingual environments, resulting in poor user experience, wasted resources, and system complexity.

Method used

The system employs a multilingual synchronous broadcasting system based on Auracast. The raw audio signal is acquired and processed by the sound pickup and processing module. The energy detection module triggers translation requests and uploads the speech segments to the cloud translation module. The cloud translation module performs real-time recognition, translation, and conversion. The broadcast transmission module allocates independent BIS channels for different language speech streams and broadcasts synchronously via the network time protocol. The listener terminal module responds to the user's selection of the target language to play.

Benefits of technology

It enables synchronous, real-time, and on-demand broadcasting of multilingual content, avoiding resource waste and management complexity caused by multiple transmitters, and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121968023A_ABST
    Figure CN121968023A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-language synchronous broadcasting system and method based on Auracast. The method comprises the steps that a pickup processing module is configured to output a voice signal to be translated; the energy detection module is configured to trigger a translation request when the voice segment starts; the cloud translation module is configured to generate a target language voice stream in response to the translation request; the broadcast transmitting module is configured to distribute an independent BIS channel for a target language voice stream and an original voice stream through a wireless broadcast signal based on an Auracast protocol, and add a network time-based protocol for each voice data packet in each voice stream; and the audience terminal module is configured to respond to a selection request of a user for a target language and output audio according to the voice data packet in the target BIS channel corresponding to the target language. Therefore, synchronous, real-time and on-demand broadcasting of multi-language content is realized through real-time pickup, segmented uploading, cloud streaming translation, synchronous broadcasting and on-demand selection playing of a user.
Need to check novelty before this filing date? Find Prior Art

Description

A Multilingual Synchronous Broadcasting System and Method Based on Auracast Technical Field

[0001] This application relates to the technical field of broadcasting, and more specifically, to a multilingual synchronous broadcasting system and method based on Auracast. Background Technology

[0002] In related technologies, providing understandable voice information to listeners of different native languages ​​has always been a key requirement in cross-border communication scenarios and multilingual environments. In broadcasting systems, traditional wireless simultaneous interpretation has wide coverage but lacks a personalized experience, while Bluetooth sharing is personalized but has narrow coverage. Moreover, transmitters usually adopt a single-language transmission mode or provide multilingual content through frequency bands and time periods. Adding languages ​​usually means adding independent wireless transmission resources, and there is a delay introduced by the translation process. Furthermore, existing distribution methods cannot guarantee lip-sync or content synchronization between multilingual streams. Although these methods can achieve limited multilingual coverage, they cannot achieve multilingual synchronization, real-time and on-demand broadcasting, resulting in poor user experience, wasted resources and system complexity. Summary of the Invention

[0003] In view of the above problems, this application proposes a multilingual synchronous broadcasting system and method based on Auracast, which can realize multilingual synchronous, real-time and on-demand broadcasting, thereby effectively improving user experience, reducing resource waste and reducing system complexity.

[0004] In a first aspect, embodiments of this application provide a multilingual synchronous broadcasting system based on Auracast. The system includes: a sound pickup and processing module configured to acquire and process raw audio signals from the environment to obtain a speech signal to be translated; an energy detection module configured to trigger a translation request when a speech segment in the speech signal to be translated begins, and to upload the speech segment to a cloud translation module in a streaming manner, and to stop uploading when the speech segment ends; a cloud translation module configured to respond to the translation request, and to recognize, translate, and convert the original speech stream corresponding to the speech segment to generate at least one target language speech stream; a broadcast transmission module configured to allocate independent BIS channels for at least one target language speech stream and the original speech stream via a wireless broadcast signal based on the Auracast protocol, and to add a network time protocol-based addition to each speech data packet in each speech stream; and a listener terminal module configured to respond to a user's request to select a target language, determine the target BIS channel corresponding to the target language, and output audio in the target language based on the speech data packets in the target BIS channel.

[0005] Secondly, embodiments of this application also provide a multilingual synchronous broadcasting method based on Auracast. The method includes: acquiring and processing raw audio signals in the environment to obtain a speech signal to be translated; triggering a translation request when a speech segment in the speech signal to be translated begins, and uploading the speech segment in a streaming manner, and stopping the uploading when the speech segment ends; responding to the translation request, recognizing, translating, and converting the original speech stream corresponding to the speech segment to generate at least one target language speech stream; based on the Auracast protocol, allocating independent BIS channels for at least one target language speech stream and the original speech stream through a wireless broadcast signal, and adding a network time protocol-based addition to each speech data packet in each speech stream; responding to a user's request to select a target language, determining the target BIS channel corresponding to the target language, and outputting the audio of the target language based on the speech data packets in the target BIS channel.

[0006] The technical solution provided in this application includes a system comprising: a sound pickup and processing module configured to acquire and process raw audio signals in the environment to obtain a speech signal to be translated; an energy detection module configured to trigger a translation request when a speech segment in the speech signal to be translated begins, and to upload the speech segment to a cloud translation module in a streaming manner, and to stop uploading when the speech segment ends; a cloud translation module configured to respond to the translation request, and to recognize, translate and convert the original speech stream corresponding to the speech segment to generate at least one target language speech stream; a broadcast transmission module configured to allocate independent BIS channels for at least one target language speech stream and the original speech stream through a wireless broadcast signal based on the Auracast protocol, and to add network time protocol-based audio to each speech data packet in each speech stream; and a listener terminal module configured to respond to a user's request to select a target language, determine the target BIS channel corresponding to the target language, and output audio in the target language based on the speech data packets in the target BIS channel. Thus, by using local real-time audio pickup, intelligent segmented uploading, cloud-based streaming translation, simultaneous broadcasting of multilingual audio streams, and user-selectable playback, synchronous, real-time, and on-demand broadcasting of multilingual content is achieved, while avoiding resource waste and management complexity caused by multiple transmitters. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0008] Figure 1 shows a schematic diagram of the structure of a multilingual synchronous broadcasting system based on Auracast according to an embodiment of this application.

[0009] Figure 2 shows a flowchart of a multilingual synchronous broadcasting method based on Auracast according to an embodiment of this application. Detailed Implementation

[0010] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0011] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0012] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0013] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0014] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0015] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.

[0016] In related technologies, providing understandable voice information to listeners of different native languages ​​has always been a key requirement in cross-border communication scenarios and multilingual environments. In broadcasting systems, traditional wireless simultaneous interpretation has wide coverage but lacks a personalized experience, while Bluetooth sharing is personalized but has narrow coverage. Moreover, transmitters usually adopt a single-language transmission mode or provide multilingual content through frequency bands and time periods. Adding languages ​​usually means adding independent wireless transmission resources, and there is a delay introduced by the translation process. Furthermore, existing distribution methods cannot guarantee lip-sync or content synchronization between multilingual streams. Although these methods can achieve limited multilingual coverage, they cannot achieve multilingual synchronization, real-time and on-demand broadcasting, resulting in poor user experience, wasted resources and system complexity.

[0017] To address the aforementioned issues, this application provides a multilingual synchronous broadcasting system and method based on Auracast. The system includes: a sound pickup and processing module configured to acquire and process raw audio signals from the environment to obtain a speech signal to be translated; an energy detection module configured to trigger a translation request when a speech segment in the speech signal to be translated begins, and to upload the speech segment to a cloud translation module in a streaming manner, stopping the upload when the speech segment ends; a cloud translation module configured to respond to the translation request, recognize, translate, and convert the original speech stream corresponding to the speech segment, generating at least one target language speech stream; a broadcast transmission module configured to allocate independent BIS channels for at least one target language speech stream and the original speech stream via a wireless broadcast signal based on the Auracast protocol, and to add a network time protocol-based addition to each speech data packet in each speech stream; and a listener terminal module configured to respond to a user's request to select a target language, determine the target BIS channel corresponding to the target language, and output audio in the target language based on the speech data packets in the target BIS channel.

[0018] Thus, by using local real-time audio pickup, intelligent segmented uploading, cloud-based streaming translation, simultaneous broadcasting of multilingual audio streams, and user-selectable playback, synchronous, real-time, and on-demand broadcasting of multilingual content is achieved, while avoiding resource waste and management complexity caused by multiple transmitters.

[0019] Please refer to Figure 1, which shows a schematic diagram of the structure of a multilingual synchronous broadcasting system based on Auracast according to an embodiment of this application. As shown in Figure 1, the multilingual synchronous broadcasting system 100 based on Auracast includes a sound pickup and processing module 110, an energy detection module 120, a cloud translation module 130, a broadcast transmission module 140, and a listener terminal module 150.

[0020] Specifically, the sound pickup processing module 110 is configured to acquire and process the raw audio signal in the environment to obtain the speech signal to be translated; the energy detection module 120 is configured to trigger a translation request when a speech segment in the speech signal to be translated begins, and upload the speech segment to the cloud translation module 130 in a streaming manner, and stop uploading when the speech segment ends; the cloud translation module 130 is configured to respond to the translation request, recognize, translate and convert the original speech stream corresponding to the speech segment, and generate at least one target language speech stream; the broadcast transmission module 140 is configured to allocate independent BIS channels for at least one target language speech stream and the original speech stream through a wireless broadcast signal based on the Auracast protocol, and add network time protocol-based audio data packets to each speech data packet in each speech stream; the listener terminal module 150 is configured to respond to the user's request to select the target language, determine the target BIS channel corresponding to the target language, and output the audio of the target language according to the speech data packets in the target BIS channel.

[0021] The original audio signal can be an unprocessed audio signal directly collected from the acoustic environment, which may contain the user's voice, background noise, reverberation and other environmental interference components.

[0022] In some embodiments, the sound pickup processing module 110 includes a microphone array, which includes a preset number of condenser microphones, and the preset number of condenser microphones are arranged in a ring with equal spacing.

[0023] A preset number of condenser microphones are arranged in a ring with equal spacing. This structure provides omnidirectional symmetrical spatial sampling, enabling the Auracast-based multilingual synchronous broadcasting system 100 to accurately estimate the location of the sound source based on the phase difference, and dynamically focus on the target speaker or target sound source through a digital beamforming algorithm, effectively suppressing environmental noise from non-target directions.

[0024] In one specific implementation, the microphone array employs eight high-sensitivity condenser microphones arranged in a ring, with a spacing of 4 cm between adjacent microphones. In some implementations, the microphone array operates in the frequency range of 100Hz-16kHz and has a sensitivity of -38dB. This allows for efficient acquisition of clear and complete speech signals in typical conference environments while suppressing irrelevant noise, providing high-quality audio input for subsequent real-time translation.

[0025] The original audio signal is acquired through a microphone array and preprocessed to obtain the speech signal to be translated. Specifically, in some embodiments, the sound pickup processing module 110 further includes a digital signal processor; the digital signal processor is configured to: perform frame-by-frame processing on the original audio signal, and based on the short-time energy and zero-crossing rate of each frame, use a dual-threshold mechanism to generate a speech activity detection result to identify whether each frame is a speech frame; based on the speech activity detection result, perform time-domain smoothing on the spectrum of the original audio signal during a silent period consisting of consecutive non-speech frames to estimate the current noise spectrum; subtract the current noise spectrum from the spectrum of the original audio signal to generate an enhanced speech spectrum, and transform the enhanced speech spectrum into an enhanced speech signal; apply automatic gain control to the enhanced speech signal to obtain the speech signal to be translated.

[0026] The sound pickup processing module 110 performs frame-by-frame processing on the original audio signal and combines the short-time energy and zero-crossing rate of each frame to perform speech activity detection using a dual-threshold mechanism, thereby accurately determining whether each frame belongs to a speech frame, so as to reliably distinguish the speech and non-speech regions in the original audio signal.

[0027] Framing is a fundamental operation on the original audio signal, enabling non-stationary speech to be perceived as a stationary signal over a short period of time. This is achieved by dividing the continuous original audio signal into short frames of fixed length (e.g., 20–30 ms) that typically overlap (e.g., 50%).

[0028] Short-time energy can be the sum of the squares of the amplitudes of the samples in a single frame of speech signal, and it is used to measure the intensity of that frame. Speech frames typically have higher energy, while silence or noise frames have lower energy.

[0029] Zero-crossing rate (ZCR) is the number of times a signal waveform crosses the zero axis within a single frame, reflecting the frequency characteristics of the signal. Speech frames have a moderate zero-crossing rate, high-frequency noise has a high zero-crossing rate, and pure silence has a zero-crossing rate close to zero.

[0030] The dual-threshold mechanism allows setting two energy thresholds, high and low, to determine the start and end of a speech segment based on the set energy thresholds.

[0031] Building upon this, the system utilizes silent periods composed of consecutive non-speech frames to perform time-domain smoothing on the spectrum of the original audio signal, thereby estimating the spectral characteristics of the current environmental noise in real time. Furthermore, since the noise model is constructed solely based on pure noise segments, it effectively avoids mistakenly estimating speech components as noise during periods of active speech.

[0032] Specifically, during the detected silence period, a Fourier transform is performed on the original audio signal to obtain its spectrum, and a noise model of the current environment is obtained through time-domain smoothing (e.g., exponential averaging). The estimated noise spectrum is subtracted from the spectrum of the original audio signal to suppress background noise, preserve the target speech components, and improve the signal-to-noise ratio of the speech. The amplitude of the audio signal is dynamically adjusted to maintain it within a suitable level range for subsequent processing, avoiding a decrease in recognition performance due to speaker distance or volume changes. Finally, a clear, balanced, and suitable speech signal for translation processing is obtained, providing high-quality audio input for the energy detection module 120 to trigger translation requests and for the cloud translation module 130 to perform real-time translation.

[0033] Therefore, the audio processing module 110 completes the acquisition, noise reduction, speech enhancement, and level stabilization of the original audio signal, outputting a high-quality speech signal to be translated. The speech signal to be translated retains the speech characteristics of the target speaker while significantly suppressing environmental noise interference, laying a reliable foundation for subsequent speech activity detection and real-time translation. Based on this, the audio processing module 110 inputs the speech signal to be translated to the energy detection module 120, which is responsible for real-time monitoring of the speech start and end status to trigger and control the start and termination of the translation process.

[0034] In some implementations, the energy detection module 120 is integrated into the sound pickup processing module 110. For example, the energy detection module 120 may be a dedicated processing unit integrated into a digital signal processor.

[0035] The energy detection module 120 is configured to: calculate the short-time energy and zero-crossing rate of the speech signal to be translated in real time; when the short-time energy is greater than or equal to a first threshold and the zero-crossing rate is less than or equal to a second threshold, determine the start of the speech segment, generate a translation request, and write the current timestamp and subsequent speech data into the buffer; when the short-time energy of a consecutive preset number of frames is less than the first threshold and the zero-crossing rate is greater than the second threshold, determine the end of the speech segment, stop writing to the buffer, and send the complete speech segment stored in the buffer along with the timestamp to the cloud translation module.

[0036] The first threshold can be a preset short-time energy threshold, used to distinguish speech activity from silence or low-intensity noise. The value of the first threshold is set based on the typical human voice energy level and is usually significantly higher than the average energy value of the ambient background noise.

[0037] The second threshold can be a preset zero-crossing rate threshold, which is used to exclude false triggering by high-frequency noise (e.g., friction noise, electronic interference). The value of the second threshold is set within the zero-crossing rate range of typical speech signals, and is lower than the zero-crossing rate level of non-speech interferences such as white noise.

[0038] By jointly judging the short-time energy and zero-crossing rate with the first threshold and the second threshold respectively, the energy detection module 120 can accurately identify the starting point of the speech segment in a complex acoustic environment, effectively avoiding false detection or missed detection caused by sudden noise or silence gaps.

[0039] The dedicated processing unit corresponding to the energy detection module 120 has a computer program embedded in it to immediately generate a translation request and initiate the caching and uploading process of the voice data when the start of a voice segment is detected. When the voice segment ends, the complete voice segment, along with a precise start timestamp, is sent to the cloud translation module 130.

[0040] The timestamp not only identifies the absolute start time of the audio content, but also serves as the benchmark for subsequent synchronous broadcasting of multilingual audio streams. The cloud translation module 130 responds to translation requests, performs streaming recognition, translation, and synthesis on the received audio segments, thereby realizing an end-to-end low-latency processing link from local audio capture to real-time multilingual output.

[0041] Specifically, in some implementations, the cloud translation module 130 is configured to: upon receiving voice data of a preset duration, process the voice data through a speech recognition engine to generate source language text of the voice data; translate the source language text into at least one target language text through a neural machine translation model; and convert each target language text into a corresponding language speech stream using a speech synthesis engine to obtain at least one target language speech stream.

[0042] In some implementations, the preset duration can refer to the basic time unit for streaming processing by the cloud translation module 130, which is typically set to a fixed value between 50 milliseconds and 200 milliseconds (e.g., 100 milliseconds). The preset duration is sufficient to include the short-term steady-state characteristics of the speech to ensure recognition accuracy, while being short enough to maintain low latency response, making it suitable for real-time scenarios such as meetings and speeches.

[0043] A speech recognition engine is an automatic speech recognition engine based on deep learning, configured to convert input raw speech streams into corresponding source language text. Speech recognition engines can employ end-to-end models (e.g., Conformer, Transducer) or traditional hybrid models (e.g., DNN-HMM), supporting incremental decoding of streaming audio segments and outputting source language text corresponding to the current speech content.

[0044] Neural machine translation (NMT) models are a type of machine translation model based on neural networks. They are used to translate source language text output by a speech recognition engine into one or more target language texts in real time. NMT models can be based on the Transformer architecture, trained on large-scale parallel corpora, and possess context-aware and streaming translation capabilities, adapting to colloquial expressions while maintaining semantic accuracy.

[0045] A speech synthesis engine is a text-to-speech (TTS) engine configured to convert translated target language text into natural and fluent speech waveforms. Speech synthesis engines can use neural network-based acoustic models (e.g., Tacotron, FastSpeech) in conjunction with vocoders (e.g., WaveNet, HiFi-GAN) to generate highly natural, low-latency target language speech streams for subsequent synchronous broadcasting.

[0046] For example, the cloud translation module 130 converts English speech into text through speech recognition, translates the English text into Chinese and French text through machine translation, and converts the translated text into Chinese and French speech streams through speech synthesis.

[0047] Therefore, it can be seen that by having the speech recognition engine, neural machine translation model, and speech synthesis engine work collaboratively in a pipeline manner, real-time conversion from speech input to multilingual speech output is achieved. However, in the streaming process, if each segment is recognized, translated, and synthesized independently, it is easy to cause problems such as semantic breaks, abrupt changes in intonation, or disjointed rhythm in the generated speech. For example, "today and tomorrow" is mistranslated as "tomorrow and today" due to lack of context.

[0048] Based on this, in some implementations, the cloud translation module 130 is configured to: cache source language text, at least one target language text, and time alignment information between the two during streaming processing; and control the speech synthesis engine to generate at least one target language speech stream based on the time alignment information.

[0049] During the streaming process, the cloud translation module 130 caches the generated source language text, at least one target language text, and the time alignment information between the two (e.g., word-level or phoneme-level timestamp mapping). Based on the time alignment information, it dynamically adjusts the pronunciation rhythm, pause positions, and intonation contours of the speech synthesis engine to ensure that the output target language speech stream maintains semantic accuracy while having a natural and fluent listening experience and temporal characteristics synchronized with the original speech.

[0050] Furthermore, in order to achieve low-latency and high-reliability distribution of multilingual audio, in some embodiments, the broadcast transmission module 140 is configured to: allocate independent BIS channels for the original voice stream and each target language voice stream based on the Bluetooth protocol; assign a unique channel ID to each BIS channel and divide the original voice stream and each target language voice stream into voice data packets according to time; add a timestamp synchronized based on the Network Time Protocol to each voice data packet; and synchronously broadcast the timestamped voice data packets in all BIS channels based on the Bluetooth protocol.

[0051] Before broadcasting, each audio stream is segmented into audio data packets at fixed time intervals (e.g., 10 milliseconds), and each data packet is embedded with a timestamp synchronized based on the Network Time Protocol (NTP). This timestamp ensures that the audio data in all BIS channels are strictly aligned on the timeline, enabling listeners with different languages ​​to hear the corresponding content at the same speaking time even if they select different channels, thus achieving a cross-language synchronous listening experience.

[0052] The Bluetooth protocol can be Bluetooth 5.2 or a later version that supports low-power audio functionality. The Bluetooth protocol introduces broadcast audio capabilities, allowing a single transmitting device to simultaneously transmit high-quality audio streams to multiple receiving devices, providing a standardized wireless transmission foundation for multilingual public address scenarios.

[0053] A BIS channel (Broadcast Isochronous Stream) is a unidirectional, time-synchronized audio broadcast channel defined by the LE Audio protocol. Each BIS channel can carry an independent audio stream and is distinguished by a unique channel ID.

[0054] In the Auracast-based multilingual synchronous broadcasting system 100, the broadcast transmission module 140 allocates a BIS channel for the original audio stream (e.g., English) and allocates an independent BIS channel for each target language audio stream (e.g., Chinese, French), thereby realizing the parallel broadcasting of multilingual content.

[0055] For example, the broadcast transmission module 140 receives the original English voice stream and the translated Chinese and French voice streams, assigns them channel IDs: 0x0001, 0x0002 and 0x0003 respectively, and broadcasts them through the BIS channel after adding a synchronization timestamp.

[0056] More specifically, in some preferred embodiments, in the Auracast-based multilingual synchronous broadcasting system 100, BIS channel management is dynamic and has a pre-synchronization mechanism. To optimize end-to-end latency, the energy detection module 120 is further configured to analyze the energy envelope change trend of the speech signal in real time. When the energy of a speech segment is detected to be continuously decreasing and below a preset end prediction threshold, the energy detection module 120 sends an "expected signal" to the broadcast transmission module 140 and the cloud translation module 130.

[0057] In response to the expected signal, the cloud translation module 130, while completing the streaming translation output of the current sentence segment, initializes the translation context of the next potential sentence segment in advance and pre-allocates internal buffer resources.

[0058] In response to the expected signal, the broadcast transmission module 140 initiates the "channel timing warm-up" process: while waiting for the translation result of the next segment, according to the Bluetooth LE Audio protocol specification, padding data packets are inserted into all active BIS channels to maintain synchronization and keep the isochronous intervals and time bases of each channel stable. This allows the original and translated speech streams of the next segment to begin broadcasting almost simultaneously in each destination BIS channel, minimizing perceptual latency.

[0059] Furthermore, to optimize wireless spectrum utilization, the system implements content-aware dynamic resource scheduling, which can adaptively allocate BIS channel resources based on the complexity of the translated content.

[0060] During real-time processing, the cloud-based translation module 130 analyzes the translation complexity metrics of the current sentence segment. These metrics include, but are not limited to, the confidence level of the machine translation model, sentence length, and the bit rate required for target language synthesis. Based on these metrics, it generates a "resource requirement alert signal" and sends it to the broadcast transmission module 140.

[0061] The broadcast transmission module 140 is configured to receive resource demand prompts and, within the limits allowed by the Bluetooth LE Audio protocol, dynamically fine-tune the communication parameters of different BIS channels. For example, for speech streams with low confidence or requiring high fidelity to ensure clarity, a higher encoding bit rate or a better encoding mode is temporarily allocated; conversely, a higher compression ratio encoding is used to save overall bandwidth. This on-demand allocation mechanism maximizes the number of languages ​​the system can simultaneously broadcast while ensuring critical voice quality.

[0062] After the broadcast transmission module 140 completes the synchronous broadcast of the multilingual voice stream, the listener terminal module 150, as the receiving end, is responsible for scanning and accessing the broadcast signal. Specifically, in some embodiments, the listener terminal module 150 is configured to: in response to the user's selection of the target language, determine the target BIS channel corresponding to the target language, and synchronously receive the voice data packets in the target BIS channel; decode and play the data packets based on the network time protocol synchronization timestamp carried by the data packets, and output the audio in the target language.

[0063] In some implementations, the listener terminal module 150 can be an Auracast-enabled Bluetooth headset. In one specific implementation, the Bluetooth headset has a built-in Bluetooth 5.2 chip that supports BIS reception and decoding. The listener terminal module 150 supports switching BIS channels via physical buttons or voice commands. For example, the listener terminal module 150 includes a physical button with multiple functions; a short press switches to the next channel, and a long press for a preset duration switches to the previous channel. As another example, voice commands support instructions such as "switch to Chinese" and "switch to English," which are implemented through the voice recognition engine built into the listener terminal module 150.

[0064] To improve the accuracy of multilingual synchronization and system robustness, the system introduces bidirectional timestamp alignment and cross-channel fault tolerance compensation.

[0065] First, the energy detection module 120 generates a high-precision start timestamp for each voice segment, which is not only uploaded to the cloud translation module 130, but also directly transmitted to the broadcast transmission module 140 as an absolute time reference.

[0066] Secondly, during streaming translation and speech synthesis, the cloud translation module 130 generates a "time-aligned metadata" file. This data accurately records the mapping relationship between each semantic unit (such as a word or phoneme) in the target language speech and the original speech timestamp. This metadata is sent to the broadcast transmission module 140 along with the target language speech stream.

[0067] The broadcast transmission module 140 is configured to receive the original voice stream, voice streams of each target language, and their corresponding "time-aligned metadata." Using the original voice timeline as a reference, it dynamically adjusts the encapsulation timing of data packets in the non-original language BIS channel using the metadata to ensure semantic synchronization. Simultaneously, it performs forward error correction coding on the broadcast data.

[0068] The listener terminal module 150 is further configured to: during decoding and playback, if a brief data packet loss is detected in the current BIS channel, use the timestamp and / or alignment metadata in the data packet to extract audio features from the corresponding time period of other BIS channels (such as the original voice channel), and generate brief compensation audio through an interpolation algorithm to avoid playback interruption, maintain listening continuity, and enhance fault tolerance.

[0069] In some implementations, the listener terminal module 150 also supports channel list display. For example, all currently available language channels and their signal strengths can be displayed via a mobile app that accompanies the listener terminal module 150.

[0070] In some implementations, the listener terminal module 150 also supports automatic channel switching. For example, when the signal strength of the current channel is detected to be lower than a preset value (e.g., -85dBm) for a preset duration (e.g., 3 seconds), it automatically searches for and switches to other channels in the same language (if there are multiple transmitters covering the channel).

[0071] In some implementations, the listener terminal module 150 also supports reconnection after a disconnection. For example, when the signal is interrupted, the listener terminal module 150 automatically attempts to reconnect, with a maximum reconnection interval of 5 seconds.

[0072] For example, visitors wearing the listener terminal module 150 can switch between Chinese, English, and Japanese audio guides by pressing a button on the listener terminal module 150 while touring the museum. When visitors move to areas with weak signal, the listener terminal module 150 automatically switches to the same language channel provided by a backup transmitter to ensure the continuity of the visitor's listening experience.

[0073] During normal playback, the listener terminal module 150 continuously receives voice data packets from the corresponding BIS channel based on the user's initial target language selection, and decodes and plays them according to the embedded NTP synchronization timestamps to ensure that the audio output is consistent with the speaking time. However, in actual use, users may switch target languages ​​midway due to changes in understanding needs (e.g., from Chinese to French). If the user switches directly and immediately plays the latest data packets of the new language stream, the playback content will jump to a point after the current time, resulting in the loss of key sentences and disrupting auditory continuity.

[0074] Based on this, in some implementations, the listener terminal module 140 is configured to: in response to a user's request to switch target languages, align the playback of the newly selected audio stream based on timestamps.

[0075] Upon receiving a user's request to switch target languages, the system obtains the current system playback time (or references the time base of the original audio stream). Based on the NTP timestamps carried in the audio data packets of each BIS channel, it locates the data packet in the newly selected language stream that is closest to the current time and begins decoding and playback from that position. Through this timestamp-based playback alignment mechanism, even if the user switches languages ​​midway through a speech, they can still hear content synchronized with the current speaking progress, achieving a seamless and coherent multilingual listening experience.

[0076] More specifically, in some preferred embodiments, in the Auracast-based multilingual simultaneous broadcasting system 100, to improve the accuracy of speech endpoint detection in complex acoustic environments and to provide richer context for translation, the energy detection module 120 can be upgraded to a "multimodal speech state detection engine".

[0077] The engine is configured to receive and fuse audio signals from the sound pickup processing module 110 and signals from at least one auxiliary sensor (such as a lavalier microphone signal associated with the speaker, a status signal from the meeting reservation system, or a camera visual signal processed by edge computing) to determine the start and end of a valid speech with higher accuracy.

[0078] When a translation request is triggered, the engine can generate a "speech scene tag" (such as "keynote speech", "free discussion", "background music playing") and upload it to the cloud translation module 130. The cloud translation module 130 can use this tag to select an appropriate translation model or post-processing strategy (for example, enabling a more conversational and interference-resistant translation model in the "free discussion" scenario).

[0079] Meanwhile, the broadcast transmission module 140 can use the received scene tags as broadcast metadata and broadcast them through the auxiliary data field of the BIS. The listener terminal module 150 can receive and parse this metadata and automatically adjust the playback parameters accordingly (such as appropriately increasing the voice gain in the "background music playing" scenario).

[0080] Please refer to Figure 2, which shows a flowchart of a multilingual synchronous broadcasting method based on Auracast according to an embodiment of this application. As shown in Figure 2, the multilingual synchronous broadcasting method based on Auracast includes steps 210 to 250.

[0081] In step 210, the raw audio signal in the environment is acquired and processed to obtain the speech signal to be translated.

[0082] The raw audio signal can be an unprocessed audio signal directly collected from the acoustic environment, which may contain the user's voice, background noise, reverberation and other environmental interference components.

[0083] The speech signal to be translated can be an optimized audio signal output from the original audio signal after preprocessing. Preprocessing typically includes operations such as noise reduction, beamforming, automatic gain control (AGC), and spectrum enhancement, which aim to suppress environmental noise and reverberation interference and improve the clarity and signal-to-noise ratio of the target speaker's speech.

[0084] The speech signal to be translated retains the semantic content and temporal features of the original audio signal, while having higher recognizability, serving as a high-quality input source for subsequent speech activity detection and cloud translation.

[0085] In step 220, a translation request is triggered when a speech segment in the speech signal to be translated begins, and the speech segment is uploaded in a streaming manner. The upload is stopped when the speech segment ends.

[0086] The short-time energy and zero-crossing rate of the speech signal to be translated are calculated in real time. When the short-time energy is greater than or equal to the first threshold and the zero-crossing rate is less than or equal to the second threshold, the speech segment is determined to start, a translation request is generated, and the current timestamp and subsequent speech data are written to the buffer.

[0087] When the short-time energy of a preset number of consecutive frames is less than the first threshold and the zero-crossing rate is greater than the second threshold, the speech segment is determined to be over, writing to the buffer is stopped, and the complete speech segment stored in the buffer, along with the timestamp, is sent to the cloud translation module.

[0088] To achieve dynamic pre-synchronization, step 221 is also included: when the end of a speech segment is detected, an expected end signal is generated; based on the signal, the broadcast channel is pre-warmed up in time and resources are pre-allocated for the translation processing of the next segment.

[0089] In step 230, in response to the translation request, the original speech stream corresponding to the speech segment is identified, translated and converted to generate at least one target language speech stream.

[0090] Upon receiving voice data of a preset duration, the voice data is processed by a speech recognition engine to generate source language text; the source language text is translated into at least one target language text using a neural machine translation model; and each target language text is converted into a corresponding speech stream using a speech synthesis engine to obtain at least one target language speech stream.

[0091] To ensure that timestamps are aligned as much as possible and to increase the fault tolerance of translation, step 231 is also included: passing the start timestamp of the speech segment as an absolute reference to the broadcast unit; dynamically adjusting the encapsulation timing of non-original language broadcast data packets using time-aligned metadata generated during the translation process; and at the receiving end, when data loss of the current channel is detected, extracting features from other channels based on timestamps for audio compensation.

[0092] Furthermore, it can dynamically adjust the wireless communication parameters of different language broadcasting channels based on the real-time complexity of the translated content, thereby enabling intelligent on-demand allocation of spectrum resources.

[0093] In step 240, based on the Auracast protocol, independent BIS channels are allocated for at least one target language speech stream and the original speech stream via a wireless broadcast signal, and a network time protocol is added to each speech data packet in each speech stream.

[0094] Before broadcasting, each audio stream is segmented into audio data packets at fixed time intervals (e.g., 10 milliseconds), and each data packet is embedded with a timestamp synchronized based on the Network Time Protocol (NTP). This timestamp ensures that the audio data in all BIS channels are strictly aligned on the timeline, enabling listeners with different languages ​​to hear the corresponding content at the same speaking time even if they select different channels, thus achieving a cross-language synchronous listening experience.

[0095] In step 250, in response to the user's request to select a target language, the target BIS channel corresponding to the target language is determined, and the audio of the target language is output according to the voice data packets in the target BIS channel.

[0096] Based on the target language initially selected by the user, the system continuously receives voice data packets from the corresponding BIS channel and decodes and plays them according to the embedded NTP synchronization timestamp, ensuring that the audio output is consistent with the speaking time.

[0097] In addition, audio signals can be fused with signals from at least one auxiliary sensor to improve the accuracy of effective speech detection; based on the detected speech scenario, an appropriate translation strategy is selected, and the scenario information is sent as metadata with the broadcast for the receiving end to make adaptive playback adjustments.

[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above description of the specific working process of the multilingual synchronous broadcasting method based on Auracast can be referred to the corresponding process in the aforementioned system embodiments, and will not be repeated here.

[0099] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A multilingual synchronous broadcasting system based on Auracast, characterized in that, The system includes: a sound pickup and processing module configured to acquire and process raw audio signals from the environment to obtain a speech signal to be translated; an energy detection module configured to trigger a translation request when a speech segment in the speech signal to be translated begins, and upload the speech segment to a cloud translation module in a streaming manner, and stop uploading when the speech segment ends; the cloud translation module configured to respond to the translation request, and to recognize, translate and convert the original speech stream corresponding to the speech segment to generate at least one target language speech stream; a broadcast transmission module configured to allocate independent BIS channels for the at least one target language speech stream and the original speech stream through a wireless broadcast signal based on the Auracast protocol, and to add a network time protocol-based addition to each speech data packet in each speech stream; and a listener terminal module configured to respond to a user's request to select a target language, determine the target BIS channel corresponding to the target language, and output the audio of the target language based on the speech data packets in the target BIS channel.

2. The multilingual synchronous broadcasting system based on Auracast according to claim 1, characterized in that, The sound pickup processing module includes a microphone array, which includes a preset number of condenser microphones, and the preset number of condenser microphones are arranged in a ring with equal spacing. The microphone array is used to acquire the raw audio signal.

3. The multilingual synchronous broadcasting system based on Auracast according to claim 2, characterized in that, The sound pickup processing module further includes a digital signal processor; the digital signal processor is configured to: perform frame-by-frame processing on the original audio signal, and based on the short-time energy and zero-crossing rate of each frame, use a dual-threshold mechanism to generate a speech activity detection result to identify whether each frame is a speech frame; based on the speech activity detection result, within a silent period consisting of consecutive non-speech frames, perform time-domain smoothing on the spectrum of the original audio signal to estimate the current noise spectrum; subtract the current noise spectrum from the spectrum of the original audio signal to generate an enhanced speech spectrum, and transform the enhanced speech spectrum into an enhanced speech signal; Automatic gain control is applied to the enhanced speech signal to obtain the speech signal to be translated.

4. The multilingual synchronous broadcasting system based on Auracast according to claim 1, characterized in that, The energy detection module is integrated into the sound pickup processing module. The energy detection module is configured to: calculate the short-time energy and zero-crossing rate of the speech signal to be translated in real time; when the short-time energy is greater than or equal to a first threshold and the zero-crossing rate is less than or equal to a second threshold, determine the start of the speech segment, generate the translation request, and write the current timestamp and subsequent speech data into the buffer. When the short-time energy of a consecutive preset number of frames is less than a first threshold and the zero-crossing rate is greater than a second threshold, the speech segment is determined to have ended, writing to the buffer is stopped, and the complete speech segment stored in the buffer, along with the timestamp, is sent to the cloud translation module.

5. The multilingual synchronous broadcasting system based on Auracast according to claim 1 or 4, characterized in that, The energy detection module is configured to predict the imminent end of the current speech segment based on the changing trend of the original audio signal and generate a prediction signal; the broadcast transmission module and / or the cloud translation module are configured to perform a pre-synchronization operation in response to the prediction signal to reduce the end-to-end broadcast delay of the next speech segment.

6. The multilingual synchronous broadcasting system based on Auracast according to claim 1, characterized in that, The cloud translation module is configured to: upon receiving voice data of a preset duration, process the voice data using a speech recognition engine to generate source language text of the voice data; translate the source language text into at least one target language text using a neural machine translation model; and convert each of the target language texts into a corresponding language speech stream using a speech synthesis engine to obtain the at least one target language speech stream.

7. The multilingual synchronous broadcasting system based on Auracast according to claim 6, characterized in that, The cloud translation module is configured to: cache the source language text, the at least one target language text, and the time alignment information between them during streaming processing; and control the speech synthesis engine to generate the at least one target language speech stream based on the time alignment information.

8. The multilingual synchronous broadcasting system based on Auracast according to claim 1, characterized in that, The broadcast transmission module is configured to: allocate independent BIS channels for the original voice stream and each target language voice stream based on the Bluetooth protocol; assign a unique channel ID to each BIS channel and divide the original voice stream and each target language voice stream into voice data packets according to time; and add a timestamp synchronized based on the Network Time Protocol to each voice data packet. Based on the Bluetooth protocol, timestamped voice data packets in all BIS channels are broadcast synchronously.

9. The multilingual synchronous broadcasting system based on Auracast according to claim 1, characterized in that, The listener terminal module is configured to: in response to the user's selection of a target language, determine the target BIS channel corresponding to the target language, and synchronously receive voice data packets in the target BIS channel; Based on the network time protocol synchronization timestamp it carries, it is decoded and played to output the audio of the target language.

10. A multilingual synchronous broadcasting method based on Auracast, characterized in that, The method includes: acquiring and processing raw audio signals in the environment to obtain a speech signal to be translated; triggering a translation request when a speech segment in the speech signal to be translated begins, and uploading the speech segment in a streaming manner, and stopping the uploading when the speech segment ends; responding to the translation request, recognizing, translating and converting the original speech stream corresponding to the speech segment to generate at least one target language speech stream; allocating independent BIS channels for the at least one target language speech stream and the original speech stream through a wireless broadcast signal based on the Auracast protocol, and adding a network time protocol-based addition to each speech data packet in each speech stream; responding to a user's request to select a target language, determining the target BIS channel corresponding to the target language, and outputting the audio of the target language based on the speech data packets in the target BIS channel.