Call method, device and call platform suitable for multi-user multi-language translation

By employing dynamic role switching and multi-channel anti-interference transmission technology, the scalability bottleneck and transmission vulnerability of existing technologies in multi-person, multi-language translation have been resolved. This enables efficient and reliable translation and voice interaction in multi-person, multi-language scenarios, supports dynamic role switching and language selection, and improves translation quality and synchronization between devices.

CN120980171APending Publication Date: 2025-11-18JI LI MI (SHEN ZHEN) JI SHU YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511087737.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing technologies for multi-person, multilingual translation suffer from several problems, including scalability bottlenecks, high costs, uncontrollable latency, one-way fixed roles, mechanical language switching, strong network dependence, complex operation, fragile transmission, need to interrupt the session for language switching, and hardware-software separation. As a result, they cannot achieve efficient and reliable real-time multi-person, multilingual translation and voice interaction.

Method used

By employing dynamic role switching, multi-channel anti-interference transmission, and intelligent language routing technologies, real-time translation and voice interaction are achieved in multi-person, multi-language scenarios. A lavalier microphone is used for language recognition and labeling, and neural machine translation and speech synthesis are combined to generate multi-target language translation audio. Independent transmission links are established through dynamic channel preemption, and the receiving device can switch modes and select languages ​​to ensure the real-time performance and stability of the translation.

Benefits of technology

It enables seamless integration and efficient and reliable cross-language communication in multi-person, multi-language scenarios, supports dynamic role switching, reduces operational complexity, improves translation quality and device synchronization, adapts to complex acoustic environments, and ensures real-time translation and anti-interference capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980171A_ABST
    Figure CN120980171A_ABST
Patent Text Reader

Abstract

The invention relates to a conversation method, device and conversation platform suitable for multi-user multi-language translation, and the method comprises the steps: obtaining a first-language audio signal inputted by a speaker through a speaker device; performing real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language; distributing the translated audio signal to a plurality of receiving ends; recognizing a language selection operation of each receiving end, and outputting a translation voice signal of a target language version corresponding to the language selection operation; responding to a mode switching operation of the receiving end equipment, and switching the receiving end equipment into a sending end mode; acquiring a second language audio signal input by the auxiliary speaker through the equipment with the switched mode; and performing real-time translation processing on the second-language audio signal and distributing the second-language audio signal to each receiving end so as to realize seamless connection of real-time translation and voice interaction in a multi-user and multi-language scene and construct efficient and reliable full-duplex cross-language communication.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of voice communication and real-time translation, in particular to a call method and device suitable for multi-person multi-language translation and a call platform. BACKGROUND

[0002] The deep development of global cooperation makes higher requirements for real-time translation technology in scenarios such as international conferences, cross-border business negotiations and multi-language teaching. Traditional solutions have significant limitations at the technical generation level: although the artificial simultaneous interpretation system can handle complex contexts, it relies on professional interpreters to translate voice through soundproof cabins, which has an expansion bottleneck, for example, a single conference requires a multi-language interpreter team, the cost increases geometrically, the delay is uncontrollable, such as the inherent delay of several seconds caused by the cognitive processing of interpreters, which breaks the coherence of the conversation and has poor scene adaptability, such as the inability to respond to spontaneous questions and multi-person debates and other high-dynamic interactions. As for existing electronic translation devices, single-machine offline translators use embedded speech recognition (ASR) and machine translation (MT) technologies, which are limited by one-way solidification of roles, only support one-way transmission mode of "speaker-listener", mechanical language switching, rely on physical sound channel segmentation, cannot support more than three languages output, and are weak in anti-interference, single-channel transmission has high packet loss rate in WiFi / Bluetooth coexistence environment, while online conference translation systems based on cloud servers improve translation quality but introduce network dependency, such as audio asynchronization caused by cross-border transmission jitter, privacy leakage risk, sensitive conference content transmission through public network, and lack of hardware cooperation, translation terminals, microphones and earphones are independent, increasing the operation complexity. The industry problems are mainly manifested in the speaker-listener role gap hindering natural conversation, transmission vulnerability, lack of audio stability in high-density wireless environment, operation anti-human, language switching needs to interrupt the conversation for manual adjustment, and system closure, which leads to fragmented experience due to separation of hardware, software and service. SUMMARY

[0003] The main purpose of the present application is to provide a call method, device and platform suitable for multi-person multi-language translation, which realizes seamless connection of real-time translation and voice interaction in multi-person multi-language scenarios through dynamic role switching, multi-channel anti-interference transmission and intelligent language routing technology, and builds an efficient and reliable full-duplex cross-language communication.

[0004] To achieve the above purpose, the present application provides a call method suitable for multi-person multi-language translation, comprising the following steps:

[0005] obtaining a first language audio signal input by a speaker through a speaker device;

[0006] performing real-time translation processing on the first language audio signal to generate a translation audio signal containing at least one target language;

[0007] distributing the translated audio signal to a plurality of receiving ends;

[0008] identifying a language selection operation of each receiving end, and outputting a translated voice signal of a target language version corresponding to the language selection operation;

[0009] switching the receiving end device to a sending end mode in response to a mode switching operation of the receiving end device;

[0010] obtaining a second language audio signal input by a secondary speaker through the device after switching the mode;

[0011] performing real-time translation processing on the second language audio signal and distributing the second language audio signal to each receiving end.

[0012] Further, the step of obtaining a first language audio signal input by a primary speaker through a primary speaker device includes:

[0013] collecting primary speaker speech and identifying an original language;

[0014] transmitting the first language audio signal carrying a language tag to a translation host.

[0015] Further, the step of performing real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language includes:

[0016] selecting a translation engine based on the language tag to generate a translated text in at least one target language;

[0017] performing voice synthesis on the translated text to generate a translated audio signal in a corresponding target language.

[0018] Further, the step of distributing the translated audio signal to a plurality of receiving ends includes:

[0019] scanning a plurality of frequency band wireless channel occupation states and identifying an idle channel;

[0020] occupying the idle channel to establish an independent transmission link with each receiving end;

[0021] copying the translated audio signal into an independent data stream with the same number of receiving ends;

[0022] sending the data stream to the corresponding receiving end through each independent transmission link.

[0023] Further, the step of identifying a language selection operation of each receiving end and outputting a translated voice signal of a target language version corresponding to the language selection operation includes:

[0024] obtaining a target language selection instruction sent by the receiving end device;

[0025] Based on the target language selection instruction of each receiving end device, the translation host is used to screen the translation audio signals of corresponding languages;

[0026] The translation audio signals are sent to the corresponding receiving end devices.

[0027] Further, in response to the mode switching operation of the receiving end device, the step of switching the receiving end device into the sending end mode comprises:

[0028] Receiving a speech request sent by the receiving end device with a unique ID;

[0029] Generating an approval pop-up window containing the ID of the receiving end device on the interface of the translation host;

[0030] If there are multiple speech requests at the same time, a priority queue is generated according to the receiving time, and only the approval pop-up window of the first request in the queue is displayed;

[0031] In response to the confirmation operation of the translation host on the approval pop-up window, the communication mode of the receiving end device initiating the speech request is switched to the sending end mode, and the audio acquisition function of the receiving end device initiating the speech request is activated.

[0032] Further, the step of obtaining the second language audio signal input by the secondary speaker through the device after switching the mode comprises:

[0033] Verifying the consistency of the ID of the receiving end device initiating the speech request and the ID of the device after switching the mode;

[0034] Obtaining the audio signal of the secondary speaker through the audio acquisition unit built in the receiving end device initiating the speech request.

[0035] Further, the step of performing real-time translation processing on the second language audio signal and distributing it to each receiving end comprises:

[0036] Identifying the language characteristics of the second language audio signal;

[0037] Generating a secondary speech translation audio signal containing at least one target language;

[0038] Distributing to all devices in the receiving mode through the preempted idle channel.

[0039] The application also provides a call device suitable for multi-person multi-language translation, comprising:

[0040] A main speech audio acquisition module is used to acquire the first language audio signal input by the main speaker through the main speaker device;

[0041] A first translation processing module is used to perform real-time translation processing on the first language audio signal, and generate a translation audio signal containing at least one target language;

[0042]

[0042] a voice distribution module configured to distribute the translated audio signal to a plurality of receiving ends;

[0043] a voice output module configured to identify a language selection operation of each receiving end and output a translated voice signal in a target language version corresponding to the language selection operation;

[0044] a mode switching module configured to switch the receiving end device into a sending end mode in response to a mode switching operation of the receiving end device;

[0045] a secondary speaker audio acquisition module configured to acquire a second language audio signal input by a secondary speaker through the device after switching the mode;

[0046] a second translation processing module configured to perform real-time translation processing on the second language audio signal and distribute the second language audio signal to each receiving end.

[0047] The application further provides a call platform suitable for multi-person multi-language translation, comprising:

[0048] a multi-language interaction management system configured to acquire a first language audio signal input by a primary speaker through a primary speaker device and perform real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language;

[0049] a dynamic wireless distribution system configured to distribute the translated audio signal to a plurality of receiving ends, the dynamic wireless distribution system comprising:

[0050] a channel scanning module configured to scan a multi-band wireless channel occupation state and identify an idle channel;

[0051] a link preemption module configured to preempt the idle channel to establish an independent transmission link with each receiving end;

[0052] a data stream replication module configured to replicate the translated audio signal into independent data streams in the same number as the receiving ends;

[0053] a physically isolated transmission module configured to send the data streams to the corresponding receiving ends through the independent transmission links;

[0054] a user language adaptation system configured to identify a language selection operation of each receiving end and output a translated voice signal in a target language version corresponding to the language selection operation;

[0055] a role switching control system configured to switch the receiving end device into a sending end mode in response to a mode switching operation of the receiving end device;

[0056] a secondary speaker voice acquisition system configured to acquire a second language audio signal input by a secondary speaker through the device after switching the mode;

[0057] A multi-lingual parallel translation system is used to perform real-time translation processing on the second language audio signal and distribute to each receiving end.

[0058] The application provides a talk method, device and platform suitable for multi-person multi-lingual translation, which has the following beneficial effects: the application constructs a dynamic role conversion mechanism, realizes the role migration from a listener to a speaker based on hardware-level state conversion and approval management, and breaks the fixed boundary of one-way output of a speaker and a listener in a traditional conference. The receiving end device can be seamlessly switched to the sending end after permission approval, and the device ID security verification technology is combined to ensure the orderly handover of the speaking right, so that the cross-language dialogue realizes real two-way interaction. Moreover, the language is recognized in real time based on acoustic characteristics, and a multi-target translation stream is generated, the physical sound channel and the logical language are decoupled through instruction direct control technology, and the user can obtain the target language version as needed by single-key operation, and completely get rid of the mechanical restriction of traditional sound channel switching. In addition, a multi-channel cooperative transmission architecture is designed, the optimal idle channel is dynamically occupied through spectrum sensing, and an anti-interference transmission link is established in the coexistence environment of 2.4G / WiFi / Bluetooth by combining independent data stream replication and regional coding strategy, so that the high stability of multi-device synchronous reception within a range of 200 meters is ensured. Through the closed loop of voice collection-translation-distribution-interaction-retransmission, the application improves the cross-language cooperation efficiency, and provides good technical support for building an accessible cross-language communication environment. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 is a flowchart of a talk method suitable for multi-person multi-lingual translation in an embodiment of the application;

[0060] Figure 2 is a structural block diagram of a talk device suitable for multi-person multi-lingual translation in an embodiment of the application;

[0061] Figure 3 is a structural schematic block diagram of a talk platform suitable for multi-person multi-lingual translation in an embodiment of the application.

[0062] The implementation of the application, functional features and advantages will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0063] In order to make the purpose, technical scheme and advantages of the application more clear, the application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and not to limit the application.

[0064] Reference Figure 1 is a flowchart of a talk method suitable for multi-person multi-lingual translation in an embodiment of the application, which includes the following steps: is a flowchart of a talk method suitable for multi-person multi-lingual translation in an embodiment of the application, which includes the following steps:

[0065] S1, obtaining, by a host device, a first language audio signal input by a host speaker;

[0066] S2, performing real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language;

[0067] S3, distributing the translated audio signal to a plurality of receiving ends;

[0068] S4, identifying a language selection operation of each receiving end, and outputting a translated voice signal of a target language version corresponding to the language selection operation;

[0069] S5, in response to a mode switching operation of a receiving end device, switching the receiving end device to a sending end mode;

[0070] S6, obtaining, by the device after switching the mode, a second language audio signal input by a secondary speaker;

[0071] S7, performing real-time translation processing on the second language audio signal and distributing it to each receiving end.

[0072] As described in step S1 above, the first language audio signal input by the host speaker is obtained by the host device. The host speaker audio signal is obtained, the original speech of the host speaker is accurately collected, and the language is identified to provide input for translation. By using a lapel microphone (high signal-to-noise ratio) to collect the host speaker's speech, and performing real-time language recognition (such as Chinese, English, etc.) and marking the audio signal, the problem of translation deviation caused by incorrect language recognition is solved. By forming a closed loop of collection-identification-packaging-transmission, structured input data is provided for host translation, and the combination of the hardware characteristics of the lapel microphone and the software optimization of the language recognition algorithm overcomes the problems of inaccurate language recognition and signal transmission packet loss in traditional solutions in complex acoustic environments.

[0073] As described in step S2 above, the first language audio signal is processed in real time to generate a translated audio signal containing at least one target language. The host speaker's speech is quickly converted into a multilingual translation result, and the key technologies include neural machine translation (NMT), i.e., selecting a dedicated translation engine based on language tags, and speech synthesis (TTS) to generate natural and fluent target language speech, solving the problem that traditional translation devices cannot balance low latency and multilingual support.

[0074] As described in step S3 above, the translated audio signal is distributed to a plurality of receiving ends. In a multi-device environment, the translated audio is stably transmitted by dynamically occupying channels, i.e., scanning 2.4G / WiFi / Bluetooth channels and selecting the optimal idle channel, and by duplicating data streams, i.e., sending data independently to each receiving end to avoid packet loss, solving the problem of audio interruption or delay caused by wireless signal interference.

[0075] As described in the above step S4, the language selection operation of each receiving end is identified, and a translated voice signal in a target language version corresponding to the language selection operation is output. The technical purpose of language selection and output is to allow users to switch the translated language version on demand, to send a language selection instruction through a physical key or an APP through the receiving end, and to dynamically filter the corresponding language signal by the translation host and send it to the designated device, thereby solving the problem of manual switching of sound channels in traditional devices and the inconvenience of operation.

[0076] As described in the above step S5, the receiving end device is switched to a sending end mode in response to a mode switching operation of the receiving end device. The technical purpose of switching the receiving end to the sending end mode is to realize the dynamic conversion of the "listener to speaker" role, to switch only after the host pop-up confirmation through the approval mechanism (to avoid grabbing the microphone), and to close the receiving circuit and activate the microphone circuit through the hardware state switching, thereby solving the problem of chaotic speaking rights in a multi-person meeting.

[0077] As described in the above step S6, the second language audio signal input by the secondary speaker is obtained through the device after mode switching. The secondary speaker audio signal is obtained to ensure that the secondary speaker voice comes from a legal device, which is verified through the device ID, i.e., matching the requesting device with the authorized device, and is collected by the built-in microphone to ensure the clarity of the voice, thereby solving the risk of malicious access by unauthorized devices.

[0078] As described in the above step S7, the second language audio signal is translated in real time and distributed to each receiving end. The technical purpose of the secondary speaker audio translation and distribution is to support independent translation and synchronous distribution of multiple speakers, through independent translation of the secondary speaker voice (isolated from the main speaker signal), and through channel occupation and distribution to all receiving ends (including the original main speaker device), thereby solving the translation conflict problem when multiple people speak alternately.

[0079] In one embodiment, the step S1 of obtaining the first language audio signal input by the main speaker through the main speaker device includes:

[0080] S11, collecting the voice of the main speaker and identifying the original language;

[0081] S12, transmitting the first language audio signal carrying the language tag to the translation host.

[0082] In a specific implementation, in a multilingual classroom of an international school, a lightweight lapel microphone device is worn by the main speaker to deliver a lecture. The lapel microphone device is built-in with a high-sensitivity MEMS microphone, which collects the original speech signal of the main speaker through directional sound collection technology with a sampling rate of 16 kHz. For example, a French teacher explains "the impact of climate change on the ecosystem" in a classroom of an international school. The microphone captures the French sentence "Le réchauffement climatique menace la biodiversité" (climate change threatens biodiversity). Meanwhile, a real-time noise reduction algorithm based on spectral subtraction is built-in to suppress the environmental noise in the classroom (such as student discussion and projector operation). The noise reduction algorithm eliminates the sound of turning pages by attenuating 12 dB in the frequency band of 200-500 Hz, and focuses on ± 15° sound sources through beamforming technology to ensure the clarity of the teacher's voice. The collected audio signal is input into a language recognition unit in real time. By extracting the MFCC features of the voice and combining a pre-trained probability classification model (supporting commonly used languages in the classroom such as English, Chinese, and French), the original language used by the teacher is dynamically identified. For example, when the teacher switches to bilingual explanation, the system generates a language tag (such as "zh-CN" or "en-US") in real time, and binds and packages it with the noise-reduced audio data to form a structured audio package. The frame header of the audio package is 0xAA, which contains the language tag fr_FR (French), noise-reduced audio data, and CRC check code. The audio package is transmitted to the translation host through a proprietary low-latency wireless protocol (compatible with 2.4 GHz and Bluetooth 5.2). Forward error correction (FEC) technology is used to deal with multi-device interference in the classroom during transmission, supporting 200m stable communication and ensuring data integrity. This embodiment solves the problem of manually switching language configurations in traditional multilingual communication through automatic language recognition and dynamic packaging, providing accurate input for subsequent multilingual translation, and adapting to the high-concurrency connection needs of the voice communication scene. The near-field sound collection characteristics of the lapel microphone combined with the adaptive noise reduction algorithm improve the clarity of the voice in a noisy environment, laying a solid foundation for the translation process.

[0083] In one embodiment, the step S2 of performing real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language includes:

[0084] S21, selecting a translation engine based on the language tag to generate a translated text in at least one target language;

[0085] S22, performing voice synthesis on the translated text to generate a translated audio signal in the corresponding target language.

[0086] In a specific implementation, in a multilingual classroom of an international school, a teacher (the main speaker) wears a lavalier microphone to explain "the impact of climate change on ecosystems" in French. The French voice signal (labeled language tag fr FR) is transmitted to the translation host through the S1 step, triggering the fine translation process. The translation host first parses the language tag fr FR (French) encapsulated in S1, and calls the "French-English" "French-Chinese" "French-Spanish" three-channel translation model from the pre-set engine library. For example, the French sentence "Le réchauffement climatique menace la biodiversité" (Climate warming threatens biodiversity) is converted into a translated text by an end-to-end speech recognition module, combined with a professional term library in the field of climate (such as "biodiversité"→"biodiversity" / "biodiversity"), and outputs the double-target language text through a neural machine translation engine: English translation "Climate warming threatens biodiversity", Chinese translation "Climate warming threatens biodiversity", and Spanish translation "El calentamiento global amenaza la biodiversidad". The translated text enters the speech synthesis stage, and the system extracts the acoustic fingerprint of the French teacher from the pre-stored voiceprint library, including the fundamental frequency (F0) distribution range 125-150Hz (stable voice quality), formant (F1-F3) and other key parameters. Through cross-lingual acoustic migration technology, the nasal resonance characteristics of the original French voice are adapted to the target language: when synthesizing English, the low-frequency formant (125-150Hz) is retained, and the aspiration intensity of the explosive sound / p / / t / is enhanced to meet the English pronunciation rules; when synthesizing Chinese, the nasal resonance strength is reduced, and the fundamental frequency fluctuation of the four-tone type is increased (such as the fourth tone is lowered to 110Hz); when synthesizing Spanish, the nasal resonance characteristics of French are retained, but the vowel tension is adjusted (such as the opening degree of / a / sound is increased to 85% to adapt to the Spanish pronunciation), and the consonant concatenation is strengthened (such as the transition of / m / + / n / in "amenaza"). The synthesis process uses an improved WaveGrad vocoder to generate translated audio with human voice temperature at a sampling rate of 24kHz, and finally outputs three independent signals: English channel (amplitude -3dB, speech rate 120 words / minute), Chinese channel (amplitude -2.5dB, speech rate 110 words / minute), and Spanish channel (amplitude -2.8dB, speech rate 115 words / minute).The embodiment solves the multilingual scheduling bottleneck through a dynamic routing mechanism: the language tag fr_FR as a control signal activates a dedicated translation pipeline, translating the French long sentence (e.g. "Les écosystèmes marins sont particulièrement vulnérables") into English, Chinese, and Spanish in parallel, reducing the processing time to within 200ms. The voiceprint library stores the French teacher's acoustic fingerprint (extracted by calibrating the voice 5 minutes before class), which is used for timbre migration in the synthesis stage: when synthesizing English, the pitch fluctuation pattern is maintained, and only the syllable stress position is adjusted; when synthesizing Chinese, the nasal cavity resonance strength is increased while the vocal cord vibration characteristics are preserved, so that the Chinese translation output not only conforms to the pronunciation specifications of the target language, but also retains the teacher's calm tone; when synthesizing Spanish, the formant ratio (F1:F2:F3 = 600:1800:2600Hz) is adjusted to restore the unique Spanish alveolar fricative sound (e.g. the / s / sound energy is increased by 20%). In addition, the translation host ensures classroom content privacy through the localization engine, and adapts to the needs of multilingual students (e.g. English, Chinese, and Spanish-speaking students receiving translation simultaneously) using parallel processing mechanisms, ensuring smooth classroom rhythm without interruption.

[0087] In one embodiment, the step S3 of distributing the translated audio signal to multiple receiving ends comprises:

[0088] S31, scanning the occupation state of the multi-band wireless channel and identifying the idle channel;

[0089] S32, occupying the idle channel to establish an independent transmission link with each receiving end;

[0090] S33, copying the translated audio signal into as many independent data streams as the number of receiving ends;

[0091] S34, sending data streams to corresponding receiving ends through independent transmission links.

[0092] In a specific implementation, in a French class at the Berlin International School, after the translation host generates a translated audio signal containing target languages such as English, Chinese, Spanish, etc., the intelligent distribution process is started. The system synchronously monitors the real-time spectrum energy distribution of 2.4GHz band channels, 5GHz band WiFi channels and Bluetooth frequency points through a software-defined radio module. Fast Fourier transform is used to identify high-quality idle channels with background noise below -90dBm and occupancy rate <15%. When it is detected that 2.4GHz channel 6 has no burst pulse interference for 300ms, the translation host immediately sends a handshake signal to occupy the channel and establish a dedicated low-latency link with 32 student devices (including 30 earphones, 1 teacher tablet, and 1 assistant device). The master chip copies the original audio stream into 32 independent data copies, each data injected with target device identification code and time synchronization stamp. Dynamic coding optimization is implemented according to the characteristics of the classroom devices: 15 earphones in the European region use LDAC high-fidelity encoding (256kbps), 12 tablets in the American region use aptX HD medium-rate scheme (192kbps), and 5 in-ear devices in the African region use SBC+ forward error correction anti-packet loss mode (128kbps). After the data stream is encrypted by AES-128, it is distributed in parallel through three physically isolated transmission channels: 2.4GHz channel 6 uses time slot division technology to ensure a constant delay of 15ms for English audio, WiFi channel 48 uses frame aggregation to improve French transmission throughput, and Bluetooth frequency point 37 uses adaptive frequency hopping to resist mobile phone signal crosstalk. When a student in the back row of the classroom turns on a mobile hotspot, causing congestion in the 5.8GHz band, the real-time spectrum sensing system captures the mobile interference pulse (center frequency 2442MHz) of channel 48 within 120ms, triggering a seamless switching mechanism: English region devices automatically migrate to DFS channel 165, and historical data packets are retransmitted through regional ARQ protocol, with French audio stream interruption time compressed to within 80ms. Student devices wearing the receiving end automatically perform coding scheme switching, and when the signal strength drops to -75dBm, the device downgrades to an anti-interference mode, ensuring clear recognition of Spanish consonant syllables in a fluctuating signal environment. This embodiment avoids same-frequency device conflicts (such as classroom projector interference) through dynamic channel occupation mechanism, matches terminal performance differences (such as teacher tablet and student earphone rate adaptation) through regional coding strategy, and realizes zero-aware switching within a 30-meter radius through a real-time spectrum sensing system. For example, when the French teacher explains "Les écosystèmes marins sont particulièrement vulnérables" (marine ecosystems are particularly vulnerable), the English transmission link accurately preserves the teacher's throat vibration characteristics, the French channel restores the nasal cavity vowel resonance, and the Spanish audio maintains a high packet integrity rate even in a fluctuating signal environment. The time difference between the three audio signals reaching each student terminal is controlled within 5ms, ensuring the synchronization and fluency of classroom content.

[0093] In one embodiment, the step S4 of identifying the language selection operation of each receiving end and outputting a translated voice signal of a target language version corresponding to the language selection operation comprises:

[0094] S41, obtaining a target language selection instruction sent by a receiving end device;

[0095] S42, based on the target language selection instruction of each receiving end device, filtering a translated audio signal of a corresponding language through a translation host;

[0096] S43, sending each translated audio signal to a corresponding receiving end device.

[0097] In a specific implementation, a receiving end earphone worn by a Canadian student is provided with a physical language switching button on the side. When the student short-presses the button to select French, a target language selection instruction code (0xFR) is generated through a semiconductor chip built in the device, the target language selection instruction code contains a unique ID (such as DEV7A83E) of the device and a time stamp, and is sent to the translation host through a special 2.4G protocol three times at an interval of 5 ms. The host instruction analysis module monitors the data packet in real time, uses a cyclic redundancy check to ensure the integrity of the instruction, and immediately activates the filtering pipeline after detecting a valid instruction. The host maintains a dynamic language mapping table, and binds the device ID with the target language. When the 0xFR instruction of DEV7A83E is detected, a French independent audio track is extracted from the multi-language mixed stream audio (the audio track has been equalized and optimized in step S2). In a high-concurrency scenario where multiple students switch languages at the same time, 500 ms is divided into 100 5 ms micro-slots through a time slot allocation algorithm, and the sending time slot is determined based on the device ID hash value (for example, the last digit of the hash value of DEV7A83E is 3, and the third time slot is occupied). Adaptive audio distribution is performed, and the filtered French audio is subjected to dynamic code rate adjustment: signal strength > -60 dBm uses 24 bit / 48 kHz lossless encoding; -60 dBm to -70 dBm switches to 16 bit / 44.1 kHz lossy encoding; < -70 dBm enables 8 kHz narrowband encoding and forward error correction. The data packet adds a device-specific IP header, and is accurately delivered through a pre-established independent link. The present embodiment compresses the end-to-end response to within 45 ms through direct control of the instruction and independent stream distribution.

[0098] In one embodiment, the step S5 of switching the receiving end device to a sending end mode in response to a mode switching operation of the receiving end device comprises:

[0099] S51, receiving a speech request sent by carrying a unique ID of the receiving end;

[0100] S52, generating an approval pop-up window containing the ID of the receiving end on the interface of the translation host;

[0101] S53, if multiple speech requests exist at the same time, generate a priority queue according to the receiving time and only display the approval pop-up window of the first request in the queue;

[0102] S54, in response to the confirmation operation of the host on the approval pop-up window, switch the communication mode of the receiving end device initiating the speech request to the sending end mode, and activate the audio acquisition function of the receiving end device initiating the speech request.

[0103] In a specific implementation, when the Chinese student (device ID is DEV5C0D3) selects Chinese translation through S4 step, and needs to ask questions, directly long press the earphone side key to trigger mode switching. The physical button distinguishes the operation type at the hardware level: short press (<500ms) is the language selection instruction (i.e. S41), and long press (>1s) is the speaking request instruction (i.e. S51). After the student (device ID is DEV5C0D3) long presses the side key for 1.2 seconds, the device generates an encrypted request package and sends it to the host through the established 2.4G channel 9 (the original Chinese audio transmission link), and the system detects the change of the instruction type. The language selection (S4) and the speaking request (S5) share the same physical link, and the channel 9 bound to the device ID (DEV5C0D3) is automatically converted into a control channel, which can avoid the 150ms delay caused by re-handshake. When the host simultaneously receives the French selection request of the German student (device ID is DEV8E1 F2) and the speaking request of the Chinese student (DEV5C0D3), the priority queue linkage is set up, and the time stamp is (12:15:30.100) and (12:15:30.205) respectively, as shown in Table 1:

[0104] Request Type Device ID Timestamp Queue Bit Language Selection 8E1F2 30.100ms Q1 first Speech Request 5C0D3 30.205ms Q2 first

[0105] The teacher (the main speaker) control panel displays in two columns: the left column is the language approval, and the highlight is 8E1 F2 switching French; the right column is the speaking approval, and 5C0D3 is requesting to speak.

[0106] After the teacher approves the Chinese student to speak, the radio frequency is switched, the channel 9 is converted from the receiving mode to the sending mode (the impedance matching network is automatically tuned to 50Ω), the acoustic link is reconstructed, the original receiving link (the host audio output→the earphone loudspeaker) is switched to the new sending link (the microphone→the preamplifier→the ADC→the host), the host continuously caches the Chinese translation stream during the speaking, and the pushing is automatically restored within 3 seconds after the speaking is finished.

[0107] In one embodiment, the step S6 of obtaining the second language audio signal input by the secondary speaker through the switched device after the mode switching comprises:

[0108] S61, verifying the consistency of the device ID of the receiving end device initiating the speech request and the device ID of the device after the mode switching;

[0109] S62, obtaining the audio signal of the secondary speaker through an audio acquisition unit built in the receiving end device initiating the request.

[0110] In a specific implementation, when the teacher approves the speaking request of the Chinese student (device ID: DEV5C0D3), the security verification and voice acquisition double processes are immediately started. First, the device ID security verification is performed. The translation host calls the dynamic authorization list to compare the device ID initiating the speaking request. The authorized device queue is [DEV5C0D3 (China), DEV8E1 F2 (Germany)], and the current request is ID: DEV_5C0D3. The verification algorithm performs three matching: digital signature decryption (decrypting the elliptic curve signature in the request packet using ECC-256), timestamp validity (detecting that the request time and the approval time difference < 800 ms (anti-replay attack)), and channel binding verification (confirming that the request comes from the same 2.4G channel 9 as the approval time). After the verification is passed, the host sends a security authentication pulse (a 125 kHz carrier signal with a width of 50 ms) to the device initiating the request, triggering the hardware cascade mechanism. The collection of directional voice is performed. After the device receives the authentication pulse, the audio acquisition unit is activated, the MEMS microphone array is powered on (the main microphone applies a +2.8V bias voltage, two groups of auxiliary microphones are turned on, the beamforming algorithm is started, and ±30° pickup beams are generated), the adaptive filter is used to analyze the classroom background noise spectrum in real time (focus on eliminating 200-500Hz air conditioner sound), and the spectral subtraction method is used to weaken the noise energy. When the Chinese student (device ID: DEV5C0D3) finishes speaking and releases the key, the system turns off the microphone bias voltage, releases the beamforming processor resources, and sends a collection termination signal to the host within 0.3 seconds. The host immediately resumes the Chinese translation stream push of the device, and completes the seamless conversion of "question-answer-continue listening to the class".

[0111] In one embodiment, the step S7 of performing real-time translation processing on the second language audio signal and distributing to each receiving end includes:

[0112] S71, identifying the language characteristics of the second language audio signal;

[0113] S72, generating a secondary speaker translation audio signal containing at least one target language;

[0114] S73, distributing to all devices in the receiving mode through the preempted idle channel.

[0115] In a specific implementation, when the Spanish representative (original receiving end device ID ESDEVF2A9) is authorized to speak through S6, the system starts the full-process processing of the secondary speech. When the translation host receives the Spanish representative speech stream, real-time acoustic fingerprint analysis is performed, the fundamental frequency contour is extracted, the typical Spanish intonation pattern (falling tone to 120 Hz at the end of a statement, rising tone to 220 Hz in a question) is detected, and phoneme probability analysis is performed. If the frequency of fricative / θ / is greater than 8 times per second, it is determined to be a Spanish feature (such as the / c / pronunciation in "ciudad"); if the duration of the uvular trill / R / is 120 ms, it is confirmed to be a non-French or Italian language. Perform prosodic model matching based on the syllable timing pattern specific to Spanish (stress interval 350±50 ms) to output the language label within 380 ms. Based on the dynamically identified Spanish label, the engine starts three processing pipelines for English translation, French translation, and German translation in parallel, and according to the speaker's laryngeal resonance characteristics (fundamental frequency range 108-142 Hz), the / g / / k / energy is strengthened in German synthesis. Multiplexing channel scanning data of step S3, dynamically selecting the optimal path, intelligently distributing channels, English stream occupying 5GHz channel 112 (time delay sensitive, transmission jitter <3ms), French stream binding 2.4G channel 9 (high fault tolerance mode, forward error correction ratio 30%), German stream enabling Bluetooth channel 38 (low power transmission, peak current ≤18mA), and the podium device (ID PRES01) is marked as the highest priority, with a bandwidth allocation ratio of 40%, the Spanish representative's own device is automatically muted to avoid acoustic feedback. When the WiFi scanner is disturbed (2.412 GHz noise surge), the system migrates the French stream to 5GHz channel 165 within 80ms, dynamically compresses the German stream code rate to 64kbps, and keeps the English stream in dedicated channel transmission. When the speaking end device switches back to receiving mode, the system automatically releases the 5GHz / 2.4G / Bluetooth channel resources, restores the original language settings of the Spanish device (Spanish translation stream), and updates the chairman's console speaking queue, automatically activates the lower requester.

[0116] Reference Figure 2 , the structure block diagram of the call device suitable for multi-person multi-language translation in an embodiment of the present application comprises:

[0117] The main speaker audio acquisition module is configured to acquire the first language audio signal input by the main speaker through the main speaker device.

[0118] The first translation processing module is configured to perform real-time translation processing on the first language audio signal to generate a translation audio signal containing at least one target language.

[0119] The speech distribution module is configured to distribute the translation audio signal to multiple receiving ends.

[0120] The voice output module is configured to identify a language selection operation of each receiving end and output a translated voice signal in a target language version corresponding to the language selection operation.

[0121] The mode switching module is configured to switch the receiving end device into a sending end mode in response to a mode switching operation of the receiving end device.

[0122] The secondary speaker audio acquisition module is configured to acquire a second language audio signal input by a secondary speaker through the switched device.

[0123] The second translation processing module is configured to perform real-time translation processing on the second language audio signal and distribute the second language audio signal to each receiving end.

[0124] The specific implementation of each module in the above device example is described in the above method embodiment, which will not be described here.

[0125] Referring to Figure 3 The structure schematic block diagram of a call platform suitable for multi-person multi-language translation according to an embodiment of the present application comprises:

[0126] The multi-language interaction management system is configured to acquire a first language audio signal input by a primary speaker through a primary speaker device and perform real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language.

[0127] The dynamic wireless distribution system is configured to distribute the translated audio signal to multiple receiving ends, and the dynamic wireless distribution system comprises:

[0128] The channel scanning module is configured to scan the occupation state of multiple frequency band wireless channels and identify an idle channel.

[0129] The link preemption module is configured to preempt the idle channel to establish an independent transmission link with each receiving end.

[0130] The data stream replication module is configured to replicate the translated audio signal into independent data streams in the same number as the receiving ends.

[0131] The physically isolated transmission module is configured to send the data streams to the corresponding receiving ends through the independent transmission links.

[0132] The user language adaptation system is configured to identify a language selection operation of each receiving end and output a translated voice signal in a target language version corresponding to the language selection operation.

[0133] The role switching control system is configured to switch the receiving end device into a sending end mode in response to a mode switching operation of the receiving end device.

[0134] The speech collection system of the assistant speaker is used to obtain the second language audio signal input by the assistant speaker through the device after switching mode.

[0135] The multi-language parallel translation system is used to perform real-time translation processing on the second language audio signal and distribute to each receiving end.

[0136] In summary, the present application obtains the first language audio signal input by the main speaker through the main speaker device, performs real-time translation processing on the first language audio signal to generate a translation audio signal containing at least one target language, distributes the translation audio signal to a plurality of receiving ends, identifies the language selection operation of each receiving end, outputs the translation speech signal of the target language version corresponding to the language selection operation, responds to the mode switching operation of the receiving end device to switch the receiving end device to the sending end mode, obtains the second language audio signal input by the assistant speaker through the device after switching mode, and performs real-time translation processing on the second language audio signal and distributes to each receiving end, to realize seamless connection of real-time translation and voice interaction in a multi-person multi-language scene, and build an efficient and reliable full-duplex cross-language communication purpose.

[0137] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium provided by the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchl ink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.

[0138] It is to be understood that the terminology "including", "comprising", or any other variation thereof, is intended to cover a non-exclusive inclusion such that process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0139] The above description is merely the preferred embodiments of the present application, and is not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made according to the content of the present application specification and drawings, or directly or indirectly applied to other related technical fields, is also included in the patent protection scope of the present application.

Claims

1. A call translation method suitable for multiple users and multiple languages, characterized in that, Includes the following steps: The first language audio signal input by the speaker is obtained through the main speaking device; The audio signal in the first language is translated in real time to generate a translated audio signal containing at least one target language; The translated audio signal is distributed to multiple receiving terminals; Identify the language selection operation of each receiving end and output the translated speech signal corresponding to the target language version of the language selection operation; In response to the mode switching operation of the receiving device, the receiving device is switched to the transmitting mode; The device, after switching modes, acquires the second language audio signal input by the co-speaker; The audio signal in the second language is translated in real time and distributed to each receiving end.

2. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The step of acquiring the first language audio signal input by the speaker through the speaker device includes: Collect the speaker's voice and identify the original language; The first language audio signal carrying the language tag is transmitted to the translation host.

3. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The step of performing real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language includes: Based on the language tags, a translation engine is selected to generate translated text in at least one target language; The translated text is then subjected to speech synthesis to generate a translated audio signal in the corresponding target language.

4. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The step of distributing the translated audio signal to multiple receivers includes: Scan the occupancy status of multi-band wireless channels and identify idle channels; Preempting the idle channel to establish an independent transmission link with each receiving end; The translated audio signal is copied into the same number of independent data streams as the receiving end; Data streams are sent to the corresponding receiving end through each independent transmission link.

5. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The step of identifying the language selection operation of each receiving end and outputting the translated speech signal corresponding to the target language version of the language selection operation includes: Obtain the target language selection instruction sent by the receiving device; Based on the target language selection instructions of each receiving device, the translation host filters the corresponding language translation audio signals; Each translated audio signal is sent to its corresponding receiving device.

6. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The response receiving device mode switching operation, the step of switching the receiving device to the transmitting device mode, includes: Receive a speech request that carries the receiver's unique ID; Generate an approval pop-up window containing the receiver's ID on the translation host interface; If multiple speech requests exist simultaneously, a priority queue is generated based on the time of receipt, and only the approval pop-up of the first request in the queue is displayed; In response to the translation host's confirmation of the approval pop-up, the communication mode of the receiving device that initiated the speech request is switched to the sending mode, and the audio acquisition function of the receiving device that initiated the speech request is activated.

7. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The step of acquiring the second language audio signal input by the co-speaker through the device after switching modes includes: Verify the consistency between the receiving device ID that initiated the speech request and the device ID that has switched modes; The audio signal of the co-speaker is acquired through the built-in audio acquisition unit of the receiving device that initiates the speech request.

8. The call method for multi-person, multilingual translation according to claim 1, characterized in that, The step of performing real-time translation processing on the second language audio signal and distributing it to each receiving end includes: Identify the linguistic features of second language audio signals; Generate a sub-narration audio signal containing at least one target language; Distributed to all devices in receive mode via preempted idle channels.

9. A communication device suitable for multi-person, multilingual translation, characterized in that, include: The main speaker audio acquisition module is used to acquire the first language audio signal input by the speaker through the main speaker device; The first translation processing module is used to perform real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language; A voice distribution module is used to distribute the translated audio signal to multiple receiving terminals; The voice output module is used to identify the language selection operation of each receiving end and output the translated voice signal corresponding to the target language version of the language selection operation; The mode switching module is used to respond to the mode switching operation of the receiving device and switch the receiving device to the transmitting mode. The co-presenter audio acquisition module is used to acquire the second language audio signal input by the co-presenter through the device after switching modes; The second translation processing module is used to perform real-time translation processing on the audio signal in the second language and distribute it to each receiving end.

10. A call platform suitable for multi-person, multilingual translation, characterized in that, include: A multilingual interactive management system is used to acquire the first language audio signal input by the speaker through the speaker device, and to perform real-time translation processing on the first language audio signal to generate a translated audio signal containing at least one target language. A dynamic wireless distribution system is used to distribute the translated audio signal to multiple receiving terminals, the dynamic wireless distribution system comprising: The channel scanning module is used to scan the occupancy status of multi-band wireless channels and identify idle channels; The link preemption module is used to preempt the idle channel to establish independent transmission links with each receiving end; The data stream copying module is used to copy the translated audio signal into the same number of independent data streams as the receiving end; The physically isolated transmission module is used to send data streams to the corresponding receiving end through each independent transmission link; The user language adaptation system is used to identify the language selection operation of each receiving end and output the translated speech signal corresponding to the target language version of the language selection operation. A role switching control system is used to respond to the mode switching operation of the receiving device and switch the receiving device to the transmitting mode. The assistant speaker's voice acquisition system is used to acquire the second language audio signal input by the assistant speaker through a device that has switched modes. A multilingual parallel translation system is used to perform real-time translation processing on audio signals in the second language and distribute them to various receiving terminals.

Citation Information

Patent Citations

  • Network-based simultaneous interpretation method and system

    CN110677406A

  • Multifunctional audio intelligent optimization conference system

    CN119135832A