Voice processing method and system for multi-site parallel speech

By combining voiceprint separation and identity binding, speech prosody feature extraction and multilingual synthesis with temporal scheduling and alignment technology, the problems of speech mixing and translation style distortion in parallel speaking in multi-venue remote conferencing systems have been solved. This has enabled semantically accurate cross-venue information transmission and natural cross-lingual interaction, thereby improving the efficiency of meeting communication.

CN122135727APending Publication Date: 2026-06-02SOYO TECH DEV CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOYO TECH DEV CO LTD
Filing Date
2026-04-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing multi-venue remote conferencing systems suffer from several problems when handling parallel speeches, including difficulty in separating mixed speech, inability to preserve the rhythm of the source language in multilingual translation leading to awkward interaction, weak correlation between speaker identity and content, and chaotic parallel information transmission. These issues affect the accuracy and efficiency of cross-venue communication.

Method used

By performing voiceprint separation and identity binding on audio streams from multiple venues, and combining speech prosody feature extraction and multilingual synthesis, semantic accuracy and speech style fidelity are achieved in cross-venue multilingual translation scenarios. In terms of venue language, the audio streams and texts of the reference language are scheduled and aligned in a temporal manner to achieve orderly playback and synchronous presentation of parallel speeches from different venues and speakers in the same target language.

Benefits of technology

It effectively solves problems such as mixed speech in parallel speaking, distorted translation style, and unclear association between identity and content, and significantly improves the clarity of information transmission, the naturalness of cross-language interaction, and the efficiency of meeting communication in multi-venue parallel speaking scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135727A_ABST
    Figure CN122135727A_ABST
Patent Text Reader

Abstract

This application discloses a speech processing method and system for parallel speaking in multiple meeting rooms. The method includes: acquiring multiple audio streams from multiple meeting rooms through multiple meeting boxes; extracting style features of the audio streams; translating the audio streams according to the style features and preset meeting room language requirements to generate multiple language text and audio streams; detecting that multiple target audio streams are speaking in parallel; determining multiple reference language text and audio streams corresponding to the multiple target audio streams; performing a timing scheduling operation on the multiple reference language audio streams to obtain time-interleaved multiple target language audio streams, and performing a timing alignment operation on the multiple reference language texts to obtain multiple target language texts; and pushing the multiple target language audio streams and texts to the multiple meeting boxes respectively. This application can solve the problems of speech confusion and translation style distortion in multi-meeting parallel speaking scenarios, improving the communication efficiency and experience of remote meetings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of speech or sound processing technology, or the fields of electronic digital data processing technology, or the fields of digital information transmission technology, and in particular to a speech processing method and system for multi-venue parallel speaking. Background Technology

[0002] Multi-level, multi-venue remote conferencing needs to support two-way interaction, multilingual communication, and multi-person discussions. Existing multi-venue remote conferencing systems suffer from problems when handling parallel speeches, such as difficulty in separating mixed speech, inability to preserve the rhythm of the source language in multilingual translation leading to awkward interaction, weak correlation between speaker identity and content, and chaotic parallel information transmission. These issues seriously affect the accuracy and efficiency of cross-venue communication. Summary of the Invention

[0003] In view of this, embodiments of this application provide a speech processing method and system for parallel speaking in multiple venues. By performing voiceprint separation and identity binding on audio streams from multiple venues, combined with speech prosody feature extraction and multilingual synthesis, semantic accuracy and speech style fidelity are achieved in cross-venue multilingual translation scenarios. Furthermore, by performing temporal scheduling and alignment of audio streams and texts in reference languages ​​at the venue language level, orderly playback and synchronous presentation of parallel speaking content from different venues and speakers in the same target language are achieved. Additionally, while ensuring the main speaker's speaking rights and audio-visual experience, the entire process of multi-venue parallel speaking conference content is traceable. This effectively solves problems such as mixed speech in parallel speaking, translation style distortion, and ambiguous association between identity and content, significantly improving the clarity of information transmission, the naturalness of cross-language interaction, and the efficiency of conference communication in multi-venue parallel speaking scenarios.

[0004] In a first aspect, embodiments of this application provide a voice processing method for parallel speaking across multiple meeting rooms, applied to a server of a business management platform in a remote conferencing system. The system further includes multiple conferencing terminals across multiple meeting rooms, each conferencing terminal comprising a conferencing box, the conferencing box being communicatively connected to the server. The method includes: Multiple audio streams from the multiple meeting rooms are acquired through multiple conference boxes; For each of the plurality of audio streams, the following processing is performed: Extracting style features from the audio stream, the style features characterizing the speech prosody of the audio stream; and, The audio stream is translated based on the style characteristics and the preset venue language requirements information to generate multi-language text and multi-language audio streams for the one-to-one requirements of the multiple venues. Multiple target audio streams were detected to be speaking in parallel sequence among the multiple audio streams. Determine the multi-reference language text and multi-reference language audio streams corresponding to the multiple target audio streams; A timing scheduling operation is performed on the multi-reference language audio streams to obtain a timing-interleaved multi-target language audio stream, and a timing alignment operation is performed on the multi-reference language text to obtain a multi-target language text; The multi-channel target language audio streams and the multi-channel target language text streams are respectively pushed to the multiple meeting boxes.

[0005] Secondly, this application also provides a remote conferencing system, which includes a server of a business management platform and multiple conferencing terminals in multiple meeting rooms. Each conferencing terminal includes a display screen, a playback device, and a conferencing box. The conferencing box is communicatively connected to the server, wherein the server is used to execute the steps in the first aspect of the embodiments of this application.

[0006] Thirdly, embodiments of this application provide an electronic device, including a processing module, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processing module, and the programs include instructions for performing the steps in the first aspect of embodiments of this application.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of embodiments of this application.

[0008] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of embodiments of this application. The computer program product may be a software installation package.

[0009] As can be seen, the speech processing method and system for parallel speaking in multiple venues provided in this application involves the server acquiring multiple audio streams from multiple venues through multiple conference boxes; performing the following processing on each audio stream: extracting style features of the audio stream; translating the audio stream based on the style features and preset venue language requirements to generate multi-language text and multi-language audio streams for one-to-one requirements of multiple venues; detecting that multiple target audio streams are speaking in parallel in the speaking time sequence; determining the multi-reference language text and multi-reference language audio streams corresponding to the multiple target audio streams; performing a timing scheduling operation on the multi-reference language audio streams to obtain a timing-interleaved multi-target language audio stream, and performing a timing alignment operation on the multi-reference language text to obtain multi-target language text; and pushing the multi-target language audio streams and multi-target language texts to the multiple conference boxes respectively. Thus, this application effectively solves the problems of mixed voices in parallel speeches, distorted translation style, and unclear association between identity and content, and significantly improves the clarity of information transmission, the naturalness of cross-language interaction, and the efficiency of meeting communication in multi-venue parallel speaking scenarios. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic diagram of the architecture of a remote conferencing system provided in an embodiment of this application; Figure 2 This is a schematic flowchart of a speech processing method for parallel speaking in multiple venues provided in an embodiment of this application; Figure 3 This is a flowchart illustrating a timing scheduling operation on a language audio stream, as provided in an embodiment of this application. Figure 4 This is a schematic diagram illustrating a scenario of performing timing scheduling operations on a language audio stream, as provided in an embodiment of this application. Figure 5 This is a flowchart illustrating a method for performing temporal alignment operations on language text, as provided in an embodiment of this application. Figure 6 This is a schematic diagram of a conference terminal display interface provided in an embodiment of this application; Figure 7 This is a block diagram of the functional units of a remote conferencing system provided in an embodiment of this application; Figure 8This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0013] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0014] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article indicates that the preceding and following related objects have an "or" relationship.

[0015] In this application's embodiments, "multiple" refers to two or more. In this application's embodiments, "connection" refers to various connection methods, such as direct or indirect connections, to achieve communication between devices; this application's embodiments do not impose any limitations on this.

[0016] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0017] The following describes the relevant content, concepts, meanings, technical issues, technical solutions, and beneficial effects involved in the embodiments of this application.

[0018] Existing multi-venue remote conferencing systems suffer from several problems when handling parallel speeches, including difficulty in separating mixed speech, inability to preserve the rhythm of the source language in multilingual translation leading to awkward interaction, weak correlation between speaker identity and content, and chaotic parallel information transmission. These issues severely affect the accuracy and efficiency of cross-venue communication.

[0019] To address the aforementioned issues, this application provides a speech processing method and system for multi-venue parallel speaking. It aims to achieve semantic accuracy and speech style fidelity in cross-venue multilingual translation scenarios by performing voiceprint separation and identity binding on multi-venue audio streams, combined with speech prosody feature extraction and multilingual synthesis. Furthermore, by temporally scheduling and aligning the audio streams and text of reference languages ​​at the venue language level, it enables the orderly playback and synchronous presentation of parallel speaking content from different venues and speakers in the same target language. Finally, it ensures the main speaker's speaking rights and audio-visual experience while achieving full traceability of the conference content during multi-venue parallel speaking. This effectively solves problems such as mixed speech in parallel speaking, translation style distortion, and ambiguous association between identity and content, significantly improving the clarity of information transmission, the naturalness of cross-language interaction, and the efficiency of conference communication in multi-venue parallel speaking scenarios.

[0020] First, combined Figure 1 The remote conferencing system in the embodiments of this application will be described. Figure 1 This is a schematic diagram of the architecture of a remote conferencing system provided in an embodiment of this application, such as... Figure 1 As shown, the system includes a main venue and at least one sub-venue. The main venue is equipped with a conference terminal 100 and a business management platform 200. Each sub-venue is equipped with a conference terminal 100. The business management platform 200 is communicatively connected to multiple conference terminals 100.

[0021] Specifically, the conference terminal 100 includes a display screen 110, a playback device 120, and a conference box 130, and the business management platform 200 includes an operation and maintenance console 210 and a server 220, with the server 220 communicating with the conference terminal 100 and the operation and maintenance console 210 respectively.

[0022] The conference terminal 100 is the core sensing and presentation terminal deployed in each meeting room, used to collect audio and video data from the meeting room and execute playback and display control commands issued by the server 220. The display screen 110 is the visual presentation terminal for the meeting room, used to display multi-channel target language text, speaker identity information, split-screen layout content, and meeting status information in real time, and supports dynamic adjustment of the display area and content according to the partition display commands from the server 220. The playback device 120 is the audio output unit for the meeting room, used to play multi-channel target language audio streams processed by the server 220 through timing scheduling, including the main speaker's audio and interleaved audio from other speakers, and supports local execution and locking of volume parameters. The playback device 120 may include simultaneous interpretation headphones and speakers. The conference box 130 is the edge processing and communication node of the conference terminal 100. It is responsible for collecting the initial audio stream of the conference room, performing preliminary noise reduction and format conversion, and uploading the audio data to the server 220. At the same time, it receives playback control instructions and split-screen display instructions from the server 220, and drives the playback device 120 and the display screen 110 to perform synchronous playback and partitioned display.

[0023] Among them, the business management platform 200 is usually deployed in the main venue and is the core management and processing hub for multi-organization and multi-venue remote conferencing. The server 220 is responsible for core operations such as unified acquisition of audio streams from multiple venues, style extraction, multilingual translation, parallel speaking detection, audio scheduling, text alignment, split-screen decision-making, and abnormal behavior intervention. The operation and maintenance console 210 includes a human-computer interaction interface, which is used by meeting administrators to configure meeting parameters, manage participants, view system status, adjust language requirements, set volume thresholds and interleaving timing, and issue meeting strategies and control commands to the server 220.

[0024] As can be seen, in this embodiment, the remote conferencing system supports a multi-person interactive discussion mode. In this mode, the meeting box 130 of each meeting room collects the voices of the current participants, the server 220 aggregates the voices of all participants, processes them, and pushes them to the meeting box 130 of each meeting room according to the language requirements of each meeting room. The meeting room playback device 120 then plays the audio, and the display screen 110 displays the audio and text synchronously. This achieves high-definition, high-synchronization, and high-naturalness in information transmission, effectively ensuring smooth communication and interactive experience in cross-language, multi-speaker, and multi-meeting meetings.

[0025] The following is combined with Figure 2 The following section further explains the speech processing method for parallel speaking in multiple venues provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating a speech processing method for parallel speaking in multiple venues, provided in an embodiment of this application. It is applied to... Figure 1 Server 220 in the middle, such as Figure 2 As shown, the method includes the following steps: Step S210: Acquire multiple audio streams from the multiple meeting rooms through multiple conference boxes.

[0026] The audio stream consists of digitized audio data representing the speech signals within each meeting room, collected and uploaded by the meeting boxes in each room. The audio stream exists in time-series format, and its content originates from real-time speech within the meeting room. It can cover continuous speech segments from a single speaker or mixed speech signals from multiple speakers.

[0027] In this embodiment, before acquiring the audio stream, the conference terminal registration and parameter configuration operations must be performed first. The server receives the hardware identifier, affiliated organization, preset language requirements, and network connection information of each conference terminal. It verifies the legitimacy of the terminal through identity authentication mechanisms such as encryption key verification. After completing the terminal registration, it configures basic parameters such as default transmission protocol and audio encoding format, and pushes operating instructions such as microphone array sensitivity and speaker volume reference value. It establishes an encrypted communication link between the terminal and the server, and generates a list of registered terminals containing the conference room hierarchy. This provides a legitimate, connected, and parameter-adapted foundation for the server to stably acquire multi-conference audio streams through the conference box.

[0028] In one possible embodiment, acquiring multiple audio streams from the multiple meeting rooms through multiple conference boxes includes: acquiring multiple initial audio streams from the multiple meeting rooms through the multiple conference boxes, each initial audio stream including a mixed audio stream from multiple speakers; preprocessing the mixed audio streams in the multiple initial audio streams according to a pre-recorded participant voiceprint database to obtain multiple audio streams corresponding one-to-one with the multiple speakers, wherein the preprocessing includes speech separation processing and voiceprint identity binding, and each audio stream is associated with a single speaker's identity identifier, meeting room information, voiceprint feature identifier, speaking start timestamp, and style features.

[0029] The multiple initial audio streams consist of raw, unprocessed audio data collected from different meeting rooms by multiple meeting boxes using their built-in microphone arrays. The single initial audio stream is the audio of all speakers within a single meeting room, similar to the superposition of multiple people speaking simultaneously in the same meeting room.

[0030] In one possible embodiment, after acquiring multiple initial audio streams from multiple meeting rooms through the multiple conference boxes, the method further includes: performing standardized preprocessing and format encapsulation on the multiple initial audio streams. Specifically, this includes uniformly performing noise reduction, echo cancellation, and audio format standardization processing on the multiple initial audio streams acquired from each meeting room, removing background noise and redundant signals, adjusting parameters such as sampling rate and quantization bit depth, and adding meeting room identifiers, timestamps, and sound source localization information to each frame of audio.

[0031] The participant voiceprint database mainly includes unique identifiers for participants (such as participant ID, name, position / department information), voiceprint feature data (such as speech spectrum features, Mel-frequency cepstral coefficients (MFCC), fundamental frequency (pitch), energy distribution), association information corresponding to the venue and speaking scenario (such as venue number, venue location), and collection and update timestamps. By binding the speaker's identity, voiceprint features, venue location, and time information, it achieves accurate speech recognition and association management with one voiceprint per person and traceable audio.

[0032] Specifically, speech separation processing refers to processing a single initial audio stream (mixed audio) to separate the mixed signal into independent audio streams corresponding to each speaker, thus achieving separation from a multi-person mixed audio stream to a single-person audio stream. Speech separation processing techniques include, but are not limited to, those based on deep learning models, such as the improved Conv-TasNet model.

[0033] Specifically, voiceprint identity binding refers to extracting voiceprint features from the separated single-person audio stream, comparing them with multiple voiceprint features of multiple speakers in the participant voiceprint database, and identifying and binding the speaker identity information corresponding to the single-person audio stream.

[0034] In one possible embodiment, after obtaining the plurality of audio streams corresponding one-to-one with the plurality of speakers, the method further includes: performing audio purification calibration processing on the plurality of audio streams. Specifically, this includes performing secondary noise reduction on the separated individual speech streams to completely filter out residual background noise and speech reverberation; simultaneously, based on a preset volume benchmark, performing volume equalization calibration on each audio stream to unify the loudness and gain of each audio stream, and finally outputting a standardized audio stream with a compliant signal-to-noise ratio and uniform volume.

[0035] For example, a single audio stream is associated with a single speaker's identity, meeting room information, voiceprint feature identifier, speaking start timestamp, and style features. For example, "Speaker A001" (corresponding to participant Zhang San) - "Sub-venue 3 - Meeting Room 5" - "Voiceprint Feature ID005" - "2026-02-06 10:20:30.500" - "Smooth tone, moderate speaking speed".

[0036] As can be seen, in this embodiment, the initial audio streams of multiple meeting rooms are collected by the conference box, and the voiceprint database is used to complete the voice separation and identity binding. This achieves the accurate association between a single audio stream and a single speaker, and realizes the one-to-one binding of the audio stream with the speaker's identity, meeting room information, timestamp, and style characteristics. This effectively solves the problems of difficulty in distinguishing identities, difficulty in tracing audio sources, and inconsistent volume affecting subsequent processing when multiple people are mixed in multiple meeting rooms.

[0037] Step S220: For each of the multiple audio streams, perform the following processing: extract the style features of the audio stream; and translate the audio stream according to the style features and the preset venue language requirement information to generate multi-language text and multi-language audio streams for the one-to-one requirements of the multiple venues.

[0038] The style features described herein represent the phonological prosody of the audio stream. For example, the intonation, rhythm, and pause patterns of American / British pronunciation.

[0039] In one possible embodiment, translating the audio stream according to the style features and preset venue language requirement information to generate multi-language text and multi-language audio streams for the one-to-one requirements of multiple venues includes: converting the audio stream into first speech text; determining multiple target languages ​​for the one-to-one requirements of multiple venues according to the preset venue language requirement information; calling a pre-trained multilingual translation model to process the first speech text, generating multi-language texts corresponding one-to-one with the multiple target languages; and fusing the style features with the multi-language texts based on a speech generation algorithm to generate multi-language audio streams corresponding one-to-one with the multi-language texts, wherein the multi-language audio streams are multiple audio streams that mimic the style features of the audio stream and are in the form of the multiple target languages, and the style features include intonation, timbre, rhythm, and pause patterns.

[0040] The preset language requirements for each meeting venue are core configuration parameters that are pre-configured and linked to the meeting venue and participants. These parameters include, but are not limited to, four main categories: target language, meeting venue scenario, language priority, and output format. Specifically, this includes the target output language for each venue (e.g., Chinese, English, German), the meeting venue's business scenario (e.g., conference speeches, news broadcasts, interviews), language usage priority (primary language / backup language), and audio output format requirements such as sampling rate and speaking speed. This information is used to accurately match the language service needs of different meeting venues.

[0041] In addition to intonation, timbre, rhythm, and pause patterns, stylistic features may also include speech rate, stress distribution, tone strength, accent characteristics (unique regional or personal accent characteristics of the speaker), vibrato / breathiness, and other vocal characteristics, as well as speech rhythm features such as the mode of resonance (differences in oral, nasal, and chest resonance).

[0042] Understandably, the process involves first converting the original audio into text, then translating it into the target language based on the needs of the meeting. Next, stylistic features such as intonation, timbre, rhythm, and pauses are extracted from the original audio. These features are then integrated with the translated text using a speech generation algorithm to generate an audio stream that retains the original audio's tone but uses the target language's pronunciation. This ensures semantic accuracy while restoring the emotion and rhythm of the original speech. For example, if speaker A suddenly switches from American to British pronunciation in English, the audio data in the Chinese audio channel can reflect this difference. The Chinese audio channel can demonstrate the transition from American to British pronunciation through features such as rhythm and intonation, allowing Chinese listeners to indirectly perceive the change in the speaker's voice style and enhancing the naturalness of the interaction.

[0043] As can be seen, in this embodiment, by extracting the style features of the audio stream and combining them with the language requirements of the venue, multilingual text translation is completed. The style features are then fused with the target language text using a speech generation algorithm to generate a multilingual audio stream that restores the rhythm and emotion of the audio. This enables precise one-to-one speech delivery across venues, ensuring semantic accuracy while fully preserving the expressive style and emotional characteristics of the original speech. This solves the problems of stiff tone and loss of rhythm in traditional translation.

[0044] Step S230: It is detected that multiple target audio streams are in parallel in terms of speaking time among the multiple audio streams.

[0045] In this embodiment, the speaking time sequence of multiple audio streams is monitored in real time using audio stream timestamp alignment and sound source localization technology, identifying overlapping segments of multiple speakers in the time dimension. Specifically, the system first assigns precise timestamps and sound source location markers to each audio stream, and then distinguishes speech segments of different speakers using a voiceprint recognition model. Next, through frame-level time sequence comparison, it determines whether there are two or more valid speech frames of audio streams overlapping in the same time window. If the overlap duration exceeds a preset threshold, it is determined to be "parallel speaking time sequence".

[0046] Step S240: Determine the multi-reference language text and multi-reference language audio streams corresponding to the multiple target audio streams.

[0047] In this embodiment of the application, based on the multiple reference language texts and audio streams that correspond one-to-one with the original audio streams generated in step S220, the multiple target audio streams are audio streams that are parallel in speaking time among multiple (original) audio streams. It is only necessary to accurately filter the multiple reference language texts and audio streams that correspond one-to-one with the original audio streams according to the identifier / timestamp / participant identity of the target audio streams, and extract the reference language texts and audio streams that correspond one-to-one with the target audio streams.

[0048] Step S250: Perform a timing scheduling operation on the multi-reference language audio streams to obtain a timing-interleaved multi-target language audio stream, and perform a timing alignment operation on the multi-reference language text to obtain a multi-target language text.

[0049] Understandably, in parallel speaking scenarios (such as when A speaks while B interrupts), there are timestamp discrepancies and differences in speaking rhythm in the audio recordings of different speakers, and there are also issues with the synchronization between the text content and the audio. At the same time, the lack of a unified timeline anchor point between multilingual text and the corresponding language audio leads to problems such as audio and text misalignment and ambiguity in language attribution when identifying languages ​​and associating content in the future, which directly affects the meeting experience.

[0050] Specifically, performing temporal interleaving processing on multiple reference language audio streams essentially involves reconstructing and arranging the audio streams of different speakers and languages ​​based on a globally unified timeline benchmark. This breaks the linear playback logic of a single audio stream and, according to the real time distribution characteristics of parallel speaking scenarios, interleaves and merges multiple audio streams along the time dimension, ultimately generating multiple target language audio streams that retain the original vocal characteristics of each speaker and conform to the actual speaking rhythm.

[0051] Specifically, performing time-series alignment processing on multi-reference language texts essentially involves calibrating and associating the text streams of different speakers and languages ​​based on a globally unified timeline anchor point. This breaks the linear association logic of single-stream texts and precisely binds the multi-stream texts to their corresponding audio streams and speaker identities according to the real time distribution characteristics of parallel speaking scenarios. This solves the problems of text and audio time sequence misalignment and ambiguous content attribution in parallel speaking scenarios, ultimately generating multi-stream target language texts that are semantically coherent, clearly attributed, and strictly synchronized with the audio stream.

[0052] Step S260: Push the multi-target language audio streams and the multi-target language text streams to the multiple conference boxes respectively.

[0053] The conference box is used to control the playback device to play the target language audio stream and to control the display screen to display the text content corresponding to different speakers in the target language text in partitions.

[0054] In one possible embodiment, the step of pushing the multi-channel target language audio streams and the multi-channel target language text streams to the multiple conference boxes respectively includes: determining the main speaker who speaks earlier and other speakers who speak later based on the speaking time sequence of the speakers; determining the first voice-text content corresponding to the main speaker and other voice-text content corresponding to the other speakers in the multi-channel target language text streams; determining target partition display information based on the number of speakers and the speaking time sequence, wherein the target partition display information includes the first voice-text content and the other voice-text content displayed in partitions on the display screen. The system determines the regional location, area, and display rules; generates multiple first control commands corresponding one-to-one with the multiple meeting rooms. Each first control command carries multiple target language audio streams and multiple target language text streams pointing to a single meeting room and matching the target language requirements of that single meeting room, as well as the target partition display information. The first control command is used to instruct the playback device of the meeting room to play the multiple target language audio streams and to instruct the display screen to synchronously display the multiple target language text streams according to the target partition display information; and sends the multiple first control commands to the multiple meeting boxes corresponding to the multiple meeting rooms.

[0055] Specifically, the target partition display information is determined based on the number of speakers and the speaking sequence, including: first, dividing the speaking roles according to the parallel relationship of the speaking sequence to determine the hierarchical relationship between the main speaker and other speakers; then, matching the corresponding display partition layout based on the number of speakers: when there is only one speaker, a full-screen single-area display is used to suit single-person presentation scenarios; when there are 2-3 speakers, a double / triple partition layout with left and right columns or top and bottom columns is used to clarify the display hierarchy between the main speaker and other speakers; when there are 4 or more speakers, a grid partition layout is used, allocating the size and position of each area according to the speaking sequence priority. Simultaneously, the partitioning rules are dynamically adjusted based on the speaking sequence. For example, the main speaker corresponds to the core display area, occupying a larger display area first, while other speakers are allocated secondary areas according to their speaking order or importance. The area position and area are dynamically updated with the speaking sequence to ensure that the core speaking content is always in the visual focus, while clearly presenting the text information of all speakers, achieving synchronous matching and visualization of the audio stream and text stream.

[0056] In one possible embodiment, before determining the target partition display information based on the number of different speakers and the speaking sequence, the method further includes: performing text enhancement processing on the first speech text content and the other speech text content, the text enhancement processing including key information highlighting and contextual semantic annotation.

[0057] Specifically, key information is highlighted, including extracting core semantic keywords (such as core viewpoints, key data, and instructions / actions) from the text content of different speakers, and setting differentiated display styles according to the speaker's identity. For example, interactive keywords (such as "disagree," "supplement," "agree," and "refute" when B interjects) and scenario-related keywords (such as "key points," "pending confirmation," and "execution" in a meeting) are highlighted to clearly distinguish the speaker's attitude, actions, and core demands.

[0058] It should be noted that this application only provides one way to determine the target partition display information, including but not limited to dividing the speaking roles according to the parallel relationship of speaking time to clarify the hierarchical relationship, matching the full-screen single area / double / triple area / grid partition display layout based on the number of speakers, and dynamically adjusting the position and area of ​​the core display area and the secondary area in combination with the speaking time, so as to finally achieve synchronous matching and visualization synchronization of audio stream and text stream.

[0059] In one possible embodiment, the method further includes: detecting that at least one of the other speakers has a gradually increasing speaking volume and the absolute value of the volume difference between the speaker's current speaking volume and the main speaker's volume is less than a first preset threshold, and / or detecting that the at least one speaker has engaged in malicious volume adjustment, the malicious volume adjustment including lowering the main speaker's audio playback volume without obtaining meeting management permissions, the audio playback volume being the playback volume of the multi-channel target language audio stream corresponding to the main speaker transmitted from the meeting box to the playback device; sending a second control command to the meeting box, the second control command being used to instruct the playback device to play the main speaker and / or the multi-channel target language audio stream corresponding to the at least one speaker at a preset locked volume.

[0060] The preset volume lock playback feature involves the system pre-setting fixed volume thresholds and playback curves for different speakers' audio streams. This instructs playback devices to output audio according to these fixed values, avoiding auditory interference from fluctuating volume and preventing volume imbalance through a unified volume benchmark. Locking the volume also prevents malicious adjustments during playback, ensuring stable and consistent audio volume between the main speaker and other speakers, thus guaranteeing the consistency and controllability of conference audio playback. A pop-up window on the display screen will indicate "Abnormal volume adjustment detected" when volume lock playback is activated.

[0061] The malicious volume adjustment operations also include increasing the audio playback volume of at least one speaker without obtaining meeting management privileges, frequently and quickly switching volume levels, repeatedly triggering volume mute / unmute, and intentionally decreasing the audio playback volume of the main speaker or intentionally increasing the audio playback volume of at least one speaker when obtaining meeting management privileges.

[0062] This application does not restrict compliant volume adjustment behaviors, including but not limited to: adjusting the global volume according to meeting rules after obtaining management privileges, adjusting one's own speaking volume within the scope of one's speaking privileges, temporarily balancing the volume of each channel according to the needs of the meeting scenario, and compliant mute / unmute operations (such as unmute after speaking, or temporarily mute to avoid background noise). The above compliant operations all fall within the scope of normal meeting audio management and are not subject to the locking and restriction mechanisms in this solution; however, operations without authorization, intended to disrupt audio balance, or interfere with the normal speaking of the main speaker are all considered malicious volume adjustment operations as defined in this solution and will trigger the preset volume locking and control mechanisms.

[0063] In one possible embodiment, after pushing the multi-target language audio streams and the multi-target language text streams to the multiple meeting boxes respectively, the method further includes: generating a complete meeting record with identity annotations based on the multiple audio streams, the speaking sequence, the multi-target language audio streams and the multi-target language text streams, wherein the complete meeting record includes meeting text content, speaker identity annotations, audio segments and interaction relationship annotations.

[0064] Specifically, generating a complete meeting record with identity annotations involves the server taking the separated individual audio streams output in step S210, the target language text output in step S220, and the speaking time sequence and identity association data output in step S250 as inputs. It then integrates all meeting information along a global timeline: accurately associating each speaker's identity with their corresponding original audio, target language audio, and text content, forming a timestamped identity-audio-text correspondence; dividing speaking segments based on pauses and topic changes; and annotating interactive relationships such as interruptions and responses in parallel speaking scenarios to clarify the interaction logic between speakers; ultimately generating a structured meeting record that supports multi-dimensional retrieval by speaker, speaking time, and core keywords, outputting a complete meeting record with identity annotations, audio clips, and interactive relationship annotations, achieving full-dimensional traceability and convenient retrieval of meeting content.

[0065] As can be seen, this embodiment achieves semantic accuracy and faithfulness of speech style in cross-venue multilingual translation scenarios by performing voiceprint separation and identity binding on multi-venue audio streams, combined with speech prosody feature extraction and multilingual synthesis. Furthermore, by temporally scheduling and aligning the audio streams and texts of the reference language at the venue language level, it enables the orderly playback and synchronous presentation of parallel speeches from different venues and speakers in the same target language. Finally, while ensuring the main speaker's speaking rights and audio-visual experience, it achieves full traceability of the conference content from multiple venues. This effectively solves problems such as mixed speech in parallel speeches, distorted translation style, and ambiguous association between identity and content, significantly improving the clarity of information transmission, the naturalness of cross-language interaction, and the efficiency of conference communication in multi-venue parallel speaking scenarios.

[0066] Please refer to details. Figure 3 , Figure 3 This is a flowchart illustrating a timing scheduling operation on a language audio stream, as provided in an embodiment of this application. Figure 3 As shown, the timing scheduling operation on the multi-reference language audio streams includes the following steps: S301, Align the multiple reference language audio streams according to the speaking sequence.

[0067] Among them, the single-channel reference language audio stream is a single audio stream generated by the translation operation of the initial audio stream of a single speaker, which matches the language required by a single meeting venue.

[0068] Specifically, using the system's global clock as a unified timeline, the effective speech frame timestamps of speakers in each reference language audio stream are extracted. The multiple reference language audio streams from different speakers and different venues are precisely aligned on the timeline according to the actual speaking time sequence and overlap, so that the speech segments of each audio stream are completely matched with the real speaking scenario in the time dimension.

[0069] It is understood that the multi-reference language audio stream comprises multiple target audio streams from multiple speakers, which are translated and pre-transmitted (directed) to multiple meeting rooms. For example, there are a first meeting room, a second meeting room, and a third meeting room. In terms of speaking sequence, there is a first target audio stream corresponding to the first speaker and a second target audio stream corresponding to the second speaker. Therefore, the multi-reference language audio stream includes a first reference language audio stream (translated from the first target audio stream and directed to the first meeting room), a second reference language audio stream (directed to the second meeting room), a third reference language audio stream (directed to the third meeting room), a fourth reference language audio stream (translated from the second target audio stream and directed to the first meeting room), a fifth reference language audio stream (directed to the second meeting room), and a sixth reference language audio stream (directed to the third meeting room). Of course, these six reference language audio streams include audio streams from and directed to the speaker's current meeting room that do not require language translation.

[0070] S302, at least two reference language audio streams that point to the same venue and match the language required by the same venue are selected from the aligned multi-reference language audio streams, and the at least two reference language audio streams correspond to different speakers.

[0071] Specifically, relying on the venue pointing identifier and language attribute identifier bound to each reference language audio stream during the generation and timing alignment process, a single venue is used as the filtering unit. First, the preset required language of the venue is matched. Then, all reference language audio streams pointing to the venue and matching the language are filtered from the aligned audio streams. Since the filtering object is multiple target audio streams in the parallel speaking scenario, the filtering result is at least two audio streams belonging to different speakers.

[0072] As can be seen, in this embodiment, by removing languages ​​and audio streams that are irrelevant to the current venue, it is ensured that each venue only obtains the audio content in the language it needs, while also preserving the relationship between the parallel audio streams of different speakers in that venue. This provides a suitable processing object for subsequent audio interleaving playback for a single venue, achieving precise dimensionality reduction from global multi-channel audio streams to single-venue dedicated multi-channel audio streams.

[0073] S303, the at least two reference language audio streams are divided into multiple audio streams according to a preset playback interleaving sequence, and an audio identification identifier and a start playback timestamp are added to each audio stream. The preset playback interleaving sequence includes the alternating playback order of the at least two reference language audio streams and the preset playback duration corresponding to each audio stream during alternating playback. The preset playback interleaving sequence is used to determine the start playback timestamp.

[0074] Specifically, based on the alternating playback order of the at least two reference language audio streams and the preset playback duration of each alternation, the audio streams are divided into multiple audio segments. Simultaneously, a unique audio identification identifier is added to each audio segment to accurately associate it with information such as the speaker's identity and the original audio stream, preventing confusion in segmentation. Furthermore, based on a globally unified timeline and a preset playback interleaving sequence, a start playback timestamp is calculated and added to each audio segment. This timestamp is determined by the alternating playback order and the preset playback duration, ensuring that the terminal playback device plays each audio segment in an orderly manner according to the timestamp. This not only restores the true interactive rhythm of parallel speaking but also solves the problem of sound mixing when parallel audio streams are played directly, guaranteeing the continuity and orderliness of playback.

[0075] For example, suppose the Chinese audio stream in the same venue contains two parallel audio streams from speakers A and B. The preset playback interleaving sequence is that speakers A and B play alternately, with each playback lasting 5 seconds. The starting reference based on the globally unified timeline is 10:00:00. According to this sequence, the audio streams of speakers A and B are divided into multiple audio streams of 5 seconds each. An identification identifier ID001-A is added to each audio stream of speaker A, and an identification identifier ID002-B is added to each audio stream of speaker B. A start playback timestamp is matched to each audio stream, where the timestamp of the first audio stream of speaker A is 10:00:00, the timestamp of the first audio stream of speaker B is 10:00:05, the timestamp of the second audio stream of speaker A is 10:00:10, the timestamp of the second audio stream of speaker B is 10:00:15, and so on. The terminal playback device plays each audio stream sequentially according to the timestamp, looping until one stops, thus achieving orderly alternation of audio streams A and B. Participants can clearly distinguish the content of their speeches without any sound mixing.

[0076] In one possible embodiment, the plurality of meeting rooms includes a first meeting room and a second meeting room. The first meeting room requires a first language, and the second meeting room requires a second language. The plurality of target audio streams include a first audio stream from a first speaker located in the first meeting room and a second audio stream from a second speaker located in the second meeting room. The first audio stream has an earlier start time. The multi-reference language audio streams include the first audio stream, the second audio stream, a third audio stream translated from the first audio stream and matching the second language, and a fourth audio stream translated from the second audio stream and matching the first language. The timing scheduling operation on the multi-reference language audio streams includes: aligning the multi-reference language audio streams according to the speaking time sequence; and filtering from the aligned multi-reference language audio streams those pointing to the first meeting room and matching the first language. The first audio stream and the fourth audio stream, and the second audio stream and the third audio stream that point to the second venue and match the second language are filtered out; the first audio stream and the fourth audio stream are divided into multiple audio streams according to the preset playback interleaving sequence, and the second audio stream and the third audio stream are divided into multiple audio streams, and an audio identification identifier and a start playback timestamp are added to each audio stream. The preset playback interleaving sequence includes a first alternating playback order corresponding to the first audio stream and the fourth audio stream, a second alternating playback order corresponding to the second audio stream and the third audio stream, and a preset playback duration corresponding to the first audio stream, the second audio stream, the third audio stream and the fourth audio stream respectively during alternating playback. The preset playback interleaving sequence is used to determine the start playback timestamp.

[0077] Specifically, please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating a scenario for performing timing scheduling operations on a language audio stream, as provided in an embodiment of this application. Figure 4As shown, the entire process of multilingual audio stream timing scheduling and distribution in a cross-venue parallel speaking scenario is presented. The first speaker in the first venue (Chinese) and the second speaker in the second venue (English) are speaking in parallel. The conference box 130 in the first venue collects the first speaker's first audio stream and uploads it to the server 220, and the conference box 130 in the second venue collects the second speaker's second audio stream and uploads it to the server 220. After server 220 completes the translation processing of the audio stream, it generates a third audio stream translated from the first audio stream into English, and a fourth audio stream translated from the second audio stream into Chinese. Then, it performs timing scheduling, dividing the first and fourth audio streams (which point to the first venue and match Chinese) and the second and third audio streams (which point to the second venue and match English) into multiple segments according to a preset playback interleaving sequence and adding identification marks and timestamps to obtain the time-interleaved first and fourth audio streams, and the time-interleaved second and third audio streams. Finally, the time-interleaved audio streams are pushed to the conference box 130 in the corresponding venue. The conference box 130 controls the display screen 110 to display text in sections, and the playback device 120 (simultaneous interpretation headphones worn by the speaker) to play the interleaved audio stream in the corresponding language, realizing the synchronous transmission and orderly presentation of multilingual parallel speeches across venues.

[0078] As can be seen, in this embodiment, the timestamps of each audio stream are unified based on a global clock. Then, parallel audio streams matching the language requirements of each meeting room are precisely selected. Finally, the audio streams are divided into segments that can be played alternately according to preset timing rules, and identification marks and timestamps are added to ensure that the terminal can play them in sequence and clearly distinguish the content of different speakers. This not only solves the sound mixing problem caused by direct playback of parallel audio, but also ensures that participants in different meeting rooms can obtain a synchronized audio-visual experience adapted to their own language, realizing efficient collaboration and content traceability for parallel speaking in multiple scenarios.

[0079] It is important to note that the speech processing method for parallel speaking in multiple venues provided in this application is not only applicable to the parallel speaking time sequence of speakers in different venues, but also to the parallel speaking time sequence of speakers in the same venue. Whether it is a remote collaboration scenario across venues or a local discussion scenario in which multiple people interrupt each other in the same venue, the core capabilities such as global timeline alignment, multilingual audio translation and interleaving playback, and partitioned text display can be used to achieve accurate processing and orderly presentation of parallel speaking audio. This effectively solves the problems of mixed sound, misaligned text, and language adaptation in parallel speaking in multiple languages ​​and scenarios, providing a unified speech processing solution for various meeting scenarios and ensuring that participants have a clear and synchronized audiovisual experience.

[0080] Please refer to details. Figure 5 , Figure 5 This is a flowchart illustrating a process for performing temporal alignment operations on language text, as provided in an embodiment of this application. Figure 5 As shown, the step of performing a time-series alignment operation on the multi-reference language texts to obtain multi-target language texts includes the following steps: S501, the multi-reference language texts are arranged according to the speaking sequence.

[0081] Among them, the single-channel reference language text is the speech text corresponding to a single audio stream that matches the language required by a single meeting venue, generated after the initial audio stream of the single speaker has undergone the translation operation.

[0082] S502, determine multiple speech texts that correspond one-to-one with the multiple audio streams from the at least multiple reference language texts.

[0083] S503, perform a timing alignment operation on the multiple audio streams based on the multiple start playback timestamps corresponding to the multiple audio streams to obtain the multiple target language texts.

[0084] The timing alignment operation includes adjusting the arrangement order of the multiple audio text segments according to the timing of the multiple start playback timestamps, and / or adding multiple start display timestamps to the multiple audio text segments according to the multiple start playback timestamps. The multiple start playback timestamps and the multiple start display timestamps are synchronized in timing. The start display timestamps are used to set the start time of display of the multiple audio text segments on the display screen.

[0085] It is understood that the timing alignment operation is used to perform timing alignment operation on the multi-segment speech text that has been segmented and bound, with the precise start playback timestamp of the multi-segment audio stream as the only timing reference, so that the display timing of the text is completely synchronized with the playback timing of the audio. The final output multi-channel target language text has a timing attribute that corresponds one-to-one with the target language audio stream, and can be directly pushed to the meeting boxes in each venue to achieve synchronous presentation.

[0086] It should be noted that the embodiments of this application only provide several core implementation methods for timing alignment operations, including but not limited to adjusting the arrangement order of multiple audio text segments according to the start playback timestamp, adding a start display timestamp synchronized with the start playback timestamp for multiple audio text segments, and adding display duration markers, display level markers, synchronization blanking markers, timing offset calibration markers, etc., to multiple audio text segments based on the start playback timestamp of the audio stream. All implementation methods take audio-text synchronization as the core objective and can be flexibly combined and configured according to the actual display needs of the meeting.

[0087] As can be seen, in this embodiment, by arranging the reference language text according to the speaking sequence, binding the segmented audio streams with the corresponding text fragments, and performing text timing alignment operations based on the audio start playback timestamp, the timing of multilingual text and corresponding audio streams is accurately matched, ensuring the timing synchronization of text display and audio playback. This effectively solves the problem of timing misalignment and asynchronous display of multilingual text and audio in parallel speaking scenarios, and improves the audiovisual collaboration experience of cross-venue multilingual conferences.

[0088] Please see Figure 6 , Figure 6 This is a schematic diagram of a conference terminal display interface provided in an embodiment of this application, such as... Figure 6 As shown, this demonstrates the partitioned display effect of the conference terminal in a parallel speaking scenario. The interface is divided into left and right columns, corresponding to the parallel speakers in the first and second conference rooms respectively. Each partition clearly displays the speaker's identity information, the conference room they belong to, and the real-time text content of their speech. This layout not only meets the need for simultaneous visualization of multiple speakers' content in a parallel speaking scenario, but also allows participants to quickly distinguish the speaking affiliation of different conference rooms, avoiding information confusion caused by multiple texts being mixed together.

[0089] Building upon the zoned display, text enhancement effects have been incorporated. For example, the keyword "oppose" in the text of the second session's presentation is highlighted, allowing participants to quickly grasp the speaker's core viewpoint and interaction intent. Furthermore, the text display in each zone is perfectly synchronized with the playback sequence of the corresponding audio stream.

[0090] As can be seen, in this embodiment, by using a left-right column layout combined with text enhancement and audio-text synchronization technology, the clear presentation and efficient transmission of content from multiple venues and speakers in parallel speaking scenarios are achieved. By highlighting key information, participants can quickly grasp the core viewpoints. Furthermore, the text display, which is precisely synchronized with the audio timing, significantly improves the information reception efficiency and meeting interaction experience in parallel speaking scenarios. This application embodiment can divide the electronic device into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0091] Please see Figure 7 , Figure 7This is a functional block diagram of a remote conferencing system provided in an embodiment of this application. The remote conferencing system includes: an acquisition unit 701 and a processing unit 702; wherein, the acquisition unit 701 is used to acquire multiple audio streams from multiple meeting rooms through multiple meeting boxes; the processing unit 702 is used to perform the following processing on each of the multiple audio streams: extracting style features of the audio stream, the style features representing the speech prosody of the audio stream; and translating the audio stream according to the style features and preset meeting room language requirements information to generate the multiple meeting rooms. The system handles one-to-one requests for multiple language text and audio streams; detects that multiple target audio streams are speaking in parallel; determines the corresponding multiple reference language text and audio streams for each target audio stream; performs timing scheduling operations on the multiple reference language audio streams to obtain time-interleaved multiple target language audio streams, and performs timing alignment operations on the multiple reference language texts to obtain multiple target language texts; and pushes the multiple target language audio streams and texts to the multiple meeting boxes respectively.

[0092] In one possible embodiment, the audio stream is translated according to the style features and preset venue language requirement information to generate multi-language text and multi-language audio streams for the one-to-one requirements of the multiple venues. The processing unit 702 is specifically used to: convert the audio stream into first speech text; determine multiple target languages ​​for the one-to-one requirements of the multiple venues according to the preset venue language requirement information; call a pre-trained multilingual translation model to process the first speech text, and generate the multi-language text corresponding one-to-one with the multiple target languages; and based on a speech generation algorithm, fuse the style features with the multi-language text to generate the multi-language audio stream corresponding one-to-one with the multi-language text. The multi-language audio stream is a multiple audio stream that imitates the style features of the audio stream and is in the form of the multiple target languages. The style features include intonation, timbre, rhythm, and pause patterns.

[0093] In one possible embodiment, a timing scheduling operation is performed on the multi-reference language audio streams. The processing unit 702 is specifically used to: align the multi-reference language audio streams according to the speaking time sequence, wherein a single reference language audio stream is a single audio stream generated after the translation operation from the initial audio stream of a single speaker, and matches the language required by a single meeting venue; select at least two reference language audio streams from the aligned multi-reference language audio streams that point to the same meeting venue and match the language required by the same meeting venue, wherein the at least two reference language audio streams correspond to different speakers; divide the at least two reference language audio streams into multiple audio streams according to a preset playback interleaving time sequence, and add an audio identification identifier and a start playback timestamp to each audio stream segment, wherein the preset playback interleaving time sequence includes the alternating playback order of the at least two reference language audio streams and the preset playback duration corresponding to each audio stream during alternating playback, and the preset playback interleaving time sequence is used to determine the start playback timestamp.

[0094] In one possible embodiment, multiple meeting rooms include a first meeting room and a second meeting room. The first meeting room requires a first language, and the second meeting room requires a second language. The multiple target audio streams include a first audio stream from a first speaker located in the first meeting room and a second audio stream from a second speaker located in the second meeting room. The first audio stream has an earlier start time. The multi-reference language audio streams include the first audio stream, the second audio stream, a third audio stream translated from the first audio stream and matching the second language, and a fourth audio stream translated from the second audio stream and matching the first language. A timing scheduling operation is performed on the multi-reference language audio streams. The processing unit 702 is specifically used to: align the multi-reference language audio streams according to the speaking time sequence; and filter out audio streams from the aligned multi-reference language audio streams that point to the first meeting room and match the first language. The system includes the first and fourth audio streams of the specified language, and the second and third audio streams that are filtered to point to the second venue and match the second language. The first and fourth audio streams are divided into multiple audio streams according to a preset playback interleaving sequence, and the second and third audio streams are also divided into multiple audio streams. Each audio stream is assigned an audio identification identifier and a start playback timestamp. The preset playback interleaving sequence includes a first alternating playback order corresponding to the first and fourth audio streams, a second alternating playback order corresponding to the second and third audio streams, and preset playback durations for each of the first, second, third, and fourth audio streams during alternating playback. The preset playback interleaving sequence is used to determine the start playback timestamp.

[0095] In one possible embodiment, a timing alignment operation is performed on the multiple reference language texts to obtain multiple target language texts. The processing unit 702 is specifically used to: arrange the multiple reference language texts according to the speaking time sequence, where a single reference language text is the speech text corresponding to a single audio stream generated after the initial audio stream of a single speaker is translated, matching the language required by a single meeting venue; determine multiple speech texts that correspond one-to-one with the multiple audio streams from the at least multiple reference language texts; perform a timing alignment operation on the multiple speech texts according to multiple start playback timestamps corresponding to the multiple audio streams to obtain the multiple target language texts. The timing alignment operation includes adjusting the arrangement order of the multiple speech texts according to the timing of the multiple start playback timestamps, and / or adding multiple start display timestamps to the multiple speech texts according to the multiple start playback timestamps. The multiple start playback timestamps and the multiple start display timestamps are synchronized in timing, and the start display timestamps are used to set the display start time of the multiple speech texts on the display screen.

[0096] In one possible embodiment, multiple audio streams from the multiple meeting rooms are acquired through multiple conference boxes. The acquisition unit 701 is specifically used to: acquire multiple initial audio streams from the multiple meeting rooms through the multiple conference boxes, where each initial audio stream includes a mixed audio stream from multiple speakers; preprocess the mixed audio streams in the multiple initial audio streams according to a pre-recorded participant voiceprint database to obtain multiple audio streams that correspond one-to-one with the multiple speakers. The preprocessing includes speech separation processing and voiceprint identity binding, whereby a single audio stream is associated with a single speaker's identity identifier, meeting room information, voiceprint feature identifier, speaking start timestamp, and style features.

[0097] In one possible embodiment, a single conference terminal further includes a display screen and a playback device, with the conference box connected to the display screen and the playback device respectively; the multi-channel target language audio streams and the multi-channel target language text streams are respectively pushed to the multiple conference boxes, and the processing unit 702 is specifically used for: determining the main speaker who speaks earlier and other speakers who speak later based on the speaking time sequence of the speakers; determining the first voice text content corresponding to the main speaker and other voice text content corresponding to the other speakers in the multi-channel target language text streams; determining target partition display information based on the number of speakers and the speaking time sequence, the target partition display information including the first voice text stream. The content and other audio-visual text content are displayed in partitioned areas on the display screen, including their locations, areas, and display rules. Multiple first control commands are generated, each corresponding to one of the multiple meeting rooms. Each first control command carries multiple target language audio streams and multiple target language text streams that are directed to a single meeting room and match the language requirements of that single meeting room, as well as the target partition display information. The first control command instructs the playback device in the meeting room to play the multiple target language audio streams and instructs the display screen to synchronously display the multiple target language text streams according to the target partition display information. The multiple first control commands are then sent to the multiple meeting boxes corresponding to the multiple meeting rooms.

[0098] In one possible embodiment, the processing unit 702 is further configured to: detect that at least one of the other speakers has a gradually increasing volume and the absolute value of the volume difference between the speaker's current volume and the main speaker's volume is less than a first preset threshold, and / or detect that the at least one speaker has engaged in malicious volume adjustment, the malicious volume adjustment including lowering the main speaker's audio playback volume without obtaining meeting management permissions, the audio playback volume being the playback volume of the multi-channel target language audio stream corresponding to the main speaker transmitted from the meeting box to the playback device; and send a second control instruction to the meeting box, the second control instruction being used to instruct the playback device to play the main speaker and / or the multi-channel target language audio stream corresponding to the at least one speaker at a preset locked volume.

[0099] As can be seen, in this embodiment, by collecting audio streams from multiple venues, extracting speech style features, and combining them with the language requirements of each venue to complete multilingual translation, parallel speaking is accurately detected, and audio timing is scheduled and text timing is aligned. Finally, the appropriate audio and text are pushed to each venue and displayed in separate zones. This effectively solves problems such as mixed speech in parallel speaking, distorted translation style, and unclear association between identity and content. It significantly improves the clarity of information transmission, the naturalness of cross-language interaction, and the efficiency of meeting communication in multi-venue parallel speaking scenarios.

[0100] It is understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the method embodiment section in this application should be adapted to the device embodiment section in a synchronous manner, and will not be repeated here.

[0101] Figure 8 This is a structural block diagram of an electronic device provided in an embodiment of this application. For example... Figure 8 As shown, the electronic device 800 may include one or more of the following components: a processing module 801 and a memory 802 coupled to the processing module 801, wherein the memory 802 may store one or more computer programs, which may be configured to implement the methods described in the examples above when executed by one or more processing modules 801. The electronic device 800 may be as follows: Figure 1 The server shown is 220.

[0102] The processing module 801 may include one or more processing cores. The processing module 801 connects to various parts within the electronic device 800 using various interfaces and lines. It executes various functions and processes data of the electronic device 800 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 802, and by calling data stored in the memory 802. Optionally, the processing module 801 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processing module 801 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. It is understood that the aforementioned modem may also not be integrated into the processing module 801 and may be implemented separately through a communication chip.

[0103] The memory 802 may include random access memory (RAM) or read-only memory (ROM). The memory 802 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 802 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method examples described above. The data storage area may also store data created by the electronic device 800 during use.

[0104] It is understood that the electronic device 800 may include more or fewer structural elements than those shown in the above block diagram, such as a power module, physical buttons, WiFi (Wireless Fidelity) module, speaker, Bluetooth module, sensor, etc., without limitation.

[0105] This application also provides a computer storage medium storing a computer program / instructions thereon, which, when executed by a processor, implements some or all of the steps of any of the methods described in the above method embodiments.

[0106] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments.

[0107] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0108] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and there may be other division methods in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0110] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can be physically comprised separately, or two or more units can be integrated into one unit. The integrated unit described above can be implemented in hardware or in the form of hardware plus software functional units.

[0111] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute partial steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, volatile memory, or non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM), etc., which are various media capable of storing program code.

[0112] While the present invention has been disclosed above, it is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions without departing from the spirit and scope of the present invention, and various modifications and alterations can be made, including combinations of the different functions and implementation steps described above, as well as software and hardware implementation methods, all of which are within the protection scope of the present invention.

Claims

1. A speech processing method for parallel speaking in multiple venues, characterized in that, A server is used in a business management platform of a remote conferencing system. The system also includes multiple conferencing terminals in multiple meeting rooms. Each conferencing terminal includes a conferencing box, and the conferencing box is communicatively connected to the server. The method includes: Multiple audio streams from the multiple meeting rooms are acquired through multiple conference boxes; For each of the plurality of audio streams, the following processing is performed: Extracting style features from the audio stream, the style features characterizing the speech prosody of the audio stream; and, The audio stream is translated based on the style characteristics and the preset venue language requirements information to generate multi-language text and multi-language audio streams for the one-to-one requirements of the multiple venues. It was detected that multiple target audio streams were speaking in parallel sequence among the multiple audio streams. Determine the multi-reference language text and multi-reference language audio streams corresponding to the multiple target audio streams; A timing scheduling operation is performed on the multi-reference language audio streams to obtain a timing-interleaved multi-target language audio stream, and a timing alignment operation is performed on the multi-reference language text to obtain a multi-target language text; The multi-channel target language audio streams and multi-channel target language text streams are respectively pushed to the multiple conference boxes; wherein... The timing scheduling operation performed on the multi-reference language audio streams includes: The multi-reference language audio streams are aligned according to the speaking sequence. The single-reference language audio stream is a single audio stream generated by the translation operation of the initial audio stream of a single speaker, which matches the language required by a single meeting venue. From the aligned multi-reference language audio streams, at least two reference language audio streams that point to the same venue and match the language required by the same venue are selected, and the at least two reference language audio streams correspond to different speakers. The at least two reference language audio streams are divided into multiple audio streams according to a preset playback interleaving sequence, and an audio identification identifier and a start playback timestamp are added to each audio stream. The preset playback interleaving sequence includes the alternating playback order of the at least two reference language audio streams and the preset playback duration corresponding to each audio stream during alternating playback. The preset playback interleaving sequence is used to determine the start playback timestamp.

2. The method according to claim 1, characterized in that, The process of translating the audio stream based on the style characteristics and preset venue language requirements, and generating multi-language text and multi-language audio streams to meet the one-to-one requirements of multiple venues, includes: Convert the audio stream into first speech text; Based on the preset venue language requirement information, determine multiple target languages ​​for the one-to-one requirements of the multiple venues; The first speech text is processed by calling a pre-trained multilingual translation model to generate multilingual texts that correspond one-to-one with the multiple target languages. Based on the speech generation algorithm, the style features are fused with the multilingual text to generate a multilingual audio stream that corresponds one-to-one with the multilingual text. The multilingual audio stream is a multiple audio stream that imitates the style features of the audio stream and is in the form of the multiple target languages. The style features include intonation, timbre, rhythm, and pause patterns.

3. The method according to claim 1, characterized in that, The multiple meeting rooms include a first meeting room and a second meeting room. The first meeting room requires a first language, and the second meeting room requires a second language. The multiple target audio streams include a first audio stream from a first speaker located in the first meeting room and a second audio stream from a second speaker located in the second meeting room. The first audio stream starts speaking earlier. The multiple reference language audio streams include the first audio stream, the second audio stream, a third audio stream translated from the first audio stream and matching the second language, and a fourth audio stream translated from the second audio stream and matching the first language. The timing scheduling operation performed on the multi-reference language audio streams includes: Align the multi-reference language audio streams according to the speaking sequence; The first audio stream and the fourth audio stream that point to the first venue and match the first language are selected from the aligned multi-reference language audio streams; the second audio stream and the third audio stream that point to the second venue and match the second language are selected. The first audio stream and the fourth audio stream are divided into multiple audio streams according to the preset playback interleaving sequence, and the second audio stream and the third audio stream are divided into multiple audio streams. An audio identification identifier and a start playback timestamp are added to each audio stream. The preset playback interleaving sequence includes a first alternating playback order corresponding to the first audio stream and the fourth audio stream, a second alternating playback order corresponding to the second audio stream and the third audio stream, and a preset playback duration corresponding to the first audio stream, the second audio stream, the third audio stream and the fourth audio stream during alternating playback. The preset playback interleaving sequence is used to determine the start playback timestamp.

4. The method according to claim 1, characterized in that, The step of performing a time-series alignment operation on the multi-reference language texts to obtain multi-target language texts includes: The multi-reference language texts are arranged according to the speaking sequence, and the single-reference language text is the speech text corresponding to a single audio stream that matches the language required by a single venue, generated after the initial audio stream of a single speaker is processed by the translation operation. Determine multiple speech text segments that correspond one-to-one with the multiple audio streams from the at least multiple reference language texts; The timing alignment operation is performed on the multiple audio streams according to the multiple start playback timestamps to obtain the multiple target language texts. The timing alignment operation includes adjusting the arrangement order of the multiple audio streams according to the timing of the multiple start playback timestamps, and / or adding multiple start display timestamps to the multiple audio streams according to the multiple start playback timestamps. The multiple start playback timestamps and the multiple start display timestamps are synchronized in timing. The start display timestamps are used to set the display start time of the multiple audio streams on the display screen.

5. The method according to any one of claims 1-4, characterized in that, The process of acquiring multiple audio streams from the multiple meeting rooms through multiple conference boxes includes: Multiple initial audio streams from multiple meeting rooms are acquired through the multiple conference boxes, and each initial audio stream includes a mixed audio stream from multiple speakers; Based on the pre-recorded participant voiceprint database, the mixed audio stream in the multiple initial audio streams is preprocessed to obtain multiple audio streams that correspond one-to-one with the multiple speakers. The preprocessing includes speech separation processing and voiceprint identity binding. Each audio stream is associated with a single speaker's identity identifier, venue information, voiceprint feature identifier, speaking start timestamp, and style features.

6. The method according to claim 5, characterized in that, A single conference terminal also includes a display screen and a playback device, and the conference box is connected to the display screen and the playback device respectively; the step of pushing the multi-channel target language audio streams and the multi-channel target language text streams to the multiple conference boxes respectively includes: Based on the speaking order of the speakers, the main speaker who spoke earlier and the other speakers who spoke later were identified; Determine the first speech text content corresponding to the main speaker and the other speech text content corresponding to the other speakers in the multi-path target language text; The target partition display information is determined based on the number of speakers and the speaking sequence. The target partition display information includes the region location, region area, and display rules for the first voice text content and the other voice text content to be displayed on the display screen. Generate multiple first control instructions corresponding one-to-one with the multiple meeting venues. Each first control instruction carries multiple target language audio streams and multiple target language text streams that are directed to a single meeting venue and match the language required by the single meeting venue, as well as the target partition display information. The first control instructions are used to instruct the playback device of the meeting venue to play the multiple target language audio streams and to instruct the display screen to synchronously display the multiple target language text streams according to the target partition display information. The plurality of first control commands are respectively sent to the plurality of meeting boxes in the respective plurality of meeting venues.

7. The method according to claim 6, characterized in that, The method further includes: The system detects that at least one of the other speakers has a gradually increasing volume and the absolute value of the volume difference between the speaker and the main speaker's current volume is less than a first preset threshold, and / or detects that at least one speaker has engaged in malicious volume adjustment operations, wherein the malicious volume adjustment operations include lowering the main speaker's audio playback volume without obtaining meeting management permissions, wherein the audio playback volume is the playback volume of the multi-channel target language audio stream corresponding to the main speaker transmitted by the meeting box to the playback device; Send a second control command to the conference box, the second control command being used to instruct the playback device to play the multi-target language audio streams corresponding to the main speaker and / or at least one speaker at a preset locked volume.

8. A remote conferencing system, characterized in that, The remote conferencing system includes a server of a business management platform and multiple conferencing terminals in multiple venues. Each conferencing terminal includes a display screen, a playback device, and a conferencing box. The conferencing box is communicatively connected to the server. The server is used to perform the steps in the method as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, The computer program / instructions are executed by the processor to implement the steps of the method according to any one of claims 1-7.