Device and method for generating live streaming in different languages, and computer readable medium

By detecting, converting, and translating the broadcaster's audio stream, the problem of inconsistent translations in multilingual live streaming was solved, achieving consistency in timelines and semantic accuracy across different language versions, thus improving the efficiency and quality of live translation.

CN122053871APending Publication Date: 2026-05-15SQ TECH (SHANGHAI) CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610255680.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies suffer from translation inconsistencies and mistranslations of proper nouns or terms in multilingual live streaming, especially when the broadcaster switches languages ​​or mixes in foreign languages, leading to speech recognition errors and inconsistent translations.

Method used

By detecting the language of the broadcaster's audio stream, converting it into text information in the broadcaster's language, filtering out sensitive words, obtaining core information and translating it into the target language, and combining the streaming time information to generate the target audio stream, multilingual live content is synthesized to ensure consistency.

Benefits of technology

It achieves consistency in timeline and semantic accuracy across different language versions, improves the efficiency and consistency of multilingual translation, and ensures that global audiences receive synchronized and language-compliant live streaming services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053871A_ABST
    Figure CN122053871A_ABST
Patent Text Reader

Abstract

The invention relates to a device and a method for converting anchor voice to generate live streaming in different languages, and a computer readable medium. The method comprises the following steps of: detecting an anchor language type in an anchor voice streaming and converting the anchor voice streaming into anchor text information belonging to the anchor language type; the method comprises the following steps: translating anchor text information into target text information corresponding to a target language, converting the target text information into a target voice streaming according to streaming time information of the target text information, and synthesizing a live video streaming and a target voice streaming. According to the invention, the multilingual live streaming with synchronously corresponding translation content and correct translation can be generated, and the technical effect of improving the translation semantic consistency and generation efficiency of the multilingual streaming is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of live streaming generation, and in particular to an apparatus, method, and computer-readable medium for converting a broadcaster's voice to generate live streaming streams in different languages. Background Technology

[0002] With the widespread adoption of video streaming platforms, live video content is now widely used in various scenarios, including sports events, esports broadcasts, news broadcasts, online education, and entertainment programs. To expand the audience base, existing technologies have proposed solutions for real-time translation of live video content into different languages ​​and its distribution to multiple streaming platforms.

[0003] The common approach in existing technologies is to first perform automatic speech recognition (ASR) on the speaker's or broadcaster's voice, then convert the recognized text into the target language text using machine translation (MT), and finally use text-to-speech (TTS) to generate the corresponding language audio content. However, such technologies mostly assume that the speaker's language is known or fixed. When the speaker switches languages, uses foreign languages, or uses proper nouns during a live broadcast, it can easily lead to speech recognition errors or inconsistent translations.

[0004] In addition, existing technologies, when performing multilingual translation, usually only convert the language of the entire text and lack a mechanism for cross-language consistency control of key information (such as names of people, team names, scores, proper nouns or numerical information). This may cause discrepancies or conflicts in the core information of live broadcast content in different language versions, affecting the audience's understanding.

[0005] In summary, it is evident that existing technologies have long suffered from issues such as asymmetric translations and mistranslations of proper nouns or terminology during multilingual live streaming. Therefore, it is necessary to propose improved technical methods to address this problem. Summary of the Invention

[0006] In view of the problems of translation asymmetry and mistranslation of proper nouns or terminology that often occur in multilingual live streaming in existing technologies, this invention discloses an apparatus, method, and computer-readable medium for converting the broadcaster's speech to generate live streams in different languages, wherein: The apparatus disclosed in this invention for converting broadcaster voice to generate live streams in different languages ​​includes at least: a data transmission module for receiving live audio / video streams and receiving broadcaster voice streams synchronized with the live audio / video streams; a language detection module for detecting the type of broadcaster language in the broadcaster voice stream; a voice conversion module for recognizing the broadcaster voice stream to generate broadcaster text information belonging to the broadcaster language type, wherein the broadcaster text information includes streaming time information corresponding to the broadcaster voice stream; a word filtering module for filtering sensitive words in the broadcaster text information; and an information translation module for... The system extracts source core information from the anchor's text information, and obtains target core information corresponding to the target language from the source core information. It also translates the anchor's text information into target text information in the corresponding target language. The target text information includes target core information and streaming time information. The text conversion module is used to convert the target text information into a target speech stream corresponding to the target language based on the streaming time information. The streaming synthesis module is used to synthesize the live audio and video stream and the target speech stream to generate the target audio and video stream, and the data transmission module pushes the target audio and video stream to the target streaming platform.

[0007] The method for converting broadcaster speech to generate live streams in different languages ​​disclosed in this invention includes at least the following steps: receiving a live audio / video stream and receiving a broadcaster speech stream synchronized with the live audio / video stream; detecting the broadcaster language type in the broadcaster speech stream; identifying the broadcaster speech stream to generate broadcaster text information belonging to the broadcaster language type, the broadcaster text information including streaming time information corresponding to the broadcaster speech stream; filtering sensitive words in the broadcaster text information; obtaining source core information from the broadcaster text information and obtaining target core information corresponding to the target language from the source core information; translating the broadcaster text information into target text information in the corresponding target language, the target text information including target core information and streaming time information; converting the target text information into a target speech stream corresponding to the target language based on the streaming time information; synthesizing the live audio / video stream and the target speech stream to generate a target audio / video stream, and pushing the target audio / video stream to the target streaming platform.

[0008] The computer-readable medium disclosed in this invention stores a computer program thereon, which, when executed by a device, enables the device to implement the above-described method for converting broadcaster speech to generate live streams in different languages.

[0009] The apparatus, method, and computer-readable medium disclosed in this invention are as described above. The difference between this invention and the prior art lies in that this invention detects the type of the broadcaster's language in the broadcaster's speech stream and converts the broadcaster's speech stream into broadcaster text information belonging to the broadcaster's language type. Then, it translates the broadcaster's text information into target text information in the corresponding target language. Based on the streaming time information of the target text information, it converts the target text information into a target speech stream and then synthesizes the live video stream and the target speech stream. This solves the problems existing in the prior art and can achieve the technical effect of improving the semantic consistency and generation efficiency of multilingual translation streams. Attached Figure Description

[0010] Figure 1A This is a flowchart of the method for converting broadcaster's voice to generate live streams in different languages, as proposed in this invention.

[0011] Figure 1B This is a flowchart of the method for obtaining target core information from the anchor's text information proposed in this invention.

[0012] Figure 1C This is a flowchart of the method for editing live video streams to generate a target video stream, as proposed in this invention.

[0013] Figure 1D This is a flowchart of the method for adding an advertising video stream to a target video stream according to the present invention.

[0014] Figure 1E This is a flowchart of the method for generating chapter index tags for a target audio-visual stream according to the present invention.

[0015] Figure 2 This is a schematic diagram of data association proposed in an embodiment of the present invention.

[0016] Figure 3 This is a schematic diagram of the module of the device proposed in this invention for converting the broadcaster's voice to generate live streams in different languages.

[0017] Figure 4 This is a schematic diagram of the computer system of the electronic device proposed in this invention.

[0018] Explanation of reference numerals in the attached figures: Step 110: Receive the live video and audio stream and the broadcaster's audio stream synchronized with the live video and audio stream. Step 120: Detect the type of language spoken by the broadcaster in the broadcaster's audio stream. Step 125: Recognize the broadcaster's audio stream to generate broadcaster text information belonging to the broadcaster's language. Step 130: Filter sensitive words in the broadcaster's text messages Step 150: Obtain the source core information from the anchor's text information, and then obtain the target core information in the target language corresponding to the source core information. Step 151: Identify the core source information contained in the information mapping table from the anchor's text information. Step 153: Read the target core information corresponding to the source core information from the information mapping table. Step 155: Translate the anchor's text information into target text information in the corresponding target language. Step 160: Convert the target text information into a target speech stream corresponding to the target language based on the streaming time information. Step 170: Synthesize the live video stream and the target audio stream to generate the target video stream. Step 171: Edit the live video stream according to the publishing strategy corresponding to the streaming platform to generate heterogeneous video streams. Step 173: Obtain the heterogeneous audio stream that is time-synchronized with the heterogeneous audio-visual stream from the target audio stream. Step 175: Synthesize heterogeneous audio / video streams and heterogeneous audio streams to generate the target audio / video stream. Step 181: Obtain the advertising audio-visual stream corresponding to the target language and / or target streaming platform. Step 185: Synthesize the live video stream, the target audio stream, and the advertising video stream to generate the target video stream. Step 190: Push the target audio / video stream to the target streaming platform Step 195: Generate chapter index tags for the corresponding target audio / video stream based on the streaming time information. 210: Live video streaming 220: Anchor's audio streaming 230: Target speech stream 300: Device 303: Communication Interface 304: Storage Media 307: Processor 310: Data transmission module 320: Language Detection Module 330: Voice conversion module 340: Word Filtering Module 350: Information Translation Module 360: Text Conversion Module 370: Stream combining module 380: Video Editing Module 390: Index Building Module 400: Computer Systems 401: CPU 402: ROM 403: RAM 404: Bus 405: I / O Interface 406: Input section 407: Output Section 408: Storage Section 409: Communications Section 410: Driver 411: Removable media Detailed Implementation The features and implementation methods of the present invention will be described in detail below with reference to the accompanying drawings and embodiments. The content is sufficient to enable anyone skilled in the art to easily and fully understand the technical means used by the present invention to solve the technical problem and to implement it accordingly, thereby achieving the effects that the present invention can achieve.

[0019] It should be noted that the drawings are incorporated into and constitute a part of this specification, illustrating embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention. Obviously, the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings without any creative effort.

[0020] The diagrams provided in the following embodiments are only schematic representations of the basic concept of the present invention. The diagrams only show the elements related to the present invention and are not drawn according to the actual number, shape and size of the elements in the actual implementation. In the actual implementation, the form, quantity and proportion of each element can be arbitrarily changed, and the layout of the elements may also be more complex.

[0021] In this invention, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0022] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, known structures and devices are shown schematically rather than in detail to avoid obscuring embodiments of the invention.

[0023] This invention can convert the broadcaster's speech to a live video stream into different languages, and combine the broadcaster's speech with the live video stream to generate a target video stream corresponding to different languages.

[0024] The following is a preliminary step. Figure 1A The flowchart of the method for converting broadcaster speech to generate live streams in different languages ​​proposed in this invention illustrates the system operation process of this invention. Please refer to the following: Figure 2 A schematic diagram of data association proposed in this invention. The steps of the method for converting broadcaster speech to generate live streams in different languages ​​proposed in this invention are as follows: Step 110: Receive the live video stream 210 and the synchronized broadcaster voice stream 220. In this invention, the broadcaster voice stream 220 can be a series of voice signals emitted by a real broadcaster in response to the live video stream 210, such as the voice emitted by the broadcaster when performing broadcasting, explanation, or supplementary explanations of the live video stream, but this invention is not limited thereto.

[0025] In some embodiments, step 110 may further include dynamically correcting the time offset between the broadcaster's voice stream and the live video stream based on the buffering state and frame rate of the audio signal in the live video stream, or the buffering state and frames per second (FPS) of the audio signal in the live video stream. For example, if the frame rate of the video stream is 60 FPS, meaning each frame is displayed for approximately 16.67 milliseconds, and network latency is detected causing one second of unprocessed data to accumulate in the audio signal buffer of the video stream, the time axis of the audio signal can be compensated. For example, the playback time of the broadcaster's voice stream can be delayed by one second to ensure that the goal mentioned by the broadcaster accurately corresponds to the time when the goal occurs in the video frame of the live video stream, avoiding a situation where the player in the video frame of the live video stream has just received the ball when the broadcaster mentions the goal.

[0026] Step 120: Detect the type of broadcaster language in the broadcaster voice stream 220 obtained in step 110.

[0027] In step 120, language identification (LID) can be performed in real time based on changes in the physical characteristics of the broadcaster's audio stream 220 to detect the broadcaster's language. These physical characteristics include, but are not limited to, phoneme distribution, prosodic rhythm, and spectral characteristics. Therefore, step 120 can detect more than one broadcaster language simultaneously.

[0028] More specifically, in step 120, the language type of the broadcaster's speech can be determined by analyzing statistical characteristics such as the set of phonemes and phoneme combination rules appearing in the broadcaster's speech stream, analyzing differences in temporal rhythm and prosodic structure (such as stress distribution, syllable length patterns and pause patterns), and analyzing differences in spectral energy distribution and formant structure. For example, an acoustic feature vector can be extracted from the broadcaster's audio stream using Mel-Frequency Cepstral Coefficients (MFCCs) within a time window (e.g., 0.5 to 2 seconds). This extracted acoustic feature vector (MFCC) can then be converted into a phoneme sequence or a posterior probability distribution of phonemes using a Universal Phone Recognizer model. This allows for the establishment of a phoneme statistical model that describes the phonological structure of the broadcaster's speech. Furthermore, classifiers such as Probability Linear Discriminant Analysis (PLDA) or Deep Neural Networks (DNNs) can be used to determine the probability that the broadcaster's speech within that time window belongs to a specific language, thereby determining the language of each word in the broadcaster's audio stream. The aforementioned language-independent phoneme model includes, but is not limited to, combining Deep Neural Networks (DNNs) with Hidden Markov Models (HMMs). The present invention may include DNN-HMM acoustic models (HMM) or acoustic models using CTC (Connectionist temporal classification) as the loss function, as well as the aforementioned phoneme statistical models, such as unigram, bigram, or trigram models of phonemes, but is not limited thereto.Voice activity detection (VAD) and syllable or beat boundary detection can also be performed on the broadcaster's audio stream to obtain rhythmic features (hereinafter referred to as "rhythm features"), including but not limited to average syllable length, syllable length variation, fundamental frequency (F0) variation amplitude, fundamental frequency slope, energy contour variation, and pause duration distribution. Hidden Markov Models (HMM), Long Short-Term Memory (LSTM) networks, or one-dimensional Temporal Convolutional Neural Networks (TempCNN) can be used to capture the temporal dependencies of rhythmic features, thereby converting audio segments of a certain length into rhythmic feature sequences. Sequence classification models can also be used to determine the probability that the broadcaster's audio belongs to a specific language.Alternatively, the broadcaster's speech stream can be converted into spectral features, such as Mel-frequency cepstral coefficients (MFCC), Log-Mel spectrogram, filter bank features, or Linear Predictive Coding (LPC) parameters (but this invention is not limited thereto). Formant tracking can be performed on the converted spectral features to ensure that the generated spectral features contain statistical differences in vowel distribution, consonant plosive bands, and energy structure of the broadcaster's speech stream. A deep neural network (DNN) including a statistical pooling layer can be used to learn the generated spectral features to generate a fixed-dimensional language embedding vector. Then, a fully connected layer and a normalized exponential function (such as softmax) are used to determine the language category of the broadcaster's speech stream. The aforementioned deep neural network includes, but is not limited to, the x-vector model based on a time-delay neural network (TDNN) architecture, and advanced speaker recognition models based on time-delay neural networks (Emphasized Channel Attention, Propagation, and...). Aggregation in TDNN (ECAPA-TDNN), or Transformer models based on self-attention. In practice, the language type can usually be quickly and initially determined by the spectral characteristics of the broadcaster's audio stream. Then, weighted calculations based on phoneme distribution and speech rhythm can be performed to confirm or correct the determined language type. This allows for real-time detection of code-switching and switching of processing logic when the broadcaster suddenly intersperses English terminology in Chinese broadcasting.

[0029] Step 125: Identify the broadcaster's voice stream 220 obtained in step 110 to generate broadcaster text information belonging to the broadcaster's language type detected in step 120.

[0030] In step 125, a recognition dictionary corresponding to the broadcaster's language can be loaded, and relevant processing parameters can be dynamically configured. This allows the broadcaster's speech stream to be converted into text information using the loaded recognition dictionary through Automatic Speech Recognition (ASR) technology combined with Speech-to-Text (STT) technology (model, engine, or algorithm). The text information generated in step 125 includes the streaming time information corresponding to the broadcaster's speech stream.

[0031] In some embodiments, step 125 can dynamically switch between different speech-to-text technologies based on the language identification result of step 120, thereby maintaining the real-time conversion accuracy of cross-language broadcast voice streams. For example, multiple automatic speech recognition (ASR) models can be executed simultaneously, and dynamically fused according to the confidence score of each ASR model to generate a language identification result. Alternatively, step 125 can also utilize a large language model (LLM) to automatically correct speech segments whose language was incorrectly identified in step 120 based on the preceding semantics. For example, if the preceding text discusses American professional basketball (NBA), and step 120 incorrectly identifies a term or player name (such as Curry) in the broadcast voice stream as Chinese, step 125 can correct it to the corresponding English.

[0032] Step 130: Filter sensitive words in the anchor text information generated in step 125. In some embodiments, the sensitive word list may vary depending on the country or region where the target language is used, the location of the target streaming platform, or the location where the playback is provided. For example, different sensitive word lists may be established for different countries, regions, and target streaming platforms, or the applicable scope of each sensitive word may be set in the sensitive word list.

[0033] In step 130, a pre-established sensitive word list can be used to filter sensitive words in the broadcaster's text message. For example, sensitive words can be deleted or replaced with other words with similar meanings based on the content of the sensitive word list. For instance, if inappropriate discriminatory language appears in the broadcaster's text message, it can be replaced with symbols or neutral words in step 130 to ensure that the broadcaster's text message complies with the platform's security policy.

[0034] Step 150: Obtain the source core information from the anchor's text information after filtering sensitive words in Step 130, and then obtain the target core information of the target language. For example, the words in the pre-established information mapping table can be compared with each word in the anchor's text information to identify the source core information from the anchor's text information.

[0035] like Figure 1B As shown in the flowchart, step 150 may also include the following steps: Step 151: Identify the core source information contained in the pre-established information mapping table from the anchor's text information. For example, the words in the information mapping table (such as match player rosters, team codes, and terminology database) can be compared with each word in the anchor's text information to identify the core source information. For example, the team name Indiana Pacers can be identified as the core source information.

[0036] Step 153: Read out the target core information corresponding to the source core information from the information mapping table. For example, the source core information extracted from the host text information can be associated with a unique identifier (Unique Identifier, UID) according to the information mapping table, and the corresponding target core information can be read out from the information mapping table based on the unique identifier to obtain the target core information corresponding to the source core information. For example, by querying the information mapping table with the unique identifier, it is known that the target core information in traditional Chinese for the team name Indiana Pacers is 溜马, and the target core information in simplified Chinese is 步行者.

[0037] The following will return to Figure 1A .

[0038] Step 155: Translate the host text information after filtering sensitive words in Step 130 into the target text information in the corresponding target language. In the present invention, Step 155 can use machine translation (Machine Translation, MT) to convert the host text information into the target text information according to the target core information corresponding to the language source core information obtained in Step 150, so that the target text information contains the target core information. At the same time, in Step 155, the generated target text information can also contain streaming time information in addition to the target core information, so as to force the target text information to be synchronized and consistent in translations in multiple different target languages.

[0039] It should be noted that in Step 155, the numerical information (such as score or time) detected from the host text information can also be encapsulated as an immutable token and forced to be embedded into the grammatical structure of the target text information, so that the detected numerical information is not incorrectly changed by the translation model, ensuring the absolute correctness of the numerical information in multilingual versions.

[0040] Step 160: Convert the target text information into a target speech stream 250 corresponding to the target language according to the streaming time information in the target text information generated in Step 155. In the present invention, Step 160 can use text-to-speech (Text-to-Speech, TTS) technology to generate the target speech stream 250 according to the streaming time information in the target text information.

[0041] In step 160, the similarity between the target speech stream 250 and the voiceprint features of the broadcaster's speech can be compared. If the similarity between the target speech stream 250 and the voiceprint features of the broadcaster's speech is lower than a predetermined value, the target speech stream 250 can be regenerated to ensure that the generated target speech stream 250 matches the broadcaster's voice features. In addition, in step 160, when generating the target speech stream 250, emotional parameters such as speech rate and pitch can be adjusted to dynamically adjust the tone of the synthesized target speech based on semantic intensity. For example, if the broadcaster's text message contains "This is an incredible game-winning shot!" and the target text message is "This is an incredible game-winning shot!", it can be determined that "incredible" and "game-winning shot" are words with high semantic intensity. In this way, the speech rate and pitch center value of the text-to-speech engine can be increased when converting the target text message into the target speech, such as increasing the speech rate by 15% and shifting the pitch center value upward by 20Hz, to generate target speech with high emotion and fast rhythm.

[0042] Step 170: Combine the live video stream 210 obtained in step 110 with the target audio stream 250 generated in step 160 to generate the target video stream.

[0043] Step 190: Push the target audio / video stream generated in step 170 to the target streaming platform.

[0044] Combining steps 110 to 190 above, this invention can first receive a live video stream and a synchronized broadcaster's voice stream, and in real time detect the broadcaster's language in the voice stream. Then, it performs speech-to-text processing based on the broadcaster's language to generate broadcaster text information containing time information. Subsequently, it filters the broadcaster text information for sensitive words and extracts the source core information, further obtaining target core information corresponding to the target language, thereby assisting machine translation to generate target text information with consistent core information. Next, based on the time information in the target text information, it converts the target text information into a target voice stream, and combines it with the live video stream to generate a target video stream, which is then pushed to the target streaming platform, thus completing the real-time generation of a multilingual live video stream. Thus, through this invention, the broadcaster's voice can be converted into target audio streams in multiple different target languages ​​in real time without pre-defining the broadcaster's language, and the live audio and video streams in different target languages ​​can be kept consistent on the timeline, thereby improving the accuracy and consistency of the translation content of multilingual live streams, so that viewers of different languages ​​around the world can receive live services with highly consistent information content that conforms to the habits of each language at the same time.

[0045] In step 170 above, it is also possible to... Figure 1C As shown in the process diagram, different target audio and video streams are generated based on different target streaming platforms. The steps are as follows: Step 171: Edit the live video stream 210 obtained in step 110 according to the publishing strategy corresponding to the target streaming platform to generate a heterogeneous video stream. In step 171, the core events in the live video stream can be converted into language-independent data structures to serve as the basis for generating the heterogeneous video stream. The publishing strategy is usually predefined and includes the video length, editing rhythm, or subtitle presentation method corresponding to the target streaming platform. For example, the publishing strategy for TikTok is to edit short videos with a length of 15-60 seconds and an aspect ratio of 9:16, and the publishing strategy for Instagram Reels or Facebook Stories is to edit short videos with a length of less than 90 seconds and an aspect ratio of 9:16, etc., but the present invention is not limited to this.

[0046] In step 171, the core events, types of core events, event audio segments of core events, and event time data (such as start time and end time or start time and duration) of the live audio-visual stream obtained in step 110 can be determined based on the acoustic energy peak (such as the anchor screaming) in the anchor voice stream obtained in step 110, and / or the semantic structure, tone intensity changes, semantic weight change rate, semantic clustering degree (semantic density), sentence structure transitions, etc. in the anchor text information generated in step 155. Based on the determined core event type and core event event time data, heterogeneous audio-visual streams can be obtained from the live audio-visual stream. More specifically, step 171 can dynamically set the backtracking time based on the semantic inference results, and reduce the backtracking time and the end time or duration of the core event forward based on the start time of the core event to determine the time range in the live video stream that is related to the core event; step 171 can also combine the image features and sound features of the start time to verify whether the start time of the backtracking meets the conditions for the occurrence of the exciting event, wherein the image features are changes in motion and sudden changes in the composition of the picture (such as violent shaking of the picture), and the sound features are the peak volume energy (such as the cheers of the audience).

[0047] Step 173: Obtain a heterogeneous audio stream whose time is synchronized with the heterogeneous audio-visual stream generated in step 170 from the target audio stream 250 generated in step 160.

[0048] Step 175: Combine the heterogeneous video / audio stream generated in Step 171 with the heterogeneous audio stream generated in Step 173 based on their timestamps to generate the target video / audio stream. This ensures that the timestamps of the different heterogeneous video / audio streams are aligned with those of the heterogeneous audio streams, for example, generating a short summary version of a TikTok video, Instagram Reels video, or Facebook Stories video, or a full-length podcast video.

[0049] Furthermore, in the above description, the present invention can also be described as follows: Figure 1D As shown in the flowchart, the steps for inserting an advertisement into the target image stream are as follows: Step 181: Obtain the advertising video stream corresponding to the target language and / or the target streaming platform. Step 181 can select the corresponding advertising video stream based on the country or region where the target language is used, the location of the target streaming platform, or the region where the broadcast is provided. For example, for viewers in the United States, select a fast food advertisement authorized for that region to achieve geographically targeted advertising replacement.

[0050] Step 185: Combine the live video stream obtained in step 110, the target audio stream generated in step 160, and the advertising video stream obtained in step 181 to generate the target video stream.

[0051] Furthermore, in the above description, the present invention can also be described as follows: Figure 1E As shown in the flowchart, step 195 involves generating chapter index tags for the target audio-visual stream based on the streaming time information generated in step 170 or step 185. For example, it automatically marks time indices such as the start of the first half and the first goal in the live audio-visual stream on the corresponding YouTube platform for user convenience.

[0052] The following continues with Figure 3 The schematic diagram of the device for converting broadcaster speech to generate live streams in different languages, as proposed in this invention, illustrates the device for implementing this invention. (See attached diagram.) Figure 3 As shown, the device 300 of the present invention includes a communication interface 303, a storage medium 304, and a processor 307. The processor 307 is interconnected with the communication interface 303 and the storage medium 304 via a bus (not shown).

[0053] The communication interface 303 can be connected to network storage devices or servers (not shown in the figure) outside the device 300, and can request and download data (such as video signals, audio signals, metadata, etc.) from the connected network devices.

[0054] Storage medium 304 is typically a large-capacity storage area of ​​device 300. It can store data or signals downloaded by communication interface 303 (such as pre-buffered streaming frames), data or signals provided to processor 307 (such as sensitive word library, information mapping table), and data or signals generated by processor 307 (such as target audio and video stream).

[0055] The processor 307 may include modules such as a data transmission module 310, a language detection module 320, a speech conversion module 330, a word filtering module 340, an information translation module 350, a text conversion module 360, and a streaming synthesis module 370. It may also include an attachable audio-visual editing module 380 and an index building module 390. Specifically, the data transmission module 310 is connected to the language detection module 320, the speech conversion module 330, the streaming synthesis module 370, and the index building module 390; the language detection module 320 is connected to the speech conversion module 330; the information translation module 350 is connected to the word filtering module 340 and the text conversion module 360; and the streaming synthesis module 370 is connected to the audio-visual editing module 380 and the index building module 390. It should be noted that if the audio-visual editing module 380 exists, the text conversion module 360 ​​is connected to the audio-visual editing module 380; otherwise, it is connected to the streaming synthesis module 370.

[0056] In some embodiments, processor 307 can execute one or more sets of computer instructions stored in memory (not shown), and can generate the included modules after executing the computer instructions; in other embodiments, the modules included in processor 307 can be generated by one or more circuits and / or complete or partial chips and other hardware elements, that is, processor 307 includes hardware elements that make up the included modules. In other words, the modules included in processor 307 can be software modules or hardware modules, and the present invention has no particular limitations.

[0057] The data transmission module 310 is responsible for receiving live video and audio streams and the broadcaster's voice stream synchronized with the broadcaster's video and audio streams through the communication interface 303. The data transmission module 310 can also monitor transmission delays and perform initial timing synchronization processing on the video and audio streams.

[0058] The data transmission module 310 is also responsible for pushing the target audio and video stream generated by the stream synthesis module 370 to the target streaming platform through the communication interface 303. The data transmission module 310 can support a variety of communication protocols, such as Real-Time Messaging Protocol (RTMP) and HTTP Live Streaming (HLS).

[0059] The data transmission module 310 can also acquire advertising audio and video streams corresponding to the target language and / or streaming platform.

[0060] The data transmission module 310 can also encapsulate the chapter index tags generated by the index building module 390 in a transmission packet or Manifest file and push them synchronously to the target audio-visual stream to the target streaming platform.

[0061] The language detection module 320 is responsible for detecting the type of broadcaster language in the broadcaster's voice stream received by the data transmission module 310. The language detection module 320 can detect one or more types of broadcaster languages ​​contained in the broadcaster's voice stream in real time, thereby providing the speech conversion module 330 with a decision on whether code-switching is required.

[0062] The speech conversion module 330 is responsible for recognizing the broadcaster's speech stream received by the data transmission module 310 to generate broadcaster text information belonging to the broadcaster's language. The broadcaster text information generated by the speech conversion module 330 includes the streaming time information corresponding to the broadcaster's speech stream. The speech conversion module 330 can use a large language model for semantic correction.

[0063] The word filtering module 340 is responsible for filtering sensitive words in the text information generated by the speech conversion module 330. The list of sensitive words used by the word filtering module 340 for filtering sensitive words can be pre-stored in the storage medium 304, or stored in a storage device or network device such as a data server outside the device 300 (not shown in the figure), but the present invention is not limited thereto.

[0064] The information translation module 350 is responsible for obtaining the source core information from the anchor's text information after the word filtering module 340 has filtered out sensitive words, and obtaining the target core information in the target language to provide a consistency verification function, which can ensure that proper nouns and professional terms are correctly translated in all translations. The information translation module 350 can directly read the target core information from the information mapping table pre-stored in the storage medium 304; the information translation module 350 can also obtain the information mapping table stored in the external storage device 300 through the data transmission module 310, and read the target core information from the obtained information mapping table.

[0065] The information translation module 350 is also responsible for translating the anchor's text information after the word filtering module 340 filters out sensitive words into target text information in the corresponding target language. The target text information translated by the information translation module 350 includes the target core information and streaming time information.

[0066] The text conversion module 360 ​​is responsible for converting the target text information into a target speech stream corresponding to the target language, based on the streaming time information in the target text information generated by the information translation module 350. The text conversion module 360 ​​may have an emotion parameter adjustment function, which can dynamically adjust the tone of the synthesized speech according to the semantic intensity.

[0067] The streaming synthesis module 370 is responsible for synthesizing the live audio and video stream received by the data transmission module 310 and the target speech stream generated by the text conversion module 360 ​​to produce the target audio and video stream.

[0068] The streaming synthesis module 370 can also synthesize the live audio and video stream and advertising audio and video stream received by the data transmission module 310 with the target speech stream generated by the text conversion module 360 ​​to generate the target audio and video stream, thereby achieving the function of dynamic advertising insertion.

[0069] The streaming synthesis module 370 can also edit heterogeneous audio and video streams from the live audio and video stream through the audio and video editing module 380, and can obtain heterogeneous audio streams that are synchronized with the heterogeneous audio and video streams in time from the target speech stream generated by the text conversion module 360, and synthesize the heterogeneous audio and video streams and heterogeneous audio streams to generate the target audio and video stream.

[0070] The audio-visual editing module 380 can edit the live audio-visual stream received by the data transmission module 310 according to the publishing strategy corresponding to the target streaming platform to generate a heterogeneous audio-visual stream. More specifically, the audio-visual editing module 380 can determine the core event, the type of the core event, and the event time data of the core event based on the acoustic energy peak in the broadcaster's voice stream received by the data transmission module 310 and / or the semantic weight change rate or semantic aggregation degree of the broadcaster's text information generated by the speech conversion module 330. Based on the type of the core event, the event time data, and the publishing strategy corresponding to the target streaming platform, it can edit a heterogeneous audio-visual stream from the live audio-visual stream. In this way, the audio-visual editing module 380 can automatically identify highlights and perform conscious content backtracking editing.

[0071] The index building module 390 can generate chapter index tags for the target audio-visual stream based on the streaming time information in the target audio-visual stream generated by the streaming synthesis module 370.

[0072] It should be noted that the functions of the above modules can be implemented by a single module or multiple modules working together, and the boundaries between the modules are for illustrative purposes only and are not intended to limit the implementation of the present invention. In different embodiments, the functions of some modules can also be completed collaboratively by cloud servers, edge computing devices, or distributed systems without affecting the technical effects to be achieved by the present invention.

[0073] In summary, the difference between this invention and existing technologies lies in its ability to detect the type of the broadcaster's language in the broadcaster's audio stream, convert the broadcaster's audio stream into text information belonging to the broadcaster's language, translate the broadcaster's text information into target text information in the corresponding target language, convert the target text information into a target audio stream based on the streaming time information of the target text information, and then synthesize the live audio-visual stream and the target audio stream. This technology can solve the problems of asymmetric translation content and mistranslation of proper nouns or terms that often occur in multilingual live broadcasts in existing technologies, thereby achieving the technical effect of improving the semantic consistency and generation efficiency of multilingual translation streams.

[0074] Figure 4 This is a schematic diagram of an electronic device provided in one embodiment of the present invention. It should be noted that... Figure 4 The computer system 400 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0075] like Figure 4 As shown, the computer system 400 includes a Central Processing Unit (CPU) 401, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 402 or programs loaded from storage portion 408 into Random Access Memory (RAM) 403, such as performing the methods described in the above embodiments. The RAM 403 also stores various programs and data required for system operation. The CPU 401, ROM 402, and RAM 403 are interconnected via a bus 404. An Input / Output (I / O) interface 405 is also connected to the bus 404.

[0076] The following components are connected to the input / output interface 405: an input section 406 including a keyboard, mouse, etc.; an output section 407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface controller such as a local area network (LAN) card, modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output interface 405 as needed. Removable media 411, such as floppy disks, optical disks, magneto-optical disks, semiconductor memory, etc., are installed on the drive 410 as needed so that computer programs read from them can be installed into the storage section 408 as needed.

[0077] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing computer programs for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 409, and / or installed from removable medium 411. When the computer program is executed by central processing unit 401, it performs various functions defined in the system of the present invention.

[0078] It should be noted that the computer-readable medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or element, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a floppy disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disc read-only memory (CD-ROM), optical storage, magnetic random access memory (MRAM), or any suitable combination thereof. In the present invention, a computer-readable signal medium may include a data signal propagated in a baseband frequency or as part of a carrier wave, wherein a computer-readable computer program is carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or element. A computer program contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0079] The flowcharts and schematic diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or schematic diagram may represent a module, a program segment, or a portion of program code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than indicated in the figures. For example, two consecutively indicated blocks may actually be executed simultaneously, or sometimes in reverse order, depending on the functions involved. It should also be noted that each block in a schematic diagram or flowchart, and combinations of blocks in a schematic diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0080] The units described in the embodiments of the present invention can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0081] Another aspect of the present invention provides a computer-readable medium having a computer program stored thereon. When executed by a processor of a device (such as a computer) in a distributed network environment, the computer program enables the device to implement the method described above for converting broadcaster speech to generate live streams in different languages. This computer-readable medium may be included in the electronic device described in the above embodiments, or it may exist independently and not incorporated into the electronic device.

[0082] Another aspect of the present invention provides a computer program product or computer program including computer instructions stored in a computer-readable medium. A processor of a computer device reads the computer instructions from the computer-readable medium and executes the computer instructions, causing the computer device to perform the method provided in the above embodiments for converting broadcaster speech to generate live streams in different languages.

[0083] While the embodiments disclosed in this invention are as described above, the content is not intended to directly limit the scope of patent protection for this invention. Any modifications or refinements to the form and details of the implementation of this invention made by those skilled in the art, without departing from the spirit and scope disclosed herein, shall fall within the scope of patent protection for this invention. The scope of patent protection for this invention shall still be determined by the scope defined in the appended claims.

Claims

1. A method for converting a broadcaster's speech to generate live streams in different languages, applied to a device, the method comprising at least the following steps: Receive live video and audio streams, and receive the broadcaster's voice stream synchronized with the live video and audio streams; Detect the type of language spoken by the broadcaster in the broadcaster's audio stream; The broadcaster's voice stream is identified to generate broadcaster text information belonging to the broadcaster's language type. The broadcaster text information includes the streaming time information corresponding to the broadcaster's voice stream. Filter out sensitive words and phrases from the broadcaster's text messages; The source core information is obtained from the anchor's text information, and the source core information is then mapped to at least one target language core information. Translate the anchor's text information into target text information corresponding to at least one target language. The target text information includes the target core information and the streaming time information. Based on the streaming time information, the target text information is converted into a target speech stream corresponding to the at least one target language; and The live video stream and the target audio stream are combined to generate the target video stream, and the target video stream is pushed to the target streaming platform.

2. The method for converting broadcaster speech to generate live streams in different languages ​​as described in claim 1, wherein the step of synthesizing the live audio-visual stream and the target speech stream to generate the target audio-visual stream further includes the step of obtaining an advertising audio-visual stream corresponding to the at least one target language and / or the target streaming platform, and synthesizing the live audio-visual stream, the target speech stream, and the advertising audio-visual stream to generate the target audio-visual stream.

3. The method for converting broadcaster speech to generate live streams in different languages ​​as described in claim 1, wherein the step of detecting the broadcaster's language type in the broadcaster's speech stream is to detect the broadcaster's language type based on changes in phoneme distribution, speech rhythm, or spectral characteristics in the broadcaster's speech stream.

4. The method for converting broadcaster voice to generate live streams in different languages ​​as described in claim 1, wherein the step of obtaining the source core information from the broadcaster's text information and obtaining the target core information corresponding to the source core information in the at least one target language is to identify the source core information contained in a pre-established information mapping table from the broadcaster's text information and read the target core information corresponding to the source core information from the information mapping table.

5. The method for converting broadcaster speech to generate live streams in different languages ​​as described in claim 1, wherein the step of synthesizing the live audio-visual stream and the target audio stream to generate the target audio-visual stream further includes the step of editing the live audio-visual stream according to the publishing strategy corresponding to the target streaming platform to generate a heterogeneous audio-visual stream, and obtaining a heterogeneous audio stream whose time is synchronized with the heterogeneous audio-visual stream from the target audio stream, and synthesizing the heterogeneous audio-visual stream and the heterogeneous audio stream to generate the target audio-visual stream.

6. The method for converting broadcaster voice to generate live streams in different languages ​​as described in claim 5, wherein the step of editing the live audio-visual stream to generate the heterogeneous audio-visual stream according to the publishing strategy corresponding to the streaming platform is to determine the core event and the event time data of the core event based on the acoustic energy peak in the broadcaster voice stream and / or the semantic weight change rate or semantic aggregation degree in the broadcaster text information, and to obtain the heterogeneous audio-visual stream from the live audio-visual stream according to the type of the core event and the event time data.

7. The method for converting broadcaster speech to generate live streams in different languages ​​as described in claim 1, wherein after the step of synthesizing the live audio-visual stream and the target speech stream to generate the target audio-visual stream, the method further includes the step of generating chapter index tags corresponding to the target audio-visual stream based on the stream time information.

8. An apparatus for converting a broadcaster's voice to generate live streams in different languages, the apparatus comprising at least: The data transmission module is used to receive live video and audio streams, as well as to receive the broadcaster's voice stream synchronized with the live video and audio streams. The language detection module is used to detect the language spoken by the broadcaster in the broadcaster's audio stream; The speech conversion module is used to recognize the broadcaster's speech stream to generate broadcaster text information belonging to the broadcaster's language type. The broadcaster text information includes the streaming time information corresponding to the broadcaster's speech stream. The word filtering module is used to filter sensitive words in the broadcaster's text messages; The information translation module is used to obtain source core information from the anchor's text information, obtain target core information corresponding to the source core information in at least one target language, and translate the anchor's text information into target text information corresponding to the at least one target language, wherein the target text information includes the target core information and the streaming time information. The text conversion module is used to convert the target text information into a target speech stream corresponding to the at least one target language based on the streaming time information. and The streaming synthesis module is used to synthesize the live video stream and the target audio stream to generate the target video stream, and the data transmission module pushes the target video stream to the target streaming platform.

9. A computer-readable medium storing a computer program that, when executed by a device, causes the device to perform a method for converting broadcaster speech to generate live streams in different languages ​​as described in any one of claims 1 to 7.