Cross-language live broadcast interaction method and device, electronic equipment and storage medium
By using a streaming translation engine based on a large language model and semantic association analysis, the problems of speech translation latency and interaction disconnect in cross-language live streaming have been solved, enabling real-time and accurate cross-language interaction and improving the interactive efficiency and user experience of live streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-19
- Publication Date
- 2026-04-14
AI Technical Summary
In existing cross-language live streaming, real-time voice translation suffers from high latency, failing to meet the needs of real-time dialogue and instant interaction. Furthermore, machine translation lacks a relevant understanding of the real-time context of the live stream, resulting in disconnected interaction, a heavy workload for the broadcaster, and low interaction efficiency.
A streaming translation engine using a large language model is used to recognize and translate audio segments in real time. Combined with the text stream of questions in the bullet comments, semantic relevance analysis is performed to filter out questions that are highly relevant to the content being explained and to automatically generate response text. Through sensitive word detection and cultural adaptation processing, a complete interactive loop is formed.
It achieves extremely low latency voice translation output, accurately filters bullet comments related to the content being explained, reduces the burden on the broadcaster, improves the accuracy and efficiency of interaction, and ensures the timeliness and accuracy of responses.
Smart Images

Figure CN121865028A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interactive technology for cross-language live streaming, and more particularly to an interactive method, apparatus, electronic device, and storage medium for cross-language live streaming. Background Technology
[0002] In cross-border live streaming scenarios, existing technical solutions typically operate real-time voice translation and bullet screen interaction systems independently, resulting in significant defects in the overall interaction process. First, traditional voice translation tools suffer from excessively high processing latency, with end-to-end delays from voice input to translated output often exceeding several seconds, failing to meet the demands of real-time dialogue and instant interaction, severely impacting user experience. Second, mechanical translation results lack a connection to the real-time context of the live stream, making it impossible to effectively identify and respond to questions in the audience's bullet screen that are highly relevant to the content being explained by the streamer, leading to a disconnect in interaction. Furthermore, the streamer also needs to handle translated content and a massive number of bullet screen questions simultaneously, resulting in a heavy operational burden and low interaction efficiency. Summary of the Invention
[0003] Based on this, it is necessary to address the existing interactive issues in cross-language live streaming by proposing an interactive method, device, electronic equipment, and storage medium for cross-language live streaming.
[0004] An interactive method for cross-language live streaming, the method comprising: The source language speech stream of the first user is acquired in real time, and the source language speech stream is divided into multiple speech segment streams of preset duration; A streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, and outputs a target language text stream. The system receives multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; Based on the correlation value, target question texts with a correlation degree higher than a preset threshold are filtered out from the bullet screen question text stream; For each of the target bullet screen question texts, generate corresponding reply text.
[0005] Furthermore, the streaming translation engine based on a large language model performs real-time recognition and translation of each of the speech segment streams and outputs a target language text stream, including: The streaming translation engine based on a large language model identifies speech segments in the currently acquired speech segment stream to obtain the current source language text; wherein, the streaming translation engine is configured to output the corresponding target language text segment synchronously for each input speech segment stream; The current source language text is translated using the streaming translation engine to obtain a current target language text fragment; The previously output historical target language text fragments are corrected based on the current target language text fragment; The target language text stream is obtained by combining the corrected historical target language text fragments and the current target language text fragments.
[0006] Further, the step of performing semantic correlation analysis on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream includes: By using a comparative learning algorithm, semantic features of the target language text and the bullet screen question text are extracted respectively, and a first semantic feature vector and a second semantic feature vector are obtained respectively. Calculate the similarity between the first semantic feature vector and the second semantic feature vector, and use it as the association value.
[0007] Furthermore, the step of generating corresponding response text for each of the target bullet screen question texts includes: The intent of the target bullet screen question text is identified to determine its compliant query category; Based on the compliance query category, retrieve the corresponding structured compliance data from the pre-built cross-border commodity compliance database, and obtain the corresponding target response template from the pre-set response template; The structured compliance data is populated into the target response template to generate the response text.
[0008] Furthermore, after the step of generating corresponding response text for each of the target bullet screen question texts, the method further includes: Obtain the sensitive word database corresponding to the target language text stream; The response text is subjected to sensitive word detection using the aforementioned sensitive word database; When a sensitive word from the sensitive word library is detected in the response text, an equivalent word that conforms to the cultural habits of the target market is obtained and replaced.
[0009] Further, the step of filtering target question texts with a relevance higher than a preset threshold from the bullet screen question text stream based on the correlation value includes: Add bullet screen question texts with associated values greater than a preset threshold to the pending processing queue; The bullet screen question texts in the queue to be processed are sorted according to a preset priority order; Based on the sorting results, the bullet screen question texts are processed as the target bullet screen question texts in sequence.
[0010] Furthermore, prior to the step of receiving the real-time text stream of questions from multiple second users, the method further includes: Obtain the language of the original bullet screen text stream; Determine whether the language of the original bullet screen text stream is the same as the language of the target language text stream; If they are different, the original bullet screen text stream is translated into the same language as the target language text stream to obtain the bullet screen question text stream.
[0011] An interactive device for cross-language live streaming, the device comprising: The acquisition module is used to acquire the source language speech stream of the first user in real time and divide the source language speech stream into multiple speech segment streams of preset duration; The translation module is used in a streaming translation engine based on a large language model to perform real-time recognition and translation of each of the aforementioned speech segments and output the target language text stream. The receiving module is used to receive multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. The analysis module is used to perform semantic correlation analysis on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; The filtering module is used to filter out target question texts with a relevance higher than a preset threshold from the bullet screen question text stream based on the correlation value; The generation module is used to generate corresponding response text for each of the target bullet screen question texts.
[0012] An electronic device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps: The source language speech stream of the first user is acquired in real time, and the source language speech stream is divided into multiple speech segment streams of preset duration; A streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, and outputs a target language text stream. The system receives multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; Based on the correlation value, target question texts with a correlation degree higher than a preset threshold are filtered out from the bullet screen question text stream; For each of the target bullet screen question texts, generate corresponding reply text.
[0013] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the following steps: The source language speech stream of the first user is acquired in real time, and the source language speech stream is divided into multiple speech segment streams of preset duration; A streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, and outputs a target language text stream. The system receives multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; Based on the correlation value, target question texts with a correlation degree higher than a preset threshold are filtered out from the bullet screen question text stream; For each of the target bullet screen question texts, generate corresponding reply text.
[0014] The beneficial technical effects of this invention are as follows: By segmenting the source language speech stream in real time and performing incremental processing based on a streaming translation engine, extremely low latency output from speech to target language text stream is achieved, effectively ensuring the real-time performance of cross-language live streaming. Then, by performing online semantic association analysis between the bullet screen question text stream and the real-time generated target language text stream, questions highly relevant to the current explanation content can be accurately and automatically filtered from a massive amount of bullet screens, solving the problem of multimodal information fragmentation. Reply text is automatically generated for the filtered relevant questions, forming a complete interactive closed loop from real-time translation, intelligent question filtering to automatic response, greatly reducing the interactive burden on the anchor and significantly improving the accuracy and overall efficiency of live streaming interaction. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] in: Figure 1 This is a diagram illustrating the application environment of an interactive method for cross-language live streaming in one embodiment. Figure 2 A flowchart of an interactive method for cross-language live streaming in one embodiment; Figure 3 This is a structural block diagram of an interactive device for cross-language live streaming in one embodiment; Figure 4 This is a structural block diagram of an electronic device in one embodiment. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Figure 1 This is a diagram illustrating an interactive application environment for cross-language live streaming in one embodiment. (Refer to...) Figure 1 This cross-language live streaming interactive method is applied to a cross-language live streaming interactive system. The system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal; a mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers. The terminal 110 is used to acquire the source language audio stream, and the server 120 is used to generate response text.
[0019] like Figure 2 As shown, in one embodiment, an interactive method for cross-language live streaming is provided. This method can be applied to both terminals and servers; this embodiment illustrates its application to terminals. The interactive method for cross-language live streaming specifically includes the following steps: S1: Real-time acquisition of the source language speech stream of the first user, and segmenting the source language speech stream into multiple speech segment streams of preset duration; S2: A streaming translation engine based on a large language model, which performs real-time recognition and translation of each of the aforementioned speech segments and outputs a target language text stream; S3: Receive multiple bullet screen question text streams from second users in real time; wherein, the target language text stream is in the same language as the bullet screen question text stream; S4: Perform semantic correlation analysis on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; S5: Based on the correlation value, filter out target question texts with a correlation degree higher than a preset threshold from the bullet screen question text stream; S6: Generate corresponding response text for each of the target bullet screen question texts.
[0020] As described in step S1 above, the source language speech stream of the first user is acquired in real time and segmented into multiple speech segments of preset duration. The source language speech stream of the first user (usually a broadcaster) is acquired in real time via a smart device or microphone. Specifically, acoustic capture can be performed using a device with high-sensitivity speech recognition capabilities. The source language speech stream is then segmented, with the segmentation timing parameters preset, for example, segmenting into a speech segment every 150 milliseconds. This minimizes latency and ensures the system can respond quickly to the broadcaster's speech. Instead of waiting for the input of entire sentences or segments, a "listen-and-process" mode is achieved, enabling real-time translation. Furthermore, the segmented speech segments have a degree of independence, allowing the system to process each segment in parallel, further improving overall processing efficiency.
[0021] As described in step S2 above, the streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, outputting a target language text stream. The received speech segment stream is fed into the streaming translation engine based on a large language model, such as using a translation model with an architecture like Transformer-XL. This engine performs real-time recognition of each speech segment, first converting the audio signal into text form. Specifically, it employs an incremental encoder-decoder architecture based on Transformer, supporting segment-by-segment input and output. This process is accomplished through Automatic Speech Recognition (ASR) technology, aiming to convert the audio into text content in the source language. Subsequently, the streaming translation engine translates the text information based on a specific machine learning model, converting it into target language text. The unique feature of this model is its ability to process streaming data; that is, while parsing each segment, it continuously generates and outputs a target language text stream. This makes the translation process continuous, closely following the speaker's delivery, greatly improving real-time performance and interactivity. Furthermore, the streaming translation system can support dynamic updates, making corresponding corrections to the preceding text based on contextual information to ensure the accuracy and consistency of the translation results, thereby better meeting the needs of interaction.
[0022] As described in step S3 above, multiple bullet screen question text streams from second users are received in real time. The target language text stream and the bullet screen question text stream are in the same language. Second users (usually viewers) can send instant questions through the input box during the live broadcast. The system will receive these bullet screen texts in real time to ensure that viewer feedback can be quickly integrated into the interaction. This is very important for improving the interactivity of the live broadcast, because viewer questions are often based on the real-time content of the host. If they can be responded to in a timely manner, it will significantly improve the experience satisfaction. In order to ensure efficient interaction, the system must ensure that the received bullet screen question text stream is in the same language as the currently generated target language text stream. In other words, if the host uses language A for the explanation, then the received question text must also be in language A. Since live broadcasts are generally aimed at a specific region, the language used is usually uniform. If there are mixed languages, such as language A and language B in a certain area, language B can be translated into language A before being received. This mechanism helps maintain the consistency and fluency of the interaction and further simplifies the subsequent semantic analysis process. This is because directly comparing data in the same language can more clearly measure the relevance and avoid ambiguity caused by language differences.
[0023] As described in step S4 above, semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain a correlation value between each question text and the target language text stream. The analysis process can employ a machine learning model. Through contrastive learning and natural language processing techniques, the bullet screen question texts are converted into semantic feature vectors and compared with the feature vectors of corresponding segments in the target language text stream. By calculating their similarity (e.g., cosine similarity or Euclidean distance), a numerical value representing the correlation can be obtained. This value directly reflects the relevance between the bullet screen question and the currently being explained content.
[0024] As described in step S5 above, target question texts with a relevance higher than a preset threshold are filtered from the bullet screen question text stream based on the correlation value. The bullet screen question text stream is filtered based on the obtained correlation value. When the calculated correlation value is higher than a pre-set threshold, the question text is considered a target bullet screen question text and can proceed to subsequent processing. This threshold is usually set based on historical data, user behavior analysis, and system experience to ensure relevance and accuracy. Through this method, the system can perform efficient information filtering, avoiding processing redundant information or irrelevant questions, and focusing attention on questions highly relevant to the host's speech. This process not only improves the system's interactivity but also retains the most valuable information for subsequent response generation, making the communication between the host and the audience more precise, thereby improving the quality and efficiency of the live stream. The preset threshold can be set based on historical data analysis and user behavior habits, and typically represents the relevance of the question to the live stream content, for example, set to 0.7.
[0025] As described in step S6 above, corresponding response text is generated for each target bullet screen question text. Specifically, the system performs intent analysis on the target bullet screen question text to clarify its specific content and identify the category or question the user wants to know (such as product attributes, ingredients, price, etc.). Then, the system retrieves information related to the question from a dynamic knowledge base or a pre-built compliance database. The retrieved information is then formatted using a pre-set response template to generate response text that conforms to the target language. The entire process is automated, significantly reducing the need for manual intervention while ensuring the accuracy and timeliness of the response information. Finally, the response text is sent to the live stream interface for viewers to view or hear instantly, further enhancing the interactivity and engagement of the live stream and improving the user experience.
[0026] In one embodiment, step S2 of the streaming translation engine based on a large language model, which performs real-time recognition and translation of each speech segment stream and outputs a target language text stream, includes: S201: The streaming translation engine based on a large language model identifies the speech segments in the currently acquired speech segment stream to obtain the current source language text; wherein, the streaming translation engine is configured to output the corresponding target language text segment synchronously for each input speech segment stream; S202: Use the streaming translation engine to translate the current source language text to obtain the current target language text fragment; S203: Correct the previously output historical target language text segments based on the current target language text segment; S204: Combine the corrected historical target language text fragments and the current target language text fragments to obtain a target language text stream.
[0027] As described in step S201 above, the streaming translation engine based on a large language model identifies the speech segments in the currently acquired speech segment stream to obtain the current source language text. The streaming translation engine is configured to simultaneously output the corresponding target language text segment for each input speech segment stream. The streaming translation engine converts the input source language speech segment stream into text format, i.e., the current source language text. The streaming translation engine is built as a real-time processing structure, enabling the system to immediately recognize the speech signal upon acquiring each new speech segment without waiting for the complete input of the entire sentence or paragraph. This means that the engine can immediately extract text information upon each speech segment input, laying the foundation for real-time translation in subsequent steps. Simultaneously, this engine has incremental output characteristics: while processing the current speech segment, it can simultaneously output the corresponding target language text segment. This approach greatly improves the system's response speed and overall smoothness, ensuring rapid information transmission and interaction during live streaming.
[0028] As described in step S202 above, the current source language text is translated using the streaming translation engine to obtain the current target language text fragment. The current source language text fragment obtained in the previous step is then fed into the streaming translation engine for translation. The streaming translation engine is based on a pre-trained deep learning model and adopts a large language model architecture, such as Transformer or a similar structure. This allows the engine to receive and process text segments in real time, continuously generating corresponding target language text fragments. It is important to note that the translation is not a one-time processing of the entire information, but rather an incremental translation of each input source text fragment. This ensures that the system's output is highly synchronized with the broadcaster's real-time speech. The translation engine also possesses context-aware capabilities, enabling it to optimize the current translation using previously input information, ensuring that the translation result is natural, fluent, and accurately culturally and linguistically appropriate. It should be noted that the streaming translation engine can be trained on a pre-built first neural network model based on a preset sample set. The first neural network model uses a Transformer structure, where each sample data in the preset sample set includes sample text data (including multiple source text fragments) and the corresponding translated text. When training the pre-built first neural network model, the sample text data in each sample data is used as the input of the first neural network model, and the corresponding translated text in each sample data is used as the output of the first neural network model. Through training, the first neural network model can learn the correspondence between all possible sample text data and translated text. The trained first neural network model is used as a streaming translation engine.
[0029] As described in step S203 above, the previously output historical target language text fragments are corrected based on the current target language text fragment. The streaming translation engine utilizes its powerful context analysis capabilities to comprehensively consider the new translation fragments with the previously generated historical text. Specifically, if the current target language text fragment provides new information or context changes, the translation of the historical text needs to be adjusted. The correction process includes consistency adjustments to already translated terms, and even word rearrangement or rephrasing to conform to the grammar and conventions of the target language. This ensures that the expression of all relevant content remains consistent and accurate throughout long conversations or explanations. This dynamic correction capability enables the system to avoid misunderstandings or confusion caused by inconsistencies or ambiguities in translation when handling complex real-time language communication, thereby improving the overall translation quality.
[0030] As described in step S204 above, the corrected historical target language text fragments and the current target language text fragments are combined to obtain the target language text stream. The corrected historical target language text fragments are integrated with the currently generated target language text fragments to form an overall information stream referred to as the "target language text stream." The information is logically arranged considering the context. During this process, the system can also apply additional context algorithms to evaluate the importance of each text fragment in a given scenario, ensuring that the final target language text stream includes the most relevant and necessary content. Furthermore, the integrated text stream also needs to consider user feedback and interaction. By extracting context provided by user questions, the system can guide adjustments to the text order and optimize the final format to ensure it is suitable for real-time reading or listening by the audience.
[0031] In one embodiment, step S4, which involves performing semantic correlation analysis on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream, includes: S401: Using a contrastive learning algorithm, extract the semantic features of the target language text and the bullet screen question text respectively to obtain the first semantic feature vector and the second semantic feature vector respectively; S402: Calculate the similarity between the first semantic feature vector and the second semantic feature vector as the association value.
[0032] As described in step S401 above, semantic features of the target language text and the bullet screen question text are extracted using a contrastive learning algorithm, resulting in a first semantic feature vector and a second semantic feature vector. The contrastive learning algorithm is used to extract semantic features from the target language text stream and the bullet screen question text stream. Sentence-BERT or SimCSE can be used for semantic vector extraction. First, the target language text (i.e., the content translated by the anchor) and the bullet screen question text (related questions entered by the audience) are used as input data. These are then converted into semantic feature vectors by a pre-trained language model. For example, a text encoding model from natural language processing can be used to map each text segment into a high-dimensional space for subsequent similarity comparison. First, the system processes the target language text to generate a first semantic feature vector, which contains rich contextual information and semantic relationships within the text. Then, the system performs similar processing on the bullet screen question text to generate a second semantic feature vector. Through this method, the system can capture deep semantic information between texts, ensuring that even with different expressions or semantic changes, it can still be converted into a usable vector representation, laying the foundation for subsequent similarity calculations. Contrastive learning algorithms can effectively improve the discriminative power of text features, making the vector distance between similar texts closer and the vector distance between unrelated texts farther.
[0033] As described in step S402 above, the similarity between the first semantic feature vector and the second semantic feature vector is calculated as the association value. The similarity between the first semantic feature vector (representing the target language text) and the second semantic feature vector (representing the bullet screen question text) is calculated. This process is implemented using various similarity calculation algorithms, such as cosine similarity, Euclidean distance, or Manhattan distance. The most commonly used cosine similarity algorithm quantifies the similarity between two vectors by calculating the cosine of the angle between them. The value typically ranges from -1 to 1; the closer the value is to 1, the higher the similarity between the two. The calculated result is called the association value, which is a measure of a given pair of potentially similar content. Specifically, the system takes the extracted semantic feature vectors as input, performs mathematical operations to obtain a similarity value, and determines whether the bullet screen question text is related to the current target language text based on a set threshold.
[0034] In one embodiment, step S6, which generates corresponding response text for each of the target bullet screen question texts, includes: S601: Perform intent recognition on the target bullet screen question text to determine its compliant query category; S602: Based on the compliance query category, retrieve the corresponding structured compliance data from the pre-built cross-border commodity compliance database, and obtain the corresponding target response template from the pre-set response template; S603: Fill the structured compliance data into the target response template to generate the response text.
[0035] As described in step S601 above, intent recognition is performed on the target bullet screen question text to determine its corresponding compliance query category. Intent recognition is performed on the identified target bullet screen question text to determine the specific intent or query category of each question. Natural Language Processing (NLP) techniques, such as classifier models or deep learning models, can be used as input, with the target bullet screen question text as input. The system first parses the text, including word segmentation, stop word removal, and syntactic analysis, converting the text into an analyzable data format. Next, a trained model (e.g., based on support vector machines, decision trees, or neural networks) is used to extract features from the text content, followed by intent classification, mapping the question to a preset compliance query category. The compliance query categories in the compliance database include product ingredient queries, price inquiries, and halal certification requirements.
[0036] As described in step S602 above, based on the compliance query category, the system retrieves corresponding structured compliance data from a pre-built cross-border commodity compliance database and obtains a corresponding target response template from a pre-set response template library. Based on the confirmed compliance query category, data is retrieved from the pre-built cross-border commodity compliance database. This database contains relevant compliance information for commodities and typical user query data, aiming to ensure that each published response complies with the laws, regulations, and cultural habits of a specific market. During this process, the system sends a query request containing the key elements mentioned in the target bullet screen question text to accurately find the structured compliance data matching the query from the database. For example, if a user inquires about the halal certification of a product ingredient, the system will retrieve the product's certification status, relevant institutions, and valid certificate number. Simultaneously, to generate accurate and customer service-compliant responses, the system also needs to select a response template corresponding to the query category from a pre-set response template library. These templates are typically pre-designed with standardized formats to ensure information consistency and accuracy, and to optimize the user's reading experience. By combining compliance data with suitable response templates, the system lays a solid foundation for subsequent text generation.
[0037] As described in step S603 above, the structured compliance data is populated into the target response template to generate the response text. The retrieved structured compliance data is then combined with the selected response template to generate the final response text. This process typically involves multiple stages: First, the system identifies the placeholder positions in the standardized target response template. These placeholders usually represent data fields to be filled, such as product name, halal certification status, or certificate number. Then, the system dynamically populates the previously retrieved compliance data into the corresponding placeholders, generating a fluent and coherent text. For example, if a user inquires about the halal certification status of a lipstick, the generated response might be, "This product has obtained the relevant certification, number 561, and all ingredients comply with halal standards." This ensures that the response is both compliant and provides clear information to the user. The final generated response text is then sent back to the live stream interface for viewers to view in real time. Through this automated processing, the system not only improves the efficiency of response generation but also ensures the accuracy and relevance of information, making live stream interaction more efficient and user-friendly.
[0038] In one embodiment, after step S6 of generating corresponding response text for each of the target bullet screen question texts, the method further includes: S701: Obtain the sensitive word database corresponding to the target language text stream; S702: Perform sensitive word detection on the reply text using the aforementioned sensitive word database; S703: When the response text is detected to contain sensitive words from the sensitive word library, equivalent words that conform to the cultural habits of the target market are obtained and replaced.
[0039] As described in step S701 above, a sensitive word database corresponding to the target language text stream is obtained. This sensitive word database is a structured dataset that specifically records sensitive words that may cause negative reactions, cultural inappropriateness, or legal risks in a particular market or region. The content of the sensitive word database may vary for different target language markets to ensure appropriate respect for local culture, festivals, and social customs. In practice, the system automatically selects the corresponding sensitive word database version based on the target language text stream currently being processed. This enhances the compliance of interactive content and avoids the use of potentially misunderstanding-inducing or offensive language during live broadcasts. Furthermore, by obtaining this sensitive word database, the system can provide necessary basic data support for subsequent sensitive word detection and replacement processes, ensuring that the generated response text complies with local cultural customs and legal regulations.
[0040] As described in step S702 above, the response text is subjected to sensitive word detection using the sensitive word database. The generated response text is rigorously compared and searched against the corresponding sensitive word database to identify whether the text contains sensitive words. Specifically, Natural Language Processing (NLP) technology can be used to achieve efficient and accurate sensitive word detection. In practice, the system segments the response text into words and matches each word against words in the sensitive word database, using AC automata or Trie trees for multi-pattern matching. During the matching process, the system identifies whether there are expressions in the text that are the same as or similar to the sensitive words recorded in the database. This allows for the timely discovery and identification of sensitive words that may cause cultural conflicts or legal risks, providing a basis for subsequent modifications. For example, if the response mentions sensitive content, the system can automatically capture this key information. Through this mechanism, the system ensures that the generated content undergoes compliance review before publication, thereby protecting the reputation of the broadcaster and the platform and reducing complaints and negative impacts caused by inappropriate wording.
[0041] As described in step S703 above, when a sensitive word from the sensitive word library is detected in the response text, an equivalent word conforming to the cultural habits of the target market is obtained and replaced. Obtaining an equivalent expression corresponding to the sensitive word and conforming to the cultural habits of the target market ensures that the generated text maintains the accuracy of the information while avoiding unnecessary offense or cultural misunderstanding. The system first searches for equivalent words associated with the detected sensitive word in the sensitive word library. These are usually pre-prepared alternative expressions, ensuring they are semantically consistent with the sensitive word. For example, the original sensitive word "alcohol" might be replaced with "naturally fermented ingredients." During the replacement process, the system ensures that the concatenated text is fluent and grammatically correct, avoiding misleading or reduced clarity. This replacement mechanism not only plays a positive role in improving the viewer experience but also helps maintain the compliance of platform content, avoiding legal issues caused by language errors. Therefore, this step plays an important role in review and protection throughout the interaction process, making the live broadcast content more culturally adaptable and socially friendly.
[0042] In one embodiment, step S5, which filters target question texts with a relevance higher than a preset threshold from the bullet screen question text stream based on the correlation value, includes: S501: Add bullet screen question texts with associated values greater than a preset threshold to the pending processing queue; S502: Sort the bullet screen question texts in the queue to be processed according to a preset priority order; S503: Process the bullet screen question texts as the target bullet screen question texts according to the sorting results.
[0043] As described in step S501 above, bullet screen questions with a correlation value greater than a preset threshold are added to a processing queue. The preset threshold can be set based on historical data analysis and user behavior habits, and typically represents the relevance of the question to the live stream content, for example, set to 0.7. In this way, the system can filter out questions with low relevance that may distract attention or cause misunderstandings, concentrating resources on processing more important information, thereby improving viewer satisfaction and experience quality. Setting up a processing queue enables rapid response to high-priority questions, while low-priority questions may be processed later, further improving processing efficiency, reducing the system's computational burden, and ensuring that interactions during the live stream are not only timely but also content-relevant, promoting a positive viewer experience and engagement.
[0044] As described in step S502 above, the bullet screen questions in the queue to be processed are prioritized according to a preset priority order. The priority can be set based on various factors, including but not limited to the compliance of the question content, the urgency of the question, the frequency of repetition of the question, and the questioner's interaction history. The system may use a simple sorting algorithm or a complex machine learning model to analyze the relevant information of each question, thereby generating a comprehensive score. In live streaming scenarios, some queries are more critical than others; for example, questions about product safety, compliance, or allergens are often of higher priority than questions about product color or type. By prioritizing, the system can fully consider the diversity of audience information needs when processing questions to be fed back to the audience, thereby improving interaction efficiency and quality. The effectiveness and flexibility of this sorting can significantly improve the audience's viewing experience, ensuring that every audience member's question receives a timely response during the live stream.
[0045] As described in step S503 above, the bullet screen question texts are processed sequentially as target bullet screen question texts according to the sorting results. Based on the determined priority sorting results, the bullet screen question texts in the processing queue are processed sequentially. In specific implementation, the system will perform intent recognition, compliance query, and final response generation based on the current target bullet screen question text. By prioritizing the processing of highly relevant and high-priority bullet screen questions, the system can quickly respond to the audience's concerns while ensuring the accuracy and relevance of the response content. This process not only improves the interactivity of the live stream but also helps the broadcaster maintain a better sense of rhythm and interaction quality when responding to audience questions.
[0046] In one embodiment, before step S3 of receiving the real-time text stream of questions from multiple second users, the method further includes: S211: Obtain the language of the original bullet screen text stream; S212: Determine whether the language type of the original bullet screen text stream is the same as the language type of the target language text stream; S213: If they are different, the original bullet screen text stream is translated into the same language as the target language text stream to obtain the bullet screen question text stream.
[0047] As described in step S211 above, the language type of the original bullet screen text stream is obtained. This is achieved through a language recognition algorithm in Natural Language Processing (NLP). This algorithm analyzes the features of the bullet screen text, such as vocabulary, syntactic structure, and contextual information, to determine the language type. The importance of this step lies in its ability to quickly and accurately identify the language involved, laying the foundation for subsequent steps and ensuring accurate processing. In cross-language live streaming scenarios, viewers' questions may come from different countries and use different languages. Therefore, the system needs strong multilingual processing capabilities to promptly identify the language of the bullet screen text, ensuring the accuracy of subsequent processing. Once identification is complete, the system retains this information as an important variable, providing a basis for determining whether language conversion or translation is necessary.
[0048] As described in step S212 above, it is determined whether the language of the original bullet screen text stream is the same as the language of the target language text stream. If the system finds that the language of the original bullet screen text stream is inconsistent with the target language, it indicates that the viewer's question may be in a different language than the live broadcast content. This is common in cross-language communication, especially in international live broadcasts, where viewers may ask questions in different languages. Therefore, determining whether the language matches and selecting the appropriate language in this step directly affects the timeliness and accuracy of subsequent responses. If the determination is yes, the system can continue to process the bullet screen using the target language text stream; if the determination is no, translation processing is performed.
[0049] As described in step S213 above, if the languages are different, the original bullet screen text stream is translated into the same language as the target language text stream to obtain the bullet screen question text stream. This language conversion can be performed using an efficient machine translation model, such as a neural network translation model. This model can process single or multiple text segments, accurately translating text information from the source language to the target language through methods such as sentence splitting, word segmentation, and contextual understanding. In this process, the characteristics of streaming translation play a crucial role, allowing the system to translate any bullet screen text immediately after acquisition, ensuring that the generated content matches the real-time target language text stream. Through this step, audience questions in different languages can be converted into the target language designed by the system, maintaining coherence and consistency between the two, thus ensuring the interactivity and effectiveness of the live stream. This meets the needs of the audience, enabling each viewer to understand and participate in the discussion, enhancing the interactive effect and satisfaction of the live stream. Ultimately, the translated question text stream becomes the basis for subsequent semantic analysis and responses, ensuring that the system can provide more flexible and professional interaction based on audience feedback.
[0050] Reference Figure 3 The present invention also provides an interactive device for cross-language live streaming, the device comprising: The acquisition module 902 is used to acquire the source language speech stream of the first user in real time and divide the source language speech stream into multiple speech segment streams of preset duration; Translation module 904 is used for a streaming translation engine based on a large language model to perform real-time recognition and translation of each of the aforementioned speech segments and output the target language text stream. The receiving module 906 is used to receive multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Analysis module 908 is used to perform semantic correlation analysis on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; The filtering module 910 is used to filter out target question texts with a correlation degree higher than a preset threshold from the bullet screen question text stream based on the correlation value. The generation module 912 is used to generate corresponding response text for each of the target bullet screen question texts.
[0051] In one embodiment, the translation module 904 includes: The speech segment recognition submodule is used by the streaming translation engine based on a large language model to recognize speech segments in the currently acquired speech segment stream and obtain the current source language text; wherein, the streaming translation engine is configured to output the corresponding target language text segment synchronously for each input speech segment stream; The translation submodule is used to translate the current source language text using the streaming translation engine to obtain the current target language text fragment; The correction submodule is used to correct previously output historical target language text segments based on the current target language text segment; The synthesis submodule is used to synthesize the corrected historical target language text fragments and the current target language text fragments to obtain a target language text stream.
[0052] In one embodiment, the analysis module 908 includes: The semantic feature extraction submodule is used to extract the semantic features of the target language text and the bullet screen question text respectively through a contrastive learning algorithm, and obtain the first semantic feature vector and the second semantic feature vector respectively; The similarity calculation submodule is used to calculate the similarity between the first semantic feature vector and the second semantic feature vector, as the association value.
[0053] In one embodiment, the generation module 912 includes: The intent recognition submodule is used to perform intent recognition on the target bullet screen question text and determine its corresponding compliant query category; The structured compliance data retrieval submodule is used to retrieve the corresponding structured compliance data from the pre-built cross-border commodity compliance database according to the compliance query category, and to obtain the corresponding target response template from the pre-set response template; The structured compliance data population submodule is used to populate the structured compliance data into the target response template to generate the response text.
[0054] In one embodiment, the interactive device for cross-language live streaming further includes: The sensitive word database acquisition module is used to acquire the sensitive word database corresponding to the target language text stream; The sensitive word detection module is used to detect sensitive words in the response text using the sensitive word database. The equivalent vocabulary acquisition module is used to acquire and replace equivalent vocabulary that conforms to the cultural habits of the target market when the response text is detected to contain sensitive words from the sensitive word library.
[0055] In one embodiment, the filtering module 910 includes: The bullet screen question text addition submodule is used to add bullet screen question texts with associated values greater than a preset threshold to the pending queue; The priority sorting submodule is used to sort the bullet screen question texts in the queue to be processed according to a preset priority order. The target bullet screen question text processing submodule is used to process the bullet screen question text as the target bullet screen question text according to the sorting result.
[0056] In one embodiment, the interactive device for cross-language live streaming further includes: The language acquisition module is used to obtain the language of the original bullet screen text stream; The language type determination module is used to determine whether the language type of the original bullet screen text stream is the same as the language type of the target language text stream; The original bullet screen text translation module is used to translate the original bullet screen text stream into the same language as the target language text stream if they are different, thus obtaining the bullet screen question text stream.
[0057] Figure 4 An internal structural diagram of an electronic device in one embodiment is shown. This electronic device can specifically be a terminal or a server, and more specifically, a computer device. Figure 4 As shown, the electronic device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a cross-language live interactive method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement a cross-language live interactive method. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0058] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps: The source language speech stream of the first user is acquired in real time, and the source language speech stream is divided into multiple speech segment streams of preset duration; A streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, and outputs a target language text stream. The system receives multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; Based on the correlation value, target question texts with a correlation degree higher than a preset threshold are filtered out from the bullet screen question text stream; For each of the target bullet screen question texts, generate corresponding reply text.
[0059] By segmenting the source language speech stream in real time and performing incremental processing based on a streaming translation engine, extremely low latency output from speech to target language text stream is achieved, effectively ensuring the real-time nature of cross-language live streaming. Then, by performing online semantic association analysis between the bullet screen question text stream and the real-time generated target language text stream, questions highly relevant to the current explanation content can be accurately and automatically filtered from massive bullet screens, solving the problem of multimodal information fragmentation. Reply text is automatically generated for the filtered relevant questions, forming a complete interactive closed loop from real-time translation, intelligent question filtering to automatic response, greatly reducing the interactive burden on the anchor and significantly improving the accuracy and overall efficiency of live streaming interaction.
[0060] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps: The source language speech stream of the first user is acquired in real time, and the source language speech stream is divided into multiple speech segment streams of preset duration; A streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, and outputs a target language text stream. The system receives multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; Based on the correlation value, target question texts with a correlation degree higher than a preset threshold are filtered out from the bullet screen question text stream; For each of the target bullet screen question texts, generate corresponding reply text.
[0061] By segmenting the source language speech stream in real time and performing incremental processing based on a streaming translation engine, extremely low latency output from speech to target language text stream is achieved, effectively ensuring the real-time nature of cross-language live streaming. Then, by performing online semantic association analysis between the bullet screen question text stream and the real-time generated target language text stream, questions highly relevant to the current explanation content can be accurately and automatically filtered from massive bullet screens, solving the problem of multimodal information fragmentation. Reply text is automatically generated for the filtered relevant questions, forming a complete interactive closed loop from real-time translation, intelligent question filtering to automatic response, greatly reducing the interactive burden on the anchor and significantly improving the accuracy and overall efficiency of live streaming interaction.
[0062] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0063] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0064] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An interactive method for cross-language live streaming, characterized in that, The method includes: The source language speech stream of the first user is acquired in real time, and the source language speech stream is divided into multiple speech segment streams of preset duration; A streaming translation engine based on a large language model performs real-time recognition and translation of each of the aforementioned speech segments, and outputs a target language text stream. The system receives multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. Semantic correlation analysis is performed on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; Based on the correlation value, target question texts with a correlation degree higher than a preset threshold are filtered out from the bullet screen question text stream; For each of the target bullet screen question texts, generate corresponding reply text.
2. The interactive method for cross-language live streaming according to claim 1, characterized in that, The streaming translation engine based on a large language model performs real-time recognition and translation of each speech segment stream and outputs a target language text stream, including the following steps: The streaming translation engine based on a large language model identifies speech segments in the currently acquired speech segment stream to obtain the current source language text; wherein, the streaming translation engine is configured to output the corresponding target language text segment synchronously for each input speech segment stream; The current source language text is translated using the streaming translation engine to obtain a current target language text fragment; The previously output historical target language text fragments are corrected based on the current target language text fragment; The target language text stream is obtained by combining the corrected historical target language text fragments and the current target language text fragments.
3. The interactive method for cross-language live streaming according to claim 1, characterized in that, The step of performing semantic correlation analysis between each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream includes: By using a comparative learning algorithm, semantic features of the target language text and the bullet screen question text are extracted respectively, and a first semantic feature vector and a second semantic feature vector are obtained respectively. Calculate the similarity between the first semantic feature vector and the second semantic feature vector, and use it as the association value.
4. The interactive method for cross-language live streaming according to claim 1, characterized in that, The step of generating corresponding response text for each of the target bullet screen question texts includes: The intent of the target bullet screen question text is identified to determine its compliant query category; Based on the compliance query category, retrieve the corresponding structured compliance data from the pre-built cross-border commodity compliance database, and obtain the corresponding target response template from the pre-set response template; The structured compliance data is populated into the target response template to generate the response text.
5. The interactive method for cross-language live streaming according to claim 1, characterized in that, After the step of generating corresponding reply text for each of the target bullet screen question texts, the method further includes: Obtain the sensitive word database corresponding to the target language text stream; The response text is subjected to sensitive word detection using the aforementioned sensitive word database; When a sensitive word from the sensitive word library is detected in the response text, an equivalent word that conforms to the cultural habits of the target market is obtained and replaced.
6. The interactive method for cross-language live streaming according to claim 1, characterized in that, The step of filtering target question texts with a relevance higher than a preset threshold from the bullet screen question text stream based on the correlation value includes: Add bullet screen question texts with associated values greater than a preset threshold to the pending processing queue; The bullet screen question texts in the queue to be processed are sorted according to a preset priority order; Based on the sorting results, the bullet screen question texts are processed as the target bullet screen question texts in sequence.
7. The interactive method for cross-language live streaming according to claim 1, characterized in that, Before the step of receiving the real-time text stream of questions from multiple second users, the method further includes: Obtain the language of the original bullet screen text stream; Determine whether the language of the original bullet screen text stream is the same as the language of the target language text stream; If they are different, the original bullet screen text stream is translated into the same language as the target language text stream to obtain the bullet screen question text stream.
8. An interactive device for cross-language live streaming, characterized in that, The device includes: The acquisition module is used to acquire the source language speech stream of the first user in real time and divide the source language speech stream into multiple speech segment streams of preset duration; The translation module is used in a streaming translation engine based on a large language model to perform real-time recognition and translation of each of the aforementioned speech segments and output the target language text stream. The receiving module is used to receive multiple bullet screen question text streams from second users in real time; wherein the target language text stream is in the same language as the bullet screen question text stream. The analysis module is used to perform semantic correlation analysis on each question text in each of the bullet screen question text streams and the target language text stream to obtain the correlation value between each question text and the target language text stream; The filtering module is used to filter out target question texts with a relevance higher than a preset threshold from the bullet screen question text stream based on the correlation value; The generation module is used to generate corresponding response text for each of the target bullet screen question texts.
9. A computer-readable storage medium, characterized in that, The system contains a computer program that, when executed by a processor, causes the processor to perform the steps of the interactive method for cross-language live streaming as described in any one of claims 1 to 7.
10. An electronic device, characterized in that, The device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the interactive method for cross-language live streaming as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Bullet screen management method and device, terminal equipment and storage medium
CN112765336A
Real-time interaction method and terminal based on large language model
CN120768868A
Multi-language voice content recognition method and system
CN121662048A