Method, device, equipment, medium and product for restoring variant words in live broadcast scenarios
By training live sample sentences and performing high concurrency deployment of dynamic batch inference, a variant word restoration model and interface API are built, which solves the problem of difficulty in restoring variant words in live broadcast scenarios, and improves the accuracy and system performance of the restoration.
Patent Information
- Application Number
- CN202510449369.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The reproduction of variant words in live broadcast scenarios is difficult and lacks robustness. The existing technology cannot effectively identify and restore the location of variant words, making false propaganda difficult to detect.
By training the variant word restoration model based on multiple live sample sentences, and performing high concurrency deployment of dynamic batch inference, the inference interface API is built to achieve efficient restoration of variant words.
It significantly improves the system's concurrency performance and resource utilization, improves the accuracy and robustness of variant word restoration, and realizes efficient and accurate variant word restoration in live broadcast scenarios.
Smart Images

Figure CN119962533B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a method, device, equipment, medium and product for restoring variant words in a live broadcast scenario. Background Art
[0002] To boost sales and attract customers during livestreams, many businesses circumvent platform censorship by using word variants, thereby engaging in false advertising. For example, one business used a phrase like "If you use our products, you won't need to go to a certain hospital or pay for a doctor's appointment." Using word variants like "a certain hospital" and "a white coat" to replace "hospital" and "doctor," the business implied that its products had medicinal properties. These word variants are widely used in various scenarios and are flexible and diverse, posing a significant challenge to standardized management in the industry. Therefore, resolving word variants in livestreams is crucial for protecting consumer rights and promoting standardized industry development.
[0003] However, current research on variant words primarily focuses on social media and the underground industry, while livestreaming, as an emerging context, has received insufficient attention. Variant words in livestreaming differ significantly from those in social media and the underground industry, both in form and subject matter. Furthermore, the location of variant words is currently unknown, making them less robust in practical applications. Summary of the Invention
[0004] The embodiments of the present application provide a method, device, equipment, medium and product for restoring variant words in a live broadcast scenario, so as to at least solve the problems in the related art of difficulty in restoring variant words in a live broadcast scenario and lack of robustness in practical applications.
[0005] In a first aspect, an embodiment of the present application provides a method for restoring variant words in a live broadcast scenario, comprising:
[0006] Receiving live variant word restoration requests respectively sent by multiple requesting parties, wherein the time interval between sending the multiple live variant word restoration requests satisfies a preset condition, and the live variant word restoration requests include live variant word texts to be restored;
[0007] Calling an inference API to infer the live broadcast variant word text to be restored to obtain a live broadcast variant word restoration result. The inference API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live broadcast sample sentence pairs;
[0008] Determining position information corresponding to the variant word restoration based on the live variant word restoration result and the live variant word text to be restored;
[0009] The live broadcast variant word restoration result and location information are sent to the corresponding requesting party through the API.
[0010] In a second aspect, an embodiment of the present application provides a device for restoring variant words in a live broadcast scenario, the device comprising:
[0011] a receiving module configured to receive live variant word restoration requests respectively sent by a plurality of requesting parties, wherein the time interval between sending the plurality of live variant word restoration requests satisfies a preset condition, and the live variant word restoration requests include live variant word texts to be restored;
[0012] A calling module is configured to call an inference API to perform inference on the live broadcast variant word text to be restored, thereby obtaining a live broadcast variant word restoration result. The inference API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live broadcast sample sentence pairs.
[0013] A determination module, configured to determine position information corresponding to the variant word restoration based on the live variant word restoration result and the live variant word text to be restored;
[0014] The sending module is used to send the live broadcast variant word restoration result and location information to the corresponding requesting party through the API.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the processor implements the steps of the method for restoring variant words in a live broadcast scene as described in any one of the embodiments of the first aspect.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the steps of the method for restoring variant words in a live broadcast scene as described in any one of the embodiments of the first aspect are implemented.
[0017] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the variant word restoration method for a live broadcast scenario provided in the first aspect of the embodiment of the present application.
[0018] The variant word restoration method, device, equipment, medium and product for live broadcast scenarios in the embodiments of the present application are based on a variant word restoration model obtained by training multiple live broadcast sample sentences, and an inference interface API obtained by encapsulating the variant word restoration model after high-concurrency deployment based on dynamic batch reasoning. The inference interface API is called to process batches of live broadcast variant word restoration requests, and the live broadcast variant word restoration results and position information corresponding to the live broadcast variant word text to be restored in the live broadcast variant word restoration request are obtained, and the information is fed back to the request sender. It can be seen that the variant word model deployment based on dynamic batch reasoning significantly improves the system's concurrency performance and resource utilization, and improves the model's variant word restoration performance and robustness, because the inference batch is dynamic, thereby realizing efficient and accurate variant word restoration processing in live broadcast scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 This is a flow chart of a method for restoring variant words in a live broadcast scenario provided by an embodiment of the present application;
[0021] Figure 2 Schematic diagram of the process of training and deploying the variant word restoration model provided in the embodiment of the present application;
[0022] Figure 3 This is a structural diagram of a variant word restoration device for a live broadcast scenario provided by an embodiment of the present application;
[0023] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present application.
[0024] Reference numerals:
[0025] The variant word restoration device 300 of the live broadcast scene includes a receiving module 301, a calling module 302, a determining module 303, and a sending module 304.
[0026] Electronic device 400 , processor 401 , memory 402 , communication interface 403 , bus 410 . DETAILED DESCRIPTION
[0027] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0028] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, the elements defined by the phrase "comprising..." do not exclude the presence of other identical elements in the process, method, article, or device comprising the elements.
[0029] To boost sales and attract customers during livestreams, many businesses circumvent platform censorship by using word variants, thereby engaging in false advertising. For example, one business used a phrase like "If you use our products, you won't need to go to a certain hospital or pay for a doctor's appointment." Using word variants like "a certain hospital" and "a white coat" to replace "hospital" and "doctor," the business implied that its products had medicinal properties. These word variants are widely used in various scenarios and are flexible and diverse, posing a significant challenge to standardized management in the industry. Therefore, resolving word variants in livestreams is crucial for protecting consumer rights and promoting standardized industry development.
[0030] The current research on variant words mainly focuses on social media and underground industries. As an emerging scenario, the live broadcast scenario has not received sufficient attention. The variant words in the live broadcast scenario are quite different from those in social media and underground industries. Firstly, the forms of variant words are different: the scenarios mentioned above are all visual scenarios, so the variant words in this scenario mainly look similar to the original words visually. For example, "muqipai" and "2 xiaoricun" are used to represent "qipai" and "2 hours"; while the variant words in the live broadcast scenario mainly insert some meaningless words while maintaining the integrity of the rhythm. For example, "mouyi mouyuan" and "mian shenme yi" are used to represent "hospital" and "immune". Secondly, the involved themes are different: the variant words in social media mainly involve current events and politics, the variant words in underground industries mainly involve illegal industries, while the variant words in the live broadcast scenario mainly involve the health and medical industries.
[0031] Currently, for the research on restoring variant words in scenarios such as social media and underground industries, pre-trained models such as BERT are applied to restore complex words. BERT uses the Transformer architecture, and the pre-training tasks include the Masked Language Model (MLM) task and the Next Sequence Prediction (NSP) task; when training the Masked Language Model MLM, some words are randomly masked and the model is required to predict the masked parts; the Next Sentence Prediction task NSP is to enable BERT to learn the relationships between sentences; therefore, when the masked part is a variant word, based on the pre-training task of the Masked Language Model, BERT will try to predict the replacement word for the masked part, that is, the original word of the variant word. However, this method has significant drawbacks. Firstly, since the positions of variant words are unknown, an additional judgment module needs to be trained to identify which words in the sentence are variant words, which greatly increases the complexity and computational cost of the system. Secondly, this paradigm is only applicable to the case where the length of the variant word is equal to that of the original word and cannot handle the case where the lengths are different, resulting in a lack of robustness in practical applications.
[0032] To solve the problems of related technologies, the embodiments of this application provide a method, device, equipment, medium and product for restoring variant words in the live broadcast scenario.
[0033] The following will combine the accompanying drawings to detail the method for restoring variant words in the live broadcast scenario provided by the embodiments of this application through specific embodiments and their application scenarios.
[0034] Figure 1 The flowchart of the method 100 for restoring variant words in the live broadcast scenario according to the embodiments of this application is shown. As Figure 1 shown, the method 100 for restoring variant words in the live broadcast scenario may specifically include the following steps:
[0035] S101: Receive live variant word restoration requests respectively sent by multiple requesting parties, where the time interval between sending the multiple live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored;
[0036] S102: Calling an inference API to perform inference on the live broadcast variant word text to be restored to obtain a live broadcast variant word restoration result. The inference API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live broadcast sample sentence pairs.
[0037] S103, determining position information corresponding to the variant word restoration based on the live broadcast variant word restoration result and the live broadcast variant word text to be restored;
[0038] S104: Send the live broadcast variant word restoration result and location information to the corresponding requesting party through the API.
[0039] Therefore, based on the variant word restoration model obtained by training multiple live sample sentences, and the inference interface API obtained by encapsulating the variant word restoration model after high-concurrency deployment based on dynamic batch reasoning, the inference interface API is called to process batches of live variant word restoration requests, and the live variant word restoration results and position information corresponding to the live variant word text to be restored in the live variant word restoration request are obtained, and the information is fed back to the request sender. It can be seen that the variant word model deployment based on dynamic batch reasoning significantly improves the system's concurrency performance and resource utilization, and improves the model's variant word restoration performance and robustness, because the inference batch is dynamic, and realizes efficient and accurate variant word restoration processing in live broadcast scenarios.
[0040] The specific implementation methods of the above steps are introduced below.
[0041] In some embodiments, in S101, the time interval for sending multiple live variant word restoration requests meets a preset condition, which includes: multiple requesting parties corresponding to these requests send the requests simultaneously, or send them at an extremely short time interval (for example, 0.1 seconds).
[0042] It should be noted that the live broadcast variant text to be restored in the request body is obtained by performing speech recognition on the acquired live broadcast video. Specifically, the live broadcast content on the live broadcast platform is crawled to obtain multiple live broadcast video data. The live broadcast video data is then recognized using a speech recognition model to obtain the text data corresponding to the live broadcast video data.
[0043] In some embodiments, before S102, a variant word restoration model is obtained by training multiple live sample sentences, and then the variant word restoration model is deployed with high concurrency based on dynamic batch reasoning and then encapsulated to obtain an inference interface API. Specifically, the model training and deployment process provided in the embodiment of this application can be referred to Figure 2 .
[0044] like Figure 2 As shown, the training and deployment method 200 of the variant word restoration model provided in the embodiment of the present application may include the following steps: S201 to S206.
[0045] S201. Acquire a plurality of live broadcast sample sentence pairs, where the live broadcast sample sentence pairs include two mutually corresponding live broadcast sample sentences, and the data type of the live broadcast sample sentences is text data.
[0046] In some embodiments, a plurality of live video data and live text data corresponding to each of the live video data are obtained; the live text data are semantically segmented into a plurality of first live sample sentences; it is determined whether the first live sample sentence includes a first live variant word; if the first live sample sentence does not include the first live variant word, the first live sample sentence is determined as a live negative sample; if the first live sample sentence includes the first live variant word, the first live variant word in the first live sample sentence is marked and the corresponding first live original word is annotated to obtain a live positive sample; the live positive sample and the live negative sample are determined as a live sample sentence pair.
[0047] During the specific implementation, a crawler tool is used to crawl the live content on the live broadcast platform and obtain live video clips. After the crawler tool is running, it will cyclically read the live link pool in the configuration file, open a thread for each live link, and download the live video content in parallel. The live link pool stores the live link currently being crawled. It contains two fields: the live broadcast name and the live link, both of which are text type. To ensure the diversity and coverage of the data, different time periods, different anchors, and different types of health products are selected for data crawling. A total of multiple (for example, 6273) live broadcast clips are crawled, and each live broadcast clip is set to 60 seconds. This length can not only ensure that a sufficient amount of information is captured, but also reduce the complexity of data processing to a certain extent.
[0048] In specific implementations, automatic speech recognition technology is used to transcribe live video clips into text data and screen it. Optionally, an end-to-end speech recognition model is built and trained on a manually annotated Mandarin speech recognition dataset. Furthermore, after transcribing the video into text, manual screening is performed to remove useless data such as background music, ultimately resulting in multiple (e.g., 6,223) high-quality live text data sets.
[0049] In specific implementation, sentence segmentation technology is used to semantically segment the transcribed live broadcast text fragments into sentences. After the text fragments are divided into sentences, meaningless short sentences or repeated sentences are screened out to improve the accuracy and efficiency of subsequent analysis. In other words, multiple first live broadcast sample sentences are obtained.
[0050] In the specific implementation, manual annotation is adopted. First, multiple first live broadcast sample sentences are checked to determine whether they contain variant words. If variant words are contained, they are marked and the corresponding original words are annotated. In other words, manual annotation is adopted to provide the annotators with ASR transcription text, and combine it with the corresponding live broadcast clip as auxiliary information. In the annotation process, the annotators need to mark the variant words in the sentences and add the corresponding original words as annotations. In the end, 8236 positive samples containing variant words and 78554 negative samples that do not contain variant words are obtained. In this way, through this annotation process, the live broadcast sample sentence pairs required for subsequent model training and the variant word vocabulary required for subsequent data enhancement can be generated.
[0051] Furthermore, in some embodiments, when the first live broadcast sample sentence includes a first live broadcast variant word, the first live broadcast variant word in the first live broadcast sample sentence is marked and the corresponding first live broadcast original word is annotated. After obtaining a live broadcast positive sample, a live broadcast variant word vocabulary is constructed based on the first live broadcast variant word and the first live broadcast original word; a number of second live broadcast variant words and second live broadcast original words matching the second live broadcast variant word are randomly extracted from the live broadcast variant word vocabulary, and the second live broadcast variant word is any one of the multiple first live broadcast variant words in the live broadcast variant word vocabulary; a large language model is used to generate multiple second live broadcast sample sentences containing the second live broadcast original word; the second live broadcast original word in the second live broadcast sample sentence is replaced with a second live broadcast variant word matching it to obtain a third live broadcast sample sentence; the second live broadcast sample sentence and the third live broadcast sample sentence are determined as a live broadcast sample sentence pair.
[0052] Specifically, a dictionary of variant words is constructed based on the annotation results, for example, containing 431 original words and their 2,688 variant forms. Next, several variant words and their corresponding original words are randomly extracted from the variant word dictionary to generate fluent text that fits the topic and contains all the original words. Subsequently, the original words in the text generated by the large language model are replaced with the corresponding variant words to form sentence pairs.
[0053] In specific implementation, several variant words are randomly selected from a dictionary of variant words. Based on the dictionary, the original words corresponding to the variant words are obtained. A large language model is then used to generate multiple sentences containing the original words that match the topic. The original words are then replaced with the corresponding variant words. This constructs a set of sentence pairs consisting of source and target texts. The target text is the source text with the variant words replaced by the corresponding original words. This method generates 11,280 positive samples and 2,155 negative samples containing variant words.
[0054] Taking the example of using a large language model to generate five sentences that contain the original word and are consistent with the topic, the large language model prompt template can be as follows: "You are a live broadcast host, responsible for promoting products. You need to generate five promotional sentences that contain the target word. Below are some real promotional sentences for you to imitate. The generated sentences should not have repeated meanings. The target word should remain unchanged. The length of the sentence should be as consistent as possible with the provided example."
[0055] In this way, a high-quality dataset of variant words in live broadcast scenarios was constructed through a combination of manual annotation and large-scale model data augmentation, addressing the scarcity of datasets for this scenario. Specifically, in subsequent model training, the manually annotated training set and the data augmentation dataset can be mixed as the training set. The live broadcast dataset contains multiple live video text sentence pairs, each consisting of a source text and a target text.
[0056] Furthermore, in some embodiments, data preprocessing is performed on the live sample sentence pairs: the source sentences input to the model are defined ,in is the sentence length, For the first The maximum length of the sentence is 128 characters; if the sentence length exceeds 128, the sentence is divided into clauses in sequence, and each clause does not exceed the specified maximum length.
[0057] S202: Segment the live sample sentences using BPE to obtain corresponding token sequences.
[0058] In some embodiments, the source sentence of the model input is passed through T5Tokenizer Through BPE word segmentation, the unit is token, and each token has a separate ID. Get the sequence ,in is the length of the token sequence, is the i-th token in the token sequence. It should be noted that the target sentence in the live sample sentence pair also undergoes the same preprocessing and word segmentation as the source sentence. In this way, the pretrained language model tokenizes the input text, converting it into a form that the model can understand.
[0059] S203: Divide the plurality of live sample sentence pairs according to a preset ratio to obtain a training set, a validation set, and a test set.
[0060] That is, based on the live broadcast variant word dataset, a training set, a validation set, and a test set are constructed, for example, in a ratio of 8:1:1, and during the training process, a mixture of the manually annotated training set and the data augmentation dataset is used as the training set.
[0061] S204: Input the token sequence corresponding to the training set into the preset model for training, then input the token sequence corresponding to the validation set into the preset model for verification, and input the token sequence corresponding to the test set into the preset model for testing.
[0062] In some embodiments, the preset model includes an encoder component and a decoder component. The embedding layer of the preset model performs token embedding and position embedding on the token sequence to obtain a corresponding dense vector. The encoder component outputs a contextual representation corresponding to the dense vector. The contextual representation is used to guide the output of the decoder component at each time step.
[0063] In specific implementation, the token sequence of the source sentence Input to the end-to-end model mengzi-T5-base model, first it will pass through the embedding layer to obtain token embedding and position embedding A dense vector of .
[0064] Mengzi-T5-base uses the T5 model training paradigm and is retrained on a large Chinese corpus. It includes encoder and decoder components, each consisting of 12 Transformer layers with 12 attention heads to capture the dependencies between input and output sequences. This design enables Mengzi-T5-base to excel in Chinese natural language processing tasks, especially in generative tasks, with high flexibility and accuracy.
[0065] In specific implementation, the dense vector passes through the Encoder component of the model, which consists of 12 layers of multi-head Attention layers to obtain the contextual representation of the target sentence .
[0066] In the specific implementation, the input start symbol is input in the first time step of the Decoder stage. As the input of the first time step, it passes through the embedding layer and the Decoder component. The Decoder component consists of 12 layers of multi-head Attention layers. In the process of Attention calculation, the cross attention mechanism is used to calculate the input of each time step of the Decoder. , capturing the contextual information of the sequence. At each time step of the model, the output passes through a linear layer to map the Decoder output to the dimension of the target vocabulary. Subsequently, the SoftMax function is applied to the output of the linear layer to convert these values into a probability distribution. The output is recorded as ,in Is the output of the first time step, as the input of the second time step. Until the last time step output terminator The sentence probability output by the model is:
[0067] = ; (1)
[0068] S205. Determine whether at least one of the number of iterative training times or the test loss value of the preset model meets a preset training stop condition. If not, adjust the weight parameters of the preset model until the preset training stop condition is met, thereby obtaining a variant word restoration model.
[0069] Optionally, the preset training stop condition includes at least one of the following: the number of iterative training reaches a first preset threshold, and the number of times the test loss value remains continuously without decreasing reaches a second preset threshold.
[0070] The test loss value is obtained based on the test set and the loss function.
[0071] The loss function is:
[0072] ; (2)
[0073] in, represents the batch size, and T represents the maximum length of the token sequence (target sequence); Represents the mask matrix, ignoring the loss of padding tokens; Represents the token predicted by the model's i-th sample at the t-th time step probability.
[0074] Optionally, in this example the initial learning rate is set to , using Adam optimizer, epoch is set to 20, Set to 32 and warm_steps to 200.
[0075] In this way, by training the model on the training set, calculating the model loss, performing back propagation, updating the model weight parameters, and performing iterative operations, the variant word restoration model M can be obtained by training the variant word dataset.
[0076] S206: Perform high-concurrency deployment based on dynamic batch reasoning on the variant word restoration model, and then encapsulate it to obtain an inference interface API.
[0077] When implementing, create a global queue Collect data from multiple requests, each request will be pushed into the queue and wait to be processed. Each element in the queue is ,in To request the data you want to process, use to mark each request.
[0078] When implementing, set the maximum batch size and maximum waiting time The maximum waiting time is 40ms. Before the maximum waiting time is reached, the request data will be collected as much as possible and divided into batches according to the maximum batch size. If the number of requests in the queue reaches the maximum batch size, or the waiting time reaches the maximum waiting time, the batch is divided. . For the inference batches, including Requests, is the first Requests, Less than the maximum batch size.
[0079] In specific implementation, the divided batches are segmented and fed into the model in order. Inference results are obtained , defined as follows:
[0080] ; (3)
[0081] in, For the Inference batches, For the The inference results of the inference batches.
[0082] In specific implementation, the inference results Contains The request output inference results and corresponding ,according to Split and return the results to the corresponding request. That is, assign an ID to each request and return the results to the corresponding request by query ID.
[0083] In this way, the variant word model deployment method based on dynamic batch inference can significantly improve the system's concurrency performance and resource utilization by dynamically adjusting the batch size and optimizing the inference process.
[0084] The above is a specific implementation of the variant word restoration model trained and then encapsulated to obtain the inference interface API provided by the embodiments of this application. Using an end-to-end model, the variant word detection and correction tasks are unified into a generation task based on the seq2seq paradigm, improving the model's variant word restoration performance and robustness.
[0085] Back to Figure 1 The variant word restoration method 100 for the live broadcast scenario shown performs variant word restoration on actual business text.
[0086] For S102 to S104, the specific implementation method may be as follows:
[0087] Towards Send a request The request body contains the variant word text that needs to be restored , the input obtained by the model , defined as follows:
[0088] ; (4)
[0089] in, Data required for model inference, including text And other necessary preprocessing information.
[0090] Next, the model Data sent upon request Make inferences, reasoning results pass Return, defined as follows:
[0091] ; (5)
[0092] in, Contains the results of restoring variant words in text S, as well as information such as the position of variant words in the sentence.
[0093] As an optional embodiment, the difference between the input and output sequences can be determined by string comparison, that is, the corresponding token number. In this way, by comparing the difference between the input and output sequences, the word position can be recorded and the content can be corrected.
[0094] In addition, in order to verify the performance of the variant word restoration method for the live broadcast scenario of the embodiment of the present application, the superiority of this method is illustrated below through automatic indicator evaluation.
[0095] Specifically, the evaluation is performed using test sets from different sources. Test set 1 is the data from the same live link pool as the training set. Its data has no duplication with the training set and the validation set. It contains a total of 1,600 samples, which are manually annotated, including 400 positive samples and 400 negative samples. In addition, each positive sample contains the manually annotated result as a reference. Test set 2 is a new live link pool that has no overlap with the live link pool in the training set. It contains a total of 800 samples, including 400 positive samples and 400 negative samples. Each positive sample contains the manually annotated result as a reference.
[0096] This method uses The value is evaluated. It is defined as follows:
[0097] ; (6)
[0098] ; (7)
[0099] ; (8)
[0100] ; (9)
[0101] in, To correctly restore the positive sub-sentences of all variant words; is an uncorrected counterexample; There are positive examples where variant words are not restored; is a counterexample of incorrect restoration.
[0102] Compare the method of the embodiment of the present application with the other methods: (Large Language Model), (N-gram language model), (Showing sequence-to-sequence modeling editing operations), (Sequence-to-Sequence Models for Convolutional Networks), (Pre-trained language model with self-noising encoding) Tables 1 and 2 below show the comparative experimental results for test set 1 and test set 2, respectively.
[0103] Table 1
[0104] method Acc Precision Recall F1 LLM 0.605 0.660 0.529 0.587 Kenlm 0.583 0.607 0.372 0.537 Seq2Edit 0.651 0.968 0.361 0.526 Convseq2seq 0.740 0.978 0.527 0.685 BART 0.708 0.701 0.767 0.738 This embodiment 0.928 0.937 0.927 0.932
[0105] Table 2
[0106] method Acc Precision Recall F1 LLM 0.373 0.450 0.459 0.346 Kenlm 0.516 0.515 0.513 0.514 Seq2seq 0.702 0.987 0.408 0.588 Convseq2seq 0.687 0.898 0.421 0.573 BART 0.656 0.670 0.611 0.639 This embodiment 0.863 0.929 0.787 0.852
[0107] As shown in Table 1 and Table 2, the method of this embodiment achieves the highest The experimental results presented show that character-level error correction methods (such as Seq2Edit and the statistical language model Kenlm) are insufficient when handling morphological variations in live broadcast scenarios. In contrast, Seq2seq models (such as Convseq2seq, BART, and T5) perform better in managing inconsistent output lengths. Notably, the T5 model in this example achieved the highest F1 score on both test sets, demonstrating its effectiveness in this task.
[0108] Furthermore, Table 3 below provides a method for efficient deployment of the model. In order to verify its beneficial indicators, the efficiency of the method of the present application is illustrated through performance testing.
[0109] Table 3. Concurrency test comparison
[0110]
[0111] As shown in Table 3, a stress test was conducted using a tool for API testing and performance evaluation, with a concurrency of 10 and a round number of 5. The memory usage is the maximum memory usage recorded during the stress test, in MB; the test duration is the test time, in seconds (s); the throughput is the number of requests processed per second, in R / s; and the error rate is the percentage of requests sent that resulted in errors. The dynamic inference-based method consumes additional memory compared to the thread pool due to the dynamic inference batches, but significantly saves inference time. Specifically, dynamic inference increases the system's concurrency by 789%, greatly optimizing the system's concurrency capabilities. The dynamic inference method significantly outperforms the thread pool in performance, especially in terms of throughput and time. Despite the higher memory usage, it significantly improves the system's concurrency capabilities.
[0112] Thus, the embodiments of this application, by constructing a high-quality variant word dataset and developing a variant word restoration method based on an end-to-end architecture, are of great significance in the field of variant word restoration technology in live broadcast scenarios, enriching the technical system in this field. This method not only demonstrates excellent performance but also exhibits strong generalization capabilities, effectively handling complex and changing variant word scenarios.
[0113] In addition, the high-performance concurrent deployment algorithm for the variant word restoration model in the embodiment of the present application is based on a dynamic batch deployment strategy, which greatly improves the system's concurrency capability, real-time processing efficiency and resource utilization.
[0114] The embodiments of the present application overcome the problems of difficulty in detecting variant words, low efficiency, and lack of training data in current live broadcast scenarios. Compared with traditional rule-based variant word restoration methods, the robustness and accuracy of restoration are significantly improved.
[0115] It should be noted that the above description is limited to some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0116] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a variant word restoration device 300 for a live broadcast scenario.
[0117] like Figure 3 As shown, the variant word restoration device 300 of the live broadcast scene may include:
[0118] A receiving module 301 is configured to receive live variant word restoration requests respectively sent by multiple requesting parties, wherein the time interval between sending the multiple live variant word restoration requests satisfies a preset condition, and the live variant word restoration requests include live variant word texts to be restored;
[0119] Calling module 302 is configured to call an inference API to perform inference on the live broadcast variant word text to be restored, thereby obtaining a live broadcast variant word restoration result. The inference API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live broadcast sample sentence pairs.
[0120] A determination module 303 is configured to determine location information corresponding to the variant word restoration based on the live variant word restoration result and the live variant word text to be restored;
[0121] The sending module 304 is used to send the live broadcast variant word restoration result and location information to the corresponding requesting party through the API.
[0122] In some embodiments, the variant word restoration device 300 for live broadcast scenarios further includes a training deployment module ( Figure 3 Specifically, the training deployment module includes the following units:
[0123] An acquiring unit, configured to acquire a plurality of live sample sentence pairs, wherein the live sample sentence pairs include two mutually corresponding live sample sentences, and the data type of the live sample sentences is text data;
[0124] A word segmentation unit is used to segment the live sample sentences through BPE to obtain a corresponding token sequence;
[0125] A division unit, configured to divide the plurality of live sample sentence pairs according to a preset ratio to obtain a training set, a validation set, and a test set;
[0126] An input unit, configured to input the token sequence corresponding to the training set into a preset model for training, and then input the token sequence corresponding to the validation set into the preset model for verification, and the token sequence corresponding to the test set for testing;
[0127] an adjustment unit, configured to determine whether at least one of the number of iterative training times or the test loss value of the preset model satisfies a preset training stop condition; if not, adjust a weight parameter of the preset model until the preset training stop condition is satisfied, thereby obtaining a variant word restoration model;
[0128] The encapsulation unit is used to perform high-concurrency deployment of the variant word restoration model based on dynamic batch reasoning, and then encapsulate it to obtain an inference interface API.
[0129] In some optional embodiments, the acquisition unit is specifically used to: acquire multiple live video data and live text data corresponding to each of the live video data; segment the live text data into multiple first live sample sentences according to semantics; determine whether the first live sample sentence includes a first live variant word; if the first live sample sentence does not include the first live variant word, determine the first live sample sentence as a live negative sample; if the first live sample sentence includes the first live variant word, mark the first live variant word in the first live sample sentence and annotate the corresponding first live original word to obtain a live positive sample; determine the live positive sample and the live negative sample as a live sample sentence pair.
[0130] In some optional embodiments, the acquisition unit is further specifically used to: construct a live broadcast variant word library based on the first live broadcast variant word and the first live broadcast original word; randomly extract a number of second live broadcast variant words and second live broadcast original words matching the second live broadcast variant word from the live broadcast variant word library, where the second live broadcast variant word is any one of the multiple first live broadcast variant words in the live broadcast variant word library; use a large language model to generate multiple second live broadcast sample sentences containing the second live broadcast original word; replace the second live broadcast original word in the second live broadcast sample sentence with the second live broadcast variant word matching it to obtain a third live broadcast sample sentence; determine the second live broadcast sample sentence and the third live broadcast sample sentence as a live broadcast sample sentence pair.
[0131] Optionally, the preset model includes an Encoder component and a Decoder component.
[0132] In some optional embodiments, the input unit is specifically used to: perform token embedding and position embedding on the token sequence through the embedding layer of the preset model to obtain a corresponding dense vector; output the context representation corresponding to the dense vector through the Encoder component; and use the context representation to guide the output of each time step of the Decoder component.
[0133] Optionally, the test loss value is obtained based on the test set and the loss function;
[0134] The loss function is:
[0135] ;
[0136] in, Indicates the batch size, T indicates the maximum length of the token sequence; represents the mask matrix; Represents the token predicted by the model's i-th sample at the t-th time step probability.
[0137] It should be noted that, for the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0138] The device of the above embodiment is used to implement the variant word restoration method of the corresponding live broadcast scenario in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0139] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an electronic device.
[0140] Figure 4 A more specific hardware structure diagram of an electronic device provided by this embodiment is shown.
[0141] The electronic device 400 may include a processor 401 and a memory 402 storing computer program instructions.
[0142] Specifically, the processor 401 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0143] Memory 402 may include a large-capacity memory for data or instructions. By way of example and not limitation, memory 402 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 402 may include removable or non-removable (or fixed) media. Where appropriate, memory 402 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, memory 402 is a non-volatile solid-state memory.
[0144] In certain embodiments, the memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present application.
[0145] The processor 401 reads and executes the computer program instructions stored in the memory 402 to implement the variant word restoration method for any live broadcast scenario in the above embodiments.
[0146] In some examples, the electronic device 400 may further include a communication interface 403 and a bus 410. Figure 4 As shown, the processor 401 , the memory 402 , and the communication interface 403 are connected via a bus 410 and communicate with each other.
[0147] The communication interface 403 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0148] Bus 410 includes hardware, software, or both, and couples the components of the online data traffic metering device to each other. By way of example, and not limitation, bus 410 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industrial Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industrial Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Area Network (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 410 may include one or more buses. Although the embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.
[0149] Illustratively, the electronic device 400 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0150] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a non-transitory computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by the processor, the variant word restoration method of any of the live broadcast scenarios in the above-mentioned embodiments is implemented. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, etc.
[0151] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the method for restoring variant words in a live broadcast scenario. Corresponding to the execution subject corresponding to each step in each embodiment of the method for restoring variant words in a live broadcast scenario, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0152] It should be understood that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present application.
[0153] The functional blocks shown in the block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they may be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments may be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or communication link. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memory, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. Code segments may be downloaded via a computer network such as the Internet or an intranet.
[0154] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps. In other words, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0155] Aspects of the present application have been described above with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each block in the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine such that execution of these instructions by the processor of the computer or other programmable data processing device enables the implementation of the functions / actions specified in one or more blocks in the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It should also be understood that each block in the block diagrams and / or flowcharts, as well as combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware that performs the specified functions or actions, or by a combination of dedicated hardware and computer instructions.
[0156] The above description is only a specific embodiment of the present application. Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application.
Claims
1. A method for restoring variant words in a live broadcast scene, characterized in that: include: Receiving live variant word restoration requests respectively sent by multiple requesting parties, wherein the time interval between sending the multiple live variant word restoration requests satisfies a preset condition, and the live variant word restoration requests include live variant word texts to be restored; Calling an inference API to infer the live broadcast variant word text to be restored to obtain a live broadcast variant word restoration result. The inference API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live broadcast sample sentence pairs; Determining position information corresponding to the variant word restoration based on the live variant word restoration result and the live variant word text to be restored; Sending the live broadcast variant word restoration result and location information to the corresponding requesting party via the API; Before calling the inference interface API to infer the live broadcast variant word text to be restored to obtain the live broadcast variant word restoration result, the method further includes: The trained variant word restoration model is deployed with high concurrency based on dynamic batch inference, and then encapsulated to obtain the inference interface API, which specifically includes: Create a queue, each element in the queue is ,in To request the data you want to process, use To mark each request; Set the maximum batch size and maximum waiting time. If the number of requests in the queue reaches the maximum batch size, or the waiting time reaches the maximum waiting time, divide the batch and get , For the inference batches, including Requests, is the first Requests, Smaller than the maximum batch size; Segment the divided batches and feed them into the model in sequence Inference results are obtained , For the Inference results of inference batches; The inference results are Split the query and return the results to the corresponding requests by query ID.
2. The method according to claim 1, characterized in that Before performing high-concurrency deployment based on dynamic batch inference on the trained variant word restoration model and then encapsulating it to obtain an inference interface API, the method further includes: Acquire a plurality of live broadcast sample sentence pairs, wherein the live broadcast sample sentence pairs include two mutually corresponding live broadcast sample sentences, and the data type of the live broadcast sample sentences is text data; The live sample sentences are segmented by BPE to obtain the corresponding token sequence; Dividing the plurality of live sample sentence pairs according to a preset ratio to obtain a training set, a validation set, and a test set; Input the token sequence corresponding to the training set into the preset model for training, then input the token sequence corresponding to the validation set into the preset model for verification, and the token sequence corresponding to the test set for testing; Determine whether at least one of the iterative training times or the test loss value of the preset model meets the preset training stop condition; if not, adjust the weight parameters of the preset model until the preset training stop condition is met, thereby obtaining a variant word restoration model.
3. The method according to claim 2, characterized in that The obtaining of multiple live broadcast sample sentence pairs includes: Acquire multiple live video data and live text data corresponding to each of the live video data; Segmenting the live broadcast text data into a plurality of first live broadcast sample sentences according to semantics; Determining whether the first live broadcast sample sentence includes a first live broadcast variant word; In a case where the first live broadcast sample sentence does not include a first live broadcast variant word, determining the first live broadcast sample sentence as a live broadcast negative sample; In a case where the first live broadcast sample sentence includes a first live broadcast variant word, marking the first live broadcast variant word in the first live broadcast sample sentence and annotating the corresponding first live broadcast original word to obtain a live broadcast positive sample; The live broadcast positive sample and the live broadcast negative sample are determined as a live broadcast sample sentence pair.
4. The method according to claim 3, characterized in that In a case where the first live broadcast sample sentence includes a first live broadcast variant word, marking the first live broadcast variant word in the first live broadcast sample sentence and annotating the corresponding first live broadcast original word to obtain a live broadcast positive sample, the method further includes: Constructing a vocabulary of live broadcast variant words based on the first live broadcast variant words and the first live broadcast original words; Randomly extracting a plurality of second live broadcast variant words and second live broadcast original words matching the second live broadcast variant words from the live broadcast variant word library, wherein the second live broadcast variant word is any one of the plurality of first live broadcast variant words in the live broadcast variant word library; generating a plurality of second live broadcast sample sentences including the second live broadcast original words using the large language model; Replacing the second live broadcast original word in the second live broadcast sample sentence with a second live broadcast variant word that matches it, to obtain a third live broadcast sample sentence; The second live broadcast sample sentence and the third live broadcast sample sentence are determined as a live broadcast sample sentence pair.
5. The method according to claim 2, characterized in that The preset model includes an Encoder component and a Decoder component; The step of inputting the token sequence corresponding to the training set into the preset model for training includes: Performing token embedding and position embedding on the token sequence through the embedding layer of the preset model to obtain a corresponding dense vector; Outputting the context representation corresponding to the dense vector through the Encoder component; The context representation is used to guide the output of each time step of the Decoder component.
6. The method according to claim 2, characterized in that The test loss value is obtained based on the test set and the loss function; The loss function is: ; in, Indicates the batch size, T indicates the maximum length of the token sequence; represents the mask matrix; Represents the token predicted by the model's i-th sample at the t-th time step probability.
7. A device for restoring variant words in a live broadcast scene, characterized in that: The device comprises: a receiving module configured to receive live variant word restoration requests respectively sent by a plurality of requesting parties, wherein the time interval between sending the plurality of live variant word restoration requests satisfies a preset condition, and the live variant word restoration requests include live variant word texts to be restored; A calling module is configured to call an inference API to perform inference on the live broadcast variant word text to be restored, thereby obtaining a live broadcast variant word restoration result. The inference API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live broadcast sample sentence pairs. A determination module, configured to determine position information corresponding to the variant word restoration based on the live variant word restoration result and the live variant word text to be restored; A sending module, configured to send the live broadcast variant word restoration result and location information to a corresponding requesting party via an API; An inference module is used to perform high-concurrency deployment based on dynamic batch inference on the trained variant word restoration model before calling the inference interface API to infer the live variant word text to be restored and obtaining the live variant word restoration result, and then encapsulate the inference interface API; The reasoning module is specifically used to: Create a queue, each element in the queue is ,in To request the data you want to process, use To mark each request; Set the maximum batch size and maximum waiting time. If the number of requests in the queue reaches the maximum batch size, or the waiting time reaches the maximum waiting time, divide the batch and get , For the inference batches, including Requests, is the first Requests, Smaller than the maximum batch size; Segment the divided batches and feed them into the model in sequence Inference results are obtained , For the Inference results of inference batches; The inference results are Split the query and return the results to the corresponding requests by query ID.
8. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; when the processor calls the computer program instructions, it implements the variant word restoration method for the live broadcast scene as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer program instructions, which, when called by a processor, implement the method for restoring variant words in a live broadcast scenario as described in any one of claims 1-6.
10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the variant word restoration method for a live broadcast scene as described in any one of claims 1-6.
Citation Information
Patent Citations
Variant word processing method and device, electronic equipment and readable storage medium
CN116402053A