Live broadcast scene variant word restoration method and device, equipment, medium and product
By calling the high-concurrent deployment of inference interface API based on dynamic batch inference in the live broadcast scenario, the variant word restoration model is reasoned, and the problem of difficulty in restoring variant words in the live broadcast scenario is solved, and efficient and accurate variant word restoration processing is achieved.
Patent Information
- Application Number
- CN202510449369.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-04-11
AI Technical Summary
It is difficult to restore variant words in live broadcast scenarios and lacks robustness in practical applications.
By calling the highly concurrently deployed inference interface API based on dynamic batch inference, the pre-trained variant word restoration model is inferred, and batch live variant word restoration requests are processed to obtain variant word restoration results and location information.
It significantly improves the system's concurrency performance and resource utilization, improves the model's variant word restoration performance and robustness, and realizes efficient and accurate variant word restoration processing in live broadcast scenarios.
Smart Images

Figure CN119962533A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to a method, device, equipment, medium and product for restoring variant words in a live broadcast scenario. Background Art
[0002] In order to increase sales and attract customers during live broadcasts, many businesses use variant words to circumvent platform review and engage in false advertising. For example, a business used a sentence like "If you use our products, you don't need to go to a certain hospital or spend money to see a doctor" in its promotion, using variant words such as "a certain hospital" and "white coat" to replace "hospital" and "doctor", implying that its products have medicinal effects. These variant words are widely used in various scenarios, with flexible and changeable forms, which brings huge challenges to the standardized management of the industry. Therefore, restoring the variant words in the live broadcast scenario is crucial to protecting consumer rights and promoting the standardized development of the industry.
[0003] However, current research on variant words mainly focuses on social media and underground industries, and live broadcast scenarios, as an emerging scenario, have not received enough attention. Variant words in live broadcast scenarios are quite different from those in social media and underground industries, with different forms and topics. In addition, the location of variant words is currently unknown, which lacks robustness in practical applications. Summary of the invention
[0004] The embodiments of the present application provide a method, device, equipment, medium and product for restoring variant words in a live broadcast scenario, so as to at least solve the problem in the related art that variant words in a live broadcast scenario are difficult to restore and lack robustness in practical applications.
[0005] In a first aspect, an embodiment of the present application provides a method for restoring variant words in a live broadcast scenario, comprising: Receiving live variant word restoration requests respectively sent by multiple requesting parties, wherein the sending time interval of the multiple live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored; Calling an inference interface API to infer the live variant word text to be restored to obtain a live variant word restoration result, wherein the inference interface API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference, and the variant word restoration model is obtained by training based on multiple live sample sentence pairs; Determine the position information corresponding to the variant word restoration according to the live variant word restoration result and the live variant word text to be restored; The live broadcast variant word restoration result and location information are sent to the corresponding requesting party through the API.
[0006] In a second aspect, an embodiment of the present application provides a device for restoring variant words in a live broadcast scenario, the device comprising: A receiving module, configured to receive live variant word restoration requests respectively sent by a plurality of requesting parties, wherein the sending time interval of the plurality of live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored; A calling module is used to call an inference interface API to infer the live variant word text to be restored to obtain a live variant word restoration result. The inference interface API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live sample sentence pairs; A determination module, used for determining the position information corresponding to the variant word restoration according to the live variant word restoration result and the live variant word text to be restored; The sending module is used to send the live broadcast variant word restoration result and location information to the corresponding requesting party through the API.
[0007] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the processor implements the steps of the method for restoring variant words in a live broadcast scene as described in any one of the embodiments of the first aspect.
[0008] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the steps of the method for restoring variant words in a live broadcast scene as described in any one of the embodiments of the first aspect are implemented.
[0009] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of a method for restoring variant words in a live broadcast scenario as provided in the first aspect of an embodiment of the present application.
[0010] The variant word restoration method, device, equipment, medium and product for live broadcast scenarios in the embodiments of the present application are based on a variant word restoration model obtained by training multiple live broadcast sample sentences, and an inference interface API obtained by encapsulating the variant word restoration model after high-concurrency deployment based on dynamic batch reasoning. The inference interface API is called to process batches of live broadcast variant word restoration requests, and the live broadcast variant word restoration results and location information corresponding to the live broadcast variant word text to be restored in the live broadcast variant word restoration request are obtained, and the information is fed back to the request sender. It can be seen that the variant word model deployment based on dynamic batch reasoning significantly improves the system's concurrency performance and resource utilization, and improves the model's variant word restoration performance and robustness, because the inference batches are dynamic, thereby realizing efficient and accurate variant word restoration processing in live broadcast scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 It is a flowchart of a method for restoring variant words in a live broadcast scenario provided by an embodiment of the present application; Figure 2 It is a flowchart of a method for training and deploying a variant word restoration model provided in an embodiment of the present application; Figure 3 It is a structural schematic diagram of a variant word restoration device for a live broadcast scenario provided by an embodiment of the present application; Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application.
[0013] Reference numerals: A device 300 for restoring variant words in a live broadcast scene, a receiving module 301, a calling module 302, a determining module 303, a sending module 304, Electronic device 400 , processor 401 , memory 402 , communication interface 403 , bus 410 . DETAILED DESCRIPTION
[0014] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0015] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0016] In order to increase sales and attract customers during live broadcasts, many businesses use variant words to circumvent platform review and engage in false advertising. For example, a business used a sentence like "If you use our products, you don't need to go to a certain hospital or spend money to see a doctor" in its promotion, using variant words such as "a certain hospital" and "white coat" to replace "hospital" and "doctor", implying that its products have medicinal effects. These variant words are widely used in various scenarios, with flexible and changeable forms, which brings huge challenges to the standardized management of the industry. Therefore, restoring the variant words in the live broadcast scenario is crucial to protecting consumer rights and promoting the standardized development of the industry.
[0017] Current research on variant words mainly focuses on social media and the underground industry. As an emerging scenario, the live broadcast scenario has not received sufficient attention. The variant words in the live broadcast scenario are quite different from those in social media and the underground industry. Firstly, the forms of variant words are different: the scenarios mentioned above are all visual scenarios, so the variant words in this scenario mainly look similar to the original words visually. For example, "muqipai" and "2 xiaoshi" are used to represent "qipai" and "2 hours"; while the variant words in the live broadcast scenario mainly maintain the integrity of the rhythm by inserting some meaningless words. For example, "mouyi mouyuan" and "mian shenme yi" are used to represent "hospital" and "immunization". Secondly, the themes involved are different: the variant words in social media mainly involve current events and politics, the variant words in the underground industry mainly involve illegal industries, and the variant words in the live broadcast scenario mainly involve the health and medical industries.
[0018] Currently, for the research on restoring variant words in scenarios such as social media and the underground industry, pre-trained models such as BERT are applied to restore complex words. BERT uses the Transformer architecture, and the pre-training tasks include the Masked Language Model (MLM) task and the Next Sequence Prediction (NSP) task; when training the Masked Language Model MLM, some words are randomly masked for the model to predict the masked parts; the Next Sentence Prediction task NSP is to enable BERT to learn the relationships between sentences; therefore, when the masked part is a variant word, based on the pre-training task of the Masked Language Model, BERT will try to predict the alternative word for the masked part, that is, the original word of the variant word. However, this method has significant drawbacks. Firstly, since the positions of variant words are unknown, an additional judgment module needs to be trained to identify which words in the sentence are variant words, which greatly increases the complexity and computational cost of the system. Secondly, this paradigm is only applicable to the case where the length of the variant word is equal to that of the original word and cannot handle the case where the lengths of the variant word and the original word are different, resulting in a lack of robustness in practical applications.
[0019] To solve the problems of related technologies, the embodiments of this application provide a method, device, device, medium and product for restoring variant words in the live broadcast scenario.
[0020] Next, in conjunction with the accompanying drawings, through specific embodiments and their application scenarios, the method for restoring variant words in the live broadcast scenario provided by the embodiments of this application will be described in detail.
[0021] Figure 1 The flowchart of the method 100 for restoring variant words in the live broadcast scenario according to the embodiments of this application is shown. As Figure 1 shown, the method 100 for restoring variant words in the live broadcast scenario may specifically include the following steps: S101, receiving live variant word restoration requests respectively sent by multiple requesting parties, wherein the sending time interval of the multiple live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored; S102, calling the inference interface API to infer the live variant word text to be restored, and obtaining the live variant word restoration result, wherein the inference interface API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference, and the variant word restoration model is obtained by training based on multiple live sample sentence pairs; S103, determining the position information corresponding to the variant word restoration according to the live variant word restoration result and the live variant word text to be restored; S104, sending the live broadcast variant word restoration result and location information to the corresponding requesting party through an API.
[0022] Therefore, based on the variant word restoration model trained on multiple live sample sentences and the inference interface API encapsulated after high-concurrency deployment of the variant word restoration model based on dynamic batch reasoning, the inference interface API is called to process batches of live variant word restoration requests, and the live variant word restoration results and position information corresponding to the live variant word text to be restored in the live variant word restoration request are obtained, and the information is fed back to the request sender. It can be seen that the variant word model deployment based on dynamic batch reasoning, because the inference batch is dynamic, significantly improves the system's concurrency performance and resource utilization, and improves the model's variant word restoration performance and robustness, and realizes efficient and accurate variant word restoration processing in live broadcast scenarios.
[0023] The specific implementation methods of the above steps are introduced below.
[0024] In some embodiments, in S101, the time interval for sending multiple live variant word restoration requests meets a preset condition, and the preset condition includes: multiple requesting parties corresponding to these requests send the requests simultaneously, or send them at an extremely short time interval (for example, 0.1 seconds).
[0025] It should be noted that the live variant word text to be restored in the request body is obtained by performing speech recognition on the acquired real-time live video. Specifically, the live content on the live broadcast platform is crawled to obtain multiple real-time live video data, and then the real-time live video data is recognized using a speech recognition model to obtain the text data corresponding to the real-time live video data.
[0026] In some embodiments, before S102, a variant word restoration model is obtained by training multiple live sample sentences, and then the variant word restoration model is deployed with high concurrency based on dynamic batch reasoning and encapsulated to obtain the reasoning interface API. Specifically, the model training and deployment process provided in the embodiment of the present application can refer to Figure 2 .
[0027] like Figure 2 As shown, the training and deployment method 200 of the variant word restoration model provided in the embodiment of the present application may include the following steps: S201 to S206.
[0028] S201. Acquire a plurality of live broadcast sample sentence pairs, wherein the live broadcast sample sentence pairs include two corresponding live broadcast sample sentences, and the data type of the live broadcast sample sentences is text data.
[0029] In some embodiments, a plurality of live video data and live text data corresponding to each of the live video data are obtained; the live text data are semantically segmented into a plurality of first live sample sentences; it is determined whether the first live sample sentence includes a first live variant word; if the first live sample sentence does not include the first live variant word, the first live sample sentence is determined as a live negative sample; if the first live sample sentence includes the first live variant word, the first live variant word in the first live sample sentence is marked and the corresponding first live original word is annotated to obtain a live positive sample; the live positive sample and the live negative sample are determined as a live sample sentence pair.
[0030] In the specific implementation, the crawler tool is used to crawl the live content on the live platform and obtain live video clips. After the crawler tool is running, it will cyclically read the live link pool in the configuration file, open a thread for each live link, and download the live video content in parallel. The live link pool stores the live link currently being crawled, which contains two fields: live name and live link, both of which are text type. In order to ensure the diversity and coverage of the data, different time periods, different anchors and different types of health products are selected for data crawling. A total of multiple (for example, 6273) live clips are crawled, and each live clip is set to 60 seconds. This length can not only ensure that enough information is captured, but also reduce the complexity of data processing to a certain extent.
[0031] In the specific implementation, automatic speech recognition technology is used to transcribe the live video clips into text data and screen them. Optionally, an end-to-end speech recognition model is built and trained with a manually annotated Mandarin speech recognition dataset. In addition, after transcribing the video into text information, useless data with only background music and other noise is manually screened to remove, and finally multiple (for example, 6223) high-quality live text data are obtained.
[0032] In specific implementation, the sentence segmentation technology is used to segment the live text segments obtained after transcription into sentences according to semantics, and then after the text segments are cut into sentences, short sentences or repeated sentences with no real meaning are screened out to improve the accuracy and efficiency of subsequent analysis. That is, multiple first live sample sentences are obtained.
[0033] In the specific implementation, a manual annotation method is adopted to first check multiple first live sample sentences to determine whether they contain variant words. If variant words are contained, they are marked and the corresponding original words are annotated for them. In other words, a manual annotation method is adopted to provide the annotators with ASR transcription texts, and combine them with the corresponding live clips as auxiliary information. In the annotation process, the annotators need to mark the variant words in the sentences and add the corresponding original words as annotations. In the end, 8236 positive samples containing variant words and 78554 negative samples that do not contain variant words are obtained. In this way, through this annotation process, the live sample sentence pairs required for subsequent model training and the variant word thesaurus required for subsequent data enhancement can be generated.
[0034] Further, in some embodiments, when the first live broadcast sample sentence includes a first live broadcast variant word, the first live broadcast variant word in the first live broadcast sample sentence is marked and the corresponding first live broadcast original word is annotated. After obtaining a live broadcast positive sample, a live broadcast variant word library is constructed based on the first live broadcast variant word and the first live broadcast original word; a number of second live broadcast variant words and second live broadcast original words matching the second live broadcast variant word are randomly selected from the live broadcast variant word library, and the second live broadcast variant word is any one of the multiple first live broadcast variant words in the live broadcast variant word library; a plurality of second live broadcast sample sentences containing the second live broadcast original word are generated using a large language model; the second live broadcast original word in the second live broadcast sample sentence is replaced with a second live broadcast variant word matching it to obtain a third live broadcast sample sentence; and the second live broadcast sample sentence and the third live broadcast sample sentence are determined as a live broadcast sample sentence pair.
[0035] That is, a variant word dictionary is constructed based on the annotation results, for example, containing 431 original words and their 2,688 variant forms. Next, a number of variant words and their corresponding original words are randomly selected from the variant word dictionary to generate a fluent text that fits the topic and contains all the original words. Subsequently, the original words in the text generated by the large language model are replaced with the corresponding variant words to form sentence pairs.
[0036] In the specific implementation, several variant words are randomly selected from the variant word dictionary, and the original words corresponding to the variant words are obtained based on the dictionary. A large language model is used to generate multiple sentences containing the original words and matching the topic. Then the original words are replaced with the corresponding variant words. In this way, a set of sentence pairs containing source text and target text is constructed, where the target text refers to the text after the variant words are replaced with the corresponding original words based on the source text. In this way, 11,280 positive samples and 2,155 negative samples containing variant words are constructed.
[0037] Taking the example of using a large language model to generate five sentences that contain the original word and are consistent with the topic, the large language model prompt template can be as follows: "Our role is a live broadcast anchor who is responsible for promoting products. You need to generate five promotional sentences containing the target word. Here are some real promotional sentences for you to imitate. The generated sentences should not have repeated meanings. The target word should remain unchanged. The length of the sentence should be as consistent as possible with the provided example." In this way, a high-quality dataset of variant words in live broadcast scenarios was constructed through manual annotation and large-model data enhancement, solving the problem of scarcity of datasets in this scenario. That is, in the subsequent model training process, the manually annotated training set and the data enhancement dataset can be mixed as the training set, and the live broadcast dataset contains multiple live video text sentence pairs, each of which contains the source text and the target text.
[0038] Furthermore, in some embodiments, data preprocessing is performed on the live sample sentence pairs: the source sentences input into the model are defined ,in is the sentence length, For the first The maximum length of the sentence is 128 characters; if the sentence length exceeds 128, the sentence is divided into clauses in order, and each clause does not exceed the specified maximum length.
[0039] S202. Segment the live sample sentences through BPE to obtain corresponding token sequences.
[0040] In some embodiments, the source sentence input to the model is passed through the T5Tokenizer Through BPE word segmentation, the unit is token, and each token has a separate ID. Get the sequence ,in is the length of the token sequence, is the i-th token in the token sequence. It should be noted that the target sentence in the live sample sentence pair is also subjected to the same preprocessing and word segmentation operations as the source sentence. In this way, the input text is tokenized by the pre-trained language model so that the text data is converted into a form that the model can understand.
[0041] S203, dividing the plurality of live sample sentence pairs according to a preset ratio to obtain a training set, a validation set and a test set.
[0042] That is, according to the live broadcast variant word dataset, a training set, a validation set, and a test set are constructed in a ratio of 8:1:1, for example, and during the training process, a mixture of the manually annotated training set and the data augmentation dataset is used as the training set.
[0043] S204, input the token sequence corresponding to the training set into the preset model for training, then input the token sequence corresponding to the validation set into the preset model for verification, and the token sequence corresponding to the test set for testing.
[0044] In some embodiments, the preset model includes an Encoder component and a Decoder component. Then, the token sequence is token-embedded and position-embedded through the embedding layer of the preset model to obtain a corresponding dense vector; the context representation corresponding to the dense vector is output through the Encoder component; and the context representation is used to guide the output of each time step of the Decoder component.
[0045] In specific implementation, the token sequence of the source sentence Input to the end-to-end model mengzi-T5-base model, first it will pass through the embedding layer to get token embedding and position embedding A dense vector of .
[0046] Among them, mengzi-T5-base is a training paradigm using the T5 model, which is retrained on a large amount of Chinese corpus. It contains Encoder and Decoder components, each of which consists of 12 Transformer layers, each with 12 attention heads to capture the dependencies between the input sequence and the output sequence. With this design, Mengzi-T5-base can perform well in Chinese natural language processing tasks, especially in generation tasks with high flexibility and accuracy.
[0047] In the specific implementation, the dense vector passes through the Encoder component of the model, which consists of 12 layers of multi-head Attention layers, to obtain the contextual representation of the target sentence .
[0048] In the specific implementation, the input start symbol is input in the first time step of the Decoder stage. As the input of the first time step, it passes through the embedding layer and the Decoder component. The Decoder component consists of 12 layers of multi-head Attention layers. In the process of Attention calculation, the cross attention mechanism is used to calculate the Decoder input and , capturing the contextual information of the sequence. At each time step of the model, the output passes through a linear layer to map the Decoder output to the dimension of the target vocabulary. Subsequently, the SoftMax function is applied to the output of the linear layer to convert these values into a probability distribution. The output is recorded as ,in is the output of the first time step and serves as the input of the second time step. Until the last time step outputs the terminator The sentence probability output by the model is: = ; (1) S205. Determine whether at least one of the iterative training times or the test loss value of the preset model meets the preset training stop condition. If not, adjust the weight parameters of the preset model until the preset training stop condition is met to obtain a variant word restoration model.
[0049] Optionally, the preset training stop condition includes at least one of the following: the number of iterative training reaches a first preset threshold, and the number of times the test loss value remains continuously without decreasing reaches a second preset threshold.
[0050] The test loss value is obtained based on the test set and the loss function.
[0051] The loss function is: ; (2) in, represents the batch size, T represents the maximum length of the token sequence (target sequence); Represents the mask matrix, ignoring the calculation loss of padding tokens; Represents the token predicted by the model's i-th sample at the t-th time step probability.
[0052] Optionally, in this example the initial learning rate is set to , using Adam optimizer, epoch is set to 20, Set to 32 and warm_steps to 200.
[0053] In this way, by training the model on the training set, calculating the model loss, performing back propagation, updating the model weight parameters, and performing iterative operations, the variant word restoration model M can be obtained by training the variant word dataset.
[0054] S206: Perform high-concurrency deployment of the variant word restoration model based on dynamic batch reasoning, and then encapsulate it to obtain a reasoning interface API.
[0055] When implementing, create a global queue Collect data from multiple requests, each request will be pushed into a queue and wait to be processed. Each element in the queue is ,in To request the data you want to process, use to mark each request.
[0056] When implementing, set the maximum batch size and the maximum waiting time The maximum waiting time is 40ms. Before the maximum waiting time is reached, the request data will be collected as much as possible and divided into batches according to the maximum batch size. If the number of requests in the queue reaches the maximum batch size, or the waiting time reaches the maximum waiting time, the batch is divided. . For the inference batches, including Requests, is the first Requests, Less than the maximum batch size.
[0057] In the specific implementation, the divided batches are segmented and sent to the model in order. In the inference result , defined as follows: ; (3) in, For the Inference batches, For the Inference results of inference batches.
[0058] In specific implementation, the inference results Contains The request output inference results and corresponding ,according to The results are split and returned to the corresponding requests. That is, an ID is assigned to each request so that the results are returned to the corresponding requests by querying the ID.
[0059] In this way, the variant word model deployment method based on dynamic batch reasoning can significantly improve the system's concurrency performance and resource utilization by dynamically adjusting the batch size and optimizing the reasoning process.
[0060] The above is a specific implementation method for training a variant word restoration model provided in the embodiment of the present application and encapsulating the variant word restoration model to obtain an inference interface API after high-concurrency deployment based on dynamic batch reasoning. Using an end-to-end model, the variant word detection and correction tasks are unified into a generation task based on the seq2seq paradigm, which improves the variant word restoration performance and robustness of the model.
[0061] Back to Figure 1 The variant word restoration method 100 for the live broadcast scenario shown performs variant word restoration on actual business text.
[0062] For S102 to S104, the specific implementation method may be as follows: Towards Send Request The request body contains the variant word text that needs to be restored , the model gets the input , defined as follows: ; (4) in, Data required for model reasoning, including text And other necessary preprocessing information.
[0063] Next, the model Data sent upon request Make inferences, and the results of inferences pass Returns, defined as follows: ; (5) in, Contains the results of restoring variant words in the text S, as well as information such as the position of the variant words in the sentence.
[0064] As an optional embodiment, the difference between the input and output sequences can be determined by character string comparison, that is, the corresponding token. In this way, by comparing the difference between the input and output sequences, the word position can be recorded and the content can be corrected.
[0065] In addition, in order to verify the performance of the variant word restoration method for the live broadcast scenario of the embodiment of the present application, the superiority of the present method is illustrated below through automatic indicator evaluation.
[0066] Specifically, the evaluation is performed using test sets from different sources. Test set 1 is the data from the same live link pool as the training set. Its data is not repeated in the training set or the validation set. It contains a total of 1,600 samples, which are manually annotated, including 400 positive samples and 400 negative samples; and each positive sample contains the manually annotated result as a Reference. Test set 2 is a new live link pool, which has no overlap with the live link pool in the training set. It contains a total of 800 samples, including 400 positive samples and 400 negative samples. Each positive sample contains the manually annotated result as a Reference.
[0067] This method uses The value is evaluated. It is defined as follows: ; (6) ; (7) ; (8) ; (9) in, To correctly restore the positive subspecies of all variant words; is a counterexample that has not been corrected; There are positive examples where variant words have not been restored; is a counterexample of incorrect restoration.
[0068] Compare the method of the embodiment of the present application with other methods: (Large Language Model), (N-gram language model), (Showing sequence-to-sequence modeling editing operations), (Sequence-to-Sequence Models for Convolutional Networks), (Pre-trained language model with self-noise removal encoding) The following Tables 1 and 2 show the comparative experimental results of test set 1 and test set 2 respectively.
[0069] Table 1 method Acc Precision Recall F1 LLM 0.605 0.660 0.529 0.587 Kenlm 0.583 0.607 0.372 0.537 Seq2Edit 0.651 0.968 0.361 0.526 Convseq2seq 0.740 0.978 0.527 0.685 BART 0.708 0.701 0.767 0.738 This embodiment 0.928 0.937 0.927 0.932 Table 2 method Acc Precision Recall F1 LLM 0.373 0.450 0.459 0.346 Kenlm 0.516 0.515 0.513 0.514 Seq2seq 0.702 0.987 0.408 0.588 Convseq2seq 0.687 0.898 0.421 0.573 BART 0.656 0.670 0.611 0.639 This embodiment 0.863 0.929 0.787 0.852 As shown in Table 1 and Table 2, the method of this embodiment achieves the highest ; The experimental results presented show that character-level error correction methods such as Seq2Edit and the statistical language model Kenlm are insufficient in handling morphological changes in live broadcast scenarios. In contrast, Seq2seq models such as Convseq2seq, BART, and T5 perform better in managing output length inconsistencies. It is worth noting that the T5 model in this example achieved the highest F1 score on both test sets, demonstrating its effectiveness in this task.
[0070] Furthermore, Table 3 below provides a method for efficient deployment of the model. In order to verify its beneficial indicators, the efficiency of the method of the present application is illustrated through performance testing.
[0071] Table 3. Concurrency test comparison
[0072] As shown in Table 3, a stress test was performed using a tool for API testing and performance evaluation, with the number of concurrent connections set to 10 and the number of rounds set to 5. The video memory usage is the maximum video memory usage recorded during the stress test, in MB; the time consumed is the time consumed by the test, in seconds (s); the throughput is the number of requests processed per second, in R / s; and the error rate is the percentage of requests with errors among the requests sent. The dynamic reasoning-based method will take up additional video memory compared to the thread pool because the inference batches are dynamic, but it significantly saves inference time. Specifically, dynamic reasoning increases the system's concurrency by 789%, greatly optimizing the system's concurrency capabilities. The dynamic reasoning method is significantly better than the thread pool in performance, especially in terms of throughput and time consumption. Although the video memory usage is high, its improvement in the system's concurrency capabilities is significant.
[0073] It can be seen that the embodiment of the present application has important significance in the technical field of variant word restoration in live broadcast scenarios by constructing a high-quality variant word dataset and a variant word restoration method based on an end-to-end architecture, and enriches the technical system in this field. The method not only shows excellent performance, but also shows strong generalization ability, and can effectively cope with complex and changeable variant word scenarios.
[0074] In addition, the high-performance concurrent deployment algorithm for the variant word restoration model in the embodiment of the present application is based on a dynamic batch processing deployment strategy, which greatly improves the system's concurrency capability, real-time processing efficiency and resource utilization.
[0075] The embodiments of the present application overcome the problems of difficulty in detecting variant words, low efficiency, lack of training data, etc. in current live broadcast scenarios. Compared with the traditional rule-based variant word restoration method, the robustness and accuracy of the restoration are significantly improved.
[0076] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0077] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides a variant word restoration device 300 for a live broadcast scenario.
[0078] like Figure 3 As shown, the variant word restoration device 300 of the live broadcast scene may include: The receiving module 301 is used to receive live variant word restoration requests respectively sent by multiple requesting parties, wherein the sending time interval of the multiple live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored; A calling module 302 is used to call an inference interface API to infer the live variant word text to be restored to obtain a live variant word restoration result. The inference interface API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live sample sentence pairs; A determination module 303 is used to determine the position information corresponding to the variant word restoration according to the live variant word restoration result and the live variant word text to be restored; The sending module 304 is used to send the live broadcast variant word restoration result and location information to the corresponding requesting party through the API.
[0079] In some embodiments, the variant word restoration device 300 for live broadcast scenarios further includes a training deployment module ( Figure 3 Specifically, the training deployment module includes the following units: An acquisition unit, used for acquiring a plurality of live broadcast sample sentence pairs, wherein the live broadcast sample sentence pairs include two mutually corresponding live broadcast sample sentences, and the data type of the live broadcast sample sentences is text data; A word segmentation unit, used for segmenting the live sample sentences through BPE to obtain a corresponding token sequence; A division unit, used for dividing the plurality of live sample sentence pairs according to a preset ratio to obtain a training set, a validation set and a test set; An input unit, used to input the token sequence corresponding to the training set into a preset model for training, and then input the token sequence corresponding to the validation set into the preset model for verification, and the token sequence corresponding to the test set for testing; An adjustment unit, used to determine whether at least one of the number of iterative training times or the test loss value of the preset model meets a preset training stop condition, and if not, adjust a weight parameter of the preset model until the preset training stop condition is met, thereby obtaining a variant word restoration model; The encapsulation unit is used to perform high-concurrency deployment of the variant word restoration model based on dynamic batch reasoning, and then encapsulate it to obtain an inference interface API.
[0080] In some optional embodiments, the acquisition unit is specifically used to: acquire multiple live video data and live text data corresponding to each of the live video data; segment the live text data into multiple first live sample sentences according to semantics; determine whether the first live sample sentence includes a first live variant word; if the first live sample sentence does not include the first live variant word, determine the first live sample sentence as a live negative sample; if the first live sample sentence includes the first live variant word, mark the first live variant word in the first live sample sentence and annotate the corresponding first live original word to obtain a live positive sample; determine the live positive sample and the live negative sample as a live sample sentence pair.
[0081] In some optional embodiments, the acquisition unit is further specifically used to: construct a live broadcast variant word library based on the first live broadcast variant word and the first live broadcast original word; randomly extract a number of second live broadcast variant words and second live broadcast original words matching the second live broadcast variant word from the live broadcast variant word library, the second live broadcast variant word being any one of the multiple first live broadcast variant words in the live broadcast variant word library; generate a plurality of second live broadcast sample sentences containing the second live broadcast original word by using a large language model; replace the second live broadcast original word in the second live broadcast sample sentence with a second live broadcast variant word matching it to obtain a third live broadcast sample sentence; determine the second live broadcast sample sentence and the third live broadcast sample sentence as a live broadcast sample sentence pair.
[0082] Optionally, the preset model includes an Encoder component and a Decoder component.
[0083] In some optional embodiments, the input unit is specifically used to: perform token embedding and position embedding on the token sequence through the embedding layer of the preset model to obtain a corresponding dense vector; output the context representation corresponding to the dense vector through the Encoder component; and use the context representation to guide the output of each time step of the Decoder component.
[0084] Optionally, the test loss value is obtained based on the test set and the loss function; The loss function is: ; in, represents the batch size, and T represents the maximum length of the token sequence; represents the mask matrix; Represents the token predicted by the model's i-th sample at the t-th time step probability.
[0085] It should be noted that, for the convenience of description, the above device is described in various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0086] The device of the above embodiment is used to implement the variant word restoration method of the corresponding live broadcast scenario in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0087] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an electronic device.
[0088] Figure 4 A more specific schematic diagram of the hardware structure of an electronic device provided by this embodiment is shown.
[0089] The electronic device 400 may include a processor 401 and a memory 402 storing computer program instructions.
[0090] Specifically, the processor 401 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0091] The memory 402 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 402 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 402 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 402 is a non-volatile solid-state memory.
[0092] In certain embodiments, the memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present application.
[0093] The processor 401 reads and executes the computer program instructions stored in the memory 402 to implement the variant word restoration method for any live broadcast scenario in the above embodiments.
[0094] In some examples, the electronic device 400 may further include a communication interface 403 and a bus 410. Figure 4 As shown, the processor 401, the memory 402, and the communication interface 403 are connected via a bus 410 and communicate with each other.
[0095] The communication interface 403 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0096] Bus 410 includes hardware, software or both, and couples the components of online data traffic billing equipment to each other. For example, but not limitation, bus 410 may include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front-side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnect (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. Where appropriate, bus 410 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the present application considers any suitable bus or interconnection.
[0097] Exemplarily, the electronic device 400 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0098] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a non-transitory computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by the processor, the variant word restoration method of any of the above-mentioned live broadcast scenarios is implemented. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, etc.
[0099] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the variant word restoration method for the live broadcast scene. Corresponding to the execution subject corresponding to each step in each embodiment of the variant word restoration method for the live broadcast scene, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0100] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0101] The functional blocks shown in the structural block diagram described above can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0102] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0103] The above describes various aspects of the present application with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0104] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A method for restoring variant words in a live broadcast scene, characterized in that: include: Receiving live variant word restoration requests respectively sent by multiple requesting parties, wherein the sending time interval of the multiple live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored; Calling an inference interface API to infer the live variant word text to be restored to obtain a live variant word restoration result, wherein the inference interface API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference, and the variant word restoration model is obtained by training based on multiple live sample sentence pairs; Determine the position information corresponding to the variant word restoration according to the live variant word restoration result and the live variant word text to be restored; The live broadcast variant word restoration result and location information are sent to the corresponding requesting party through the API.
2. The method according to claim 1, characterized in that: Before calling the inference interface API to infer the live variant word text to be restored to obtain the live variant word restoration result, the method further includes: Acquire a plurality of live broadcast sample sentence pairs, wherein the live broadcast sample sentence pairs include two mutually corresponding live broadcast sample sentences, and the data type of the live broadcast sample sentences is text data; The live sample sentences are segmented by BPE to obtain corresponding token sequences; Dividing the plurality of live sample sentence pairs according to a preset ratio to obtain a training set, a validation set, and a test set; Input the token sequence corresponding to the training set into the preset model for training, then input the token sequence corresponding to the validation set into the preset model for verification, and the token sequence corresponding to the test set for testing; Determine whether at least one of the number of iterations of training or the test loss value of the preset model meets a preset training stop condition; if not, adjust the weight parameter of the preset model until the preset training stop condition is met, thereby obtaining a variant word restoration model; The variant word restoration model is deployed with high concurrency based on dynamic batch reasoning, and then encapsulated to obtain a reasoning interface API.
3. The method according to claim 2, characterized in that The obtaining of a plurality of live broadcast sample sentence pairs comprises: Acquire multiple live video data and live text data corresponding to each of the live video data; Segmenting the live broadcast text data into a plurality of first live broadcast sample sentences according to semantics; Determining whether the first live broadcast sample sentence includes a first live broadcast variant word; In a case where the first live broadcast sample sentence does not include a first live broadcast variant word, determining the first live broadcast sample sentence as a live broadcast negative sample; In the case where the first live broadcast sample sentence includes a first live broadcast variant word, marking the first live broadcast variant word in the first live broadcast sample sentence and annotating the corresponding first live broadcast original word to obtain a live broadcast positive sample; The live broadcast positive sample and the live broadcast negative sample are determined as a live broadcast sample sentence pair.
4. The method according to claim 3, characterized in that: In the case where the first live broadcast sample sentence includes a first live broadcast variant word, marking the first live broadcast variant word in the first live broadcast sample sentence and annotating the corresponding first live broadcast original word, and after obtaining a live broadcast positive sample, the method further includes: Constructing a vocabulary of live broadcast variant words according to the first live broadcast variant words and the first live broadcast original words; Randomly extracting a plurality of second live broadcast variant words and second live broadcast original words matching the second live broadcast variant words from the live broadcast variant word library, wherein the second live broadcast variant word is any one of the plurality of first live broadcast variant words in the live broadcast variant word library; Generate a plurality of second live sample sentences including the second live original words by using the large language model; The second live broadcast original word in the second live broadcast sample sentence is replaced with a second live broadcast variant word that matches it, to obtain a third live broadcast sample sentence; The second live broadcast sample sentence and the third live broadcast sample sentence are determined as a live broadcast sample sentence pair.
5. The method according to claim 2, characterized in that: The preset model includes an Encoder component and a Decoder component; The step of inputting the token sequence corresponding to the training set into the preset model for training includes: Performing token embedding and position embedding on the token sequence through the embedding layer of the preset model to obtain a corresponding dense vector; Outputting the context representation corresponding to the dense vector through the Encoder component; The context representation is used to guide the output of the Decoder component at each time step.
6. The method according to claim 2, characterized in that The test loss value is obtained based on the test set and the loss function; The loss function is: ; in, represents the batch size, and T represents the maximum length of the token sequence; represents the mask matrix; Represents the token predicted by the model's i-th sample at the t-th time step probability.
7. A device for restoring variant words in a live broadcast scene, characterized in that: The device comprises: A receiving module, configured to receive live variant word restoration requests respectively sent by a plurality of requesting parties, wherein the sending time interval of the plurality of live variant word restoration requests meets a preset condition, and the live variant word restoration requests include live variant word texts to be restored; A calling module is used to call an inference interface API to infer the live variant word text to be restored to obtain a live variant word restoration result. The inference interface API is obtained by encapsulating a pre-trained variant word restoration model after high-concurrency deployment based on dynamic batch inference. The variant word restoration model is trained based on multiple live sample sentence pairs; A determination module, used for determining the position information corresponding to the variant word restoration according to the live variant word restoration result and the live variant word text to be restored; The sending module is used to send the live broadcast variant word restoration result and location information to the corresponding requesting party through the API.
8. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; when the processor calls the computer program instructions, it implements the variant word restoration method for the live broadcast scene as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are called by the processor, the method for restoring variant words in a live broadcast scene as described in any one of claims 1-6 is implemented.
10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the method for restoring variant words in a live broadcast scene as described in any one of claims 1-6.
Citation Information
Patent Citations
Short text auditing method and device fusing variant word recognition
CN112287684A
Semantic recognition model training method and device based on comparative learning and medium
CN114722834A
Variant word processing method and device, electronic equipment and readable storage medium
CN116402053A
Generative large model-based malicious short message variant word restoration method
CN117010328A
Live broadcast illegal behavior detection method, device and equipment based on variant word recognition
CN118967153A