Live broadcast violation detection method, device and equipment based on variant word recognition

Through the multi-level variant word recognition method, combined with regular matching, statistical language model and large language model, the problem of insufficient real-time and accuracy of variant word recognition in the live broadcast platform is solved, efficient and accurate detection of violations is achieved, and the standardization of the live broadcast market and consumer interests are guaranteed.

CN118967153BActive Publication Date: 2025-08-19CHINA ACADEMY OF ELECTRONICS AND INFORMATION TECHNOLOGY OF CHINA ELECTRONICS TECHNOLOGY GROUP CORPORATION +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410979366.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-22
Publication Date
2025-08-19
Estimated Expiration
2044-07-22

AI Technical Summary

Technical Problem

When detecting variant words, the live broadcast platform has problems of insufficient real-timeness and accuracy, which makes it difficult to effectively identify and prevent violations.

Method used

A multi-level variant word recognition method is adopted, combining regular matching, statistical language model and large language model, text data is obtained through speech recognition and optical character recognition, and multi-level variant word recognition is carried out, combined with time priority hierarchical filtering strategy, sensitive vocabulary is identified and matched, and evidence of violation is preserved.

Benefits of technology

It realizes efficient and accurate identification of variant words in live broadcast scenarios, ensures a balance between real-time and accuracy, protects the order of live broadcast market and consumer interests, and prevents the dissemination of bad information and commodities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118967153B_ABST
    Figure CN118967153B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for detecting live broadcast violations based on variant word recognition, including: based on a speech recognition model and an optical character recognition model, obtaining the audio and visual text of the live broadcast room and converting them into text data; extracting text data and performing multi-level variant word recognition, including: variant word recognition based on regular matching, variant word recognition based on a statistical language model, and variant word recognition based on a large language model; based on the recognized variant word, obtaining the original word of the variant word, and matching the original word with a sensitive word library to determine whether the original word exists; if the original word exists, retrieving the video data of a set length before and after the variant word, and saving it as violation evidence. This application adopts different recognition and detection methods to deal with different types of variant words, and adopts variant word recognition methods of different degrees of sophistication at different time granularities, achieving a balance between real-time and accuracy in live broadcast violation detection technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of language data processing technology, and in particular to a method, device, equipment and storage medium for detecting live broadcast violations based on variant word recognition. Background Art

[0002] All major live streaming platforms have taken corresponding measures to detect violations, effectively reducing violations in live streaming e-commerce.

[0003] However, in order to circumvent the live broadcast platform's censorship of banned words, some live broadcast rooms have adopted a new language strategy: using variant words, that is, some banned words in live broadcasts are transformed into non-standard words that are auditorily or visually similar to the original words, thereby bypassing the automatic detection mechanism of most live broadcast platforms. Summary of the Invention

[0004] The present invention provides a live broadcast violation detection method, device, equipment and storage medium based on variant word recognition, which solves the problem of insufficient real-time and accuracy of variant word recognition in live broadcast violation detection technology.

[0005] To achieve the above objectives, this application adopts the following technical solutions:

[0006] In a first aspect, a method for detecting live broadcast violations based on variant word recognition is provided, comprising:

[0007] Based on the speech recognition model and optical character recognition model, the audio and visual text of the live broadcast room are obtained and converted into text data;

[0008] Extracting the text data and performing multi-level variant word recognition;

[0009] Based on the identified variant words, the original word of the variant word is obtained, and the original word is matched with a pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library;

[0010] If the original word exists in the sensitive word library, retrieve the video data of the set length before and after the variant word and save it as violation evidence;

[0011] The multi-level variant word recognition includes: variant word recognition based on regular matching, variant word recognition based on statistical language model and variant word recognition based on large language model.

[0012] In a first possible implementation of the first aspect, the variant word is used to circumvent automatic review, including:

[0013] Structural variants, which include variants that change the physical structure of the original word while maintaining its visual or auditory similarity;

[0014] Variant words, including variant words that replace characters with similar pronunciations or written forms;

[0015] Semantic variant words, including variant words with different structures but similar semantics or that can refer to other words;

[0016] The variant word recognition based on regular expression matching is used to recognize the structural variant words;

[0017] The variant word recognition based on the statistical language model is used to identify the phonetic and morphological variant words;

[0018] The variant word recognition based on the large language model is used to recognize the semantic variant words.

[0019] Based on the first possible implementation of the first aspect, in a second possible implementation of the first aspect, the performing of multi-level variant word identification includes:

[0020] Based on text data, structural variant words are identified through preset regular expressions;

[0021] Analyze the language characteristics of text data based on statistical language models, correct spelling errors and identify phonetic and morphological variants;

[0022] Perform semantic understanding and context analysis on text data based on large models to identify semantic variant words.

[0023] Based on any possible implementation of the first aspect, in a third possible implementation of the first aspect, if the original word exists in the sensitive word library, the following steps are further performed:

[0024] The original word of the identified variant word is compared with the sensitive word library. If the original word exists in the sensitive word library, the variant word is stored in the variant word library.

[0025] Based on the first possible implementation of the first aspect, in a fourth possible implementation of the first aspect, the multi-level variant word recognition is configured with a time priority hierarchical filtering strategy, namely:

[0026] In the real-time stream of live video content, the time priority of variant word recognition based on regular matching is greater than the time priority of variant word recognition based on the statistical language model, and the time priority of variant word recognition based on the statistical language model is greater than the time priority of variant word recognition based on the large language model.

[0027] Based on the fourth possible implementation manner of the first aspect, in a fifth possible implementation manner of the first aspect, after the multi-level variant word recognition is started, the variant word recognition based on regular expression matching is fed back;

[0028] After each set first time period, the variant word recognition based on the statistical language model recognizes and feeds back the data of the previous first time period;

[0029] After each set second period, the variant word recognition based on the large language model recognizes and feeds back the data of the previous second period;

[0030] The duration of the first time period is shorter than the duration of the second time period.

[0031] In a second aspect, a live broadcast violation detection device based on variant word recognition is provided, comprising: a text data recognition and conversion module for acquiring audio and visual text in the live broadcast room based on a speech recognition model and an optical character recognition model, and converting the data into text data;

[0032] A multi-level variant word recognition module, configured to extract the text data and perform multi-level variant word recognition;

[0033] A variant word and sensitive word matching module is used to obtain the original word of the variant word based on the identified variant word, and match the original word with a pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library;

[0034] an evidence-building module, configured to retrieve video data of a set length of time before and after the variant word if the original word exists in the sensitive word library, and save it as evidence of violation;

[0035] The multi-level variant word recognition includes: variant word recognition based on regular matching, variant word recognition based on statistical language model and variant word recognition based on large language model.

[0036] In a first possible implementation of the second aspect, the variant word is used to circumvent automatic review, including:

[0037] Structural variants, which include variants that change the physical structure of the original word while maintaining its visual or auditory similarity;

[0038] Variant words, including variant words that replace characters with similar pronunciations or written forms;

[0039] Semantic variant words, including variant words with different structures but similar semantics or that can refer to other words;

[0040] The variant word recognition based on regular expression matching is used to recognize the structural variant words;

[0041] The variant word recognition based on the statistical language model is used to identify the phonetic and morphological variant words;

[0042] The variant word recognition based on the large language model is used to recognize the semantic variant words.

[0043] According to a third aspect, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the live broadcast violation detection method based on variant word recognition as described in the first aspect are implemented.

[0044] In a fourth aspect, a readable storage medium is provided, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the live broadcast violation detection method based on variant word recognition as described in the first aspect are implemented.

[0045] Beneficial effects:

[0046] This application adopts different identification and detection methods to deal with different types of variant words, and solves the problem of variant word detection in live broadcasts in a targeted manner; at the same time, based on the constructed multi-level variant word recognition system, variant word recognition methods of different degrees of sophistication are adopted at different time granularities, achieving a balance between real-time and accuracy; through efficient and accurate detection methods and systems, this application provides technical guarantees for standardizing the market order of live broadcast sales and protecting the interests of consumers, effectively preventing the spread of bad information and products. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 A flowchart of a method for detecting live broadcast violations based on variant word recognition provided in an embodiment of the present application;

[0048] Figure 2 A schematic diagram of a process for performing multi-level variant word identification steps provided in an embodiment of the present application;

[0049] Figure 3 A schematic diagram of another process for performing multi-level variant word identification steps provided in an embodiment of the present application;

[0050] Figure 4 A schematic diagram of a Spark large model prompt provided in an embodiment of the present application;

[0051] Figure 5 A schematic diagram of the structure of a live broadcast violation detection device based on variant word recognition provided in an embodiment of the present application;

[0052] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] To further illustrate the technical means and effects of the present invention to achieve its intended purpose, the technical solutions in the embodiments of this application are clearly described. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of this application are within the scope of protection of this application.

[0054] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0055] The description of the method flow in the specification of this application and the steps in the flowcharts in the drawings of the specification of this application do not necessarily need to be strictly executed according to the step numbers. The method steps can be executed in a different order. In addition, some steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.

[0056] The following is a detailed description of the live broadcast violation detection method, device, equipment and storage medium based on variant word recognition provided by the embodiments of the present application in combination with the accompanying drawings and preferred embodiments.

[0057] First, the application scenarios of the live broadcast violation detection method, device, equipment and storage medium based on variant word recognition in the embodiment of the present application are described in detail.

[0058] To ensure their message reaches every viewer and avoid being filtered by the platform's automated censorship system for sensitive or restricted terms, some livestreamers often convert sensitive terms into variations. For example, relevant laws stipulate that "advertisements must not contain false or misleading content, nor deceive or mislead consumers." Therefore, overly absolute terms like "100%" are considered extreme and sensitive terms that shouldn't appear in livestream sales. However, to fully demonstrate the superiority of their products and enhance their promotional effectiveness, livestreamers may choose to use variations with similar meanings but completely different structures, such as "99%+1" or "One Zero Zero." This approach avoids triggering platform censorship while also generating interest among consumers, sparking curiosity and a desire to buy.

[0059] The platform's censorship system is supposed to safeguard the healthy development of livestreaming e-commerce and prevent the spread of harmful information and products. However, the use of variant words effectively evades this responsibility, potentially leading to the proliferation of prohibited products and false advertising, disrupting market order and harming consumer interests. Furthermore, this practice not only erodes the accuracy and seriousness of language but also renders previously standard and clear business communications ambiguous and casual.

[0060] In the context of live e-commerce, the recognition of variant words still has some shortcomings and defects in terms of real-time and accuracy.

[0061] One of the hallmarks of live e-commerce is its strong real-time nature, requiring systems to respond quickly and process large amounts of data. However, identifying variant words often requires complex natural language processing. To achieve high-precision variant word recognition, the system must run complex models, placing high demands on computing resources and memory. In live streaming scenarios, if every livestream room required real-time processing of these models, it would significantly increase server load, impacting overall system performance and potentially preventing timely implementation of regulatory measures.

[0062] Variant words come in many forms, including homophones, near-phonetic words, misspellings, pinyin abbreviations, and homophones. This diversity, coupled with the continuous evolution of online language, makes it extremely difficult to construct a comprehensive dictionary covering all variants, thus impacting recognition accuracy. Furthermore, live streamers may intentionally use vague or ambiguous vocabulary to circumvent censorship, or employ language features like speech speed and intonation to mask the true meaning of variant words. This also complicates variant word recognition and reduces accuracy.

[0063] Therefore, to address the above-mentioned issues of live broadcast violation detection, especially the real-time and accuracy issues of variant word recognition, an embodiment of the present application provides a live broadcast violation detection method based on variant word recognition. It uses a multi-level variant word recognition method to detect variant words, including three modules: variant word recognition based on regular expressions, variant word recognition based on statistical language models, and variant word recognition based on large models. It also adopts a time priority hierarchical filtering strategy, using variant word recognition methods of different precision at different time nodes. In the specific implementation process, the live broadcast room audio is converted into text data and visual text is converted into text data through a pre-trained speech recognition model (Automatic Speech Recognition, ASR) and optical character recognition model (Optical Character Recognition, OCR), respectively. The extracted text is handed over to the three variant word recognition modules for parallel processing. When a variant word is detected, the original word of the variant word is matched with a pre-constructed sensitive word library. If the original word exists in the sensitive word library, the system will save the video clips and variant words within three minutes before and after for verification.

[0064] See Figure 1 , the embodiment of the present application provides a method for automatically generating preference data for secure alignment of large language models, such as Figure 1 As shown, the method for automatically generating preference data in an embodiment of the present application includes:

[0065] Step S1: Based on the speech recognition model and the optical character recognition model, the audio and visual text of the live broadcast room are obtained and converted into text data.

[0066] Using pre-trained Automatic Speech Recognition (ASR) and Optical Character Recognition (OCR) models, we convert livestream audio and visual text into text, respectively. The resulting text is cleaned to remove noise, punctuation, and special characters to improve the accuracy of subsequent processing.

[0067] Step S2: extract the text data and perform multi-level variant word recognition.

[0068] The multi-level variant word recognition includes: variant word recognition based on regular matching, variant word recognition based on statistical language model and variant word recognition based on large language model.

[0069] This application conducts a large amount of statistical analysis on the variant words commonly used by live-streaming anchors on different platforms. Based on the statistical analysis results of the variant words commonly used by live-streaming anchors on different platforms, the variant words are divided into structural variant words, phonetic variant words, and semantic variant words.

[0070] Structural variants, including those that alter the physical structure of the original word while maintaining visual or auditory similarities, primarily circumvent censorship by altering the original word's internal structure, such as by inserting, deleting, or replacing characters. They are visually or auditorily similar to the original word, but with structural changes. This strategy maintains visual or auditory similarity to the original word, allowing people to understand its true intent while circumventing machine automatic recognition. For example, in live broadcasts of non-pharmaceuticals or medical devices, due to relevant legal provisions that "except for advertisements for medical treatments, drugs, and medical devices, no other advertisements involving disease treatment functions are allowed, and medical terms or terms that could easily confuse the promoted products with drugs or medical devices are prohibited," broadcasters may replace the banned word "clinical" with "Clinical Bed" and "enhance resistance" with "enhance and strengthen resistance" to evade platform review. The following table lists some of the most frequently occurring structural variants and their corresponding original words:

[0071] Original word Variants Hospital A certain hospital clinical Lin Mou Bed prevention Prevention and Control Improve digestion Promote digestion

[0072] Phonetic variants, including variants that replace characters with similar pronunciations or similar writing forms; relying on Chinese homophones or characters with similar shapes, they avoid being recognized by the automatic review system by replacing them with characters with similar sounds or shapes. They are similar to the original word in pronunciation or writing form, but achieve the effect of confusion through replacement. Such variant words mainly exist in information such as the background and title of the live broadcast room. For example, in the live broadcast room of health products, when the anchor promotes the "anti-sugar" function of the product, they may use the homophonic "K sugar" as a variant in the background of the live broadcast room to avoid censorship.

[0073] Semantic variants, including variants with different structures but similar semantics or that can be used to refer to; by changing the semantic context of the word or using words with the same or similar meanings to replace the original word, the purpose is to avoid detection while maintaining information transmission. Such variant words may have completely different structures from the original word, but can achieve the same expression purpose through semantic association. For example, some anchors will replace the prohibited word "doctor" with "white coat", and "tablet" will be replaced with "white tablets" to prevent the live broadcast room from being blocked by the detection system. Specific examples of semantic variants are shown in the following table:

[0074] Original word Variants doctor White coat pill White flakes age spots Old age Blood vessel Red pipes

[0075] See Figure 2-3 , the multi-level variant word recognition described in this application, that is, for the above three categories of variant words, the extracted text is processed in parallel through three targeted variant word recognition modules. Specifically, it includes:

[0076] Step S201, based on the text data, identify structural variant words through a preset regular expression.

[0077] Exemplarily, through the statistical analysis of the data of structural variant words on each live broadcast platform, the vast majority of anchors use the form of inserting characters and follow specific rules. Here, for the four most frequently occurring variant methods, regular expressions are set manually according to the form of inserting the words "certain", "what", "little" and "and" to extract variant words from the original text. The following are the targeted regular expression designs for the four structural variants of inserting the words "certain", "what", "little" and "and":

[0078] r‘([\u4e00-\u9fa5]+)certain([\u4e00-\u9fa5]+)’ is in the form of ‘*certain*’,

[0079] r‘([\u4e00-\u9fa5]+)what([\u4e00-\u9fa5]+)’ is in the form of ‘*what*’,

[0080] r‘([\u4e00-\u9fa5]+) small ([\u4e00-\u9fa5]+)’ is in the form of ‘* small *’.

[0081] r‘([\u4e00-\u9fa5]+) and ([\u4e00-\u9fa5]+)’ is in the form of ‘* and *’.

[0082] When the text detected by the OCR and ASR modules contains text with these four inserted words, the characters before and after the inserted words are combined and matched with the designed sensitive word library. If the combined word exists in the sensitive word library, the detected word is stored as a variant word of the sensitive word in the variant word library.

[0083] The following is an example of variant word recognition based on regular expressions in this application:

[0084] "The original input is: This product is the most popular;

[0085] Sensitive word detected: Most popular;

[0086] Variant word saved: Most what popular;

[0087] The original input is: This has a therapeutic effect on our cardiovascular and cerebrovascular diseases;

[0088] Sensitive word detected: Cardiovascular;

[0089] Variant word saved: Cardiovascular and cerebrovascular;

[0090] The original input is: The effect of this has been verified clinically;

[0091] Sensitive word detected: Clinically;

[0092] Variant word saved: Clinically".

[0093] Based on this, the variant word recognition method based on regular expressions can quickly and accurately detect structural variant words in the text according to the pre-set regular expressions.

[0094] Step S202, analyze the language features of the text data based on the statistical language model, correct spelling mistakes and identify phonetic variant words.

[0095] Considering the recognition requirements of phonetic variant words, this application uses a Chinese N-gram language model trained by the KenLM statistical language model tool for recognition. The Chinese NGram language model trained based on KenLM combines the rule method and the confusion set, which can not only correct Chinese spelling mistakes, but also identify homophones and near-homophone variants, effectively dealing with the phonetic and morphological changes of variant words.

[0096] Statistical language models are good at implementing variant word recognition based on homophones and near-homophones, and can handle more situations when combined with regular expression-based variant word models. The following are several examples of variant word recognition based on statistical language models:

[0097] “{'source': '这种药物的林床笑果非常好', 'target': '这种药物的临床效果非常好', 'errors': [('林床', '临床', 5), ('笑果', '效果', 7)]}

[0098] {'source': '咱们医院里的伊生都在推荐啊', 'target': '咱们医院里的医生都在推荐啊', 'errors': [('伊生', '医生', 6)]}

[0099] {'source': '咱们的欣闹学管疾病等等都可以治疗啊', 'target': '咱们的心脑血管疾病等等都可以治疗啊', 'errors': [('欣闹', '心脑', 3), ('学管', '血管', 5)]}”.

[0100] As can be seen from the third example, the statistical language model can replace the phonetic variant word "欣闹" in the text with the prohibited word "心脑". As can be seen from the first two examples, in addition to detecting variant words, the statistical language model can also detect typos in OCR and ASR texts. Therefore, combining the statistical language model with regular expression-based variant word recognition can detect more variant words hidden in the text.

[0101] Step S203: Based on the large model, perform semantic understanding and context analysis on the text data to identify semantic variant words.

[0102] Although the first two methods can solve most variant words, it is difficult to detect and identify emerging semantic variant words well. Moreover, over time, more and more semantic variant words will appear. Therefore, it is necessary to借助 a large language model that can understand deep semantics and also continuously learn and update as time goes by.

[0103] It should be noted that there is an incorrect Chinese character "借助" in the original text which should be in English. I have translated it as "借助" for now as it seems to be a non-standard or incorrect usage in this context. If it was meant to be something else, please correct the original text for a more accurate translation.Large language models are particularly good at understanding the deep meaning and context of language and can take into account the context information around words. This means that even if variant words are structurally very different from the original words, as long as they are semantically related, the large model can recognize this connection. This is crucial for identifying and understanding semantic variant words, which often avoid detection by replacing words with the same or similar meanings. Moreover, for application scenarios that need to handle semantic variant words in multiple languages or dialects, large language models usually have cross-linguistic understanding capabilities, which enable them to work effectively in a wider cultural and geographical context. Most importantly, even though new semantic variant words keep emerging, large language models that continuously learn and update can adapt to this change because they can learn new language patterns and vocabulary usage from a large amount of text.

[0104] The Spark large model can understand complex language structures and deep semantic information and can handle most cases where the intention is cleverly hidden in semantic variant words. More importantly, compared with other large models, the Spark large model has a huge advantage in the accumulation of Chinese language corpora and can detect Chinese semantic variant words more accurately and efficiently. This application uses the Spark large model as an important component of the variant word recognition module based on the large model.

[0105] By calling the Spark large model API interface and designing reasonable prompts, the semantic-level understanding of text data is achieved to identify words with deep semantic associations. Exemplarily, the prompt design is shown in Figure 4 .

[0106] The following are examples of the Spark large model processing semantic variant words:

[0107] "I: It's much taller, and I even have to go to see the white coat to take white pills. Then you should take the item in Link 6.

[0108] Spark:

[0109] The variant words in this passage are paired with their corresponding original words as follows:

[0110] Variant word: white coat, Original word: doctor

[0111] Variant word: take white pills, Original word: take medicine".

[0112] It can be seen that the Spark large model can detect the semantic-level variant words "white coat" and "white pills" and give the correct original words according to the semantics. Thanks to the powerful semantic understanding ability of the Spark large model and the large accumulation of Chinese language corpora, the Spark large model has proven its effectiveness in processing semantic variant words.

[0113] It is worth noting that this application demonstrates the feasibility of variant word recognition from three modules: variant word recognition based on regular expressions, variant word recognition based on statistical language models, and variant word recognition based on large models. Among them, variant word recognition based on regular expressions consumes the least computing power, but requires manual setting of the rule base and is highly dependent on experience; variant word recognition based on statistical language models is executed locally and is good at identifying variant words of the same or similar sound types; variant word recognition based on large models is implemented by calling remote API interfaces, which can achieve semantic-level understanding.

[0114] In order to coordinate the three recognition methods and form an overall variant word recognition process, this application comprehensively utilizes the above three different technical strategies when building a multi-level variant word recognition system, and adopts a time priority hierarchical filtering strategy, namely: the time priority of variant word recognition based on regular expressions > the time priority of variant word recognition based on statistical language models > the time priority of variant word recognition based on large models.

[0115] After the multi-level variant word recognition is started, the variant word recognition based on regular matching provides feedback; after each set first time period, the variant word recognition based on the statistical language model recognizes and provides feedback on the data of the previous first time period; after each set second time period, the variant word recognition based on the large language model recognizes and provides feedback on the data of the previous second time period; wherein the length of the first time period is less than the length of the second time period.

[0116] Specifically, in the real-time stream of live video content, regular expression-based variant word recognition can quickly identify and filter out obvious structural variant words according to preset regular expressions; and it can continuously update the variant word library, which improves the update frequency to a certain extent; the low latency and high efficiency of regular expression-based variant word recognition make it an ideal choice for preliminary filtering, providing the first level of real-time variant word recognition for multi-level variant word recognition.

[0117] As live content accumulates, after a set period of time, the statistical language model identifies variant words, including homophones and near-phonetic variants. For example, after one minute, the system will initiate the variant word identification process based on the statistical language model, specifically conducting in-depth analysis of variants of homophones and near-phonetic variants. The statistical language model accurately identifies variant words that have undergone changes in sound and form through pre-trained language models and confusion matrices, addressing the limitations of regular expression-based recognition methods, which are limited by the comprehensiveness and update frequency of the rule base. It also provides a mid-level latency layer for multi-level variant word recognition, balancing accuracy with variant word recognition.

[0118] Furthermore, after a set period of time, the large model-based recognition processes deeper semantic variant words. For example, after 5 minutes of cumulative analysis, the system will initiate variant word recognition based on the large model. This step is the most sophisticated part of the entire recognition process. By calling the large model through a remote API interface, the system can understand the deep semantics of the text and identify complex variant words that have undergone subtle changes in structure and sound. This process provides strong support for identifying semantic variant words, especially sensitive words that are neither structural nor sound variants, providing the highest accuracy variant word recognition at a long-term cumulative level for multi-level variant word recognition.

[0119] The time priority graded filtering strategy of multi-level identification in this application ensures that variant word identification methods of different degrees of precision are adopted at different time nodes. Each method plays a unique role in the variant word identification process, aiming to take into account both real-time and accuracy. Such a design not only takes into account the need for real-time response, but also ensures in-depth analysis of long-term accumulated text, achieving a balance between efficiency and accuracy, and providing a content review solution that is both efficient and reliable for live streaming platforms.

[0120] In some possible implementations, after performing multi-level variant word identification, the following step S204 is further included:

[0121] The original word of the identified variant word is compared with the sensitive word library. If the original word exists in the sensitive word library, the variant word is stored in the variant word library.

[0122] Step S3: Based on the identified variant word, the original word of the variant word is obtained, and the original word is matched with a pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library.

[0123] When a variant word is detected, the original word of the variant word will be matched with the pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library;

[0124] Step S4: If the original word exists in the sensitive word library, the video data of the set time before and after the variant word is retrieved and saved as violation evidence.

[0125] For example, when the system detects a violation, it saves video clips and variant words within three minutes before and after the variant word as evidence. This automatically detected violation data can be manually reviewed to ensure detection accuracy. Based on the review results, the violating content can be handled accordingly (such as warnings, bans, etc.), and feedback can be provided to relevant parties (such as streamers and platform administrators).

[0126] To sum up, this application adopts different identification and detection methods to deal with different types of variant words, and solves the problem of variant word detection in live broadcasts in a targeted manner; at the same time, based on the constructed multi-level variant word recognition system, variant word recognition methods of different degrees of sophistication are adopted at different time granularities, achieving a balance between real-time and accuracy; through an efficient and accurate detection system, this application provides technical guarantees for regulating the market order of live broadcast sales and protecting the interests of consumers, effectively preventing the spread of bad information and products.

[0127] See also Figure 5 Corresponding to the above-mentioned embodiment of the live broadcast violation detection method based on variant word recognition, the embodiment of the present application provides a live broadcast violation detection device based on variant word recognition, the live broadcast violation detection device comprising:

[0128] The text data recognition and conversion module 1001 is used to obtain the audio and visual text of the live broadcast room based on the speech recognition model and the optical character recognition model, and convert them into text data;

[0129] A multi-level variant word identification module 1002 is used to extract the text data and perform multi-level variant word identification;

[0130] The variant word and sensitive word matching module 1003 is used to obtain the original word of the variant word based on the identified variant word, and match the original word with a pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library;

[0131] Evidence collection module 1004, configured to retrieve video data of a set length of time before and after the variant word if the original word exists in the sensitive word library, and save it as violation evidence;

[0132] The multi-level variant word recognition includes: variant word recognition based on regular matching, variant word recognition based on statistical language model and variant word recognition based on large language model.

[0133] Furthermore, the variant words used to circumvent automated censorship include:

[0134] Structural variants, which include variants that change the physical structure of the original word while maintaining its visual or auditory similarity;

[0135] Variant words, including variant words that replace characters with similar pronunciations or written forms;

[0136] Semantic variant words, including variant words with different structures but similar semantics or that can refer to other words;

[0137] The variant word recognition based on regular expression matching is used to recognize the structural variant words;

[0138] The variant word recognition based on the statistical language model is used to identify the phonetic and morphological variant words;

[0139] The variant word recognition based on the large language model is used to recognize the semantic variant words.

[0140] The above-mentioned live broadcast violation detection device based on variant word recognition implements the steps and various processes of the above-mentioned live broadcast violation detection method based on variant word recognition, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0141] See also Figure 6 Corresponding to the above-mentioned embodiment of the live broadcast violation detection method based on variant word recognition, the embodiment of the present application provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the steps and various processes of the above-mentioned embodiment of the live broadcast violation detection method based on variant word recognition are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0142] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 1009 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0143] Processor 1010 may include one or more processing units. Optionally, processor 1010 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1010.

[0144] Corresponding to the above-mentioned embodiment of the live broadcast violation detection method based on variant word recognition, the embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps and various processes of the above-mentioned embodiment of the live broadcast violation detection method based on variant word recognition are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0145] The processor is the processor in the electronic device described in the embodiment of the present application. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0146] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0147] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0148] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A live broadcast violation detection method based on variant word recognition is characterized by: include: Based on the speech recognition model and optical character recognition model, the audio and visual text of the live broadcast room are obtained and converted into text data; Extracting the text data and performing multi-level variant word recognition; Based on the identified variant words, the original word of the variant word is obtained, and the original word is matched with a pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library; If the original word exists in the sensitive word library, retrieve the video data of the set length before and after the variant word and save it as violation evidence; The multi-level variant word recognition includes: variant word recognition based on regular matching, variant word recognition based on statistical language model and variant word recognition based on large language model; The multi-level variant word recognition is configured with a time priority hierarchical filtering strategy, namely: In a real-time stream of live video content, the temporal priority of variant word identification based on regular expression matching is higher than the temporal priority of variant word identification based on a statistical language model, and the temporal priority of variant word identification based on a statistical language model is higher than the temporal priority of variant word identification based on a large language model; After the multi-level variant word recognition is started, the variant word recognition based on regular expression matching is fed back; After each set first time period, the variant word recognition based on the statistical language model recognizes and feeds back the data of the previous first time period; After each set second time period, the variant word recognition based on the large language model recognizes the data of the previous second time period and feeds back the data.

2. The method for detecting live broadcast violations based on variant word recognition according to claim 1 is characterized in that: These variations are used to circumvent automated censorship and include: Structural variants, which include variants that change the physical structure of the original word while maintaining its visual or auditory similarity; Variant words, including variant words that replace characters with similar pronunciations or written forms; Semantic variant words, including variant words with different structures but similar semantics or that can refer to other words; The variant word recognition based on regular expression matching is used to recognize the structural variant words; The variant word recognition based on the statistical language model is used to identify the phonetic and morphological variant words; The variant word recognition based on the large language model is used to recognize the semantic variant words.

3. The method for detecting live broadcast violations based on variant word recognition according to claim 2, characterized in that: The multi-level variant word identification includes: Based on text data, structural variant words are identified through preset regular expressions; Analyze the language characteristics of text data based on statistical language models, correct spelling errors and identify phonetic and morphological variants; Perform semantic understanding and context analysis on text data based on large models to identify semantic variant words.

4. The live broadcast violation detection method based on variant word recognition according to any one of claims 1 to 3 is characterized in that: If the original word exists in the sensitive word library, the following steps are also performed: The original word of the identified variant word is compared with the sensitive word library. If the original word exists in the sensitive word library, the variant word is stored in the variant word library.

5. The method for detecting live broadcast violations based on variant word recognition according to claim 1, characterized in that: The duration of the first time period is shorter than the duration of the second time period.

6. A live broadcast violation detection device based on variant word recognition, characterized in that: The method for detecting live broadcast violations based on variant word recognition according to claim 1 comprises: The text data recognition and conversion module is used to obtain the audio and visual text of the live broadcast room based on the speech recognition model and the optical character recognition model, and convert them into text data; A multi-level variant word recognition module, configured to extract the text data and perform multi-level variant word recognition; A variant word and sensitive word matching module is used to obtain the original word of the variant word based on the identified variant word, and match the original word with a pre-constructed sensitive word library to determine whether the original word exists in the sensitive word library; an evidence-building module, configured to retrieve video data of a set length of time before and after the variant word if the original word exists in the sensitive word library, and save it as evidence of violation; The multi-level variant word recognition includes: variant word recognition based on regular matching, variant word recognition based on statistical language model and variant word recognition based on large language model.

7. The live broadcast violation detection device based on variant word recognition according to claim 6 is characterized in that: These variations are used to circumvent automated censorship and include: Structural variants, which include variants that change the physical structure of the original word while maintaining its visual or auditory similarity; Variant words, including variant words that replace characters with similar pronunciations or written forms; Semantic variant words, including variant words with different structures but similar semantics or that can refer to other words; The variant word recognition based on regular expression matching is used to recognize the structural variant words; The variant word recognition based on the statistical language model is used to identify the phonetic and morphological variant words; The variant word recognition based on the large language model is used to recognize the semantic variant words.

8. An electronic device, characterized in that: The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the live broadcast violation detection method based on variant word recognition as described in any one of claims 1 to 5 are implemented.

9. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, which, when executed by a processor, implements the steps of the live broadcast violation detection method based on variant word recognition as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Multi-level natural language anti-junk text method and system

    CN109977416A

  • Variant sensitive word recognition method and device, electronic equipment and storage medium

    CN117574887A