A live speech sensitive word monitoring method and system
By using intelligent monitoring methods to detect sensitive words in live-streaming dialogue, a set of timestamped facial and voice clips is generated. Specific expressions are filtered and converted into text, and combined with a sensitive word database for detection. This solves the problems of incomplete monitoring and low efficiency in existing technologies, and achieves efficient and accurate sensitive word monitoring.
Patent Information
- Application Number
- CN202111296886.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-02
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2041-11-02
AI Technical Summary
Existing live streaming platforms suffer from incomplete coverage, high labor costs, and low detection efficiency when monitoring sensitive words in live streaming content. In particular, manual monitoring and direct text detection methods are not highly targeted.
The system employs intelligent monitoring methods to acquire live video and audio, generate timestamped sets of facial images and voice clips, filter specific expressions and convert them into text fragments, and combine this with a sensitive word database for detection and deduction, thus forming an intelligent monitoring system.
It improved the targeting and efficiency of sensitive word monitoring, shortened the monitoring time, reduced manpower costs, and improved the translation accuracy of homophones by building a domain corpus.
Smart Images

Figure CN114022933B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio and video detection technology, and in particular to a method and system for monitoring sensitive words in live broadcast speech. Background Art
[0002] As the live streaming e-commerce business becomes more and more popular, industry supervision is gradually becoming stricter. The control of the content of live streaming language has become a pain point that urgently needs to be solved in the NLP field. Whether from the perspective of brands or MCNs, the supervision of live streaming language is imminent.
[0003] At present, most platforms use manual monitoring methods to monitor sensitive words, that is, manually checking the live broadcast rooms to monitor whether there are sensitive words in the live broadcast language. The disadvantage of this monitoring method is that it cannot cover all live broadcast rooms and has the defect of high labor costs; some platforms also use the method of intercepting the voice text and text text output by the anchor user in the live broadcast client, converting the obtained voice text into converted text (text format), matching the converted text and text with sensitive words, and obtaining the sensitive word values to detect sensitive words. Although this method can greatly reduce the workload of administrators on the live broadcast platform, it blindly and directly detects the text segment set generated by the voice segment set, and has the problems of non-targeted detection and low detection efficiency. Summary of the Invention
[0004] In response to the above problems, an object of the present invention is to provide a method for monitoring sensitive words in live broadcast speech, which intelligently monitors and warns of sensitive words (politically sensitive words, words that violate the Advertising Law, platform-violating words, mentions of competitor names, etc.) in live broadcast speech; and the method uses targeted detection of text fragments (speech) corresponding to negative expressions, which effectively improves monitoring efficiency and shortens monitoring time.
[0005] The second object of the present invention is to provide a system for monitoring sensitive words in live broadcast speech.
[0006] The first technical solution adopted by the present invention is: a method for monitoring sensitive words in live broadcast speech, comprising the following steps:
[0007] S100: Acquire live video and live audio, pre-process the live video to generate a set of face images with timestamps, and pre-process the live audio to generate a set of voice clips with timestamps;
[0008] S200: Filtering specific expression images from the face image set with timestamps, and generating a text segment set with timestamps based on the voice segment set with timestamps;
[0009] S300: searching for a text segment corresponding to the specific emoticon image based on the timestamp, and generating a speech text;
[0010] S400: Detecting whether the speech text contains sensitive words based on the sensitive word library. If sensitive words are detected, calculating the hit rate of the sensitive words, and deducting points from the anchor corresponding to the speech text based on the hit rate of the sensitive words;
[0011] S500: Count the points deducted from each anchor to obtain the total points deducted from each anchor, and supervise any anchor when the total points deducted from the anchor are greater than a threshold.
[0012] Preferably, the pre-processing of the live video in step S100 includes: generating a set of face pictures with timestamps by slicing the live video according to time.
[0013] Preferably, the pre-processing of the live audio in step S100 includes: generating a set of voice segments with timestamps for the live audio by time slicing.
[0014] Preferably, the step S200 of screening specific expression pictures from the face picture set with timestamps includes the following sub-steps:
[0015] S211: Recognize the expression of each face image and assign a corresponding expression label;
[0016] S212: Filtering facial images with specific expression types based on the expression tags, where the specific expression types include one or more preset expression types.
[0017] Preferably, the step S211 includes:
[0018] The emotion-recognition algorithm is used to perform real-time expression recognition on facial images, and corresponding expression labels are assigned to each facial image according to different expression types.
[0019] Preferably, generating a text segment set with timestamps based on the speech segment set with timestamps in step S200 includes:
[0020] An automatic speech recognition algorithm is used to convert each speech segment in the speech segment set into a corresponding text segment, thereby obtaining a text segment set with a timestamp.
[0021] Preferably, the method further includes pre-training the automatic speech recognition algorithm with corpus, including the following steps:
[0022] Assign TF-IDF weights to the discourse data in the corpus to be trained;
[0023] Sort all the assigned words by score from high to low to get the pre-training score of the speech technique;
[0024] When automatic speech recognition encounters homophones, it matches them from high to low based on the pre-trained scores of the speech script.
[0025] Preferably, the S400 includes:
[0026] Match the words in the dialogue text based on the sensitive word library. If a sensitive word is detected, find the anchor corresponding to the dialogue text based on the timestamp corresponding to the dialogue text and the schedule of the live broadcast; count the hit rate of the sensitive words in the dialogue text, and deduct points from the anchor corresponding to the dialogue text based on the hit rate.
[0027] Preferably, the method further includes S600: forming a sensitive word intelligent dashboard based on the statistical results of the sensitive word hit rate, and displaying the points deduction situation of each anchor in real time.
[0028] The second technical solution adopted by the present invention is: a system for monitoring sensitive words in live broadcast speech, including a preprocessing module, a screening module, a text segment set generation module, a matching module, a detection module and an early warning module;
[0029] The preprocessing module is used to obtain live video and live audio, preprocess the live video to generate a set of face pictures with timestamps, and preprocess the live audio to generate a set of voice clips with timestamps;
[0030] The screening module is used to screen specific expression pictures from a collection of face pictures with timestamps;
[0031] The text segment set generation module is used to generate a text segment set with a timestamp based on the speech segment set with a timestamp;
[0032] The matching module is used to search for a text segment corresponding to the specific emoticon image based on a timestamp and generate a speech text;
[0033] The detection module is used to detect whether the speech text contains sensitive words based on the sensitive word library. If sensitive words are detected, the hit rate of the sensitive words is calculated, and the anchor corresponding to the speech text is deducted based on the hit rate of the sensitive words;
[0034] The early warning module is used to count the points deducted by each anchor to obtain the total points deducted by each anchor, and to supervise the anchor when the total points deducted by the anchor are greater than a threshold.
[0035] Beneficial effects of the above technical solution:
[0036] (1) The present invention discloses a method for monitoring sensitive words in live broadcast speech, which intelligently monitors and issues early warnings for sensitive words (politically sensitive words, words that violate the Advertising Law, platform-violating words, mentions of competitor names, etc.) in live broadcast speech.
[0037] (2) The present invention adopts targeted detection of text fragments (speech) corresponding to negative expressions, which effectively improves the monitoring efficiency and shortens the monitoring time.
[0038] (3) In order to more accurately output the correct text for "homophones with different characters", the present invention constructs a large number of corpora in the field of fast-moving consumer goods; taking the beauty industry as an example, a corpus of the beauty industry is constructed, and the ASR is pre-trained based on the corpus in the beauty industry corpus. When "homophone" characters need to be translated into text, they are translated into the text in the beauty industry corpus first, thereby improving the accuracy of ASR live translation in the field of fast-moving consumer goods. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A flowchart of a method for monitoring sensitive words in live broadcasting provided by one embodiment of the present invention;
[0040] Figure 2 A flowchart of a method for monitoring sensitive words in live broadcasting provided by one embodiment of the present invention;
[0041] Figure 3 A schematic diagram of a speech text provided by one embodiment of the present invention;
[0042] Figure 4 A schematic diagram of the results of sensitive word detection in a speech text provided by one embodiment of the present invention;
[0043] Figure 5 A schematic diagram of a sensitive word smart signboard provided by one embodiment of the present invention;
[0044] Figure 6 A schematic diagram of the structure of a system for monitoring sensitive words in live broadcast speech provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0045] The following detailed description of the embodiments of the present invention is provided in conjunction with the accompanying drawings and examples. The following detailed description of the embodiments and the accompanying drawings are intended to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention. That is, the present invention is not limited to the preferred embodiments described, and the scope of the present invention is defined by the claims.
[0046] In the description of the present invention, it should be noted that, unless otherwise specified, “plurality” means two or more; the terms “first”, “second”, etc. are used for descriptive purposes only and cannot be understood as indicating or implying relative importance; for ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to the specific circumstances.
[0047] Example 1
[0048] like Figure 1 and Figure 2 As shown, this embodiment discloses a method for monitoring sensitive words in live broadcast speech, including the following steps:
[0049] S100: Acquire live video and live audio; pre-process the live video to generate a set of face images with timestamps, and pre-process the live audio to generate a set of voice clips with timestamps;
[0050] In the same live broadcast, the live video (image) and live audio (sound) will be stored separately according to the timeline. That is, the live video will be stored in the video database, and the live audio will be stored in the voice database. The live video and live audio will be stored one by one according to the timeline for easy subsequent slicing processing.
[0051] Obtaining live video means obtaining all past live videos (all historical live videos) from the video database, and obtaining live audio means obtaining all past live audio (all historical live audio) from the voice database;
[0052] The live video preprocessing (slicing) is specifically as follows: slicing the live video by time to generate a set of face images with timestamps; for example, taking screenshots of any live video frame by frame every 30 seconds to complete the live video image slicing;
[0053] The pre-processing (slicing) of the live audio is specifically as follows: slicing the live audio by time to generate a set of voice segments with timestamps; for example, slicing any live audio into segments every 30 seconds to complete the live audio slicing.
[0054] S200: Filtering specific expression images (time-stamped face images with specific expressions) from the time-stamped face image set, and generating a time-stamped text segment set based on the time-stamped voice segment set;
[0055] Filtering time-stamped facial images with specific expression types includes the following sub-steps:
[0056] S211: assigning a corresponding expression label to each timestamped face image in the timestamped face image set, specifically: using an emotion-recognition algorithm to perform real-time expression recognition on the timestamped face image set, and assigning a corresponding expression label to each timestamped face image according to different expression types (expression types include, for example, anger, disgust, fear, happiness, sadness, surprise, and neutrality); the emotion-recognition algorithm is an open source expression recognition algorithm that can recognize the host's expression category in real time, thereby forming a real-time expression label; expression labels include (but are not limited to) Angry, Disgust, Scared, Happy, Sad, Surprise, and Neutral;
[0057] S212: Filtering time-stamped facial images with specific expression types based on expression tags, where specific expressions include negative expressions, such as anger, disgust, fear, and surprise;
[0058] The above expression tags are all text data, and the computer can automatically filter out the "face pictures" and corresponding timestamps behind these "text tags".
[0059] Generating a set of text segments with timestamps based on a set of speech segments with timestamps includes the following sub-steps:
[0060] An automatic speech recognition algorithm is used to convert each speech segment in the speech segment set with timestamps into a corresponding text segment, thereby obtaining a text segment set with timestamps.
[0061] Furthermore, in one embodiment, the automatic speech recognition algorithm is, for example, Automatic Speech Recognition (ASR). For the fast-moving consumer goods industry in various fields, in order to more accurately output the correct text for "homophones with different characters", the present invention constructs a large number of corpora in the fast-moving consumer goods field; taking the beauty industry as an example, a beauty industry corpus is constructed, and the ASR is pre-trained based on the corpus in the beauty industry corpus. When "homophone" characters need to be translated into text, they are translated into text in the beauty industry corpus first, thereby improving the accuracy of ASR live translation in the fast-moving consumer goods field.
[0062] The specific steps for ASR corpus pre-training are as follows: TF-IDF (term frequency–inverse document frequency) weights are assigned to industry jargon corpora in any fast-moving consumer goods field, that is, each word is assigned a value based on its frequency of occurrence and its actual meaning. For example, words with high frequency of occurrence and actual meaning are assigned higher scores. All assigned words are sorted from high to low according to the scores to obtain the industry jargon pre-training scores. During the ASR homophone translation process, matching is performed from high to low based on the industry jargon pre-training scores; for example, "yan xian an" prioritizes matching "niacinamide" rather than "eyeliner press".
[0063] S300: matching a specific expression picture (a face picture with a time stamp and a specific expression type) with a corresponding text segment based on the timestamp to generate a speech text;
[0064] All text fragments in the timestamped text fragment set are timestamped. Timestamped face pictures with specific expression types also have timestamps. According to the timestamps corresponding to the filtered pictures (timestamped face pictures with specific expression types), the text fragments with the same timestamp as the filtered pictures are found from the timestamped text fragment set to obtain the speech text (the speech text is the text fragment corresponding to the specific expression); the speech text is as follows: Figure 3 As shown, the speech text includes but is not limited to text corresponding to expressions of anger, disgust, fear, and surprise (negative expressions).
[0065] There is a high probability that some negative expressions may contain uncivilized behaviors such as swearing and descriptions prohibited by advertising laws such as exaggerating product descriptions. Therefore, after these negative expressions appear, the present invention uses the corresponding timestamp to check the corresponding anchor's speech to detect whether there are any violations in the speech. Compared with blindly detecting the text segment set generated by the voice segment set, the present invention uses targeted detection of the text segments (speech) corresponding to the negative expressions, which effectively improves the monitoring efficiency and shortens the monitoring time.
[0066] S400: Detecting whether the speech text contains sensitive words based on the sensitive word (banned word) library. If sensitive words are detected, deducting points from the host corresponding to the speech text based on the hit rate of sensitive words;
[0067] like Figure 4As shown, the words in the dialogue text are matched based on the sensitive word (banned word) library. If a sensitive word is detected (a sensitive word is hit), the anchor corresponding to the dialogue text is found based on the timestamp corresponding to the dialogue text and the live broadcast schedule (the schedule records the names of the anchors corresponding to different time periods and different live broadcast rooms), and points are deducted from the anchor corresponding to the dialogue text (for example, 1 point is deducted). One point is deducted each time a sensitive word is triggered (a sensitive word is hit). The hit rate of the sensitive words in the dialogue text is counted and points are deducted from the anchor corresponding to the dialogue text.
[0068] like Figure 5 As shown, based on the above-mentioned sensitive word hit rate results, a sensitive word smart dashboard is formed. Through the sensitive word smart dashboard, the violation records triggered by the anchor in each game can be reviewed to achieve continuous improvement; it is convenient for brands to understand what happened in the live broadcast room to achieve timely warnings and prevent the live broadcast platform from deducting points from the live broadcast room or even banning the account.
[0069] Sensitive word types include competitor product names (Type 1), spokespersons / fans (Type 2), platform violation words (Type 3), product description violations (Type 4), and advertising law violations (Type 5);
[0070] The sensitive word (banned word) library for platform violation words (type 3) and advertising law violations (type 5) includes but is not limited to "most series", "related to one", "related to ultimate / ultimate", "related to country / first / family", "related to brand", "related to false advertising", "related to authority", "related to fraud (inducing consumers), "related to time", and common beauty and cosmetics words; there is overlap between platform violation words (sensitive words) and advertising law violation words (sensitive words), and there may be situations where both violation information is triggered at the same time; the sensitive word (banned word) library for product description violations (type 4) includes but is not limited to words related to false advertising of cosmetics; different brands define different competitors in competitor names (type 1), so there is no fixed sensitive word (banned word) library; based on regulatory requirements, different brands will have specific words that are triggered in spokespersons / fans (type 2), and live broadcast rooms cannot incite fans to place orders for "idols" or make impulsive purchases.
[0071] Among them, the most series of sensitive words include most (best, highest-end, cheapest, first to enjoy, most advanced), best (best, most luxurious, lowest price in history, most suitable, last), most (largest, lowest, most popular, most comfortable, last wave), favorite (maximum, lowest level, most popular, first, latest), most profitable (highest, lowest price, most fashionable, most advanced technology, latest technology), and best (superlative, bottom, most gathered, latest model); sensitive words can be replaced with available related replacement words. Related replacement words for the most series of sensitive words include (but are not limited to) relatively concealing, comfortable, very, special, rock-bottom price, another wave, and omitting the word "most";
[0072] Sensitive words related to "one" include "China's No. 1," "No. 1," "Only this time," "Only this day," "No. 1 on the entire network," "TOP 1," "The last wave," "No. 1 in sales," "Unrivaled," "One of the top X brands on the entire network," "Ranked first," "Top brand," "Only," "First-rate," and "One day" (but not limited to these). Alternative words to "one" include "Excellent sales," "Fighter product," "Leading product," "Specific event date," and "High ranking" (but not limited to these).
[0073] Sensitive words related to "ultimate" include "national level", "top level", "top quality", "top craftsmanship", "national product", "global top", "cosmic level", "high-end", "ultimate", "cutting-edge" and "absolute" (but not limited to these). Alternative words to "ultimate" include "national" and "global";
[0074] Sensitive words related to "national first" include "first," "network-wide first," "China's well-known trademark," "X network first," "first choice," "national first," "first time," "international quality," "exclusive," "first," "exclusive formula," "national leader," "national inspection exemption," and "filling a domestic gap" (but not limited to these). Alternative words to sensitive words related to "national first" include "first," "priority," and "few" (but not limited to these).
[0075] Sensitive words related to brands include, but are not limited to, big brand, luxury, leading brand, supreme, leader, gold medal, world leader, peak, king, first to market, huge volume, champion, king, trump card, excellent, leading brand, head, creator, and senior; related alternative words to sensitive words related to brands include, but are not limited to, well-known brand, very popular, and very popular;
[0076] Sensitive words related to false advertising include (but are not limited to) special effects, unprecedented, invincible, unprecedented, permanent, all-natural, universal, and genuine leather; alternative words related to sensitive words related to false advertising include (but are not limited to) "free of XX additives" and "guaranteed authenticity and authenticity";
[0077] Sensitive words related to authority include "time-honored brand," "expert recommendation," "XX special person recommendation," "quality inspection exemption," "national well-known trademark," "special supply," and "exclusive supply" (but not limited to these). Related alternative words for sensitive words related to authority include "qualified" (but not limited to these).
[0078] Sensitive words related to fraud (inducing consumers) include "free nationwide," "click for a chance to win," "congratulations on winning," "free for all," "click to flip," "click to win," "it doesn't get any cheaper than this," "it's gone now," "crazy rush," and "flash sale" (but not limited to these). Alternative words to sensitive words related to fraud (inducing consumers) include "cannot induce clicks," "special price," "low stock, don't miss out," and "sold out as soon as it's on the shelves" (but not limited to these).
[0079] Sensitive words related to time include, but are not limited to, limited time, countdown, only for today, weekend, brand group, boutique group, several days and nights, strictly prohibited, restore to original price, immediate price reduction, price increase at any time, and end at any time; related replacement words for sensitive words related to time include, but are not limited to, must have a specific time and order time countdown;
[0080] Sensitive words commonly used in the beauty and makeup category include whitening, freckle removal, blackheads, pores, acne, wrinkle removal, pregnant women, firming and lifting, and anti-aging (but not limited to these); replacement words related to sensitive words commonly used in the beauty and makeup category include white, small spots on the face, a little dark on the nose, small wrinkles that want to be improved, expectant mothers, firming and lifting action demonstrations, and worry about looking like an old aunt at a young age (but not limited to these).
[0081] Sensitive words related to false advertising of cosmetics include special effects, strong effects, high efficiency, fast effects, whitening in one go, effective in X days, all-round, comprehensive, safe, non-toxic, fat burning, slimming, face slimming, weight loss, longevity, memory improvement, dissolving / removing dead cells, improving skin resistance to irritation, removing wrinkles, restoring broken elastic fibers, using new coloring textures, destroying melanin, blocking melanin formation, breast enhancement, breast enlargement, preventing sagging, and promoting sleep; related replacement words for sensitive words related to false advertising of cosmetics include descriptions of effects cannot be exaggerated, efficacy details page cannot be mentioned if it is not available, periodic words cannot be mentioned, and improvement will gradually occur after XX days.
[0082] S500: Count the points deducted for each anchor by live broadcast session, day, week or month to obtain the total points deducted for each anchor per session, day, week or month; supervise the anchor when the total points deducted are greater than a threshold; specifically, when the total points deducted are greater than a first threshold, issue a warning to the anchor; when the total points deducted are greater than a second threshold, order the anchor to stop broadcasting and make rectifications, etc.
[0083] The total deduction points of each anchor are sorted from high to low to obtain the anchor ranking, and based on the anchor ranking, the anchors ranked at the top (with more deductions and more violations) are warned.
[0084] Example 2
[0085] like Figure 6 As shown, the present invention provides a system for monitoring sensitive words in live broadcast speech, including a preprocessing module, a screening module, a text segment set generation module, a matching module, a detection module and an early warning module;
[0086] The preprocessing module is used to obtain live video and live audio, preprocess the live video to generate a set of face images with timestamps, and preprocess the live audio to generate a set of voice clips with timestamps;
[0087] The filtering module is used to filter specific expression images from the face image set with timestamps;
[0088] The text segment set generation module is used to generate a text segment set with a time stamp based on the speech segment set with a time stamp;
[0089] The matching module is used to match specific emoticon images with corresponding text fragments based on timestamps to generate speech text;
[0090] The detection module is used to detect whether the speech text contains sensitive words based on the sensitive word library. If sensitive words are detected, the anchor corresponding to the speech text will be deducted points based on the hit rate of sensitive words;
[0091] The early warning module is used to count the points deducted by each anchor to obtain the total points deducted by each anchor; when the total points deducted by the anchor are greater than the threshold, regulatory measures are taken against the anchor.
[0092] While the present invention has been described with reference to preferred embodiments, various modifications may be made and equivalent components may be substituted without departing from the scope of the present invention. In particular, the various technical features described in the various embodiments may be combined in any manner, provided no structural conflicts exist. The present invention is not limited to the specific embodiments disclosed herein, but encompasses all technical solutions within the scope of the claims. Any portions not described in detail herein are within the knowledge of those skilled in the art.
Claims
1. A method for monitoring sensitive words in live broadcasting, characterized in that: The following steps are involved: S100: Acquire live video and live audio, pre-process the live video to generate a set of face images with timestamps, and pre-process the live audio to generate a set of voice clips with timestamps; S2 00: Filter specific expression images from the time-stamped face image set, and generate a time-stamped text segment set based on the time-stamped voice segment set; S300: searching for a text segment corresponding to the specific expression image based on the timestamp, and generating a speech text; the speech text includes text corresponding to the negative expression; S400: Detecting whether the speech text contains sensitive words based on the sensitive word library. If sensitive words are detected, calculating the hit rate of the sensitive words, and deducting points from the anchor corresponding to the speech text based on the hit rate of the sensitive words; S500: Counting the points deducted from each anchor to obtain the total points deducted from each anchor, and supervising any anchor when the total points deducted from the anchor are greater than a threshold; The step S200 of filtering out specific expression pictures from the face picture set with timestamps includes the following sub-steps: S211: Recognize the expression of each face image and assign a corresponding expression label; S212: Filtering facial images with specific expression types based on expression tags, where the specific expression types include one or more pre-set expression types; and the specific expressions include negative expressions.
2. The sensitive word monitoring method according to claim 1, characterized in that: The pre-processing of the live video in step S100 includes: generating a set of face pictures with timestamps by slicing the live video according to time.
3. The sensitive word monitoring method according to claim 1, characterized in that: The pre-processing of the live audio in step S100 includes: generating a set of voice segments with timestamps for the live audio by time slicing.
4. The sensitive word monitoring method according to claim 1, characterized in that: The step S211 includes: The emotion-recognition algorithm is used to perform real-time expression recognition on facial images, and corresponding expression labels are assigned to each facial image according to different expression types.
5. The sensitive word monitoring method according to claim 1, characterized in that: The step S200 of generating a text segment set with a timestamp based on the speech segment set with a timestamp includes: An automatic speech recognition algorithm is used to convert each speech segment in the speech segment set into a corresponding text segment, thereby obtaining a text segment set with a timestamp.
6. The sensitive word monitoring method according to claim 5, characterized in that: It also includes corpus pre-training for the automatic speech recognition algorithm, including the following steps: Assign TF-IDF weights to the discourse data in the corpus to be trained; Sort all the assigned words by score from high to low to get the pre-training score of the speech technique; When automatic speech recognition encounters homophones, it matches them from high to low based on the pre-trained scores of the speech script.
7. The sensitive word monitoring method according to claim 1, characterized in that: The S400 includes: Match the words in the dialogue text based on the sensitive word library. If a sensitive word is detected, find the anchor corresponding to the dialogue text based on the timestamp corresponding to the dialogue text and the schedule of the live broadcast; count the hit rate of the sensitive words in the dialogue text, and deduct points from the anchor corresponding to the dialogue text based on the hit rate.
8. The sensitive word monitoring method according to claim 1, characterized in that: The method further includes S600: forming a sensitive word intelligent dashboard based on the statistical results of the sensitive word hit rate, and displaying the score deduction situation of each anchor in real time.
9. A system for monitoring sensitive words in live broadcasting, characterized in that: It includes pre-processing module, screening module, text segment set generation module, matching module, detection module and early warning module; The preprocessing module is used to obtain live video and live audio, preprocess the live video to generate a set of face pictures with timestamps, and preprocess the live audio to generate a set of voice clips with timestamps; The screening module is used to screen specific expression pictures from a collection of face pictures with timestamps; The text segment set generation module is used to generate a text segment set with a timestamp based on the speech segment set with a timestamp; The matching module is used to search for a text segment corresponding to the specific expression picture based on a timestamp and generate a speech text; the speech text includes text corresponding to a negative expression; The detection module is used to detect whether the speech text contains sensitive words based on the sensitive word library. If sensitive words are detected, the hit rate of the sensitive words is calculated, and the anchor corresponding to the speech text is deducted based on the hit rate of the sensitive words; The early warning module is used to count the points deducted by each anchor to obtain the total points deducted by each anchor, and supervise the anchor when the total points deducted by the anchor are greater than a threshold; The screening module performs the following operations: S211: Recognize the expression of each face image and assign a corresponding expression label; S212: Filtering facial images with specific expression types based on expression tags, where the specific expression types include one or more pre-set expression types; and the specific expressions include negative expressions.
Citation Information
Patent Citations
Live broadcast room violation detection method and device
CN112995696A
Word collection method matching emotion polarity
WO2021135140A1