Real-time silencing replacement method and device for live broadcast forbidden words and computer equipment

By real-time detection and using semantic replacement technology to process banned words in live broadcasts, the problem of banned word processing affecting content consistency and high cost in existing technologies is solved, and low-cost and efficient live broadcast banned word processing is achieved, which is suitable for a variety of Internet content dissemination scenarios.

CN120612942APending Publication Date: 2025-09-09CYPRESS SEED (HANGZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510512392.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing live broadcast banned word processing technology destroys content consistency and naturalness, lacks intelligent semantic replacement capabilities, is costly, and cannot meet the needs of small and medium-sized platforms and independent creators.

Method used

By obtaining the live broadcast platform industry selected by the user, loading the corresponding banned word library, combining it with the protection sensitivity set by the user, real-time detection and semantic replacement technology are used to process banned words, using virtual microphones, noise suppression and enhancement algorithms to convert them into text streams, using deep learning models and third-party speech recognition technology to detect banned words, and using bleep sound, silence or replacement word processing.

Benefits of technology

It achieves the fluency and continuity of live broadcast content, reduces costs, provides a low-cost and efficient banned word processing solution, and improves the content compliance of small and medium-sized platforms and independent creators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612942A_ABST
    Figure CN120612942A_ABST
Patent Text Reader

Abstract

The invention discloses a live broadcast forbidden word real-time silencing replacement method and device and computer equipment. The method comprises the steps of obtaining an industry to which a live broadcast platform selected by a user belongs; loading a corresponding prohibited word bank according to the industry; obtaining protection sensitivity and strict degree customized by a user to obtain protection strength information; acquiring live broadcast content, and converting the live broadcast content into a text stream; performing prohibited word detection on the text stream according to the protection strength information and a prohibited word library to obtain a detection result; when the detection result is that the forbidden word exists, processing the detection result by adopting a set rule to obtain a processing result; and presenting the real-time protection state, the processing result and the corresponding detection result in a visualized manner. By implementing the method provided by the invention, prohibited vocabularies can be eliminated, semantic replacement can be provided, the fluency and coherence of live broadcast contents are ensured, the cost is low, the application is easy, and the content compliance of small and medium-sized platforms and independent creators can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet content dissemination technology, and more specifically to a method, device and computer equipment for real-time silencing and replacing banned words in live broadcasts. Background Art

[0002] With the rapid development of the internet, live streaming has become a crucial form of digital content dissemination. Leveraging real-time audio and video streaming technology, live content has gained widespread adoption globally, encompassing social platforms, entertainment programs, education, and training, among other fields. While live streaming offers users greater opportunities for interaction and communication, it also carries a range of legal, ethical, and social risks, particularly regarding the use of sensitive or prohibited terms in live content.

[0003] While there are currently some solutions on the market for monitoring and silencing banned words in live broadcasts, most still have significant flaws in practical application. First, while some solutions can effectively identify banned words and mute them, they often use a crude approach, simply deleting or muting the audio segments containing banned words. This approach often disrupts the coherence and naturalness of the live broadcast content, affecting the viewer experience. In some cases, sudden muting can cause confusion among viewers, leading to misunderstandings of the host's true intentions and message, which in turn undermines user trust in the platform and content creators.

[0004] Another major issue with existing solutions is the lack of intelligent semantic replacement capabilities. While they can detect and address banned words, they don't offer the ability to replace them with the host's original voice. This means that even if banned words are removed, the coherence and fluency of the content cannot be guaranteed, and viewers may be dissatisfied with the interruptions. Furthermore, most systems fail to flexibly address the diversity of language expression and are unable to accurately determine the contextual meaning of certain utterances, sometimes misclassifying appropriate content as banned, further affecting the effectiveness of live broadcasts.

[0005] From a cost perspective, existing solutions are generally expensive, which poses a certain financial burden for small and medium-sized live streaming platforms and independent content creators. This has forced some small platforms and creators to abandon the use of these technologies when faced with content compliance requirements, thus affecting the compliance improvement of the entire industry.

[0006] Therefore, it is necessary to design a new method that can not only eliminate banned words but also provide semantic replacement to ensure the fluency and coherence of live broadcast content. It is low-cost and easy to apply, which will help improve the content compliance of small and medium-sized platforms and independent creators. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects of the prior art and provide a method, device and computer equipment for real-time silencing and replacing banned words in live broadcasts.

[0008] To achieve the above-mentioned purpose, the present invention adopts the following technical solution: a method for real-time silencing and replacing banned words in live broadcast, comprising:

[0009] Get the industry of the live streaming platform selected by the user;

[0010] Load the corresponding banned word library according to the industry;

[0011] Obtain user-defined protection sensitivity and strictness to obtain protection strength information;

[0012] Acquire live content and convert the live content into a text stream;

[0013] Performing a banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result;

[0014] When the detection result indicates that a banned word exists, the detection result is processed using a set rule to obtain a banned word processing result;

[0015] The real-time protection status, the banned word processing results and the corresponding detection results are presented visually.

[0016] A further technical solution is that the step of obtaining live content and converting the live content into a text stream includes:

[0017] Create a virtual microphone device through the Windows API to capture the audio stream of live content;

[0018] Converting the audio stream into a standard PCM format to obtain a conversion result;

[0019] Applying a noise suppression and enhancement algorithm to the conversion result to obtain a processing result;

[0020] The processing result is converted into a text stream using third-party speech recognition technology.

[0021] A further technical solution is: performing banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result, including:

[0022] A text matching algorithm is used to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result.

[0023] A further technical solution is: the set rules include at least one of bleep sound processing, silence processing, and replacement word processing.

[0024] A further technical solution is: the detection result is processed according to the set rules to obtain the processing result, including:

[0025] When using the word replacement method, the voice cloning model is used to convert the replacement word into a voice segment consistent with the host's voice to obtain a new voice segment;

[0026] Replacing the content corresponding to the detection result with the new voice segment, and performing a smooth transition between the audio segments using audio processing technology to obtain a processing result;

[0027] The voice cloning model is obtained by using personal audio samples as a sample set to train a deep learning model.

[0028] Its further technical solution is: the banned word library includes personal word library, enterprise word library and system word library. The personal word library is created and managed by the user through the front-end interactive interface, and words are added or deleted and classified; the enterprise word library is a banned word library customized according to industry needs; the system word library is the default banned word library maintained by the platform.

[0029] Its further technical solution is: the banned word library analyzes historical behaviors, violation events and big data through AI technology and machine learning algorithms, automatically identifies frequently appearing banned words and updates them.

[0030] The present invention also provides a device for real-time silencing and replacing banned words in live broadcast, comprising:

[0031] An industry acquisition unit, used to obtain the industry to which the live broadcast platform selected by the user belongs;

[0032] A loading unit, configured to load a corresponding banned word library according to the industry;

[0033] A protection strength acquisition unit is used to obtain user-defined protection sensitivity and strictness to obtain protection strength information;

[0034] A content conversion unit, configured to obtain live content and convert the live content into a text stream;

[0035] a banned word detection unit, configured to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result;

[0036] a processing unit, configured to, when the detection result indicates that a banned word exists, process the detection result using a set rule to obtain a banned word processing result;

[0037] The visual display unit is used to visually present the real-time protection status, the banned word processing results and the corresponding detection results.

[0038] The present invention further provides a computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.

[0039] The present invention also provides a storage medium, wherein the storage medium stores a computer program, and the computer program implements the above method when executed by a processor.

[0040] The beneficial effects of the present invention compared with the existing technology are: the present invention obtains the live broadcast platform industry selected by the user, loads the corresponding banned word library, and combines the protection sensitivity and strictness set by the user to intelligently detect banned words in the live broadcast content; after converting the live broadcast content into a text stream, the system monitors it in real time based on the set protection intensity information and the banned word library. After discovering banned words, it adopts semantic replacement and processing rules to ensure the fluency and coherence of the live broadcast content; through a visual interface, the protection status, processing results and detection feedback are displayed in real time, helping small and medium-sized platforms and independent creators to achieve content compliance in a low-cost and efficient manner without affecting the live broadcast experience.

[0041] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A flowchart of a method for real-time silencing and replacing banned words in live broadcasts provided by an embodiment of the present invention;

[0044] Figure 2 A schematic diagram of a sub-process of a method for real-time silencing and replacing banned words in live broadcasts provided by an embodiment of the present invention;

[0045] Figure 3 Schematic diagram of a sub-process of a method for real-time silencing and replacing banned words in live broadcasts provided by an embodiment of the present invention

[0046] Figure 4 A schematic block diagram of a device for real-time silencing and replacing banned words in live broadcasts provided by an embodiment of the present invention;

[0047] Figure 5 A schematic block diagram of a content conversion unit of a device for real-time silencing and replacing banned words in live broadcasts provided by an embodiment of the present invention;

[0048] Figure 6 A schematic block diagram of a processing unit of a device for real-time silencing and replacing banned words in live broadcasts provided by an embodiment of the present invention;

[0049] Figure 7 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0051] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0052] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0053] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0054] See also Figure 1 , Figure 1This is a schematic flow chart of a method for real-time silencing and replacing banned words in live broadcasts, provided by an embodiment of the present invention. This method is applied to a server. The server interacts with terminals and intelligently detects and replaces banned words in live broadcast content by combining a banned word library, speech recognition technology, and a deep learning model. In practice, the server first loads the banned word library corresponding to the user's selected industry and performs detection based on the user's set protection level. A virtual microphone captures the audio stream, which is converted into a text stream using noise suppression and enhancement algorithms for matching banned words. If a banned word is detected, it is processed according to predefined rules, such as bleeping, muting, or replacing it with a new word. In particular, voice cloning technology is used to convert the replacement word into a voice clip consistent with the host's voice, achieving seamless replacement. A personalized, industry-customized, and platform-maintained banned word library, combined with AI technology, automatically updates the library to ensure content compliance. This method not only eliminates banned words but also provides semantic replacement, maintaining the fluency and coherence of live broadcast content. Its low cost and ease of application help small and medium-sized platforms and independent creators improve content compliance.

[0055] This method is mainly used in the Internet content dissemination industry, especially in scenarios involving real-time audio and video streaming. The system can serve various live broadcast platforms, including but not limited to the following:

[0056] Online entertainment platforms: such as live game broadcasts, talent show broadcasts, and sports events. These platforms need to ensure that their live broadcast content complies with local laws, regulations, and social ethics.

[0057] Education and training platforms: For live events such as online courses, seminars, or lectures, this system can effectively prevent inappropriate speech and ensure the professionalism and seriousness of educational content.

[0058] Business meetings and webinars: Avoid leaking sensitive information or inappropriate remarks that could damage the company's image during enterprise-level online communications.

[0059] Social media platforms: especially service platforms dominated by UGC (user-generated content), maintain platform order by automatically detecting and processing language that may violate community norms.

[0060] Government agencies and public utilities: When releasing official information, ensure that live broadcast content is accurate and complies with policy requirements, and avoid misinformation or inappropriate remarks.

[0061] News broadcasting: During the digital transformation process, traditional media can also use this technology to conduct real-time review of sudden remarks during live news broadcasts to ensure the accuracy and compliance of information transmission.

[0062] This method uses intelligent speech recognition technology and deep learning algorithms to detect and process banned words in live broadcasts in real time, ensuring content compliance and helping platforms effectively manage and optimize the quality of their live content.

[0063] Figure 1 Schematic diagram of the process of real-time silencing and replacing banned words in live broadcast provided by the embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S150.

[0064] S110. Obtain the industry to which the live broadcast platform selected by the user belongs.

[0065] In this embodiment, the user needs to verify his / her identity through the login interface and enter the main interface of the system after logging in. On this interface, the user can select the live broadcast platform to which he / she belongs and further select the industry category of the platform.

[0066] The front-end page provides a drop-down box or multiple-selection box for industry selection, allowing users to select the corresponding industry (such as education, finance, entertainment, etc.) according to their needs.

[0067] After the system receives the industry information selected by the user, it will store it according to the selected data and prepare for subsequent loading of the banned word library. The system needs to associate the industry information with the user's account information and ensure that each account can provide an accurate banned word library based on the selected industry.

[0068] The front end uses technologies such as HTML, CSS, and JavaScript to implement interface interaction, ensuring that the system can respond promptly and provide industry-related options after the user selects an industry.

[0069] The system will receive the industry information selected by the user through the API interface, store it in the database, and perform logical processing to ensure that the correct banned word library can be loaded in the next step.

[0070] S120: Loading a corresponding banned word library according to the industry.

[0071] In this embodiment, the banned word library refers to a collection of words containing sensitive or non-compliant content, which is used to filter, monitor and prevent the spread of inappropriate information.

[0072] Specifically, the banned word library includes personal word library, enterprise word library and system word library. The personal word library is created and managed by the user through the front-end interactive interface, adding or deleting words and classifying words; the enterprise word library is a banned word library customized according to industry needs; the system word library is the default banned word library maintained by the platform.

[0073] The banned word library uses AI technology and machine learning algorithms to analyze historical behaviors, violation events and big data, and automatically identifies frequently appearing banned words for updating.

[0074] Specifically, in this step, the system will load the banned word library that matches the industry selected by the user. The banned word library includes three categories:

[0075] Personal vocabulary: A library of banned words that users can customize and manage through the interface. Users can add, delete, and categorize banned words based on their needs.

[0076] Enterprise Dictionary: A customized dictionary based on the industry needs of the enterprise. Enterprises can upload banned words for specific industries and manage them according to regulations and industry trends.

[0077] System vocabulary: A banned word library uniformly maintained by the platform, covering a wide range of banned words in various industries. The platform updates it regularly and automatically pushes it to the user system.

[0078] In this embodiment, the corresponding banned word library is loaded in the following ways:

[0079] Based on the industry selected by the user, the appropriate banned word library is searched and loaded from the predefined industry database.

[0080] During the loading process, if the industry selected is relatively special (such as an emerging industry), the system may dynamically load the relevant banned word library based on the specific needs of the industry.

[0081] If the user selects the "Education" industry, the system will load a vocabulary containing sensitive words specific to the education industry. For the "Finance" industry, the system will load a sensitive vocabulary related to finance.

[0082] The industry information selected by the user will be passed to the backend through API calls. The backend will return the corresponding banned word library based on the industry information and display it on the front-end interface.

[0083] The backend queries the preset banned word library from the database according to the industry type, loads the corresponding personal word library, corporate word library or system word library, and finally binds it to the protection engine.

[0084] The storage of banned word databases is managed through a database (such as MySQL, PostgreSQL, etc.). The banned word databases for each industry are classified and stored in different tables or databases. The system queries and loads the corresponding data based on user selections.

[0085] The system will regularly obtain the latest banned word information from external regulations, industry trends and other data sources, and automatically push updates to the user's protection system.

[0086] Enterprise users and individual users can also manually upload new banned words to ensure that the vocabulary content is updated in real time to respond to changing regulations and industry requirements.

[0087] By implementing the above steps, the system can provide a matching banned word library based on the user's selected industry, ensuring compliance with relevant laws and industry standards during live broadcasts. Furthermore, users can customize their vocabulary through personalized word library management, maintaining the timeliness and effectiveness of the vocabulary based on industry trends and regulatory updates. This multi-level, multi-dimensional vocabulary library management model enables the system to adapt to the needs of different users and improve the accuracy and efficiency of compliance management.

[0088] S130: Obtain the user-defined protection sensitivity and strictness to obtain protection strength information.

[0089] In this embodiment, protection level information refers to the protection level selected by the user through system settings, reflecting the system's rigor and accuracy in handling sensitive information. Users can adjust the sensitivity of protection using a slider or settings within the interface, and the system will adjust its banned word recognition and handling strategies accordingly.

[0090] Specifically, protection sensitivity refers to the accuracy and strictness with which the system identifies sensitive information. The system offers several levels for users to choose from, such as low, medium, and high sensitivity. Each level affects the intensity of sensitive information filtering.

[0091] Low sensitivity: The system will moderately relax the recognition of banned words and only filter out words that are obviously non-compliant or dangerous. It is suitable for some scenarios with low requirements for information flow.

[0092] Medium sensitivity: The system will conduct more stringent checks and filter out more types of sensitive words. It is suitable for industries and scenarios that require a certain degree of content supervision.

[0093] High sensitivity: The system will strictly check almost all possible sensitive words and is suitable for industries with extremely high compliance requirements, such as finance, healthcare, etc.

[0094] The strictness of protection refers to the system's response behavior when sensitive information is discovered, and is mainly divided into the following categories:

[0095] Warning: When the system identifies mildly sensitive content, a warning notification pops up to remind users that inappropriate information may exist, but no immediate enforcement action is taken.

[0096] Marking: The system marks sensitive words, which users can review and handle manually. This is suitable for medium sensitivity settings.

[0097] Automatic blocking / deletion: The system automatically handles illegal content, blocking or deleting sensitive words, suitable for high-sensitivity settings.

[0098] Dynamically adjust sensitivity based on industry and scenario:

[0099] Based on the industry or scenario selected by the user, the system automatically loads a library of banned words related to that scenario and adjusts the recognition strength of the library based on the protection sensitivity set by the user. For example, if a user selects the financial industry and sets a high sensitivity, the system will not only load the banned words specific to the financial industry but also strictly filter content related to financial fraud, insider trading, etc.

[0100] Users set the sensitivity (low, medium, high) through a slider, drop-down box, or checkbox, and select the relevant scenario or industry (such as education, finance, healthcare, etc.) through the interface. These settings will interact with the back-end interface through the interface.

[0101] The backend receives the user's sensitivity setting and loads the corresponding word filtering logic based on the selected sensitivity. For example, high sensitivity may involve more in-depth analysis and add multiple rules for content review.

[0102] The system manages its vocabulary through a tagging system that combines industry, scenario, and sensitivity tags. Each tag corresponds to different vocabulary and filtering rules. For example, the financial industry might be associated with terms like "financial fraud" and "insider trading," while the education industry might be associated with terms like "exam cheating" and "violations."

[0103] When a user selects a scenario or industry, the system automatically loads the corresponding vocabulary template based on the scenario tag and adjusts the vocabulary filtering rules based on the selected protection sensitivity. For example, in high sensitivity mode, the system may additionally load certain uncommon but potentially risky sensitive words.

[0104] The system stores the user's protection settings (such as sensitivity and strictness) in the database and associates them with the user's account information. Whenever the user logs in, the system automatically reads and applies the previous settings.

[0105] Depending on the level of protection, the system will adjust content filtering policies and rules based on specific needs. If the user sets a high sensitivity, the system will use a more refined banned word filtering engine to improve accuracy.

[0106] The system regularly updates industry lexicons based on the latest regulations, industry trends, and user feedback. For example, the financial industry may need to adjust its lexicon content as anti-fraud technology and policies are updated.

[0107] The system not only supports automated updates, but also allows users or businesses to add or delete specific banned words based on actual needs, ensuring the flexibility and timeliness of the vocabulary.

[0108] This shows that protection intensity information reflects the user's customized needs for sensitive information protection, including protection sensitivity (low, medium, high) and strictness (warning, marking, blocking). Based on the user's selection of scenarios and industries, the system dynamically loads matching vocabulary libraries and adjusts filtering strategies based on the set sensitivity and strictness. In technical implementation, the labeling system and vocabulary management ensure the accuracy and industry adaptability of protection measures.

[0109] S140: Acquire live content and convert the live content into a text stream.

[0110] In this embodiment, the live broadcast content refers to the audio stream during the live broadcast; the text stream refers to the content converted from the live broadcast content to text.

[0111] In this step, the goal is to extract information from the live audio stream and convert it into a text stream for subsequent sensitive word detection and other operations. This process typically involves multiple technical steps and processing links to ensure the real-time and accurate conversion.

[0112] In one embodiment, see Figure 2 , the above-mentioned step S140 may include steps S141 to S144.

[0113] S141. Create a virtual microphone device through the Windows API to capture the audio stream of the live content.

[0114] In this embodiment, in a real-time live broadcast scenario, it is first necessary to capture the audio stream of the live broadcast. To achieve this, a virtual microphone device can be created through the Windows API. The virtual microphone is equivalent to a simulated hardware device that can capture the audio stream from the live content.

[0115] Using interfaces provided by the operating system (such as the Windows API), software can simulate an audio input device, allowing the audio stream received by the system to originate from the audio output of the live broadcast platform or application. This allows the audio data emitted during the live broadcast to be captured in real time and transmitted to subsequent processing steps.

[0116] S142: Convert the audio stream into a standard PCM format to obtain a conversion result.

[0117] In this embodiment, the conversion result refers to the result obtained after the audio stream is converted into a format.

[0118] After capturing the audio stream, the system needs to convert the audio data. The audio stream may exist in some compressed format (such as MP3 or AAC), but for subsequent processing, especially speech recognition, it usually needs to be converted to the standard PCM (Pulse Code Modulation) format, which is a common lossless audio format that facilitates subsequent processing and analysis.

[0119] Audio data in PCM format preserves the original sound signal, enabling more accurate speech recognition. Furthermore, the standard PCM format is well compatible with other speech processing algorithms.

[0120] S143: Apply a noise suppression and enhancement algorithm to the conversion result to obtain a processing result.

[0121] In this embodiment, since the live broadcast environment may contain background noise, echo, audio distortion and other problems, the audio data must be processed to improve the accuracy of speech recognition. To this end, noise suppression algorithms and audio enhancement technologies can be used.

[0122] Noise suppression: Use algorithms to filter out irrelevant noise components in the audio (such as background sounds, noise other than human voices, etc.) to retain clear speech signals.

[0123] Enhancement algorithm: After noise suppression, speech enhancement algorithm can be further applied to make the speech signal clearer, reduce audio quality issues caused by environmental factors, and ensure that subsequent speech recognition is more accurate.

[0124] S144: Use third-party speech recognition technology to convert the processing result into a text stream.

[0125] In this embodiment, the processed audio stream (i.e., the noise-suppressed and enhanced PCM audio data) is converted into a text stream using third-party speech recognition technology (such as Google Speech-to-Text or the Microsoft Azure Speech API). This process is crucial because only accurate speech recognition can convert live broadcast content into text.

[0126] These technologies can recognize speech signals in real time and convert them into text streams corresponding to the audio content. The text streams will be used for sensitive word detection and other analysis tasks in subsequent steps.

[0127] Once the audio stream is converted into a text stream, the system will use efficient text matching algorithms such as AC automata to perform real-time sensitive word detection on the text. This means that each recognized text segment will be compared to see if it contains any banned words.

[0128] If banned words are detected, the system will immediately perform replacement, warning, or blocking actions to ensure the compliance of the live broadcast content.

[0129] To ensure high efficiency and low latency, the entire sensitive word detection process is usually performed asynchronously so as not to affect the smoothness of the live broadcast content.

[0130] The application creates a virtual microphone device by using the Windows API so that it can be used as an audio source after installation.

[0131] Through the virtual microphone device, the system can capture all audio data, whether it is from the user's voice or other sounds, ensuring that all audio that needs to be processed is obtained in real time.

[0132] The captured audio data will be converted into a standard audio format (such as PCM) for subsequent processing.

[0133] To improve speech recognition accuracy, before the audio stream enters the automatic speech recognition (ASR) module, the system applies noise suppression and speech enhancement technologies to remove ambient noise and enhance the clarity of the speech signal.

[0134] The system uses third-party speech recognition technology to convert the captured and processed audio into text. The text is then checked for banned words.

[0135] If banned words are detected, the system will take appropriate measures based on pre-set rules, such as silencing or replacing sensitive words.

[0136] The system uses the FFmpeg framework to synthesize the processed audio, ensuring that the replaced or modified audio stream remains coherent and natural.

[0137] The processed audio will be redirected to the live broadcast assistant software (such as TikTok Live Companion, etc.) for users to broadcast live.

[0138] Users will use these processed audios for live broadcasts to ensure that the live content complies with regulatory requirements.

[0139] During the entire process of audio capture, processing, and redirection, system optimization technology minimizes delays as much as possible to ensure real-time performance and avoid affecting the user's interactive experience.

[0140] All operations are performed in real time, and the delay of the entire processing process is controlled within 3 seconds, ensuring that there will be no significant delay during the live broadcast and guaranteeing user experience.

[0141] The system optimizes audio processing and voice recognition processes to ensure real-time detection and effective processing of prohibited content, thereby ensuring the compliance of live broadcast content while providing users with a smooth live broadcast experience.

[0142] The above steps (S141-S144) describe the complete process from capturing the live audio stream to converting the speech into a text stream. By combining a virtual microphone device, PCM format conversion, noise suppression and enhancement algorithms, and third-party speech recognition technology, the resulting text stream will become the basis for subsequent sensitive word detection and other content processing. This process requires special attention to processing latency and accuracy to ensure that banned words are identified and processed promptly and effectively during the live broadcast, while ensuring that the user experience is not affected.

[0143] S150: Perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result.

[0144] In this embodiment, the detection result refers to whether the current text stream contains banned words.

[0145] Specifically, a text matching algorithm is used to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result.

[0146] In this embodiment, during the live broadcast, the system needs to perform banned word detection on the real-time generated text stream to ensure that the text content meets regulations and compliance requirements. Step S150 describes how to perform this banned word detection based on the protection level information and the banned word library.

[0147] The protection strength information determines how the system handles banned words and the strictness of the detection. In practical applications, the protection strength can be divided into different levels, for example:

[0148] Low protection: Allows some less sensitive banned words to exist, and only processes content that clearly violates regulations.

[0149] Medium protection: Detects and processes banned words of moderate sensitivity, while also filtering less serious content.

[0150] High protection: All possible banned words are strictly detected and processed to ensure that every text segment meets compliance requirements.

[0151] The system adjusts the banned word detection strategy based on real-time protection strength information, thereby adopting different sensitivity processing in different scenarios.

[0152] A banned word library is a collection of banned words, against which the system compares text streams. This library typically contains sensitive words, illegal words, and potentially controversial terms. For example, it includes politically sensitive words, violent, pornographic, or racist language.

[0153] The banned word library should be kept dynamically updated to cope with new banned words and potential changes. A real-time updated banned word library can help the system identify newly emerging banned content in a timely manner.

[0154] In a multilingual environment, the banned word library should support multiple language versions and be able to automatically adjust according to the language settings of the user's region.

[0155] In order to achieve efficient banned word detection, the system usually uses an efficient text matching algorithm, AC automaton (Aho-Corasick automaton), to perform fast banned word matching.

[0156] Working principle of AC automaton:

[0157] Construct a dictionary tree: Organize all banned words in the banned word library into a dictionary tree (Trie tree). Each node represents a character or part of a word, and the path from the root to a node forms a banned word.

[0158] Constructing a mismatch table: The AC automaton also constructs a "mismatch table" that allows the algorithm to quickly jump to other appropriate states when encountering mismatched characters during the matching process, rather than starting the comparison from scratch.

[0159] Real-time detection: When a text stream is input, the AC robot will match each character in the dictionary tree and quickly identify whether there are any banned words. If a matching banned word is found, the system can mark it as banned content and take appropriate measures.

[0160] The advantage of the AC automaton is that its time complexity for character-by-character matching in a text stream is O(n), where n is the length of the text. This avoids the performance issues that may arise from traditional brute-force matching algorithms (such as brute-force comparison). Therefore, it is very suitable for real-time detection scenarios.

[0161] In step S150, the system will detect the text stream based on the protection strength information and the banned word library, combined with the AC automaton algorithm. Specifically, the detection process can be divided into the following steps:

[0162] The system obtains a real-time generated text stream from the real-time speech recognition module.

[0163] Perform AC automaton matching on each word in the text stream to find out whether there are any banned words.

[0164] Based on the current protection level information, the system decides whether to process certain banned words. For example, at a low protection level, only obvious banned words may be processed; at a high protection level, all banned words will be strictly detected.

[0165] The detection result indicates whether the current text stream contains banned words. For each text segment, the system will output an indicator based on the matching results, indicating whether the text segment contains banned words.

[0166] Presence of banned words: If banned words are found in the text, the detection result will show the "banned" status, and the system will trigger the predetermined processing action (such as replacement, warning or blocking).

[0167] No banned words detected: If the text stream does not contain banned words, the detection result is "Compliant" and the system will continue to process the next text stream.

[0168] Based on the detection results, the system will take appropriate actions:

[0169] Replacement: The system can replace banned words with legal characters (such as asterisks "*" or other symbols) to prevent banned words from appearing in the text.

[0170] Warning: If banned words appear in the text, the system may send a warning to the user, indicating that the content is non-compliant.

[0171] Blocking: In some cases, if the occurrence of banned words is serious, the system may directly block the text to ensure that the banned content cannot be spread.

[0172] To ensure that banned word detection doesn't disrupt the smoothness of live broadcast content, the system typically employs an asynchronous processing mechanism. Banned word detection and subsequent processing of the text stream occur in parallel in the background, reducing real-time detection latency. This ensures that the user's viewing experience isn't significantly disrupted during the live broadcast.

[0173] If banned words are detected but no action is taken, the system will completely skip any action on banned words and allow the text to pass through unmodified. This means that the system will not filter or modify banned words. This is suitable for scenarios that do not require real-time intervention or when users want to handle banned words themselves. Technically, the system will ignore banned word detection and pass the text stream directly without any form of processing.

[0174] If interval processing is selected, this option allows users to control the processing of banned words based on specific conditions. There are mainly the following strategies:

[0175] Time interval processing: The system can set a time window (for example, every 10 seconds). If the same banned word is detected multiple times during this time period, the system will merge them into a single event for processing. This can avoid triggering processing actions too frequently and reduce interference with real-time content.

[0176] Trigger Count: The system monitors the number of times a banned word appears within a specified time period. Action is only taken when a banned word's trigger count reaches a preset threshold. This prevents overreaction to occasional banned words, allowing only frequent ones to be addressed.

[0177] Scenario Processing: Based on different application scenarios (such as live broadcast rooms, chat rooms, etc.), the system allows users to customize the rules and strictness of banned word processing. For example, in some scenarios, more relaxed processing standards may be allowed, thereby reducing the frequency of banned word processing and avoiding excessive intervention.

[0178] Step S150 combines protection level information with a banned word library, employing a highly efficient AC automaton algorithm to perform real-time banned word detection on the text stream. This process not only ensures real-time performance but also ensures the compliance of the text content. Appropriate measures can be taken based on the level of protection to ensure that live content complies with regulations.

[0179] S160: When the detection result indicates that a banned word exists, the detection result is processed using a set rule to obtain a banned word processing result.

[0180] In this embodiment, the banned word processing result refers to the result after the banned words are processed using the set rules.

[0181] Specifically, the set rules include at least one of bleep sound processing, silence processing, and replacement word processing.

[0182] Specifically, in step S160, the detection system first detects banned words in the audio stream using audio analysis technology (such as automatic speech recognition (ASR)). Once banned words are detected, the system will perform corresponding processing according to the set rules to prevent the spread of banned content. The processing method can be one of the following three methods, or a combination of them:

[0183] Bleep sound processing: When a banned word is detected, the system replaces it with a specific sound effect called a "beep." Using an audio processing library like FFmpeg, the audio stream is processed in real time, inserting the "beep" sound to mask the original banned word.

[0184] Silence processing: When banned words appear, the system will use audio processing technology (such as streaming audio processing) to replace the banned word portion with completely silent audio. At this time, the rest of the audio stream remains unchanged, and only the banned word portion is silenced. Through streaming audio processing, the system will directly generate a silence effect at the time when the banned word is detected, maintaining synchronization with other audio streams. Silence processing involves replacing audio frames, ensuring that the replaced area is consistent with the original audio length, so as not to affect the overall time progress of the audio.

[0185] Word replacement: This is the most complex method. The system replaces banned words with pre-defined alternatives and uses voice cloning technology (such as CosyVoice) to generate voice clips that are consistent with the host's original voice, ensuring the replacement's pronunciation style matches the host's voice. Ultimately, these replacements are inserted into the audio stream, replacing the banned words.

[0186] In one embodiment, see Figure 3 , the above-mentioned step S160 may include steps S161 to S162.

[0187] S161. When using the word replacement processing method, use the voice cloning model to convert the replacement word into a voice segment consistent with the host's voice to obtain a new voice segment;

[0188] S162: replacing the content corresponding to the detection result with the new voice segment, and performing a smooth transition between the audio segments using an audio processing technology to obtain a processing result;

[0189] The voice cloning model is obtained by using personal audio samples as a sample set to train a deep learning model.

[0190] In this embodiment, the audio stream of the replacement word is generated in advance and stored in the audio library. Whenever the system detects a banned word, the corresponding replacement voice is retrieved from the library and inserted into the audio stream.

[0191] Speech snippets for these replacement words are stored in a local cache to improve real-time processing efficiency and reduce the computational burden of real-time speech generation. When a banned word is detected, the system uses the host's voice synthesis to generate speech to replace the banned word, maintaining semantic coherence. Users can choose whether to enable this feature.

[0192] The system clones the host's voice using a deep learning model (such as CosyVoice) based on a user-provided audio sample set. The voice cloning model uses the host's audio samples to analyze voice characteristics such as pitch, speaking speed, and emotional expression.

[0193] The voice cloning model generates a new voice clip based on the host's voice characteristics, with the replacement word sounding like the host's original words. To ensure this effect, the system may perform detailed adjustments, including voiceprint analysis and synchronization of emotional expression, to ensure that the replacement word pronunciation is not only accurate but also natural.

[0194] Furthermore, model training and updates are ongoing. To ensure the adaptability and accuracy of the cloning technology, the system regularly updates the host's audio samples to ensure the replacement voice remains consistent.

[0195] The system identifies the audio segments that need to be replaced based on the specific locations of the banned words it detects. Next, the system replaces the banned words with new audio segments generated using the voice cloning model. During this process, ensuring that the replacement audio segments match the duration and rhythm of the original audio stream is crucial to avoid time misalignment or audio interruptions.

[0196] After the replacement, the system applies audio processing techniques (such as crossfading and spectrum matching) to ensure a smooth transition between the old and new audio segments. This technique can reduce auditory abruptness and make the overall audio flow more coherent.

[0197] Finally, the processed audio is output and adjusted to the format and quality required by the platform, ensuring that the final audio effect can be seamlessly integrated into the original audio stream.

[0198] Users need to provide their own audio samples, which will be used as part of the training set to train the voice cloning model. To ensure the best results, the system will provide detailed recording guidelines, including recording in a quiet environment and maintaining a consistent volume and speaking speed.

[0199] The system uses deep learning frameworks (such as CosyVoice) to model the host's voice. The model learns features such as voice timbre, intonation, emotion, and speaking rate. To ensure more accurate voice cloning, the system also regularly updates the host's audio samples to adapt to changes in their voice (for example, changes in their health or mood).

[0200] The voice cloning model not only accurately generates a voice similar to the host's, but also handles complex emotional and intonation changes to ensure the replacement voice sounds natural. Through this deep learning model, the replacement voice clips maintain a coherent and natural feel.

[0201] Step S160 primarily handles the actions taken after banned word detection. The replacement process involves using a voice cloning model to replace banned words with alternative voices that are consistent with the host's voice. This process includes multiple substeps, including training the voice cloning model, generating new voice segments, replacing banned words in the audio stream, and ensuring a smooth transition between the old and new audio segments. Through deep learning, audio processing, and voice cloning technologies, the system ensures that the final audio output meets platform requirements while maintaining authenticity and preventing the spread of banned content.

[0202] S170: Visually presenting the real-time protection status, the banned word processing result, and the corresponding detection result.

[0203] In this embodiment, a visual interface is provided to the user, displaying the protection status in real time, including information such as banned words that have been processed and the current protection level. The front-end interface interacts with the back-end data via WebSocket or a real-time API, updating data such as the protection status, banned words that have been replaced, and the number of triggers in real time.

[0204] Each time a banned word event is triggered, the system will record a detailed log, including the banned word content, handling method, and processing time. All log data is stored in a distributed database for easy subsequent query and audit. The log content includes information such as user ID, banned word, handling method, and timestamp.

[0205] The system provides a cross-platform PC client that is suitable for operating systems such as Windows and Mac. Users can download and install the client to configure and manage protection settings and enjoy comprehensive functions.

[0206] For mobile devices (iOS and Android), the system provides an adapted mobile app, which can be downloaded from the App Store and Google Play. This app synchronizes the functionality of the desktop version and optimizes the interface and interaction to ensure a smooth mobile experience.

[0207] When users encounter problems, the system provides remote technical support. By integrating remote assistance software (such as TeamViewer, AnyDesk, etc.), users can directly request technical support personnel to connect remotely and solve problems in real time.

[0208] The system has a built-in user feedback function, allowing users to submit questions and suggestions during use. All feedback records will be automatically sent to the product team for regular analysis and improvement to enhance product quality and user experience.

[0209] Users are encouraged to invite others to use the product, and the inviter can receive rewards. By generating a unique invitation code, users can invite others to register and use the product. The system will track the use of each invitation code and give rewards based on the number of invitees.

[0210] Through the referral mechanism, users can recommend new users to register and pay, and earn commission rewards. The system will generate a unique referral link or invitation code for each referrer, track their conversion rate, and distribute commissions according to the set commission ratio.

[0211] These functional modules constitute an efficient and comprehensive real-time protection method for silencing banned words in live broadcasts, aiming to improve the security and compliance of live broadcast content while ensuring a smooth and natural user experience.

[0212] This embodiment's method and product not only detects banned words in live broadcasts in real time, but also uses advanced speech synthesis technology to intelligently replace these words with the host's original voice, ensuring the smoothness and integrity of the live broadcast content. Unlike direct muting, this product maintains the natural flow of conversations and avoids confusion or dissatisfaction among viewers caused by sudden silences.

[0213] The method of this embodiment supports both computer and mobile versions, allowing users to flexibly use it on different devices, improving the user experience. This enhances user convenience and enables the system to meet different live broadcast needs in a variety of scenarios. Whether on a computer or a mobile phone, users can enjoy a consistent user experience.

[0214] Users can select the appropriate protection strength and customize their own protection strategies based on different live streaming platforms and industry requirements. This provides a high degree of customization to meet specific user needs and improve protection accuracy and flexibility.

[0215] The method of this embodiment provides an intuitive visual protection interface, allowing users to clearly view the current protection status and historical records. The intuitive display of protection results helps users better understand and control the operating status of the system, improving the user experience.

[0216] The method of this embodiment generates a detailed word search report to help users understand the situation and treatment results of banned words. The report facilitates users to conduct subsequent analysis and audits, ensuring the compliance of live content and enhancing management transparency.

[0217] The method of this embodiment provides remote technical support services and collects user feedback to continuously improve the product, enhance user satisfaction, ensure that technical issues are resolved quickly, and promote continuous product optimization through feedback mechanisms.

[0218] For example: On a live streaming platform (such as game live streaming, education live streaming, or news live streaming), using this product for real-time silencing of banned words can significantly improve the compliance of live streaming content while ensuring the audience experience.

[0219] Platform administrators first log in to the system, select an appropriate industry, and set the protection level. For example, let's say the administrator selects "Online Education" as the platform industry and selects "Strict" in the "Protection Level Settings" section. The system will then be customized based on the selected industry and protection level, ensuring accurate identification and processing of banned words. This ensures that the platform's content is more compliant with industry regulations and ensures compliance in the live streaming environment.

[0220] During live broadcasts, the system monitors the host's voice content in real time and automatically identifies banned words. For example, if a host mentions the banned word "cheat" during a live broadcast, the system immediately detects the word and initiates an intelligent replacement mechanism. To maintain semantic consistency and compliance, the system replaces "cheat" with "unfair means," ensuring no interruption to the live broadcast and no impact on the viewer experience. This intelligent replacement function effectively reduces the impact of banned content on the smoothness of the live broadcast.

[0221] The system provides a clear, visual protection interface, allowing administrators to view the current live broadcast's protection status and historical records at any time. In the "Visual Protection" module, administrators can view a list of banned words that have been processed and their replacements, providing real-time visibility into protection progress. If any issues arise, administrators can make immediate adjustments to ensure smooth and efficient monitoring. This visual management approach effectively enhances the platform's real-time protection capabilities.

[0222] The system automatically generates detailed word search reports to help administrators understand the specific circumstances of banned words and the results of their handling. In the "Word Search Report" module, administrators can view all detected banned words and their specific handling methods. These detailed reports allow administrators to conduct subsequent content audits to ensure that live content strictly complies with relevant compliance requirements, and they can also conduct necessary analysis and optimization to improve the accuracy of protection.

[0223] The system also offers remote technical support. Administrators can contact the technical support team through the "Remote Assistance" module when encountering problems. The technical support team can quickly respond and resolve issues. Furthermore, the system collects administrator feedback to continuously optimize and enhance product functionality. This immediate technical support and feedback collection mechanism ensures continuous system improvement and increased user satisfaction.

[0224] The system supports seamless switching between desktop and mobile versions, allowing users to enjoy the same protection experience across multiple devices. For example, while a live streamer is broadcasting on their mobile phone, the system can still provide real-time monitoring and banned word processing, ensuring a consistent user experience across devices. This cross-platform compatibility enhances product flexibility and meets the needs of users on different platforms.

[0225] The system uses big data analysis and monitoring to update its banned word database in real time to ensure it remains responsive to the ever-changing live content landscape. Administrators don't need to manually update the database; the system regularly retrieves the latest banned word information from the "Banned Big Data" module and automatically updates the database. This ensures the system remains current on the latest illegal words, further enhancing the timeliness and accuracy of its protection.

[0226] Through the introduction of these embodiments, we can see the comprehensive advantages of this product in actual applications. It can not only help live broadcast platforms effectively maintain content compliance, but also provide a smooth and personalized live broadcast experience, meeting the multiple needs of platforms, anchors and viewers.

[0227] The above-mentioned real-time silencing and replacement method for banned words in live broadcasts obtains the live broadcast platform industry selected by the user, loads the corresponding banned word library, and intelligently detects banned words in the live broadcast content in combination with the protection sensitivity and strictness set by the user; after converting the live broadcast content into a text stream, the system conducts real-time monitoring based on the set protection intensity information and banned word library. After discovering banned words, it adopts semantic replacement and processing rules to ensure the fluency and coherence of the live broadcast content; through a visual interface, it displays the protection status, processing results and detection feedback in real time, helping small and medium-sized platforms and independent creators to achieve content compliance in a low-cost and efficient manner without affecting the live broadcast experience.

[0228] Figure 4 3 is a schematic block diagram of a real-time silencing and replacement device 300 for live broadcast banned words provided by an embodiment of the present invention. Figure 4 As shown, corresponding to the above live broadcast banned word real-time mute replacement method, the present invention also provides a live broadcast banned word real-time mute replacement device 300. The live broadcast banned word real-time mute replacement device 300 includes a unit for executing the above live broadcast banned word real-time mute replacement method, and the device can be configured in a server. Specifically, please refer to Figure 4 The live broadcast banned word real-time silencing replacement device 300 includes an industry acquisition unit 301, a loading unit 302, a protection strength acquisition unit 303, a content conversion unit 304, a banned word detection unit 305, a processing unit 306 and a visual display unit 307.

[0229] The industry acquisition unit 301 is used to obtain the industry to which the live broadcast platform selected by the user belongs; the loading unit 302 is used to load the corresponding banned word library according to the industry; the protection strength acquisition unit 303 is used to obtain the user-defined protection sensitivity and strictness to obtain protection strength information; the content conversion unit 304 is used to obtain the live broadcast content and convert the live broadcast content into a text stream; the banned word detection unit 305 is used to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result; the processing unit 306 is used to process the detection result using the set rules when the detection result shows that banned words exist, so as to obtain a banned word processing result; the visual display unit 307 is used to use visualization to present the real-time protection status, the banned word processing result and the corresponding detection result.

[0230] In one embodiment, if Figure 5 As shown, the content conversion unit 304 includes a creation subunit 3041 , a format conversion subunit 3042 , a pre-processing subunit 3043 and a conversion subunit 3044 .

[0231] The creation subunit 3041 is used to create a virtual microphone device through the Windows API to capture the audio stream of the live content; the format conversion subunit 30443042 is used to convert the audio stream into a standard PCM format to obtain a conversion result; the preprocessing subunit 3043 is used to apply noise suppression and enhancement algorithms to the conversion result to obtain a processing result; the conversion subunit 3044 is used to convert the processing result into a text stream using third-party speech recognition technology.

[0232] In one embodiment, the banned word detection unit 305 is configured to perform banned word detection on the text stream using a text matching algorithm according to the protection strength information and the banned word library to obtain a detection result.

[0233] In one embodiment, the processing unit 306 includes a conversion subunit 3061 and a replacement subunit 3062 .

[0234] The conversion subunit 3061 is used to convert the replacement word into a voice segment consistent with the host's voice using the voice cloning model when using the replacement word processing method to obtain a new voice segment; the replacement subunit 3062 is used to replace the content corresponding to the detection result with the new voice segment and perform a smooth transition between audio segments using audio processing technology to obtain a processing result;

[0235] The voice cloning model is obtained by using personal audio samples as a sample set to train a deep learning model.

[0236] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned live broadcast banned word real-time silencing replacement device 300 and each unit can refer to the corresponding description in the aforementioned method embodiment. For the convenience and brevity of description, it will not be repeated here.

[0237] The above-mentioned live broadcast banned word real-time silencing replacement device 300 can be implemented in the form of a computer program. The computer program can be used in the following ways: Figure 7 Runs on the computer equipment shown.

[0238] See also Figure 7 , Figure 7 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a server, wherein the server may be an independent server or a server cluster composed of multiple servers.

[0239] See Figure 7 The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .

[0240] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which, when executed, can cause the processor 502 to execute a method for real-time silencing and replacing banned words in live broadcasts.

[0241] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.

[0242] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a method for real-time silencing and replacing banned words in live broadcasts.

[0243] The network interface 505 is used to communicate with other devices through the network. Figure 7 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device 500 to which the solution of the present application is applied. The specific computer device 500 may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0244] The processor 502 is configured to execute a computer program 5032 stored in the memory to implement the following steps:

[0245] Obtain the industry to which the live broadcast platform selected by the user belongs; load the corresponding banned word library according to the industry; obtain the user-defined protection sensitivity and strictness to obtain protection strength information; obtain the live broadcast content and convert the live broadcast content into a text stream; perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result; when the detection result is that banned words exist, use the set rules to process the detection result to obtain a banned word processing result; use visualization to present the real-time protection status, the banned word processing result and the corresponding detection result.

[0246] The set rules include at least one of bleep sound processing, silence processing, and replacement word processing.

[0247] The banned word library includes personal word library, enterprise word library and system word library. The personal word library is created and managed by the user through the front-end interactive interface, adding or deleting words and classifying words; the enterprise word library is a banned word library customized according to industry needs; the system word library is the default banned word library maintained by the platform.

[0248] The banned word library uses AI technology and machine learning algorithms to analyze historical behaviors, violation events and big data, and automatically identifies frequently appearing banned words for updating.

[0249] In one embodiment, when the processor 502 implements the step of acquiring live content and converting the live content into a text stream, it specifically implements the following steps:

[0250] Through the Windows API, a virtual microphone device is created to capture the audio stream of the live content; the audio stream is converted into a standard PCM format to obtain a conversion result; noise suppression and enhancement algorithms are applied to the conversion result to obtain a processed result; and the processed result is converted into a text stream using third-party speech recognition technology.

[0251] In one embodiment, when the processor 502 performs the step of detecting banned words on the text stream according to the protection strength information and the banned word library to obtain a detection result, the processor 502 specifically implements the following steps:

[0252] A text matching algorithm is used to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result.

[0253] In one embodiment, when the processor 502 implements the set rule including at least one of bleep sound processing, silence processing, and replacement word processing, it specifically implements the following steps:

[0254] When using the word replacement processing method, a voice cloning model is used to convert the replacement word into a voice segment consistent with the host's voice to obtain a new voice segment; the content corresponding to the detection result is replaced with the new voice segment, and a smooth transition between audio segments is performed through audio processing technology to obtain a processing result; wherein, the voice cloning model is obtained by using personal audio samples as a sample set to train a deep learning model.

[0255] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU) 306. The processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0256] Those skilled in the art will appreciate that all or part of the steps in the method of the above-described embodiment can be implemented by instructing the relevant hardware through a computer program. The computer program includes program instructions, which can be stored in a storage medium that is computer-readable. The program instructions are executed by at least one processor in the computer system to implement the steps in the method of the above-described embodiment.

[0257] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when the computer program is executed by a processor, the processor performs the following steps:

[0258] Obtain the industry to which the live broadcast platform selected by the user belongs; load the corresponding banned word library according to the industry; obtain the user-defined protection sensitivity and strictness to obtain protection strength information; obtain the live broadcast content and convert the live broadcast content into a text stream; perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result; when the detection result is that banned words exist, use the set rules to process the detection result to obtain a banned word processing result; use visualization to present the real-time protection status, the banned word processing result and the corresponding detection result.

[0259] The set rules include at least one of bleep sound processing, silence processing, and word replacement processing.

[0260] The banned word library includes personal word library, enterprise word library and system word library. The personal word library is created and managed by the user through the front-end interactive interface, adding or deleting words and classifying words; the enterprise word library is a banned word library customized according to industry needs; the system word library is the default banned word library maintained by the platform.

[0261] The banned word library uses AI technology and machine learning algorithms to analyze historical behaviors, violation events and big data, and automatically identifies frequently appearing banned words for updating.

[0262] In one embodiment, when the processor executes the computer program to implement the steps of obtaining live content and converting the live content into a text stream, the processor specifically implements the following steps:

[0263] Through the Windows API, a virtual microphone device is created to capture the audio stream of the live content; the audio stream is converted into a standard PCM format to obtain a conversion result; noise suppression and enhancement algorithms are applied to the conversion result to obtain a processed result; and the processed result is converted into a text stream using third-party speech recognition technology.

[0264] In one embodiment, when the processor executes the computer program to implement the step of detecting banned words on the text stream based on the protection level information and the banned word library to obtain a detection result, the processor specifically implements the following steps:

[0265] A text matching algorithm is used to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result.

[0266] In one embodiment, when the processor executes the computer program to implement the step of processing the detection result using a set rule to obtain a processing result, the processor specifically implements the following steps:

[0267] When using the word replacement processing method, the voice cloning model is used to convert the replacement word into a voice segment consistent with the host's voice to obtain a new voice segment; the content corresponding to the detection result is replaced with the new voice segment, and a smooth transition between audio segments is performed using audio processing technology to obtain a processing result;

[0268] The voice cloning model is obtained by using personal audio samples as a sample set to train a deep learning model.

[0269] The storage medium may be any computer-readable storage medium that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk.

[0270] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0271] In the several embodiments provided herein, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the various units is merely a logical functional division, and actual implementation may employ other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented.

[0272] The steps in the method of the embodiment of the present invention may be adjusted in order, combined, or deleted as needed. The units in the apparatus of the embodiment of the present invention may be combined, divided, or deleted as needed. In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit 306, each unit may exist physically separately, or two or more units may be integrated into a single unit.

[0273] If this integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, terminal, or network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present invention.

[0274] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A real-time silencing and replacement method for banned words in live broadcast, characterized in that: include: Get the industry of the live streaming platform selected by the user; Load the corresponding banned word library according to the industry; Obtain user-defined protection sensitivity and strictness to obtain protection strength information; Acquire live content and convert the live content into a text stream; Performing a banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result; When the detection result indicates that a banned word exists, the detection result is processed using a set rule to obtain a banned word processing result; The real-time protection status, the banned word processing results and the corresponding detection results are presented visually.

2. The method for real-time silencing and replacing banned words in live broadcast according to claim 1, characterized in that: The acquiring of live content and converting the live content into a text stream includes: Create a virtual microphone device through the Windows API to capture the audio stream of live content; Converting the audio stream into a standard PCM format to obtain a conversion result; Applying a noise suppression and enhancement algorithm to the conversion result to obtain a processing result; The processing result is converted into a text stream using third-party speech recognition technology.

3. The method for real-time silencing and replacing banned words in live broadcast according to claim 1, characterized in that: The performing banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result includes: A text matching algorithm is used to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result.

4. The method for real-time silencing and replacing banned words in live broadcast according to claim 1, characterized in that: The set rules include at least one of bleep sound processing, silence processing, and replacement word processing.

5. The method for real-time silencing and replacing banned words in live broadcast according to claim 4, characterized in that: The detection result is processed according to the set rules to obtain a processing result, including: When using the word replacement method, the voice cloning model is used to convert the replacement word into a voice segment consistent with the host's voice to obtain a new voice segment; Replacing the content corresponding to the detection result with the new voice segment, and performing a smooth transition between the audio segments using audio processing technology to obtain a processing result; The voice cloning model is obtained by using personal audio samples as a sample set to train a deep learning model.

6. The method for real-time silencing and replacing banned words in live broadcast according to claim 1, characterized in that: The banned word library includes personal word library, enterprise word library and system word library. The personal word library is created and managed by the user through the front-end interactive interface, and words can be added or deleted and classified. The enterprise word library is a banned word library customized according to industry needs. The system word library is a default prohibited word library maintained by the platform.

7. The method for real-time silencing and replacing banned words in live broadcast according to claim 6, characterized in that: The banned word library uses AI technology and machine learning algorithms to analyze historical behaviors, violation events and big data, and automatically identifies frequently appearing banned words for updating.

8. A real-time silencing and replacement device for banned words in live broadcast, characterized in that: include: An industry acquisition unit, used to obtain the industry to which the live broadcast platform selected by the user belongs; A loading unit, configured to load a corresponding banned word library according to the industry; A protection strength acquisition unit is used to obtain user-defined protection sensitivity and strictness to obtain protection strength information; A content conversion unit, configured to obtain live content and convert the live content into a text stream; a banned word detection unit, configured to perform banned word detection on the text stream according to the protection strength information and the banned word library to obtain a detection result; a processing unit, configured to, when the detection result indicates that a banned word exists, process the detection result using a set rule to obtain a banned word processing result; The visual display unit is used to visually present the real-time protection status, the banned word processing results and the corresponding detection results.

9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.