Method and device for processing audio in compliance manner, and electronic equipment
By identifying illegal audio content in live streaming through an audio compliance detection model and processing it using an enhanced audio compliance model, the problem of the concealment and difficulty in identification of audio content supervision has been solved. This enables real-time and automated compliance processing, ensuring a balance between content security and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUZHOU ROCKCHIP SEMICON
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-08
AI Technical Summary
The regulation of audio content in online live streaming faces challenges such as the spread of illegal and harmful information, false advertising, leakage of sensitive company information, and exposure of personal privacy. In particular, the high degree of concealment and difficulty in identifying audio content leads to the lag of traditional regulatory methods, making it difficult to achieve full coverage and efficient compliance.
The audio compliance detection model identifies non-compliant content segments and their timestamps in real time, and uses an audio compliance enhancement model for precise processing, including downplaying or eliminating non-compliant content, to achieve automated and accurate compliance processing and ensure the compliance of content before it is pushed out.
It enables real-time, automated compliance testing of audio content, eliminates regulatory blind spots, ensures comprehensive compliance coverage, avoids a decline in user experience due to interrupted processing, and achieves a balance between source governance and user experience.
Smart Images

Figure CN121999802A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and more particularly to methods, apparatus, and electronic devices for compliant audio processing. Background Technology
[0002] With the rapid development of live streaming technology, the influencer live streaming industry has risen rapidly and is widely used in various technological scenarios such as social media, entertainment platforms, e-commerce live streaming, educational live streaming, product sales live streaming, news media live streaming, and product launches. The application of live streaming is becoming increasingly rich and in-depth.
[0003] However, numerous challenges remain in the regulation of audio and video content. For example, issues such as the spread of illegal and harmful information, false advertising, leakage of sensitive company information, exposure of personal privacy, weak user self-discipline, and the relative lag between technological development and regulatory measures are becoming increasingly prominent. Audio content, in particular, presents new challenges to content compliance due to its highly concealed nature and difficulty in identification. Summary of the Invention
[0004] This invention provides a method, apparatus, and electronic device for compliant audio processing, which can achieve compliant detection of audio data and enhance the ability to supervise audio and video content.
[0005] In one aspect of the present invention, a method for compliant audio processing is provided. The method includes: acquiring raw audio data to be processed; detecting the raw audio data using an audio compliance detection model to generate content information associated with a non-compliant content segment of the raw audio data; processing the non-compliant content segment of the raw audio data using an audio compliance enhancement model based on the content information associated with the non-compliant content segment to obtain compliant audio data; and outputting the compliant audio data.
[0006] In another aspect of the invention, an apparatus for compliant audio processing is provided. The apparatus includes: an audio acquisition module configured to acquire raw audio data to be processed; an audio compliance detection module configured to detect the raw audio data using an audio compliance detection model to generate content information associated with non-compliant content segments of the raw audio data; to process the non-compliant content segments of the raw audio data using an audio compliance enhancement model to obtain compliant audio data; and an audio output module configured to output the compliant audio data.
[0007] In another aspect of the invention, an electronic device is provided. The electronic device includes: a memory configured to store an executable program; and a processor configured to execute the program to perform the aforementioned method for compliant audio processing.
[0008] According to the technical solution of this invention, by acquiring raw audio data and using an audio compliance detection model to perform real-time, automatic compliance analysis on the raw audio data, it can accurately identify non-compliant content segments that are difficult to cover by traditional manual review based on semantic understanding. This eliminates regulatory blind spots caused by limitations in manpower and time, ensuring comprehensive compliance coverage. Unlike traditional one-size-fits-all blocking, it uses an audio compliance enhancement model for precise processing based on the identification results. This point-to-point processing flow ensures that the audio is converted into compliant audio before output, controlling content risks from the source. It solidifies complex compliance rules into an automatically executable model, achieving stable and objective compliance management without real-time human intervention. This reduces operating costs and the risk of human error, providing an efficient, reliable, and user-friendly automated solution that balances user experience and security. Attached Figure Description
[0009] Figure 1 A flowchart illustrating the steps of a method for compliant audio processing according to an embodiment of the present invention; Figure 2 This is an overall environmental architecture diagram of a method for compliant audio processing according to an embodiment of the present invention; Figure 3 A flowchart illustrating a method for performing compliant audio processing on an electronic device with intelligent compliant audio processing according to an embodiment of the present invention; Figure 4 A flowchart illustrating the application of the intelligent compliant audio processing method according to an embodiment of the present invention in a real-world scenario; Figure 5 A flowchart illustrating the use of an audio compliance enhancement model to enhance audio content according to an embodiment of the present invention; Figure 6 A schematic diagram of the structure of an audio processing apparatus according to an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0010] To explain in detail the technical content, objectives, and effects of the present invention, the following description is provided in conjunction with the embodiments and accompanying drawings.
[0011] To address compliance requirements in online live streaming, feasible solutions primarily include establishing a database of prohibited content, utilizing big data analytics for real-time early warning, and introducing artificial intelligence technologies such as machine learning to improve content identification and review efficiency. These measures enhance the regulatory capacity for audio and video content to some extent, but continuous optimization and upgrades are still needed to adapt to the ever-changing content ecosystem.
[0012] To address at least the aforementioned technical issues, this disclosure provides a solution for compliant audio processing. According to this disclosure, by sequentially executing two core steps—audio compliance detection and intelligent enhancement—after obtaining the illegal content fragments and their precise timestamp information from the original audio data, an audio compliance enhancement model is used to determine whether the illegal content is diluted or compliant content is implanted, thereby enhancing the automated and precise compliance processing of the original audio data. In this way, according to embodiments of this disclosure, illegal fragments in audio can be automatically identified and intelligently replaced or repaired at the content level before content is pushed, thus ensuring content compliance while effectively avoiding a decline in user experience due to interrupted processing or muting, achieving a balance between source governance and user experience.
[0013] In the following, the technical solutions according to this disclosure will be described with reference to specific embodiments and in conjunction with the accompanying drawings.
[0014] Figure 1 This is a flowchart illustrating a method 100 for compliant audio processing according to an embodiment of this disclosure. (Refer to...) Figure 1 The method 100 includes the following steps S102 to S108.
[0015] In S102, the raw audio data to be processed is obtained.
[0016] In some embodiments, the broadcaster software of the live streaming device acquires digital audio data collected by the audio acquisition module.
[0017] In S104, the original audio data is detected by an audio compliance detection model to generate content information associated with the illegal content segments of the original audio data.
[0018] In some embodiments, the broadcaster software of the live streaming device utilizes an intelligent compliance audio processing module to detect the raw audio data using an audio compliance detection model, thereby generating content information associated with the non-compliant content segments in the raw audio data. In this way, audio compliance analysis capabilities are embedded into the source of live streaming content production, enabling real-time detection and precise location of non-compliant content at the broadcaster's end. This achieves proactive and localized compliance control, effectively reducing the risks of network latency and privacy leaks caused by audio streams being transmitted to cloud servers for review. Simultaneously, it provides immediate and reliable data for the broadcaster to adjust broadcast content in real time or for the system to automatically trigger enhanced processing.
[0019] In some embodiments, detecting the original audio data using an audio compliance detection model includes: performing content understanding based on audio features and textual semantics on the original audio data using the audio compliance detection model, which serves as a bimodal audio-text model. This bimodal model integrates audio feature extraction and textual semantic analysis capabilities. In this way, the underlying acoustic features and potential high-level semantic information of the audio can be simultaneously analyzed, achieving a comprehensive understanding and identification of colloquial expressions, context-related violations, and specific violating audio clips. This breaks down the technical barriers of traditional single-modal audio analysis, constructs a more comprehensive audio content understanding framework, and improves the detection rate and accuracy of concealed and diverse violating audio segments, providing a reliable technical foundation for refined audio compliance governance.
[0020] In some embodiments, a newly added library of illegal audio samples and corresponding illegal text tags are acquired, and the pre-trained audio compliance detection model is incrementally trained based on the newly added library of illegal audio samples. In this way, an intelligent audio compliance recognition system with continuous evolution capabilities is constructed. The system can dynamically collect newly emerging illegal audio cases and their multi-dimensional semantic descriptions, and use these new samples to incrementally fine-tune the parameters of the deployed audio-text bimodal detection model, enabling it to identify new illegal expressions, variant homophones, emerging sensitive topics, and constantly changing advertising and lead generation techniques. This mechanism ensures that the audio compliance detection capability can be upgraded synchronously with changes in the content ecosystem, overcoming the limitations of traditional static dictionaries or fixed models that are prone to becoming outdated and ineffective. By performing online or periodic incremental updates to the model, the system's ability to detect complex and hidden illegal content and its warning accuracy can be continuously improved with low computational cost without retraining the entire model, thereby maintaining a high level of efficient and accurate content control in the dynamically changing online live streaming environment.
[0021] In some embodiments, detecting the original audio data using an audio compliance detection model to generate content information associated with the non-compliant content segments of the original audio data includes: converting the original audio data into a phoneme-aligned lexical sequence using an audio encoder; and inputting the lexical sequence into a large language model to perform compliance analysis on the original audio data based on semantic understanding, thereby generating content information associated with the non-compliant content segments of the original audio data. In this way, the original audio signal is mapped into a structured sequence with semantic information, and the deep semantic understanding capabilities of the large language model are utilized to achieve contextual analysis and intent recognition of non-compliant content. This overcomes the limitations of traditional audio feature matching in semantic understanding, enabling not only the identification of explicit keywords but also the understanding of complex non-compliant forms such as colloquial expressions, homophones, and implicit hints through contextual information. This improves the accuracy and coverage of audio content compliance detection. Simultaneously, transforming audio analysis into structured semantic analysis facilitates the system's precise location of non-compliant content on the timestamp, providing a reliable basis for subsequent refined processing.
[0022] In some embodiments, converting the raw audio data into a phoneme-aligned lexical sequence using an audio encoder includes: associating the audio features of the raw audio data with text semantics using the audio encoder to generate the lexical sequence corresponding to the audio features and the text semantics. In this way, a structured mapping between audio signals and semantic space is achieved. The audio encoder converts continuous audio waveforms into discrete lexical sequences with semantic representation capabilities, where each lexical not only contains acoustic feature information but also is associated with potential text semantic units, thereby constructing an alignment relationship between audio features and high-level semantics. This phoneme-aligned lexical sequence can effectively represent the dual information of audio content at both the acoustic and semantic levels, providing a structured and interpretable intermediate representation for subsequent semantic understanding and analysis based on large language models. This allows the system to bypass transcription errors in traditional speech recognition and directly and accurately judge audio compliance at the fusion representation level, improving the robustness of semantic parsing and the accuracy of violation identification for complex audio content.
[0023] In some embodiments, the content information associated with the offending content segment in the original audio data includes the timestamp of the offending content, the location of the audio segment, and the type of violation. This provides a structured, operable, and precise input for subsequent audio enhancement processing. The timestamp identifies the start and end times of the offending content in the original audio stream, providing an absolute reference for temporal localization. The location of the audio segment further clarifies the relative interval of the offending segment in the separated voice channel or the complete audio stream, ensuring the targeted nature of the enhancement operation. The type of violation indicates the risk attributes and semantic categories of the content, enabling the audio compliance enhancement model to adaptively select and execute the most suitable processing strategy based on different violation characteristics. This multi-dimensional associated information forms a crucial bridge from violation detection to precise processing, achieving an upgrade in compliance governance from broad interception to refined and differentiated operations, optimizing resource allocation and processing efficiency while ensuring processing effectiveness.
[0024] In S106, based on the content information associated with the illegal content fragment, the illegal content fragment of the original audio data is processed through an audio compliance enhancement model to obtain compliant audio data.
[0025] In some embodiments, the anchor software of the live streaming device utilizes an intelligent compliant audio processing module to process non-compliant content segments of the original audio data through an audio compliance enhancement model to obtain compliant audio data. In this way, content enhancement capabilities are deployed on the live streaming content production side, enabling non-compliant audio segments to be repaired or replaced in real time before streaming. This ensures that the broadcast content is fully compliant while maintaining the continuity and natural listening experience of the audio to the greatest extent possible, providing anchors with seamless compliant broadcast protection and improving the automation level and broadcast security of the live streaming process.
[0026] In some embodiments, based on content information associated with the violating content fragment, the violating content fragment of the original audio data is processed by an audio compliance enhancement model to obtain compliant audio data. This includes: performing voice separation processing on the original audio data to obtain independent voice and background sound channels; performing processing operations on the violating content fragment in the voice channel based on the audio segment position of the violating content fragment using the audio compliance enhancement model; and merging the processed voice channel with the background sound channel to obtain the compliant audio data. In this way, voice separation processing can be performed in parallel or in advance after obtaining the original audio data to obtain independent voice and background sound channels. When the audio compliance detection model identifies the violating content fragment and its timestamp information, the system directly locates the corresponding audio segment in the voice channel based on the timestamp information and performs targeted processing operations on that part of the voice using the audio compliance enhancement model. After processing, the enhanced voice channel is merged back with the original background sound channel to generate the final compliant audio data. By separating and processing the human voice channel from the background sound channel, the system can accurately eliminate or replace non-compliant human voice content while maintaining the continuity and naturalness of the background ambient sound. This avoids interference with compliant background sound and ensures that the final output audio has a smooth and realistic overall listening experience, thus improving the sound quality and user experience after compliance processing.
[0027] In some embodiments, processing the illegal content segment in the human voice channel based on the audio segment position of the illegal content segment using the audio compliance enhancement model includes: performing at least one of the following processing methods on the illegal content segment in the human voice channel: fade-out processing and compliance content enhancement implantation processing. The fade-out processing includes using a deep learning audio enhancement model to reduce the clarity and recognizability of the illegal content while preserving the basic characteristics of the human voice. The compliance content enhancement implantation processing includes implanting preset compliant audio content at a specific audio segment position. In this way, not only can the identified illegal content be processed flexibly, avoiding the auditory interruption and abruptness caused by traditional muting or truncation, but brand promotion or advertising information can also be proactively implanted in compliant scenarios, achieving an upgrade in content governance from passive interception to proactive guidance. Specifically, the fade-out processing uses selective frequency attenuation and phase perturbation technology to weaken the semantic clarity of the illegal content while maintaining the natural timbre and loudness dynamics of the human voice, ensuring a smooth transition between the processed segment and the preceding and following audio. The integration of compliant content utilizes acoustic environment matching and dynamic adjustment of vocal volume to seamlessly blend the embedded audio material with the original vocals and background environment. This achieves commercial promotion or content supplementation without disrupting the overall listening experience and atmosphere of the live stream. Through this dual-mode processing mechanism, this embodiment significantly improves the broadcast quality, auditory continuity, and flexibility of commercial applications of live audio while ensuring content security and compliance.
[0028] In some embodiments, employing a deep learning audio enhancement model to reduce the clarity and recognizability of inappropriate content while preserving basic human voice characteristics includes: selectively attenuating the frequency and perturbing the phase of the audio signal corresponding to the inappropriate content segment in the human voice channel. The selective frequency attenuation targets high-frequency and transient frequency components in the audio signal that characterize semantic clarity. In this way, semantic deconstruction of inappropriate speech content is achieved rather than physical deletion. By precisely attenuating key high-frequency components carrying semantic information and perturbing the phase, the intelligibility and recognizability of inappropriate content can be significantly reduced audibly, making it unclear and thus meeting compliance requirements. Simultaneously, because the processing is concentrated on a specific frequency band and uses a smooth attenuation curve, the basic timbre, fundamental frequency profile, and loudness dynamics of the human voice are preserved, avoiding the severe sound quality degradation caused by traditional low-pass filtering or full-band suppression. This technique ensures that the processed speech segment still sounds like a natural human voice, smoothly connecting with preceding and following normal speech, maximizing the continuity and naturalness of the live audio, and optimizing the listener's auditory experience while achieving content compliance.
[0029] In some embodiments, embedding preset compliant audio content at specific audio segment locations includes: performing acoustic environment matching processing on the preset compliant audio content; and dynamically adjusting the loudness of the human voice channel in the audio segment where the preset compliant audio content and human voice coexist. This achieves seamless integration of the embedded content with the original live stream audio scene. By applying reverberation, equalization, and noise characteristics that match the current live stream environment to the preset compliant audio content, it sounds as if it originates from the same spatial sound field, avoiding a jarring post-added feel. Simultaneously, during the period when the embedded content and original human voice coexist, the system analyzes the loudness profile of the human voice in real time and performs adaptive dynamic compression or volume envelope matching technology to ensure that the embedded content is clearly audible without masking or interfering with the expression of the main human voice, maintaining the prominence and intelligibility of the main speech. This refined audio processing strategy allows embedded content such as brand advertisements and compliance prompts to be naturally embedded in the live stream, achieving commercial promotion or information supplementation without interrupting the live stream rhythm or impairing the listening experience, thus enhancing the commercial value and functional expandability of the live stream content.
[0030] In S108, the compliant audio data is output.
[0031] In some embodiments, the broadcaster software on the live streaming device uses audio encoders such as AAC / MP3 to compress and encode compliant live audio streams. The compliant digital audio and video streams are then packaged into a compliant live stream using MPEG2TS, and distributed to the live streaming server via HTTP or RTSP. In this way, audio encoding and compression significantly reduce data bandwidth usage and improve network transmission efficiency while maintaining sound quality. MPEG2TS encapsulation multiplexes the processed compliant audio and synchronized video streams into an industry-standard transport stream format, ensuring audio-video synchronization, program information integrity, and cross-platform compatibility. Finally, the live stream is pushed to the live streaming server via common streaming media protocols such as HTTP or RTSP, enabling the live content, after real-time compliance enhancement processing, to seamlessly integrate with existing distribution networks and infrastructure. This process achieves efficient integration of compliance enhancement processing and broadcast-grade streaming output on the content production side. While ensuring content security and broadcast quality, it avoids increased system complexity and deployment costs caused by introducing proprietary protocols or non-standard formats, thus supporting the large-scale, smooth implementation and practical application of intelligent audio compliance processing technology.
[0032] In some embodiments, outputting the compliant audio data includes: audio encoding and streaming media packaging of the compliant audio data to generate a compliant live stream; and pushing the compliant live stream to a live streaming server. This completes a full technical loop from intelligent audio processing to standard live streaming distribution. Specifically, the system compresses the processed compliant audio data using efficient audio encoding formats such as AAC or MP3 to reduce bandwidth consumption; then, the encoded audio stream and the synchronized video stream are multiplexed and encapsulated according to standard streaming media formats such as MPEG2TS to form a transport stream conforming to common live streaming protocols such as HTTP and RTSP. Finally, the packaged compliant live stream is pushed to the live streaming server in real time via network protocols for subsequent distribution and broadcasting. This process ensures that the intelligently compliant audio content can seamlessly integrate with existing live streaming infrastructure and industry standards, guaranteeing content security and broadcast quality without introducing additional system complexity or compatibility barriers, thus achieving the smooth implementation and large-scale deployment of intelligent audio compliance processing technology.
[0033] Figure 2 This is an overall environmental architecture diagram illustrating a method for compliant audio processing according to an embodiment of the present invention. As shown in the figure, this architecture constructs a full-link embedded processing system with intelligent live streaming equipment at its core, covering the entire audio content chain from acquisition to distribution, realizing real-time and localized audio compliance governance at the production source.
[0034] Specifically, the live streaming device integrates an audio acquisition module 201 and a video acquisition module 202, which work together to simultaneously capture raw audio and video data during the live stream. The raw audio data, as the primary target for governance, is sent in real-time to the integrated intelligent compliance audio processing unit. This unit is the core functional module of this architecture, further divided into an audio compliance detection module and an audio compliance enhancement model, forming a two-stage processing pipeline that operates in series. The audio compliance detection module first analyzes the raw audio data in real-time, using a pre-loaded audio-text dual-modal detection model to identify potentially illegal content segments such as those involving pornography, violence, advertising, or copyright infringement, and outputs their precise timestamp information. Subsequently, based on the identification results, the audio compliance enhancement model performs intelligent content processing on the illegal segments: on the one hand, it can selectively attenuate the frequency and perturb the phase of the illegal voice through a deep learning enhancement model to achieve semantic-level fading and elimination; on the other hand, in compliant scenarios, it can also naturally embed pre-set brand promotion or advertising audio content into the corresponding positions through acoustic matching and loudness fusion technology. Both processing modes aim to preserve the naturalness of the human voice and auditory continuity, avoiding the experience interruption caused by traditional mute or truncation.
[0035] The enhanced, compliant audio data, along with synchronously acquired or parallel-processed video data, is fed into an integrated audio encoder. The audio encoder efficiently compresses the audio data, such as using AAC / MP3 audio encoding, and then multiplexes and encapsulates it according to standard streaming media formats like MPEG2TS, generating a structurally complete, compliant audio and video stream that conforms to general transmission specifications. This compliant stream is then directly pushed to the live streaming server via real-time streaming protocols such as RTSP or HTTP. Upon receiving the stream, the live streaming server distributes it to various video live streaming clients, thus completing the end-to-end business process from content acquisition, local intelligent compliance processing, standardized encoding and encapsulation, to platform distribution and terminal playback.
[0036] This architecture pushes complex audio compliance analysis models and enhanced processing capabilities down to electronic devices, completing all content governance work before audio data encoding and streaming. This achieves real-time and low-latency compliance processing, avoiding network latency and uncertainty caused by cloud processing; it also strengthens source control of content security and reduces reliance on platform-side review resources. Simultaneously, localized processing ensures the accuracy and flexibility of audio processing, achieving a complete governance loop from violation suppression to compliance guidance through dual-module collaboration. While ensuring content safety and broadcast compliance, it maximizes user experience and auditory smoothness, providing the industry with an integrable, evolvable, and controllable embedded audio compliance solution.
[0037] Figure 3 This is a flowchart illustrating the method for performing the above-described compliance audio processing by an electronic device with intelligent compliance audio processing according to an embodiment of the present invention. Please refer to... Figure 3 The method includes S301 to S307.
[0038] S301, The broadcast software acquires digital audio from the audio acquisition module.
[0039] In some embodiments, the audio acquisition module is integrated into the live streaming device to capture the original analog audio signal in real time with a preset sampling rate and bit depth, and convert it into a continuous digital audio stream for further processing by the broadcaster software.
[0040] S302. The broadcaster software uses an intelligent compliant audio processing module to process digital audio.
[0041] In some embodiments, the intelligent compliant audio processing module loads and runs a pre-built audio compliance detection model and an audio compliance enhancement model. First, the audio compliance detection model performs real-time analysis of the digital audio stream to identify non-compliant content segments and their corresponding timestamps. Then, based on the identification results, the audio compliance enhancement model is invoked to perform fade-out or compliant content insertion processing on the non-compliant segments, generating compliant digital audio data.
[0042] S303 and broadcast software use audio encoders such as AAC / MP3 to compress and encode compliant live audio streams.
[0043] In some embodiments, the broadcaster software inputs the processed compliant digital audio stream to the audio encoder, which encodes it using lossy compression algorithms such as AAC or MP3, significantly reducing the amount of data while ensuring auditory quality, in order to adapt to the bandwidth limitations of network transmission.
[0044] S304 The broadcasting software packages compliant digital audio and video streams into compliant live audio streams according to MPEG2TS.
[0045] In some embodiments, the encoded compliant audio stream and the synchronously acquired or processed video stream are multiplexed and encapsulated in the MPEG2TS transport stream format to form a standard composite stream containing metadata such as audio and video synchronization information and program-specific information, ensuring its integrity and resolvability during transmission and distribution.
[0046] S305. The broadcaster software distributes compliant live audio streams to the live streaming server via HTTP or RTSP.
[0047] In some embodiments, the encapsulated compliant audio and video streams are pushed to the live streaming server via network transmission protocols such as HTTP, RTSP, or RTMP in a streaming manner. This step enables content uploading from local devices to cloud services, providing the source stream for subsequent large-scale distribution.
[0048] S306. The live streaming server distributes compliant live audio streams to the live streaming client.
[0049] In some embodiments, after receiving compliant audio and video streams from the broadcaster's device, the live streaming server distributes them to one or more connected live streaming clients via a content delivery network or direct stream forwarding, supporting the real-time viewing needs of low latency and high concurrency.
[0050] S307. Live streaming clients use smart terminals to play compliant live audio streams.
[0051] In some embodiments, the live streaming client receives audio and video streams on a smart terminal (such as a mobile phone, tablet, or computer), decodes and renders them, and then plays compliant live streaming content synchronously through speakers and a display screen, completing the entire business process from content production, processing, transmission to consumption. Figure 4 This is a flowchart illustrating the application of the intelligent compliant audio processing method according to an embodiment of the present invention in a real-world scenario. Please refer to... Figure 4 The process includes steps S402 to S410.
[0052] S402, The broadcaster software loads an AI model for intelligent and compliant audio processing.
[0053] In some embodiments, the AI model includes an audio compliance detection model and an audio compliance enhancement model, which are pre-loaded and integrated into the broadcaster software as a composite model. The audio compliance detection model is preferably a bimodal audio-text architecture, integrating audio feature extraction and text semantic understanding capabilities. The audio compliance enhancement model is built on a deep learning framework, supporting the downscaling and removal of illegal audio content and the enhancement and embedding of compliant content.
[0054] S404, The broadcast software acquires digital audio from the audio acquisition module.
[0055] In some embodiments, the audio acquisition module captures analog audio signals in a live broadcast scene in real time and generates a continuous digital audio stream through analog-to-digital conversion. This digital audio stream serves as input data for intelligent compliant audio processing and is transmitted to the AI model processing pipeline already loaded in the broadcaster software.
[0056] S406. Use an audio compliance detection model to detect audio segments that violate content regulations in digital audio.
[0057] In some embodiments, the audio compliance detection model performs real-time analysis on the input digital audio stream: first, the audio signal is converted into a discrete token sequence with phono-text alignment by an audio token encoder; then, the token sequence is input into an integrated large language model, which identifies illegal content such as pornography, violence, advertising, and copyright infringement based on its semantic understanding capabilities; finally, the illegal audio segments and their precise start and end timestamp information are output.
[0058] In audio-text multimodal LLM, the goal is to convert continuous audio signals (such as speech, ambient sound, and music) into discrete audio tokens with semantic information, and map these tokens into a feature space shared with text tokens, enabling LLM to simultaneously understand and process information from both modalities. The process involves several steps: First, a discrete tokenizer for the audio signal is used. Since audio signals are continuous, high-dimensional, and have a much higher temporal resolution than text, they need to be converted into discrete tokens to fit the LLM input structure. Second, cross-modal connections and feature mapping map the features of the audio tokens to the text feature space of the LLM. Third, unified sequence modeling and inference, after processing by the encoder and adapter, unifies the semantics and dimensionality of the audio and text token embeddings. Fourth, semantic alignment between audio and text allows the model to seamlessly transfer rich linguistic and world knowledge learned from text to audio understanding tasks, improving the model's generalization ability on unknown audio tasks. The model output is customized using an LLM prohibited word detection template. Generally, the output includes the timestamp range and violation type.
[0059] According to embodiments of the present invention, through "audio-text dual-modal" fusion and LLM intelligent analysis, efficient, accurate, and real-time compliance detection of audio content is achieved. This not only solves the accuracy and real-time issues in traditional audio review but also provides detailed violation information, offering comprehensive and reliable compliance assurance for audio content platforms. This technology has extremely high application value in scenarios with explosive growth in audio content, such as live streaming, voice-based social networking, and online games, effectively reducing content risks for enterprises and improving user experience.
[0060] S408. Use the audio compliance enhancement model to either fade or remove audio segments that violate content regulations, or to enhance or embed them.
[0061] In some embodiments, based on the timestamp information of the violating segments output by S406, the audio compliance enhancement model performs corresponding content processing operations. The fade-out removal process employs deep learning-based audio restoration technology, which reduces the semantic clarity of the violating content through selective frequency attenuation and phase perturbation, while preserving the basic characteristics and naturalness of the human voice. The enhancement implantation process, through acoustic environment matching and dynamic adjustment of human voice loudness, naturally integrates preset compliant audio content, such as brand promotions and advertising slogans, into a designated position in the original human voice channel, achieving compliance enhancement at the content level.
[0062] In audio processing, basic vocal features typically refer to the underlying acoustic parameters that define a speaker's identity, emotions, and basic pronunciation attributes. These are relatively independent of the semantic content of speech. Basic vocal features include: 1) Fundamental frequency: the frequency of vocal cord vibration, which determines the pitch of the voice; 2) Formants: the resonant frequencies generated when airflow passes through the vocal tract (throat, oral cavity, nasal cavity), which determine timbre, tone, and vowel pronunciation; 3) Pace / Rhythm: the speed and pauses of speech, related to the speaker's style and emotions; and 4) Loudness / Energy: the volume of the sound, also related to emotion and emphasis.
[0063] Furthermore, semantic clarity primarily exists in the high-frequency and transient components of speech signals. By precisely attenuating these high-frequency information, the speech content becomes blurred and difficult to understand, while preserving the fundamental low-frequency pitch. Slightly perturbing the phase information of the speech causes temporal distortion, thus reducing clarity.
[0064] Furthermore, to ensure the seamless integration of embedded brand advertising content with the original voice, the key lies in achieving acoustic consistency and temporal smoothness. First, the model strives to determine the optimal timing for insertion, typically by identifying and utilizing natural pauses, breaths, or non-speech background noise segments within the voice, thus avoiding abrupt insertions during high-intensity vocalizations and ensuring a smooth temporal flow. Subsequently, to guarantee acoustic consistency, the system must perform environmental feature matching. This includes using deep learning models to analyze the reverberation characteristics of the original voice segment and the type of background noise. Once these environmental acoustic features are accurately extracted, they are synthesized and applied to the embedded advertising content. For example, if the voice sounds like it has a long room reverberation, the advertising text must also be given the same reverberation effect, making the listener feel as if the advertising text is emanating from the same physical space as the voice, greatly enhancing the realism of the integration. In segments where the embedded content and voice need to coexist, the model employs dynamic audio processing techniques rather than simple superposition. Specifically, the system analyzes the frequency overlap between the embedded content and the human voice, and utilizes the frequency masking effect of human hearing to optimize the embedding, ensuring that key human voice information is not completely masked. Simultaneously, through dynamic compression or volume envelope matching technology, the model subtly and smoothly reduces the loudness of the human voice when the ad text appears, and immediately restores it when the ad text ends, ensuring that the volume change is gentle and imperceptible, ultimately achieving the effect that the embedded content sounds as if it were natively present in the audio stream.
[0065] According to embodiments of the present invention, the audio compliance enhancement model achieves high-quality compliance processing of audio content by combining deep learning technology with accurate violation information. It can not only effectively eliminate illegal content, but also implant compliant content when necessary, ensuring the compliance, high quality and user experience of audio content.
[0066] S410 outputs the digital audio after intelligent compliance audio processing to the next level.
[0067] In some embodiments, the enhanced digital audio data is output to a subsequent processing module, such as an audio encoder or an audio-video multiplexing module, for further compression, encapsulation, and ultimately to form a compliant live audio-video stream suitable for streaming. This step completes the core processing loop from raw audio input to compliant audio generation, providing technical assurance for the secure broadcast of live content.
[0068] Figure 5 This is a flowchart illustrating the use of an audio compliance enhancement model to enhance audio content according to an embodiment of the present invention. Please refer to... Figure 5 The process includes S501 to S507.
[0069] S501, the broadcaster software loads the audio compliance enhancement model.
[0070] In some embodiments, the broadcasting software dynamically loads a pre-trained audio compliance enhancement model during the initialization phase or before processing. This model is based on a deep learning architecture and has the ability to extract audio features, repair content, and enhance fusion. It supports the fading and elimination of illegal audio segments and the natural implantation of compliant audio content.
[0071] S502, the broadcast software acquires digital audio from the audio acquisition module.
[0072] In some embodiments, the audio acquisition module continuously captures audio signals in the live streaming environment and converts them into a digital audio stream as an input source for enhancement processing.
[0073] S503. Obtain information on illegal audio content output by the superior module.
[0074] In some embodiments, the parent module is an audio compliance detection model, which outputs information about the illegal audio content, including the start and end timestamps of the illegal segment, the location of the audio segment, and the type of violation, such as pornography, violence, advertising, or copyright infringement. This information provides a basis for positioning and classification for subsequent targeted enhancement processing.
[0075] S504, Performs voice separation on digital audio.
[0076] In some embodiments, based on the timestamp information obtained in S503, the audio data in the original digital audio within the corresponding time period is subjected to voice separation processing. For example, a speech separation model based on deep neural networks or convolutional neural networks is used to separate the audio stream into independent voice channels and background sound channels so that the voice content can be independently enhanced without affecting the background sound effects.
[0077] S505. Use an audio compliance enhancement model to enhance the audio content of human voices. Optionally, audio segments that violate content regulations may be diluted or enhanced.
[0078] In some embodiments, the audio compliance enhancement model performs enhancement operations on the portion of the human voice channel corresponding to the violation timestamp: if a fade-out process is performed, the semantic clarity and recognizability of that portion of the human voice are reduced by audio enhancement technology based on generative adversarial networks or recurrent neural networks, while maintaining a natural timbre; if an enhancement implantation process is performed, preset compliant audio content, such as brand slogans or advertising audio, is implanted into the corresponding position through acoustic matching and mixing technology, and the loudness of the human voice is adjusted to achieve smooth integration.
[0079] S506: Merge the audio enhancement channels for human voice and background sound.
[0080] In some embodiments, the enhanced vocal channel and the original separated background sound channel are time-domain aligned and signal-mixed to regenerate a complete audio stream, ensuring that the vocal and background sound effects are consistent in timing, loudness, and sound field characteristics.
[0081] S507, Outputs enhanced digital audio content.
[0082] In some embodiments, the merged digital audio stream is output as a result of compliance enhancement processing to downstream modules, such as audio encoders or stream encapsulation modules, for further compression, multiplexing, and ultimately forming a compliant live audio stream that can be distributed.
[0083] According to another aspect of the invention, Figure 6 This is a schematic diagram illustrating the structure of an audio compliance processing apparatus 600 according to an embodiment of the present invention. (Refer to...) Figure 6 The device 600 includes an audio acquisition module 601, an audio compliance detection module 602, an audio compliance enhancement module 603, and an audio output module 604.
[0084] The audio acquisition module 601 is configured to acquire raw audio data to be processed.
[0085] The audio compliance detection module 602 is configured to detect the original audio data through an audio compliance detection model in order to generate content information associated with the illegal content segments of the original audio data.
[0086] The audio compliance enhancement module 603 is configured to process the illegal content fragments of the original audio data through an audio compliance enhancement model based on the content information associated with the illegal content fragments, so as to obtain compliant audio data.
[0087] The audio output module 604 is configured to output the compliant audio data.
[0088] It should be understood that the audio acquisition module 601, the audio compliance detection module 602, the audio compliance enhancement module 603, and the audio output module 604 can be further configured to perform the corresponding steps or actions in the method described above, which will not be elaborated here.
[0089] According to embodiments of the present invention, a complete and integrable hardware-level audio compliance processing system has been constructed, achieving full-process automation from audio input, violation identification, intelligent enhancement to compliance output. This device, through modular design, solidifies complex AI models and processing logic into dedicated functional units, which can be directly deployed in embedded environments such as live streaming equipment and audio / video processors. It completes compliance governance in real time within the audio signal link without relying on cloud computing power or manual intervention. This not only significantly reduces system latency and network dependence, ensuring the real-time nature and continuity of live broadcasts, but also strengthens data security and privacy protection through localized processing, providing a highly efficient, reliable, and scalable integrated compliance solution for various audio and video content production platforms.
[0090] According to another aspect of the invention, Figure 7 This is a schematic diagram illustrating an electronic device 700 according to an embodiment of the present invention. (Refer to...) Figure 7 The electronic device 700 includes a memory 702, a processor 704, and an executable program stored in the memory 702 and executable on the processor 704. When the processor 704 executes the program, it implements the various steps of the compliant audio processing method described above.
[0091] In summary, this invention provides a method, apparatus, and electronic device for compliant audio processing. By constructing a bimodal audio compliance detection model and an audio compliance enhancement model supporting multi-mode enhancement, it achieves real-time identification and intelligent processing of illegal content before live audio streams are pushed. This method can accurately locate illegal audio segments such as those involving pornography, violence, advertising, and copyright infringement based on semantic understanding. Through voice separation and independent enhancement technologies, it performs fade-out or compliant content insertion processing on illegal content while preserving the naturalness of the voice and broadcast continuity, thus overcoming the problems of experience interruption and sound quality damage caused by traditional keyword filtering or simple muting. Simultaneously, through an incremental learning mechanism, the detection model can continuously adapt to new forms of violations, establishing an evolutionary audio compliance governance system at the content source. This solution significantly improves the precision of audio processing, system adaptability, broadcast quality, and user experience while ensuring the safety and compliance of live streaming content, providing reliable technical support for the healthy development and commercial operation of the online live streaming industry.
[0092] In its implementation, the system constructs a dynamically configurable intelligent processing strategy based on the violation attributes of audio content and the timeliness requirements for processing. For different violation categories, such as "pornographic content," "advertising," and "copyrighted audio," the system can adaptively match the corresponding compliance processing mode and real-time strategy, outputting differentiated enhanced processing solutions. Simultaneously, based on the semantic characteristics of the violation content and the urgency of the situation, the system employs an intelligent scheduling mechanism to collaboratively optimize real-time detection and enhanced processing. Based on the impact weight of different violation attributes on content security and broadcast experience, it dynamically allocates computing resources and processing priorities, thereby achieving a balance between recognition accuracy and real-time processing in live streaming compliance.
[0093] Furthermore, during the compliance enhancement process, the system establishes a dual-mode complementary mechanism of de-emphasis and compliance integration, supporting intelligent switching based on contextual semantics. When high-risk violations are detected, the system prioritizes rapid de-emphasis to ensure broadcast safety; in compliance guidance or commercial promotion scenarios, it can switch to content integration mode to achieve natural integration of brand information. This method differs from traditional, singular audio processing approaches. Through a dual-dimensional decision-making process of semantic understanding and scenario adaptation, it can more accurately balance content security and broadcast quality, significantly improving the efficiency of compliance governance for live audio.
[0094] This system is applicable to various live streaming scenarios that require real-time audio content control. By building a complete intelligent compliance processing system, it effectively improves the security and controllability of live streaming content and the smoothness of broadcasting, enhances the platform's content governance and risk response capabilities, and optimizes the efficiency of computing resource utilization and the consistency of user experience.
[0095] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent modifications made based on the content of the present invention specification and drawings, or direct or indirect applications in related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for compliant audio processing, characterized in that, include: Obtain the raw audio data to be processed; The original audio data is detected using an audio compliance detection model to generate content information associated with the illegal content segments in the original audio data. Based on the content information associated with the illegal content fragments, the illegal content fragments of the original audio data are processed through an audio compliance enhancement model to obtain compliant audio data; as well as Output the compliant audio data.
2. The method according to claim 1, characterized in that, The detection of the original audio data using the audio compliance detection model includes: The audio compliance detection model, which serves as a bimodal audio-text model, performs content understanding on the original audio data based on audio features and text semantics. The bimodal audio-text model integrates audio feature extraction and text semantic analysis capabilities.
3. The method according to claim 1, characterized in that, The original audio data is inspected using an audio compliance detection model to generate content information associated with the non-compliant content segments of the original audio data, including: The raw audio data is converted into a phoneme-aligned sequence of words using an audio encoder; and The word sequence is input into a large language model to perform compliance analysis on the original audio data, generating content information associated with the illegal content segments of the original audio data.
4. The method according to claim 3, characterized in that, Converting the raw audio data into a phoneme-aligned sequence of terms using an audio encoder includes: The audio encoder associates the audio features of the original audio data with the text semantics to generate the word sequence that corresponds to the audio features and the text semantics.
5. The method according to claim 1, characterized in that, The content information associated with the offending content segment of the original audio data includes the timestamp of the offending content, the location of the audio segment, and the type of violation.
6. The method according to claim 1, characterized in that, Based on the content information associated with the violating content fragments, the violating content fragments of the original audio data are processed using an audio compliance enhancement model to obtain compliant audio data, including: The original audio data is subjected to voice separation processing to obtain independent voice channels and background sound channels; Based on the location of the audio segment containing the infringing content, the audio compliance enhancement model is used to process the infringing content segment in the human voice channel; and The processed human voice channel is merged with the background sound channel to obtain the compliant audio data.
7. The method according to claim 6, characterized in that, Based on the location of the audio segment containing the infringing content, the processing operation performed on the infringing content segment in the human voice channel by the audio compliance enhancement model includes: The audio compliance enhancement model performs at least one of the following processing steps on the illegal content segments in the human voice channel: fade-out removal and compliance content enhancement implantation. The fade-out process involves using a deep learning audio enhancement model to reduce the clarity and identifiability of the infringing content while preserving the basic characteristics of human voices. The aforementioned compliance content enhancement implantation process includes implanting preset compliance audio content at specific audio segment locations.
8. The method according to claim 7, characterized in that, Deep learning audio enhancement models are used to reduce the clarity and recognizability of inappropriate content while preserving basic human voice characteristics, including: Selective frequency attenuation and phase perturbation are applied to the audio signal in the human voice channel corresponding to the illegal content segment. The selective frequency attenuation targets the high-frequency and transient frequency components in the audio signal that characterize semantic clarity.
9. The method according to claim 7, characterized in that, Inserting pre-defined compliant audio content at specific audio segment locations includes: Perform acoustic environment matching processing on the preset compliant audio content; and In the audio segment where the preset compliant audio content and human voice coexist, the loudness of the human voice channel is dynamically adjusted.
10. The method according to claim 1, characterized in that, The output of the compliant audio data includes: The compliant audio data is encoded and packaged into streaming media to generate a compliant live stream; and The compliant live stream is pushed to the live streaming server.
11. The method according to claim 2, characterized in that, Also includes: Obtain the newly added library of illegal audio samples and their corresponding illegal text tags; as well as Based on the newly added illegal audio sample library, the pre-trained audio compliance detection model is incrementally trained.
12. An apparatus for compliant audio processing, characterized in that, include: The audio acquisition module is configured to acquire the raw audio data to be processed. The audio compliance detection module is configured to detect the original audio data through an audio compliance detection model in order to generate content information associated with the non-compliant content segments of the original audio data; An audio compliance enhancement module is configured to process the non-compliant content segment of the original audio data using an audio compliance enhancement model based on content information associated with the non-compliant content segment, to obtain compliant audio data; and The audio output module is configured to output the compliant audio data.
13. An electronic device, characterized in that, include: The memory is configured to store executable programs; as well as A processor is configured to execute the program to perform the method according to any one of claims 1 to 11.