Live streaming processing method, device, electronic device and storage medium

By performing frame processing and asynchronous sensitive word detection and replacement of live streams, the problem of poor real-time and stability of live streams in the prior art is solved, and sensitive speech intervention without delay is achieved, and the real-time stability of live streams is improved.

CN115767111BActive Publication Date: 2025-08-15GUANGZHOU BOGUAN TELECOMM TECH LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210962312.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2025-08-15
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

The existing live broadcast sensitive speech intervention technology requires three steps: tandem review, silencing and audio-visual synchronization, resulting in the real-time and stability of the live broadcast stream, and is prone to problems such as lag, disconnection and excessive delay.

Method used

By performing frame-based processing on the live stream, the position information of sensitive words in the audio stream is determined, and frame-by-frame audio is replaced asynchronously through the preset sliding window to avoid synchronous blockage between steps, and sensitive words are detected and replaced by asynchronous processing.

Benefits of technology

It improves the real-time and stability of live stream output, reduces the impact of sensitive speech intervention on live stream, and ensures the delay-free silence and stable output of live stream.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115767111B_ABST
    Figure CN115767111B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, electronic device, and storage medium for processing a live stream. The method comprises: in response to an inflow of a live stream generated by a host terminal, performing frame-by-frame processing on an audio stream in the live stream to obtain multiple audio frames; determining sensitive word position information in the audio stream, and asynchronously performing sensitive word frame-by-frame audio replacement on the audio frames through a preset sliding window based on target audio replacement; and outputting the live stream after audio replacement. In this method, after the live stream is inflowed, the audio stream in the live stream is divided into audio frames to form multiple audio frames, and then sensitive word positions are determined and sliding audio replacement is performed on an audio frame-by-audio frame basis in an asynchronous manner, so that the sensitive word determination process and the replacement process do not block each other. Furthermore, the audio replacement of the sliding window is performed with the frame level as the audio processing unit, which can ensure that the replacement audio and audio frames are silenced without delay, thereby improving the real-time performance and stability of the live stream output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of live stream processing, and in particular to a live stream processing method, device, electronic device and storage medium. Background Art

[0002] With the growing development of the live streaming industry, more and more broadcasters and users are interacting in real time through this real-time communication medium. However, broadcasters' emotional changes during live broadcasts are difficult to predict and intervene in advance. Driven by extreme emotions, they may engage in inappropriate words and deeds, causing undue impact on society and minors. Therefore, intervening in sensitive speech during live broadcasts is the premise and foundation for preventing the negative impact of live broadcasts.

[0003] Existing livestream intervention techniques typically involve first reviewing sensitive speech, then muting the audio stream within a specific timeframe. Finally, the audio and video are synchronized based on timestamp alignment, encoded, and then delivered to the user. This approach requires a cascade of review, muting, and synchronization, and the synchronization between each step is often blocked, compromising the real-time and stability of the livestream. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a live stream processing method, device, electronic device and storage medium, which are used to ensure the real-time stability of the live stream output when sensitive speech intervention is performed on the live stream.

[0005] In a first aspect, an embodiment of the present invention provides a method for processing a live stream, the method comprising: in response to an inflow of a live stream generated by a host end, performing frame processing on an audio stream in the live stream to obtain multiple audio frames; determining the position information of sensitive words in the audio stream, and based on target audio replacement, asynchronously performing sensitive word frame-by-frame audio replacement on the audio frames through a preset sliding window; and outputting the live stream after audio replacement.

[0006] The above-mentioned determination of the sensitive word position information in the audio stream, and based on the target audio replacement, asynchronously performing sensitive word frame-by-frame audio replacement on the audio frame through a preset sliding window includes: sending the audio stream to the review server through a first process; receiving the sensitive word position information sent by the review server in real time through a second process, and based on the target audio replacement, asynchronously performing sensitive word frame-by-frame audio replacement on the audio frame through a preset sliding window.

[0007] The above-mentioned sending of the audio stream to the audit server through the first process includes: decapsulating and decoding the audio stream through the first process to obtain first pulse code modulation data of the audio stream; encrypting the first pulse code modulation data, and sending the encrypted first pulse code modulation data to the audit server through the first process, so that the audit server identifies sensitive words on the encrypted first pulse code modulation data, and sends the identified sensitive word location information to the second process.

[0008] The above-mentioned target-based audio replacement method asynchronously replaces the audio frames with sensitive words frame by frame through a preset sliding window, including: asynchronously inputting the audio frames into a preset audio queue; when the number of audio frames in the preset audio queue reaches the maximum audio frame capacity, based on the sensitive word position information, slidingly taking values of the target replacement audio through a preset sliding window, and replacing the values with the sensitive word audio frames in the preset audio queue during the sliding process; the above-mentioned sensitive word audio frames are used to: indicate the audio frames where the sensitive words indicated by the sensitive word position information are located.

[0009] The above-mentioned method is based on the sensitive word position information, and the target replacement audio is slidingly valued through a preset sliding window, and during the sliding value process, the value is replaced with the sensitive word audio frame in the preset audio queue, including: according to the sensitive word position information, determining the audio frame where the sensitive word is located in the preset audio queue to obtain the sensitive word audio frame; according to the data length of the sensitive word audio frame, determining the maximum sliding distance of the preset sliding window, and sliding the target replacement audio frame by frame according to the maximum sliding distance; the size of the sliding window is the same as the length of an audio frame; during the frame-by-frame sliding process of the sliding window, the value is replaced with the sensitive word audio frame in the preset audio queue.

[0010] In the above-mentioned response to the inflow of the live stream generated by the anchor end, before the audio stream in the live stream is framed to obtain multiple audio frames, the method also includes: generating a target replacement audio with a preset maximum audio replacement length according to preset waveform parameters.

[0011] In the above-mentioned response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and before the sensitive word position information in the audio stream is determined and the audio frames are asynchronously replaced with sensitive words frame by frame through a preset sliding window based on the target replacement audio, the method also includes: extracting audio parameters from the live stream to obtain target audio parameters; decoding and resampling the target replacement audio to obtain second pulse code modulation data of the target replacement audio; setting the audio parameters of the second pulse code modulation data to the target audio parameters to obtain the target replacement audio with the same audio parameters.

[0012] In the above-mentioned response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and before the sensitive word position information in the audio stream is determined and the audio frame is asynchronously replaced with sensitive words frame by frame through a preset sliding window based on the target replacement audio, the method also includes: volume prediction through amplitude data in the first pulse code modulation data of the audio frame to obtain target volume data; setting the target volume data on the second pulse code modulation data to obtain target replacement audio with the same volume.

[0013] The above-mentioned framing processing of the audio stream in the live stream to obtain multiple audio frames includes: decapsulating and decoding the audio stream in the live stream to obtain first pulse code modulation data of the audio stream; and dividing the first pulse code modulation data into audio frames in real time according to a preset sampling rate to obtain multiple audio frames.

[0014] The above-mentioned output of the live stream after audio replacement includes: encoding the audio stream after audio replacement to obtain an encoded audio stream; encapsulating the encoded audio stream and the video stream in the live stream into a live stream, and outputting the encapsulated live stream.

[0015] In the second aspect, an embodiment of the present invention provides a live stream processing device, which includes: a response module, which is used to respond to the inflow of the live stream generated by the anchor end, and perform frame processing on the audio stream in the live stream to obtain multiple audio frames; a replacement module, which is used to determine the position information of sensitive words in the audio stream, and based on the target replacement audio, asynchronously replace the sensitive words in the audio frames frame by frame through a preset sliding window; an output module, which is used to output the live stream after audio replacement.

[0016] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned live stream processing method.

[0017] In a fourth aspect, an embodiment of the present invention provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the above-mentioned live stream processing method.

[0018] The embodiments of the present invention bring the following beneficial effects:

[0019] The live stream processing method, device, electronic device, and storage medium described above, in response to the inflow of a live stream generated by a host end, performs frame processing on the audio stream in the live stream to obtain multiple audio frames; determines the position information of sensitive words in the audio stream, and based on the target audio replacement, asynchronously performs sensitive word frame-by-frame audio replacement on the audio frames using a preset sliding window; and outputs the live stream after the audio replacement. In this method, after the live stream is inflow, the audio stream in the live stream is divided into audio frames to form multiple audio frames. The sensitive word positions are then determined and the sliding audio replacement is performed frame by frame asynchronously, so that the sensitive word determination process and the replacement process do not block each other. The sliding window audio replacement is performed using the frame level as the audio processing unit, which can ensure that the replacement audio and audio frames are silenced without delay, thereby improving the real-time performance and stability of the live stream output.

[0020] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 This is a flow chart of an embodiment of a method for processing a live stream according to an embodiment of the present invention;

[0024] Figure 2 A schematic diagram of a live stream processing device provided by an embodiment of the present invention;

[0025] Figure 3A schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0027] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0028] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 An embodiment of a method for processing a live stream in an embodiment of the present invention includes:

[0029] Step S10: In response to the inflow of the live stream generated by the anchor terminal, the audio stream in the live stream is framed to obtain multiple audio frames;

[0030] It is understandable that during a live broadcast, the interaction of live data typically involves a client and a server, wherein the client is used to receive and send live data, and the server is used to process live data. The client includes a host and an audience. In this step, the server responds to the inflow of the live stream generated by the host, replaces the audio of the live stream with sensitive words, and then outputs the live stream with audio replacement to the audience, so that the live stream played on the audience is the live stream after the sensitive speech intervention, thereby avoiding the improper impact of sensitive speech on the audience. In one embodiment, a persistent long link is established between the client and the server based on a full-duplex communication protocol, so that the client and the server can perform two-way data transmission and asynchronous data interaction.

[0031] It should be noted that most existing technologies for intervening in sensitive speech during live broadcasts are stream-level synchronous blocking schemes. First, the audio stream needs to be reviewed for sensitive speech. After the review results are determined, the audio stream is muted within the time range of the sensitive speech. Finally, the muted audio stream is timestamped with the video stream to synchronize the audio and video. In this scheme, since the review, muting, and timestamp alignment steps are connected in series, the next step can only be executed after the previous step is completed. At the same time, since the frequency and duration of sensitive speech made by the anchor vary, the processing time for the sensitive speech review is uncontrollable, making the live stream prone to synchronous blocking with uncontrollable blocking duration. Moreover, any anomaly in any of these steps may cause anomalies in the entire live stream, which in turn may lead to freezes, disconnections, excessive delays, or even crashes. Therefore, the existing technology has technical problems with poor live stream output stability and real-time performance. Based on this, the present invention proposes a frame-level asynchronous processing method for live streaming, which asynchronously performs sensitive speech review and shielding. In this method, since the intervention process of sensitive speech does not cause synchronous blocking, there is no need to align the timestamps of the audio stream and the video stream, which greatly reduces the impact of the intervention processing of sensitive speech on the live streaming, thereby improving the real-time and stability of the live streaming.

[0032] In this step, since both the live stream and the audio stream are streaming data, there is no clear concept and standard of frames. Therefore, after the live stream flows in, the audio stream in the live stream is framed according to the preset frame rate to obtain multiple audio frames, so that the subsequent processing steps are all based on the audio frame as the processing unit and processing object, achieving the live stream processing effect at the frame level. It can be understood that this step is based on the audio stream. It does not generate new data and does not change the data transmission method of the live stream. It only clarifies the length, range and standard of the audio frame in the audio stream, so that the data processing unit of the audio stream is reduced and the data processing efficiency is improved. It should be noted that the preset frame rate can be set according to the characteristics of the audio encoder and decoder, as well as the business scenario, and is not limited here.

[0033] In one embodiment, since custom rate effects are usually added to audio or video during live broadcasts, before responding to the inflow of the live stream generated by the anchor end, the system's underlying frame-by-frame processing interface is encapsulated based on the system's audio and video filter module, so that the sensitive word audio is subsequently replaced frame by frame through the encapsulated frame-by-frame processing interface. In this embodiment, since the frame-by-frame processing interface is an interface at the bottom level of the system, no data accumulation and caching are generated, thereby improving the real-time performance of live stream processing. In addition, since the frame-by-frame processing interface has strong versatility, based on the frame-by-frame processing interface, audio frames can also be subjected to frame-level processing such as audio enhancement or voice change effects, thereby improving the accuracy of sensitive word detection or the audio effect of live stream output.

[0034] In one embodiment, the audio frame is pulse code modulation (PCM) data, which is a type of coded data for digital communication, also known as raw audio data. In this embodiment, the PCM data can be obtained by decoding the audio stream through a preset decoder, or it can be directly derived from a system interface, the specific details of which are not limited here. Live stream processing based on PCM data can improve the processing efficiency of the live stream. Based on this embodiment, the above-mentioned frame processing of the audio stream in the live stream to obtain multiple audio frames includes: decapsulating and decoding the audio stream in the live stream to obtain the first pulse code modulation data of the audio stream; and dividing the first pulse code modulation data into audio frames in real time according to a preset sampling rate to obtain multiple audio frames. In this embodiment, in order to improve the efficiency of subsequent audio replacement, after the live stream flows in, the audio stream in the live stream is immediately decapsulated and decoded to obtain the PCM data of the audio stream (i.e., the first pulse code modulation data), and then dividing the first pulse code modulation data into audio frames in real time according to a preset sampling rate to obtain multiple audio frames. It can be seen that the essence of the audio frame is pulse code modulation data. The sampling rate refers to the sampling frequency, also known as the sampling speed, which is used to indicate the number of samples extracted from a continuous signal and composed of a discrete signal per unit time, such as 60Hz.

[0035] Step S20: Determine the location information of sensitive words in the audio stream, and based on the target audio replacement, asynchronously replace the sensitive words frame by frame on the audio frames through a preset sliding window;

[0036] It can be understood that frame-by-frame audio replacement of sensitive words refers to replacing the sensitive word audio frames in the audio stream with the target replacement audio frame by frame, wherein the sensitive word audio frames refer to the audio frames indicated by the sensitive word position information. For example, if the sensitive word position information indicates that the 5th and 6th frames in the audio stream are sensitive words, then the 5th and 6th frames are the sensitive word audio frames.

[0037] In this embodiment, the steps of determining the position information of sensitive words in the audio stream and sliding the audio stream frame by frame through a preset sliding window and replacing the audio of sensitive words with a target frame by frame are asynchronously executed. That is, the determination of the position information of sensitive words in the audio stream does not block the audio replacement of sensitive words. In one embodiment, an asynchronous thread is created to determine the position information of sensitive words in the audio stream; a main thread that outputs the live stream receives a real-time message sent by the asynchronous thread, wherein the real-time message includes information indicating whether each audio frame in the audio stream is a sensitive word; based on the received real-time message, the main thread slides the audio stream frame by frame through a preset sliding window and replaces the audio of sensitive words with a target frame by frame, and outputs the live stream after the audio replacement. The main thread continuously receives the sensitive word position information and performs audio replacement processing on the audio stream in the process of continuous output, thereby reducing the impact of sensitive word interference processing on the output of the live stream, preventing blocking between threads, and the processing processes do not affect each other, thereby improving the stability of the live stream output.

[0038] In one embodiment, determining the sensitive word position information in the audio stream includes: performing speech recognition on multiple audio frames through a preset speech recognition algorithm to obtain a speech recognition result, and then performing sensitive word recognition on the speech recognition result through a preset sensitive word database to obtain sensitive word position information, wherein the sensitive word position information can be used to indicate the audio frame sequence position of the sensitive word in the audio stream (relative to the sequence number of the first audio frame), and can also be used to indicate the temporal position of the occurrence of the sensitive word, such as the millisecond-level Greenwich time stamp, custom timestamp, etc. In another embodiment, sensitive word detection can also be performed on multiple audio frames through a pre-trained neural network model to obtain sensitive word position information. This embodiment can improve the accuracy and efficiency of sensitive word detection by means of a speech recognition algorithm or model, thereby improving the real-time and accuracy of sensitive speech intervention.

[0039] It should be noted that the target replacement audio can be a system preset audio, or it can be an audio file customized by the client for audio replacement. In one embodiment, the target replacement audio is also PCM data. Before responding to the inflow of the live stream generated by the anchor end, the preset replacement audio file is decoded to obtain the target replacement audio in PCM data format, so that when the sensitive word audio of the target replacement audio is subsequently replaced, the sensitive word audio replacement is performed based on the target replacement audio in PCM data format. In one embodiment, in order to improve the efficiency and accuracy of audio replacement, the target replacement audio in PCM data format can also be frame-processed to obtain multiple replacement audio frames, so that when performing sensitive word audio replacement of the target replacement audio frame by frame, the audio frame and the replacement audio frame are slid frame by frame and the sensitive word audio is replaced frame by frame through a preset sliding window to obtain the live stream after audio replacement. This embodiment can quickly replace audio frame by frame based on the basic encoding format data of the audio, thereby improving the real-time and stability of the live stream.

[0040] It is understandable that the sliding window is a flow control technology that forms a window of the size of an audio frame by pointing to different audio frames in the audio stream or target replacement audio through a double pointer. By adjusting the pointing of the double pointer, the size of the sliding window can be adjusted, making the flow control of the audio frame more flexible. In one embodiment, the size of the sliding window is the same as the length of an audio frame. For example, there are 60 audio frames in 1 second, so the size of the sliding window is 1 / 60 second. In order to improve the real-time performance of audio replacement, the target replacement audio is slid and valued through a preset sliding window to obtain a replacement audio frame, and based on the sensitive word position information, the replacement audio frame is replaced with the audio frame at the location of the sensitive word to obtain a live stream after audio replacement. During the replacement process, the replacement audio frame and the audio frame have the same data structure. In one embodiment, the replacement audio frame and the audio frame are both PCM data. This embodiment can achieve delay-free audio replacement, thereby improving the real-time performance and stability of the live stream output.

[0041] In one embodiment, the preset sliding window may include a first sliding window and a second sliding window, wherein the first sliding window is used to take values of the target replacement audio, and the second sliding window is used to limit the range of the sensitive word audio frame in the audio stream, wherein the sensitive word audio frame refers to the audio frame where the sensitive word indicated by the sensitive word position information is located. Specifically, the pointer of the second sliding window in the audio stream is determined to point to the audio frame based on the sensitive word position information, and then the maximum sliding distance of the first sliding window is determined based on the pointer pointing to the audio frame, and based on the maximum sliding distance, the target replacement audio is taken values through the first sliding window, and the replacement audio obtained by taking the values is replaced with the audio frame pointed to by the pointer of the second sliding window to obtain the live stream after the audio replacement.

[0042] Step S30: output the live stream after audio replacement.

[0043] In this step, the audio frames after audio replacement are first encoded and encapsulated to obtain an audio stream after audio replacement. The audio stream after audio replacement is then merged with the video stream to output the live stream after audio replacement. In this embodiment, since the audio stream and the video stream are always synchronized during the audio replacement process, there is no need to perform related alignment operations such as timestamp alignment during the merging of the audio stream and the video stream. This reduces the impact of the audio replacement process on the output of the live stream, making it possible to intervene in sensitive speech in the live stream while ensuring the real-time stability of the live output.

[0044] In one embodiment, step S30 includes: encoding the audio stream after the audio replacement to obtain an encoded audio stream; encapsulating the encoded audio stream and the video stream in the live stream into a live stream, and outputting the encapsulated live stream. It is understood that the audio stream after the audio replacement has been processed by replacing sensitive words. The audio stream after the audio replacement is encoded using a preset encoder, and the encoded audio stream and video stream are then encapsulated into a live stream, and the live stream can be output to the client.

[0045] The live stream processing method provided by the above embodiment divides the audio stream in the live stream into audio frames after the live stream flows in to form multiple audio frames, and then determines the position of sensitive words and performs sliding audio replacement on each audio frame in an asynchronous manner, so that the sensitive word determination program and the replacement program will not block each other, and performs sliding window audio replacement with the frame level as the audio processing unit, which can ensure delay-free silencing of the replaced audio and audio frames, thereby improving the real-time and stability of the live stream output.

[0046] One of the asynchronous processing methods is further described below. In one embodiment, the asynchronous processing is performed in a dual-process manner. The above step S20 includes the following steps:

[0047] (1) Sending the audio stream to the audit server through the first process;

[0048] In one embodiment, in response to the influx of the live stream generated by the anchor end, two processes are created: a first process and a second process. The inputs of these two processes are both audio streams, wherein the first process is the review process and the second process is the audio replacement process. It can be understood that the first process is used for sensitive word review and the second process is used for sensitive word audio replacement. In the review process, the input and output of each audio frame are consistent, no special processing is performed, and only data is sent, that is, multiple audio frames in the audio stream are sent to the review server through the first process; in the audio replacement process, the audio frame review result received from the review process by the message queue is used to determine whether the input audio frame is a sensitive word audio frame. If the input audio frame is a non-sensitive word audio frame, then the input and output of the audio replacement process are also consistent. If the input audio frame is a sensitive word audio frame, the preset sliding window is used to perform frame-by-frame audio replacement of the target replacement audio on the sensitive word audio frame, thereby obtaining a live stream after audio replacement.

[0049] In one embodiment, the first process and the second process can be created on the server side or on the client side, and the specific details are not limited here. The input of the first process can be an unprocessed audio stream or PCM data that has been decapsulated and decoded. The first process can also pre-process the audio stream. Specifically, the above step (1) includes: decapsulating and decoding the audio stream through the first process to obtain the first pulse code modulation data of the audio stream; encrypting the first pulse code modulation data, and sending the encrypted first pulse code modulation data to the audit server through the first process, so that the audit server can identify sensitive words on the encrypted first pulse code modulation data and send the identified sensitive word location information to the second process.

[0050] In this embodiment, the input of the first process is an unprocessed audio stream. The first process decapsulates and decodes the audio stream to obtain the PCM data of the audio stream (i.e., first pulse code modulation data). The first pulse code modulation data is then encrypted using a preset encryption algorithm, and the encrypted first pulse code modulation data is sent to the audit server. The encryption algorithm is a stream cipher algorithm, a symmetric encryption algorithm. The encryption and decryption ends use the same pseudo-random encrypted data stream as the key. The plaintext data and the encrypted data stream are encrypted sequentially to obtain a ciphertext data stream. The stream cipher algorithm can improve the security of audio stream transmission, thereby ensuring the security of live broadcast data.

[0051] In this embodiment, the review server can be any node server in the distributed system or any target server in the cluster system. When the live stream starts to flow in, a review server is determined through a preset load balancing algorithm to perform sensitive word review on the audio stream to improve the load capacity and processing efficiency of the live stream processing.

[0052] (2) The sensitive word location information sent by the audit server is received in real time through the second process, and based on the target audio replacement, the sensitive word audio frame is asynchronously replaced through a preset sliding window.

[0053] In one embodiment, the second process receives the sensitive word location information sent by the audit server in real time through a message queue, and completes the frame-by-frame audio replacement of the sensitive word in the second process. In this embodiment, the second process determines whether the input audio frame is a sensitive word audio frame based on the audio frame audit results received from the audit process in real time by the message queue. If the input audio frame is a non-sensitive word audio frame, then the input and output of the second process are consistent and no special processing is performed. If the input audio frame is a sensitive word audio frame, the sensitive word audio frame is subjected to frame-by-frame audio replacement of the target replacement audio through a preset sliding window, thereby obtaining a live stream after the audio replacement.

[0054] In one embodiment, the input to the second process is an audio frame, which can be an unprocessed audio frame or an audio frame in a PCM data format that has been decapsulated and decoded, and the specifics are not limited here. It is understood that there is a certain time difference between the second process and the first process, with the first process being ahead of the second process in time, and this time difference can be controlled by a preset delay duration. The longer the preset delay duration, the more accurate the detection of sensitive word location information. After testing, the preset delay duration of 7 seconds meets the accuracy requirements of sensitive word detection in most live broadcast scenarios.

[0055] One of the implementation methods of the sliding window is described below. In one implementation method, the above-mentioned target replacement audio is based on asynchronously replacing the audio frames with sensitive words frame by frame through a preset sliding window, including: asynchronously inputting audio frames into a preset audio queue; when the number of audio frames in the preset audio queue reaches the maximum audio frame capacity, based on the sensitive word position information, the target replacement audio is slidingly valued through a preset sliding window, and in the process of sliding value taking, the value is replaced with the sensitive word audio frame in the preset audio queue; the above-mentioned sensitive word audio frame is used to: indicate the audio frame where the sensitive word indicated by the sensitive word position information is located.

[0056] It can be understood that in order to improve the controllability of the delay duration and delay rate of the live stream when sensitive remarks are intervened in the live stream, the delay duration is controlled by a preset audio queue. Specifically, the maximum audio frame capacity of the preset audio queue is used to indicate the delay duration of the live stream. In this embodiment, the audio stream in the live stream is framed to obtain multiple audio frames, and then the audio frames are input into the preset audio queue. When the number of audio frames in the preset audio queue is equal to the maximum audio frame capacity, all audio frames in the preset audio queue are dequeued, and based on the sensitive word position information, the dequeued audio frames are replaced with sensitive word audio frames frame by frame in the process of sliding the sliding window to obtain the live stream after audio replacement. For example, assuming that the maximum audio frame capacity of the preset audio queue is 7 seconds, then when the audio frames asynchronously input to the preset audio queue are full of 7 seconds (the number of audio frames within 7 seconds can be calculated based on the length of each audio frame), all the audio frames of these 7 seconds are dequeued. After dequeuing, the audio frames within these 7 seconds are replaced frame by frame using a sliding window based on whether the audio frames within these 7 seconds contain sensitive word audio frames. At the same time, after all the audio frames of these 7 seconds are dequeued, the preset audio queue immediately continues to queue for the next round of audio frame accumulation, thereby ensuring the controllability of the delay of the live stream. In this embodiment, audio frames are accumulated through the audio queue, and during the accumulation process, sensitive word position information is asynchronously obtained. After the accumulation is full, the accumulated audio frames are replaced frame by frame based on the sensitive word position information, so that the delay duration of the live stream is always controlled within the stacking number of the audio queue, and no additional uncontrollable delay is generated, thereby improving the stability and real-time performance of the live stream.

[0057] Furthermore, in one embodiment, the above-mentioned method is based on the sensitive word position information, and the target replacement audio is slid and valued through a preset sliding window, and in the process of sliding value, the value is replaced with the sensitive word audio frame in the preset audio queue, including: according to the sensitive word position information, determining the audio frame where the sensitive word is located in the preset audio queue, and obtaining the sensitive word audio frame; according to the data length of the sensitive word audio frame, determining the maximum sliding distance of the preset sliding window, and sliding the target replacement audio frame by frame according to the maximum sliding distance; the size of the sliding window is the same as the length of an audio frame; in the process of sliding the sliding window frame by frame, the value is replaced with the sensitive word audio frame in the preset audio queue.

[0058] In this embodiment, after the preset audio queue is full of audio frames, all audio frames are dequeued, and then, based on the sensitive word position information, the audio frame containing the sensitive word among all the dequeued audio frames is determined to obtain the sensitive word audio frame, wherein the indication information of the sensitive word audio frame can be timestamp timing information or frame sequence information, such as the 8th audio frame is a sensitive word audio frame, the 21st audio frame is a sensitive word audio frame, etc. Then, based on the data length of the sensitive word audio frame, the target replacement audio of the same data length is read, or the maximum sliding distance of the preset sliding window is set, wherein the data length is used to indicate the time length or frame sequence length of the sensitive word audio frame, such as 2 seconds or 8 frames, which are not specifically limited here. If the target replacement audio of the same data length is read, then there is no need to set the maximum sliding distance of the sliding window. The sliding window only needs to be slid from the starting point of the target replacement audio to the end point to obtain the value. If the preset maximum sliding distance of the sliding window is set, then the sliding window is slid from the starting point of the target replacement audio by the maximum sliding distance to obtain the value. The maximum sliding distance can be a time length, such as 2 seconds, or a frame sequence length, such as 8 frames. The specific details are not limited here. It can be understood that the size of the sliding window is the same as the length of an audio frame. For example, if the length of an audio frame is 0.1 seconds, then the size of the sliding window is also 0.1 seconds. The size of the sliding window is used to indicate the unit sliding distance of the sliding window, that is, each sliding of the sliding window slides the length of an audio frame, so that the unit sliding distance of the sliding window is consistent with the size of the audio frame in the audio stream, thereby improving the real-time and delay-free performance of the audio replacement.

[0059] A preprocessing method for replacement audio is described below. In one embodiment, after responding to the influx of the live stream generated by the anchor end, the target replacement audio can be preprocessed according to the relevant audio parameters of the live stream, or the target replacement audio can be generated according to the characteristics of the replacement audio to make the target replacement audio and the audio of the live stream more smoothly connected to avoid abruptness. The replacement audio can also be flexibly generated without replacing the audio material, making the setting of the replacement audio more flexible.

[0060] In one embodiment, in response to the inflow of the live stream generated by the host end, the audio stream in the live stream is framed and processed to obtain multiple audio frames. The method also includes: generating a target replacement audio with a preset maximum audio replacement length according to preset waveform parameters; before responding to the inflow of the live stream generated by the host end, the host end can select a sensitive word to replace the sound type, such as beep audio, noise audio. In the case that the host end does not select a sensitive word to replace the sound type, the default sensitive word replacement sound type is used, and audio replacement can be performed without uploading an audio file. In this embodiment, in the case of no replacement of audio material, the target replacement audio is generated according to the characteristics of different sounds. Specifically, the waveform parameters of the target replacement sound type are obtained, wherein the waveform parameters include but are not limited to waveform, amplitude, phase and other parameters used to express sound. According to these waveform parameters, a wave with a specific waveform, amplitude, phase, etc. is generated to express the sound, thereby obtaining the target replacement audio. It can be understood that the preset maximum audio replacement length refers to the maximum continuous audio replacement duration or frame length. For example, the maximum allowed continuous audio replacement is 10 seconds. The setting of the maximum audio replacement length can determine the maximum boundary of the replacement audio. When the maximum boundary is exceeded, the audio replacement can be performed through loop playback, which can ensure the audio replacement effect while reducing the storage space of the audio file, thereby improving the reading efficiency of the replacement audio.

[0061] In one embodiment, for a specific target replacement audio, the target replacement audio is preprocessed according to the relevant audio parameters of the live stream so that the target replacement audio is consistent with the sound characteristics of the live stream, avoiding the abruptness of the audio replacement. Specifically, in response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and then the sensitive word position information in the audio stream is determined. Based on the target replacement audio, before the audio frame is asynchronously replaced with sensitive words frame by frame through a preset sliding window, the method also includes: extracting audio parameters from the live stream to obtain target audio parameters; decoding and resampling the target replacement audio to obtain second pulse code modulation data of the target replacement audio; setting the audio parameters of the second pulse code modulation data to the target audio parameters to obtain the target replacement audio with the same audio parameters.

[0062] In this embodiment, audio parameters include but are not limited to audio inherent characteristic parameters such as sampling rate, number of sampling points, audio encoding format, and number of channels. In order to keep the audio parameters of the target replacement audio consistent with those of the audio frame, after the live stream flows in, the audio parameters of the live stream are extracted to obtain the target audio parameters, and then the preset target replacement audio is decoded and resampled to obtain the second PCM data of the target replacement audio, and the audio parameters of the second PCM data are set as the target audio parameters, thereby obtaining the target replacement audio with the same audio parameters, so that the audio frames can be subsequently replaced with sensitive words of the target replacement audio with the same audio parameters frame by frame through a preset sliding window.

[0063] In one embodiment, for some non-audio inherent characteristic parameters, such as sound parameters that can be freely set, such as volume and tone, these sound parameters are predicted after the live stream flows in. Specifically, in response to the inflow of the live stream generated by the host end, the audio stream in the live stream is framed and processed to obtain multiple audio frames. The sensitive word position information in the audio stream is determined, and based on the target replacement audio, before the audio frame is asynchronously replaced with sensitive words frame by frame through a preset sliding window, the method also includes: performing volume prediction based on the amplitude data in the first pulse code modulation data of the audio frame to obtain target volume data; setting the target volume data for the second pulse code modulation data to obtain the target replacement audio with the same volume. In this embodiment, the audio frame can be volume predicted based on the amplitude data in the PCM data of the audio frame to obtain the target volume data, and then the target volume data can be set for the PCM data of the target replacement audio to obtain the target replacement audio with the same volume, so that the audio frame can be subsequently replaced with sensitive words frame by frame through a preset sliding window with the same audio parameters and the same volume.

[0064] Corresponding to the above method embodiment, see Figure 2 A schematic diagram of a live stream processing device is shown, which includes: a response module 20, which is used to respond to the inflow of the live stream generated by the anchor end, and perform frame processing on the audio stream in the live stream to obtain multiple audio frames; a replacement module 22, which is used to determine the position information of sensitive words in the audio stream, and based on the target replacement audio, asynchronously replace the audio frames with sensitive words frame by frame through a preset sliding window; an output module 24, which is used to output the live stream after the audio is replaced.

[0065] The above-mentioned live stream processing device, after the live stream flows in, divides the audio stream in the live stream into audio frames to form multiple audio frames, and then determines the sensitive word position and performs sliding audio replacement on each audio frame in an asynchronous manner, so that the sensitive word determination program and the replacement program will not block each other, and performs sliding window audio replacement with the frame level as the audio processing unit, which can ensure delay-free silencing of the replaced audio and audio frames, thereby improving the real-time and stability of the live stream output.

[0066] The above-mentioned replacement module is also used to: send the audio stream to the review server through the first process; receive the sensitive word location information sent by the review server in real time through the second process, and replace the audio based on the target, and asynchronously replace the sensitive word audio frame by frame through a preset sliding window.

[0067] The above-mentioned replacement module is also used to: decapsulate and decode the audio stream through the first process to obtain the first pulse code modulation data of the audio stream; encrypt the first pulse code modulation data, and send the encrypted first pulse code modulation data to the audit server through the first process, so that the audit server can identify sensitive words in the encrypted first pulse code modulation data, and send the identified sensitive word location information to the second process.

[0068] The above-mentioned replacement module is also used to: asynchronously input audio frames into the preset audio queue; when the number of audio frames in the preset audio queue reaches the maximum audio frame capacity, based on the sensitive word position information, the target replacement audio is slid and valued through a preset sliding window, and during the sliding value-taking process, the value is replaced with the sensitive word audio frame in the preset audio queue; the above-mentioned sensitive word audio frame is used to: indicate the audio frame where the sensitive word indicated by the sensitive word position information is located.

[0069] The above-mentioned replacement module is also used to: determine the audio frame where the sensitive word is located in the preset audio queue according to the sensitive word position information, and obtain the sensitive word audio frame; determine the maximum sliding distance of the preset sliding window according to the data length of the sensitive word audio frame, and slide the target replacement audio frame by frame according to the maximum sliding distance; the size of the sliding window is the same as the length of an audio frame; during the frame-by-frame sliding process of the sliding window, the value is replaced to the sensitive word audio frame in the preset audio queue.

[0070] The above-mentioned device also includes: a generating module, which is used to generate a target replacement audio with a preset maximum audio replacement length according to preset waveform parameters.

[0071] The above-mentioned device also includes: a parameter synchronization module, which is used to: extract audio parameters of the live stream to obtain target audio parameters; decode and resample the target replacement audio to obtain second pulse code modulation data of the target replacement audio; set the audio parameters of the second pulse code modulation data as the target audio parameters to obtain the target replacement audio with the same audio parameters.

[0072] The above-mentioned device also includes: a volume prediction module, which is used to: predict the volume through the amplitude data in the first pulse code modulation data of the audio frame to obtain target volume data; set the target volume data for the second pulse code modulation data to obtain a target replacement audio with the same volume.

[0073] The above-mentioned response module is also used to: decapsulate and decode the audio stream in the live stream to obtain the first pulse code modulation data of the audio stream; divide the first pulse code modulation data into audio frames in real time according to a preset sampling rate to obtain multiple audio frames.

[0074] The above-mentioned output module is also used to: encode the audio stream after audio replacement to obtain an encoded audio stream; encapsulate the encoded audio stream and the video stream in the live stream into a live stream, and output the encapsulated live stream.

[0075] This embodiment further provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above-mentioned live stream processing method. The electronic device can be a server or a terminal device.

[0076] See also Figure 3 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100. The processor 100 executes the machine-executable instructions to implement the above-mentioned live stream processing method.

[0077] Furthermore, Figure 3 The electronic device shown further includes a bus 102 and a communication interface 103 , and the processor 100 , the communication interface 103 and the memory 101 are connected via the bus 102 .

[0078] The memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface 103 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0079] The processor 100 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 100 or software instructions. The above processor 100 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register, which is well-known in the art. The storage medium is located in the memory 101. The processor 100 reads the information in the memory 101 and, in conjunction with its hardware, performs the steps of the method of the aforementioned embodiment, for example:

[0080] In response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is framed to obtain multiple audio frames; the sensitive word position information in the audio stream is determined, and based on the target audio replacement, the audio frames are asynchronously replaced with sensitive words frame by frame through a preset sliding window; the live stream after audio replacement is output.

[0081] The above-mentioned live stream processing electronic device, after the live stream flows in, divides the audio stream in the live stream into audio frames to form multiple audio frames, and then determines the sensitive word position and performs sliding audio replacement on each audio frame in an asynchronous manner, so that the sensitive word determination program and the replacement program will not block each other, and performs sliding window audio replacement with the frame level as the audio processing unit, which can ensure delay-free silencing of the replaced audio and audio frames, thereby improving the real-time and stability of the live stream output.

[0082] The above-mentioned determination of the sensitive word position information in the audio stream, and based on the target replacement audio, asynchronously replaces the sensitive word audio frame by frame through a preset sliding window for the audio frame, including: sending the audio stream to the review server through a first process; receiving the sensitive word position information sent by the review server in real time through a second process, and based on the target replacement audio, asynchronously replaces the sensitive word audio frame by frame through a preset sliding window for the audio frame.

[0083] The above-mentioned sending of the audio stream to the audit server through the first process includes: decapsulating and decoding the audio stream through the first process to obtain the first pulse code modulation data of the audio stream; encrypting the first pulse code modulation data, and sending the encrypted first pulse code modulation data to the audit server through the first process, so that the audit server identifies sensitive words in the encrypted first pulse code modulation data, and sends the identified sensitive word location information to the second process.

[0084] The above-mentioned target-based audio replacement asynchronously replaces audio frames with sensitive words frame by frame through a preset sliding window, including: asynchronously inputting audio frames into a preset audio queue; when the number of audio frames in the preset audio queue reaches the maximum audio frame capacity, based on the sensitive word position information, slidingly taking values of the target replacement audio through a preset sliding window, and replacing the values with the sensitive word audio frames in the preset audio queue during the sliding process; the above-mentioned sensitive word audio frames are used to: indicate the audio frames where the sensitive words indicated by the sensitive word position information are located.

[0085] The above-mentioned method is based on the sensitive word position information, and the target replacement audio is slidingly valued through a preset sliding window, and in the process of sliding value, the value is replaced with the sensitive word audio frame in the preset audio queue, including: according to the sensitive word position information, determining the audio frame where the sensitive word is located in the preset audio queue, and obtaining the sensitive word audio frame; according to the data length of the sensitive word audio frame, determining the maximum sliding distance of the preset sliding window, and sliding the target replacement audio frame by frame according to the maximum sliding distance; the size of the sliding window is the same as the length of an audio frame; in the frame-by-frame sliding process of the sliding window, the value is replaced with the sensitive word audio frame in the preset audio queue.

[0086] Before the above-mentioned frame processing of the audio stream in the live stream in response to the inflow of the live stream generated by the anchor end to obtain multiple audio frames, the method also includes: generating a target replacement audio with a preset maximum audio replacement length according to preset waveform parameters.

[0087] In the above-mentioned response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and before the above-mentioned determination of the sensitive word position information in the audio stream, and based on the target replacement audio, asynchronously replacing the audio frames with sensitive words frame by frame through a preset sliding window, the method also includes: extracting audio parameters from the live stream to obtain target audio parameters; decoding and resampling the target replacement audio to obtain second pulse code modulation data of the target replacement audio; setting the audio parameters of the second pulse code modulation data to the target audio parameters to obtain the target replacement audio with the same audio parameters.

[0088] In the above-mentioned response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and before the sensitive word position information in the audio stream is determined and the audio frame is asynchronously replaced with sensitive words frame by frame through a preset sliding window based on the target replacement audio, the method also includes: volume prediction through the amplitude data in the first pulse code modulation data of the audio frame to obtain target volume data; setting the target volume data for the second pulse code modulation data to obtain target replacement audio with the same volume.

[0089] The above-mentioned frame processing of the audio stream in the live stream to obtain multiple audio frames includes: decapsulating and decoding the audio stream in the live stream to obtain first pulse code modulation data of the audio stream; and dividing the first pulse code modulation data into audio frames in real time according to a preset sampling rate to obtain multiple audio frames.

[0090] The above-mentioned output of the live stream after audio replacement includes: encoding the audio stream after audio replacement to obtain an encoded audio stream; encapsulating the encoded audio stream and the video stream in the live stream into a live stream, and outputting the encapsulated live stream.

[0091] This embodiment further provides a machine-readable storage medium storing machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the above-mentioned live stream processing method, for example:

[0092] In response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is framed to obtain multiple audio frames; the sensitive word position information in the audio stream is determined, and based on the target audio replacement, the audio frames are asynchronously replaced with sensitive words frame by frame through a preset sliding window; the live stream after audio replacement is output.

[0093] The above-mentioned live stream processing and storage medium divides the audio stream in the live stream into audio frames after the live stream flows in, forming multiple audio frames, and then determines the sensitive word position and performs sliding audio replacement on each audio frame in an asynchronous manner, so that the sensitive word determination program and the replacement program will not block each other, and performs sliding window audio replacement with the frame level as the audio processing unit, which can ensure delay-free silencing of the replaced audio and audio frames, thereby improving the real-time and stability of the live stream output.

[0094] The above-mentioned determination of the sensitive word position information in the audio stream, and based on the target replacement audio, asynchronously replaces the sensitive word audio frame by frame through a preset sliding window for the audio frame, including: sending the audio stream to the review server through a first process; receiving the sensitive word position information sent by the review server in real time through a second process, and based on the target replacement audio, asynchronously replaces the sensitive word audio frame by frame through a preset sliding window for the audio frame.

[0095] The above-mentioned sending of the audio stream to the audit server through the first process includes: decapsulating and decoding the audio stream through the first process to obtain the first pulse code modulation data of the audio stream; encrypting the first pulse code modulation data, and sending the encrypted first pulse code modulation data to the audit server through the first process, so that the audit server identifies sensitive words in the encrypted first pulse code modulation data, and sends the identified sensitive word location information to the second process.

[0096] The above-mentioned target-based audio replacement asynchronously replaces audio frames with sensitive words frame by frame through a preset sliding window, including: asynchronously inputting audio frames into a preset audio queue; when the number of audio frames in the preset audio queue reaches the maximum audio frame capacity, based on the sensitive word position information, slidingly taking values of the target replacement audio through a preset sliding window, and replacing the values with the sensitive word audio frames in the preset audio queue during the sliding process; the above-mentioned sensitive word audio frames are used to: indicate the audio frames where the sensitive words indicated by the sensitive word position information are located.

[0097] The above-mentioned method is based on the sensitive word position information, and the target replacement audio is slidingly valued through a preset sliding window, and in the process of sliding value, the value is replaced with the sensitive word audio frame in the preset audio queue, including: according to the sensitive word position information, determining the audio frame where the sensitive word is located in the preset audio queue, and obtaining the sensitive word audio frame; according to the data length of the sensitive word audio frame, determining the maximum sliding distance of the preset sliding window, and sliding the target replacement audio frame by frame according to the maximum sliding distance; the size of the sliding window is the same as the length of an audio frame; in the frame-by-frame sliding process of the sliding window, the value is replaced with the sensitive word audio frame in the preset audio queue.

[0098] Before the above-mentioned frame processing of the audio stream in the live stream in response to the inflow of the live stream generated by the anchor end to obtain multiple audio frames, the method also includes: generating a target replacement audio with a preset maximum audio replacement length according to preset waveform parameters.

[0099] In the above-mentioned response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and before the above-mentioned determination of the sensitive word position information in the audio stream, and based on the target replacement audio, asynchronously replacing the audio frames with sensitive words frame by frame through a preset sliding window, the method also includes: extracting audio parameters from the live stream to obtain target audio parameters; decoding and resampling the target replacement audio to obtain second pulse code modulation data of the target replacement audio; setting the audio parameters of the second pulse code modulation data to the target audio parameters to obtain the target replacement audio with the same audio parameters.

[0100] In the above-mentioned response to the inflow of the live stream generated by the anchor end, the audio stream in the live stream is frame-processed to obtain multiple audio frames, and before the sensitive word position information in the audio stream is determined and the audio frame is asynchronously replaced with sensitive words frame by frame through a preset sliding window based on the target replacement audio, the method also includes: volume prediction through the amplitude data in the first pulse code modulation data of the audio frame to obtain target volume data; setting the target volume data for the second pulse code modulation data to obtain target replacement audio with the same volume.

[0101] The above-mentioned frame processing of the audio stream in the live stream to obtain multiple audio frames includes: decapsulating and decoding the audio stream in the live stream to obtain first pulse code modulation data of the audio stream; and dividing the first pulse code modulation data into audio frames in real time according to a preset sampling rate to obtain multiple audio frames.

[0102] The above-mentioned output of the live stream after audio replacement includes: encoding the audio stream after audio replacement to obtain an encoded audio stream; encapsulating the encoded audio stream and the video stream in the live stream into a live stream, and outputting the encapsulated live stream.

[0103] The computer program product of the live stream processing method, device, electronic device and storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. For specific implementation, please refer to the method embodiments and will not be repeated here.

[0104] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0105] In addition, in the description of the embodiments of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.

[0106] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0107] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0108] Finally, it should be noted that the above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A method for processing a live stream, characterized in that: The method comprises: In response to an inflow of a live stream generated by a host terminal, performing frame processing on an audio stream in the live stream to obtain a plurality of audio frames; Determine the sensitive word position information in the audio stream, and based on the target replacement audio, asynchronously replace the sensitive word frame by frame of the audio frame through a preset sliding window; Output the live stream after audio replacement; Determine the sensitive word position information in the audio stream, and based on the target replacement audio, asynchronously replace the sensitive word frame by frame on the audio frame through a preset sliding window, including: Sending the audio stream to a review server through a first process; Determine whether the input audio frame is a sensitive word audio frame based on the audio frame review result received by the second process from the first process; If the input audio frame is a sensitive word audio frame, the sensitive word audio frame is replaced frame by frame with the target replacement audio through a preset sliding window.

2. The method according to claim 1, characterized in that Sending the audio stream to the review server through a first process includes: Decapsulating and decoding the audio stream through a first process to obtain first pulse code modulation data of the audio stream; The first pulse code modulation data is encrypted and sent to the audit server through the first process, so that the audit server identifies sensitive words in the encrypted first pulse code modulation data and sends the identified sensitive word location information to the second process.

3. The method according to claim 1, characterized in that Based on the target replacement audio, asynchronously performing sensitive word frame-by-frame audio replacement on the audio frame through a preset sliding window, including: Asynchronously inputting the audio frame into a preset audio queue; When the number of audio frames in the preset audio queue reaches the maximum audio frame capacity, based on the sensitive word position information, the target replacement audio is slid and valued through a preset sliding window, and during the sliding value-taking process, the value is replaced with the sensitive word audio frame in the preset audio queue; the above-mentioned sensitive word audio frame is used to: indicate the audio frame where the sensitive word indicated by the sensitive word position information is located.

4. The method according to claim 3, characterized in that Based on the sensitive word position information, slidingly obtaining a value of the target replacement audio through a preset sliding window, and replacing the value with the sensitive word audio frame in the preset audio queue during the sliding obtaining process, including: Determine, based on the sensitive word position information, the audio frame where the sensitive word is located in the preset audio queue, and obtain the sensitive word audio frame; Determining a maximum sliding distance of a preset sliding window based on the data length of the sensitive word audio frame, and sliding the target replacement audio frame by frame according to the maximum sliding distance; the size of the sliding window is the same as the length of an audio frame; During the frame-by-frame sliding process of the sliding window, the value is replaced into the sensitive word audio frame in the preset audio queue.

5. The method according to claim 1, wherein Before, in response to the inflow of the live stream generated by the host terminal, performing frame processing on the audio stream in the live stream to obtain multiple audio frames, the method further includes: Generate target replacement audio with a preset maximum audio replacement length based on preset waveform parameters.

6. The method according to claim 1, characterized in that In response to an inflow of a live stream generated by a host terminal, after performing frame processing on an audio stream in the live stream to obtain a plurality of audio frames, determining sensitive word position information in the audio stream, and before asynchronously performing sensitive word frame-by-frame audio replacement on the audio frames through a preset sliding window based on target audio replacement, the method further includes: Extracting audio parameters from the live stream to obtain target audio parameters; Decoding and resampling the target replacement audio to obtain second pulse code modulation data of the target replacement audio; The audio parameters of the second pulse code modulation data are set as the target audio parameters to obtain the target replacement audio with the same audio parameters.

7. The method according to claim 6, characterized in that In response to an inflow of a live stream generated by a host terminal, after performing frame processing on an audio stream in the live stream to obtain a plurality of audio frames, determining sensitive word position information in the audio stream, and before asynchronously performing sensitive word frame-by-frame audio replacement on the audio frames through a preset sliding window based on target audio replacement, the method further includes: Performing volume prediction based on amplitude data in the first pulse code modulation data of the audio frame to obtain target volume data; The target volume data is set for the second pulse code modulation data to obtain a target replacement audio with the same volume.

8. The method according to claim 1, characterized in that The audio stream in the live stream is framed to obtain multiple audio frames, including: Decapsulating and decoding the audio stream in the live stream to obtain first pulse code modulation data of the audio stream; The first pulse code modulation data is divided into audio frames in real time according to a preset sampling rate to obtain multiple audio frames.

9. The method according to claim 1, characterized in that Output live stream after audio replacement, including: Encoding the audio stream after the audio is replaced to obtain an encoded audio stream; The encoded audio stream and the video stream in the live stream are encapsulated into a live stream, and the encapsulated live stream is output.

10. A live stream processing device, characterized in that: The device comprises: a response module, configured to respond to an inflow of a live stream generated by a host terminal and perform frame processing on an audio stream in the live stream to obtain a plurality of audio frames; A replacement module is configured to determine the position information of sensitive words in the audio stream and, based on the target replacement audio, asynchronously replace the sensitive words frame by frame on the audio frames through a preset sliding window; Output module, used to output the live stream after audio replacement; The replacement module is also used to: send the audio stream to the review server through the first process; determine whether the input audio frame is a sensitive word audio frame through the audio frame review result from the first process received by the second process; if the input audio frame is a sensitive word audio frame, perform frame-by-frame audio replacement of the target replacement audio on the sensitive word audio frame through a preset sliding window.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the live stream processing method according to any one of claims 1 to 9.

12. A machine-readable storage medium, characterized in that The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions prompt the processor to implement the live stream processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Sensitive word filtering method and device for speech recognition input, control panel and equipment

    CN113851132A

  • Audio prohibited word filtering method and device, electronic equipment and storage medium

    CN114694656A