Live broadcast stream processing method and device, equipment and storage medium

By extracting frames and cutting the live stream, combining image, audio and text recognition models, the integrated tree model is used to detect whether the live stream carries the target type content, which solves the problem of low detection accuracy in the prior art and achieves higher detection accuracy and efficiency.

CN120302086APending Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410034771.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the accuracy of detecting whether the target type content is carried in the live stream based on the current image frame and the audio slice of the live stream is low.

Method used

By extracting and cutting the live stream, image frames and audio slicing sequences are obtained, image recognition, audio recognition and text recognition models are used, and detection is combined with an integrated tree model to improve detection accuracy.

Benefits of technology

By introducing rich information above, the detection accuracy of target type content in the live stream is improved, the detection complexity is reduced and efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302086A_ABST
    Figure CN120302086A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a live broadcast stream processing method and device, equipment and a storage medium, and can relate to the technical field of audio and video, and the method comprises the steps: carrying out the frame extraction of an image stream of a current live broadcast stream, and obtaining a tth image frame to a (t-n) th image frame; cutting the audio stream of the current live broadcast stream to obtain a tth audio slice to a (t-n) th audio slice; wherein n is a positive integer; and based on the t-th image frame to the (t-n)-th image frame and the t-th audio slice to the (t-n)-th audio slice, detecting whether the current live broadcast stream carries the content of the target type. Therefore, the detection precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of audio and video technologies, and in particular, to a live stream processing method, apparatus, device, and storage medium. Background Art

[0002] With the rapid development of the live broadcast industry, more and more hosts have flocked to this industry. However, there are still some hosts who attract users with unhealthy and borderline live broadcast styles. Therefore, it is crucial to detect whether the live stream carries content of a target type, such as unhealthy and borderline content.

[0003] Currently, it is mainly based on the current image frame and current audio slice of the live stream to detect whether the live stream carries content of a target type. However, this method has the problem of low detection accuracy. Summary of the Invention

[0004] The embodiments of the present application provide a live stream processing method, apparatus, device, and storage medium, thereby improving the detection accuracy.

[0005] In a first aspect, the embodiments of the present application provide a live stream processing method, including: extracting frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame; cutting the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice; where n is a positive integer; detecting whether the current live stream carries content of a target type based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice.

[0006] In a second aspect, the embodiments of the present application provide a live stream processing apparatus, including: a frame extraction module, a cutting module, and a detection module. The frame extraction module is configured to extract frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame; the cutting module is configured to cut the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice; where n is a positive integer; the detection module is configured to detect whether the current live stream carries content of a target type based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice.

[0007] In a third aspect, an electronic device is provided, including: a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the method as described in the first aspect or its various implementation manners.

[0008] In a fourth aspect, a computer-readable storage medium is provided for storing a computer program, and the computer program causes a computer to execute the method as described in the first aspect or its various implementation manners.

[0009] Fifth aspect, a computer program product is provided, including computer program instructions that cause a computer to execute the method in the first aspect or its various implementation manners.

[0010] Sixth aspect, a computer program is provided, and the computer program causes a computer to execute the method in the first aspect or its various implementation manners.

[0011] Through the technical solution provided by this application, by introducing richer context information, such as the t-1th to t-nth image frames and the t-1th to t-nth audio slices, and the image frames and audio slices in the live broadcast scenario have strong continuity, the current image frame is relatively similar to the previous image frames, and the current audio slice is relatively similar to the previous audio slices. Therefore, these context information can play a role in information supplementation, and thus can improve the content detection accuracy. Description of the Drawings

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings required for the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0013] Figure 1 It is a schematic diagram of a system architecture related to the embodiments of this application;

[0014] Figure 2 It is a flowchart of a live stream processing method provided by the embodiments of this application;

[0015] Figure 3 It is a schematic diagram of a live stream processing process provided by the embodiments of this application;

[0016] Figure 4 It is a schematic diagram of a live stream processing process provided by the related art;

[0017] Figure 5 It is a schematic diagram of a model training method provided by the embodiments of this application;

[0018] Figure 6 It is a statistical chart of the effects provided by the embodiments of this application;

[0019] Figure 7 It is a schematic diagram of a live stream processing apparatus 700 provided by the embodiments of this application;

[0020] Figure 8 It is a schematic block diagram of an electronic device provided by the embodiments of this application. Detailed Embodiments

[0021] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0023] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of the module or unit.

[0024] The embodiments of the present application may relate to audio-visual technology and artificial intelligence (AI) technology, but are not limited thereto.

[0025] Among them, audio-visual technology is a technology for processing, transmitting, storing, and playing audio-visual signals. It involves processes such as the acquisition, encoding, transmission, decoding, and rendering of audio and video signals. In the embodiments of the present application, video frames can be extracted (i.e., acquired) from a video stream to obtain the t-th image frame to the t-n-th image frame. And audio slices can be cut (i.e., acquired) from an audio stream to obtain the t-th audio slice to the t-n-th audio slice.

[0026] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0027] Artificial intelligence technology is an interdisciplinary subject that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the base model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The embodiments of this application may be related to computer vision technology, speech processing technology, and natural language processing technology.

[0028] Computer Vision (CV) is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, detection, and measurement, and further performing graphic processing to make the computer process the images into a form more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. The embodiments of this application can mainly process the t-th image frame to the t-n-th image frame through computer vision technology to obtain the respective scores of the t-th image frame to the t-n-th image frame.

[0029] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and sense is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods. The large model technology has brought about a revolution in the development of speech technology. Pretrained models such as WavLM and UniSpeech that follow the Transformer architecture have strong generalization and versatility and can excellently complete speech processing tasks in various directions. In the embodiments of this application, speech technology can mainly be used to process the t-th audio slice to the (t - n)-th audio slice, obtain the respective scores of the t-th audio slice to the (t - n)-th audio slice, and can separately convert the t-th audio slice to the (t - n)-th audio slice into the t-th text to the (t - n)-th text.

[0030] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistics research; at the same time, it involves computer science and mathematics, which are important technologies for model training in the field of artificial intelligence. The pretrained model is developed from the large language model in the NLP field. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs. In the embodiments of this application, natural language processing technology can mainly be used to process the t-th text to the (t - n)-th text to obtain the respective scores of the t-th text to the (t - n)-th text.

[0031] The relevant knowledge involved in this application will be elaborated below:

[0032] First, a live stream is a technology that transmits audio and video content to the viewer's end in real time. In a live broadcast, the host collects audio and video signals through devices such as cameras and microphones, and then transmits them to the viewers in real time through the live stream technology. Viewers can watch the live broadcast through various devices (such as mobile phones, computers, tablets, etc.) and interact with the host in real time.

[0033] Second, an image stream is a special type of stream that contains still images assigned to the presentation time.

[0034] III. Audio stream is a data stream used for real-time transmission of audio data. It is transmitted through networks such as the Internet or local area network and is commonly used in scenarios such as video conferencing, music playback, and live broadcasting. Audio stream is real-time data based on network transmission and needs to be decoded and played in real-time.

[0035] IV. Frame extraction refers to extracting individual image frames from an image stream.

[0036] V. Audio slice is obtained by using audio editing software to cut and process an audio file. When performing audio slicing, the following points generally need to be noted:

[0037] Determine the slice position: Select the position to be cut in the audio editing software and determine the cutting time point.

[0038] Execute the slicing operation: According to the software's prompt or button, select to execute the cutting operation to separate the selected part from the original audio file.

[0039] Save the slice: Save the cut audio slice to the specified location for subsequent use.

[0040] VI. Flink, whose aggregation operation is used to summarize, calculate, and statistically analyze a data stream. The aggregation operation can perform various complex calculations on the data stream, such as summation, average calculation, maximum and minimum value calculation, etc.

[0041] VII. Ensemble tree models include: Random Forest, Gradient Boosting Tree, Extreme Gradient Boosting (XGBoost) ensemble tree, etc. Among them, Random Forest is an algorithm that constructs multiple trees and integrates the results. Gradient Boosting Tree is an algorithm that improves prediction accuracy by iteratively constructing better models. XGBoost is a machine learning algorithm based on gradient boosting decision trees, which improves the efficiency and accuracy of gradient boosting decision trees by introducing some special techniques.

[0042] VIII. XGBoost is an efficient gradient boosting decision tree algorithm, featuring high speed, high accuracy, and strong interpretability. It improves the efficiency and accuracy of decision trees by introducing some new techniques, such as second-order derivative pruning, column sampling, feature subsets, etc. At the same time, XGBoost also supports various tasks, such as classification, regression, sorting, etc., and can be widely applied to various machine learning tasks.

[0043] The technical problems to be solved, inventive concepts, and system architectures of the embodiments of the present application will be elaborated below:

[0044] As described above, currently, whether the live stream carries content of the target type is mainly detected based on the current image frame and the current audio slice of the live stream. However, this method has the problem of low detection accuracy.

[0045] To solve the above technical problems, the embodiments of the present application propose to detect whether the current live stream carries content of the target type based on the image frames and audio slices within a sliding time window, that is, by introducing richer context information to improve the content detection accuracy.

[0046] In some implementable ways, the system architecture of the embodiments of the present application is as Figure 1 shown.

[0047] Figure 1 It is a schematic diagram of a system architecture related to the embodiments of the present application, including a live broadcast end 110, an audience end 120, and a background server 130.

[0048] Among them, the live broadcast end 110 may be installed with a video client or a small program, etc., to conduct video live broadcasts through the video client. The background server 130 may be installed with a video client or a small program, etc., to watch the live broadcast through the video client.

[0049] In some implementable ways, the live broadcast end 110 may be a smart phone, a tablet computer, a smart watch, a virtual reality (VR) device, an augmented reality (AR) device, etc., but is not limited thereto.

[0050] In some implementable ways, the audience end 120 may be a smart phone, a tablet computer, a smart watch, a VR device, an AR device, etc., but is not limited thereto.

[0051] The live broadcast end 110 and the background server 130 may be directly or indirectly connected through a wired or wireless communication method, and the present application does not make any restrictions here.

[0052] The audience end 120 and the background server 130 may be directly or indirectly connected through a wired or wireless communication method, and the present application does not make any restrictions here.

[0053] In some implementable ways, the background server 130 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0054] In the embodiments of the present application, the background server 130 can be used to implement the live stream processing method provided by the embodiments of the present application, but is not limited thereto.

[0055] It should be noted that Figure 1 is only a schematic diagram of a system architecture provided by the embodiments of the present application. The system architecture involved in the embodiments of the present application is not limited to Figure 1 the system architecture shown. For example, in some implementable ways, based on Figure 1 the system architecture shown, the system architecture involved in the embodiments of the present application may further include: other electronic devices, which can be connected to the background server 130 to pull the live stream from the background server 130 and implement the live stream processing method provided by the embodiments of the present application.

[0056] The embodiments of the present application will be elaborated in detail below:

[0057] Figure 2 is a flowchart of a live stream processing method provided by the embodiments of the present application. This method can be applied to an electronic device, which can be Figure 1 the background server in Figure 2 but is not limited thereto. As

[0058] S210: Extract frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame;

[0059] where n is a positive integer, and n can be understood as the width of the time window corresponding to the t-th image frame to the (t - n)-th image frame.

[0060] The following is an example illustration of the frame extraction situation:

[0061] For example, n = 1, that is, the electronic device can extract frames from the image stream of the current live stream to obtain the t-th image frame and the (t - 1)-th image frame. Another example, n = 2, that is, the electronic device can extract frames from the image stream of the current live stream to obtain the t-th image frame, the (t - 1)-th image frame, and the (t - 2)-th image frame. Another example, n = 3, that is, the electronic device can extract frames from the image stream of the current live stream to obtain the t-th image frame, the (t - 1)-th image frame, the (t - 2)-th image frame, and the (t - 3)-th image frame.

[0062] In some implementable ways, the t-th image frame can be the image frame corresponding to the current moment, and the current moment can be the t-th moment, but is not limited thereto.

[0063] It should be understood that the electronic device can extract frames from the image stream of the current live stream at a preset sampling rate. For example, the preset sampling rate is 1 frame per second. Based on this, the t-th image frame to the (t - n)-th image frame can be understood as: the image frame at the t-th second to the image frame at the (t - n)-th second.

[0064] S220: Cut the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice;

[0065] Among them, n can be understood as the width of the time window corresponding to the t-th audio slice to the (t - n)-th audio slice.

[0066] The following is an example to illustrate the cutting situation:

[0067] For example, n = 1, that is, the electronic device can cut the audio stream of the current live stream to obtain the t-th audio slice and the (t - 1)-th audio slice. For another example, n = 2, that is, the electronic device can cut the audio stream of the current live stream to obtain the t-th audio slice, the (t - 1)-th audio slice, and the (t - 2)-th audio slice. For another example, n = 3, that is, the electronic device can cut the audio stream of the current live stream to obtain the t-th audio slice, the (t - 1)-th audio slice, the (t - 2)-th audio slice, and the (t - 3)-th audio slice.

[0068] In some implementable ways, the t-th audio slice can be the audio slice corresponding to the current moment, and the current moment can be the t-th moment, but not limited to this.

[0069] It should be understood that the electronic device can cut the audio stream of the current live stream at a preset cutting rate. For example, the preset cutting rate is 1 time per second. Based on this, the t-th audio slice to the (t - n)-th audio slice can be understood as: the audio slice at the t-th second to the audio slice at the (t - n)-th second.

[0070] It should be understood that when the electronic device cuts the audio stream of the current live stream, it needs to determine the cutting start time and end time of each audio slice, and perform the cutting of the audio slice based on the cutting start time and end time of each audio slice. And the moment in the audio slice at a certain moment can be the cutting end time of this audio slice, and the cutting start time of this audio slice is the cutting end time of the previous audio slice in this audio slice.

[0071] For example, for the audio slice at the t-th second, its cutting start time is the (t - 1)-th second, and its cutting end time is the t-th second.

[0072] In some implementable ways, the t-th audio slice corresponds to the t-th image frame, that is, they are respectively the audio slice and the image frame of the current live stream at the same moment. For example, the t-th audio slice and the t-th image frame are respectively the audio slice and the image frame at the t-th second. Similarly, the (t - i)-th audio slice corresponds to the (t - i)-th image frame, where i = 1, 2... n.

[0073] S230: Based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice, detect whether the current live stream carries content of the target type.

[0074] In the embodiments of the present application, the electronic device can implement S230 in any of the following implementable ways, but is not limited thereto:

[0075] Implementable way one, S230 may include:

[0076] S230-1A: Based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice, obtain the score of the current live stream;

[0077] In some implementable ways, the score of the current live stream is used to measure whether the current live stream carries content of the target type. Among them, if the score of the current live stream is higher, it means that the possibility that the current live stream carries content of the target type is greater; if the score of the current live stream is lower, it means that the possibility that the current live stream carries content of the target type is lower.

[0078] Among them, S230-1A can be implemented in any of the following implementable ways, but is not limited thereto:

[0079] In some implementable ways, S230-1A may include:

[0080] S230-1A-1a: Input the t-th image frame to the (t - n)-th image frame into the image recognition model to obtain the scores of the t-th image frame to the (t - n)-th image frame respectively;

[0081] In some implementable ways, the number of image recognition models can be one or more, and the embodiments of the present application do not limit this.

[0082] In some implementable ways, the image recognition model can be a machine learning model, for example, it can be a Residual Network (ResNet).

[0083] It should be understood that ResNet is a deep neural network architecture that is widely used in computer vision tasks, especially image classification. The main idea of ResNet is to introduce a residual structure, and by introducing one or more residual blocks, the problem of vanishing gradients caused by increasing depth is avoided. The residual structure in the residual block takes the weighted sum of the input and the output of the previous layer as the input of the current layer, thereby achieving a "shortcut" connection to the output of the previous layer, enabling the gradient to be directly passed from the previous layer to the next layer. In this way, ResNet has successfully increased the depth and performance of the network and demonstrated superior performance in various image classification tasks.

[0084] In some implementable ways, the scores of the t-th image frame to the (t - n)-th image frame are used to measure whether the t-th image frame to the (t - n)-th image frame each carry content of the target type. Among them, for any one of the t-th image frame to the (t - n)-th image frame, if the score of this image frame is higher, it means that the possibility that this image frame carries content of the target type is greater; if the score of this image frame is lower, it means that the possibility that this image frame carries content of the target type is lower.

[0085] In some implementable ways, the image recognition model can be understood as an image classification model, or even an image binary classification model. For example, the classification results of this classification can be: the category of the image frame carrying content of the target type, and the category of the image frame not carrying content of the target type.

[0086] For example, for any one of the t-th image frame to the (t - n)-th image frame, if the score of this image frame is greater than a preset score, it means that this image frame carries content of the target type; if the score of this image frame is less than or equal to the preset score, it means that this image frame does not carry content of the target type.

[0087] S230-1A-2a: Input the t-th audio slice to the (t - n)-th audio slice into the audio recognition model to obtain the scores of the t-th audio slice to the (t - n)-th audio slice respectively;

[0088] In some implementable ways, the number of audio recognition models can be one or more, and the embodiments of the present application do not limit this.

[0089] In some implementable ways, the audio recognition model can be a machine learning model. For example, it can be a Transformer-based Convolutional Neural Network (Conformer).

[0090] It should be understood that the Conformer model is a neural network model for natural language processing and speech recognition tasks. Its main advantage is that it reduces the number of model parameters without sacrificing accuracy and improves the model running speed. The Conformer model first performs downsampling through a convolutional network and then connects a series of conformer modules. Among them, the conformer module includes the following parts: a feedforward module, a multi-head self-attention module, and a convolution module.

[0091] In some implementable ways, the scores of the t-th audio slice to the t-n-th audio slice are used to measure whether the t-th audio slice to the t-n-th audio slice carry content of the target type. Among them, for any one of the t-th audio slice to the t-n-th audio slice, if the score of this audio slice is higher, it means that the possibility of this audio slice carrying content of the target type is greater; if the score of this audio slice is lower, it means that the possibility of this audio slice carrying content of the target type is lower.

[0092] In some implementable ways, the audio recognition model can be understood as an audio classification model, or even an audio binary classification model. For example, the classification results can be: the category of audio slices carrying content of the target type, and the category of audio slices not carrying content of the target type.

[0093] For example, for any one of the t-th audio slice to the t-n-th audio slice, if the score of this audio slice is greater than a preset score, it means that this audio slice carries content of the target type; if the score of this audio slice is less than or equal to the preset score, it means that this audio slice does not carry content of the target type.

[0094] It should be understood that the preset score corresponding to the score of the audio slice and the preset score corresponding to the score of the image frame can be the same or different, and the embodiments of the present application do not limit this.

[0095] S230-1A-3a: Input the t-th audio slice to the t-n-th audio slice into the speech recognition model to obtain the t-th text to the t-n-th text;

[0096] It should be understood that the speech recognition model is used to convert audio slices into text. Based on this, the speech recognition model is also called a speech-to-text model.

[0097] In some implementable manners, the speech recognition model may be a Hidden Markov Model (HMM), a Gaussian Mixture Model (GMM), a deep learning model, an end-to-end model, etc.

[0098] Among them, HMM is a statistical model used to describe the transition process of the system state. In speech recognition, HMM is used to model the temporal characteristics of the audio signal, divide the audio signal into different state sequences, and thus recognize the corresponding speech content.

[0099] GMM is a probability density model used to describe the statistical characteristics of data distribution. In speech recognition, GMM can be used to model the feature distribution of the audio signal, obtain the model parameters of different speech units through training, and thus realize speech recognition.

[0100] In some implementable manners, the deep learning model may be any of the following, but not limited to: Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN), etc. These models can extract high-level feature representations by learning a large amount of speech data, thereby improving the accuracy of speech recognition.

[0101] The end-to-end model is a model that directly maps the audio input to the text output, avoiding the traditional step-by-step processing method. Among them, the end-to-end model may be any of the following, but not limited to: a model based on the encoder-decoder architecture (such as Seq2Seq), Connectionist Temporal Classification (CTC), and Attention Mechanism, etc.

[0102] S230-1A-4a: Input the t-th text to the (t - n)-th text into the text recognition model to obtain the scores of the t-th text to the (t - n)-th text respectively;

[0103] In some implementable manners, the number of text recognition models may be one or more, and the embodiments of this application do not limit this.

[0104] In some implementable ways, the text recognition model can be a machine learning model, for example, it can be a Bidirectional Encoder Representations from Transformers (BERT) model.

[0105] It should be understood that the BERT model is a bidirectional encoder representation based on Transformer and is a pre-trained language representation model. It adopts the Masked Language Model (MLM) so as to generate deep bidirectional language representations. The goal of the BERT model is to train with a large-scale unlabeled corpus to obtain a representation of the text containing rich semantic information, that is, the semantic representation of the text, and then fine-tune the semantic representation of the text in a specific NLP task and finally apply it to this NLP task. The BERT model is usually used to implement the following tasks: text classification, question answering system, text generation, text summarization, text matching.

[0106] In the embodiments of the present application, the BERT model is mainly used for text classification.

[0107] In some implementable ways, the scores of the t-th text to the (t - n)-th text are used to measure whether the t-th text to the (t - n)-th text carry content of the target type. Among them, for any one of the t-th text to the (t - n)-th text, if the score of this text is higher, it means that the possibility that this text carries content of the target type is greater, and if the score of this text is lower, it means that the possibility that this text carries content of the target type is lower.

[0108] In some implementable ways, the text recognition model can be understood as a text classification model, or even a text binary classification model. For example, the classification results of this classification can be: the text category carrying content of the target type, the text category not carrying content of the target type.

[0109] For example, for any one of the t-th text to the (t - n)-th text, if the score of this text is greater than the preset score, it means that this text carries content of the target type, and if the score of this text is less than or equal to the preset score, it means that this text does not carry content of the target type.

[0110] It should be understood that the preset score corresponding to the score of the text, the preset score corresponding to the score of the audio slice, and the preset score corresponding to the score of the image frame can be the same or different, and the embodiments of the present application do not limit this.

[0111] S230-1A-5a: Obtain the score of the current live stream based on the scores of the t-th to the (t-n)-th image frames, the scores of the t-th to the (t-n)-th audio slices, and the scores of the t-th to the (t-n)-th texts respectively.

[0112] Among them, the electronic device can implement S230-1A-5a through any of the following implementable ways, but not limited to this:

[0113] In some implementable ways, S230-1A-5a may include:

[0114] S230-1A-5a-1a: Aggregate the scores of the t-th to the (t-n)-th image frames respectively to obtain a first aggregation result;

[0115] In some implementable ways, the electronic device can use the maximum value, average value or summation result of the scores of the t-th to the (t-n)-th image frames respectively as the first aggregation result.

[0116] For example, assume that the scores of the t-th to the (t-n)-th image frames are respectively denoted as: A_t, A_t-1... A_t-n, and they form a set A. Then the first aggregation result = MAX{A} or AVG{A} or SUM{A}, where MAX{A} represents taking the maximum value in set A, AVG{A} represents taking the average value of all elements in set A, and SUM{A} represents taking the summation result of all elements in set A.

[0117] In some implementable ways, the electronic device can use the aggregation operation of Flink to aggregate the scores of the t-th to the (t-n)-th image frames respectively, but not limited to this.

[0118] S230-1A-5a-2a: Aggregate the scores of the t-th to the (t-n)-th audio slices respectively to obtain a second aggregation result;

[0119] In some implementable ways, the electronic device can use the maximum value, average value or summation result of the scores of the t-th to the (t-n)-th audio slices respectively as the second aggregation result.

[0120] For example, assume that the scores of the t-th to the (t-n)-th audio slices are respectively denoted as: B_t, B_t-1... B_t-n, and they form a set B. Then the second aggregation result = MAX{B} or AVG{B} or SUM{B}, where MAX{B} represents taking the maximum value in set B, AVG{B} represents taking the average value of all elements in set B, and SUM{B} represents taking the summation result of all elements in set B.

[0121] In some implementable ways, the electronic device can perform an aggregation operation of Flink on the scores of each of the t-th audio slice to the t-n-th audio slice, but not limited to this.

[0122] S230-1A-5a-3a: Aggregate the scores of each of the t-th text to the t-n-th text to obtain a third aggregation result;

[0123] In some implementable ways, the electronic device can use the maximum value, average value or summation result of the scores of each of the t-th text to the t-n-th text as the third aggregation result.

[0124] For example, assume that the scores of each of the t-th text to the t-n-th text are respectively recorded as: C_t, C_t-1... C_t-n, and they form a set C. Then the third aggregation result = MAX{C} or AVG{C} or SUM{C}, where MAX{C} represents taking the maximum value in the set C, AVG{C} represents taking the average value of all elements in the set C, and SUM{C} represents taking the summation result of all elements in the set C.

[0125] In some implementable ways, the electronic device can perform an aggregation operation of Flink on the scores of each of the t-th text to the t-n-th text, but not limited to this.

[0126] S230-1A-5a-4a: Obtain the score of the current live stream based on the first aggregation result, the second aggregation result and the third aggregation result.

[0127] In some implementable ways, the electronic device can input the first aggregation result, the second aggregation result and the third aggregation result into an ensemble tree model to obtain the score of the current live stream.

[0128] In some other implementable ways, the electronic device first performs a normalization process on the first aggregation result, the second aggregation result and the third aggregation result to obtain a first normalization result corresponding to the first aggregation result, a second normalization result corresponding to the second aggregation result, and a first normalization result corresponding to the third aggregation result. Further, then input the first normalization result, the second normalization result and the third normalization result into the ensemble tree model to obtain the score of the current live stream.

[0129] For example, assume that the first aggregation result is 6, the second aggregation result is 3, and the third aggregation result is 1. Then, the first normalization result = 6 / (6 + 3 + 1) = 0.6, the second normalization result = 3 / (6 + 3 + 1) = 0.3, and the third normalization result = 1 / (6 + 3 + 1) = 0.1. Further, the electronic device can input 0.6, 0.3, and 0.1 into the integrated tree model to obtain the score of the current live stream.

[0130] In some implementable ways, the integrated tree model can be a random forest, a gradient boosting tree, an XGBoost integrated tree, etc., but not limited thereto.

[0131] In some other implementable ways, S230-1A-5a may include:

[0132] S230-1A-5a-1b: Input the scores of the t-th image frame to the t-n-th image frame, the scores of the t-th audio slice to the t-n-th audio slice, and the scores of the t-th text to the t-n-th text into the machine learning model to obtain the score of the current live stream.

[0133] In some implementable ways, the machine learning model can be a neural network model, but not limited thereto.

[0134] In one implementable way, the training device can obtain training samples and the labels of the training samples; train the machine learning model based on the training samples and the labels of the training samples. Among them, each training sample may include the scores of the t-th image frame to the t-n-th image frame, the scores of the t-th audio slice to the t-n-th audio slice, the scores of the t-th text to the t-n-th text, and the score of the live stream to which these image frames, audio slices, and texts belong. The score of the live stream is used as the label of the training sample.

[0135] Among them, in the training stage of the machine learning model, the electronic device can determine the loss based on the above training samples and the labels of the above training samples. Further, the electronic device can minimize the loss to adjust the parameters of the machine learning model until the loss converges or the number of training times reaches a preset number.

[0136] Optionally, the loss can be cross-entropy (Cross Entropy Loss) loss, L1 norm (L1) loss, mean square error (Mean Square Error, MSE), etc., but not limited thereto.

[0137] In some other implementable ways, S230-1A may include:

[0138] S230-1A-1b: Input the t-th to (t-n)-th image frames into the image recognition model to obtain the scores of the t-th to (t-n)-th image frames respectively;

[0139] It should be understood that the explanation of S230-1A-1b can refer to the above explanation of S230-1A-1a, and the embodiments of this application will not repeat it here.

[0140] S230-1A-2b: Input the t-th to (t-n)-th audio slices into the audio recognition model to obtain the scores of the t-th to (t-n)-th audio slices respectively;

[0141] It should be understood that the explanation of S230-1A-2b can refer to the above explanation of S230-1A-2a, and the embodiments of this application will not repeat it here.

[0142] S230-1A-3b: Based on the scores of the t-th to (t-n)-th image frames and the scores of the t-th to (t-n)-th audio slices respectively, obtain the score of the current live stream.

[0143] Among them, the electronic device can implement S230-1A-3b through any of the following implementable ways, but not limited to this:

[0144] In some implementable ways, S230-1A-3b may include:

[0145] S230-1A-3b-1a: Aggregate the scores of the t-th to (t-n)-th image frames respectively to obtain the first aggregation result;

[0146] It should be understood that the explanation of S230-1A-3b-1a can refer to the above explanation of S230-1A-5a-1a, and the embodiments of this application will not repeat it here.

[0147] S230-1A-3b-2a: Aggregate the scores of the t-th to (t-n)-th audio slices respectively to obtain the second aggregation result;

[0148] It should be understood that the explanation of S230-1A-3b-2a can refer to the above explanation of S230-1A-5a-2a, and the embodiments of this application will not repeat it here.

[0149] S230-1A-3b-3a: Based on the first aggregation result and the second aggregation result, obtain the score of the current live stream.

[0150] In some implementations, the electronic device may input the first aggregation result and the second aggregation result into an ensemble tree model to obtain the score of the current live stream.

[0151] In some other implementations, the electronic device first performs normalization processing on the first aggregation result and the second aggregation result to obtain the first normalization processing result corresponding to the first aggregation result and the second normalization processing result corresponding to the second aggregation result. Further, the first normalization processing result and the second normalization processing result are input into the ensemble tree model to obtain the score of the current live stream.

[0152] For example, assuming the first aggregation result is 6 and the second aggregation result is 4, then the first normalization processing result = 6 / (6 + 4) = 0.6, and the second normalization processing result = 4 / (6 + 4) = 0.4. Further, the electronic device may input 0.6 and 0.4 into the ensemble tree model to obtain the score of the current live stream.

[0153] In some other implementations, S230-1A-3b may include:

[0154] S230-1A-3b-1b: Input the scores of each of the t-th image frame to the t-n-th image frame and the scores of each of the t-th audio slice to the t-n-th audio slice into a machine learning model to obtain the score of the current live stream.

[0155] In some implementations, the machine learning model may be a neural network model, but is not limited thereto.

[0156] In one implementation, the training device may obtain training samples and labels of the training samples; train the machine learning model based on the training samples and the labels of the training samples. Wherein, each training sample may include the scores of each of the t-th image frame to the t-n-th image frame, the scores of each of the t-th audio slice to the t-n-th audio slice, and the score of the live stream to which these image frames and audio slices belong, and the score of the live stream is used as the label of the training sample.

[0157] Wherein, in the training stage of the machine learning model, the electronic device may determine the loss based on the above training samples and the labels of the above training samples. Further, the electronic device may minimize the loss to adjust the parameters of the machine learning model until the loss converges or the number of training times reaches a preset number.

[0158] Optionally, the loss may be cross-entropy (Cross Entropy Loss) loss, L1 norm (L1) loss, MSE, etc., but is not limited thereto.

[0159] S230-2A: Detect whether the current live stream carries content of a target type based on the score of the current live stream.

[0160] In some implementable ways, if the score of the current live stream is greater than a preset score, it is determined that the current live stream carries content of the target type; if the score of the current live stream is less than or equal to the preset score, it is determined that the current live stream does not carry content of the target type.

[0161] For example, assume that the preset score corresponding to the score of the current live stream is 6, and the score of the current live stream is 7. Since the score of the current live stream is greater than the preset score, it is determined that the current live stream carries content of the target type.

[0162] In some other implementable ways, if the score of the current live stream is greater than or equal to the preset score, it is determined that the current live stream carries content of the target type; if the score of the current live stream is less than the preset score, it is determined that the current live stream does not carry content of the target type.

[0163] It should be understood that the preset score corresponding to the score of the current live stream may be the same as or different from the preset scores corresponding to the scores of the above-mentioned image frames, audio slices, and texts. The embodiments of the present application do not limit this.

[0164] Implementable way two, S230 may include:

[0165] S230-1B: Input the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice into a machine learning model to detect whether the current live stream carries content of the target type.

[0166] In some implementable ways, the machine learning model may be a neural network model, but is not limited thereto.

[0167] In one implementable way, a training device may obtain training samples and labels of the training samples; train the machine learning model based on the training samples and the labels of the training samples. Each training sample may include the t-th image frame to the (t - n)-th image frame, the t-th audio slice to the (t - n)-th audio slice, and indication information on whether the current live stream carries content of the target type, where the indication information serves as the label of the training sample.

[0168] Among them, in the training stage of the machine learning model, the electronic device may determine a loss based on the above training samples and the labels of the above training samples. Further, the electronic device may minimize the loss to adjust the parameters of the machine learning model until the loss converges or the number of training times reaches a preset number.

[0169] Optionally, the loss can be Cross Entropy Loss, L1 loss, MSE, etc., but not limited thereto.

[0170] The following is an exemplary illustration of the live stream processing method provided by the embodiments of the present application through an example:

[0171] Figure 3 It is a schematic diagram of a live stream processing process provided by the embodiments of the present application. As Figure 3 shown, for the live stream, the electronic device can extract frames from the image stream of the live stream to obtain the t-th image frame to the (t - n)-th image frame; and cut the audio stream of the live stream to obtain the t-th audio slice to the (t - n)-th audio slice.

[0172] Further, the electronic device can input the t-th image frame to the (t - n)-th image frame into an image recognition model to obtain the scores of the t-th image frame to the (t - n)-th image frame respectively. Among them, the scores of the t-th image frame to the (t - n)-th image frame are respectively denoted as: A_t, A_t-1... A_t-n, and they form a set A. The electronic device can input the t-th audio slice to the (t - n)-th audio slice into an audio recognition model to obtain the scores of the t-th audio slice to the (t - n)-th audio slice respectively. Among them, the scores of the t-th audio slice to the (t - n)-th audio slice are respectively denoted as: B_t, B_t-1... B_t-n, and they form a set B. The electronic device can input the t-th audio slice to the (t - n)-th audio slice into a speech recognition model to obtain the t-th text to the (t - n)-th text; input the t-th text to the (t - n)-th text into a text recognition model to obtain the scores of the t-th text to the (t - n)-th text respectively. Among them, the scores of the t-th text to the (t - n)-th text are respectively denoted as: C_t, C_t-1... C_t-n, and they form a set C.

[0173] Furthermore, the electronic device can use the maximum value MAX{A}, average value AVG{A}, or sum result SUM{A} among the scores of the t-th image frame to the (t - n)-th image frame as the first aggregation result; use the maximum value MAX{B}, average value AVG{B}, or sum result SUM{B} among the scores of the t-th audio slice to the (t - n)-th audio slice as the second aggregation result; use the maximum value MAX{C}, average value AVG{C}, or sum result SUM{C} among the scores of the t-th text to the (t - n)-th text as the third aggregation result.

[0174] Finally, the electronic device can input the first aggregation result, the second aggregation result, and the third aggregation result into the XGBoost integrated tree model to obtain the score G of the live stream. The rule for detecting that the live stream carries content of the target type is that the score G is greater than the preset score g.

[0175] In summary, the embodiment of the present application provides a method for processing a live stream, including: extracting frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame; cutting the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice; where n is a positive integer; based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice, detecting whether the current live stream carries content of the target type. Since the embodiment of the present application introduces richer previous context information, such as the (t - 1)-th image frame to the (t - n)-th image frame, the (t - 1)-th audio slice to the (t - n)-th audio slice, and the image frames and audio slices in the live scene have strong continuity, the current image frame is similar to the previous image frames, and the current audio slice is similar to the previous audio slices. Therefore, this previous context information can play a role in information supplementation, and thus can improve the content detection accuracy.

[0176] Furthermore, the embodiment of the present application proposes to detect whether the current live stream carries content of the target type through an integrated tree model. The integrated tree model can use the image frames and audio slices before the current time t to enrich the understanding of the image frame and audio slice at time t by the integrated tree model, enabling the integrated tree model to have a greater certainty in detecting whether the current live stream carries content of the target type, and thus can improve the content detection accuracy. And by using the integrated tree model, the first aggregation result, the second aggregation result, and the third aggregation result can be aligned to a score, which is beneficial to the formulation of the final rule.

[0177] In addition, Figure 4 is a schematic diagram of a live stream processing process provided by the related art. As Figure 4 shown, for the live stream, the electronic device can obtain the current image frame and the current audio slice.

[0178] Furthermore, the electronic device can input the current image frame into the image recognition model to obtain the score of the current image frame. Assume that this score is denoted as A. The electronic device can input the current audio slice into the audio recognition model to obtain the score of the current audio slice. Assume that this score is denoted as B. The electronic device can input the current audio slice into the speech recognition model to obtain the current text; input the current text into the text recognition model to obtain the score of the current text. Assume that this score is denoted as C.

[0179] Finally, the electronic device can detect whether the live stream carries content of the target type based on the following rules:

[0180] (A > X1 OR B > Y1 OR C > Z1)

[0181] OR (A > X2 OR B > Y2)

[0182] OR (A > X3 OR C > Z2)

[0183] OR (B > Y3 OR C > Z3)

[0184] OR (A > X4 AND B > Y4 OR C > Z4)

[0185] It should be understood that in Figure 4 the provided technical solution, the audio slices and the image frames are independent. When the embodiments of the present application use an integrated tree model to detect whether a live stream carries content of a target type, compared with Figure 4 the provided technical solution, the technical solution provided by the embodiments of the present application can comprehensively utilize the audio slices and the image frames, thereby further improving the content detection accuracy.

[0186] In addition, when the embodiments of the present application use an integrated tree model to detect whether a live stream carries content of a target type, the rules involved in the embodiments of the present application are fewer than those Figure 4 in the provided technical solution. Especially when Figure 4 the provided technical solution involves multiple single models, such as multiple image recognition models, the rules involved in the embodiments of the present application are even fewer than those Figure 4 in the provided technical solution, thereby reducing the detection complexity and improving the detection efficiency.

[0187] When the embodiments of the present application use an integrated tree model to detect whether a live stream carries content of a target type, since the embodiments of the present application involve many models, including: image recognition model, audio recognition model, speech recognition model, text recognition model, integrated tree model, etc., therefore, in order to improve the model training efficiency, in the short to medium term, only the integrated tree model needs to be trained:

[0188] In some implementable ways, the electronic device can obtain the training set corresponding to the integrated tree model; based on the training set corresponding to the integrated tree model, train the integrated tree model.

[0189] In some implementable ways, the electronic device can calculate the fourth aggregation result of the historical live stream according to the calculation method of the first aggregation result; calculate the fifth aggregation result of the historical live stream according to the calculation method of the second aggregation result; calculate the sixth aggregation result of the historical live stream according to the calculation method of the third aggregation result; based on the fourth aggregation result, the fifth aggregation result and the sixth aggregation result, obtain the training set corresponding to the integrated tree model.

[0190] It should be understood that the electronic device and Figure 2 the electronic device involved in the corresponding embodiment may be the same electronic device or different electronic devices. For example, if the Figure 2 electronic device involved in the corresponding embodiment is understood as an execution device, then the electronic device here can be understood as a training device, and the execution device and the training device may be the same device or different devices.

[0191] It should be understood that the training method of the integrated tree model in the embodiments of the present application may be a supervised training method or an unsupervised training method. If the supervised training method is adopted, then the above fourth aggregation result, fifth aggregation result, sixth aggregation result, and the indication information of whether the historical live stream carries content of the target type will constitute a training sample, where the indication information serves as the label of the training sample. If the unsupervised training method is adopted, then the above fourth aggregation result, fifth aggregation result, and sixth aggregation result will constitute a training sample.

[0192] In some implementable ways, the electronic device may perform periodic training on the integrated tree model, and the period may be one day, three days, one week, etc., but is not limited thereto.

[0193] For example, Figure 5 is a schematic diagram of a model training method provided by an embodiment of the present application. As Figure 5 shown, the electronic device may obtain a new data distribution every 24 hours, train the integrated tree model based on the new data distribution, and after training the integrated tree model, effect verification may be performed. If the standard is met, the trained integrated tree model may be put on the line. If the standard is not met, manual intervention may be used to correct the integrated tree model, etc.

[0194] Among them, the above new data distribution is derived from the latest data generated by the online product, where the online product is used to detect whether the live stream carries content of the target type.

[0195] In some implementation ways, when performing effect verification, the following evaluation metrics may be used to evaluate the integrated tree model, but are not limited thereto: accuracy, precision, recall, F1 score, etc.

[0196] Among them, accuracy is the ratio of the number of correctly predicted samples to the total number of samples. Precision is the ratio of the number of truly positive samples among the samples predicted as positive to the number of samples predicted as positive. Recall is the ratio of the number of samples predicted as positive among the truly positive samples to the number of truly positive samples. The F1 score (F-Measure) is the harmonic mean of precision and recall and is used to comprehensively evaluate the performance of the model.

[0197] In summary, in the embodiments of the present application, in the short to medium term, only the integrated tree model can be considered for training. When training the integrated tree model, it can also learn the data distribution changes of the upstream unimodal model. Since the number of parameters of this model is relatively low and the update difficulty is small, the model efficiency can be improved.

[0198] Figure 6 The statistical chart of the effects provided by the embodiments of the present application is as Figure 6 shown. After adopting the live stream processing method provided by the embodiments of the present application, the coverage rate has been significantly improved compared with that before adopting the live stream processing method provided by the embodiments of the present application.

[0199] Among them, the coverage rate represents the ratio of the live streams carrying the content of the target type detected by the live stream processing method to M for M live streams carrying the content of the target type, where M is an integer greater than 1.

[0200] As Figure 6 shown. After adopting the live stream processing method provided by the embodiments of the present application, the coverage rate is stable above 80%. While adopting the Figure 4 corresponding live processing method, the corresponding coverage rate is at the level of 40%-65%. It can be seen that for the live stream processing method provided by the embodiments of the present application, compared with the live stream processing method provided by the related technology, the coverage rate has increased by 15%-40%. And if the integrated tree model is updated automatically at regular intervals, the effect is in a relatively stable state within 3 months.

[0201] The preferred embodiments of the present application have been described in detail above with reference to the accompanying drawings. However, the present application is not limited to the specific details in the above embodiments. Within the technical concept of the present application, various simple modifications can be made to the technical solutions of the present application, and these simple modifications all fall within the protection scope of the present application. For example, among the specific technical features described in the above specific embodiments, they can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present application does not separately describe various possible combination methods. Again, any combination can be made between various different embodiments of the present application as long as it does not violate the idea of the present application, and it should also be regarded as the content disclosed by the present application.

[0202] It should also be understood that in various method embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0203] The method provided by the embodiments of the present application has been described above. Next, the live stream processing device provided by the embodiments of the present application will be described.

[0204] Figure 7 It is a schematic diagram of a live stream processing device 700 provided by the embodiments of the present application. As Figure 7 shown, the device 700 includes: a frame extraction module 710, a cutting module 720, and a detection module 730. Among them, the frame extraction module 710 is used to extract frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame; the cutting module 720 is used to cut the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice; where n is a positive integer; the detection module 730 is used to detect whether the current live stream carries content of a target type based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice.

[0205] In some implementable ways, the detection module 730 is specifically used to: obtain the score of the current live stream based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice; detect whether the current live stream carries content of a target type based on the score of the current live stream.

[0206] In some implementable ways, the detection module 730 is specifically used to: input the t-th image frame to the (t - n)-th image frame into an image recognition model to obtain the scores of the t-th image frame to the (t - n)-th image frame respectively; input the t-th audio slice to the (t - n)-th audio slice into an audio recognition model to obtain the scores of the t-th audio slice to the (t - n)-th audio slice respectively; input the t-th audio slice to the (t - n)-th audio slice into a speech recognition model to obtain the t-th text to the (t - n)-th text; input the t-th text to the (t - n)-th text into a text recognition model to obtain the scores of the t-th text to the (t - n)-th text respectively; obtain the score of the current live stream based on the scores of the t-th image frame to the (t - n)-th image frame respectively, the scores of the t-th audio slice to the (t - n)-th audio slice respectively, and the scores of the t-th text to the (t - n)-th text respectively.

[0207] In some implementable ways, the detection module 730 is specifically used to: aggregate the scores of the t-th image frame to the (t - n)-th image frame respectively to obtain a first aggregation result; aggregate the scores of the t-th audio slice to the (t - n)-th audio slice respectively to obtain a second aggregation result; aggregate the scores of the t-th text to the (t - n)-th text respectively to obtain a third aggregation result; obtain the score of the current live stream based on the first aggregation result, the second aggregation result, and the third aggregation result.

[0208] In some implementable ways, the detection module 730 is specifically configured to: use the maximum value, average value, or summation result of the scores of each of the t-th image frame to the (t - n)-th image frame as the first aggregation result.

[0209] In some implementable ways, the detection module 730 is specifically configured to: use the maximum value, average value, or summation result of the scores of each of the t-th audio slice to the (t - n)-th audio slice as the second aggregation result.

[0210] In some implementable ways, the detection module 730 is specifically configured to: use the maximum value, average value, or summation result of the scores of each of the t-th text to the (t - n)-th text as the third aggregation result.

[0211] In some implementable ways, the detection module 730 is specifically configured to: input the first aggregation result, the second aggregation result, and the third aggregation result into an ensemble tree model to obtain the score of the current live stream.

[0212] In some implementable ways, the apparatus 700 further includes: an acquisition module 740 and a training module 750, where the acquisition module 740 is configured to acquire a training set corresponding to the ensemble tree model; the training module 750 is configured to train the ensemble tree model based on the training set corresponding to the ensemble tree model.

[0213] In some implementable ways, the acquisition module 740 is specifically configured to: calculate a fourth aggregation result of a historical live stream according to the calculation method of the first aggregation result; calculate a fifth aggregation result of the historical live stream according to the calculation method of the second aggregation result; calculate a sixth aggregation result of the historical live stream according to the calculation method of the third aggregation result; and obtain the training set corresponding to the ensemble tree model based on the fourth aggregation result, the fifth aggregation result, and the sixth aggregation result.

[0214] In some implementable ways, the detection module 730 is specifically configured to: input the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice into a machine learning model to detect whether the current live stream carries content of a target type.

[0215] In some implementable ways, the detection module 730 is specifically configured to: if the score of the current live stream is greater than a preset score, determine that the current live stream carries content of a target type; if the score of the current live stream is less than or equal to the preset score, determine that the current live stream does not carry content of a target type.

[0216] It should be understood that the apparatus embodiments and the method embodiments can correspond to each other, and similar descriptions can refer to the method embodiments. To avoid repetition, they are not elaborated here. Specifically, Figure 7 the illustrated apparatus 700 can execute Figure 2The corresponding method embodiments, and the foregoing and other operations and / or functions of each module in the apparatus 700 are respectively for implementing Figure 2 the corresponding processes in each of the methods in, for the sake of brevity, will not be described herein again.

[0217] The apparatus 700 of the embodiments of the present application has been described above from the perspective of functional modules with reference to the accompanying drawings. It should be understood that the functional modules can be implemented in the form of hardware, or in the form of instructions in software, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in the present application can be completed by the integrated logic circuit in hardware in the processor and / or instructions in software form. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. Optionally, the software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps in the above method embodiments.

[0218] Figure 8 is a schematic block diagram of an electronic device provided by an embodiment of the present application.

[0219] As Figure 8 shown, the electronic device may include:

[0220] A memory 810 and a processor 820. The memory 810 is used to store a computer program and transmit the program code to the processor 820. In other words, the processor 820 can call and run the computer program from the memory 810 to implement the method in the embodiments of the present application.

[0221] For example, the processor 820 can be used to execute the above method embodiments according to the instructions in the computer program.

[0222] In some embodiments of the present application, the processor 820 may include, but is not limited to:

[0223] A general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0224] In some embodiments of the present application, the memory 810 includes, but is not limited to:

[0225] Volatile memory and / or non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0226] In some embodiments of the present application, the computer program may be divided into one or more modules, and the one or more modules are stored in the memory 810 and executed by the processor 820 to complete the method provided by the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.

[0227] As Figure 8 shown, the electronic device may further include:

[0228] A transceiver 830, and the transceiver 830 may be connected to the processor 820 or the memory 810.

[0229] Among them, the processor 820 may control the transceiver 830 to communicate with other devices. Specifically, it may send information or data to other devices, or receive information or data sent by other devices. The transceiver 830 may include a transmitter and a receiver. The transceiver 830 may further include an antenna, and the number of antennas may be one or more.

[0230] It should be understood that the various components in the electronic device are connected through a bus system. Among them, the bus system includes not only a data bus, but also a power bus, a control bus, and a status signal bus.

[0231] This application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a computer, the computer can execute the methods in the above method embodiments. Or rather, the embodiments of this application also provide a computer program product containing instructions. When the instructions are executed by a computer, the computer executes the methods in the above method embodiments.

[0232] When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc.

[0233] Those of ordinary skill in the art can realize that the modules and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0234] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or modules can be in electrical, mechanical, or other forms.

[0235] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0236] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A live stream processing method, characterized in that, Including: Taking frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame; Cutting the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice; where n is a positive integer; Based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice, detecting whether the current live stream carries content of a target type.

2. The method according to claim 1, characterized in that, The detecting whether the current live stream carries content of a target type based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice includes: Based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice, obtaining a score of the current live stream; Based on the score of the current live stream, detecting whether the current live stream carries content of a target type.

3. The method according to claim 2, wherein The obtaining a score of the current live stream based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice includes: Inputting the t-th image frame to the (t - n)-th image frame into an image recognition model to obtain respective scores of the t-th image frame to the (t - n)-th image frame; Inputting the t-th audio slice to the (t - n)-th audio slice into an audio recognition model to obtain respective scores of the t-th audio slice to the (t - n)-th audio slice; Inputting the t-th audio slice to the (t - n)-th audio slice into a speech recognition model to obtain the t-th text to the (t - n)-th text; Inputting the t-th text to the (t - n)-th text into a text recognition model to obtain respective scores of the t-th text to the (t - n)-th text; Based on the respective scores of the t-th image frame to the (t - n)-th image frame, the respective scores of the t-th audio slice to the (t - n)-th audio slice, and the respective scores of the t-th text to the (t - n)-th text, obtaining a score of the current live stream.

4. The method according to claim 3, characterized in that, The obtaining a score of the current live stream based on the respective scores of the t-th image frame to the (t - n)-th image frame, the respective scores of the t-th audio slice to the (t - n)-th audio slice, and the respective scores of the t-th text to the (t - n)-th text includes: Aggregating the respective scores of the t-th image frame to the (t - n)-th image frame to obtain a first aggregation result; Aggregating the respective scores of the t-th audio slice to the (t - n)-th audio slice to obtain a second aggregation result; Aggregating the respective scores of the t-th text to the (t - n)-th text to obtain a third aggregation result; Based on the first aggregation result, the second aggregation result, and the third aggregation result, obtaining a score of the current live stream.

5. The method according to claim 4, wherein The aggregating the respective scores of the t-th image frame to the (t - n)-th image frame to obtain a first aggregation result includes: Take the maximum value, average value or summation result of the scores of the t-th image frame to the (t - n)-th image frame as the first aggregation result.

6. The method according to claim 4, characterized in that The aggregating the scores of the t-th audio slice to the (t - n)-th audio slice to obtain a second aggregation result includes: Take the maximum value, average value or summation result of the scores of the t-th audio slice to the (t - n)-th audio slice as the second aggregation result.

7. The method according to claim 4, wherein The aggregating the scores of the t-th text to the (t - n)-th text to obtain a third aggregation result includes: Take the maximum value, average value or summation result of the scores of the t-th text to the (t - n)-th text as the third aggregation result.

8. The method according to any one of claims 4 to 7, characterized in that, The obtaining the score of the current live stream based on the first aggregation result, the second aggregation result and the third aggregation result includes: Input the first aggregation result, the second aggregation result and the third aggregation result into an ensemble tree model to obtain the score of the current live stream.

9. The method according to claim 8, wherein Further includes: Obtain the training set corresponding to the ensemble tree model; Train the ensemble tree model based on the training set corresponding to the ensemble tree model.

10. The method according to claim 9, wherein The obtaining the training set corresponding to the ensemble tree model includes: Calculate the fourth aggregation result of the historical live stream according to the calculation method of the first aggregation result; Calculate the fifth aggregation result of the historical live stream according to the calculation method of the second aggregation result; Calculate the sixth aggregation result of the historical live stream according to the calculation method of the third aggregation result; Based on the fourth aggregation result, the fifth aggregation result and the sixth aggregation result, obtain the training set corresponding to the ensemble tree model.

11. The method according to claim 1, wherein The detecting whether the current live stream carries content of a target type based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice includes: Input the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice into a machine learning model to detect whether the current live stream carries content of a target type.

12. The method according to any one of claims 2-7, characterized in that, The detecting whether the current live stream carries content of a target type based on the score of the current live stream includes: If the score of the current live stream is greater than a preset score, determine that the current live stream carries the content of the target type; If the score of the current live stream is less than or equal to the preset score, determine that the current live stream does not carry the content of the target type.

13. A live stream processing device, characterized in that, Includes: A frame extraction module for extracting frames from the image stream of the current live stream to obtain the t-th image frame to the (t - n)-th image frame; A cutting module for cutting the audio stream of the current live stream to obtain the t-th audio slice to the (t - n)-th audio slice; where n is a positive integer; A detection module for detecting whether the current live stream carries content of a target type based on the t-th image frame to the (t - n)-th image frame and the t-th audio slice to the (t - n)-th audio slice.

14. An electronic device, characterized in that, Includes: A processor and a memory, the memory being configured to store a computer program, the processor being configured to call and run the computer program stored in the memory to execute the method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causing a computer to execute the method according to any one of claims 1 to 12.