A method and system for detecting abnormal behavior in live streaming
Patent Information
- Application Number
- CN202511594396.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-11-03
AI Technical Summary
[0005]为解决上述仅依靠直播画面和空播时长进行空播检测存在的准确性低的技术问题,本发明在如下的多个方面中提供方案
[0007]有益效果:通过融合视频和音频判断主播或直播间是否存在异常行为,相对于仅依靠画面进行判断,准确性更高。
Smart Images

Figure CN121415316B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of abnormal behavior detection technology, and in particular to a method and system for detecting abnormal behavior in live streaming. Background Technology
[0002] With the development of the e-commerce industry, live streaming has become one of the main ways to promote products. For businesses, ensuring the compliance of live streaming content and a good user experience is crucial, and this depends on the live streamer's performance. If abnormal behaviors such as being offline or having an empty screen occur in the live stream, it will not only damage the user's viewing experience but also negatively impact the business's development.
[0003] To detect anomalies in live streaming rooms, Chinese patent application CN111586432A discloses a method, apparatus, server, and storage medium for determining empty live streaming rooms. The method includes acquiring a first live stream frame of a target live streaming room to be detected; extracting features from the first live stream frame to obtain first image features; determining that a first streamer user is not present in the target live streaming room if the first image features do not include first facial features; calculating the empty streaming duration before the current first time period; and determining that the target live streaming room is an empty live streaming room if the empty streaming duration is greater than a target duration. While this method can detect empty streaming behavior to some extent, it relies solely on determining whether a live stream contains facial features and the empty streaming duration to decide whether it is an empty stream. It fails to consider scenarios where the streamer's facial features are reasonably not visible, such as when products briefly obscure the streamer's face or show their back, which can easily lead to misjudgments.
[0004] Therefore, improving the accuracy of detecting abnormal behavior in live streaming is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] To address the technical problem of low accuracy in detecting empty broadcasts by relying solely on live stream footage and empty broadcast duration, this invention provides solutions in the following aspects.
[0006] In a first aspect, the present invention provides a method for detecting abnormal behavior in live streaming, comprising: acquiring video data and audio data of a live streaming room to be detected; inputting the audio data into a trained human voice judgment model to obtain a first classification result; inputting the video data into a trained target detection model to obtain a second classification result; and determining whether abnormal behavior exists in the live streaming room to be detected based on the first classification result and the second classification result using a preset rule engine.
[0007] Beneficial effects: By combining video and audio, the ability to determine whether a streamer or live stream room is exhibiting abnormal behavior is more accurate than relying solely on visuals.
[0008] Furthermore, the first classification result includes no voice and voice, and the second classification result includes face and no face. A preset rule engine is used to determine whether there is abnormal behavior in the live broadcast room to be detected, including: if the first classification result is no voice and the second classification result is no face, then it is determined that there is abnormal behavior; if the first classification result is no voice and the second classification result is face, an audio-video alignment model is used to detect whether the sound matches the lip movements of the person in the picture, and if not, it is determined that there is abnormal behavior.
[0009] Beneficial effects: By further analyzing whether the sound in the video matches the lip movements of a person when the first classification result is no voice and the second classification result is a face, it is possible to avoid situations where background music masks the anchor's voice or the anchor whispers and other scenes are misjudged as abnormal, thereby improving the accuracy of abnormal behavior detection.
[0010] Furthermore, an audio-video alignment model is used to detect whether the sound matches the lip movements of a person in the video. This includes: extracting the lip movement feature sequence from the video data and the corresponding audio feature sequence from the audio data; inputting the lip movement feature sequence and the audio feature sequence into the trained audio-video alignment model to obtain the matching degree between the sound and the lip movements; if the matching degree is less than or equal to a preset matching threshold, it is determined to be a mismatch.
[0011] Furthermore, the method of using a preset rule engine to determine whether there is abnormal behavior in the live broadcast room to be detected also includes: if the first category result is that there is voice and the second category result is that there is no face, then the duration of the absence of face and the presence of voice is counted, and it is determined whether the duration exceeds the preset time range. If so, it is determined that there is abnormal behavior.
[0012] Beneficial effects: By combining the duration of voice and facelessness when the first category result is voice and the second category result is no face, the system can avoid misjudging reasonable scenarios where the streamer's facial features are not visible as abnormal, thereby improving the accuracy of detecting abnormal behavior in live streaming.
[0013] Furthermore, the method also includes: obtaining the first facial feature vector of the anchor pre-bound to the live broadcast room to be detected; extracting the second facial feature vector from the video data; determining whether the first facial feature vector and the second facial feature vector match; if not, triggering a warning of abnormal behavior.
[0014] Beneficial effects: By detecting whether the streamer in the live stream to be detected is the same as the streamer pre-bound to the live stream to be detected, it is possible to accurately detect whether there is a proxy broadcasting behavior, thus meeting different abnormal behavior monitoring needs.
[0015] Further, determining whether the first face feature vector matches the second face feature vector includes: calculating the similarity between the first face feature vector and the second face feature vector using cosine similarity or Euclidean distance; if the similarity is greater than the similarity threshold, it is determined to be a match.
[0016] Furthermore, the first classification result is obtained, including: if the probability of human voice output by the human voice determination model is less than the preset threshold for multiple consecutive times, and the volume is less than the average environmental noise, then the first classification result is no human voice; otherwise, the first classification result is human voice.
[0017] Furthermore, the voice recognition model adopts an architecture that combines CNN and LSTM models, while the object detection model is a CNN model.
[0018] Beneficial effects: By using an architecture that combines CNN and LSTM models for human voice detection, the stability and reliability of detection are higher than that of a single model in scenarios such as intermittent speech, low-volume speech, and background noise.
[0019] Furthermore, the method also includes: performing statistical analysis on the data of all live streaming rooms and displaying the obtained key indicators; the key indicators include the current number of live streams and the number of abnormal streams, the percentage of abnormal streams, and the number of abnormal streams in the past 24 hours.
[0020] Beneficial effects: By displaying these key metrics, managers can intuitively and quickly understand the live broadcast situation, thereby making accurate decisions swiftly.
[0021] In a second aspect, the present invention provides a live streaming abnormal behavior detection system, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the live streaming abnormal behavior detection method described in the first aspect is implemented.
[0022] The beneficial effect of this invention is that, compared to judging solely based on image data, the probability of false positives is lower when detecting by fusing both image and audio data. Therefore, by first detecting audio and video data, and then using a preset rule engine to make judgments based on the classification results, the accuracy of detecting abnormal behavior in live streaming is improved. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a live streaming abnormal behavior detection method according to an embodiment of the present invention; Figure 2 This is a schematic block diagram illustrating the structure of a live-streaming abnormal behavior detection system according to an embodiment of the present invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0026] Figure 1 This is a flowchart illustrating a live streaming abnormal behavior detection method according to an embodiment of the present invention.
[0027] To address the technical problem of insufficient accuracy in detecting abnormal behavior during live streaming in existing technologies, in a first aspect, the present invention provides a method for detecting abnormal behavior during live streaming. Specifically, as follows... Figure 1 As shown, the method of the present invention includes the following steps.
[0028] S101. Obtain the video and audio data of the live stream to be tested.
[0029] In one embodiment, video and audio data of the live stream to be detected can be obtained through relevant API interfaces.
[0030] S102. Input the audio data into the trained human voice judgment model to obtain the first classification result; at the same time, input the video data into the trained object detection model to obtain the second classification result.
[0031] In this embodiment, the first classification result includes both spoken and unspoken voices. The spoken voice detection model employs a fusion architecture combining a CNN (Convolutional Neural Networks) model and an LSTM (Long Short-Term Memory) model. The CNN model extracts local spatiotemporal features of the acoustic feature sequence, while the LSTM model captures the temporal dependencies of the acoustic feature sequence. The outputs of the CNN and LSTM models are fused through a fusion layer. Compared to a single model, such as a CNN or RNN (Recurrent Neural Network), this CNN / LSTM fusion architecture is more sensitive to scenarios such as intermittent or low-volume speech, thereby improving the reliability of spoken voice detection.
[0032] Specifically, the audio data undergoes preprocessing. Specifically, the audio data is divided into short-time frames to obtain audio frames, for example, frames with a length of 25ms and a frame shift of 10ms. In optional embodiments, preprocessing may further include noise reduction filtering, such as using a low-pass filter to remove high-frequency noise from the audio data, thereby improving the quality of the audio data and reducing noise interference with subsequent analysis.
[0033] Furthermore, a short-time Fourier transform is used to convert the audio frame into a spectrogram, and acoustic features are extracted from the spectrogram. In this embodiment, the acoustic features may include: Mel frequency cepstral coefficients, zero-crossing rate, volume envelope, spectral centroid, etc., which are used to distinguish between human voices, ambient noise, and background music.
[0034] Furthermore, the acoustic features and spectrograms are input into the trained human voice determination model to obtain the probability of human voice presence. If the probability of human voice presence in multiple consecutive frames is less than a preset threshold, and the volume is lower than the average ambient noise level, then the first classification result is determined to be no human voice presence; otherwise, the first classification result is determined to be human voice presence.
[0035] In one embodiment, the mean ambient noise can be determined using a silent zone estimation method. Specifically, within the time period during which no human voice segments are detected, the root mean square energy of these segments is calculated, and the average value of the root mean square energy is used as the mean ambient noise. In alternative embodiments, it can also be determined using an adaptive sliding window method. Specifically, a time sliding window, such as 30 seconds, is used to select the latest audio data segment, calculate the volume envelope within the time sliding window, and then filter out audio frames whose volume envelope is less than a preset energy threshold. The average energy of these audio frames is calculated to obtain the mean ambient noise. The mean ambient noise calculated using this time sliding window method is more consistent with reality, thereby improving the accuracy of human voice detection.
[0036] In one embodiment, when the first classification result is no voice, a timer is started. If the subsequent detection result is voice, the timer is reset to obtain the duration of the no-voice event. When the duration of the no-voice event exceeds a set threshold, such as 30 seconds, an alarm for an abnormal no-voice event is triggered.
[0037] By combining the judgment of the human voice judgment model with business rules to determine whether there is human voice, the interpretability of the business and the stability of the model output can be improved, while avoiding misjudgments caused by short pauses, background music or model detection errors.
[0038] In this embodiment, the object detection model is a CNN model, and the second classification result includes faces and no faces. Empty screens, black screens, and missing faces are all considered as no faces.
[0039] Since face detection using CNN models is an existing technology, it will not be elaborated upon here.
[0040] S103. Based on the first and second classification results, a preset rule engine is used to determine whether there is any abnormal behavior in the live broadcast room to be detected.
[0041] Specifically, if the first category result is no voice and the second category result is no face, it indicates that there is an abnormal state in the live stream. For example, the streamer is absent, the streamer is AFK, or the live stream is interrupted. This is considered abnormal behavior and triggers an alert that the streamer may be AFK.
[0042] In optional embodiments, to further improve the accuracy of detection, if the first classification result is no voice and the second classification result is no face, it can be determined whether the duration of the absence of both voice and face exceeds a preset time range. If so, it is determined that there is abnormal behavior. This time range can be determined based on historical data of the streamer during normal live streaming. Specifically, the duration of each instance where the absence of both face and voice was judged as normal is statistically analyzed in historical data. Then, the average and standard deviation of these durations are calculated. The sum of the standard deviation and the average is used as the upper limit of the time range, and the difference between the standard deviation and the average is used as the lower limit of the time range.
[0043] By combining historical data to determine the time range, and then determining whether there is abnormal behavior based on whether the duration of the absence of a face and voice is within that time range, it is possible to avoid situations where people do not appear on camera for reasonable reasons, such as looking up to drink water, looking down to tidy up, or looking for items, thereby improving the accuracy and scene adaptability of abnormal behavior detection.
[0044] If the first classification result is no voice and the second classification result is a face, it indicates that the anchor may be whispering or slacking off. In this case, the audio-video alignment model is used to detect whether the sound matches the lip movements of the person in the picture. If they match, it indicates that the anchor is whispering or that the voice is being masked by background noise, and this is considered normal behavior. If they do not match, it indicates that the sound source is inconsistent with the person in the picture, and the anchor may be slacking off or moving their lips without making a sound, and an alarm for possible slacking off is triggered.
[0045] In this embodiment, the audio-video alignment model can employ SyncNet (a deep learning model for solving audio-video synchronization problems) or AV-HuBERT (Audio-Visual HuBERT). Specifically, lip movement feature sequences are extracted from video data, and audio feature sequences are extracted from audio data. The lip movement feature sequences and audio feature sequences are input into the trained audio-video alignment model to obtain the matching degree between sound and lip movements. It is then determined whether the matching degree is greater than a preset matching threshold; if so, it is considered a match; otherwise, it is considered a mismatch. The preset matching threshold can be set to 0.7 or other values.
[0046] By performing a second detection when the first classification result is no voice and the second classification result is a face, it is possible to avoid misjudgments in scenarios such as the streamer whispering or the background music being too loud and masking the streamer's voice, thereby improving the accuracy and scenario adaptability of abnormal behavior detection in live streaming.
[0047] If the first category result is "voices" and the second category result is "no face," then the duration of the "no face, voices" condition is recorded, and it is determined whether this duration exceeds a preset time range. If so, an abnormal behavior is identified, and an alarm indicating possible absence of the broadcaster is triggered. The method for determining the time range is the same as that for the "no face, no voices" condition, and will not be repeated here.
[0048] By combining the duration of the event with the first category result indicating the presence of voices and the second category result indicating the absence of faces, the system can avoid misjudging scenarios where the streamer's face is reasonably exposed, such as when a product briefly obscures the streamer's face or when the streamer is searching for a product. This improves the accuracy and scenario adaptability of detecting abnormal behavior in live streaming.
[0049] In one embodiment, the method of the present invention further includes: performing proxy broadcasting detection on the streamer in the live broadcast room to be detected; if proxy broadcasting behavior is detected, triggering a warning of proxy broadcasting behavior. It should be noted that proxy broadcasting detection can continue throughout the entire live broadcast, or it can be triggered only when a face is detected.
[0050] Specifically, the first facial feature vector of the streamer pre-bound to the live streaming room to be detected is obtained. In one embodiment, the streamer of the live streaming room to be detected can be determined according to the live streaming schedule, etc., and then the streamer's name is matched with the facial feature vector in the preset facial feature library, and the matched facial feature vector is used as the first facial feature vector.
[0051] Furthermore, a second facial feature vector is extracted from the video data. Specifically, a high-precision face detector is used to detect the anchor's face in the video data, and the detected face is tracked across multiple frames to obtain multiple frames of face images. Then, a deep convolutional network is used to extract the second facial feature vector from each frame of face images. The face detector used is the RetinaFace detector, and the deep convolutional network used is the ArcFace network.
[0052] Furthermore, it is determined whether the first facial feature vector matches the second facial feature vector. Specifically, for any second facial feature vector, the similarity between the second facial feature vector and the first facial feature vector is calculated using cosine similarity or Euclidean distance. It is then determined whether the similarity is greater than a preset similarity threshold. If so, the second facial feature vector is considered to match the first facial feature vector; otherwise, they are considered not to match. This process clearly determines whether each second facial feature vector matches the first facial feature vector.
[0053] Furthermore, the number of matches between the second and first facial feature vectors is counted. If the number of matches exceeds the number of non-matches, the first and second facial feature vectors are determined to be a match; otherwise, they are determined not to match. If there is a non-match, an alarm for proxy broadcasting is triggered.
[0054] By performing multi-frame tracking on detected faces, false detections and duplicate counts can be reduced, thereby improving the accuracy of proxy broadcasting detection. By combining a multi-frame voting strategy, the stability of recognition can be improved, thus enhancing the reliability of proxy broadcasting detection.
[0055] In an optional embodiment, the proxy broadcasting behavior detection can also combine voice and face detection. Specifically, the first facial feature vector and the first voice feature vector of the pre-bound anchor in the live broadcast room to be detected are obtained; the second facial feature vector is extracted from the video data, and the second voice feature vector is extracted from the audio data; the matching degree between the first facial feature vector and the second facial feature vector, and the matching degree between the first voice feature vector and the second voice feature vector are calculated based on cosine similarity or Euclidean distance. If both matching degrees are greater than the corresponding matching degree threshold, it is determined that there is no proxy broadcasting behavior; otherwise, it is determined that proxy broadcasting behavior exists.
[0056] By combining audio and live video footage for comprehensive judgment, the accuracy of detecting proxy broadcasting behavior can be improved.
[0057] In one embodiment, the method of the present invention further includes: summarizing data from all live streaming rooms in real time, and statistically displaying key indicators. In this embodiment, key indicators include the current number of live streams and the number of anomalies, trends over the past 24 hours (e.g., the number of anomalies over the past 24 hours), as well as downtime, anomaly percentage, and occurrence time periods. These key indicators can be displayed using pie charts, line charts, bar charts, tables, etc. By displaying these key indicators, management can intuitively and quickly understand the anomaly distribution patterns, thereby making accurate decisions.
[0058] Furthermore, it can also collect statistics on abnormal segments and durations in each live stream, and generate summary tables and detailed tables, providing a basis for accurately tracing the time and duration of abnormal behavior.
[0059] Furthermore, it is possible to count the number and duration of abnormal behaviors of the streamer, and then generate a ranking list based on the number and duration of abnormal behaviors to quantify the streamer's work status in real time and make accurate decisions.
[0060] Figure 2 This is a schematic diagram illustrating the structural block diagram of the live broadcast abnormal behavior detection system according to this embodiment.
[0061] In a second aspect, the present invention also provides a live streaming abnormal behavior detection system. For example... Figure 2 As shown, the system includes a processor and a memory, the memory storing computer program instructions, which, when executed by the processor, implement the live streaming abnormal behavior detection method described in the first aspect of the present invention.
[0062] The system also includes other components well known to those skilled in the art, such as communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.
[0063] In this invention, the aforementioned memory can be any tangible medium containing or storing a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc., or any other medium that can be used to store desired information and can be accessed by an application, module, or both. Any such computer storage medium can be part of a device or accessible to or connected to a device. Any application or module described in this invention can be implemented using computer-readable / executable instructions that can be stored or otherwise maintained by such a computer-readable medium.
[0064] In the description of this specification, "multiple" means at least two, such as two, three or more, unless otherwise explicitly specified. Furthermore, the steps described above are for clarity only; in implementation, they can be combined into one step or some steps can be broken down into multiple steps, as long as they include the same logical relationships.
[0065] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.
Claims
1. A method for detecting abnormal behavior in live streaming, characterized in that, include: Acquire video and audio data from the live stream to be tested; The audio data is input into the trained human voice recognition model to obtain the first classification result; The video data is input into the trained target detection model to obtain the second classification result; Based on the first classification result and the second classification result, a preset rule engine is used to determine whether there is abnormal behavior in the live broadcast room to be detected; The first classification result includes no voice and voice, and the second classification result includes face and no face; A pre-defined rule engine is used to determine whether there is any abnormal behavior in the live stream to be detected, including: If the first classification result is no voice and the second classification result is no face, then it is determined that there is abnormal behavior. If the first classification result is no voice and the second classification result is a face, the audio-video alignment model is used to detect whether the sound matches the lip movements of the person in the picture. If not, it is determined that there is abnormal behavior. An audio-visual alignment model is used to detect whether the sound matches the lip movements of a person in the video, including: Extract lip movement feature sequences from video data and corresponding audio feature sequences from audio data; The lip movement feature sequence and audio feature sequence are input into the trained audio-video alignment model to obtain the matching degree between sound and lip movement; If the matching degree is less than or equal to the preset matching threshold, it is determined as a non-match; If a match is found, it indicates that the broadcaster is whispering or that their voice is being masked by background noise, which is considered normal behavior. The first classification result includes: If the probability of human voice output by the human voice determination model is less than the preset threshold for multiple consecutive times, and the volume is less than the average ambient noise, then the first classification result is no human voice; otherwise, the first classification result is human voice.
2. The live streaming abnormal behavior detection method according to claim 1, characterized in that, The system uses a pre-defined rule engine to determine whether there is abnormal behavior in the live stream to be detected, and also includes: If the first classification result is "there is voice" and the second classification result is "no face", then the duration of "no face but voice" is counted, and it is determined whether the duration exceeds the preset time range. If so, it is determined that there is abnormal behavior.
3. The live streaming abnormal behavior detection method according to claim 1, characterized in that, Also includes: Obtain the first facial feature vector of the pre-bound anchor in the live stream room to be detected; Extract the second facial feature vector from the video data; Determine whether the first face feature vector matches the second face feature vector; If not, a warning of abnormal behavior will be triggered.
4. The live streaming abnormal behavior detection method according to claim 3, characterized in that, Determining whether the first facial feature vector matches the second facial feature vector includes: The similarity between the first face feature vector and the second face feature vector is calculated using cosine similarity or Euclidean distance. If the similarity is greater than the similarity threshold, it is determined to be a match.
5. The live streaming abnormal behavior detection method according to claim 1, characterized in that, The human voice determination model adopts an architecture that combines a CNN model and an LSTM model, and the object detection model is a CNN model.
6. The live streaming abnormal behavior detection method according to claim 1, characterized in that, Also includes: Statistical analysis of data from all live streaming rooms reveals key indicators; these key indicators include the current number of live streams and the number of abnormal streams, the percentage of abnormal streams, and the number of abnormal streams in the past 24 hours.
7. A live streaming abnormal behavior detection system, characterized in that, It includes a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the live streaming abnormal behavior detection method according to any one of claims 1-6 is implemented.
Citation Information
Patent Citations
Air broadcast and live broadcast room determination method and device, server and storage medium
CN111586432A
Live broadcast room on-hook behavior detection method and device, electronic device and computer readable storage medium
CN112738538A
Sound and picture synchronization detection method and device, electronic equipment and terminal
CN119094739A
Live broadcast room detection method and device, electronic equipment and computer readable storage medium
CN120223915A