Live broadcast behavior detection method and device, electronic equipment and readable storage medium
By acquiring continuous video segments from the live streaming platform as a detection window and extending the time granularity for identifying broadcast behavior, the problem of not being able to accurately identify the start and end times of broadcasts in existing technologies is solved. This enables precise positioning and real-time processing of broadcasts, while reducing resource consumption.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU HUYA TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies cannot accurately identify the start and end times of broadcasting behavior, limiting application scenarios such as real-time alerts and penalties. Furthermore, they suffer from high inference latency and high deployment costs, making them unsuitable for large-scale real-time detection.
By acquiring continuous video segments within a preset time granularity from the start of the live stream as a detection window, the system performs hang-up behavior recognition based on video frames, and repeats the recognition within the detection window by extending the time granularity until there is no hang-up behavior or the maximum allowed duration is reached, thus determining the hang-up duration segment.
It achieves precise positioning of live streaming behavior, meets the real-time reminder and punishment needs of live streaming platforms, reduces resource consumption, and is suitable for large-scale real-time detection.
Smart Images

Figure CN121908041A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of live streaming technology, and more specifically, to a live streaming behavior detection method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] With the rapid development of the online live streaming industry, the quality management of live streaming content has become one of the key challenges in the operation of major live streaming platforms. In actual live streaming, some streamers use a "suspended broadcast" method to maintain activity in their live streams in order to extend their online time, obtain platform rewards, or circumvent regulatory review. "Suspended broadcast" refers to the behavior of streamers playing static images, looping videos, or content without substantial interaction after starting their broadcast; it is essentially a low-quality live streaming behavior. Such behavior not only consumes platform resources and reduces the distribution efficiency of the recommendation system, but also seriously affects the viewing experience of viewers and damages the healthy ecosystem of the platform.
[0003] Currently, the industry has proposed various technical means to identify low-quality content such as live stream interruptions, mainly including detection methods based on vision, audio, and multimodal fusion. However, these existing technologies generally face the following common problems: most solutions can only determine whether a live stream involves live stream interruptions, but cannot accurately identify the start and end times of the interruption, which limits their application in scenarios such as real-time reminders and penalties.
[0004] Therefore, there is an urgent need for an efficient, stable method for detecting live streams with high temporal resolution, which can accurately locate live stream time periods while ensuring recognition accuracy and adapting to the actual operational needs of live streaming platforms. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a live streaming behavior detection method, device, electronic device and readable storage medium, which can accurately locate the live streaming time period while ensuring the accuracy of live streaming behavior detection, and adapt to the actual operation needs of live streaming platforms.
[0006] To achieve the above objectives, the technical solutions adopted in the embodiments of the present invention are as follows: In a first aspect, the present invention provides a live streaming behavior detection method, the method comprising: responding to a live streaming room start operation, acquiring a continuous video segment within a preset time granularity from the start time in the live stream as a detection window; performing hang-up behavior identification based on the video frames within the detection window; if hang-up behavior exists, extending the detection window using the time granularity, and re-performing hang-up behavior identification on the extended detection window until no hang-up behavior exists or the time length of the detection window reaches the maximum allowable duration; and determining the hang-up duration segment based on the last detection window in which hang-up behavior was determined to exist.
[0007] Secondly, the present invention provides a live streaming behavior detection device, comprising: an acquisition module, configured to respond to a live streaming operation and acquire continuous video segments within a preset time granularity from the start time of the live stream as a detection window; an identification module, configured to perform hang-up behavior identification based on the video frames within the detection window; an extension module, configured to extend the detection window using the time granularity if the identification module identifies hang-up behavior, and return to the identification module based on the extended detection window, until no hang-up behavior exists or the time length of the detection window reaches the maximum allowable duration; and a detection module, configured to determine the hang-up duration segment based on the last detection window that was determined to have hang-up behavior.
[0008] Thirdly, the present invention provides an electronic device including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the live streaming behavior detection method described in any of the foregoing embodiments.
[0009] Fourthly, the present invention provides a readable storage medium having machine-executable instructions stored thereon, which, when executed by a processor, implement the live streaming behavior detection method as described in any of the foregoing embodiments.
[0010] The live streaming behavior detection method, apparatus, electronic device, and readable storage medium provided in this invention differ from existing technologies that simply and roughly determine whether there is any live streaming activity during the entire live stream. This invention, from the very beginning of the live stream, acquires continuous video segments within a preset time granularity from the start of the stream as a detection window. Then, it determines whether live streaming activity is suspended based on the video frames within the detection window. If live streaming activity is detected, the entire live stream is not immediately deemed a violation; instead, the detection window is extended by time granularity and the entire stream is re-evaluated. This process is repeated until a later, extended detection window is no longer considered a live stream activity, or the maximum allowed duration is reached. Then, based on the last detection window where live streaming activity was detected, the duration of the live stream activity is determined. This method not only determines whether live streaming activity is present but also accurately pinpoints its start and end times, overcoming the shortcomings of existing technologies and meeting the refined operational needs of live streaming platforms for real-time reminders and penalties for live streams with live streaming activity.
[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A schematic flowchart illustrating the live streaming behavior detection method provided in an embodiment of the present invention; Figure 2 This diagram illustrates the process of identifying the hanging behavior of a single detection window in an embodiment of the present invention. Figure 3 This is a functional block diagram of the live streaming behavior detection device provided in an embodiment of the present invention; Figure 4 A structural block diagram of an electronic device provided in an embodiment of the present invention is shown. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0015] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0016] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0017] Some violations may occur on live streaming platforms, such as "hanging up" or "spreading content." This means the streamer starts a live stream but isn't actually in front of the camera; instead, a static image (such as a landscape photo, a game recording, or even a black screen) is continuously played. This behavior can mislead the platform into believing the streamer is broadcasting normally, leading to continued traffic and rewards. For the platform, this wastes promotional resources; for viewers, it results in a poor experience and user churn. Therefore, live streaming platforms need a detection method that can automatically identify streams exhibiting this behavior, allowing for precise penalties, warnings, or removal.
[0018] However, most current live streaming behavior detection methods can only determine whether a live stream involves a suspended broadcast, and cannot accurately identify the start and end times of the suspended broadcast, which limits their application in scenarios such as real-time reminders and penalties; secondly, they have high inference latency and high deployment costs, making them unsuitable for large-scale real-time detection scenarios.
[0019] Therefore, in order to solve the above problems, this invention presents a fast and accurate live streaming behavior detection method that can identify live streaming behavior, accurately mark the specific time period in which the live streaming behavior occurs, and at the same time, does not consume too many resources.
[0020] Please see Figure 1 , Figure 1 This is a schematic flowchart illustrating the live streaming behavior detection method provided in an embodiment of the present invention. The execution subject of this method can be an electronic device (such as a server), and the method can include steps S101 to S104, as described below: S101: Respond to the live broadcast start operation and obtain a continuous video segment within a preset time granularity from the start time of the live broadcast as a detection window; S102: Perform playback behavior recognition based on video frames within the detection window; S103: If there is a hanging behavior, the detection window is extended by time granularity, and the hanging behavior identification is re-executed on the extended detection window until there is no hanging behavior or the time length of the detection window reaches the maximum allowed time. S104: Determine the duration of the hanging behavior based on the detection window where the last hanging behavior was detected.
[0021] Unlike existing technologies that simply and roughly determine whether there is any live streaming activity being suspended, this invention acquires continuous video segments within a preset time granularity from the start of the live stream as a detection window. It then determines whether live streaming is suspended based on the video frames within the detection window. If suspension is detected, the entire live stream is not immediately deemed a violation; instead, the detection window is extended by time granularity before a new overall assessment. This process is repeated until an extended detection window is no longer considered a suspended stream, or the maximum allowed duration is reached. Then, based on the last detection window that indicates suspension, the duration of the suspension is determined. This method not only determines whether suspension occurs but also accurately pinpoints its start and end times, overcoming the shortcomings of existing technologies and meeting the refined operational needs of live streaming platforms for real-time alerts and penalties for suspended live streams.
[0022] Next, embodiments of the present invention will be described in conjunction with relevant accompanying drawings. Figure 1 The process of detecting live streaming behavior is explained in detail.
[0023] In one embodiment of the present invention, before executing step S201, a relevant technician can pre-set a time granularity, such as 1 minute, for subsequent dynamic adjustment of the detection time range to identify live streaming suspension behavior. In step S201, when the device detects that a streamer has started live streaming, that is, after responding to the live streaming start operation, it will immediately extract a continuous video content from the start time as the initial analysis object. This video is called the detection window, and its duration is determined by the preset time granularity. For example, this time granularity can be 1 minute, then the device will first extract the live stream video frames from minute 0 to minute 1 for analysis to determine whether there is suspension behavior, see step S102.
[0024] In one embodiment of the present invention, in step S102, the device performs playback behavior recognition on the video frames within the detection window. It is understood that the detection window may contain many video frames. The device can perform recognition on all video frames or select a subset of video frames. When selecting a subset of video frames for recognition, the video frames to be used for playback behavior recognition can be obtained by equally spaced frame extraction. The specific method used can be determined based on the actual performance of the device or the actual needs of the user, and is not limited here.
[0025] Furthermore, this embodiment of the invention analyzes the image features of video frames within the detection window and identifies the hooking behavior based on the similarity between these features. Specifically, step S102 can be implemented as follows, see steps a1 to a4, as explained below: Step a1: Obtain the image features of each video frame and form an image feature sequence; As we can understand, each video frame is an image representing the content of a specific moment. Next, the device extracts an image feature from each video frame. This image feature is a high-dimensional vector representation (e.g., 256-dimensional), which can be seen as a digital representation of the frame's content, reflecting its main visual information. The image features corresponding to all video frames are arranged in chronological order to form an image feature sequence.
[0026] For example, assuming that 5 video frames within the detection window are selected for analysis, the image feature sequence formed by extracting the 256-dimensional vectors corresponding to each of these 5 video frames will contain 5 256-dimensional vectors.
[0027] Step a2: Using the first image feature as a benchmark, calculate the similarity between subsequent image features and the benchmark; In this embodiment of the invention, the device uses the image features corresponding to the first video frame as a benchmark, that is, it takes this initial frame as a reference standard. Then, it calculates the similarity between the image features of each subsequent video frame and this benchmark. The similarity can be represented, but is not limited to, using Hamming distance. Assuming there are M video frames in the image feature sequence, the similarity between the first image feature and the second to Mth image features is calculated, thus obtaining M-1 similarity scores.
[0028] The similarity score used above measures how similar two images are in terms of visual content. If two images are almost identical, such as repeatedly displaying the same picture, then their similarity score will be high; conversely, if there are obvious changes in the images, such as the anchor starting to move or the camera switching, the similarity score will decrease.
[0029] Step a3: Based on the comparison results of similarity and preset similarity threshold, determine whether the subsequent image features are similar to the baseline, and generate a sequence matching vector; wherein, the sequence matching vector includes elements representing similarity and dissimilarity; Understandably, the obtained similarity score is the similarity between the first image and each subsequent image. The device compares each similarity score with a pre-defined "preset similarity threshold" (denoted by 'a'). If the similarity between a feature of a subsequent image and the baseline is greater than or equal to 'a', it is considered "similar"; otherwise, it is judged as "dissimilar". Each such judgment result is recorded, forming a list of "similar" and "dissimilar" labels, called a "sequence matching vector". In this vector, each element indicates whether a video frame is sufficiently similar to the opening scene.
[0030] Optionally, the device can use numbers, letters, or other symbols to distinguish between "similar" and "dissimilar". For example, "1" can be used to represent similarity, and "0" can be used to represent dissimilarity. The specific marking method can be flexibly set by relevant technical personnel and is not limited here.
[0031] Step a4: When the proportion of similar elements exceeds the preset embedding threshold, embedding behavior is determined to exist; otherwise, embedding behavior does not exist.
[0032] Finally, the device calculates the proportion of "similar" elements in the entire sequence matching vector. If this proportion exceeds a pre-set "preset hanging threshold" (denoted by b), meaning that the vast majority of the images are considered very similar to the original images without substantial changes, then it can be reasonably determined that hanging behavior exists within the current detection window. Conversely, if a significant portion of the images show differences, causing the proportion of similar elements to fall below the threshold, then the live content is considered to have changed and is not considered hanging.
[0033] For example, suppose the detection window contains 60 video frames. After processing, it is found that 55 of these frames have image features highly similar to the first frame, with only 5 being slightly different. If the set hooking threshold is 90%, meaning that more than 90% of the frames must be similar to be considered hooking, then the similarity ratio is approximately 91.7%, exceeding the threshold, and the device will determine that hooking has occurred. However, if there are only 50 similar frames, accounting for approximately 83.3%, which is below the threshold, then hooking is not determined to occur.
[0034] As can be seen, the above-mentioned live streaming behavior recognition process does not rely on understanding the specific content, nor does it require complex model reasoning. Instead, it can quickly and effectively identify abnormal states in the video that remain unchanged for a long time through image feature comparison and threshold judgment. The entire process is highly automated, consumes few resources, and is suitable for real-time operation on large-scale live streaming platforms.
[0035] In one embodiment of the present invention, to improve the efficiency and recognition accuracy of subsequent image processing, the video frames can be preprocessed before performing step a1. For example, the video frames can be converted into grayscale images and histogram equalization can be performed. This reduces computational load, adjusts brightness distribution, reduces illumination differences, and improves the stability of subsequent feature extraction. The processed video frames can also be scaled to a fixed size (e.g., 64). 64) These preprocessing operations can unify the input data format without losing key visual information, speed up processing, and enhance the robustness of algorithms in different scenarios.
[0036] In one embodiment of the present invention, in order to provide a reliable basis for subsequent identification of broadcasting behavior based on image features rather than relying on complex image semantic understanding, the present invention provides the following implementation method for step a1 above: Step 1: Extract the low-frequency coefficients corresponding to the video frames and determine the mean of the low-frequency coefficients; Understandably, during live streaming, if the streamer is actually on camera, even if the background is relatively static, there will usually be slight movements, facial expressions, or lighting fluctuations. These subtle changes will be reflected in the low-frequency components of the image. However, the images played during "hanging" (or "overlaying"), such as fixed images or looping videos, have highly consistent image content between consecutive frames, and their corresponding low-frequency characteristics remain stable. Based on this difference, the device can analyze the changes in low-frequency coefficients of video frames in the frequency domain to extract image features that reflect the static characteristics of the image.
[0037] Specifically, after the device acquires each video frame within the detection window, the low-frequency coefficients can be extracted as follows: First, the video frames are transformed in the frequency domain to obtain the coefficient matrix.
[0038] The technique used in the above process can be the Discrete Cosine Transform (DCT), which can transform image data in the spatial domain into a numerical distribution in the frequency domain. Each element in the resulting coefficient matrix represents a frequency intensity, and the matrix size is consistent with the size of the preprocessed image (e.g., 64). 64). In this coefficient matrix, the element in the upper left corner corresponds to the low-frequency region, representing the overall brightness, average color patches, and large-scale structural features of the image, while the element in the lower right corner corresponds to the high-frequency region, mainly reflecting edges, textures, and minor noise. Because the changes in the broadcast image are minimal, its low-frequency component is particularly stable.
[0039] Next, low-frequency coefficients are extracted from the low-frequency region of the coefficient matrix and the mean of the low-frequency coefficients is determined.
[0040] The aforementioned low-frequency region represents the main structural features of the image, and its size can be selected based on actual conditions. For example, the low-frequency region is a 16×16 block in the upper left corner of the coefficient matrix, containing 256 low-frequency coefficients. After extracting multiple low-frequency coefficients from this region, the device calculates the average value of these coefficients, known as the "low-frequency coefficient mean," which serves as a reference benchmark for subsequent judgments.
[0041] Step 2: If a low-frequency coefficient is greater than or equal to the mean of low-frequency coefficients, then configure the first feature identifier for the low-frequency coefficient; otherwise, configure the second feature identifier. In this embodiment of the invention, if a certain low-frequency coefficient is greater than or equal to the average low-frequency coefficient of the frame, a "first feature identifier" is configured for it; otherwise, a "second feature identifier" is configured for it. Here, the "first feature identifier" and the "second feature identifier" can be understood as two kinds of marking symbols. For example, "1" is used to represent the first feature identifier and "0" is used to represent the second feature identifier, thus forming a binary sequence composed of 0 and 1.
[0042] Step 3: Generate image features from the first feature identifier and the second feature identifier, and compose an image feature sequence from all image features.
[0043] The device uses each bit of the previously generated binary sequence to generate the image features of the current video frame. For example, applying this rule to 256 low-frequency coefficients results in a sequence of 256 0s or 1s, which is the "image feature" of the video frame. This representation greatly compresses the data size of the original image while preserving its core structural information, making similarity comparison between different frames fast and reliable.
[0044] Finally, the device arranges the image features generated from each video frame in chronological order to form an "image feature sequence." This sequence can be used for subsequent similarity comparisons with other frames or a reference frame.
[0045] In one embodiment of the present invention, steps a1 to a4 described above can not only identify whether there is a live-streaming activity in the live-streaming room, but also, by dividing the degree of similarity of the images into different levels and determining the specific type of live-streaming activity based on these levels, achieve refined classification and identification of live-streaming activity. Specifically: the similarity threshold is divided into multiple intervals; a live-streaming activity type is set for each interval; and the live-streaming activity type of the detection window is determined based on the interval into which the similarity falls.
[0046] This can be understood as follows: The system first divides the "similarity threshold" used to determine whether images are similar into multiple intervals. Each interval represents a level of similarity. For example, assuming the similarity threshold a=7, it can be divided into three intervals: [0, 3], (3, 5], and (5, 7). Next, the corresponding broadcast behavior type for each interval is set, such as "static image broadcast," "minor dynamic broadcast," and "dynamic broadcast," etc. In other words, a mapping relationship is established in advance, clearly defining which similarity interval corresponds to which type of broadcast behavior. This setting can be completed during the system initialization phase or configured by platform operators according to actual regulatory needs.
[0047] Next, after each successful identification of a live-streaming behavior, the system obtains a similarity result between each video frame within the current detection window and the first video frame. This similarity can be the proportion of "similar" elements in all matching judgments, or it can be the average similarity value calculated based on the image feature sequence. The system then determines which preset interval the similarity value falls into and accordingly identifies the live-streaming behavior type of the detection window. By hierarchically processing the similarity judgment results, multi-level classification of live-streaming behaviors can be achieved, improving the semantic expression capability and operational management flexibility of the detection system without increasing reliance on complex models.
[0048] In summary, the live streaming behavior detection method provided by this invention enhances the ability of live streaming platforms to identify low-quality content. It has achieved significant results in terms of business metrics. For example, in actual verification, the live streaming behavior detection method provided by this invention can penalize thousands of live streaming rooms with suspended broadcasting behavior daily, releasing an average of 3% of invalid homepage exposure daily, improving traffic distribution efficiency, and contributing 1.5% to the homepage's view count growth. Furthermore, compared to deep learning-based feature extraction or large-model recognition solutions, this method features fast inference speed and low cost. This detection method helps promote the compliant development of live streaming platforms, ensuring adherence to relevant laws, regulations, and community guidelines, and maintaining a positive industry image and user trust.
[0049] For a better understanding of the implementation of step S102 above, please refer to [link to relevant documentation]. Figure 2 , Figure 2 This diagram illustrates the process of identifying live streaming behavior in a single detection window according to an embodiment of the present invention. It can be seen that the above implementation method extracts image features based on image preprocessing, frequency domain transformation, and binarization encoding. The entire process requires no complex model inference, consumes low computational resources, and is particularly suitable for live streaming platform environments requiring large-scale real-time monitoring.
[0050] Based on the identification results of the hanging behavior identification process in step S102 above, the specific countermeasures for whether or not there is hanging behavior will be introduced next, see step S103.
[0051] In step S103 of this embodiment of the invention, a method of gradually extending the detection window is designed to confirm whether the hanging behavior actually exists and to accurately define its duration.
[0052] Specifically, if the device determines that there is a live stream violation within the current detection window, it will not immediately classify the entire live stream as a violation. Instead, it will extend the detection window by time granularity and then re-identify the live stream violation across this longer time period. Of course, this incremental detection will not continue indefinitely but will be capped based on actual business rules. If the extended detection window no longer meets the criteria for live stream violation or reaches the capped limit, the identification process will stop.
[0053] Based on the above concept, in step S103, if it is determined that there is no hanging behavior in the current detection window, the recognition ends. After initializing the detection window, the recognition restarts until the broadcaster ends. If hanging behavior exists, the time length of the detection window will be extended using time granularity, and the hanging behavior recognition will be re-executed on the extended detection window until there is no hanging behavior or the time length of the detection window reaches the maximum allowed duration.
[0054] For example, assuming the smallest time granularity is 1 minute, the device will first detect whether there is any hanging behavior in the first minute after the live stream starts. If the device determines that there is hanging behavior in this first minute, it will extend the detection window by time granularity, that is, from the original 0 to 1 minute to 0 to 2 minutes, and then re-perform hanging behavior identification on this longer time period. If hanging behavior is still identified in this new detection window of 0 to 2 minutes, the device will continue to extend the detection window again by the same time granularity to 0 to 3 minutes, and perform hanging behavior identification again. This process can be repeated, each time increasing the detection window by one time granularity unit, and re-evaluating whether there is hanging behavior based on the extended new window.
[0055] In one embodiment of the present invention, before each extension of the detection window time length, the image features and sequence matching vectors generated within the detection window are cached, which can avoid repeated calculations and improve the detection speed.
[0056] This can be understood as follows: each time the detection window is extended, it is equivalent to including the video frames in the next time granularity in the hanging behavior recognition process. For example, if 5 video frames are selected for recognition within the detection window of 0 to 1 minute, and the detection window is extended to 0 to 2 minutes after 1 minute, then 5 video frames are selected from 1 to 2 minutes, which, together with the 5 video frames from 0 to 1 minute, participate in the subsequent recognition process.
[0057] Next, for the detection window of 0 to 2 minutes, since the image features and sequence matching vectors of 5 video frames within 0 to 1 minute have already been cached, it is only necessary to extract the image features of 5 selected video frames within 1 to 2 minutes, and then calculate the similarity between these 5 image features and the first image feature in the cached result, and then the sequence matching vectors corresponding to the video frames within 1 to 2 minutes. The sequence matching vectors of 0 to 1 minutes and 1 to 2 minutes are combined to determine whether there is a hanging broadcast behavior in the detection window of 0 to 2 minutes.
[0058] During the repeated identification process described above, if a detection window does not exhibit any hanging behavior, for example, if the period from 0 to N minutes (where N is less than or equal to the maximum allowed duration) is identified as hanging, but from 0 to N+1 minutes onwards it no longer constitutes hanging, or if the duration of the detection window reaches the maximum allowed duration, then the detection window will not be extended.
[0059] It should be understood that the maximum allowed duration can be set according to actual business rules, such as 15 minutes. That is to say, when the detection window is gradually extended from 0 to 15 minutes, even if it is still identified as a hanging broadcast at this time, it will not be extended any further.
[0060] This detection mechanism not only effectively addresses the issue of misidentification caused by brief periods of inactivity, but also dynamically adapts to different lengths of live stream interruptions while ensuring accuracy. For example, if a live stream has an interruption lasting up to 30 minutes, theoretically the device will complete a full detection process in two intervals: 0-15 minutes and 15-30 minutes, thereby identifying two independent 15-minute interruption periods and triggering corresponding alert actions.
[0061] Based on the aforementioned dynamic detection mechanism that progressively extends the detection window with a fixed time granularity, in step S204 of this embodiment, the duration of the broadcast can be determined according to the detection window where the broadcasting behavior was last determined to exist. For example, if the broadcasting is identified from minute 0 to minute 15, but no longer constitutes a broadcast from minute 0 to minute 16, the device will return to the previous broadcasting period that was still valid, i.e., minute 0 to minute 15, as the final recognized broadcasting period. In other words, the device uses the principle of the last valid hit to determine the actual end point of the broadcast.
[0062] Based on the above-mentioned livestreaming violation identification results, the device can implement a series of livestreaming violation processing measures. For example, when a livestreaming room is identified as having experienced livestreaming violation behavior within a certain period, a livestreaming room is marked with a livestreaming violation tag and the time period of violation is noted, for relevant personnel to conduct subsequent verification and content review; or, a livestreaming violation warning is directly triggered for the livestreaming room, pushing a reminder message to the broadcaster, prompting them to improve the quality of the livestreaming content to avoid affecting the user experience; or, livestreaming rooms that repeatedly display the livestreaming violation tag can be penalized or have their traffic recommendations restricted, thereby reducing the exposure opportunities of low-quality livestreaming content and guiding broadcasters to increase livestreaming activity and interactivity. Furthermore, the platform can also incorporate livestreaming violation records into the broadcaster credit evaluation system as a reference for long-term operation and management, and in some implementations, can optionally be used to adjust livestreaming permissions or eligibility to participate in activities.
[0063] Based on and Figure 1 Following the same inventive concept, this invention also provides an implementation of a live-streaming behavior detection device 30. Please refer to [link to relevant documentation]. Figure 3 , Figure 3 This is a functional block diagram of a live streaming behavior detection device provided in an embodiment of the present invention. The live streaming behavior detection device 30 includes: an acquisition module 301, an identification module 302, an extension module 303, and a detection module 304.
[0064] The acquisition module 301 is used to respond to the live broadcast start operation and acquire continuous video segments within a preset time granularity from the start time of the live broadcast as a detection window. The recognition module 302 is used to perform playback behavior recognition based on video frames within the detection window; The extension module 303 is used to extend the time length of the detection window with time granularity if the identification module identifies that there is a hooking behavior, and return to the identification module 302 based on the extended detection window until there is no hooking behavior or the time length of the detection window reaches the maximum allowed time. The detection module 304 is used to determine the duration of the hanging behavior based on the detection window that was last determined to have hanging behavior.
[0065] It is understandable that the acquisition module 301, the recognition module 302, the extension module 303, and the detection module 304 can perform in a coordinated manner. Figure 1 Each step in the process is to achieve the corresponding technical effect.
[0066] It should be noted that the live streaming behavior detection device 30 provided in this embodiment of the invention can be specific hardware on a device or software or firmware installed on the device. The implementation principle and technical effects of the device provided in this embodiment of the invention are the same as those in the foregoing method embodiments. For the sake of brevity, any parts not mentioned in the device embodiments can be referred to the corresponding content in the foregoing method embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can all be referred to the corresponding processes in the above method embodiments, and will not be repeated here.
[0067] Optionally, the above modules can be stored in the form of software or firmware. Figure 4 The memory shown is either stored in or embedded in the operating system (OS) of the electronic device 40, and can be used by... Figure 4 The processor executes the commands. Meanwhile, the data and program code required to execute these modules can be stored in memory.
[0068] Please see Figure 4 , Figure 4 The diagram illustrates the structure of an electronic device according to an embodiment of the present invention, including a memory 401, a processor 402, and a communication interface 403. The memory 401, processor 402, and communication interface 403 are electrically connected to each other directly or indirectly to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0069] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized into address buses, data buses, control buses, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0070] In this embodiment of the invention, the processor 402 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in this embodiment of the invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in this embodiment of the invention can be directly manifested as execution by the hardware processor, or execution by a combination of hardware and software modules within the processor. The software modules may reside in the memory 401, and the processor 402 reads the program instructions from the memory 401 and, in conjunction with its hardware, completes the steps of the aforementioned methods.
[0071] In this embodiment of the invention, the memory 401 can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as RAM. The memory can also be any other medium capable of carrying or storing desired executable program code having an instruction or data structure form and accessible by a computer, but is not limited thereto. The memory in this embodiment of the invention can also be a circuit or any other device capable of implementing a storage function for storing instructions and / or data.
[0072] The memory 401 can be used to store software programs and modules, such as the instructions / modules of the live streaming behavior detection device 30 provided in this embodiment of the invention. These can be stored in the memory 401 in the form of software or firmware, or embedded in the operating system (OS) of the electronic device 40. The processor 402 executes various functional applications and data processing by executing the software programs and modules stored in the memory 401. The communication interface 403 can be used to communicate with other node devices for signaling or data.
[0073] Understandable. Figure 4 The structure shown is for illustrative purposes only; the electronic device 40 may also include components that are more advanced than those shown. Figure 4 The more or fewer components shown, or having the same Figure 4 The different configurations shown. Figure 4 The components shown can be implemented using hardware, software, or a combination thereof.
[0074] Based on the above embodiments, the present invention also provides a readable storage medium storing a computer program. When the computer program is executed by a computer, it causes the computer to execute the live streaming behavior detection method provided in the above embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0075] Based on the above embodiments, the present invention also provides a program product, which includes a computer program. The processor can execute the computer program to implement the live streaming behavior detection method provided in the embodiments of the present invention. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0076] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0077] Furthermore, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the objectives of the embodiments of the present invention, depending on actual needs.
[0078] Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0079] It should be noted that if the function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0080] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for detecting live streaming behavior, characterized in that, The method includes: In response to the live broadcast start operation, the system acquires continuous video segments within a preset time granularity from the start time of the live broadcast as a detection window. Perform playback behavior recognition based on the video frames within the detection window; If a rebroadcasting behavior exists, the detection window is extended using the time granularity, and the rebroadcasting behavior identification is re-executed on the extended detection window until no rebroadcasting behavior exists or the time length of the detection window reaches the maximum allowed duration. The duration of the broadcast is determined based on the last detection window that indicated the presence of broadcasting behavior.
2. The live streaming behavior detection method according to claim 1, characterized in that, Based on the video frames within the detection window, perform ad-hoc behavior recognition, including: Obtain the image features of each video frame and form an image feature sequence; Using the first image feature as a benchmark, calculate the similarity between subsequent image features and the benchmark; Based on the comparison result between the similarity and a preset similarity threshold, it is determined whether the subsequent image features are similar to the benchmark, and a sequence matching vector is generated; wherein, the sequence matching vector includes elements representing similarity and dissimilarity; When the proportion of similar elements exceeds the preset threshold for embedding, embedding behavior is determined to exist; otherwise, embedding behavior does not exist.
3. The live streaming behavior detection method according to claim 2, characterized in that, Obtain the image features of each video frame and assemble them into an image feature sequence, including: Extract the low-frequency coefficients corresponding to the video frames and determine the mean of the low-frequency coefficients; If a certain low-frequency coefficient is greater than or equal to the average value of the low-frequency coefficients, then a first feature identifier is configured for the low-frequency coefficient; otherwise, a second feature identifier is configured. The image features are generated from the first feature identifier and the second feature identifier, and the image feature sequence is composed of all the image features.
4. The live streaming behavior detection method according to claim 3, characterized in that, Extracting the low-frequency coefficients corresponding to the video frames and determining the mean of the low-frequency coefficients, including: The video frames are frequency domain transformed to obtain a coefficient matrix; where each element of the coefficient matrix represents a frequency intensity. Extract the low-frequency coefficients from the low-frequency region of the coefficient matrix and determine the mean of the low-frequency coefficients.
5. The live streaming behavior detection method according to claim 2, characterized in that, Before obtaining the image features of each video frame and assembling an image feature sequence, the method further includes: The video frames are converted into grayscale images and then subjected to histogram equalization. The processed video frames are scaled to a fixed size.
6. The live streaming behavior detection method according to claim 2, characterized in that, The method further includes: Before each extension of the detection window duration, the image features and sequence matching vector generated within the detection window are cached.
7. The live streaming behavior detection method according to any one of claims 2 to 6, characterized in that, The method further includes: The similarity threshold is divided into multiple intervals; Set the corresponding broadcast behavior type for each interval; The type of broadcast behavior of the detection window is determined based on the interval in which the similarity falls.
8. A live streaming behavior detection device, characterized in that, include: The acquisition module is used to respond to the live broadcast start operation and acquire continuous video segments within a preset time granularity from the start time of the live broadcast as a detection window. The identification module is used to perform playback behavior identification based on the video frames within the detection window; The extension module is used to extend the detection window with the time granularity if the identification module identifies that there is a hanging broadcast behavior, and return to the identification module based on the extended detection window until there is no hanging broadcast behavior or the time length of the detection window reaches the maximum allowable duration. The detection module is used to determine the duration of the hanging behavior based on the detection window that was last determined to have hanging behavior.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor to implement the live streaming behavior detection method according to any one of claims 1-7.
10. A readable storage medium having machine-executable instructions stored thereon, characterized in that, When the machine-executable instructions are executed by the processor, the live streaming behavior detection method as described in any one of claims 1-7 is implemented.