Anchor live broadcast behavior auditing method and device, electronic equipment and readable storage medium

By decoding video frame data and analyzing various algorithm models of the live stream, the problems of automation and accuracy in monitoring the behavior of live streamers have been solved, and efficient live stream behavior review has been achieved.

CN115393929BActive Publication Date: 2026-02-17INNOVATION QIZHI (GUANGZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210962478.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-11
Publication Date
2026-02-17
Estimated Expiration
2042-08-11

AI Technical Summary

Technical Problem

Live streaming platforms struggle to effectively monitor the live streaming behavior of a large number of streamers. Existing manual monitoring methods are labor-intensive and lack objectivity and comprehensiveness.

Method used

By decoding the live stream into video frame data and using various algorithm models to analyze the frame data, including violation behavior analysis, facial expression analysis, and body behavior analysis, the system determines the anchor's behavior data and judges whether there are any abnormalities based on preset abnormal behavior rules, providing real-time alerts and scores.

Benefits of technology

It enables automated, accurate, and comprehensive review of the streamer's live broadcast behavior, reduces manual intervention, improves the efficiency and accuracy of supervision, and ensures the quality of live broadcasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393929B_ABST
    Figure CN115393929B_ABST
Patent Text Reader

Abstract

The application provides a host live broadcast behavior auditing method and device, electronic equipment and a computer readable storage medium. The host live broadcast behavior auditing method provided by the application comprises: decoding a live stream obtained into video frame data; inputting the video frame data into multiple algorithm models for analysis to obtain behavior data of a host corresponding to the live stream; wherein the algorithm model comprises a rule violation behavior analysis algorithm model; and determining whether the live broadcast behavior of the host is abnormal according to the behavior data and a preset abnormal behavior rule. In the above method, the entire verification process is completed by the host live broadcast behavior auditing system, without human intervention, reducing the workload and labor cost of staff. At the same time, all behavior data is audited according to a unified standard, and the problem of subjectivity and incompleteness caused by human monitoring does not occur, improving the accuracy and comprehensiveness of the audit.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multimedia, in particular to a host live broadcast behavior auditing method and device, electronic equipment and readable storage medium. BACKGROUND

[0002] With the rapid development of network technology, many platforms have launched live broadcast, and the current live broadcast threshold is low, anyone can become a host, so the live broadcast industry has also developed rapidly. For live broadcast platforms, due to the low threshold of hosts, the quality of hosts is uneven, and even some hosts will make some irregular behaviors in order to attract traffic. Therefore, in the face of a large number of hosts, the supervision of live broadcast behaviors and live broadcast quality has become a very important work for live broadcast platforms.

[0003] The current supervision method often relies on manual work, that is, the so-called "network manager" randomly enters some live broadcast rooms to find problems, which requires many people and brings a lot of workload, and there are problems such as subjectivity and inability to monitor the whole process. SUMMARY

[0004] Therefore, the purpose of the embodiments of the present application is to provide a host live broadcast behavior auditing method and device, electronic equipment and readable storage medium. The live broadcast behavior can be automatically audited according to a unified standard.

[0005] In the first aspect, the embodiments of the present application provide a host live broadcast behavior auditing method, including: decoding a obtained live stream into video frame data; inputting the video frame data into a plurality of algorithm models for analysis to obtain behavior data of a host corresponding to the live stream; wherein the algorithm model includes a irregular behavior analysis algorithm model; and determining whether the live broadcast behavior of the host is abnormal according to the behavior data and a preset abnormal behavior rule.

[0006] In the above implementation process, the frame data obtained by parsing the live stream is input into various algorithm models for analysis to obtain a plurality of behavior data of the live stream, and whether the live broadcast behavior of the host corresponding to the live stream is abnormal is further determined based on the behavior data. The entire verification process is completed by the host live broadcast behavior auditing system without human intervention, reducing the workload of staff. At the same time, all behavior data is audited according to a unified standard, and the problems of subjectivity and incompleteness caused by human monitoring are avoided, improving the accuracy and comprehensiveness of the audit.

[0007] In one embodiment, the behavior data includes human limb spatial position data and facial feature data, and the abnormality includes a violation behavior; the inputting the video frame data into multiple algorithm models for analysis to obtain the behavior data of the anchor corresponding to the live stream includes: inputting the video frame data into the violation behavior analysis algorithm model; identifying the frame data by using the violation behavior analysis algorithm model to determine the human limb spatial position data and the facial feature data corresponding to the frame data; wherein the violation behavior analysis algorithm model includes a violation rule library, and the violation rule library includes behavior types and violation levels corresponding to multiple behavior data; the determining whether the live behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule includes: analyzing the human limb spatial position data and the facial feature change data based on the violation rule library to determine whether the anchor corresponding to the live stream has a violation behavior; and if the anchor has a violation behavior, determining a violation level of the anchor according to the behavior data and the behavior types and violation levels corresponding to the multiple behavior data in the violation rule library.

[0008] In the above implementation process, the frame data is input into the violation behavior analysis algorithm model to determine the limb behavior and facial expression corresponding to the frame data, and then the behavior data of the anchor is determined. The behavior data and the violation behavior are matched to determine whether the anchor has a violation behavior and a violation level. This implementation realizes multi-aspect analysis of the video frame data, can accurately determine the behavior corresponding to the video frame data, and accurately audits the violation behavior and the violation level in a timely manner, prevents the anchor from violating the live broadcast, and improves the quality of the live broadcast.

[0009] In one embodiment, the algorithm model further includes a facial expression algorithm model, and the behavior data includes emotion data; the inputting the video frame data into multiple algorithm models for analysis to obtain the behavior data of the anchor corresponding to the live stream includes: inputting the video frame data into the facial expression algorithm model; performing face recognition on the frame data by using the facial expression algorithm model to determine face data corresponding to the frame data; extracting features from the face data, classifying the extracted features, and determining emotion data of the anchor corresponding to the frame data.

[0010] In the above implementation process, the facial expression algorithm model is used to identify the face data in the frame data and extract features from the face data to extract feature data in the face, classify the features, determine the emotion range to which each feature belongs, and then determine emotion data of the anchor corresponding to the frame data according to the emotion ranges of multiple feature data. Since the emotion data is determined based on multiple features, the accuracy of the emotion data can be ensured, and the accuracy of the audit on the live stream is improved.

[0011] In one embodiment, the algorithm model further comprises a limb behavior analysis model, and the behavior data comprises limb data; the inputting the video frame data into multiple algorithm models for analysis to obtain the behavior data of the anchor corresponding to the live stream comprises: inputting the video frame data into the limb behavior analysis model; detecting the limb in the frame data by the limb behavior analysis model to determine the limb data corresponding to the frame data; performing feature extraction on the limb data, and classifying the extracted features to determine the action data of the anchor corresponding to the frame data.

[0012] In the above implementation process, the limb data in the frame data is identified by the limb behavior analysis model, and feature extraction is performed on the limb data to extract feature data in the limb, and the features are classified to determine the action range to which each feature belongs, and then the action data of the anchor corresponding to the frame data is determined according to the action ranges of the multiple feature data. Since the action data is determined based on multiple features, the accuracy of the behavior data can be ensured, and the accuracy of the review is improved when the live stream is reviewed.

[0013] In one embodiment, the anomaly comprises an alarm; and the determining whether the live behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule comprises: inputting the emotion data and the action data into an abnormal alarm model; comparing the emotion data and the action data with alarm behaviors in an alarm library by the abnormal alarm model to determine whether the anchor has an alarm behavior; wherein the abnormal alarm model comprises an alarm library, and the alarm library comprises alarm behaviors and corresponding behavior data.

[0014] In the above implementation process, the emotion data and the action data are input into the abnormal alarm model to match the emotion data and the action data of the two types of data with the behaviors in the alarm library by the abnormal alarm model, which can determine the real live state of the anchor by comprehensively considering the emotion and the action of the anchor, more accurately determine the behavior of the anchor, make the obtained alarm behavior more accurate, prevent false positives, and improve the accuracy of the alarm.

[0015] In one embodiment, after the video frame data is input into multiple algorithm models for analysis to obtain the behavior data of the anchor corresponding to the live stream, the method further comprises: scoring the behavior data according to a set scoring rule to determine a live score of the anchor corresponding to the frame data; wherein the behavior data and the set scoring rule are determined according to the live field of the anchor.

[0016] In the implementation process, the live broadcast of the anchor is scored according to the behavior data of the anchor and a scoring rule, so as to determine the live broadcast quality of the anchor during live broadcast, facilitate the management of the platform on each anchor, and enable certain constraints on the live broadcast behavior of the anchor, while realizing unified management of the anchor and improving the quality of live broadcast.

[0017] In one embodiment, the behavior data includes behavior duration and behavior frequency; and the scoring of the behavior data according to the behavior data and the set scoring rule to determine the live broadcast score of the anchor corresponding to the frame data includes: determining a duration proportion of the behavior duration in the live broadcast process according to the behavior duration; determining a frequency proportion of the behavior frequency in the live broadcast process according to the behavior frequency; and determining the live broadcast score of the anchor according to the duration proportion and the frequency proportion.

[0018] In the implementation process, the live broadcast score of the anchor in the live broadcast process is calculated according to the proportion of the behavior data and the behavior frequency after removing the score of the alarm. The real live broadcast situation of the anchor in the live broadcast can be more comprehensively reflected, so as to improve the accuracy of determining the live broadcast quality of the anchor.

[0019] In a second aspect, the embodiments of the present application also provide an anchor live broadcast behavior auditing device, which includes: a decoding module configured to decode the obtained live stream into video frame data; an analysis module configured to input the video frame data into a plurality of algorithm models for analysis to obtain behavior data of an anchor corresponding to the live stream; wherein the algorithm models include a rule violation behavior analysis algorithm model; and a determination module configured to determine whether the live broadcast behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule.

[0020] In a third aspect, the embodiments of the present application also provide an electronic device, which includes a processor and a memory, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the machine readable instructions are executed by the processor to perform the steps of the method of the first aspect or any possible implementation manner of the first aspect.

[0021] In a fourth aspect, the embodiments of the present application also provide a computer readable storage medium, which stores a computer program, when the computer program is run by a processor, the steps of the anchor live broadcast behavior auditing method of the first aspect or any possible implementation manner of the first aspect are executed.

[0022] In order to make the above purposes, features and advantages of the present application more obvious and easy to understand, the following embodiments are described in detail below, and the accompanying drawings are described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as a limitation to the scope. Other related drawings can also be obtained by those of ordinary skill in the art without creative labor.

[0024] Figure 1 The anchor live broadcast behavior auditing system schematic diagram provided by the embodiments of the present application;

[0025] Figure 2 The block schematic diagram of the video analysis box provided by the embodiments of the present application;

[0026] Figure 3 The flowchart of the anchor live broadcast behavior auditing method provided by the embodiments of the present application;

[0027] Figure 4 The functional module schematic diagram of the anchor live broadcast behavior auditing device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.

[0029] It should be noted that: similar reference numbers and letters represent similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. Meanwhile, in the description of the present application, the terms "first", "second", etc. are only used for differentiation, and cannot be understood as indicating or implying relative importance.

[0030] With the advent of the 4G and 5G era, short videos have changed from a new thing to a new media way universally recognized and loved by the public, thus also achieving the development of various short video platforms. The live broadcast of goods that was once popular on shopping platforms is gradually moving to short video social platforms. The live broadcast of goods based on short video platforms makes the threshold of this industry very low, and almost anyone can join in. Therefore, many live broadcast of goods companies have emerged, and these companies often hire many live broadcast anchors to perform live broadcast of goods in a batch and large-scale manner in the form of template words and industrial words. For such live broadcast companies, for such a large number of anchors, some anchors may make some irregular behaviors for the purpose of seeking traffic, and other anchors may be in a situation of playing the fool to do their work, thus causing great problems in management and supervision.

[0031] Therefore, the present application provides a method for auditing the behavior of a live streamer. One or more video analysis boxes are set up. After accessing the video stream data of a live stream, the video stream is analyzed in real time using various algorithm models to find violations, laziness and other problems, and real-time warning information is issued. At the same time, the number of times of looking at a mobile phone, the number of times of being out of the screen, the time, the number of times of lowering the head, the number of times of smiling, the number of times of body movement, the number of times of interactive behavior and the like in the live stream are counted. Then, all the statistical indicators are comprehensively scored based on the model to provide quantifiable index references for the management personnel. The automatic auditing of the behavior of a live stream is realized, the work burden of the staff is reduced, and the accuracy and efficiency of the auditing are improved.

[0032] To facilitate the understanding of the present embodiment, first, the operating environment of the method for auditing the behavior of a live streamer disclosed in the present application is introduced in detail.

[0033] As shown in Figure 1 , it is a schematic diagram of a system for auditing the behavior of a live streamer provided by the present application.

[0034] The system comprises a live stream device end 120, a video analysis box 110, a cloud end subsystem 130 and a background management subsystem 140. The video analysis box 110 is in communication connection with one or more live stream device ends 120 for data communication or interaction. The video analysis box 110 is also in communication connection with the cloud end subsystem 130 and the background management subsystem 140 for data communication or interaction.

[0035] Optionally, the video analysis box 110 can be one or more. If the video analysis box 110 is one, the video analysis box 110 is connected with multiple live stream device ends 120. If the video analysis box 110 is multiple, multiple live stream device ends 120 can be connected with each video analysis box 110 in a certain proportion.

[0036] In some embodiments, the cloud end subsystem 130 can also be in communication connection with the background management subsystem 140 for data communication or interaction.

[0037] The live stream device end 120 here can be a personal computer (PC), a tablet computer, a smart phone, a personal digital assistant (PDA) and the like. The live stream device end 120 is used to acquire video pictures in real time through a camera, push the stream to a live streaming platform to form a live stream, and copy the live stream to the video analysis box 110. As shown in Figure 2As shown, it is a block schematic diagram of the video analysis box. The video analysis box 100 can include a memory 111, a storage controller 112, a processor 113, a peripheral interface 114, and an input / output unit 115. Those skilled in the art can understand that, Figure 2 The structure shown is only schematic, which does not limit the structure of the video analysis box 100. For example, the video analysis box 100 can also include more or less components than those shown in the figure, or have a different configuration from that shown in the figure. Figure 2 The video analysis box 100 can also include more or less components than those shown in the figure, or have a different configuration from that shown in the figure. Figure 2 The video analysis box 100 can also include more or less components than those shown in the figure, or have a different configuration from that shown in the figure.

[0038] The above-mentioned memory 111, storage controller 112, processor 113, peripheral interface 114 and input / output unit 115 are directly or indirectly electrically connected to each other to realize data transmission or interaction. For example, these elements can be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 113 is used to execute the executable modules stored in the memory.

[0039] The memory 111 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) and the like. The memory 111 is used to store programs, and the processor 113 executes the programs after receiving execution instructions. The method executed by the video analysis box 100 defined by the processes disclosed in any embodiment of the present application can be applied to the processor 113 or implemented by the processor 113.

[0040] The processor 113 can be an integrated circuit chip having a processing capability of signals. The processor 113 can be a general processor, including a central processing unit (CPU), a network processor (NP), etc. The processor 113 can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The processor 113 can implement or execute the disclosed methods, steps and logic block diagrams in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor.

[0041] The peripheral interface 114 is coupled to the processor 113 and the memory 111 to couple various input / output devices. In some embodiments, the peripheral interface 114, the processor 113 and the memory controller 112 can be implemented in a single chip. In other embodiments, they can be implemented by independent chips respectively.

[0042] The input / output unit 115 is used to provide input data to the user. The input / output unit 115 can be, but is not limited to, a mouse, a keyboard, etc.

[0043] The video analysis box 110 decodes the live stream into video frame data after obtaining the live stream. The video analysis box 110 analyzes the video frame data with an algorithm model to obtain behavior data of the anchor corresponding to the live stream, and sends the behavior data to the cloud subsystem 130 and the background management subsystem 140. The video analysis box 110 can also determine whether the live behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule, and score the behavior data according to a set scoring rule to determine a live score of the anchor corresponding to the frame data, and send the result to the cloud subsystem 130 and the background management subsystem 140.

[0044] The cloud subsystem 130 is used to obtain the behavior data of the anchor, and aggregate, store and push the alarm signal of the behavior data.

[0045] In some embodiments, the cloud subsystem 130 can also determine whether the live streaming behavior of the host is abnormal according to the behavior data and a preset abnormal behavior rule after obtaining the behavior data of the host, score the behavior data according to a preset scoring rule, determine the live streaming score of the host corresponding to the frame data, and send the abnormal result and the live streaming score to the background management subsystem 140.

[0046] The background management subsystem 140 here is used to display and manage the pages of the live streaming room real-time alarm, statistics, violation analysis and scoring data of the host corresponding to the live streaming.

[0047] In some embodiments, the background management subsystem 140 can also determine whether the live streaming behavior of the host is abnormal according to the behavior data and a preset abnormal behavior rule after obtaining the behavior data of the host, score the behavior data according to a preset scoring rule, determine the live streaming score of the host corresponding to the frame data, and send the abnormal result and the live streaming score to the cloud subsystem 130.

[0048] Please refer to Figure 3 is a flowchart of the host live streaming behavior auditing method provided by the embodiments of the present application. The specific process shown in Figure 3 will be described in detail below.

[0049] Step 201, the obtained live streaming is decoded into video frame data.

[0050] Optionally, the live streaming can be decoded by using h264, h265, etc.

[0051] The live streaming here is the video picture of each host live streaming pushed to the live streaming platform by each live streaming device end.

[0052] Since the data amount of the live streaming obtained by the live streaming is large, it is not conducive to the transmission of the live streaming. In order to ensure the transmission efficiency, the live streaming is compressed by using the encoding mode. When the live streaming is analyzed, the complete original live streaming data is needed. Therefore, the compressed live streaming is decoded by using the decoding mode to obtain the frame data, so as to analyze the real and complete live streaming, and make the analysis result more accurate.

[0053] Step 202, the video frame data is input into multiple algorithm models for analysis to obtain the behavior data of the host corresponding to the live streaming.

[0054] The algorithm model includes a violation behavior analysis algorithm model.

[0055] Optionally, the algorithm model can also include a facial expression algorithm model, a body behavior analysis model, a voice algorithm model, a clothing algorithm model, an environment algorithm model, etc.

[0056] The behavior data obtained here can include expression data, action data, voice data, clothing data, etc. The expression data can include smiling, surprise, crying, etc. The action data can include waving, clapping, turning around, shaking head, running, jumping, etc. The voice data can include civilized language and uncivilized language, etc. The clothing data can include normal clothing and abnormal clothing, etc. It can be understood that the specific division rules of civilized language and uncivilized language, normal clothing and abnormal clothing here can be determined according to the rules of each platform or the related prior art or the relevant provisions, and the present application does not make specific limitations.

[0057] In step 203, it is determined whether the live streaming behavior of the host exists abnormity according to the behavior data and the preset abnormal behavior rule.

[0058] The preset abnormal behavior rule here is the behavior data belonging to abnormal behavior in different live streaming scenes according to the live streaming scene setting.

[0059] The abnormity described above can include different types of abnormity such as violation and warning.

[0060] In the implementation process described above, the frame data obtained by parsing the live streaming is input into various algorithm models for analysis to obtain multiple behavior data of the live streaming, and whether the live streaming behavior of the host corresponding to the live streaming exists abnormal behavior is further determined based on the behavior data. The whole verification process is completed by the host live streaming behavior auditing system, without human participation, reducing the workload of staff. At the same time, all behavior data is audited according to unified standards, which will not cause the problems of subjectivity and incompleteness caused by human monitoring, improving the accuracy and comprehensiveness of the audit.

[0061] In one possible implementation, step 202 includes: inputting the video frame data into a violation behavior analysis algorithm model; identifying the frame data by the violation behavior analysis algorithm model to determine human limb spatial position data and face feature data corresponding to the frame data.

[0062] Among them, the violation behavior analysis algorithm model includes a violation rule library, and the violation rule library includes behavior types and violation levels corresponding to multiple behavior data.

[0063] The human limb spatial position data here is the data of the limbs of the host at each position in the video. The human limb spatial position data can include data of all limbs of the host, or data of part of the limbs of the host. For example, the host is a fitness host, and the human limb spatial position data needs to include spatial position data of all limbs of the host. If the host is a chat host, the human limb spatial position data only needs to include spatial position data of the hands of the host.

[0064] It can be understood that the human limb spatial position data can be determined according to the specific scene of the host. For example, when the host is in a sitting state during live streaming, the human limb spatial position data only includes the spatial position data of the hands of the host. When the host is in a standing state during live streaming, the human limb spatial position data needs to include the spatial position data of all limbs of the host.

[0065] The facial feature data herein is data of facial feature changes of the host during live streaming. For example, eye feature changes, mouth feature changes, eyebrow feature changes, etc.

[0066] In an embodiment, after the video frame data is input into the violation behavior analysis algorithm model, the violation behavior analysis algorithm can also determine the live streaming environment data corresponding to the frame data.

[0067] In a possible implementation, step 203 includes: analyzing the human limb spatial position data and the facial feature change data based on the violation rule library to determine whether the host has a violation behavior corresponding to the live streaming; and if the host has a violation behavior, determining the violation level of the host according to the behavior data and the behavior type and violation level corresponding to the behavior data in the violation rule library.

[0068] The violation behavior herein includes but is not limited to a mild violation behavior, a general violation behavior, a serious violation behavior, etc. It can be understood that the specific classification rules of the mild violation behavior, the general violation behavior, and the serious violation behavior herein can be determined according to platform rules or related prior art solutions or relevant regulations, and the present application does not make specific limitations.

[0069] It can be understood that the analysis of the human limb spatial position data is to determine the position and movement of the limbs of the host corresponding to the frame data according to the human limb spatial position data, and then determine the body movement of the host. The analysis of the facial feature change data is to determine the facial expression and related changes of the facial expression of the host corresponding to the frame data according to the facial feature change data, and then determine the facial expression of the host. The behavior data of the host at the moment is determined based on the body movement and facial expression of the host, and then whether the host has a violation behavior is determined according to the behavior data.

[0070] The violation level is determined according to the size of the impact caused by the violation behavior. For example, the violation level corresponding to the "mild violation behavior" can be set to level 1, the violation level corresponding to the "general violation behavior" can be set to level 2, the violation level corresponding to the "serious violation behavior" can be set to level 3, etc. The violation level can be set according to the actual situation, and the present application does not make specific limitations.

[0071] In the implementation process, the frame data is input into the rule violation analysis algorithm model to determine the limb behavior and facial expression corresponding to the frame data, and then the behavior data of the anchor is determined. The behavior data is matched with the rule violation to determine whether the anchor has a rule violation and the level of the rule violation. The video frame data is analyzed in multiple aspects, the behavior corresponding to the video frame data is accurately determined, the rule violation and the rule violation level are accurately audited in time, the anchor rule violation live streaming is prevented, and the quality of live streaming is improved.

[0072] In a possible implementation, step 202 includes: inputting the video frame data into a facial expression algorithm model; performing facial recognition on the frame data by the facial expression algorithm model to determine the facial data corresponding to the frame data; performing feature extraction on the facial data, and classifying the extracted features to determine the emotion data of the anchor corresponding to the frame data.

[0073] It can be understood that the facial expression algorithm model is provided with a facial recognition algorithm, and the facial recognition algorithm recognizes the face in the video data to determine the facial data. The facial recognition algorithm can be opencv, convolutional neural network, Fisherfaces, PCA, SVM, Haar Cascade, etc.

[0074] The facial data includes multiple data such as eyes, nose, mouth, eyebrows, and face blank area. Further, in order to determine the emotion data of the anchor, the data related to the emotion data in the facial data needs to be extracted. For example, eye data, mouth data, and facial muscle data.

[0075] It can be understood that the feature extraction method can be HOG feature extraction, Dlib library, convolutional neural network feature extraction, etc.

[0076] After the feature extraction of the facial data, multiple feature data are obtained, the feature data are classified according to types, and each type of feature data is analyzed to determine the emotion data of the anchor corresponding to the frame data.

[0077] Exemplarily, the feature data is divided into four types of feature data, i.e., eye data, mouth data, eyebrow data, and crow's feet data. For the eye data, the emotion range corresponding to the eyes can be determined according to the opening and closing degree of the eyes in the eye data. For the mouth data, the emotion range corresponding to the mouth can be determined according to the upwarping or downwarping degree of the mouth and the opening and closing degree of the mouth. For the eyebrow data, the emotion range corresponding to the eyebrows can be determined according to the relaxation or tension degree of the eyebrows. For the crow's feet data, the emotion range corresponding to the crow's feet can be determined according to the length and depth range of the crow's feet. After determining the emotion range corresponding to each type of feature data, the emotion range corresponding to each type of feature data is further matched to determine the emotion data corresponding to the live stream.

[0078] In the implementation process described above, the face expression algorithm model is used to recognize the face data in the frame data and perform feature extraction on the face data to extract the feature data in the face, and the features are classified to determine the emotion range to which each feature belongs, and then the emotion data of the host corresponding to the frame data is determined according to the emotion ranges of the multiple feature data. Since the emotion data is determined based on multiple features, the accuracy of the emotion data can be ensured, and the accuracy of the review can be improved when the live stream is reviewed.

[0079] In a possible implementation, step 202 includes: inputting the video frame data into a body behavior analysis model; detecting the body in the frame data by the body behavior analysis model to determine the body data corresponding to the frame data; performing feature extraction on the body data, and classifying the extracted features to determine the action data of the host corresponding to the frame data.

[0080] It can be understood that the body behavior analysis model is provided with a body behavior analysis algorithm, which is used to recognize the body in the video data to determine the body data. The body behavior analysis algorithm can be an unsupervised learning-based behavior recognition algorithm, a convolutional neural network-based behavior recognition algorithm, a recurrent neural network algorithm, etc.

[0081] The body data includes left hand data, right hand data, left leg data, right leg data, etc. Further, in order to determine the action data of the host, the data related to the body data in the body data also needs to be extracted. For example, left hand data, right hand data, left leg data, and right leg data.

[0082] It can be understood that the feature extraction method can be HOG feature extraction, Dlib library, convolutional neural network feature extraction, etc.

[0083] After the limb data is subjected to feature extraction, a plurality of feature data is obtained, the feature data is classified according to types, and each type of feature data is analyzed to determine the action data of the anchor corresponding to the frame data.

[0084] Exemplarily, the feature data is classified into left hand data, right hand data, left leg data, and right leg data. For the left hand data, the action corresponding to the left hand can be determined according to the stretching degree of the left hand arm and the action of the left hand palm in the left hand data. For the right hand data, the action corresponding to the right hand can be determined according to the stretching degree of the right hand arm and the action of the right hand palm in the right hand data. For the left leg data, the action corresponding to the left leg can be determined according to the bending degree and bending direction of the left leg in the left leg data. For the right leg data, the action corresponding to the right leg can be determined according to the bending degree and bending direction of the right leg in the right leg data. After the action corresponding to each type of feature data is determined, the actions corresponding to each type of feature are further matched (for example, the action of the left hand and the action of the right hand are matched, and the action of the left leg and the action of the right leg are matched), to determine the action data corresponding to the live stream.

[0085] In the above implementation process, the limb data in the frame data is identified by the limb behavior analysis model, and the feature extraction is performed on the limb data to extract the feature data in the limb, and the features are classified to determine the action range of each feature, and then the action data of the anchor corresponding to the frame data is determined according to the action range of the plurality of feature data. Since the action data is determined based on a plurality of features, the accuracy of the behavior data can be ensured, and the accuracy of the review is improved when the live stream is reviewed.

[0086] In a possible implementation, step 203 includes: inputting the emotion data and the limb data into an abnormal alarm model; and comparing the emotion data and the limb data with alarm behaviors in an alarm library by the abnormal alarm model to determine whether the anchor has an alarm behavior.

[0087] The abnormal alarm model includes an alarm library, and the alarm library includes alarm behaviors and corresponding behavior data. The alarm library also includes alarm levels corresponding to the behavior data.

[0088] It can be understood that different alarm levels correspond to different alarm modes. For example, if the alarm level is I, the corresponding alarm mode is to send alarm information to the anchor's mobile phone message or the background notification bar alarm of the live streaming platform. If the alarm level is II, the corresponding alarm mode is to alarm in the live streaming interface. If the alarm level is III, the corresponding alarm mode is to alarm in the live streaming interface. If the alarm level is III, the corresponding alarm mode is to alarm in the live streaming interface.

[0089] Exemplarily, as shown in Table 1, Table 1 is an exemplary alarm library in which alarm behaviors, alarm levels and corresponding behavior data. Those skilled in the art can understand that the alarm behaviors, alarm levels and behavior data in Table 1 are exemplary only, and those skilled in the art can adjust and replace them according to actual conditions, which are not limited by the present application.

[0090] Table 1

[0091]

[0092] In some embodiments, the abnormal alarm model can also be obtained through multiple model training and deep learning. The alarm library can also not be set in the abnormal alarm model. After the emotion data and the body data are input into the abnormal alarm model, the abnormal alarm model directly determines the alarm behavior corresponding to the emotion data and the body data according to the deep learning result.

[0093] In the above implementation process, the emotion data and the action data are input into the abnormal alarm model to match the emotion data and the action data of two types of data with the behaviors in the alarm library through the abnormal alarm model, which can determine the real live streaming state of the host by comprehensively considering the emotion and the action of the host, and more accurately determine the behavior of the host, so that the obtained alarm behavior is more accurate, the false alarm is prevented, and the accuracy of the alarm is improved.

[0094] In a possible implementation, after step 202, the method further includes: scoring the behavior data according to the behavior data and a set scoring rule to determine a live streaming score of the host corresponding to the frame data.

[0095] Among them, the behavior data and the set scoring rule are determined according to the live streaming field of the host.

[0096] The live streaming field here can include live streaming fitness, live streaming singing, live streaming goods selling, live streaming chatting, etc. For example, if the live streaming field is live streaming fitness, the scoring rule can be determined according to the proportion of the exercise time and the live streaming time of the host. If the live streaming field is live streaming singing, the scoring rule can be determined according to the facial expression of the host.

[0097] The live streaming score described above is used to reflect the live streaming quality of the host. For example, the higher the score, the better the quality of the host. Or the lower the score, the higher the quality of the host. The relationship between the live streaming quality of the host and the score can be determined according to the specific scoring rule, which is not specifically limited by the present application.

[0098] In the implementation process, the live broadcast quality of the anchor during the live broadcast is determined by scoring the live broadcast of the anchor according to the behavior data of the anchor and the scoring rule, so as to facilitate the management of the platform on each anchor, and to constrain the live broadcast behavior of the anchor, thereby improving the quality of the live broadcast while realizing the unified management of the anchor.

[0099] In a possible implementation, scoring the behavior data according to the behavior data and the set scoring rule to determine the live broadcast score of the anchor corresponding to the frame data includes: determining a time length proportion of the behavior time length in the live broadcast process according to the behavior time length; determining a number proportion of the behavior number in the live broadcast process according to the behavior number; and determining the live broadcast score of the anchor according to the time length proportion and the number proportion.

[0100] It can be understood that the time length proportion and the number proportion can each account for 50% of the live broadcast score, or the time length proportion can account for 30% of the live broadcast score and the number proportion can account for 70% of the live broadcast score, etc. The ratio of the time length proportion and the number proportion to the live broadcast score can be adjusted according to actual conditions, and the present application does not make specific limitations.

[0101] Here, the behavior time length is the total time length of each behavior data. The number proportion is the ratio of the number of occurrences of each behavior data to the total number of occurrences allowed in the live broadcast.

[0102] For example, if the behavior data is that the four limbs and the head of the anchor are not detected, the number of occurrences of the behavior data is 5, and the total time length of 5 occurrences of the behavior data is 10 minutes. The live broadcast time length is 120 minutes, and the total number of abnormal behaviors allowed in the live broadcast time length is 10. The time length proportion is 8.3%, and the number proportion is 50%. At the same time, the time length proportion and the number proportion each account for 50% of the live broadcast score, and the total score of the live broadcast is 100 points. Therefore, the score of the live broadcast can be calculated to be 71 points.

[0103] For example, if the behavior data is that the four limbs and the head of the anchor are not detected, the number of occurrences of the behavior data is 5, and the total time length of 5 occurrences of the behavior data is 10 minutes. The live broadcast time length is 120 minutes, and the total number of abnormal behaviors allowed in the live broadcast time length is 10. The time length proportion is 8.3%, and the number proportion is 50%. At the same time, the time length proportion and the number proportion each account for 50% of the live broadcast score, and the total score of the live broadcast is 100 points. Therefore, the score of the live broadcast can be calculated to be 71 points.

[0104] In the implementation process, the anchor's live streaming score in the live streaming process is calculated according to the proportion of the behavior data and the behavior frequency data after removing the deduction of the alarm generation. The real live streaming situation of the anchor in the live streaming can be comprehensively reflected, so as to improve the accuracy of determining the live streaming quality of the anchor.

[0105] Based on the same application concept, the application embodiment also provides an anchor live streaming behavior auditing device corresponding to the anchor live streaming behavior auditing method. Since the principle of solving problems in the device in the application embodiment is similar to the aforementioned anchor live streaming behavior auditing method embodiment, the implementation of the device in the embodiment can be referred to the description in the aforementioned method embodiment, and the repeated parts will not be described here.

[0106] Please refer to Figure 4 FIG. 1 is a functional module schematic diagram of an anchor live streaming behavior auditing device provided in the application embodiment. Each module in the anchor live streaming behavior auditing device in the embodiment is used to execute each step in the aforementioned method embodiment. The anchor live streaming behavior auditing device includes a decoding module 301, an analysis module 302, and a determination module 303.

[0107] The decoding module 301 is configured to decode the acquired live streaming into video frame data.

[0108] The analysis module 302 is configured to input the video frame data into multiple algorithm models for analysis to obtain behavior data of an anchor corresponding to the live streaming; wherein the algorithm models include a rule violation behavior analysis algorithm model.

[0109] The determination module 303 is configured to determine whether the live streaming behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule.

[0110] In a possible implementation, the analysis module 302 is further configured to: input the video frame data into the rule violation behavior analysis algorithm model; and identify the frame data through the rule violation behavior analysis algorithm model to determine human body limb spatial position data and face feature data corresponding to the frame data; wherein the rule violation behavior analysis algorithm model includes a rule violation rule library, and the rule violation rule library includes behavior types and rule violation levels corresponding to multiple behavior data.

[0111] In a possible implementation, the determination module 303 is further configured to: analyze the human body limb spatial position data and the face feature change data based on the rule violation rule library to determine whether the anchor corresponding to the live streaming has a rule violation behavior; and if the anchor has a rule violation behavior, determine a rule violation level of the anchor according to the behavior data and the behavior types and rule violation levels corresponding to the multiple behavior data in the rule violation rule library.

[0112] In a possible implementation, the analysis module 302 is specifically configured to: input the video frame data into the facial expression algorithm model; perform facial recognition on the frame data by using the facial expression algorithm model to determine facial data corresponding to the frame data; perform feature extraction on the facial data, and classify the extracted features to determine emotional data of the anchor corresponding to the frame data.

[0113] In a possible implementation, the analysis module 302 is specifically configured to: input the video frame data into the body behavior analysis model; detect a body in the frame data by using the body behavior analysis model to determine body data corresponding to the frame data; perform feature extraction on the body data, and classify the extracted features to determine body data of the anchor corresponding to the frame data.

[0114] In a possible implementation, the analysis module 302 is specifically configured to: input the emotional data and the body data into the abnormal alarm model; compare the emotional data and the body data with alarm behaviors in an alarm library by using the abnormal alarm model to determine whether the anchor has an alarm behavior; wherein the abnormal alarm model includes an alarm library, and the alarm library includes alarm behaviors and corresponding behavior data.

[0115] In a possible implementation, the anchor live broadcast behavior auditing apparatus further includes a scoring module configured to score the behavior data according to a set scoring rule to determine a live broadcast score of the anchor corresponding to the frame data; wherein the behavior data and the set scoring rule are determined according to a live broadcast field of the anchor.

[0116] In a possible implementation, the anchor live broadcast behavior auditing apparatus further includes a scoring module, and the scoring module is further configured to: determine a time length proportion of the behavior time length in the live broadcast process according to the behavior time length; determine a frequency proportion of the behavior frequency in the live broadcast process according to the behavior frequency; and determine the live broadcast score of the anchor according to the time length proportion and the frequency proportion.

[0117] In addition, the embodiment of the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is run by a processor, the steps of the anchor live broadcast behavior auditing method described in the above method embodiment are executed.

[0118] The computer program product of the anchor live broadcast behavior auditing method provided by the embodiment of the present application includes a computer readable storage medium storing program codes. The instructions included in the program codes can be used to execute the steps of the anchor live broadcast behavior auditing method described in the above method embodiment. For details, refer to the above method embodiment, which will not be described here.

[0119] It should be understood that all the functional units in the embodiments of the present application can be integrated into one processing unit, or each can exist as separated independent unit. Accordingly, all the units integrated into one processing unit can be considered as a module. The unit integrating the modules can be implemented in the form of hardware embedded in a chip, or in the form of software functional module.

[0120] In addition, each functional module in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0121] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the part that contributes to the prior art, or part of the technical solutions. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes. It should be noted that in this paper, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to the process, method, article or device. Without more limitations, the elements defined by the statement "include" do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0122] The above is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0123] The above is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. The above is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application. It should be noted that similar reference numbers and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

Claims

1. A method for auditing a host live streaming behavior, characterized in that, The method comprises: decoding the obtained live stream into video frame data; inputting the video frame data into multiple algorithm models for analysis to obtain behavior data of the anchor corresponding to the live stream; wherein the algorithm models comprise a rule violation behavior analysis algorithm model; and determining whether the live streaming behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule; wherein the algorithm models further comprise a facial expression algorithm model, and the behavior data comprises emotion data; the step of inputting the video frame data into multiple algorithm models for analysis to obtain behavior data of the anchor corresponding to the live stream comprises: inputting the video frame data into the facial expression algorithm model; performing face recognition on the frame data through the facial expression algorithm model to determine the face data corresponding to the frame data; extracting features from the face data and classifying the extracted features to determine the emotion data of the anchor corresponding to the frame data.

2. The method of claim 1, wherein, wherein the behavior data comprises human limb spatial position data and facial feature data; and the abnormality comprises a rule violation behavior; the step of inputting the video frame data into multiple algorithm models for analysis to obtain behavior data of the anchor corresponding to the live stream comprises: inputting the video frame data into the rule violation behavior analysis algorithm model; identifying the frame data through the rule violation behavior analysis algorithm model to determine the human limb spatial position data and the facial feature data corresponding to the frame data; wherein the rule violation behavior analysis algorithm model comprises a rule violation rule library, and the rule violation rule library comprises behavior types and rule violation levels corresponding to multiple behavior data; the step of determining whether the live streaming behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule comprises: analyzing the human limb spatial position data and the facial feature change data based on the rule violation rule library to determine whether the anchor corresponding to the live stream has a rule violation behavior; and if the anchor has a rule violation behavior, determining the rule violation level of the anchor according to the behavior data and the behavior types and rule violation levels corresponding to multiple behavior data in the rule violation rule library.

3. The method of claim 1, wherein, wherein the algorithm models further comprise a limb behavior analysis model, and the behavior data comprises limb data; the step of inputting the video frame data into multiple algorithm models for analysis to obtain behavior data of the anchor corresponding to the live stream comprises: inputting the video frame data into the limb behavior analysis model; detecting the limbs in the frame data through the limb behavior analysis model to determine the limb data corresponding to the frame data; extracting features from the limb data and classifying the extracted features to determine the action data of the anchor corresponding to the frame data.

4. The method of claim 3, wherein, wherein the abnormality comprises an alarm; and the step of determining whether the live streaming behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule comprises: inputting the emotion data and the action data into an abnormal alarm model; comparing the emotion data and the action data with alarm behaviors in an alarm library through the abnormal alarm model to determine whether the anchor has an alarm behavior; and The abnormal alarm model includes an alarm library, and the alarm library includes alarm behaviors and corresponding behavior data.

5. The method of claim 1, wherein, After the video frame data is input into multiple algorithm models for analysis to obtain the behavior data of the anchor corresponding to the live stream, the method further includes: According to the behavior data and a set scoring rule, the behavior data is scored to determine a live score of the anchor corresponding to the frame data; The behavior data and the set scoring rule are determined according to the live field of the anchor.

6. The method of claim 5, wherein, The behavior data includes a behavior duration and a behavior frequency. According to the behavior data and a set scoring rule, the behavior data is scored to determine a live score of the anchor corresponding to the frame data, including: According to the behavior duration, a duration proportion of the behavior duration in the live process is determined; According to the behavior frequency, a frequency proportion of the behavior frequency in the live process is determined; According to the duration proportion and the frequency proportion, the live score of the anchor is determined. The method includes:

7. An anchor live broadcast behavior auditing apparatus, characterized in that, A decoding module is configured to decode the obtained live stream into video frame data; An analysis module is configured to input the video frame data into multiple algorithm models for analysis to obtain behavior data of an anchor corresponding to the live stream; wherein the algorithm models include a rule violation behavior analysis algorithm model; A determination module is configured to determine whether the live behavior of the anchor is abnormal according to the behavior data and a preset abnormal behavior rule. The algorithm models further include a facial expression algorithm model, and the analysis module is specifically configured to: input the video frame data into the facial expression algorithm model; perform facial recognition on the frame data through the facial expression algorithm model to determine face data corresponding to the frame data; extract features from the face data, and classify the extracted features to determine emotional data of the anchor corresponding to the frame data. The method includes:

8. An electronic device, comprising: A processor and a memory, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the machine readable instructions are executed by the processor to perform the steps of the method of any one of claims 1 to 6. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the steps of the method of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, ​

Citation Information

Patent Citations

  • Method and device for detecting violation behavior of live streaming room, electronic equipment and storage medium

    CN113705370A