Live broadcast illegal behavior prediction method and system

Through the lightweight multimodal information integration model, combining voice, text and image information, live broadcast violations are predicted, and the problems of poor early warning effects and high false detection rates in the existing technology are solved, and efficient and accurate predictions and reminders of violations are achieved.

CN120111282APending Publication Date: 2025-06-06XIAMEN MEIYABAIKE INFORMATION SECURITY RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510099376.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing live broadcast violation detection technology mainly relies on in-process detection, with poor early warning effect and high error detection rate. It also uses standard action templates as reference information, and the setting premise is too strong, which is not conducive to improving the accuracy of the model.

Method used

A lightweight multimodal information integration model structure is proposed, through synchronous integration of voice, text, and images, the violation process is modeled, so as to effectively predict upcoming violations by relying solely on video images, and remind them in the form of overview of text in the theme.

Benefits of technology

It realizes efficient prediction of live broadcast violations, reduces the false alarm rate, and improves the attention concentration of auditors. The overall efficiency is high, the structure is simple, and the display information is refined and accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111282A_ABST
    Figure CN120111282A_ABST
Patent Text Reader

Abstract

The invention discloses a live broadcast violation behavior prediction method and system, and the method comprises the steps: S1, collecting violation live broadcast video data of various different violation behavior types, constructing a violation live broadcast video training set and a violation live broadcast video verification set, extracting the corpus of the violation live broadcast video data, and constructing a violation live broadcast corpus set; s2, data of the violation live broadcast corpus set is preprocessed and then input into a first neural network model for training, and the first neural network model comprises a sentence encoder, a process analyzer and a sentence prediction decoder; s3, freezing the sentence prediction decoder, replacing the sentence encoder with a video encoder to obtain a second neural network model, and inputting the illegal live video training set into the second neural network model for training; and S4, inputting the illegal live video verification set into a trained second neural network model to predict the possibility and type of illegal behaviors, thereby performing early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network detection technology, and in particular to a method and system for predicting live broadcast violation behavior. Background Art

[0002] Live broadcast violation detection is an important means to ensure that the current live video content is compliant and legal, but it consumes a huge amount of human resources and the workflow is mechanical and cumbersome, which makes the majority of live broadcast platform companies overwhelmed. In recent years, with the increasing advancement of computer vision technology, AI intelligent review has become a tool to liberate the labor of live broadcast review and reduce corporate costs.

[0003] The current technical solutions for live broadcast violation detection based on deep learning are mainly in-process detection, and most of them rely on image classification models constructed for picture frames, with poor warning effects and high false detection rates. Among them, in the existing technologies, some warning methods simply integrate the recognition results of multiple service providers for judgment, and the fusion judgment of the output results of the existing live broadcast filtering system is more inclined to a scientific application method, and the cost is relatively high; the basic principles of other technologies can be summarized as building an abnormal and normal feature library for live broadcast personnel, and using feature matching to achieve action behavior analysis, and it is also impossible to predict possible violations. In essence, it is a kind of in-process processing technology, and it uses standardized action templates as reference information, setting too strong premises, which is not conducive to improving the accuracy of the model.

[0004] Although the current relevant algorithms have played an obvious role, they also have corresponding drawbacks, that is, they are mainly based on real-time monitoring of content and handling after violations. Such handling is often delayed before the violation occurs, bringing greater risks. Effectively predicting possible violations and clarifying the focus and direction of audits in real time are urgent needs in the industry technology field. Summary of the invention

[0005] The present invention proposes a method and system for predicting live broadcast violation behavior. The method is based on a lightweight multimodal information integration model structure, can model the violation behavior process, and synchronously integrates voice, text, and image information. Finally, in the reasoning stage, it is possible to effectively predict upcoming violation events only by relying on video images, and give reminders in the form of topic summary text. The method has high overall efficiency, simple structure, concise and accurate display of information, and can effectively guide auditors to focus on the predicted content.

[0006] According to one aspect of the present invention, a method for predicting live broadcast violation behavior is proposed, comprising the following steps:

[0007] S1. Collect illegal live video data of various illegal behavior types, construct an illegal live video training set and an illegal live video verification set, and extract the corpus of the illegal live video data to construct an illegal live corpus set;

[0008] S2, preprocessing the data of the illegal live broadcast corpus, and then inputting it into a first neural network model for training, wherein the first neural network model includes a sentence encoder, a flow analyzer, and a sentence prediction decoder;

[0009] S3, freezing the sentence prediction decoder, and replacing the sentence encoder with a video encoder to obtain a second neural network model, inputting the illegal live video training set into the second neural network model for training, inputting the illegal live video training set into the video encoder to extract video hidden layer features, then performing sequence modeling based on the process analyzer, and finally decoding through the sentence prediction decoder;

[0010] S4. Input the illegal live video verification set into the trained second neural network model to predict the possibility and type of illegal behavior, so as to issue an early warning.

[0011] Preferably, the step S1 extracts the corpus of the illegal live video data, specifically including: extracting subtitle corpus or barrage corpus through OCR and extracting dialogue corpus and self-talk corpus through speech recognition algorithm. By collecting corpus from different characters in different illegal scenes during the live broadcast, the subsequent neural network can comprehensively and comprehensively learn the characteristics of illegal behaviors in different scenes during the live broadcast, thereby improving the accuracy of neural network classification and prediction.

[0012] Preferably, the multiple different violation types described in S1 specifically include pornographic violation types, political violation types, violent behavior types and smoking / drug violation types.

[0013] Preferably, the preprocessing of the data of the illegal live broadcast corpus set in S2 specifically includes inputting the data of the illegal live broadcast corpus set into the dialogue summary model for compression and translation, and then segmenting it through the role separation algorithm. Through the dialogue summary model, the corpus information is condensed and the main information is extracted, which helps to lightweight the neural network and reduce the training time for subsequent neural network prediction and classification. Then, the role separation algorithm is used to distinguish the corpus texts sent by different roles (such as anchor speech and barrage text) during the illegal live broadcast, so that the neural network model can comprehensively consider the influence of corpus from different sources on the proportion of illegal behavior.

[0014] Preferably, the sentence encoder is constructed by a local sensitive hash attention encoder, and the sentence prediction decoder is constructed by a local sensitive hash attention decoder. The local sensitive hash attention encoder-decoder greatly reduces the number of queries and helps to make the neural network lightweight.

[0015] Preferably, the illegal live video training set in S3 is input into the video encoder to extract video hidden layer features, specifically including: assuming that the jth video segment c of the illegal live video training set j With L frames, the video segment c is extracted by the video encoder j Video frame embedding vector for each frame Then pass through the bidirectional Transformer of the video encoder v_e The structure is vector mapped to obtain:

[0016]

[0017] Then, the maximum pooling is performed to take the first d dimensions of the largest value in the time step t to obtain the hidden layer feature of the video (q j ) d ,Right now

[0018] Further preferably, the process analyzer is an LSTM neural network model, and the sequence modeling based on the process analyzer in S3 specifically includes: the video hidden layer feature (q j ) d Input the process analyzer, the process analyzer models the ordering information of the live broadcast violation process steps, and processes the output Right now By performing end-to-end joint training through the sentence encoder, process analyzer, and sentence prediction decoder, and then replacing the sentence encoder with a video decoder and performing secondary training, the process LSTM can receive signal feeds from text and visual features and use them to better interpret the sentence semantics.

[0019] According to one aspect of the present invention, a live broadcast violation prediction system is proposed, comprising the following modules:

[0020] Illegal live broadcast data collection module: collects illegal live broadcast video data of various illegal behavior types, constructs illegal live broadcast video training set and illegal live broadcast video verification set, and extracts the corpus of the illegal live broadcast video data to construct an illegal live broadcast corpus set;

[0021] The first neural network model training module is used to pre-process the data of the illegal live broadcast corpus and then input the data into the first neural network model for training, wherein the first neural network model includes a sentence encoder, a flow analyzer, and a sentence prediction decoder;

[0022] Second neural network model training module: freeze the sentence prediction decoder, and replace the sentence encoder with a video encoder to obtain a second neural network model, input the illegal live video training set into the second neural network model for training, input the illegal live video training set into the video encoder to extract video hidden layer features, then perform sequence modeling based on the process analyzer, and finally decode through the sentence prediction decoder;

[0023] Illegal live broadcast behavior prediction module: the illegal live broadcast video verification set is input into the trained second neural network model to predict the possibility and type of illegal behavior, so as to issue an early warning.

[0024] According to one aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the method according to any one of the first aspects is implemented.

[0025] According to one aspect of the present invention, a computing system is provided, comprising a processor and a memory, wherein the processor is configured to execute the method as described in any one of the first aspects.

[0026] Different from the prior art, the present invention is beneficial in that:

[0027] ① The method proposed in the present invention realizes the "pre-processing" solution for live broadcast violations, adopts multiple modes to analyze live broadcast information, and proposes a multi-modal framework to integrate information, which can predict possible behaviors in live broadcasts;

[0028] ② In addition, the present invention proposes a lightweight multimodal model structure by combining transformer and LSTM, which reduces the parameter requirements of the multimodal model in the dedicated field and greatly improves the practicality;

[0029] ③Finally, since the present invention uses natural language text description as output, it expands the information expression of the label space and can effectively reduce the false alarm rate. The method in the present invention improves the processing efficiency and speed without increasing the amount of calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and are used together with the description to explain the principles of the present invention. It will be easy to recognize other embodiments and many expected advantages of the embodiments because they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding similar parts.

[0031] Figure 1A schematic diagram showing a flow chart of a live broadcast violation prediction method according to the present invention is shown;

[0032] Figure 2 A schematic diagram of the structure of a live broadcast violation prediction system according to the present invention is shown;

[0033] Figure 3 A schematic diagram of the computer system structure of an electronic device suitable for implementing an embodiment of the present application is shown. DETAILED DESCRIPTION

[0034] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It should also be noted that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0035] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0036] Figure 1 According to the present invention

[0037] S1. Collect illegal live video data of various illegal behavior types, construct an illegal live video training set and an illegal live video verification set, and extract the corpus of the illegal live video data to construct an illegal live corpus set;

[0038] S2, preprocessing the data of the illegal live broadcast corpus, and then inputting it into a first neural network model for training, wherein the first neural network model includes a sentence encoder, a flow analyzer, and a sentence prediction decoder;

[0039] S3, freezing the sentence prediction decoder, and replacing the sentence encoder with a video encoder to obtain a second neural network model, inputting the illegal live video training set into the second neural network model for training, inputting the illegal live video training set into the video encoder to extract video hidden layer features, then performing sequence modeling based on the process analyzer, and finally decoding through the sentence prediction decoder;

[0040] S4. Input the illegal live video verification set into the trained second neural network model to predict the possibility and type of illegal behavior, so as to issue an early warning.

[0041] In a specific embodiment, the multiple different violation types described in S1 specifically include pornographic violation types, political violation types, violent behavior types, and smoking / drug violation types. i≥3000 segments, i=1...5. The corresponding characteristics are shown in Table 1:

[0042] Table 1 Corresponding characteristics of different types of violations

[0043]

[0044]

[0045] It is worth noting that the predictable behavior types of the prediction method proposed by the present invention include but are not limited to the above range. In addition to the above, the live broadcast violations specified by other platforms can be collected and trained, predicted, and classified according to the method disclosed by the present invention, which should be covered within the scope of protection of this application.

[0046] In a specific embodiment, the step S1 extracts the corpus of the illegal live video data, specifically including: extracting subtitle corpus or barrage corpus through OCR and extracting dialogue corpus and self-talk corpus through speech recognition algorithm. Through the dialogue summary model, the extracted huge original live illegal corpus text is described through condensed sentences to obtain the main information in the original text, and then the condensed sentences are segmented through the role separation algorithm to distinguish the corpus text issued by different roles (such as anchor speech, barrage text) during the illegal live broadcast. When the subsequent neural network is learning, it can comprehensively consider the influence of corpus from different sources on the proportion of illegal behaviors, thereby achieving better prediction and classification effects.

[0047] In a specific implementation, the sentence encoder is constructed by a local sensitive hashing attention encoder, and the sentence prediction decoder is constructed by a local sensitive hashing attention decoder.

[0048] The principle of the local sensitive hashing attention encoder and decoder is described as follows: first, the key value (or attention score) data in the attention mechanism is organized into a matrix A, and then A is mapped to a key hash signature matrix C by the simhash function. Then, C is hashed using the LSH operator so that each data point is assigned to a bucket, and similar data points are concentrated in the same bucket as much as possible. In this way, the hash bucket is equivalent to establishing a clustering index, and the calculated attention score of a query and other tokens mainly depends on the tokens with the highest similarity. Therefore, approximate nearest neighbor search can be achieved, greatly reducing the number of queries. The local sensitive hashing attention can be formally defined as follows:

[0049]

[0050] Among them, i represents the attention score of the i-th token, Dot represents the inner product operation, and qi represents the query vector, k j represents the key vector, v j represents the value vector, p i represents the set that the query vector at position i focuses on, and c represents the set that does not belong to p i The position of is penalized. z is the partition function.

[0051] The sentence encoder converts the sentences segmented by the role separation algorithm into a vector representation of a fixed length. Set the jth sentence s as the input to the sentence encoder j It is represented by M words, where M is a positive integer greater than 1, that is, and It's a word For each sentence j, the bidirectional Transformer of the sentence encoder s_e Hidden features of structured output corpus in:

[0052]

[0053] The present invention defines that the expression of the entire sentence is represented by Performing maximum pooling on the sequence yields:

[0054]

[0055] where max t∈{1,...,M} Indicates taking the maximum value in time step t, () d It means taking the first d dimensions, where d is a positive integer.

[0056] Next, the summary of the live broadcast violation dialogue is summarized and modeled in terms of process. In one embodiment, LSTM is used to model the ordering information of the steps of the live broadcast violation process. 0 , ..., p j} as input. LSTM can convert this vector sequence into a vector sequence containing the sorting information of the steps in the live broadcast violation process, which can serve as the basis for the prediction of the next step. Each future step prediction is based on the previous step, and the formal expression is as follows:

[0057]

[0058] It is worth noting that LSTM is only a preferred embodiment of the present invention. LSTM selectively retains and forgets information through forget gates, input gates and output gates, effectively processes time series data, and models the sorting information of the steps of the live broadcast violation process. The process analyzer can also choose other neural network models for processing time series data, such as RNN, etc., which should not be understood as a limitation to the present invention.

[0059] Then decode it through the sentence prediction decoder, and the corresponding expression is as follows:

[0060]

[0061] is the decoded sentence sequence, Represents the predicted Mth word.

[0062] In order to achieve directional integration of video coding in the reasoning stage, the present invention hopes that the process LSTM can accept dual signal feeds from text and visual features, and use this to better interpret the semantics of the sentence.

[0063] Therefore, in a specific embodiment, after the first neural network performs end-to-end joint training, the sentence prediction decoder is frozen, and the sentence encoder is replaced with a video encoder to obtain a second neural network model. The second neural network includes a video encoder, a flow analyzer, and a sentence decoding predictor. The illegal live video training set in S3 is input into the video encoder to extract video hidden layer features, specifically including: assuming that the jth video segment c of the illegal live video training set j With L frames, the video segment c is extracted by the video encoder j The video frame embedding vector of each frame is obtained Then pass through the bidirectional Transformer of the video encoder v_e The structure is vector mapped to obtain:

[0064]

[0065] Then, the maximum pooling is performed to take the first d dimensions of the largest value in the time step t to obtain the hidden layer feature of the video (q j ) d ,Right now

[0066] In one embodiment, the video encoder is constructed by CNN to decode the video segment c. j Each frame can be mapped into a vector using CNN and then passed through a bidirectional Transformer v_e structure

[0067] Further preferably, the process analyzer is an LSTM neural network model, and the sequence modeling based on the process analyzer in S3 specifically includes: the video hidden layer feature (q j ) d Input the process analyzer, the process analyzer models the ordering information of the live broadcast violation process steps, and processes the output Right now By performing end-to-end joint training through the sentence encoder, process analyzer, and sentence prediction decoder, and then replacing the sentence encoder with a video decoder and performing secondary training, the process LSTM can receive signal feeds from text and visual features and use them to better interpret the sentence semantics.

[0068] The second neural network model is used to perform reasoning, that is, to demonstrate the ability to predict the next illegal behavior scenario and possibility based on the live video clip, and then based on the output of the second neural network model, the prediction description of the judgment algorithm output is judged based on the scoring rule to determine whether it contains illegal content. The scoring rule can be the process consistency of the output of the second neural network model. Inertia means that the same warning is continuously issued in several consecutive time steps.

[0069] For example, in one embodiment, the second neural network model predicts and outputs: "It is expected that the anchor will strip, pornographic violation warning" and the corresponding possibility. If the continuous prediction process results meet the label of the live broadcast violation type within N steps, and the possibility continues to be higher than the set threshold, an early warning is issued, and the relevant early warning information is sent to the auditor or the handling link for further handling.

[0070] According to one aspect of the present invention, a live broadcast violation prediction system is proposed, comprising the following modules:

[0071] Illegal live broadcast data collection module 201: collects illegal live broadcast video data of various illegal behavior types, constructs an illegal live broadcast video training set and an illegal live broadcast video verification set, and extracts the corpus of the illegal live broadcast video data to construct an illegal live broadcast corpus set;

[0072] The first neural network model training module 202 pre-processes the data of the illegal live broadcast corpus and then inputs it into the first neural network model for training, wherein the first neural network model includes a sentence encoder, a flow analyzer, and a sentence prediction decoder;

[0073] Second neural network model training module 203: freeze the sentence prediction decoder, and replace the sentence encoder with a video encoder to obtain a second neural network model, input the illegal live video training set into the second neural network model for training, input the illegal live video training set into the video encoder to extract video hidden layer features, then perform sequence modeling based on the process analyzer, and finally decode through the sentence prediction decoder;

[0074] The illegal live broadcast behavior prediction module 204: inputs the illegal live broadcast video verification set into the trained second neural network model to predict the possibility and type of illegal behavior, so as to issue an early warning.

[0075] Reference below Figure 3, which shows a schematic diagram of the structure of a computer system 700 suitable for implementing an electronic device of an embodiment of the present application. Figure 3 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0076] like Figure 3 As shown, the computer system 300 includes a central processing unit (CPU) 301, which performs various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage part 309 into a random access memory (RAM) 304. In the RAM 304, various programs and data required for the operation of the system 300 are also stored. The CPU 301, the ROM 302, the ROM 303, and the RAM 304 are connected to each other through a bus 305. An input / output (I / O) interface 306 is also connected to the bus 305.

[0077] The following components are connected to the I / O interface 306: an input section 307 including a keyboard, a mouse, etc.; an output section 308 including a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 309 including a hard disk, etc.; and a communication section 310 including a network interface card such as a LAN card, a modem, etc. The communication section 310 performs communication processing via a network such as the Internet. A drive 311 is also connected to the I / O interface 306 as needed. A removable medium 312, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 311 as needed so that a computer program read therefrom is installed into the storage section 309 as needed.

[0078] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart is implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program is downloaded and installed from the network through the communication part 310, and / or installed from the removable medium 312. When the computer program is executed by the central processing unit (CPU) 301, the above-mentioned functions defined in the method of the present application are executed.

[0079] It should be noted that the computer-readable storage medium of the present application is a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium is, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium is any tangible medium containing or storing a program that is used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium includes a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal takes a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium is any computer-readable storage medium other than a computer-readable storage medium that sends, propagates or transmits a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable storage medium is transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0080] Computer program code for performing the operations of the present application is written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code is executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer is connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or is connected to an external computer (e.g., via the Internet using an Internet service provider).

[0081] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram represents a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the box also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession are actually executed substantially in parallel, and they are sometimes also executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, are implemented by a dedicated hardware-based system that performs the specified function or operation, or are implemented by a combination of dedicated hardware and computer instructions.

[0082] The modules involved in the embodiments of the present application are implemented by software and hardware.

[0083] As another aspect, the present application further provides a computer-readable storage medium, which is included in the electronic device described in the above embodiment; the computer-readable storage medium also exists independently and is not assembled into the electronic device. The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: S1, collects illegal live video data of multiple different types of illegal behaviors, constructs an illegal live video training set and an illegal live video verification set, and extracts the corpus of the illegal live video data to construct an illegal live corpus set; S2, pre-processes the data of the illegal live corpus set, and then inputs it into the first neural network model for training, wherein the first neural network model includes a sentence encoder, a process analyzer, and a sentence prediction decoder; S3, freezes the sentence prediction decoder, and replaces the sentence encoder with a video encoder to obtain a second neural network model, inputs the illegal live video training set into the second neural network model for training, inputs the illegal live video training set into the video encoder to extract video hidden layer features, and then performs sequence modeling based on the process analyzer, and finally decodes it through the sentence prediction decoder; S4, inputs the illegal live video verification set into the trained second neural network model to predict the possibility and type of illegal behaviors, thereby issuing an early warning.

[0084] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above invention concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other.

Claims

1. A method for predicting live broadcast violation behavior, characterized in that: The following steps are involved: S1. Collect illegal live video data of various illegal behavior types, construct an illegal live video training set and an illegal live video verification set, and extract the corpus of the illegal live video data to construct an illegal live corpus set; S2, preprocessing the data of the illegal live broadcast corpus, and then inputting it into a first neural network model for training, wherein the first neural network model includes a sentence encoder, a flow analyzer, and a sentence prediction decoder; S3, freezing the sentence prediction decoder, and replacing the sentence encoder with a video encoder to obtain a second neural network model, inputting the illegal live video training set into the second neural network model for training, inputting the illegal live video training set into the video encoder to extract video hidden layer features, then performing sequence modeling based on the process analyzer, and finally decoding through the sentence prediction decoder; S4. Input the illegal live video verification set into the trained second neural network model to predict the possibility and type of illegal behavior, so as to issue an early warning.

2. The prediction method according to claim 1, characterized in that: The step of extracting the corpus of the illegal live video data in S1 specifically includes: extracting subtitle corpus or barrage corpus through OCR and extracting dialogue corpus and self-talk corpus through speech recognition algorithm.

3. The prediction method according to claim 1, characterized in that: The various violation types described in S1 specifically include pornographic violation types, political violation types, violent behavior types, and smoking / drug violation types.

4. The prediction method according to claim 1, characterized in that: The preprocessing of the data of the illegal live broadcast corpus set in S2 specifically includes inputting the data of the illegal live broadcast corpus set into a dialogue summary model for compression and translation, and then segmenting the data using a role separation algorithm.

5. The prediction method according to claim 1, characterized in that: The sentence encoder is constructed by a local sensitive hashing attention encoder, and the sentence prediction decoder is constructed by a local sensitive hashing attention decoder.

6. The prediction method according to claim 1, characterized in that: The illegal live video training set in S3 is input into the video encoder to extract the video hidden layer features, specifically including: assuming that the jth video segment c of the illegal live video training set j With L frames, the video segment c is extracted by the video encoder j Video frame embedding vector for each frame Then pass through the bidirectional Transformer of the video encoder v_e The structure is vector mapped to obtain: Then, the maximum pooling is performed to take the first d dimensions of the largest value in the time step t to obtain the hidden layer feature of the video (q j ) d ,Right now 7. The prediction method according to claim 6, characterized in that: The process analyzer is an LSTM neural network model. The sequence modeling based on the process analyzer in S3 specifically includes: j ) d Input the process analyzer, the process analyzer models the ordering information of the live broadcast violation process steps, and processes the output Right now 8. A live broadcast violation prediction system, characterized in that: Includes the following modules: Illegal live broadcast data collection module: collects illegal live broadcast video data of various illegal behavior types, constructs illegal live broadcast video training set and illegal live broadcast video verification set, and extracts the corpus of the illegal live broadcast video data to construct an illegal live broadcast corpus set; The first neural network model training module is used to pre-process the data of the illegal live broadcast corpus and then input the data into the first neural network model for training, wherein the first neural network model includes a sentence encoder, a flow analyzer, and a sentence prediction decoder; Second neural network model training module: freeze the sentence prediction decoder, and replace the sentence encoder with a video encoder to obtain a second neural network model, input the illegal live video training set into the second neural network model for training, input the illegal live video training set into the video encoder to extract video hidden layer features, then perform sequence modeling based on the process analyzer, and finally decode through the sentence prediction decoder; Illegal live broadcast behavior prediction module: the illegal live broadcast video verification set is input into the trained second neural network model to predict the possibility and type of illegal behavior, so as to issue an early warning.

9. A computer program product having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 7.

10. A computing system comprising a processor and a memory, wherein the processor is configured to execute the method according to any one of claims 1 to 7.