E-commerce live broadcast violation detection method, device, equipment, medium and product
By dividing the process of e-commerce live broadcast violation detection into two stages, using pre-trained classification models to filter compliance data in the first stage, and focusing on handling potential violation content in the second stage, it solves the problems of difficulty, low accuracy and low efficiency in e-commerce live broadcast violation detection, and achieves more efficient and accurate detection results.
Patent Information
- Application Number
- CN202510449370.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
E-commerce live broadcast violation detection is difficult, has low accuracy and low efficiency.
The pre-trained classification model is used to divide the violation detection process into two stages: the first stage determines whether the content is compliant, and the second stage determines whether the content is serious violation. By filtering out most compliance data in the first phase, reducing the computational burden in the second phase, making the model more focused on handling potential violations.
The efficiency and accuracy of illegal detection have been improved, and it is more efficient and accurate than traditional manual detection.
Smart Images

Figure CN119996741A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and in particular, relates to a method, device, equipment, medium and product for detecting violations in e-commerce live broadcasts. Background Art
[0002] E-commerce live streaming, where merchants (anchors) introduce their products through e-commerce platforms and persuade viewers to buy them, has become a popular form of marketing. However, due to the short time and rapid growth of the e-commerce live streaming industry, it has provided a breeding ground for a series of live streaming chaos, such as false propaganda and the use of inappropriate terms. In addition, due to the huge traffic and information dissemination speed of live streaming e-commerce platforms, false information can spread rapidly, seriously threatening the network information security of users. Therefore, how to effectively supervise and govern these live streaming chaos is of great social significance.
[0003] Video review is an important measure of supervision. An intuitive review strategy is to allow viewers to report video sources containing abnormal content and protect themselves by using block lists. However, the effectiveness of this approach relies on user participation in the system, which may disrupt the normal use of the video platform and affect the user experience. In addition, it may be difficult for viewers to identify some obscure illegal information.
[0004] Another review strategy is "filter first, then review". The platform first builds a sensitive word library, and conducts preliminary screening by searching whether the live broadcast text contains any illegal information. Specialized reviewers then determine whether the content is illegal. However, this method has low detection accuracy and efficiency. Summary of the invention
[0005] The embodiments of the present application provide an e-commerce live broadcast violation detection method, device, equipment, medium and product, which are used to at least solve the problems of difficulty, low accuracy and low efficiency in e-commerce violation detection in related technologies.
[0006] In a first aspect, an embodiment of the present application provides an e-commerce live broadcast violation detection method, comprising: Acquire real-time live broadcast data, where the real-time live broadcast data is obtained by performing voice recognition on the acquired real-time live broadcast video, and the data type of the real-time live broadcast data is text data; Using a preset model to tokenize the real-time live broadcast data to obtain a corresponding text vector; Inputting the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of whether the first violation is compliant; In the case where the first violation detection result is non-compliant, inputting the text vector corresponding to the real-time live broadcast data into a pre-trained two-stage classification model to obtain a second violation detection result, wherein the second violation detection result includes suspected violation or serious violation; The one-stage classification model and the two-stage classification model are obtained by training based on multiple live data samples and the violation detection category label corresponding to each live data sample.
[0007] In a second aspect, an embodiment of the present application provides an e-commerce live broadcast violation detection device, the device comprising: An acquisition module is used to acquire real-time live broadcast data, wherein the real-time live broadcast data is obtained by performing voice recognition on the acquired real-time live broadcast video, and the data type of the real-time live broadcast data is text data; A tokenization module, used to tokenize the real-time live broadcast data using a preset model to obtain a corresponding text vector; A first input module, used to input the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of whether the violation is compliant; A second input module is used for inputting the text vector corresponding to the real-time live broadcast data into a pre-trained two-stage classification model to obtain a second violation detection result when the first violation detection result is non-compliant, wherein the second violation detection result includes suspected violation or serious violation; The one-stage classification model and the two-stage classification model are obtained by training based on multiple live data samples and the violation detection category label corresponding to each live data sample.
[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the steps of the e-commerce live broadcast violation detection method as described in any one of the embodiments of the first aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the steps of the e-commerce live broadcast violation detection method as described in any embodiment of the first aspect are implemented.
[0010] In a fifth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the e-commerce live broadcast violation detection method provided in the first aspect of the embodiment of the present application.
[0011] The e-commerce live broadcast violation detection method, device, equipment, medium and product of the embodiments of the present application use a pre-trained classification model to divide the violation detection process of the live broadcast data into two stages, where the first stage is used to determine whether the content is compliant, and the second stage is used to determine whether the content is seriously in violation. By filtering out most of the compliant data in the first stage, the computational burden of the second stage is reduced, allowing the model to focus more on processing potential illegal content in the second stage, thereby improving the model detection efficiency and improving the detection accuracy compared to traditional manual detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 It is a flow chart of a method for detecting violations in an e-commerce live broadcast provided in an embodiment of the present application; Figure 2 It is a flowchart of a training method for a one-stage classification model and a two-stage classification model provided in an embodiment of the present application; Figure 3 It is a structural schematic diagram of an e-commerce live broadcast violation detection device provided in an embodiment of the present application; Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application.
[0014] Reference numerals: E-commerce live broadcast violation detection device 300, acquisition module 301, tokenization module 302, first input module 303, second input module 304, Electronic device 400 , processor 401 , memory 402 , communication interface 403 , bus 410 . DETAILED DESCRIPTION
[0015] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0016] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0017] E-commerce live streaming, where merchants (anchors) introduce their products through e-commerce platforms and persuade viewers to buy them, has become a popular form of marketing. However, due to the short time and rapid growth of the e-commerce live streaming industry, it has provided a breeding ground for a series of live streaming chaos, such as false propaganda and the use of inappropriate terms. In addition, due to the huge traffic and information dissemination speed of live streaming e-commerce platforms, false information can spread rapidly, seriously threatening the network information security of users. Therefore, how to effectively supervise and govern these live streaming chaos is of great social significance.
[0018] Video review is an important measure of supervision. An intuitive review strategy is to allow viewers to report video sources containing abnormal content and protect themselves by using block lists. However, the effectiveness of this approach relies on user participation in the system, which may disrupt the normal use of the video platform and affect the user experience. In addition, it may be difficult for viewers to identify some obscure illegal information.
[0019] Another audit strategy is "filter first, then audit". The platform first builds a sensitive word library, and conducts preliminary screening by searching whether the live broadcast text contains illegal information. Then, a dedicated auditor determines whether the content violates the rules. However, this method has the following problems: First, it is difficult to exhaust all sensitive words in advance, resulting in omissions when searching for matching texts; second, the live broadcaster may disguise sensitive words to make the text difficult to detect, such as mixing in pinyin (variant word: K old, original word: anti-aging), splitting words (original word: hospital, variant word: a certain hospital), feature substitution (original word: white coat, variant word: doctor), etc.; third, the number of sensitive words may be large, and there is a problem of low detection efficiency; fourth, the inclusion of sensitive words does not mean that it is sensitive information. For example, although the content contains sensitive words, but the emotional tendency is negative, it is not considered sensitive information. For example, diabetes is a sensitive word, but diabetes cannot be eaten, which is not considered sensitive information.
[0020] In order to solve the problems of related technologies, the embodiments of the present application provide an e-commerce live broadcast violation detection method, device, equipment, medium and product.
[0021] The e-commerce live broadcast violation detection method provided by the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in combination with the accompanying drawings.
[0022] Figure 1 The flowchart of the e-commerce live broadcast violation detection method 100 of the embodiment of the present application is shown. Figure 1 As shown, the e-commerce live broadcast violation detection method 100 may specifically include the following steps: S101, obtaining real-time live broadcast data, wherein the real-time live broadcast data is obtained by performing voice recognition on the obtained real-time live broadcast video, and the data type of the real-time live broadcast data is text data; S102, tokenizing the real-time live broadcast data using a preset model to obtain a corresponding text vector; S103, inputting the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of whether the first violation is compliant; S104: when the first violation detection result is non-compliant, input the text vector corresponding to the real-time live broadcast data into a pre-trained two-stage classification model to obtain a second violation detection result, where the second violation detection result includes suspected violation or serious violation; The one-stage classification model and the two-stage classification model are obtained by training based on multiple live data samples and the violation detection category label corresponding to each live data sample.
[0023] Therefore, the pre-trained classification model is used to divide the violation detection process of live broadcast data into two stages. The first stage is used to determine whether the content is compliant, and the second stage is used to determine whether the content is seriously in violation. By filtering out most of the compliant data in the first stage, the computational burden of the second stage is reduced, allowing the model to focus more on processing potential illegal content in the second stage, thereby improving the model detection efficiency and improving the detection accuracy compared to traditional manual detection.
[0024] The specific implementation methods of the above steps are introduced below.
[0025] In some embodiments, in S101, the real-time live data is obtained by performing voice recognition on the acquired real-time live video, and the data type of the real-time live data is text data. Specifically, by crawling the live content on the live platform, multiple real-time live video files are obtained, and the video segment lengths of the multiple real-time live video files are all preset; the real-time live video files are recognized by the voice recognition model, and the text data corresponding to the real-time live video files are obtained. Among them, the setting of the video segment length should not only ensure that enough information is captured, but also reduce the complexity of data processing to a certain extent, for example but not limited to: each live segment is set to 60 seconds.
[0026] In the specific implementation, the crawler tool is used to obtain the live video. After the crawler tool is running, it will cyclically read the live link pool in the configuration file, open a thread for each live link, and download the live video content in parallel. The live link pool stores the live link currently being crawled, which contains two fields: live name and live link, both of which are text type. In order to ensure the diversity and coverage of the data, different time periods, different anchors, and different types of products (such as health products) are selected for data crawling.
[0027] In the specific implementation, based on the non-autoregressive end-to-end speech recognition model, automatic speech recognition technology is used to transcribe the live video clips into text data and filter them. That is, after the video is transcribed into text information using the speech recognition model, useless data with only background music and other noise is filtered out, and finally multiple high-quality live text data are obtained. Among them, the speech recognition model is trained using a manually annotated Mandarin speech recognition dataset.
[0028] Furthermore, for high-quality text data corresponding to real-time live video files, there may be several variant words in the text data. Although the variant words are similar to the original words in hearing or vision, they can cleverly avoid the recognition of illegal review. Therefore, in this embodiment, the variant word correction model is used to correct the variant words of the live text data. In this way, based on the end-to-end model, by applying the variant word correction technology, the original meaning of the text can be restored and the text data can be corrected.
[0029] Specifically, for the variant word correction model in this embodiment, by fine-tuning on the collected variant word training set, it is able to learn and establish a mapping relationship between variant words and standard vocabulary, and perform end-to-end training and inference directly from the original input to the final output.
[0030] Furthermore, in some embodiments, in S102, the corrected text data is converted into a form understandable by the pre-trained classification model, that is, a text vector The specific implementation method is the same as the specific implementation method of tokenizing the live data sample to obtain the corresponding sample text vector (ie, S202 below) during the model training process, which will be described in detail later.
[0031] In some embodiments, before S103, a one-stage classification model and a two-stage classification model are constructed and trained to detect violations and classify violation levels of live broadcast data. Specifically, the training process of the one-stage classification model and the two-stage classification model provided in the embodiment of the present application can be referred to Figure 2 .
[0032] like Figure 2 As shown, the training method 200 of the one-stage classification model and the two-stage classification model provided in the embodiment of the present application may include the following steps: S201 to S204.
[0033] S201. Obtain multiple live data samples and a violation detection category label corresponding to each live data sample, where the violation detection categories include compliance, suspected violation, and serious violation.
[0034] In some embodiments, multiple live video data are obtained, and the duration of the video segments of the multiple live video data are all preset durations; the live video data are recognized using a speech recognition model to obtain live text data corresponding to the live video data; and the live text data are corrected for variant words using a variant word correction model to obtain a live data sample.
[0035] Furthermore, the live broadcast data samples are categorized and labeled to obtain corresponding violation detection category labels, including compliance, suspected violation, and serious violation. Optionally, manual labeling can be used to divide the corrected live broadcast text content into three clear categories: compliance, suspected violation, and serious violation. During the labeling process, two labelers independently label the content. For content labeled as suspected violation or serious violation, the labelers are required to mark the corresponding violation sentences, which are then reviewed by a third labeler to ensure the consistency and accuracy of the labeling results.
[0036] In another embodiment, a large language model is used to generate a live data sample containing preset violation words, the data type of the live data sample is text data, and the violation detection category label of the live data sample is severe violation.
[0037] Specifically, taking the e-commerce products of health care products as an example, the exemplary large model prompt template is as follows: You are an e-commerce live broadcast salesperson, and you are given a 350-word live broadcast content of health care products. You only need to introduce the product and its efficacy, and do not include any opening and closing words (such as hello, etc.). The content must contain the following phrases: [(target words)]. By inputting known descriptions of violations, such as the mandatory presence of illegal words such as "sales No. 1" (descriptions of violations can be found in official documents), let the model automatically generate new illegal data samples. These generated samples will simulate the language style and expression of real live broadcasts as much as possible, while maintaining the characteristics of illegal behaviors.
[0038] In this way, a high-quality e-commerce live broadcast violation dataset was constructed through technologies such as crawlers, variant word correction, and large model data enhancement, and the live broadcast violation dataset was expanded using a large language model. Violation instances were generated through a large model, which can increase the scale of the dataset and balance the dataset, solving the current problem of scarcity of e-commerce live broadcast violation data.
[0039] S202: tokenize the live broadcast data sample using a preset model to obtain a corresponding sample text vector.
[0040] In some embodiments, the live data sample is defined as , and add special tokens at both ends of the live data sample , ,get ,in is the text length; when the text length of the live data sample is less than the preset length, The live data sample is padded so that the text length is equal to the preset length; when the text length of the live data sample is greater than the preset length, the live data sample is truncated so that the text length is equal to the preset length; each Token in the live data sample is mapped to an integer ID to obtain an ID sequence; the ID sequence is input into the embedding layer of the preset model, and a word vector, a position vector and a segment embedding vector are obtained by word embedding, position embedding and segment embedding respectively, and the word vector, position vector and segment embedding vector are embedded by element-by-element addition to obtain a final sample text vector, and the embedding dimensions of the word vector, position vector and segment embedding vector are the same.
[0041] In the specific implementation, taking the preset length (i.e., the maximum input length) of 512 as an example, each Token in the live data sample is mapped to an integer ID through Tokenizer to obtain an ID sequence: , the dimension is (512).
[0042] In specific implementation, Input to the embedding layer, each input token is converted into a 768-dimensional vector through three independent embedding processes to obtain the corresponding word vector , The dimension is (512,768), where correspond The word vector of the sample text is (1,768). The embedding layer is a technique for mapping discrete inputs (such as words, characters, or other categorical variables) to a continuous vector space. The input token is embedded in the word, position, and segment respectively to obtain the word vector, position vector, and segment embedding vector. The embedding dimensions of the three vectors are all 768. The embedding vectors are fused by element-by-element addition to obtain the final embedding vector (i.e., the sample text vector). This mapping allows the model to understand the semantic relationship between input elements and capture the implicit information in the text.
[0043] S203, taking the sample text vector corresponding to the live broadcast data sample as the input of the preset model, taking the violation detection category label as the output of the preset model, training the preset model, wherein the preset model includes a one-stage classification model and a two-stage classification model.
[0044] In some embodiments, a plurality of live broadcast data samples are divided according to a preset ratio to obtain a training set and a validation set; a sample text vector corresponding to the training set is input into the preset model for training, and then a sample text vector corresponding to the validation set is input into the preset model for verification, and the preset model is trained with the violation detection category label as the output of the preset model.
[0045] S204. Determine whether at least one of the iterative training times or the verification loss value of the preset model meets the preset training stop condition. If not, adjust the weight parameters of the one-stage classification model and the two-stage classification model respectively, and train the adjusted one-stage classification model and the two-stage classification model respectively until the preset training stop condition is met, and determine the one-stage classification model and the two-stage classification model with the highest value as the trained one-stage classification model and the two-stage classification model.
[0046] Optionally, the preset training stop condition includes at least one of the following: the number of iterative training reaches a first preset threshold, and the number of times the verification loss value remains continuously without decreasing reaches a second preset threshold.
[0047] The verification loss value is obtained based on the verification set and the loss function.
[0048] The loss function is: ; (1) ; (2) ; (3) ; (4) in, Indicates the number of live data samples; Indicates Violation detection category labels for live data samples; Indicates live sample The corresponding label probability distribution; express The corresponding sample text vector; represents the weight matrix dimension, Indicates the bias vector dimension.
[0049] As an optional embodiment, the two stages are trained separately using a pre-trained language model BERT plus a fully connected network, and the BERT model is fine-tuned using the constructed sample data set to obtain two trained classification models, and the best model parameters are saved. The sample data set includes multiple live data samples and the violation detection category label corresponding to each live data sample.
[0050] In specific implementation, the word vector Input into the BERT model, input into the bidirectional Transformer encoder, output a 768-dimensional vector at each position, and take out The corresponding vector As a vector representation of live content.
[0051] Optionally, the BERT model is a multi-layer bidirectional Transformer encoder structure, including 12 layers of Transformer encoders, 12 hidden layers, 12 attention heads, and a parameter size of 110M. Compared with unidirectional language models, bidirectional language models can better understand the semantics and context of text when performing text classification tasks. In the BERT model, each input token interacts with other tokens through a multi-layer self-attention mechanism. After being processed by the multi-layer Transformer encoder, it is able to capture the global information of the entire sequence. This global representation is crucial for understanding the overall meaning of the text.
[0052] When implementing it, use Representation As the input of the fully connected network, the label probability distribution is obtained. The input of the fully connected network is a 768-dimensional vector, and the output is a 1-dimensional vector .
[0053] ; (4) in, is the learnable parameter of the fully connected layer, The weight matrix dimension is (768,1), The bias vector dimension is (1,).
[0054] In specific implementation, the fully connected network output is passed through the activation function Then we get the probability distribution of the label, namely: ; (2) . (3) Taking the first stage as an example, if If the text is not in compliance, it is judged as non-compliant; otherwise, it is compliant. For non-compliant texts, the second stage is entered to further determine whether it is a serious violation.
[0055] In specific implementation, the sample data set can be used to construct a training set and a validation set in a ratio of 9:1, and the model can be trained on the training set to calculate the model loss. , perform back propagation, update the model's weight parameters, and perform iterative operations. Loss function The definition is as follows: ; (1) in, Indicates the number of live data samples; Indicates Violation detection category labels for live broadcast data samples.
[0056] Optionally, in this example the initial learning rate is set to , using the Adam optimizer. The Adam optimizer is an adaptive learning rate optimization algorithm that can effectively adjust the learning rate of each parameter. The parameter update process of the Adam optimizer is as follows: ; (5) in, is the initial learning rate, and Consistency; is a very small constant, let ; is the bias-corrected first-order moment (gradient mean) estimate; is the bias-corrected second-order moment estimate (the uncentered variance of the gradient); and are the two initial hyperparameters of the Adam optimizer, which change with the step size , which represent the first-order moment estimation and the second-order matrix estimation respectively. .in and The step length Status and These two parameters together determine the degree to which the optimizer relies on historical gradient information when updating model parameters, thus affecting the convergence speed and stability of the model.
[0057] In the specific implementation, a classification model is trained for each stage. After completing an epoch, the respective models are selected according to their performance on the validation set. The model with the highest value is taken as the final model . The definition is as follows: ; (6) in, is a positive example that is correctly predicted; is a counterexample that is correctly predicted; is a positive example that was predicted incorrectly; For the counterexample that was predicted incorrectly, the number of iterations epoch = 10.
[0058] The above is a specific implementation method of the one-stage classification model and the two-stage classification model obtained by training provided in the embodiment of the present application, which can improve the classification ability of the model. Therefore, the live broadcast content can be automatically classified by using the trained model connected to the fully connected network, which can improve the efficiency of traditional manual classification and reduce the influence of human subjectivity on the accuracy of the classification results.
[0059] Back to Figure 1 The e-commerce live broadcast violation detection method 100 is shown.
[0060] In some embodiments, in S103, the text vector to be classified corresponding to the real-time live data is Input to the trained one-stage classification model In the output , which is the first violation detection result. Its definition is as follows: ; (7) If the model output If it is 0, it is compliant; otherwise, further execute S104.
[0061] In some embodiments, in S104, the text vector to be classified is Input to the trained two-stage classification model In the output , which is the second violation detection result. Its definition is as follows: ; (8) If the model output If it is 0, it is suspected of violating the rules; otherwise, it is a serious violation. It should be understood that for live broadcast data suspected of violating the rules, further manual judgment is required to determine whether it is a violation.
[0062] In addition, in order to verify the performance of the e-commerce live broadcast violation detection method of the embodiment of the present application, the superiority of the method is illustrated below through automatic indicator evaluation.
[0063] Specifically, the evaluation is performed by using test sets from different sources: test set 1 is data from the same live link pool as the training set, that is, it is obtained from the same live broadcast room as the training set, and is the same-source live broadcast data. There is no duplication between its data and the training set verification set, and a total of 100 samples are manually labeled, of which 50 are compliant, 25 are suspected of violating regulations, and 25 are seriously violating regulations. Each sample contains a live text data and label information; test set 2 is a new live link pool, which has no overlap with the live link pool in the training set, and a total of 300 samples, of which 204 are compliant, 82 are suspected of violating regulations, and 14 are seriously violating regulations. Each sample contains a live text data and label information to further verify the dissemination of the algorithm in this embodiment.
[0064] This method uses The value is evaluated. It is defined as follows: ; (6) ; (9) ; (10) ; (11) in, is a positive example that is correctly predicted; is a counterexample that is correctly predicted; is a positive example that was predicted incorrectly; is a counterexample that was predicted incorrectly.
[0065] Contrast this approach with the following: (Large Language Model), (prompt learning), -3 (using BERT to directly perform three-class classification), the following Tables 1 and 2 show the comparative experimental results of test set 1 and test set 2 respectively.
[0066] Table 1 method Acc Precision Recall F1 LLM 0.460 0.483 0.453 0.415 Prompt-learning 0.830 0.803 0.800 0.799 BERT-3 0.790 0.776 0.767 0.768 This embodiment 0.900 0.883 0.880 0.881 Table 2 method Acc Precision Recall F1 LLM 0.373 0.450 0.459 0.346 Prompt-learning 0.883 0.831 0.813 0.822 BERT-3 0.840 0.732 0.777 0.750 This embodiment 0.897 0.833 0.879 0.851 As shown in Table 1 and Table 2, the method of this embodiment achieves the highest This method reduces the computational burden of the second stage by filtering out most of the compliant data in the first stage, so that the model can focus more on processing potential illegal content in the second stage, thereby improving the classification ability of the model, and then improving the efficiency of violation detection, and improving the detection accuracy, solving the current problem of difficulty in e-commerce violation detection.
[0067] It should be noted that the above describes some embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0068] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an e-commerce live broadcast violation detection device 300.
[0069] like Figure 3 As shown, the e-commerce live broadcast violation detection device 300 may include: The acquisition module 301 is used to acquire real-time live broadcast data, wherein the real-time live broadcast data is obtained by performing voice recognition on the acquired real-time live broadcast video, and the data type of the real-time live broadcast data is text data; A tokenization module 302 is used to tokenize the real-time live broadcast data using a preset model to obtain a corresponding text vector; The first input module 303 is used to input the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of whether the first violation is compliant; A second input module 304 is used to input the text vector corresponding to the real-time live broadcast data into a pre-trained two-stage classification model to obtain a second violation detection result when the first violation detection result is non-compliant, wherein the second violation detection result includes suspected violation or serious violation; The one-stage classification model and the two-stage classification model are obtained by training based on multiple live data samples and the violation detection category label corresponding to each live data sample.
[0070] In some embodiments, the e-commerce live broadcast violation detection device 300 further includes a training module ( Figure 3 Specifically, the training module includes the following units: An acquisition unit, used to acquire a plurality of live data samples and a violation detection category label corresponding to each live data sample, wherein the violation detection category includes compliance, suspected violation, and serious violation; A tokenization unit, used to tokenize the live broadcast data sample using a preset model to obtain a corresponding sample text vector; A training unit, used to use the sample text vector corresponding to the live data sample as the input of a preset model and the violation detection category label as the output of the preset model to train the preset model, wherein the preset model includes a one-stage classification model and a two-stage classification model; An adjustment unit is used to determine whether at least one of the iterative training times or the verification loss value of the preset model meets the preset training stop condition. If not, the weight parameters of the one-stage classification model and the two-stage classification model are adjusted respectively, and the adjusted one-stage classification model and the two-stage classification model are trained respectively until the preset training stop condition is met, and the one with the highest value is determined as the trained one-stage classification model and the two-stage classification model.
[0071] In some optional embodiments, the acquisition unit is specifically used to: acquire multiple live video data, each of which has a preset duration; identify the live video data using a speech recognition model to obtain live text data corresponding to the live video data; perform variant word correction on the live text data using a variant word correction model to obtain live data samples; and perform category labeling on the live data samples to obtain corresponding violation detection category labels, wherein the categories include compliance, suspected violation, and serious violation.
[0072] In some optional embodiments, the acquisition unit is further specifically used to: use a large language model to generate a live data sample containing preset violation words, the data type of the live data sample is text data, and the violation detection category label of the live data sample is severe violation.
[0073] In some optional embodiments, the tokenization unit is specifically used to: define the live data sample as , and add special tokens at both ends of the live data sample , ,get ,in is the text length; when the text length of the live data sample is less than the preset length, The live data sample is padded so that the text length is equal to the preset length; when the text length of the live data sample is greater than the preset length, the live data sample is truncated so that the text length is equal to the preset length; each Token in the live data sample is mapped to an integer ID to obtain an ID sequence; the ID sequence is input into the embedding layer of the preset model, and a word vector, a position vector and a segment embedding vector are obtained by word embedding, position embedding and segment embedding respectively, and the word vector, position vector and segment embedding vector are embedded by element-by-element addition to obtain a final sample text vector, and the embedding dimensions of the word vector, position vector and segment embedding vector are the same.
[0074] In some optional embodiments, the training unit is specifically used to: divide the multiple live data samples according to a preset ratio to obtain a training set and a verification set; input the sample text vector corresponding to the training set into the preset model for training, and then input the sample text vector corresponding to the verification set into the preset model for verification, and train the preset model with the violation detection category label as the output of the preset model.
[0075] The verification loss value is obtained based on the verification set and the loss function.
[0076] Optionally, the loss function is: ; ; ; ; in, Indicates the number of live data samples; Indicates Violation detection category labels for live data samples; Indicates live sample The corresponding label probability distribution; express The corresponding sample text vector; represents the weight matrix dimension, Indicates the bias vector dimension.
[0077] It should be noted that, for the convenience of description, the above device is described in various modules according to their functions. Of course, when implementing this application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0078] The device of the above-mentioned embodiment is used to implement the corresponding e-commerce live broadcast violation detection method in any of the aforementioned embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.
[0079] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present application also provides an electronic device.
[0080] Figure 4 A more specific schematic diagram of the hardware structure of an electronic device provided by this embodiment is shown.
[0081] The electronic device 400 may include a processor 401 and a memory 402 storing computer program instructions.
[0082] Specifically, the processor 401 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0083] The memory 402 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 402 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 402 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 402 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 402 is a non-volatile solid-state memory.
[0084] In certain embodiments, the memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present application.
[0085] The processor 401 implements any one of the e-commerce live broadcast violation detection methods in the above embodiments by reading and executing computer program instructions stored in the memory 402.
[0086] In some examples, the electronic device 400 may further include a communication interface 403 and a bus 410. Figure 4As shown, the processor 401, the memory 402, and the communication interface 403 are connected via a bus 410 and communicate with each other.
[0087] The communication interface 403 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0088] Bus 410 includes hardware, software or both, and couples the components of online data traffic billing equipment to each other. For example, but not limitation, bus 410 may include accelerated graphics port (AGP) or other graphics bus, enhanced industry standard architecture (EISA) bus, front-side bus (FSB), hypertransport (HT) interconnection, industry standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnect (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. Where appropriate, bus 410 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the present application considers any suitable bus or interconnection.
[0089] Exemplarily, the electronic device 400 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0090] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a non-transitory computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by the processor, any of the e-commerce live broadcast violation detection methods in the above-mentioned embodiments is implemented. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, etc.
[0091] Based on the same technical concept, corresponding to any of the above-mentioned embodiments, the present application also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer so that the computer and / or the processor execute the e-commerce live broadcast violation detection method. Corresponding to the execution subject corresponding to each step in each embodiment of the e-commerce live broadcast violation detection method, the processor that executes the corresponding step may belong to the corresponding execution subject.
[0092] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0093] The functional blocks shown in the structural block diagram described above can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier. "Machine-readable medium" may include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0094] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0095] The above describes various aspects of the present application with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the functions / actions specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0096] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A method for detecting violations in e-commerce live streaming, characterized in that: include: Acquire real-time live broadcast data, where the real-time live broadcast data is obtained by performing voice recognition on the acquired real-time live broadcast video, and the data type of the real-time live broadcast data is text data; Using a preset model to tokenize the real-time live broadcast data to obtain a corresponding text vector; Inputting the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of whether the first violation is compliant; In the case where the first violation detection result is non-compliant, inputting the text vector corresponding to the real-time live broadcast data into a pre-trained two-stage classification model to obtain a second violation detection result, wherein the second violation detection result includes suspected violation or serious violation; The one-stage classification model and the two-stage classification model are obtained by training based on multiple live data samples and the violation detection category label corresponding to each live data sample.
2. The method according to claim 1, characterized in that Before inputting the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of compliance, the method further includes: Obtain multiple live data samples and a violation detection category label corresponding to each live data sample, wherein the violation detection category includes compliance, suspected violation, and serious violation; Tokenizing the live broadcast data sample using a preset model to obtain a corresponding sample text vector; The sample text vector corresponding to the live data sample is used as the input of the preset model, and the violation detection category label is used as the output of the preset model to train the preset model, wherein the preset model includes a one-stage classification model and a two-stage classification model; Determine whether at least one of the iterative training times or the validation loss value of the preset model meets the preset training stop condition. If not, adjust the weight parameters of the one-stage classification model and the two-stage classification model respectively, and train the adjusted one-stage classification model and the two-stage classification model respectively until the preset training stop condition is met. The one-stage classification model and the two-stage classification model after training are determined to have the highest value.
3. The method according to claim 2, characterized in that The obtaining of multiple live data samples and a violation detection category label corresponding to each live data sample includes: Acquire multiple live video data, wherein the duration of video segments of the multiple live video data are all preset durations; Recognize the live video data using a speech recognition model to obtain live text data corresponding to the live video data; Performing variant word correction on the live text data using a variant word correction model to obtain a live data sample; The live broadcast data samples are categorized and labeled to obtain corresponding violation detection category labels, where the categories include compliance, suspected violation, and serious violation.
4. The method according to claim 3, characterized in that: The obtaining of multiple live data samples and a violation detection category label corresponding to each live data sample also includes: A live broadcast data sample containing preset violation words is generated by using a large language model, the data type of the live broadcast data sample is text data, and the violation detection category label of the live broadcast data sample is severe violation.
5. The method according to claim 2, characterized in that: The method of tokenizing the live broadcast data sample by using a preset model to obtain a corresponding sample text vector includes: The live data sample is defined as , and add special tokens at both ends of the live data sample , ,get ,in is the length of the text; When the text length of the live data sample is less than the preset length, The live data sample is padded so that the text length is equal to the preset length; if the text length of the live data sample is greater than the preset length, the live data sample is truncated so that the text length is equal to the preset length; Map each Token in the live data sample to an integer ID to obtain an ID sequence; The ID sequence is input into the embedding layer of the preset model, and the word vector, position vector and segment embedding vector are obtained through word embedding, position embedding and segment embedding respectively, and the word vector, position vector and segment embedding vector are fused by element-by-element addition to obtain the final sample text vector, and the embedding dimensions of the word vector, position vector and segment embedding vector are the same.
6. The method according to claim 5, characterized in that The method of taking the sample text vector corresponding to the live broadcast data sample as the input of the preset model and taking the violation detection category label as the output of the preset model to train the preset model includes: Dividing the plurality of live data samples according to a preset ratio to obtain a training set and a validation set; Inputting the sample text vector corresponding to the training set into the preset model for training, then inputting the sample text vector corresponding to the verification set into the preset model for verification, and training the preset model with the violation detection category label as the output of the preset model; The validation loss value is obtained based on the validation set and the loss function; The loss function is: ; ; ; ; in, Indicates the number of live data samples; Indicates Violation detection category labels for live data samples; Indicates live sample The corresponding label probability distribution; express The corresponding sample text vector; represents the weight matrix dimension, Indicates the bias vector dimension.
7. An e-commerce live broadcast violation detection device, characterized in that: The device comprises: An acquisition module is used to acquire real-time live broadcast data, wherein the real-time live broadcast data is obtained by performing voice recognition on the acquired real-time live broadcast video, and the data type of the real-time live broadcast data is text data; A tokenization module, used to tokenize the real-time live broadcast data using a preset model to obtain a corresponding text vector; A first input module, used to input the text vector corresponding to the real-time live broadcast data into a pre-trained one-stage classification model to obtain a first violation detection result of whether the violation is compliant; A second input module is used for inputting the text vector corresponding to the real-time live broadcast data into a pre-trained two-stage classification model to obtain a second violation detection result when the first violation detection result is non-compliant, wherein the second violation detection result includes suspected violation or serious violation; The one-stage classification model and the two-stage classification model are obtained by training based on multiple live data samples and the violation detection category label corresponding to each live data sample.
8. An electronic device, characterized in that: The device includes: a processor and a memory storing computer program instructions; when the processor calls the computer program instructions, it implements the e-commerce live broadcast violation detection method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, which, when called by the processor, implement the e-commerce live broadcast violation detection method as described in any one of claims 1-6.
10. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the e-commerce live broadcast violation detection method as described in any one of claims 1-6.
Citation Information
Patent Citations
Data classification method, device and equipment for staged quality inspection, and storage medium
CN112668857A
Voice quality detection method and device, computer equipment and storage medium
CN112669850A
Financial live broadcast violation detection method, device and equipment and readable storage medium
CN113038153A
Method and device for processing network live broadcast data
CN113949887A
Illegal video detection method based on text and video fusion
CN115775363A
Cited By
Real-time data-based live broadcast content risk control processing system and supervision method
CN120151568A
Live broadcast information analysis method based on knowledge graph
CN120429448A
E-commerce competition violation identification system based on distributed crawler and semantic analysis
CN122451196A