Playback data processing method and medium

By obtaining and standardizing the multi-dimensional playback data sets, combined with time window analysis, the problem of low malicious information detection efficiency in online education platforms is solved, more efficient and flexible malicious information detection is achieved, and the platform security is improved.

CN111985345BActive Publication Date: 2025-08-26TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010730839.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-27
Publication Date
2025-08-26
Estimated Expiration
2040-07-27

AI Technical Summary

Technical Problem

The existing online education platform is inefficient in detecting malicious information in playback services and relies on manual review, so it is unable to determine malicious attributes in a timely and effective manner.

Method used

By obtaining play data sets of at least two dimensions, performing standardization and adding preset tags, combined with time window analysis, the target data set is extracted to determine the malicious properties of the play service.

Benefits of technology

It realizes a more efficient and flexible solution for detecting malicious information in playback services, improves the accuracy and efficiency of data processing, and enhances the security of playback platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111985345B_ABST
    Figure CN111985345B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and medium for processing playback data. The method includes: obtaining a playback data set of at least two dimensions; processing the playback data set of at least two dimensions into a standard data set; adding at least one label to each standard data in the standard data set based on a preset label set; determining a target time period, and extracting a target data set corresponding to the target time period from the standard data set after adding labels; and determining the malicious attributes of the corresponding playback service based on the label carried by each target data in the target data set. The present invention provides a more efficient and flexible solution for detecting malicious information on playback services. Taking into account the different dimensions of playback data, playback data of at least two dimensions are standardized. The malicious attributes of the playback service are judged in combination with the time window, thereby improving the accuracy and efficiency of data processing and improving the security of the playback platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet communication technology, and in particular to a playback data processing method and medium. Background Art

[0002] With the development of internet communication technology, online education has emerged. It now accounts for an increasingly significant portion of today's education landscape. Compared to traditional education, online education offers significant advantages, such as flexibility, convenience, and ease of management. Furthermore, it offers significant time savings, eliminating the need to worry about missing a class.

[0003] As a key form of online education, live streaming focuses on the interactive effects of courses. Since live educational content is presented to users in real time, it's inevitable that malicious actors could exploit it to spread malicious content. Related technologies often rely on manual review by online education platforms to detect such malicious content. This heavily relies on manual effort and fails to effectively and timely determine the malicious nature of streaming services. Therefore, a more timely and effective solution for detecting malicious content in streaming services is needed. Summary of the Invention

[0004] In order to solve the problems of low efficiency and high labor cost in the application of existing technologies for malicious information detection in playback services, the present invention provides a playback data processing method and medium:

[0005] According to one aspect of the present application, a method for processing playback data is provided, the method comprising:

[0006] Obtain a playback dataset of at least two dimensions;

[0007] Processing the playback data set of at least two dimensions into a standard data set;

[0008] Add at least one label to each standard data in the standard data set based on a preset label set;

[0009] Determine a target time period, and extract a target data set corresponding to the target time period from the labeled standard data set;

[0010] The malicious attribute of the corresponding playback service is determined based on the label carried by each target data in the target data set.

[0011] According to another aspect of the present application, a playback data processing device is provided, the device comprising:

[0012] Acquisition module: used to obtain playback data sets of at least two dimensions;

[0013] A processing module: configured to process the playback data set of at least two dimensions into a standard data set;

[0014] Adding module: used to add at least one label to each standard data in the standard data set based on a preset label set;

[0015] Extraction module: used to determine the target time period and extract the target data set corresponding to the target time period from the labeled standard data set;

[0016] Determination module: used for determining the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set.

[0017] According to another aspect of the present application, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the playback data processing method as described above.

[0018] According to another aspect of the present application, a computer-readable storage medium is provided, in which at least one instruction or at least one program is stored. The at least one instruction or the at least one program is loaded and executed by a processor to implement the playback data processing method as described above.

[0019] According to another aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described playback data processing method.

[0020] The present invention provides a method and medium for processing playback data, which has the following technical effects:

[0021] This invention provides a more efficient and flexible solution for detecting malicious information in playback services. Taking into account the different dimensions of playback data, the solution standardizes playback data across at least two dimensions. The solution also uses a time window to determine the malicious nature of playback services, improving the accuracy and efficiency of data processing and enhancing the security of the playback platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 is a schematic diagram of an application environment provided by an embodiment of the present invention;

[0024] Figure 2 1 is a flowchart of a playback data processing method provided by an embodiment of the present invention;

[0025] Figure 3 This is a flow chart of processing the at least two-dimensional playback data set into a standard data set provided by an embodiment of the present invention;

[0026] Figure 4 It is also a schematic diagram of a process for processing the at least two-dimensional playback data set into a standard data set provided by an embodiment of the present invention;

[0027] Figure 5 It is also a flowchart of a playback data processing method provided by an embodiment of the present invention;

[0028] Figure 6 is a schematic diagram of data processing using a (time) sliding window provided by an embodiment of the present invention;

[0029] Figure 7 This is a block diagram of a playback data processing device provided by an embodiment of the present invention;

[0030] Figure 8 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0033] See also Figure 1 , Figure 1 This is a schematic diagram of an application environment provided by an embodiment of the present invention, which may include a first-type client 01, a server 02, and second-type clients 031, 032 to 03n. The clients (including first-type and second-type clients) and the server may be connected directly or indirectly via wired or wireless communication. The first-type client and the second-type client exchange information, and the server determines the malicious attributes of the corresponding business service based on the relevant information. It should be noted that Figure 1 This is just an example, and the application environment may also include a server and a client of the same type.

[0034] Clients may include physical devices such as smartphones, desktop computers, tablets, laptops, augmented reality (AR) / virtual reality (VR) devices, digital assistants, smart speakers, and smart wearable devices. They may also include software running on physical devices, such as computer programs. Operating systems running on client 01 may include, but are not limited to, Android, iOS (a mobile operating system developed by Apple), Linux, and Microsoft Windows.

[0035] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can include network communication units, processors, and memory, among other things. The server can provide backend services for the aforementioned clients.

[0036] In cloud computing applications, Cloud Computing Education (CCEDU) refers to an educational platform service based on a cloud computing business model. On this cloud platform, all educational institutions, training programs, enrollment services, publicity agencies, industry associations, regulatory bodies, industry media, and legal structures are integrated into a centralized resource pool. These resources interact and display with each other, enabling on-demand communication and reaching consensus, thereby reducing educational costs and improving efficiency.

[0037] In practical applications, the first type of client can be a live streamer client, and the second type of client can be a viewer client. A live streamer client can establish a live streaming room (room, live broadcast channel; live streaming: the simultaneous production and release of information as events occur and develop, a two-way information network publishing method) online to share live video with viewer clients connected to the room. Accordingly, the server determines the malicious attributes of the corresponding live streaming service (the playback service) based on the live streaming data (the playback data). Live streaming service types include, but are not limited to, educational live streaming, gaming live streaming, e-commerce live streaming, and show live streaming.

[0038] Compared to live streaming services, playback services can also be on-demand services. Specifically, live streaming refers to the simultaneous production and broadcast of videos as events occur and develop, with production and broadcast occurring simultaneously. On the other hand, on-demand video refers to the playback of completed videos based on user demand, with production and broadcast occurring at different times.

[0039] The following describes a specific embodiment of a playback data processing method of the present invention. Figure 2 、 5 It is a flowchart of a playback data processing method provided by an embodiment of the present invention. This specification provides the method operation steps described in the embodiment or flowchart, but may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 2 、 5 As shown, the method may include:

[0040] S201: Obtain a playback data set of at least two dimensions;

[0041] In an embodiment of the present invention, a playback dataset can refer to an ongoing live broadcast in a live studio (corresponding to a live broadcast dataset), or to a completed live broadcast in a live studio (corresponding to a video-on-demand dataset). A playback dataset can correspond to a data source within a specific live broadcast session. Live streaming platforms often provide a rich set of interactive tools, including support for audio / video live scenes, screen sharing, courseware sharing, electronic whiteboards, and interactive communication. Courseware formats include, but are not limited to, .ppt (corresponding to presentations), .pdf (portable document format), and .png (a bitmap format using a lossless compression algorithm). Interactive communication formats include, but are not limited to, bullet screens, chat rooms, attendance roll calls / sign-ins, quizzes, Q&A roll calls / active speaking. The data involved in these interactions includes, but is not limited to, video data, audio data, image data (picture data), text data, and behavioral data.

[0042] At least two dimensions can indicate at least two natural language dimensions, such as a video corpus dimension, an audio corpus dimension, an image corpus dimension, and a text corpus dimension. The at least two dimensions include a natural language dimension and a user operation behavior dimension. The natural language dimension here can include at least one selected from the group consisting of a video corpus dimension, an audio corpus dimension, an image corpus dimension, and a text corpus dimension. Of course, the at least two dimensions can also include a dimension indicating the device where the client is located. The playback data of this dimension can be relevant information about the device where the client is located, such as IP address (Internet Protocol Address) information and operating system information. The server obtains playback data sets of at least two dimensions. The sources of playback data corresponding to the same dimension can be different. For example, audio data (corresponding to the audio corpus dimension) can come from audio reality (audio track in video reality), chat room (such as voice), and courseware (such as sound effects).

[0043] For a live broadcast, the server can obtain live datasets of at least two dimensions (as playback datasets) in real time, after a target interval, or offline for on-demand playback of at least two dimensions (as playback datasets). In other words, the playback datasets of at least two dimensions point to the playback data stream of the live broadcast. The playback datasets of at least two dimensions can correspond to the entire playback data stream or a portion of it.

[0044] S202: Processing the playback data set of at least two dimensions into a standard data set;

[0045] In an embodiment of the present invention, considering the complex strategies involved in subsequently determining malicious attributes when using playback data corresponding to different dimensions, the concept of standards is introduced, processing playback data sets of at least two dimensions into standard datasets. This converts playback data corresponding to different dimensions into standard data under the same standard, enabling rapid and accurate malicious attribute determination based on simpler strategies. Standard datasets can correspond to the same data format, such as a string or vector. Standard datasets can also correspond to the same dimension, such as a natural language dimension or a user operation behavior dimension.

[0046] In a specific embodiment, Figure 3 As shown, when the at least two dimensions include at least two natural language dimensions, processing the playback data sets of the at least two dimensions into a standard data set includes:

[0047] S301: Selecting a natural language dimension with a higher information expression concentration from the at least two natural language dimensions as a reference dimension;

[0048] S302: Divide the playback data set of the at least two dimensions into a first data set and a second data set based on the reference dimension, where the first data set corresponds to the reference dimension;

[0049] S303: Processing the second data set into a third data set corresponding to the reference dimension;

[0050] S304: Obtain the standard dataset based on the first dataset and the third dataset.

[0051] Information expression concentration indicates whether the functions of the channels used to express information in the natural language dimension are concentrated and single. The fewer functions a channel has, the higher the information expression concentration, while the more functions a channel has, the lower the information expression concentration. For example, the image channel (corresponding to the image corpus dimension) can be used to display shapes, patterns, colors, and other objects, as well as text. The text channel (corresponding to the text corpus dimension) is generally used only for text. Accordingly, the information expression concentration of the image corpus dimension is lower than that of the text dimension.

[0052] A natural language dimension with a high concentration of information expression, such as a text corpus dimension, is used as a reference dimension. The playback data corresponding to the reference dimension can be used to construct a first dataset, while the remaining playback data can be used to construct a second dataset. The second dataset is processed into a third dataset corresponding to the reference dimension, and then the first and third datasets are used to construct a standard dataset. This standard dataset corresponds to the same natural language dimension.

[0053] In the application of generating the third data set, you can refer to the following execution process:

[0054] (1) When the at least two natural language dimensions include a video corpus dimension and an image corpus dimension, since the information expression concentration of the video corpus dimension is less than that of the image corpus dimension, the image corpus dimension with a higher information expression concentration is selected as the reference dimension from the video corpus dimension and the image corpus dimension. Then, the first data set is the playback data corresponding to the image corpus dimension, that is, the first data set is an image data set. The second data set is the playback data corresponding to the video corpus dimension, that is, the second data set is a video data stream.

[0055] Processing the second data set into a third data set corresponding to the reference dimension includes: intercepting a video image from the video data stream based on a preset time interval; and obtaining the third data set based on the intercepted video image.

[0056] The preset time interval can be a threshold for taking screenshots. When the threshold is 1 minute, a dynamic screenshot is taken every minute of the video stream corresponding to the previous minute to obtain the corresponding video image. If the video data stream is 10 minutes long, 10 video images can be obtained through dynamic screenshots. For high-dimensional video data, dynamic screenshots can be used to reduce the dimensionality of dynamic screenshots by setting a threshold, reducing the continuous video data into independent image data.

[0057] (2) When the at least two natural language dimensions include an image corpus dimension and a text corpus dimension, since the information expression concentration of the image corpus dimension is lower than that of the text corpus dimension, the text corpus dimension with a higher information expression concentration is selected from the image corpus dimension and the text corpus dimension as the reference dimension. Then, the first data set is the playback data corresponding to the text corpus dimension, that is, the first data set is a text data set. The second data set is the playback data corresponding to the image corpus dimension, that is, the second data set is an image data set.

[0058] Processing the second data set into a third data set corresponding to the reference dimension includes: extracting text data from each image data in the image data set using character recognition technology; and obtaining the third data set based on the text data extracted from each image data. Optical Character Recognition (OCR) technology can be used to reduce the dimensionality of the image data into text data.

[0059] (3) When the at least two natural language dimensions include an audio corpus dimension and a text corpus dimension, since the information expression concentration of the audio corpus dimension is lower than that of the text corpus dimension, the text corpus dimension with a higher information expression concentration is selected from the audio corpus dimension and the text corpus dimension as the reference dimension. Then, the first data set is the playback data corresponding to the text corpus dimension, that is, the first data set is a text data set. The second data set is the playback data corresponding to the audio corpus dimension, that is, the second data set is an audio data stream.

[0060] Processing the second data set into a third data set corresponding to the reference dimension includes: processing the audio data stream using speech recognition technology to obtain corresponding text data, and using the corresponding text data as the third data set. The audio data is processed into text data based on Automatic Speech Recognition (ASR) technology.

[0061] In another specific embodiment, Figure 4 As shown, when the at least two dimensions include a natural language dimension and a user operation behavior dimension, processing the playback data set of the at least two dimensions into a standard data set includes:

[0062] S401: When the natural language dimension is a preset natural language dimension, obtaining first type of standard data based on playback data corresponding to the natural language dimension;

[0063] S402: When the natural language dimension is not the preset natural language dimension, processing the playback data corresponding to the natural language dimension into playback data corresponding to the preset natural language dimension, and obtaining second-type standard data based on the processed playback data;

[0064] S403: extracting behavior feature data from the playback data corresponding to the user operation behavior dimension, and obtaining third-category standard data based on the behavior feature data;

[0065] S404: Obtain the standard data set based on the first type of standard data, the second type of standard data, and the third type of standard data;

[0066] When standardizing playback data corresponding to the natural language dimension and the playback data corresponding to the user operation behavior dimension, it is necessary not only to consider that the standard data sets correspond to the same data format (the first-category standard data, the second-category standard data, and the third-category standard data correspond to the same data format), but also to consider the preservation of relevant features of each dimension during the processing process.

[0067] When processing playback data corresponding to a natural language dimension, a preset natural language dimension is first determined (the preset natural language dimension often corresponds to a natural language dimension with a higher concentration of information expression). Then, 1) if the dimension matches, the playback data to be processed can be directly processed into first-category standard data in the target data format; 2) if the dimension does not match, the playback data to be processed can be dimensionality-reduced to playback data corresponding to the preset natural language dimension, and then processed into second-category standard data in the target data format. In practical applications, for example, the preset natural language dimension is a text corpus dimension, and a high-dimensional dataset (which may include video data streams, audio data streams, and image datasets) is dimensionality-reduced to a text dataset. A word segmentation preprocessing scheme can be used to filter non-keywords from the text dataset obtained from the dimensionality reduction process and the original text dataset. The data source of the original text dataset can be the host client and all viewer clients, or just the host client. The data source of the original video data stream, audio data stream, and image dataset can also be the host client and all viewer clients, or just the host client. Keywords can be words corresponding to tags in a preset tag set (see the relevant description of step S203 below).

[0068] When processing playback data corresponding to the user operation behavior dimension, behavioral feature data can be first extracted from the playback data corresponding to the user operation behavior dimension. The behavioral feature data can include basic data and extended data associated with the basic data. The basic data can point to specific operation events, such as click events, touch events, sliding events, etc. The basic data can also include specific object information in the above operation events (such as "button" as the click object) and condition information (such as the click condition "determining whether two click operations on the click object were detected within a preset time period"). The extended data can point to the operation path (recording multiple consecutive operation events) and operation feedback (such as page jumps triggered by "click events"). Then, the behavioral feature data is processed into third-category standard data in the form of target data. The data source of the original playback data corresponding to the user operation behavior dimension can be the host client and all viewer clients, or only the host client.

[0069] In practical applications, playback data related to user behavior focuses on the path, operation, and feedback, which can be captured through tracking and monitoring. User signaling can be the primary source of playback data related to user behavior, and relevant signaling data can be obtained through pre-tracking on the front-end and client. Signaling data can indicate frequent changes in IP addresses, frequent switching between cameras, PowerPoint presentations, and other tools.

[0070] S203: Add at least one label to each standard data in the standard data set based on a preset label set;

[0071] In an embodiment of the present invention, the labels in the preset label set can indicate content that does not comply with laws and administrative regulations, content that violates social morality, and content that harms national interests, social public interests, or the interests of third parties. Content that harms social public interests and third-party interests can include harassment, fraudulent content, advertising and promotional content, and so on. The labels in the preset label set can be divided into different categories based on the content they indicate, such as Class A labels, Class B labels, Class C labels, Class D labels, and so on. At least one label can be added to each standard data. The added label can indicate that the standard data has a positive correlation with the corresponding classification category (for example, Class A label flag = 1) or a negative correlation with the corresponding classification category (for example, Class A label flag = 0). Of course, each of the above-mentioned labels can also be classified at a more granular level.

[0072] In practical applications, for playback data corresponding to the natural language dimension, after filtering out non-keywords through the word segmentation preprocessing solution (see the relevant descriptions in steps S401-S404 above), the word segmentation results can be tagged. For playback data corresponding to the user operation behavior dimension, standard data reflecting behavioral characteristics can be tagged based on decision analysis.

[0073] S204: Determine a target time period, and extract a target data set corresponding to the target time period from the labeled standard data set;

[0074] In an embodiment of the present invention, the standard data set comes from the data source of a specific live broadcast, and the standard data set points to the playback data stream of the live broadcast. The standard data set can correspond to the entire playback data stream or a portion of the playback data stream. Each live broadcast corresponds to a certain duration, which can be currently determined (the live broadcast has ended) or currently pending (in the process of live broadcast). The target time period can be regarded as a time window for determining malicious attributes, and the time window is used to realize the segmentation of the tagged standard data set, especially the tagged standard data set corresponding to a longer time. Of course, when the duration indicated by the target time period is greater than or equal to the duration corresponding to the tagged standard data set, the tagged standard data set can be directly used as the target data set. When the duration indicated by the target time period is less than the duration corresponding to the tagged standard data set, the tagged standard data set can be segmented into at least two target data sets based on the target data segment.

[0075] In a specific embodiment, when the data volume of a target dataset exceeds a data volume threshold, the duration of the target time period can be shortened, and then the target dataset can be updated based on the shortened target time period. For example, if the current target time period is 10 minutes long, and the duration of the tagged standard dataset is 27 minutes, three target datasets can be obtained, with durations of 10 minutes, 10 minutes, and 7 minutes, respectively. When the data volume of the first 10-minute target dataset (the one furthest from the current time point) exceeds the data volume threshold, the duration of the current target time period can be shortened by half, to 5 minutes. This prevents the accumulation of large amounts of data and improves data calculation accuracy. Accordingly, the tagged standard dataset is segmented based on the shortened target time period (5 minutes), resulting in six target datasets, with durations of 5 minutes, 5 minutes, 5 minutes, 5 minutes, 5 minutes, and 2 minutes, respectively. Of course, the percentage by which the duration of the current target time period is shortened is not limited to the aforementioned 50%. The duration of the current target time period can be pre-set, for example, based on historical malicious information detection experience from the broadcast service.

[0076] S205: Determine the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set.

[0077] In an embodiment of the present invention, the malicious attributes of a playback service can be manifested as whether a specific live broadcast room is involved in the dissemination of harmful information, or whether a specific live broadcast session is involved in the dissemination of harmful information. Combined with the introduction of the aforementioned time window concept, standard data carrying different types of labels can be collaboratively analyzed within the corresponding time window to conduct more accurate qualitative analysis and increase the accuracy of rule (policy) judgments. This can avoid the high rate of misjudgment caused by determining the malicious attributes of a corresponding playback service based solely on playback data from a single dimension, and avoid adversely affecting the user experience, such as determining the malicious attributes of a corresponding playback service based solely on screenshots.

[0078] When determining the malicious attributes of the corresponding playback service based on the label carried by each target data in the target data set, you can refer to the following execution process:

[0079] 1) In the case of only one target data set, the labels carried by each target data can be counted, and the malicious attributes of the corresponding playback service can be determined based on the statistical results and preset rules;

[0080] 2) In the case of at least two target datasets:

[0081] 2.1) The labels carried by each target data item in the first target data set (corresponding to the target time period furthest from the current time point) can be counted. Based on the statistical results and preset rules, the malicious attributes of the corresponding playback service can be determined. If there is a situation of spreading harmful information, the labels corresponding to the subsequent target data sets are no longer counted. Otherwise, the labels carried by each target data item in the next target data set (corresponding to the target time period further from the current time point) are counted. The above steps of "determining the malicious attributes of the corresponding playback service based on the statistical results and preset rules until otherwise, counting the labels carried by each target data item in the next target data set" are repeated.

[0082] 2.2) The labels carried by each target data item in the last target data set (corresponding to the target time period closest to the current time point) can be counted. Based on the statistical results and preset rules, the malicious attributes of the corresponding playback service can be determined. If there is a situation of spreading harmful information, the labels corresponding to the previous target data set are no longer counted. Otherwise, the labels carried by each target data item in the previous target data set (corresponding to the target time period next closest to the current time point) are counted. The above steps of "determining the malicious attributes of the corresponding playback service based on the statistical results and preset rules until otherwise, counting the labels carried by each target data item in the previous target data set" are repeated.

[0083] 2.3) The labels corresponding to each target dataset can be counted separately, and then the sub-malicious attributes corresponding to each target dataset can be determined based on the statistical results and preset rules. The malicious attributes of the corresponding playback service can then be determined by combining the sub-malicious attributes corresponding to each target dataset.

[0084] The statistical results mentioned in 1) and 2) above may include the total number of tags, the number and distribution of categories indicated by all tags, the number and proportion of tags in a category, and the positive and negative ratios of tags in a category. Preset rules may include thresholds for the number and proportion of tags in a category corresponding to the spread of harmful information (e.g., the proportion of tags in a category among all tags, or the proportion of positive tags among tags in the same category).

[0085] In practical applications, such as Figure 6As shown, based on a labeled standard dataset and a time sliding window (sliding window), a comprehensive assessment is performed on data from new dimensions (categories indicated by the labels) within a continuous time period, improving the accuracy of handling the dissemination of harmful information, such as pornography detection. The characteristics of the time sliding window are leveraged to collaboratively group data from different (new) dimensions, with data within the same window grouped together. The size of the time sliding window is divided by time, and the amount of data falling within the time sliding window is M. Each data item is labeled (e.g., malicious or non-malicious). Data within the time sliding window is selected to determine the malicious attributes of the streaming service corresponding to the time period (including the user malicious probability and attributes for the host client and the user malicious probability and attributes for the viewer client). Collaborative processing of data from different new and old dimensions, primarily through collaborative organization and grouping based on the time sliding window, can reduce data processing volume and complexity. Data can be dynamically grouped according to collaborative criteria, and thresholds such as the dynamic time sliding window can be set based on the malicious attribute determination results, improving data processing accuracy.

[0086] In a specific embodiment, corresponding candidate data groups can be created for the target dataset based on the classification categories indicated by the labels carried by each target data item. The number of corresponding candidate data groups is equal to the number of the indicated classification categories. For example, the target dataset includes target data A, target data B, and target data C. Target data A carries a class A label (flag=1) and a class C label (flag=0), target data B carries a class A label (flag=0) and a class D label (flag=1), and target data C carries a class A label (flag=1) and a class C label (flag=1). This involves three categories of labels (class A, class C, and class D). Accordingly, the three candidate data groups are: candidate data group 1 (class A: target data A, target data B, and target data C), candidate data group 2 (class C: target data A and target data C), and candidate data group 3 (class D: target data B). Then, based on the positive and negative attributes of the labels carried by each target data item in the candidate data groups, the malicious attributes of the target dataset under the corresponding classification categories are determined. The malicious attributes of the corresponding playback service are then determined based on the malicious attributes of the target dataset under the corresponding classification category. Results-related items and pre-set rule-related items (see the aforementioned description) can be statistically analyzed to determine the malicious attributes corresponding to the candidate data set, that is, the malicious attributes of the target dataset under the corresponding classification category. This is then summarized to obtain the malicious attributes of the corresponding playback service.

[0087] In another specific embodiment, different courses are conducted in a live educational broadcast room. When determining the malicious attributes of the corresponding broadcast service, the malicious attribute determination can be optimized by combining the course attributes and the host's historical data (which can serve as a reference for horizontal collaborative analysis). For example, if the malicious attribute determination indicates the dissemination of certain types of harmful information, and the course attributes fall into the historical misjudgment category of "dissemination of certain types of harmful information," and the host's historical data indicates no such violations, the threshold for determining "dissemination of certain types of harmful information" in the preset rules (see the aforementioned description) can be lowered. For example, if the statistical results (see the aforementioned description) indicate that the number of relevant category tags is 30 and the standard threshold is 25, the lowered standard threshold can be 40. This not only effectively determines the dissemination of harmful information and promptly prevents the relevant content from having a negative impact on viewers, but also reduces the false positive rate, ensuring the effective delivery of the course and a positive learning experience for viewers.

[0088] In another specific embodiment, after determining the malicious attributes of the corresponding playback service, a corresponding release instruction or interception instruction can be generated based on the malicious attributes of the corresponding playback service; then, in response to the received feedback message, the duration of the target time period is adjusted. If the malicious attribute determination result indicates that there is a situation of spreading harmful information, an interception instruction is generated (which can be used to offline the host client's login account, offline the corresponding live broadcast, close the corresponding live broadcast room, and send a reminder notification to the corresponding viewer client); otherwise, a release instruction is generated. If a complaint-type feedback message is received after the interception instruction is executed, and verification (such as a detailed offline audit) shows that it is a reasonable complaint, it indicates that the generation of the interception instruction is inaccurate. Alternatively, if a report-type feedback message is received after the release instruction is executed, and verification (such as a detailed offline audit) shows that it is a reasonable report, it indicates that the generation of the release instruction is inaccurate. Accordingly, the target time can be used to adjust the corresponding duration and the preset time interval (the time interval threshold for taking screenshots in the aforementioned steps S301-S304) as control variables, which can guide the relevant parameter settings in the data dimensionality reduction processing and the relevant parameter settings in the word segmentation preprocessing scheme. The accuracy of determining the malicious attributes of the corresponding playback service can be improved by continuously iterating and optimizing the relevant models (adding misjudgment data and expanding the non-keyword vocabulary).

[0089] As can be seen from the technical solutions provided in the above embodiments of this specification, the embodiments of this specification provide a more efficient and flexible solution for malicious information detection on playback services. Taking into account the different dimensions of playback data, the playback data of at least two dimensions are standardized. By combining the time window to judge the malicious attributes of the playback service, problems can be located more quickly, the accuracy and efficiency of data processing can be improved, and the security of the playback platform can be improved. It can help staff locate and analyze problems more quickly, become familiar with the attributes related to security information, and formulate rules for determining malicious attributes.

[0090] The embodiment of the present invention also provides a playback data processing device, such as Figure 7 As shown, the device includes:

[0091] Acquisition module 710: used to acquire playback data sets of at least two dimensions;

[0092] Processing module 720: configured to process the playback data set of at least two dimensions into a standard data set;

[0093] Adding module 730: used to add at least one label to each standard data in the standard data set based on a preset label set;

[0094] Extraction module 740: used to determine a target time period and extract a target data set corresponding to the target time period from the labeled standard data set;

[0095] The determination module 750 is configured to determine the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set.

[0096] It should be noted that the device and method embodiments in the device embodiment are based on the same inventive concept.

[0097] An embodiment of the present invention provides an electronic device, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the playback data processing method provided in the above method embodiment.

[0098] Furthermore, Figure 8 The hardware structure diagram of an electronic device for implementing the playback data processing method provided by an embodiment of the present invention is shown. The electronic device may participate in constituting or include the playback data processing device provided by an embodiment of the present invention. Figure 8As shown, the electronic device 80 may include one or more processors 802 (illustrated as 802a, 802b, ..., 802n in the figure) (the processor 802 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 804 for storing data, and a transmission device 806 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that Figure 8 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 8 More or fewer components than shown, or with Figure 8 Different configurations shown.

[0099] It should be noted that the one or more processors 802 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." This data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the electronic device 80 (or mobile device). As discussed in the embodiments of this application, this data processing circuitry serves as a processor control (e.g., selecting a variable resistor terminal path connected to an interface).

[0100] The memory 804 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the playback data processing method described in the embodiments of the present invention. The processor 802 executes various functional applications and data processing by running the software programs and modules stored in the memory 804, thereby implementing the above-mentioned playback data processing method. The memory 804 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 804 may further include memory remotely located relative to the processor 802, and these remote memories may be connected to the electronic device 80 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0101] Transmission device 806 is used to receive or transmit data via a network. A specific example of such a network may include a wireless network provided by the telecommunications provider of electronic device 80. In one embodiment, transmission device 806 includes a network interface controller (NIC), which can connect to other network devices via a base station to enable communication with the Internet. In one embodiment, transmission device 806 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0102] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 80 (or mobile device).

[0103] An embodiment of the present invention also provides a storage medium, which can be set in an electronic device to store at least one instruction or at least one program related to a playback data processing method in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the playback data processing method provided by the above method embodiment.

[0104] Optionally, in this embodiment, the storage medium may be located in at least one of a plurality of network servers in the computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard drive, a magnetic disk, or an optical disk, among other media capable of storing program code.

[0105] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not represent the superiority or inferiority of the embodiments. The above description is of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0106] The various embodiments in this specification are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device and electronic device embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0107] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.

[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A playback data processing method, characterized in that: The method comprises: Obtain a playback dataset of at least two dimensions; Processing the playback data set of at least two dimensions into a standard data set; Add at least one label to each standard data in the standard data set based on a preset label set; Determine a target time period, and extract a target data set corresponding to the target time period from the labeled standard data set; Determining a malicious attribute of a corresponding playback service based on a label carried by each target data in the target data set; When the at least two dimensions include at least two natural language dimensions, processing the playback data sets of the at least two dimensions into a standard data set includes: Selecting a natural language dimension with a higher information expression concentration from the at least two natural language dimensions as a reference dimension, wherein the information expression concentration indicates whether the functions of natural language channels for expressing information in the natural language dimension are concentrated, and the information expression concentration is negatively correlated with the number of functions of the natural language channels; dividing the playback data set of the at least two dimensions into a first data set and a second data set based on the reference dimension, wherein the first data set corresponds to the reference dimension; processing the second data set into a third data set corresponding to the reference dimension; The standard data set is obtained based on the first data set and the third data set.

2. The method according to claim 1, characterized in that When the at least two natural language dimensions include a video corpus dimension and an image corpus dimension, the information expression concentration of the video corpus dimension is less than the information expression concentration of the image corpus dimension, and selecting the natural language dimension with higher information expression concentration from the at least two natural language dimensions as the reference dimension includes: Selecting the image corpus dimension with a higher information expression concentration from the video corpus dimension and the image corpus dimension as the reference dimension; Accordingly, the second data set is a video data stream corresponding to the video corpus dimension, and processing the second data set into a third data set corresponding to the reference dimension includes: intercepting a video image from the video data stream based on a preset time interval; The third data set is obtained based on the captured video image.

3. The method according to claim 1, characterized in that When the at least two natural language dimensions include an image corpus dimension and a text corpus dimension, the information expression concentration of the image corpus dimension is less than the information expression concentration of the text corpus dimension, and selecting the natural language dimension with higher information expression concentration from the at least two natural language dimensions as a reference dimension includes: Selecting the text corpus dimension with a higher information expression concentration from the image corpus dimension and the text corpus dimension as the reference dimension; Accordingly, the second data set is an image data set corresponding to the image corpus dimension, and processing the second data set into a third data set corresponding to the reference dimension includes: extracting text data from each image data in the image data set using character recognition technology; The third data set is obtained based on the text data extracted from each image data.

4. The method according to claim 1, wherein When the at least two natural language dimensions include an audio corpus dimension and a text corpus dimension, the information expression concentration of the audio corpus dimension is less than the information expression concentration of the text corpus dimension, and selecting the natural language dimension with higher information expression concentration from the at least two natural language dimensions as a reference dimension includes: Selecting the text corpus dimension with a higher information expression concentration from the audio corpus dimension and the text corpus dimension as the reference dimension; Accordingly, the second data set is an audio data stream corresponding to the audio corpus dimension, and processing the second data set into a third data set corresponding to the reference dimension includes: The audio data stream is processed using speech recognition technology to obtain corresponding text data, and the corresponding text data is used as the third data set.

5. The method according to claim 1, wherein When the at least two dimensions include a natural language dimension and a user operation behavior dimension, processing the playback data set of the at least two dimensions into a standard data set includes: When the natural language dimension is a preset natural language dimension, obtaining first-type standard data based on playback data corresponding to the natural language dimension; When the natural language dimension is not the preset natural language dimension, processing the playback data corresponding to the natural language dimension into playback data corresponding to the preset natural language dimension, and obtaining second-type standard data based on the processed playback data; Extracting behavior feature data from the playback data corresponding to the user operation behavior dimension, and obtaining third-category standard data based on the behavior feature data; Obtaining the standard data set based on the first type of standard data, the second type of standard data, and the third type of standard data; The first type of standard data, the second type of standard data and the third type of standard data correspond to the same data format.

6. The method according to claim 1, characterized in that The determining the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set includes: Based on the classification category indicated by the label carried by each target data, creating corresponding candidate data groups for the target data set, the number of the corresponding candidate data groups being equal to the number of the indicated classification categories; Determining the malicious attributes of the target data set under the corresponding classification category based on the positive and negative attributes of the labels carried by each target data in the candidate data set; The malicious attribute of the corresponding playback service is determined based on the malicious attribute of the target data set under the corresponding classification category.

7. The method according to claim 1, characterized in that After extracting the target data set corresponding to the target time period from the labeled standard data set, the method further includes: When the data volume of the target data set is greater than the data volume threshold, shortening the duration corresponding to the target time period; The target data set is updated based on the shortened target time period.

8. The method according to claim 1, characterized in that After determining the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set, the method further includes: Generate a corresponding release instruction or interception instruction based on the malicious attribute of the corresponding playback service; In response to the received feedback message, the duration corresponding to the target time period is adjusted.

9. A playback data processing device, characterized in that: The device comprises: Acquisition module: used to obtain playback data sets of at least two dimensions; A processing module: configured to process the playback data set of at least two dimensions into a standard data set; Adding module: used to add at least one label to each standard data in the standard data set based on a preset label set; Extraction module: used to determine the target time period and extract the target data set corresponding to the target time period from the labeled standard data set; A determination module: configured to determine the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set; Among them, when the at least two dimensions include at least two natural language dimensions, the processing module is also used to: select a natural language dimension with a higher information expression concentration from the at least two natural language dimensions as a reference dimension, the information expression concentration indicates whether the functions of the natural language channels used to express information in the natural language dimension are concentrated, and the information expression concentration is negatively correlated with the number of functions of the natural language channels; divide the playback data sets of the at least two dimensions into a first data set and a second data set based on the reference dimension, the first data set corresponding to the reference dimension; process the second data set into a third data set corresponding to the reference dimension; and obtain the standard data set based on the first data set and the third data set.

10. The device according to claim 9, characterized in that When the at least two natural language dimensions include a video corpus dimension and an image corpus dimension, the information expression concentration of the video corpus dimension is less than the information expression concentration of the image corpus dimension, and selecting the natural language dimension with higher information expression concentration from the at least two natural language dimensions as the reference dimension includes: Selecting the image corpus dimension with a higher information expression concentration from the video corpus dimension and the image corpus dimension as the reference dimension; Accordingly, the second data set is a video data stream corresponding to the video corpus dimension, and processing the second data set into a third data set corresponding to the reference dimension includes: intercepting a video image from the video data stream based on a preset time interval; The third data set is obtained based on the captured video image.

11. The device according to claim 9, characterized in that When the at least two natural language dimensions include an image corpus dimension and a text corpus dimension, the information expression concentration of the image corpus dimension is less than the information expression concentration of the text corpus dimension, and selecting the natural language dimension with higher information expression concentration from the at least two natural language dimensions as a reference dimension includes: Selecting the text corpus dimension with a higher information expression concentration from the image corpus dimension and the text corpus dimension as the reference dimension; Accordingly, the second data set is an image data set corresponding to the image corpus dimension, and processing the second data set into a third data set corresponding to the reference dimension includes: extracting text data from each image data in the image data set using character recognition technology; The third data set is obtained based on the text data extracted from each image data.

12. The device according to claim 9, characterized in that When the at least two natural language dimensions include an audio corpus dimension and a text corpus dimension, the information expression concentration of the audio corpus dimension is less than the information expression concentration of the text corpus dimension, and selecting the natural language dimension with higher information expression concentration from the at least two natural language dimensions as a reference dimension includes: Selecting the text corpus dimension with a higher information expression concentration from the audio corpus dimension and the text corpus dimension as the reference dimension; Accordingly, the second data set is an audio data stream corresponding to the audio corpus dimension, and processing the second data set into a third data set corresponding to the reference dimension includes: The audio data stream is processed using speech recognition technology to obtain corresponding text data, and the corresponding text data is used as the third data set.

13. The device according to claim 9, characterized in that When the at least two dimensions include a natural language dimension and a user operation behavior dimension, processing the playback data set of the at least two dimensions into a standard data set includes: When the natural language dimension is a preset natural language dimension, obtaining first-type standard data based on playback data corresponding to the natural language dimension; When the natural language dimension is not the preset natural language dimension, processing the playback data corresponding to the natural language dimension into playback data corresponding to the preset natural language dimension, and obtaining second-type standard data based on the processed playback data; Extracting behavior feature data from the playback data corresponding to the user operation behavior dimension, and obtaining third-category standard data based on the behavior feature data; Obtaining the standard data set based on the first type of standard data, the second type of standard data, and the third type of standard data; The first type of standard data, the second type of standard data and the third type of standard data correspond to the same data format.

14. The device according to claim 9, characterized in that The determining the malicious attribute of the corresponding playback service based on the label carried by each target data in the target data set includes: Based on the classification category indicated by the label carried by each target data, creating corresponding candidate data groups for the target data set, the number of the corresponding candidate data groups being equal to the number of the indicated classification categories; Determining the malicious attributes of the target data set under the corresponding classification category based on the positive and negative attributes of the labels carried by each target data in the candidate data set; The malicious attribute of the corresponding playback service is determined based on the malicious attribute of the target data set under the corresponding classification category.

15. The device according to claim 9, characterized in that The device is further configured to: shorten the duration corresponding to the target time period when the data volume of the target data set is greater than a data volume threshold; and update the target data set based on the shortened target time period.

16. The device according to claim 9, characterized in that The device is further configured to: generate a corresponding release instruction or an interception instruction based on the malicious attribute of the corresponding playback service; and adjust the duration corresponding to the target time period in response to a received feedback message.

17. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the playback data processing method as described in any one of claims 1-8.

18. A computer-readable storage medium, characterized in that The storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the playback data processing method according to any one of claims 1 to 8.

19. A computer program product, characterized in that The computer program includes computer instructions, which are loaded and executed by a processor to implement the playback data processing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Video violation content detection method and device and storage medium

    CN110798703A