Method and system for detecting illegal content in free online streaming media service
By simulating video playback behavior and extracting multimodal content features in free online streaming services, the problem of bad content detection on mobile is solved, and more efficient and accurate identification of illegal content is achieved.
Patent Information
- Application Number
- CN202510754344.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-18
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-15
AI Technical Summary
The existing technology is difficult to effectively identify and detect the spread of bad content in free online streaming services, especially the convenience of operation on mobile terminals increases the difficulty of supervision, making it difficult for traditional filtering technologies to deal with and users face security risks.
By searching mainstream media platforms based on preset keyword collections, hierarchical processing and classification, simulating video playback behavior, extracting multimodal content features and performing fusion analysis, and using illegal content detection models to output detection results.
It improves the accuracy and reliability of illegal content detection, reduces misjudgment and misjudgment, improves detection coverage and efficiency, and adapts to the operation convenience of the mobile terminal and dynamic changes in the platform.
Smart Images

Figure CN120499449A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to a method and system for detecting illegal content in a free online streaming service. Background Art
[0002] With the rapid development of the internet, mainstream media platforms have become a vital hub for global users to access entertainment, information, and social interaction. However, the methods used to disseminate some inappropriate content are constantly evolving, and its spread on mainstream media platforms is becoming increasingly insidious. For example, some links are packaged as legitimate video covers, but when users click on the video playback page, they are presented with non-compliant content. This non-compliant content falls into two main categories: pop-up pages containing inappropriate and misleading information; and inappropriate content within embedded video playback windows on the site. This content may not only violate relevant regulations but also negatively impact the healthy development of young people. It also poses the risk of compromising user privacy and poses numerous risks to the user's digital security.
[0003] The problem of harmful pop-up ads is particularly prominent on mobile devices. Because users operate more frequently and trigger actions more easily on mobile devices, this undoubtedly increases the difficulty of regulation. The uncertainty surrounding the emergence of harmful content makes traditional filtering technologies difficult to effectively address, posing significant security risks to users. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for detecting illegal content in a free online streaming service, thereby improving the accuracy and reliability of detection.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows: In a first aspect, a method for detecting illegal content in a free online streaming service is provided, the method comprising: Search and crawl mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set; Performing layered and classified processing on the target mainstream media platform set to generate a layered and classified platform set; Obtain content data of corresponding website platforms based on the hierarchically classified platform set, and select target platforms by calculating content differences; Simulating video playback behavior on the target platform, parsing the playback page, and generating a page processing result; Extracting multimodal content of text, image, and video stream from the page processing result to generate feature data of each modality; The feature data of each modality is fused and analyzed, and the detection results are output through the illegal content detection model.
[0006] Furthermore, based on the preset keyword set, mainstream media platforms are searched and crawled to obtain a target mainstream media platform set, including: By collecting keywords to build a keyword set; Generating a search result page through multiple search engines based on the keyword set; Performing accessibility verification on the search result page to filter out accessible pages; Analyze the content of the accessible pages and calculate their matching degree with the original search keywords; Based on the accessibility verification results and the matching degree, eligible mainstream media platforms are captured to generate a target mainstream media platform set.
[0007] Furthermore, the target mainstream media platform set is subjected to layering and classification processing to generate a layered and classified platform set, including: Obtaining the HTML document of a platform in the target mainstream media platform set; Traversing the URLs in the HTML document and generating feature values according to the URL suffix conversion; By analyzing the DOM tree structure of the HTML document, the hierarchical value corresponding to each URL feature value is located; Extracting HTML structural features, string features, suffix features, and access features of URLs at the same level based on the level value; Based on the URL features and the level value, the URL is classified as the website homepage, the first-level page or the second-level page.
[0008] Furthermore, content data of corresponding website platforms is obtained based on the hierarchically classified platform set, and target platforms are screened out by calculating content differences, including: Based on the hierarchically classified platform set, HTML texts from desktop and mobile terminals are obtained separately; Extracting core elements of the page from the HTML text and converting the core elements into text vectors; Calculating the similarity between the text vectors of the desktop and mobile terminals to generate a content difference value; The difference value is compared with a preset threshold, and the end with richer information is selected as the target content website platform.
[0009] Furthermore, the target platform simulates the video playback behavior, and parses the playback page to generate a page processing result, including: For the target content website platform, locate the video playback window and collect the target URL in the secondary page; Formulate a segmented playback strategy based on the video duration distribution characteristics corresponding to the target URL; Loading the target URL in the playback window and performing a non-continuous playback operation according to the segmented playback strategy; Monitor DOM structure changes in real time during playback, capturing triggered pop-up ads and hidden page elements; Record dynamic page data and abnormal elements triggered by playback, and generate page processing results including timing behavior.
[0010] Furthermore, multimodal content of text, image and video stream is extracted from the page processing result to generate feature data of each modality, including: Parsing dynamically loaded images and pop-up window content based on the dynamic page data and abnormal elements in the page processing result; Separate illegal links and their disguised page elements from the parsed content; For the separated illegal links and disguised elements, the following feature extraction operations are performed in parallel: Extract the DOM tree hierarchical relationship as HTML structural features, parse the link jump path to generate URL topology features, capture the link anchor text and advertising copy to generate text semantic features, and identify visual objects in disguised elements to generate image visual features.
[0011] Furthermore, the feature data of each modality is fused and analyzed, and the illegal content detection model outputs the detection results, including: Construct cross-modal association rules based on the extracted HTML structural features, URL topological features, text semantic features, and image visual features; According to the association rules, the four types of features are mapped to a unified dimensional space to generate a joint feature vector; Inputting the joint feature vector into a pre-trained multimodal illegal content detection model and analyzing the feature contribution through a dynamic weight allocation mechanism; Based on the matching results of feature contribution and historical violation pattern library, the illegal content probability value and disguise type mark are output.
[0012] In a second aspect, a system for detecting illegal content in a free online streaming service includes: An acquisition module is used to search and crawl mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set; A screening module is used to perform hierarchical processing and classification on the target mainstream media platform set to generate a hierarchical and classified platform set; obtain content data of corresponding website platforms based on the hierarchical and classified platform set, and screen out target platforms by calculating content differences; The processing module is used to simulate the video playback behavior on the target platform, parse and process the playback page, and generate a page processing result; extract the multimodal content of text, images and video streams from the page processing result to generate feature data of each modality; perform fusion analysis on the feature data of each modality, and output the detection result through the illegal content detection model.
[0013] According to a third aspect, a computing device includes: one or more processors; The storage device is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method.
[0014] In a fourth aspect, a computer-readable storage medium stores a program, which implements the method when executed by a processor.
[0015] The above solution of the present invention includes at least the following beneficial effects: Automatically searching and capturing mainstream media platforms through preset keyword sets can quickly and comprehensively obtain target platforms, avoiding omissions and inefficiencies in manual screening, greatly improving the coverage and efficiency of illegal content detection, and enabling real-time dynamic monitoring of massive platform resources.
[0016] Layering and categorizing the target set of mainstream media platforms makes subsequent content analysis more targeted. Differentiating content data acquisition based on the different tiers and categories of the platforms effectively reduces irrelevant data interference, focuses on high-risk or key platforms, and improves the efficiency of detection resource utilization.
[0017] By screening target platforms based on content differences, we can accurately identify platforms with potential risks. Compared with traditional indiscriminate detection of all platforms, this greatly narrows the detection scope, reduces detection costs, and significantly improves the accuracy of locating abnormal platforms.
[0018] Simulating video playback behavior on the target platform and parsing the page can restore the user's actual browsing scenario, obtain more authentic and effective content data, avoid misjudgments or missed judgments due to static detection, and make the detection results more in line with the user's actual situation.
[0019] Extracting and integrating multimodal content such as text, images, and video streams for analysis can capture the characteristics of inappropriate content from multiple dimensions compared to single-modality detection. Combined with illegal content detection models, it can effectively identify various types of hidden non-compliant content, thereby improving the accuracy and reliability of detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1The present invention is a flowchart of a method for detecting illegal content in a free online streaming service provided by an embodiment of the present invention.
[0021] Figure 2 The figure is a schematic diagram of an illegal content detection system in a free online streaming service provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.
[0023] like Figure 1 As shown, an embodiment of the present invention provides a method for detecting illegal content in a free online streaming service, the method comprising the following steps: Step 11: Search and crawl mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set; Step 12: performing layered and classified processing on the target mainstream media platform set to generate a layered and classified platform set; Step 13: Obtain content data of corresponding website platforms based on the hierarchically classified platform set, and select target platforms by calculating content differences; Step 14: simulating video playback behavior on the target platform, parsing the playback page, and generating a page processing result; Step 15: extracting multimodal content of text, image, and video stream from the page processing result, and generating feature data of each modality; Step 16: Perform fusion analysis on the feature data of each modality and output the detection results through the illegal content detection model.
[0024] In an embodiment of the present invention, by automatically searching and crawling mainstream media platforms through a preset keyword set, the target platform can be quickly and comprehensively acquired, avoiding omissions and inefficiencies in manual screening, greatly improving the coverage and efficiency of illegal content detection, and enabling real-time dynamic monitoring of massive platform resources. The target mainstream media platform set is layered and classified, making subsequent content analysis more targeted. Differentiated acquisition of content data based on different levels and categories of the platform can effectively reduce interference from irrelevant data, focus on high-risk or key platforms, and improve the efficiency of detection resource utilization. Screening target platforms based on content differences can accurately identify platforms with potential risks. Compared with traditional indiscriminate detection of all platforms, it greatly reduces the scope of detection, reduces detection costs, and significantly improves the accuracy of locating abnormal platforms. Simulating video playback behavior on the target platform and parsing the page can restore the user's actual browsing scene, obtain more real and effective content data, avoid misjudgment or omission due to static detection, and make the detection results more in line with the user's actual encounter. Extracting and integrating multimodal content such as text, images, and video streams for analysis can capture the characteristics of inappropriate content from multiple dimensions compared to single-modality detection. Combined with illegal content detection models, it can effectively identify various types of hidden non-compliant content, thereby improving the accuracy and reliability of detection.
[0025] In a preferred embodiment of the present invention, the above step 11, searching and crawling mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set, may include: By collecting keywords to build a keyword set; Generating a search result page through multiple search engines based on the keyword set; Performing accessibility verification on the search result page to filter out accessible pages; Analyze the content of the accessible pages and calculate their matching degree with the original search keywords; Based on the accessibility verification results and the matching degree, eligible mainstream media platforms are captured to generate a target mainstream media platform set.
[0026] In the embodiment of the present invention, when the above steps are specifically applied, the above steps can be implemented by the following steps, for example: First, requests were sent to multiple target streaming platforms to retrieve page content. After extracting the text, words with a length of less than 10 were truncated by spaces, and duplicates and invalid symbols were removed. The frequency of each keyword was calculated by counting the number of times it appeared on different platforms. From a vocabulary representing the total platform sample, high-frequency keywords were selected and included in the vocabulary database. These keywords were then divided into general keywords (such as "video" and "live broadcast," which are common on the platform), dynamic and popular keywords (such as those related to recent hot online events), and target keywords (a random combination of the first two high-frequency categories).
[0027] Multi-engine retrieval and result acquisition: submit the constructed keyword set to multiple search engines, and obtain the search result pages returned by each search engine by sending requests. Different search engines will provide page information from different sources based on their own algorithms and index libraries.
[0028] Accessibility verification involves sending an access request to the retrieved search results page to determine whether the page responds normally. If the page cannot be opened, returns an error code, or times out, it is marked as inaccessible. Pages that can display content normally are considered accessible, and inaccessible pages are excluded.
[0029] Matching calculation: clean the accessible page content, remove irrelevant information such as HTML tags and advertisements, retain only plain text, analyze the correlation between the page plain text and the original search keywords, determine the frequency of keyword appearance on the page, context and other factors, evaluate the degree of match between the page content and the keywords, and give a corresponding matching score.
[0030] Platform crawling and collection generation, combining page accessibility verification results and matching scores, selects mainstream media platforms corresponding to accessible and highly matching pages for crawling, and aggregates these platforms to generate the target mainstream media platform collection.
[0031] In an embodiment of the present invention, by constructing a hierarchical and classified keyword set and combining it with multi-search engine retrieval, mainstream platforms related to streaming media can be accurately located from massive network resources, avoiding interference from irrelevant platforms and improving the accuracy of target platform acquisition. By utilizing the differences between multiple search engines, the results of different information sources are integrated to expand the information coverage and increase the diversity of acquired content; at the same time, multi-source information complements and verifies each other, improving the reliability of the mainstream media platform information finally acquired. Through accessibility verification, invalid pages are eliminated in advance, reducing the resource consumption of subsequent processing of invalid links; based on the matching degree screening platform, highly relevant pages are prioritized to avoid invalid analysis of low-value pages, significantly improving the efficiency of system resource utilization and overall processing speed. Dynamic hot keywords are included in the keyword set, which can follow the changes in network hot spots, capture emerging or popular streaming media platforms in a timely manner, and ensure the timeliness and integrity of the target mainstream media platform set.
[0032] In a preferred embodiment of the present invention, the above step 12, performing layered processing and classification processing on the target mainstream media platform set to generate a layered and classified platform set, may include: Obtaining the HTML document of a platform in the target mainstream media platform set; Traversing the URLs in the HTML document and generating feature values according to the URL suffix conversion; By analyzing the DOM tree structure of the HTML document, the hierarchical value corresponding to each URL feature value is located; Extracting HTML structural features, string features, suffix features, and access features of URLs at the same level based on the level value; Based on the URL features and the level value, the URL is classified as the website homepage, the first-level page or the second-level page.
[0033] In the embodiment of the present invention, when the above steps are specifically applied, the above steps can be implemented by the following steps, for example: Simulate desktop browser behavior, configure automation tools (such as Selenium) to open the target streaming platform URL, set a reasonable waiting time to ensure that all elements of the page (including dynamically loaded content) are fully rendered, and obtain the complete HTML document content.
[0034] URL feature conversion: Suffix feature extraction: Analyze the URL suffix type and map the .html suffix to a feature value of 0, the .js suffix to 1, and other suffixes to 2.
[0035] DOM level positioning: Taking the HTML root tag as level 0, analyze the tag nesting relationship layer by layer to determine the level value of the tag where each URL is located (that is, the number of external nested tag layers plus 1).
[0036] Statistics of the number of URLs at the same level: Calculates the total number of URLs at the same DOM level, reflecting the link density of that level.
[0037] Attribute feature extraction: Check whether the URL contains the title and target attributes, and whether the URL string contains the play keyword or a number-number pattern substring. If so, the corresponding feature value is 1, otherwise it is 0.
[0038] Accessibility verification: Send an HTTP request to each URL and determine whether it is accessible based on the returned status code (status code 200 corresponds to a feature value of 1, otherwise it is 0).
[0039] Feature integration and classification: The extracted suffix features, DOM level values, number of URLs at the same level, attribute features, and accessibility features are integrated and input into the trained CART decision tree model. Based on the feature combination rules, the URLs are classified as website homepages (such as root directory URLs), first-level pages (such as main channel pages), or second-level pages (such as specific video playback pages). The pages to be tested are stored in the target URL library.
[0040] This invention achieves refined classification of platform URLs through multi-dimensional feature extraction (suffix, DOM hierarchy, attributes, etc.) and a decision tree model, accurately identifying pages of varying importance and functionality. Based on the page classification results, high-value pages (such as playback pages and user-generated content pages) are prioritized, optimizing detection resource allocation, improving overall processing efficiency, and reducing unnecessary detection overhead for low-risk pages (such as static introduction pages). By using automated tools to capture HTML documents and analyze the DOM structure in real time, the system can adapt to dynamic changes in the platform's page structure (such as navigation reconstruction and content module updates) while maintaining classification accuracy. The feature extraction method supports expansion (such as adding URL parameter features and response time features), enhancing the adaptability of the classification model based on actual needs and addressing new page types or complex content layouts. Classification is based on a comprehensive approach of multi-dimensional features rather than a single rule, reducing misclassifications due to superficial URL similarities (such as misclassifying advertising links as content pages) and improving classification reliability.
[0041] In a preferred embodiment of the present invention, the above step 13, obtaining content data of corresponding website platforms according to the hierarchically classified platform set and screening out target platforms by calculating content difference, may include: Based on the hierarchically classified platform set, HTML texts for desktop and mobile terminals are obtained separately; Extracting core elements of the page from the HTML text and converting the core elements into text vectors; Calculating the similarity between the text vectors of the desktop and mobile terminals to generate a content difference value; The difference value is compared with a preset threshold, and the end with richer information is selected as the target content website platform.
[0042] In the embodiment of the present invention, when the above steps are specifically applied, the above steps can be implemented by the following steps, for example: Multi-terminal content collection: Use automated tools to simulate desktop (such as window size 1920×1080) and mobile (such as iPhone12 device model) access to the hierarchical and classified platforms to obtain complete HTML text.
[0043] Preprocess HTML documents: Use regular expressions to remove irrelevant content such as advertising tags and repeated modules, and retain core elements (such as pop-up div tags, href links, img tag src attributes, and body text).
[0044] Core element vectorization: Extract key information from the preprocessed HTML text, such as pop-up identifiers, potential illegal links, image addresses, text content, etc., and convert these core elements into text vectors (e.g., mapping text content into numerical feature vectors through word frequency statistics, keyword extraction, etc.).
[0045] Difference calculation and analysis: Compare the text vectors of desktop and mobile terminals, and analyze the differences between the two in terms of keyword distribution, link structure, image resources, etc.
[0046] By evaluating indicators such as text length and feature vector overlap, the degree of difference can be quantified (for example, text length difference rate, key element frequency difference, etc.).
[0047] Threshold determination and target screening: A preset difference threshold (such as 2%) is set. If the difference is higher than the threshold, it means that there is a significant difference between the contents at both ends.
[0048] Compare the information content of the texts at both ends (such as the total number of characters, the number of unique keywords, the density of effective links, etc.), and select the end with richer information (such as the terminal with longer text and more complex elements) as the target content platform.
[0049] The present invention compares the content differences between desktop and mobile terminals to identify the illegal content dissemination paths caused by device characteristics (such as hidden pop-up windows are more likely to appear on mobile terminals), thereby improving the targeted detection; gives priority to terminals with rich information content for detection to avoid missed detections due to incomplete content (such as mobile terminals may hide some illegal content due to screen limitations), thereby improving the comprehensiveness of detection; simulates the access behavior of real users on different terminals, fits the actual scenarios of illegal content dissemination (such as the spread of illegal information caused by the convenience of mobile terminal operation), and enhances the practicality of the detection model; screens target terminals through difference, reduces invalid analysis of low-information-content pages, concentrates computing resources to process high-risk content, reduces system load and improves detection efficiency; and can adapt to the detection accuracy requirements under different regulatory scenarios by adjusting the difference threshold and information content evaluation standard (such as lowering the threshold to expand the detection range under strict regulatory scenarios).
[0050] In a preferred embodiment of the present invention, the above step 14 may include: For the target content website platform, locate the video playback window and collect the target URL in the secondary page; Formulate a segmented playback strategy based on the video duration distribution characteristics corresponding to the target URL; Loading the target URL in the playback window and performing a non-continuous playback operation according to the segmented playback strategy; Monitor DOM structure changes in real time during playback, capturing triggered pop-up ads and hidden page elements; Record dynamic page data and abnormal elements triggered by playback, and generate page processing results including timing behavior.
[0051] In the embodiment of the present invention, when the above steps are specifically applied, the above steps can be implemented by the following steps, for example: Parse the target platform's HTML structure and identify video player-related div tags (e.g., containers with characteristic class names like video-js and youtube-player). Extract target URLs (e.g., video source links, play button URLs) from secondary pages, filtering out irrelevant navigation or advertising links.
[0052] Analyze the video duration corresponding to the target URL and divide the playback process into key intervals (such as 0%-10%, 90%-100%) and regular intervals (10%-90%); set high-frequency sampling for key intervals (such as detecting DOM changes every 2 seconds) and low-frequency sampling for regular intervals (such as detecting every 10 seconds) to form a non-continuous playback strategy.
[0053] Simulation playback and dynamic monitoring: Load the target URL in the playback window and execute playback according to the segmentation strategy (such as playing for 5 seconds → pause → jumping to 90% progress to continue playing); monitor DOM change events in real time (such as DOMSubtreeModified, element addition / deletion), and capture newly appearing pop-up ads (such as floating divs, iframes) or hidden elements (such as mask layers dynamically loaded during playback).
[0054] Data recording and result generation: Record the timing data during playback (such as the time when pop-up windows appear, the order in which page elements are loaded), and the characteristics of abnormal elements (such as link text containing sensitive keywords); deduplicate and correlate the original data (such as binding pop-ups to trigger operations), and generate structured page processing results (such as logs containing timestamps, element types, and content summaries).
[0055] The present invention uses a non-continuous playback strategy to focus on monitoring high-risk periods at the beginning and end of the video (such as pre-roll ads and misleading content at the end of the video), accurately capture illegal pop-ups or hidden links, improve detection coverage, avoid full video playback, and only intensively sample key periods, concentrating computing resources on high-value detection points, significantly reducing processing time and bandwidth usage; simulating real playback behavior to trigger dynamic content (such as automatically loaded ads), discovering illegal content that cannot be identified in static page analysis, and improving detection authenticity; by recording the time series of DOM changes, the triggering path of illegal content can be traced; full-process automated processing supports large-scale video detection, and the segmentation strategy can be dynamically adjusted according to the content type (such as using a more intensive sampling frequency for live streams) to adapt to diverse streaming media scenarios.
[0056] In a preferred embodiment of the present invention, the above step 15, extracting multimodal content of text, image, and video stream from the page processing result and generating feature data of each modality, may include: Parsing dynamically loaded images and pop-up window content based on the dynamic page data and abnormal elements in the page processing result; Separate illegal links and their disguised page elements from the parsed content; For the separated illegal links and disguised elements, the following feature extraction operations are performed in parallel: Extract the DOM tree hierarchical relationship as HTML structural features, parse the link jump path to generate URL topology features, capture the link anchor text and advertising copy to generate text semantic features, and identify visual objects in disguised elements to generate image visual features.
[0057] In the embodiment of the present invention, when the above steps are specifically applied, the above steps can be implemented by the following steps, for example: Dynamic content parsing and element separation: Page scrolling and loading: Simulate user scrolling behavior and trigger dynamic page loading mechanisms (such as infinite scrolling and click-to-load more) to ensure that dynamic elements such as pop-ups and images are fully presented.
[0058] Element positioning and extraction: By parsing HTML tag attributes (such as href, id, class), locate the a tag containing the jump link, the pop-up container div and the associated img tag.
[0059] Combine the base URL and relative path to generate a complete image address, extract the text content (such as advertising copy) and link target address in the pop-up window.
[0060] Illegal element separation: Based on preset rules (such as the link domain name is not in the whitelist, the text contains abnormal keywords), suspected illegal links and their associated images and text packaging elements are screened out and stored in the detection resource library.
[0061] Parallel extraction of multimodal features: HTML structural features: Calculate the element level in the DOM tree (with the html tag as level 0 and the nesting level increasing gradually), and record the hierarchical path of the target element (such as html→body→div→a).
[0062] Detect specific tag combinations (such as div[class*="popup"] wrapping the a tag) and mark it as feature value 1; analyze element attributes (such as style="display:none"), and mark it as feature value 1 if there is a hidden attribute.
[0063] Calculate the relative position of the pop-up window on the page (such as the ratio of the vertical position to the page height, whether the horizontal position is close to the corner), and map it to a region code (such as top = 1, bottom right corner = 5).
[0064] URL topology features: parse the URL structure, extract the main domain name (such as example.com), path hierarchy (such as the number of levels of / page1 / subpage2), and record the forward and backward paths of link jumps (such as the jump chain from the homepage → pop-up link → landing page).
[0065] Text semantic features: Perform word segmentation on the page text (including image text extracted by OCR), filter out stop words and generate a keyword list.
[0066] Detect the matching degree between keywords and preset risk vocabulary (such as high-frequency inducement words), count the word frequency distribution and context association (such as the association between "get it now" and the promotion path in the URL).
[0067] Image visual features: Use convolutional neural networks (CNNs) to extract features from images, identify visual objects (such as people, icons, and scenes) in the images, and generate feature vectors that include color distribution, texture patterns, and target categories (such as "promotional icons" and "warning signs").
[0068] Feature fusion and structured storage: Associate the features of each modality (structural feature encoding, URL domain name, text keywords, image feature vectors) by timestamp to generate a feature dataset containing multi-dimensional information for subsequent detection model calls.
[0069] This invention utilizes dynamic parsing and multimodal feature extraction, covering multiple dimensions including page structure, link behavior, text semantics, and image visuals, avoiding single-modal under-detection (e.g., failing to identify illegal elements in an image based solely on text). Utilizing DOM hierarchical analysis and element position encoding, it can identify hidden pop-ups nested deep within the page structure or in the corners of the page, addressing the drawback of traditional detection that only scans surface content. Linking URL redirect paths with text keywords and image themes (e.g., combining links with unusual domain names with "limited-time offer" text and promotional images) enhances the detection of disguised illegal content. Parallel multimodal feature extraction reduces the cost of manual feature design and allows model updates (e.g., replacing CNN architectures) to adapt to new content formats (e.g., illegal images generated by AIGC). Structured storage of feature data and timestamps allows for the reproduction of illegal content presentation scenarios (e.g., a pop-up triggered at a specific moment in the lower right corner of the page, accompanied by a specific image and link), providing comprehensive evidence support for content review and violation tracing.
[0070] In a preferred embodiment of the present invention, the above step 16 may include: Construct cross-modal association rules based on the extracted HTML structural features, URL topological features, text semantic features, and image visual features; According to the association rules, the four types of features are mapped to a unified dimensional space to generate a joint feature vector; Inputting the joint feature vector into a pre-trained multimodal illegal content detection model and analyzing the feature contribution through a dynamic weight allocation mechanism; Based on the matching results of feature contribution and historical violation pattern library, the illegal content probability value and disguise type mark are output.
[0071] In the embodiment of the present invention, when the above steps are specifically applied, the above steps can be implemented by the following steps, for example: Cross-modal association rule construction: Analyze the co-occurrence patterns of various modal features in historical data (such as deeply nested div tags in HTML structure + URLs containing unfamiliar subdomains + high-frequency misleading words in text + images containing promotional icons), and establish cross-modal association rules (such as "if the URL domain name is not on the whitelist and the image is identified as a warning sign, mark it as high risk").
[0072] Define feature association weights (e.g., the weight of URL topology features for illegal link detection is 0.3, and the weight of image visual features for disguised content detection is 0.4) to form a multimodal feature mapping logic.
[0073] Joint eigenvector generation: Standardize HTML structural features (hierarchical path encoding, tag combination feature values), URL topology features (primary domain name, number of path levels), text semantic features (keyword frequency vectors), and image visual features (deep feature vectors extracted by CNN) (e.g., normalize to the [0,1] range).
[0074] According to the association rule weights, the four types of features are concatenated into a joint feature vector of unified dimension (such as a vector of length N, the first N1 dimensions are structural features, N1-N2 dimensions are URL features, and so on).
[0075] Multimodal model analysis: The joint feature vector is input into a pre-trained multimodal model (such as a Transformer-based cross-modal encoder). The model dynamically calculates the contribution of each modal feature through the self-attention mechanism (for example, the contribution of text features in identifying inductive content is increased to 60%).
[0076] Compare the feature vector with typical feature combinations in the historical violation pattern library (such as a combination pattern containing "Limited-time discount" text + promotional image + unfamiliar domain name) and calculate the matching score.
[0077] Test result output: Based on the matching score and the preset threshold, the illegal content probability value is output (e.g. 0.85 indicates a high probability of violation).
[0078] Based on the modal type with the highest feature contribution, mark the camouflage type (such as "image camouflage", "text inducement", and "link jump camouflage").
[0079] This invention fuses multimodal features through association rules, leveraging text semantics to interpret image content and URL behavior to verify structural anomalies. This enables complementary verification between features and avoids single-modality misjudgments (e.g., misjudging an image as an advertisement simply because it contains an icon, while combining it with the URL domain name can exclude legitimate content). The model automatically focuses on key features through a weight distribution mechanism (e.g., increasing the weight of video stream features when detecting live streaming platforms), adapting to the main dissemination forms of illegal content in different scenarios and improving detection generalization capabilities. It also outputs the probability of illegal content while labeling the disguise type (e.g., "induced links disguised as pop-up images"), providing clear guidance for subsequent handling (e.g., prioritizing blocking the DOM structure of such pop-ups), thereby improving content governance efficiency. Matching analysis based on a library of historical violation patterns can quickly identify variants of known violation types (e.g., links using similar domain names). Combined with dynamic feature contribution analysis, it simultaneously discovers new violation patterns, balancing detection efficiency and innovation. Multimodal analysis is accomplished through a single feature fusion and model inference, reducing computational steps by at least 30% compared to traditional methods that independently process each modality and then aggregate them, significantly improving the real-time performance of large-scale content detection.
[0080] like Figure 2 As shown, an embodiment of the present invention further provides a system for detecting illegal content in a free online streaming service, comprising: An acquisition module is used to search and crawl mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set; A screening module is used to perform hierarchical processing and classification on the target mainstream media platform set to generate a hierarchical and classified platform set; obtain content data of corresponding website platforms based on the hierarchical and classified platform set, and screen out target platforms by calculating content differences; The processing module is used to simulate the video playback behavior on the target platform, parse and process the playback page, and generate a page processing result; extract the multimodal content of text, images and video streams from the page processing result to generate feature data of each modality; perform fusion analysis on the feature data of each modality, and output the detection result through the illegal content detection model.
[0081] It should be noted that this system is a system corresponding to the above method, and all implementation methods in the above method embodiment are applicable to this embodiment and can achieve the same technical effects.
[0082] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for detecting illegal content in a free online streaming service, characterized in that: The method comprises: Search and crawl mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set; Performing layered and classified processing on the target mainstream media platform set to generate a layered and classified platform set; Obtain content data of corresponding website platforms based on the hierarchically classified platform set, and select target platforms by calculating content differences; Simulating video playback behavior on the target platform, parsing the playback page, and generating a page processing result; Extracting multimodal content of text, image, and video stream from the page processing result to generate feature data of each modality; The feature data of each modality is fused and analyzed, and the detection results are output through the illegal content detection model.
2. The method for detecting illegal content in a free online streaming service according to claim 1, wherein: Search and crawl mainstream media platforms based on the preset keyword set to obtain the target mainstream media platform set, including: By collecting keywords to build a keyword set; Generating a search result page through multiple search engines based on the keyword set; Performing accessibility verification on the search result page to filter out accessible pages; Analyze the content of the accessible pages and calculate their matching degree with the original search keywords; Based on the accessibility verification results and the matching degree, eligible mainstream media platforms are captured to generate a target mainstream media platform set.
3. The method for detecting illegal content in a free online streaming service according to claim 2, wherein: The target mainstream media platform set is subjected to layering and classification processing to generate a layered and classified platform set, including: Obtaining the HTML document of a platform in the target mainstream media platform set; Traversing the URLs in the HTML document and generating feature values according to the URL suffix conversion; By analyzing the DOM tree structure of the HTML document, the hierarchical value corresponding to each URL feature value is located; Extracting HTML structural features, string features, suffix features, and access features of URLs at the same level based on the level value; Based on the URL features and the level value, the URL is classified as the website homepage, the first-level page or the second-level page.
4. The method for detecting illegal content in a free online streaming service according to claim 3, wherein: Obtain content data for corresponding website platforms based on the hierarchically classified platform set, and filter out target platforms by calculating content differences, including: Based on the hierarchically classified platform set, HTML texts for desktop and mobile terminals are obtained separately; Extracting core elements of the page from the HTML text and converting the core elements into text vectors; Calculating the similarity between the text vectors of the desktop and mobile terminals to generate a content difference value; The difference value is compared with a preset threshold, and the end with richer information is selected as the target content website platform.
5. The method for detecting illegal content in a free online streaming service according to claim 4, wherein: Simulate video playback behavior on the target platform, parse and process the playback page, and generate page processing results, including: For the target content website platform, locate the video playback window and collect the target URL in the secondary page; Formulate a segmented playback strategy based on the video duration distribution characteristics corresponding to the target URL; Loading the target URL in the playback window and performing a non-continuous playback operation according to the segmented playback strategy; Monitor DOM structure changes in real time during playback, capturing triggered pop-up ads and hidden page elements; Record dynamic page data and abnormal elements triggered by playback, and generate page processing results including timing behavior.
6. The method for detecting illegal content in a free online streaming service according to claim 5, wherein: Extracting multimodal content of text, image, and video stream from the page processing result and generating feature data of each modality, including: Parsing dynamically loaded images and pop-up window content based on the dynamic page data and abnormal elements in the page processing result; Separate illegal links and their disguised page elements from the parsed content; For the separated illegal links and disguised elements, the following feature extraction operations are performed in parallel: Extract the DOM tree hierarchical relationship as HTML structural features, parse the link jump path to generate URL topology features, capture the link anchor text and advertising copy to generate text semantic features, and identify visual objects in disguised elements to generate image visual features.
7. The method for detecting illegal content in a free online streaming service according to claim 6, wherein: The feature data of each modality is integrated and analyzed, and the illegal content detection model outputs the detection results, including: Construct cross-modal association rules based on the extracted HTML structural features, URL topological features, text semantic features, and image visual features; According to the association rules, the four types of features are mapped to a unified dimensional space to generate a joint feature vector; Inputting the joint feature vector into a pre-trained multimodal illegal content detection model and analyzing the feature contribution through a dynamic weight allocation mechanism; Based on the matching results of feature contribution and historical violation pattern library, the illegal content probability value and disguise type mark are output.
8. A system for detecting illegal content in a free online streaming service, characterized in that: include: An acquisition module is used to search and crawl mainstream media platforms based on a preset keyword set to obtain a target mainstream media platform set; A screening module, configured to perform hierarchical processing and classification on the target mainstream media platform set to generate a hierarchical and classified platform set; Obtain content data of corresponding website platforms based on the hierarchically classified platform set, and select target platforms by calculating content differences; The processing module is used to simulate the video playback behavior on the target platform, parse and process the playback page, and generate a page processing result; extract the multimodal content of text, images and video streams from the page processing result to generate feature data of each modality; perform fusion analysis on the feature data of each modality, and output the detection result through the illegal content detection model.
9. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.