Webpage advertisement detection method based on multi-modal fusion, electronic device, medium
Patent Information
- Application Number
- CN202611133360.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]然而,基于启发式规则的方法依赖人工建立和维护规则库,面对广告域名、链接、关键词及页面样式的频繁变化时更新滞后,且容易因规则覆盖不足产生漏检,或因规则过于宽泛产生误检;该类方法通常只能对单个特征进行匹配,难以刻画广告多个组成节点之间的关联
本发明通过设置分别针对文本数据、图像数据以及HTML属性或资源请求属性的第一、第二和第三广告检测器,使文本、图像、HTML容器及外部资源能够分别采用适配的检测方式进行处理,能够从内容语义、视觉特征及页面属性等多个维度对网页节点进行广告检测,充分利用不同模态检测信息之间的互补性,提高广告检测的覆盖范围。
Smart Images

Figure CN122841016A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of web page advertising detection technology, and particularly relates to a web page advertising detection method, electronic device, and medium based on multimodal fusion. Background Technology
[0002] With the development of internet technology, advertising has become a common way to disseminate information and monetize business in the web content ecosystem and internet business. Web page ads are no longer limited to static images or text in fixed positions, but can be formed by text, images, videos, HTML elements, JavaScript scripts, network resources, and page components such as iframes and canvases, and can be loaded and rendered through asynchronous requests, script execution, and dynamic creation or modification of nodes.
[0003] In such web pages, different components of an advertisement may be scattered across multiple page nodes and external resources. Some nodes themselves may not have obvious advertising semantics, but they participate in advertisement display through loading, creation, insertion, or nesting relationships with other nodes. Therefore, to achieve tasks such as web page ad blocking, content security review, ad compliance analysis, and ad dissemination evidence collection, it is necessary to automate and fine-grainedly identify the dynamic loading process of web pages and ad-related components, and comprehensively utilize different content modalities and page structure relationships to form ad detection conclusions.
[0004] Existing webpage ad detection technologies mainly include heuristic rule-based detection methods, single-content-modality-based detection methods, and methods based on webpage structure or page loading process modeling. Heuristic rule-based methods typically match candidate ads on webpages based on pre-defined rules such as domain name, URL, keywords, DOM attributes, or style features; single-content-modality-based methods typically categorize one type of content from webpage text, images, or screenshots; and methods based on webpage structure or page loading process modeling analyze the webpage's DOM structure, resource requests, or node changes to determine the presence of ads and their related components.
[0005] However, heuristic rule-based methods rely on manually building and maintaining rule bases, which leads to update delays when faced with frequent changes in ad domains, links, keywords, and page styles. They are also prone to missed detections due to insufficient rule coverage or false positives due to overly broad rules. These methods typically only match single features, making it difficult to characterize the relationships between multiple components of an ad. Methods based on a single content modality fail to fully utilize heterogeneous information such as text, images, and page structure, making it difficult to identify ads that are expressed through multiple modalities, have text embedded in images, dynamically generated content, or undergo visual and semantic deformation. Detection results are also easily affected by missing content, modal changes, or adversarial perturbations.
[0006] While methods based on webpage structure or page loading process modeling can reflect the compositional relationships and dynamic behavior of webpages to some extent, relying solely on static page snapshots, DOM structures, or resource request records may miss asynchronous loading, script-driven node creation and modification, cross-iframe associations, and visual content generated using methods such as canvas. Even if a page loading graph is constructed, without further combining the text, image, and other content features of nodes and performing correlation propagation of detection results, it is difficult to extend directly hit local advertising evidence to other nodes that jointly constitute the advertisement, thus making it difficult to cover the complete composition range of advertisements and novel loading forms.
[0007] Therefore, existing technologies still suffer from problems in dynamic webpage scenarios, such as insufficient utilization of multimodal information, incomplete modeling of advertising composition relationships, limited detection coverage, and insufficient detection accuracy and robustness. Summary of the Invention
[0008] To address the shortcomings of existing technologies, this invention provides a web page advertising detection method, electronic device, and medium based on multimodal fusion.
[0009] In a first aspect, embodiments of the present invention provide a web page advertisement detection method based on multimodal fusion, the method comprising: Obtain a first ad detector, a second ad detector, and a third ad detector based on ad filtering rules. The first ad detector is used to detect ads in text data, the second ad detector is used to detect ads in image data, and the third ad detector is used to detect ads based on the HTML attributes corresponding to the HTML container node or the resource request attributes corresponding to the external resource node. Obtain the dynamic loading process of the webpage to be tested; based on the dynamic loading process record, abstract the page entities as nodes and the page loading actions as edges, and construct the page loading graph corresponding to the webpage to be tested; For nodes containing text data in the page loading graph, the text data is input into the first ad detector to obtain the text detection result; for nodes containing image data, the image data is input into the second ad detector to obtain the image detection result; for HTML container nodes or external resource nodes, the corresponding HTML attributes or resource request attributes are input into the third ad detector to obtain the rule detection result. The text detection results, image detection results, and / or rule detection results corresponding to each node are propagated between nodes that are associated through incoming or outgoing edges, thereby determining whether each node participates in the formation of a web page advertisement.
[0010] In a second aspect, embodiments of the present invention provide an electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the above-described multimodal fusion-based web page advertisement detection method.
[0011] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described web page advertisement detection method based on multimodal fusion.
[0012] Fourthly, embodiments of the present invention provide a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-described multimodal fusion-based webpage advertising detection method.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention sets up first, second, and third ad detectors for text data, image data, and HTML attributes or resource request attributes, respectively. This allows text, images, HTML containers, and external resources to be processed using appropriate detection methods. It can detect ads on web page nodes from multiple dimensions such as content semantics, visual features, and page attributes, making full use of the complementarity between different modal detection information to improve the coverage of ad detection.
[0014] This invention propagates the text detection results, image detection results, and / or rule detection results corresponding to each node between related nodes through incoming or outgoing edges. This allows nodes that do not contain obvious advertising text or image features but have loading or structural relationships with advertising nodes to obtain advertising association information. By fusing multimodal detection results with node association relationships in the page loading graph, it can make an overall judgment on web page advertisements composed of multiple page entities, thereby identifying multiple nodes that jointly participate in the formation of web page advertisements and reducing missed detections caused by the lack of features of a single node or the failure of a single detection method. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1A flowchart illustrating the webpage advertisement detection method based on multimodal fusion provided in this embodiment of the invention; Figure 2 This is a schematic diagram of the structure of a second advertising detector for image data provided in an embodiment of the present invention; Figure 3 A schematic diagram of the tag delivery algorithm provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] It should be noted that, unless otherwise specified, the features in the following embodiments and implementation methods can be combined with each other.
[0019] HTML containers are logical structural units used to hold and organize a group of related HTML elements on a webpage. They are used for grouping, laying out, and styling child elements on the page. HTML containers are usually represented by specific tag pairs (such as...). 、 <section> 、 <article>HTML containers are defined by their hierarchical structure, where each container can contain several child elements, text content, or other nested containers. These containers form a modular division of the page. Browsers can then use these HTML containers to style page areas, bind events, and dynamically update content.
[0020] A Uniform Resource Locator (URL) is a sequence of characters used to identify and locate a specific resource on the Internet. It describes information such as the access protocol, host address, port, and resource path of the network resource. URLs typically follow a standard syntax structure, including a protocol scheme, network authority, resource path, and optional query parameters and fragment identifiers. These components are separated by specific delimiters, collectively forming complete resource location information. Browsers can use URLs to initiate network requests, enabling the retrieval of web page content, resource loading, and navigation.
[0021] PageGraph is a graph representation method used to describe the loading and dynamic rendering process of web pages. It constructs a page graph containing web page nodes and their relationships by analyzing page behavior in the browser's HTML parsing engine and JavaScript execution engine. Nodes in a PageGraph can represent DOM elements, text content, network resources, and other objects, while edges represent nesting relationships within HTML containers, node creation relationships, node insertion relationships, and DOM relationships across iframes. By analyzing the node and edge relationships in the PageGraph, it is possible to model and detect web page structure, resource loading processes, and dynamic page behavior.
[0022] The LGC algorithm (Label Propagation with Local and Global Consistency) is a semi-supervised learning algorithm based on graph structures. It utilizes the proximity relationships between nodes in a graph to propagate known label information and predict the category of unlabeled samples. This algorithm considers both local neighborhood relationships and global structural consistency, causing similar nodes to tend to receive the same or similar labels. It is widely used in tasks such as classification, clustering, and graph data analysis.
[0023] like Figure 1 As shown in the figure, this invention provides a web page advertisement detection method based on multimodal fusion, the method comprising the following steps: Step S1: Obtain a first ad detector, a second ad detector, and a third ad detector based on ad filtering rules. The first ad detector is used to detect ads in text data, the second ad detector is used to detect ads in image data, and the third ad detector is used to detect ads based on the HTML attributes corresponding to the HTML container node or the resource request attributes corresponding to the external resource node.
[0024] Specifically, step S1 includes the following sub-steps: Step S101, obtaining the first advertising detector for text data, including: A text training dataset is constructed based on text data labeled with advertising categories. The text data includes text extracted from web pages and text extracted from advertising images through optical character recognition. The first advertising detector is obtained by training a text classification model, including a text encoder (e.g., Sentence Transformer) and a classifier (e.g., Multilayer Perceptron MLP), using the text training dataset.
[0025] Step S102, obtaining a second advertising detector for image data, including: An image training dataset was constructed based on image data labeled with advertising categories; An image classification model, including a residual network and a channel attention module, is used to perform feature learning and classification training on the image training dataset to obtain the second advertising detector; as follows: Figure 2 As shown, this example uses an image classification model based on the ResNet50-SE architecture; The input data for the second advertising detector includes the image to be detected, as well as the image's position and size data within the webpage.
[0026] Step S103, for web page resources and HTML containers within web page, deploy a rule-based third-party ad detector, including: For web page resources (such as external images, scripts, etc.) and HTML containers within web page pages, a detection process is launched outside the browser to run a rule-based third-party ad detector. When the process starts, it loads a set of ad detection rules. For example, the detection process integrates the ad filtering rule parsing engine library libadblockplus, which is equipped with the open-source ad filtering rule sets EasyList and UBlockOrigin provided and maintained by the community. A high-speed channel is established between the page rendering process in the browser kernel and the detection process outside the browser through a communication process; for example, the mojo communication module provided by the Brave browser is used to establish a channel between the rendering process and the detection process.
[0027] It should be noted that this example uses a "rendering-detection" dual-process architecture to achieve low-latency browser-based deployment of a rule-based ad detector, effectively improving the efficiency of the rule detector.
[0028] Step S2: Obtain the dynamic loading process of the webpage to be detected; based on the dynamic loading process record, abstract the page entities as nodes and the page loading actions as edges to construct the page loading graph corresponding to the webpage to be detected. Specifically, in this example, page entities, including the HTML parsing engine, script execution engine, document object model tree root, HTML container, text content, scripts, and external resources, are abstracted as nodes; Page loading actions, including node creation, node insertion, HTML container nesting, resource request, cross-document object model association, and script execution, are abstracted as edges to construct the page loading graph corresponding to the webpage to be detected.
[0029] Step S3: For nodes containing text data in the page loading graph, input the text data into the first ad detector to obtain the text detection result; for nodes containing image data, input the image data into the second ad detector to obtain the image detection result; for HTML container nodes or external resource nodes, input the corresponding HTML attributes or resource request attributes into the third ad detector to obtain the rule detection result.
[0030] Specifically, for HTML container nodes or external resource nodes, the process of inputting the corresponding HTML attributes or resource request attributes into the third-party ad detector to obtain the rule detection results includes: In this example, for HTML container nodes in the page loading graph, the attribute information of the corresponding HTML container is parsed, and the triggering conditions of the element-hidden ad filtering rules are referenced to detect ad containers on the page. These rules detect ad-related HTML containers based on the HTML attributes of graph nodes (e.g., the element ID value is "ad"), and their triggering conditions can be the HTML container's category, ID attribute, CLASS attribute, etc.
[0031] For example, firstly, the tag category, ID attribute, and CLASS attribute corresponding to the HTML container node are read from the page loading graph; then, the element hiding class rules in the ad filtering rule base are parsed, and the CSS selectors are extracted. During the detection process, the node attributes are matched with the CSS selectors. When a node attribute matches any CSS selector, the HTML container node is determined to be an ad node.
[0032] Furthermore, in this embodiment, for external resource nodes with URLs in the page loading graph, the advertising resources are detected by analyzing the URL string corresponding to the resource request and referring to the triggering conditions of resource-blocking advertising filtering rules. These rules are used to detect advertising resources that meet specific URL patterns, request types, or request context conditions; their triggering conditions are typically represented as URL strings in regular expression format.
[0033] For example, firstly, the URL of the requested resource, the resource type, and the URL of the page initiating the request are read from the page loading graph; then, the above information is input into the ad filtering rule engine libadblockplus, and the trigger conditions in the resource blocking class rules are used for detection. When the request URL, resource type, and page URL simultaneously meet the trigger conditions of any filtering rule, the corresponding resource node is marked as an ad resource node.
[0034] Step S4: Propagate the text detection results, image detection results, and / or rule detection results corresponding to each node between nodes associated through incoming or outgoing edges, thereby determining whether each node participates in forming a web page advertisement.
[0035] Specifically, such as Figure 3 As shown, step S4 includes the following steps: Step S401: Determine the initial positive score and initial negative score of each node based on the text detection results, image detection results and / or rule detection results corresponding to each node; construct an initial node score matrix based on the initial positive score and initial negative score of all nodes; wherein, the positive score is used to characterize the confidence that the node belongs to the advertisement, and the negative score is used to characterize the confidence that the node does not belong to the advertisement.
[0036] Furthermore, the page loading graph is represented as a directed graph G=(V,E), where V represents the set of nodes and E represents the set of edges. For each node v∈V, a positive score is defined. With negative scores These scores are quantized by comprehensively considering text detection results, image detection results, and / or rule detection results. An initial node score matrix is constructed based on the positive and negative scores of all nodes in the graph. .
[0037] Step S402: Prune the page loading graph by removing nodes and edges that are obviously irrelevant to the construction of the advertisement, and obtain the pruned page loading graph.
[0038] In this example, the page loading graph is pruned, retaining only nodes and edges related to webpage structure building, resource loading, and script execution: for edges, only edge types representing DOM hierarchy relationships, node creation relationships, node insertion relationships, cross-DOM relationships, and script execution relationships are retained; for nodes, after edge pruning, all isolated nodes without connected edges are deleted.
[0039] Step S403: Construct a graph weight matrix based on the connection relationships and distances between nodes in the page loading graph; iteratively update the positive and negative scores of each node based on the graph weight matrix and the initial node score matrix, so that the text detection results, image detection results and / or rule detection results corresponding to each node propagate between nodes associated through incoming or outgoing edges until the convergence condition is met, and obtain the comprehensive positive and comprehensive negative scores of each node.
[0040] For any two nodes and A weight matrix is constructed based on whether there are direct edges connecting the nodes. Among them, nodes With nodes Weights between The expression is as follows:
[0041]
[0042] In the formula, Represents a node With nodes Weights between Represents a node With nodes The distance between, Represents a node With nodes Whether there is a direct connection between them, sc represents the preset propagation range.
[0043] For any node with a non-zero positive score, within a preset propagation range, perform an influence enhancement operation on its neighboring nodes, as shown in the following expression:
[0044] In the formula, Represents a node Positive scores, Represents a node Positive score.
[0045] It should be noted that this example utilizes the local clustering characteristics of web page advertising-related nodes in the page loading graph. Nodes with non-zero positive scores are identified as suspected advertising source nodes. Starting from these nodes, the example further discovers surrounding nodes related to advertising loading, resource requests, script execution, or page rendering, thereby improving the completeness of advertising node cluster identification.
[0046] The weight matrix is normalized to obtain the normalized matrix S; the expression is as follows:
[0047] In the formula, D represents a diagonal matrix.
[0048] Run the improved LGC algorithm (Learning with Local and Global Consistency) on the pruned page loading graph. In each iteration, let:
[0049]
[0050] In the formula, This represents the weighting coefficient, used to control the ratio between the propagation score and the node's initial score.
[0051] After the LGC algorithm converges iteratively, based on the converged iteration matrix... Determine the overall positive score and overall negative score for each node. For each node... Its overall positive score Overall negative score .
[0052] Step S404: If the overall positive score of a node is greater than the overall negative score, then the node is determined to participate in the formation of web page advertisement.
[0053] In summary, this invention, by setting up first, second, and third advertising detectors for text data, image data, and HTML attributes or resource request attributes respectively, enables text, images, HTML containers, and external resources to be processed using appropriate detection methods. It can detect advertisements on web page nodes from multiple dimensions such as content semantics, visual features, and page attributes, and fully utilize the complementarity between different modal detection information to improve the coverage of advertising detection.
[0054] This invention propagates the text detection results, image detection results, and / or rule detection results corresponding to each node between related nodes through incoming or outgoing edges. This allows nodes that do not contain obvious advertising text or image features but have loading or structural relationships with advertising nodes to obtain advertising association information. By fusing multimodal detection results with node association relationships in the page loading graph, it can make an overall judgment on web page advertisements composed of multiple page entities, thereby identifying multiple nodes that jointly participate in the formation of web page advertisements and reducing missed detections caused by the lack of features of a single node or the failure of a single detection method.
[0055] According to embodiments of the present invention, the present invention also provides an electronic device and a readable storage medium.
[0056] Figure 4 A schematic block diagram of an electronic device that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0057] The electronic device includes a computing unit 101, which can perform various appropriate actions and processes according to a computer program stored in ROM 102 or a computer program loaded into RAM 103 from storage unit 108. RAM 103 may also store various programs and data required for the operation of the electronic device. The computing unit 101, ROM 102, and RAM 103 are interconnected via bus 104. I / O interface 105 is also connected to bus 104.
[0058] Multiple components in the electronic device are connected to the I / O interface 105, including: an input unit 106, such as a keyboard, mouse, etc.; an output unit 107, such as various types of displays, speakers, etc.; a storage unit 108, such as a disk, optical disk, etc.; and a communication unit 109, such as a network card, modem, wireless transceiver, etc. The communication unit 109 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0059] The computing unit 101 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 101 performs the various methods and processes described above. For example, in some embodiments, the web advertising detection method based on multimodal fusion can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 108. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 102 and / or communication unit 109. When the computer program is loaded into RAM 103 and executed by the computing unit 101, one or more steps of the web advertising detection method based on multimodal fusion described above can be performed. Alternatively, in other embodiments, the computing unit 101 can be configured to perform the web advertising detection method based on multimodal fusion by any other suitable means (e.g., by means of firmware).
[0060] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0061] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0062] In the context of this invention, a readable storage medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A readable storage medium can be a machine-readable signal medium or a machine-readable storage medium. A readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0063] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).
[0064] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0065] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0066] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only.
[0067] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.< / article> < / section>
Claims
1. A webpage advertisement detection method based on multimodal fusion, characterized in that, The method includes: Obtain a first ad detector, a second ad detector, and a third ad detector based on ad filtering rules. The first ad detector is used to detect ads in text data, the second ad detector is used to detect ads in image data, and the third ad detector is used to detect ads based on the HTML attributes corresponding to the HTML container node or the resource request attributes corresponding to the external resource node. Obtain the dynamic loading process of the webpage to be tested; based on the dynamic loading process record, abstract the page entities as nodes and the page loading actions as edges, and construct the page loading graph corresponding to the webpage to be tested; For nodes containing text data in the page loading graph, the text data is input into the first ad detector to obtain the text detection result; for nodes containing image data, the image data is input into the second ad detector to obtain the image detection result; for HTML container nodes or external resource nodes, the corresponding HTML attributes or resource request attributes are input into the third ad detector to obtain the rule detection result. The text detection results, image detection results, and / or rule detection results corresponding to each node are propagated between nodes that are associated through incoming or outgoing edges, thereby determining whether each node participates in the formation of a web page advertisement.
2. The webpage advertisement detection method based on multimodal fusion according to claim 1, characterized in that, The process of obtaining the first ad detector includes: A text training dataset is constructed based on text data labeled with advertising categories. The text data includes text extracted from web pages and text extracted from advertising images through optical character recognition. The first advertising detector is obtained by training a text classification model, which includes a text encoder and a classifier, using the text training dataset.
3. The webpage advertisement detection method based on multimodal fusion according to claim 1, characterized in that, The process of obtaining the second ad detector includes: An image training dataset was constructed based on image data labeled with advertising categories; The second advertising detector is obtained by using an image classification model that includes a residual network and a channel attention module to perform feature learning and classification training on the image training dataset. The input data for the second advertising detector includes the image to be detected, as well as the image's position and size data within the webpage.
4. The webpage advertisement detection method based on multimodal fusion according to claim 1, characterized in that, The process of deploying a rule-based third-party ad detector for web page resources and HTML containers within web pages includes: For web page resources and HTML containers within web pages, a detection process is launched outside the browser to run a rule-based third-party ad detector. When the process starts, it loads a set of ad detection rules. A high-speed channel is established between the browser kernel's page rendering process and the browser's external detection process through a communication process.
5. The webpage advertisement detection method based on multimodal fusion according to claim 1, characterized in that, The process of constructing the page loading graph corresponding to the webpage to be detected includes: Abstract page entities, including the HTML parsing engine, script execution engine, document object model root, HTML container, text content, scripts, and external resources, into nodes; Page loading actions, including node creation, node insertion, HTML container nesting, resource request, cross-document object model association, and script execution, are abstracted as edges to construct the page loading graph corresponding to the webpage to be detected.
6. The webpage advertisement detection method based on multimodal fusion according to claim 1, characterized in that, The process of propagating the text detection results, image detection results, and / or rule detection results corresponding to each node between nodes associated through incoming or outgoing edges to determine whether each node participates in the formation of a webpage advertisement includes: Based on the text detection results, image detection results, and / or rule detection results corresponding to each node, determine the initial positive score and initial negative score of each node; construct an initial node score matrix based on the initial positive score and initial negative score of all nodes; where the positive score is used to characterize the confidence that a node belongs to an advertisement, and the negative score is used to characterize the confidence that a node does not belong to an advertisement. Construct a graph weight matrix based on the connection relationships and distances between nodes in the page loading graph; The positive and negative scores of each node are iteratively updated based on the graph weight matrix and the initial node score matrix, so that the text detection results, image detection results and / or rule detection results corresponding to each node are propagated between nodes associated through incoming or outgoing edges until the convergence condition is met, and the comprehensive positive and comprehensive negative scores of each node are obtained. If a node's overall positive score is greater than its overall negative score, then that node is determined to participate in the creation of a webpage advertisement.
7. The webpage advertisement detection method based on multimodal fusion according to claim 6, characterized in that, The process of iteratively updating the positive and negative scores of each node includes: For any node with a non-zero positive score, within a preset propagation range, perform an influence enhancement operation on its neighboring nodes, as shown in the following expression: ; In the formula, Represents a node With nodes Weights between Represents a node With nodes The distance between, Represents a node Positive scores, Represents a node The positive score, sc represents the preset propagation range.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, and the one or more computer programs are executed by the at least one processor to enable the at least one processor to perform the web page advertising detection method based on multimodal fusion as described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the web page advertising detection method based on multimodal fusion as described in any one of claims 1-7.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the web page advertising detection method based on multimodal fusion as described in any one of claims 1-7.