AI content review method, system and equipment based on browser plug-in and medium
By combining browser plugins with adversarial training and model distillation, an AI detection model has been developed that addresses the issues of low efficiency, false positives and false negatives, and poor adaptability in AI-generated content review. This results in efficient and accurate AI content review, adapting to diverse website content and ensuring the authenticity and originality of the content.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-10
AI Technical Summary
In existing technologies, the identification and review of AI-generated content suffers from low efficiency, numerous false positives and false negatives, and a lack of versatility and adaptability, making it difficult to meet the needs of large-scale content review.
AI content moderation is achieved through a browser plugin. It employs an AI detection model that combines adversarial training and model distillation. By combining the sensitivity level set by the user with the industry scenario, it collects website feature data in real time for adaptive detection, directly labels the AI-generated areas, and outputs the moderation results.
It achieves efficient and accurate AI content review, reduces false positives and false negatives, provides a convenient user experience, adapts to diverse website content, and ensures the authenticity and originality of the content.
Smart Images

Figure CN121836607A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to AI content review methods, systems, devices and media based on browser plugins. Background Technology
[0002] With the rapid development of artificial intelligence technology, AI-generated content (AIGC) is becoming increasingly popular on various websites and platforms. AIGC, with its efficiency and high quality in certain scenarios, has brought numerous conveniences to information dissemination and content creation. However, in specific application scenarios, such as news reporting, academic research, and creative works presentations, the authenticity and originality of the content are crucial, making the identification and review of AIGC particularly critical.
[0003] Currently, content moderation mainly relies on manual review or tools provided by specific platforms. Manual review has significant drawbacks, requiring substantial investment of human resources and time. Faced with massive amounts of web content, manual review is not only inefficient but also prone to errors due to fatigue, making it unsuitable for large-scale content moderation.
[0004] Existing platform-specific content moderation tools also have numerous problems. Firstly, these tools lack versatility, only applicable to specific platforms. Users needing to moderate web content across different platforms must frequently switch between tools, a cumbersome process that significantly impacts user experience. Secondly, these tools are clearly inadequate in terms of AI content detection accuracy, with false positives and false negatives being common. False positives incorrectly classify genuine content as AI-generated content, interfering with normal content use; false negatives allow some AI-generated content to evade moderation, failing to effectively guarantee the authenticity and originality of the content.
[0005] Furthermore, existing technologies often fail to adapt to the content characteristics of different websites during the review process. Different websites vary in content style and presentation, making it difficult for a uniform review model to accommodate diverse website content, further limiting the accuracy and effectiveness of the review process. Summary of the Invention
[0006] The purpose of this invention is to provide an AI content review method, system, device, and medium based on a browser plugin, which enables one-click review of web page content without the need for manual review, greatly saving time and human resources. It can quickly process a large amount of web page content, meet the needs of large-scale content review, and solve at least one of the aforementioned problems of the prior art.
[0007] In a first aspect, the present invention provides an AI content review method based on a browser plugin, the method specifically comprising: Install and activate the browser plugin to obtain the webpage content of the target webpage. Based on the webpage content, select the review intensity configuration according to the sensitivity level set by the user; Based on the review intensity configuration, the review strategy is configured according to the industry or scenario preset detection parameters by matching the review strategy template library; Based on the censorship strategy, an AI detection model trained adversarially and distilled by the model is used to analyze whether web page content is generated by AI. Based on the output of the AI detection model, the AI-generated areas are directly annotated on the target webpage through a browser plugin, and the review results are output. During the analysis process of the AI detection model, the model parameters are dynamically adjusted based on the website feature data collected in real time by the browser plugin, and adaptive detection is performed on different website content features.
[0008] Secondly, the present invention provides an AI content review system based on a browser plugin, the system specifically comprising: The first review module is used to install and activate the browser plugin, and obtain the webpage content of the target webpage based on the browser plugin; The second review module is used to select the review intensity configuration based on the webpage content and the sensitivity level set by the user. The third review module is used to configure review strategies based on review intensity configuration, by matching the review strategy template library and according to the preset detection parameters of the industry or scenario. The fourth review module is used to analyze whether web page content is generated by AI based on the review strategy, using an AI detection model that has been adversarially trained and model distilled. The fifth review module is used to directly annotate the AI-generated areas on the target webpage based on the output of the AI detection model, and output the review results. The sixth review module is used to dynamically adjust model parameters based on website feature data collected in real time by browser plugins during the analysis process of the AI detection model, and to perform adaptive detection of different website content features.
[0009] Thirdly, the present invention provides a computer device, including: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the AI content review method based on a browser plugin as described in any of the above methods.
[0010] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the AI content review method based on a browser plugin as described in any of the above methods.
[0011] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention enables one-click review of web page content through a browser plugin, eliminating the need for manual review and greatly saving time and manpower. It can quickly process large amounts of web page content and meet the needs of large-scale content review.
[0012] 2. This invention utilizes the universality of browser plugins, allowing it to be used on any webpage. Users do not need to switch tools between different platforms, making it convenient to operate and providing users with a smooth and consistent review experience.
[0013] 3. This invention employs an AI detection model that has undergone adversarial training and model distillation, combined with a review strategy configured based on multiple factors such as webpage content, user settings, and industry scenarios. This effectively improves the accuracy of AI content detection, reduces false positives and false negatives, and more reliably ensures the authenticity and originality of the content.
[0014] 4. This invention dynamically adjusts model parameters based on website feature data collected in real time by browser plugins, enabling adaptive detection of different website content features, adapting to diverse website content, and further improving the accuracy and effectiveness of review.
[0015] 5. This invention can not only directly annotate AI-generated areas on target web pages, but also output content classification and review results together, providing users with more comprehensive and detailed information to facilitate subsequent processing and decision-making. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating an AI content review method based on a browser plugin, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of an AI content review system based on a browser plugin, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0018] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0019] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0020] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0021] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0022] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0023] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0024] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating an AI content review method based on a browser plugin, according to an embodiment of the present invention, is shown below in detail: S101, Install and activate the browser plugin, and obtain the webpage content of the target webpage based on the browser plugin.
[0025] In this embodiment, the user first needs to install the browser plugin of this invention in their browser. Taking a common browser as an example, the user opens the browser's app store (e.g., Google Chrome Web Store, Firefox Add-ons Manager, etc.). In the app store's search bar, enter the plugin name or keywords related to this invention to find the corresponding browser plugin through the search function. After finding the plugin, the user clicks the "Install" button on the plugin details page, and the browser will automatically download and install the plugin. After installation, the installed plugin can be seen in the browser's plugin management interface, and the plugin status is displayed as "Enabled". If the plugin is not automatically enabled, the user can manually click the enable button next to the plugin icon to ensure that the plugin is in a running state.
[0026] After the plugin is installed and enabled, the user launches their browser and opens any target webpage. At this point, the user can activate the plugin by clicking the plugin icon in the browser toolbar. Clicking the plugin icon will bring up an interface in the browser, allowing the user to configure settings and display the plugin's running status. Once activated, the plugin establishes a communication connection with the browser to obtain information about the currently accessed webpage.
[0027] After the plugin starts and establishes a communication connection with the browser, it begins to acquire the content of the target webpage. The plugin first sends a request to the browser to retrieve information about the current webpage. Upon receiving this request, the browser returns the HTML source code and other relevant information of the currently accessed target webpage to the plugin. After receiving the webpage information returned by the browser, the plugin parses this information. During parsing, the plugin identifies various elements in the webpage, such as text content, image links, video links, and table data. For text content, the plugin extracts and organizes the text from different paragraphs, headings, etc., according to the webpage's hierarchical structure and tag rules. For multimedia content such as images and videos, the plugin obtains their corresponding link addresses for further analysis and processing. Through these steps, the plugin successfully acquires the complete content of the target webpage, providing foundational data for subsequent AI content review.
[0028] S102, based on webpage content, selects the review intensity configuration according to the sensitivity level set by the user.
[0029] In this embodiment, after the browser plugin is launched and successfully obtains the webpage content of the target webpage, the plugin will provide the user with sensitivity level setting options on its interface. These options are typically presented in an intuitive graphical interface, such as a slider, with one end labeled "Low Sensitivity" and the other end labeled "High Sensitivity," with multiple scales corresponding to different sensitivity levels. Users can select the appropriate sensitivity level by dragging the slider according to their desired level of scrutiny of the webpage content. Besides a slider, a drop-down menu can also be used, listing different sensitivity levels such as "Low," "Medium," and "High" for the user to choose from. After the user completes their selection, the plugin will record the sensitivity level information set by the user.
[0030] The plugin pre-configures a set of rules for configuring review intensity corresponding to different sensitivity levels. These rules are based on the analysis and research of content review needs in different application scenarios. For example, when a user selects the "low sensitivity" level, the corresponding review intensity configuration will be relatively lenient, mainly focusing on content that is obviously generated by AI and has prominent features, such as text paragraphs with overly regular formatting, monotonous language style, and lack of emotional expression. At this time, some subtle and less obvious traces of AI generation may be ignored during the review process to reduce the possibility of misjudgment and improve review efficiency. When a user selects the "medium sensitivity" level, the review intensity will be moderate, and a more in-depth analysis will be conducted on the language logic and semantic coherence of the webpage content to identify AI-generated content with certain subtleties. When a user selects the "high sensitivity" level, the review intensity will reach its highest level. Not only will the language features of the text be strictly analyzed, but also multimedia content such as images and videos on the webpage will be reviewed in detail to check for signs of AI generation, such as whether the pixel distribution of images and the inter-frame changes of videos conform to the characteristics of natural generation.
[0031] After obtaining the user's set sensitivity level, the plugin automatically selects a matching review intensity configuration based on the pre-defined association rules. Specifically, the plugin accesses its internal association rule database, using the user's sensitivity level as a query condition to find the corresponding review intensity configuration. Once a matching configuration is found, the plugin applies it to subsequent review processes. For example, if the user sets a "high sensitivity" level, the plugin selects a set of strict review parameters and rules, including more detailed grammatical and semantic analysis of the text and deeper feature extraction and comparison of multimedia content, to ensure the most accurate identification of AI-generated parts in the webpage content. In this way, the plugin can flexibly adjust the review intensity according to different user needs, improving the targeting and effectiveness of the review.
[0032] S103, based on the review intensity configuration, configures the review strategy by matching the review strategy template library and according to the industry or scenario preset detection parameters.
[0033] In this embodiment, a comprehensive review strategy template library is pre-built. This template library covers review strategy templates for various industries and scenarios. For example, for the news reporting industry, a large amount of news web page content is collected, and its characteristics such as language style, structural features, and information sources are analyzed to develop review strategy templates specifically for news reporting. These templates specify in detail the content elements that need to be focused on in news content review, such as the objectivity of the title, the logical coherence of the body text, and the accuracy of citations, as well as the corresponding detection methods and judgment criteria. For academic research scenarios, the content characteristics of academic papers and research reports are studied to develop review strategy templates for academic content, including review requirements for the standardization of references, the rationality of research methods, and the scientific nature of conclusions. At the same time, other common industries and scenarios, such as creative work displays and commercial promotions, are also considered, and corresponding review strategy templates are developed and stored in the review strategy template library for subsequent matching and use.
[0034] Once the review intensity configuration is determined based on the webpage content, the plugin will perform a detailed analysis of that configuration. Review intensity configurations are typically divided into three levels: low, medium, and high, each corresponding to a different level of review rigor. The plugin analyzes the specific parameters and requirements involved in the review intensity configuration. For example, low-intensity review might primarily focus on obvious AI-generated features in the webpage content, such as repetitive sentence patterns and unnatural word collocations; medium-intensity review, building upon low-intensity review, further analyzes the logical structure and semantic coherence of the content; high-intensity review conducts a comprehensive and in-depth review of the webpage content, including analysis of multimedia content and detection of potential AI-generated algorithm traces. By analyzing the review intensity configuration, the plugin can clarify the general direction and focus of the subsequent review strategy.
[0035] The plugin performs a matching operation in the review strategy template library based on the parsed review intensity configuration. It uses the key parameters and requirements in the review intensity configuration as matching conditions to search for the most suitable review strategy template in the library. For example, if the review intensity is configured as high and the target webpage belongs to an academic research scenario, the plugin will search the template library for a high-intensity review strategy template specifically for academic research. During the matching process, the plugin comprehensively considers multiple factors, such as industry characteristics, scenario requirements, and review intensity, to ensure that the found template can best meet the review requirements of the current webpage content. If multiple partially matching templates exist, the plugin will select the one that most closely matches the review intensity configuration and webpage scenario according to preset priority rules.
[0036] After matching a suitable review strategy template, the plugin further refines the review strategy configuration based on preset detection parameters for the industry or scenario. Different industries and scenarios have different characteristics and needs, thus requiring the setting of corresponding detection parameters to ensure the accuracy and effectiveness of the review. For example, in a news reporting scenario, preset detection parameters might include evaluation indicators for the reliability of news sources and criteria for judging the timeliness of news. The plugin adjusts and improves the matched review strategy template based on these parameters. If the review strategy template specifies the review of news headlines, the plugin sets detection parameters such as headline length and keyword frequency based on the characteristics of the news industry to ensure that the headlines comply with the norms of news dissemination. In an academic research scenario, preset detection parameters might involve strict requirements for reference formatting and methods for verifying the authenticity of experimental data. The plugin configures the relevant parts of the review strategy in detail based on these parameters, such as specifying that when reviewing academic papers, it is necessary to check whether the references are marked according to specific academic norms and whether the analysis of experimental data conforms to scientific methods. By configuring review strategies based on pre-set detection parameters for the industry or scenario, the review strategies can be made more aligned with actual application needs, thereby improving the quality and efficiency of the review.
[0037] S104, based on censorship strategies, uses an AI detection model that has been adversarially trained and model distilled to analyze whether web page content is generated by AI.
[0038] In this embodiment, after obtaining the review strategy, the browser plugin performs in-depth analysis. The review strategy includes detailed review points and judgment criteria for different types of web page content (such as text, images, and videos). The plugin extracts these key parameters. For example, for text content, it extracts review parameters related to language style (such as sentence complexity and vocabulary frequency) and logical structure (such as paragraph connection and argumentation); for image content, it extracts parameters related to image clarity, color distribution, and the rationality of object shapes; and for video content, it focuses on parameters such as the naturalness of inter-frame changes and the synchronization between audio and video. These parameters serve as important bases for subsequent AI detection model analysis, ensuring that the model can perform detection according to the established review direction and standards.
[0039] The AI detection model used in this embodiment has undergone special adversarial training and model distillation. During the adversarial training phase, the model learns adversarially against the generator in a Generative Adversarial Network (GAN). The generator continuously generates various simulated AI-generated content, attempting to deceive the detection model, while the detection model continuously adjusts its parameters to improve its ability to distinguish between real and AI-generated content. This adversarial training enables the model to learn some latent features and patterns of AI-generated content, enhancing its ability to recognize complex AI-generated content.
[0040] Model distillation involves transferring the knowledge and experience of a large, complex teacher model to a small, efficient student model. Teacher models typically possess high accuracy and powerful feature extraction capabilities, but they consume significant computational resources. Through model distillation, student models can maintain high accuracy while reducing computational load and model size, thus improving detection efficiency. During the preparation phase, it is ensured that the AI detection model has undergone sufficient adversarial training and model distillation, and is in optimal usability to accurately and efficiently analyze web page content.
[0041] Before inputting webpage content into the AI detection model, the browser plugin preprocesses the acquired content. For text content, it performs operations such as word segmentation, stop word removal, and stemming to transform the text into a format suitable for model processing. For example, long texts are segmented into appropriate paragraphs or sentences so that the model can better analyze their linguistic features. For image content, it performs operations such as resizing and color space conversion to meet the model's input requirements. It also performs basic feature extraction on images, such as calculating image histograms and extracting edge features, providing more auxiliary information for the model. For video content, it performs frame extraction, breaking the video down into a series of image frames, and then performing similar preprocessing operations on each frame as on images. Furthermore, it extracts audio information from the video and performs audio feature extraction, such as Mel-frequency cepstral coefficients (MFCC), so that the model can comprehensively analyze multiple aspects of the video.
[0042] The pre-processed webpage content is input into the AI detection model according to the parameters extracted from the review strategy. Simultaneously, to enable the model to more accurately analyze whether the webpage content was generated by AI, relevant parameters from the review strategy are fused with the features of the pre-processed webpage content. For example, if the review strategy emphasizes the frequency of specific words in the text's language style, then the use of these specific words will be highlighted when inputting text features into the model. For image content, if the review strategy focuses on the morphological rationality of objects in the image, then features related to object morphology (such as the object's outline and proportions) will be input into the model along with other image features. Through this feature fusion, the model can more comprehensively consider the requirements of the review strategy, improving the targeting of detection.
[0043] After receiving the input webpage content features and the fused review strategy parameters, the AI detection model begins analysis. Multiple neural network layers within the model work collaboratively to extract and transform the input features layer by layer. For example, convolutional neural network (CNN) layers extract features from images and video frames, identifying objects, textures, and other information within the images; recurrent neural networks (RNNs) or their variants (such as Long Short-Term Memory (LSTM) and gated recurrent units (GRUs) process text and audio sequences, analyzing their temporal features and semantic information. During the analysis, the model uses knowledge learned from previous adversarial training to determine whether the input content exhibits AI-generated features. For instance, if the text contains numerous repetitive sentence patterns, unnatural word collocations, or if the shapes of objects in the image do not conform to natural laws, the model will consider this content potentially AI-generated.
[0044] After analysis by the model, an analysis result is generated regarding whether the webpage content was generated by AI. This result is usually presented as a probability value, indicating the likelihood that the webpage content was generated by AI. The browser plugin sets a threshold; when the probability value in the analysis result exceeds this threshold, the webpage content is determined to be AI-generated; when the probability value is below the threshold, the webpage content is determined to be real content. For example, if the threshold is set to 0.7, and the model analysis yields a probability of 0.8, the webpage content is determined to be AI-generated; when the probability is 0.6, it is determined to be real content. Simultaneously, the plugin also records key information during the model analysis process, such as which features were analyzed that led to the final judgment, for subsequent interpretation and verification of the review results.
[0045] This embodiment can accurately analyze whether web page content is generated by AI based on the review strategy using an AI detection model that has undergone adversarial training and model distillation, providing effective technical support for ensuring the authenticity and originality of web page content.
[0046] S105, based on the output of the AI detection model, directly annotates the AI-generated areas on the target webpage through a browser plugin and outputs the review results.
[0047] In this embodiment, after the AI detection model completes its analysis of the target webpage content, it generates corresponding output results. These output results include information determining whether the webpage content was generated by AI, as well as relevant data such as the specific location and features of the AI-generated content. The browser plugin receives these output results in real time through a communication interface established with the AI detection model. This communication interface ensures the stability and accuracy of data transmission, enabling the plugin to obtain the latest detection information promptly.
[0048] After receiving the output from the AI detection model, the browser plugin performs in-depth analysis. It extracts key information about the AI-generated regions from the output, such as their start and end positions on the webpage, and the range of rows and columns they occupy. Simultaneously, it acquires feature descriptions related to these regions, such as the linguistic style characteristics of the text (e.g., the presence of numerous repetitive sentence structures, unnatural word collocations, etc.) and the image features of the images (e.g., the presence of abnormal color distribution, illogical object shapes, etc.). Through this analysis, the plugin can clearly define the specific content and scope that needs to be annotated on the target webpage. For example, if the output shows that a paragraph on the webpage is AI-generated text, the plugin will extract the start and end positions of that paragraph, as well as any anomalous features in its linguistic style.
[0049] Based on the parsed annotation information, the browser plugin performs precise positioning on the target webpage. Utilizing the webpage's Document Object Model (DOM) structure, the plugin finds the specific location of the AI-generated area within the webpage through corresponding API interfaces. For text content, the plugin can quickly locate the corresponding text area based on the identifiers of elements such as paragraphs and sentences; for multimedia content such as images and videos, the plugin can locate it using attributes such as element ID and class name. For example, if the AI-generated area is an image on the webpage, the plugin will find the image node in the webpage's DOM tree based on the image element's unique identifier, thus determining its specific location.
[0050] After locating the AI-generated area, the browser plugin will visually annotate it. The plugin can use various annotation methods, such as adding a highlighted border around text areas, changing the text color or background color; for image and video areas, it can add a prominent marker box around them, displaying a prompt message inside the box. The color and style of the annotations can be customized according to user needs to improve the clarity and recognizability of the annotations. For example, users can choose to annotate AI-generated text areas with a red background and display the prompt text "AI Generated" above; for image areas, a yellow dashed box can be used for annotation, with the same prompt message displayed in the corner of the box.
[0051] In addition to labeling AI-generated areas, the browser plugin also generates detailed review results. These results are presented in concise text, including an overall assessment of whether the webpage content was AI-generated, the specific number and location of AI-generated areas, and potential risk warnings. The plugin displays the review results in a designated location on the target webpage, such as the top, bottom, or sidebar. It also provides the ability to export the review results as a text file or image for easy saving and sharing. For example, the review result might read, "Detected, this webpage contains 2 AI-generated content items: the third text and the second image. Use with caution."
[0052] To enhance user experience, the browser plugin also provides an interaction and feedback mechanism. Users can view and confirm the annotation and review results. If they find inaccurate annotations or incorrect review results, they can report the issue to the developer through the feedback portal provided by the plugin. The developer will promptly optimize and improve the plugin based on user feedback to improve the accuracy of annotation and review. For example, users can right-click on the annotation area, select the "Report an issue" option, and then describe the problem in detail in the pop-up feedback window.
[0053] This embodiment can directly annotate AI-generated areas on the target webpage through a browser plugin based on the output of the AI detection model, and output accurate and detailed review results, providing users with convenient and efficient AI content review services.
[0054] S106, during the analysis process of the AI detection model, dynamically adjusts the model parameters based on the website feature data collected in real time by the browser plugin, and performs adaptive detection of different website content features.
[0055] In this embodiment, when a user accesses a target webpage using a browser with a specific browser plugin installed, the plugin immediately initiates the collection of website feature data. The plugin interacts with the browser kernel to acquire various feature information about the webpage. For example, it collects the overall layout features of the webpage, including the positional distribution and size ratio of different sections (such as the navigation bar, main text area, sidebar, etc.); it also acquires the text features of the webpage, such as the font type, font size, color scheme, and paragraph format (such as line spacing, paragraph spacing). For multimedia elements on the webpage, the plugin collects information such as image resolution, color mode, compression format, and video encoding format, frame rate, and duration. Simultaneously, the plugin also records the interactive features of the webpage, such as the response methods of clickable elements (such as button click effects, link redirection methods), and the dynamic loading mechanism of the page (such as whether asynchronous loading is used). This feature data comprehensively reflects the content style and presentation format of the target website, providing rich data for subsequent model parameter adjustments.
[0056] After collecting website feature data, the browser plugin performs preprocessing. First, it cleans the data, removing noise and outliers. For example, if a text font size significantly exceeds the normal range, the plugin identifies it as an outlier and corrects or removes it. Next, it standardizes the data, unifying different types and dimensions of feature data into appropriate ranges. For instance, for image resolution and text font size—two different dimensions—the plugin uses conversion methods to make them comparable in subsequent analysis. Furthermore, the plugin extracts and selects features from a large pool of raw features, identifying those crucial for adjusting model parameters. For example, if analysis reveals that the layout of sections and the font type of text significantly impact AI content detection, the plugin will prioritize retaining these features to improve the efficiency and accuracy of subsequent processing.
[0057] After data preprocessing, the browser plugin establishes a mapping between website features and AI detection model parameters. This mapping is determined based on extensive experimental data and expert experience. For example, experiments have shown that when the color modes of images on a webpage are complex, the AI detection model needs to adjust the parameters of its image feature extraction part to improve its ability to recognize AI-generated content in the image. Specifically, if the webpage image uses a high dynamic range (HDR) color mode, the parameters related to color analysis in the model (such as color saturation threshold, color contrast weight, etc.) need to be adjusted accordingly. Similarly, for text features, if the font type of the webpage text is unusual, the parameters related to text semantic analysis in the model (such as word weight allocation, grammatical structure analysis weight, etc.) also need to be changed. The plugin stores these mappings in the form of a rule base so that they can be quickly and accurately invoked during subsequent parameter adjustments.
[0058] During the AI detection model's analysis of webpage content, the browser plugin dynamically adjusts the model's parameters based on real-time collected and pre-processed website feature data, combined with previously established mapping relationships. The plugin first matches the current webpage's feature data with the mapping relationships in the rule base to find the corresponding model parameters that need adjustment. For example, if the current webpage's layout features show that its content area to sidebar ratio differs from a typical webpage, the plugin will determine, based on the mapping relationship, the parameters related to content area division in the model that need adjustment. Then, the plugin modifies the model's parameters according to the adjustment methods and magnitudes specified in the mapping relationship. The adjusted model parameters are immediately applied to the AI detection model, enabling it to better adapt to the current website's content features, thereby improving the accuracy of detecting AI-generated content.
[0059] After dynamically adjusting the model parameters, the browser plugin verifies the model's performance. The plugin uses pre-prepared test data containing content samples from different types of websites, including both known AI-generated content and real content. The plugin applies the adjusted model to this test data, observing changes in metrics such as accuracy, false positive rate, and false negative rate for AI-generated content. If the verification results show improved accuracy and reduced false positive and false negative rates, the parameter adjustments are effective, and the model better adapts to the content characteristics of different websites. Conversely, if the verification results are unsatisfactory, the plugin re-analyzes website feature data and mapping relationships, further optimizing and adjusting the model parameters until satisfactory detection results are achieved.
[0060] This embodiment can dynamically adjust model parameters based on website feature data collected in real time by browser plugins during the analysis process of AI detection model, and perform adaptive detection of different website content features, thereby improving the accuracy and effectiveness of AI content review.
[0061] In some embodiments, step S101 above, specifically obtaining the webpage content of the target webpage according to the browser plugin, includes: By using browser plugins to intelligently inject scripts, the optimal content script injection strategy is dynamically selected based on the characteristic attributes of the target webpage. Based on the content script injection strategy, the semantic features of the DOM structure of the target webpage are analyzed by a semantically aware content extraction engine to identify and prioritize the extraction of core content areas in the target webpage. During the content extraction process, the DOM change monitoring interface is used to listen for and capture dynamically loaded content in the target webpage, and incremental content is captured selectively. Based on content priority, the extracted core content areas and incremental content are processed and transmitted to the browser plugin's background script through a hierarchical content transmission controller.
[0062] In this embodiment, when a user launches a browser and accesses the target webpage, the browser plugin immediately activates its built-in smart injection script. This smart injection script possesses deep analytical capabilities regarding the target webpage's characteristic attributes, comprehensively evaluating the target webpage from multiple dimensions. For example, it analyzes the complexity of the webpage's code structure to determine whether it uses a simple HTML layout or a complex framework structure (such as React, Vue, etc.); it examines the webpage's interactivity to determine if it contains a large number of dynamic elements (such as animations, real-time data updates, etc.); and it assesses the webpage's security mechanisms to check whether a Content Security Policy (CSP) is enabled. Based on the analysis results of these characteristic attributes, the smart injection script dynamically selects the optimal one from a variety of preset content script injection strategies. For example, if the target webpage is a highly interactive page built using the React framework, the smart injection script might choose an injection strategy that can be deeply integrated into the React lifecycle to ensure the accuracy and stability of content extraction.
[0063] After determining the optimal content script injection strategy, the browser plugin will launch a semantically aware content extraction engine based on that strategy. This engine will perform in-depth analysis of the target webpage's Document Object Model (DOM) structure, mining its semantic features. It will identify the semantic information of each node in the DOM tree, such as determining whether a node is a title, paragraph, list, or image. Through semantic analysis, the content extraction engine can accurately identify the core content areas of the target webpage. For example, in a news webpage, the engine can identify the article title...<h1>The nodes and paragraphs in the text are located The nodes and related images are located The engine identifies nodes and defines these areas as core content regions. Simultaneously, it prioritizes different content regions based on semantic importance, extracting only the content that offers the highest value to the user and best reflects the webpage's theme.
[0064] During content extraction, the target webpage may contain dynamically loaded content, such as data asynchronously loaded via AJAX or elements dynamically generated by JavaScript. To ensure that this important content is not overlooked, the browser plugin utilizes a DOM change monitoring interface to monitor the target webpage in real time. This interface is highly sensitive to any changes in the DOM structure. Once it detects that a new node has been added to the DOM tree, or that the attributes or content of an existing node have changed, it immediately triggers a capture mechanism. When capturing dynamically loaded content, the plugin performs a filtering operation, capturing only incremental content that is relevant to the core content or meaningful for review. For example, on an e-commerce webpage, when a user scrolls the page and triggers the asynchronous loading of the product list, the plugin will capture the newly loaded product information while ignoring some irrelevant pop-up ads or dynamic changes to page decoration elements.
[0065] After extracting and capturing the core content area and incremental content, the browser plugin prioritizes the content and uses a hierarchical content delivery controller to process and transmit it to the plugin's backend script. The controller prioritizes the core content area, ensuring the backend script receives the webpage's key information immediately. Incremental content is transmitted secondary based on its importance and relevance to the core content. During transmission, the controller performs appropriate compression and optimization to reduce data transmission volume and improve efficiency. For example, images are compressed using suitable algorithms, and unnecessary formatting tags are removed from text. The controller also monitors and manages the transmission process to ensure data integrity and accuracy, promptly retransmitting any errors.
[0066] This embodiment can efficiently and accurately obtain the webpage content of the target webpage based on the browser plugin, providing a reliable data foundation for subsequent AI content review.
[0067] In some embodiments, step S102 above, which involves selecting the review intensity configuration based on the webpage content and the sensitivity level set by the user, specifically includes: The sensitivity level set by the user is received through the browser plugin's user interface; By analyzing users' historical review records, feedback data, and operating patterns, user profiles are constructed that include their level of expertise, content preferences, and risk tolerance. Collect contextual factor features and automatically identify the key contextual factors that have the greatest impact on the review intensity configuration through a contextual importance assessment algorithm. The contextual factor features include access time, type of website accessed, topic of content viewed, and device type. By combining sensitivity level, user profile, and key contextual factors, the browser plugin's pre-set review intensity configuration mapping table is invoked to query and retrieve the corresponding review intensity configuration.
[0068] In this embodiment, when a user launches a browser with a specific browser plugin installed, the plugin presents an intuitive and easy-to-use user interface on the browser screen. This user interface features a dedicated sensitivity level setting area, typically presented as a slider, drop-down menu, or radio button. Users can freely set the sensitivity level within this area according to their desired level of scrutiny of webpage content. For example, if a user wishes to rigorously review webpage content to exclude any possible AI-generated or inaccurate content, they can set the sensitivity level to a higher level; conversely, if the user has relatively lenient requirements for content authenticity and only wishes to conduct basic review, they can set the sensitivity level to a lower level. The browser plugin captures the user's actions on the user interface in real time, accurately receiving and recording the sensitivity level information set by the user.
[0069] Browser plugins possess powerful data analysis capabilities, collecting user-related data through various channels to build comprehensive and accurate user profiles. On one hand, the plugin delves into the user's historical review records, understanding their past review practices and handling methods for different types of web content. For example, if a user frequently conducts detailed reviews of academic web content and marks potentially AI-generated references, it indicates a high level of expertise and review requirements in that field. On the other hand, the plugin collects user feedback data, which may include evaluations of review results and suggestions for plugin functionality. By analyzing this feedback data, the plugin can understand user preferences for different types of content; for instance, a user might prioritize the authenticity of news content while having relatively lower review requirements for entertainment content. Furthermore, the plugin observes user operating patterns, such as processing speed when encountering suspected AI-generated content and whether they frequently request secondary verification, thereby assessing the user's risk tolerance. By combining data from all these aspects, the plugin constructs a user profile encompassing expertise, content preferences, and risk tolerance, providing a personalized basis for subsequent review intensity settings.
[0070] During web browsing, the browser plugin collects various contextual factors in real time. These contextual factors cover aspects such as access time, website type, content theme, and device type. For example, access time reflects a user's level of attention to content moderation at different times; during the day at work, users may prefer quick and efficient moderation, while at night during leisure time, they may require a higher level of detail. Website types include news websites, academic websites, and social networking sites, and different types of websites have significantly different requirements for content authenticity. Content themes include technology, healthcare, and finance, and different themes have varying degrees of AI generation potential and technical difficulty. Device types include computers, mobile phones, and tablets, and different device screen sizes and operating methods may affect the user's desired level of moderation. After collecting these contextual factors, the plugin automatically analyzes them using a context importance assessment algorithm. This algorithm comprehensively considers the correlation and influence between various factors and moderation intensity, identifying the key contextual factors that have the greatest impact on moderation intensity configuration. For example, algorithm analysis revealed that when users access social networking sites on their mobile devices at night, the requirement for the authenticity of the content is relatively low. In this case, the type of website visited and the type of device become key contextual factors.
[0071] After obtaining the user's sensitivity level, constructing a user profile, and identifying key contextual factors, the browser plugin invokes its internally pre-set review intensity configuration mapping table. This mapping table, derived from extensive experimentation and data analysis, meticulously records the review intensity configurations corresponding to different combinations of sensitivity levels, user profile characteristics, and key contextual factors. The plugin inputs the user's sensitivity level, the user's expertise level, content preferences, and risk tolerance information from the user profile, along with the identified key contextual factors, into the review intensity configuration mapping table as query criteria. Through precise matching and querying, the plugin can retrieve the corresponding review intensity configuration from the mapping table. For example, if the user's sensitivity level is high, the user profile indicates a high level of expertise in the academic field and a low risk tolerance, and the key contextual factors include accessing academic websites and using computer devices, then the plugin will retrieve the corresponding strict review intensity configuration from the mapping table. This configuration may include more detailed AI detection algorithms, higher false positive tolerance, etc., to ensure comprehensive and in-depth review of webpage content.
[0072] This embodiment can accurately select the appropriate review intensity configuration based on the sensitivity level set by the user according to the webpage content, providing users with personalized and efficient AI content review services.
[0073] In some embodiments, step S103 above, which involves configuring the review strategy based on the review intensity configuration by matching the review strategy template library and configuring the review strategy according to the industry or scenario preset detection parameters, specifically includes: Based on the review intensity configuration, the target webpage is automatically identified by analyzing multiple feature dimensions of the target webpage through an industry identifier. Based on the industry type and review intensity configuration, multi-level matching retrieval is performed in the preset strategy template library to obtain a set of candidate templates with suitability scores; Based on the candidate template set, the technical parameters of the candidate templates are intelligently weighted and fused using a reinforcement learning fusion engine to generate an optimized parameter set adapted to the current scenario; Based on the optimized parameter set, a complete review strategy is constructed that includes industry characteristics, intensity requirements, and scenario adaptability.
[0074] In this embodiment, once the browser plugin obtains the review intensity configuration information, it activates an industry identifier to analyze the target webpage. The industry identifier possesses multi-dimensional feature analysis capabilities, comprehensively considering multiple feature dimensions of the target webpage. For example, from the perspective of the webpage's content structure, financial industry webpages typically contain detailed financial data reports, market trend charts, and professional financial terminology; medical industry webpages contain a wealth of medical knowledge, disease diagnosis information, and drug instructions. From the perspective of webpage interaction design, e-commerce industry webpages often feature rich product displays, shopping cart functions, and online payment entry points; education industry webpages may include course lists, online learning modules, and teacher evaluation systems. Furthermore, the industry identifier also analyzes the webpage's metadata information, such as keywords in the webpage title and business areas mentioned in the description. By comprehensively analyzing these different dimensions of features, the industry identifier can automatically and accurately identify the industry type of the target webpage, providing a basis for subsequent review strategy configuration.
[0075] After successfully identifying the industry type of the target webpage, the browser plugin combines the industry type information with the previously acquired review intensity configuration and performs a multi-level matching search operation in a pre-set strategy template library. The strategy template library pre-stores a large number of review strategy templates for different industries and review intensity levels, each containing specific detection parameters and review rules. The multi-level matching search process first performs an initial screening based on industry type to find all templates related to the target webpage's industry; then, it combines the review intensity configuration to further filter out templates that meet the current review intensity requirements. After these two layers of screening, a set of candidate templates with a suitability score is obtained. The suitability score is determined based on the degree of matching between the template and the current industry type and review intensity configuration; a higher score indicates that the template is more suitable for the current scenario. For example, for the financial industry with a high review intensity, templates that include parameters such as detailed financial data authenticity detection and risk assessment model review will receive a higher suitability score.
[0076] After obtaining the candidate template set, the browser plugin invokes a reinforcement learning fusion engine to intelligently weight and fuse the technical parameters of the candidate templates. The reinforcement learning fusion engine employs advanced reinforcement learning algorithms, enabling it to continuously adjust the parameter fusion strategy based on historical data and real-time feedback. During the fusion process, the fusion engine comprehensively considers the suitability score of each candidate template, the rationality of its technical parameters, and the mutual influence between different parameters. For example, if a candidate template has unique advantages in financial data detection but is relatively weak in user interaction review, the reinforcement learning fusion engine will assign higher weights to the financial data detection-related parameters in that template, while appropriately reducing the weights of the user interaction review parameters, based on the overall review requirements. Through this intelligent weighted fusion method, the fusion engine can generate an optimized parameter set adapted to the current scenario. This parameter set fully utilizes the advantages of each candidate template while avoiding the limitations of a single template.
[0077] Finally, based on the generated set of optimized parameters, the browser plugin constructs a complete review strategy that includes industry characteristics, intensity requirements, and scenario adaptability. This review strategy specifies in detail the review content and standards for each target webpage. For example, regarding industry characteristics, for a medical industry webpage, the review strategy will focus on the accuracy of medical information and the compliance of drug instructions; regarding intensity requirements, based on the previously determined review intensity configuration, high-intensity reviews may employ stricter detection algorithms and lower false positive tolerance; regarding scenario adaptability, considering that users may be quickly browsing webpages on mobile devices, the review strategy will optimize the review process, reduce unnecessary interaction steps, and improve review efficiency. By constructing such a complete review strategy, the browser plugin can provide accurate and effective review services for webpage content in different industries, with different review intensities, and in different scenarios.
[0078] This embodiment can configure appropriate review strategies for different industries or scenarios based on the review intensity and by matching the review strategy template library, thereby improving the accuracy and efficiency of AI content review.
[0079] In some embodiments, step S104 above, which involves analyzing whether webpage content is generated by AI using an AI detection model trained on adversarial principles and model distillation based on a review strategy, specifically includes: Based on the latest AI generation technology that simulates web page content using generative adversarial networks, diverse adversarial samples are dynamically generated and mixed with real content to form an enhanced training dataset. The initial AI detection model is iteratively adversarially trained using the augmented training dataset to obtain the target AI detection model. The target AI detection model is used as the teacher model, and model distillation is performed on the teacher model to obtain a lightweight AI detection model. A lightweight student model is used to detect the content of the target webpage and calculate the probability score that the webpage content is generated by AI. By combining probability scores and confidence thresholds in the review strategy, a content generation attribute determination result is generated.
[0080] In this embodiment, a Generative Adversarial Network (GAN) is used to simulate the latest AI generation technologies in the current web content domain. The GAN consists of a generator and a discriminator. The generator is responsible for generating simulated AI-generated web content, while the discriminator attempts to distinguish the generated content from real web content. Through adversarial competition between the two, the generator continuously optimizes its generation strategy to produce more realistic and diverse adversarial examples. These adversarial examples cover various possible AI generation styles and features, such as different text writing styles and image generation patterns. Simultaneously, a large amount of real web content is collected, and the generated adversarial examples and real content are mixed in a certain proportion to form an augmented training dataset. This augmented training dataset has rich sample types, which can effectively improve the generalization ability of subsequent model training and its adaptability to new AI generation technologies. For example, if a new deep learning-based text generation technology emerges on the market, the GAN can quickly simulate the content generated by this technology and incorporate it into the augmented training dataset, enabling the model to learn this new generation pattern in a timely manner.
[0081] After obtaining the augmented training dataset, the initial AI detection model is iteratively trained adversarially using this dataset. In each training round, samples from the augmented training dataset are input into the initial AI detection model, which predicts whether a sample is AI-generated. Simultaneously, a loss function is calculated based on the prediction results and the sample's true label, and the model's parameters are updated using backpropagation. During training, the generative adversarial network continuously generates new adversarial examples and adds them to the training data, constantly subjecting the initial AI detection model to new challenges. This iterative adversarial training method prompts the model to continuously adjust its detection strategy, improving its ability to recognize various AI-generated content. After multiple rounds of iterative training, training stops when the initial AI detection model's performance metrics on the validation set reach a preset target, resulting in the target AI detection model. This target AI detection model exhibits high accuracy and robustness, effectively detecting different types of AI-generated web page content.
[0082] To meet the requirements of model efficiency and resource consumption in practical applications, the trained target AI detection model is used as the teacher model, and model distillation is performed on it. The purpose of model distillation is to transfer the knowledge from the teacher model to a lightweight student model. Specifically, the lightweight student model learns the output probability distribution of the teacher model instead of directly learning the true labels of the samples. During the distillation process, a soft target (the probability output by the teacher model) is used to guide the training of the student model because it contains more information and can help the student model better learn the decision boundaries of the teacher model. At the same time, an appropriate loss function is used to measure the difference between the student model and the teacher model, and the parameters of the student model are continuously adjusted through optimization algorithms. After model distillation, a lightweight AI detection model is obtained, which maintains high detection performance while having a smaller model size and lower computational complexity, making it suitable for running on resource-constrained devices.
[0083] After obtaining the lightweight AI detection model, it is applied to detect web page content. When a user visits a web page, the browser plugin automatically extracts the web page content, including various forms of information such as text, images, and videos, and inputs it into the lightweight student model. The lightweight student model analyzes and processes the input content, calculating a probability score for whether the web page content is AI-generated through its internal neural network structure. This probability score reflects the model's confidence in whether the web page content is AI-generated; a higher score indicates that the model believes the content is more likely to be AI-generated. For example, if a web page article has a very regular text structure and overly formulaic wording, the lightweight student model may give a high probability score, suggesting that the article may be AI-generated.
[0084] Finally, by combining the probability score calculated by the lightweight student model with the preset confidence threshold in the review strategy, the content generation attribute determination result of the webpage content is generated. The confidence threshold in the review strategy is set according to the actual application scenario and user needs, reflecting the user's requirement for the credibility of the detection results. If the probability score is higher than the preset confidence threshold, the webpage content is determined to be AI-generated; if the probability score is lower than the confidence threshold, the webpage content is determined to be non-AI-generated. For example, in a news website review scenario with extremely high requirements for content authenticity, the confidence threshold can be set higher. Only when the probability score exceeds this higher threshold is the content considered AI-generated, ensuring that genuine news reports are not misjudged. In this way, the generation attribute of webpage content can be accurately determined according to different review strategies and actual needs.
[0085] This embodiment can accurately analyze whether web page content is generated by AI based on the review strategy using an AI detection model that has undergone adversarial training and model distillation, providing strong support for the review of the authenticity and reliability of web page content.
[0086] In some embodiments, in step S105 above, the output result based on the AI detection model is used to directly annotate the AI-generated area on the target webpage through a browser plugin and output the review result, specifically including: Based on the output of the AI detection model, a visual annotation layer with multi-level visual encoding is created on the content of the target webpage through an intelligent annotation engine; The output results are aggregated and analyzed in multiple dimensions by the results aggregator to generate a review report that includes content classification and confidence distribution; The review results are integrated and displayed in the browser plugin's user interface by combining a visual annotation layer and a review report. The system receives user interactions with the visual annotation layer and review report through the user interface, and dynamically adjusts the displayed content and analytical perspective.
[0087] In this embodiment, once the browser plugin receives the output of the AI detection model, it activates the intelligent annotation engine. The intelligent annotation engine first parses the output to determine which parts of the webpage content are identified as AI-generated. Then, based on factors such as the type and confidence level of the AI-generated content, it creates a visual annotation layer on the target webpage content using multi-level visual encoding. For example, high-confidence AI-generated text paragraphs are annotated with a prominent red background and a small label next to them with the text "High-confidence AI-generated"; low-confidence AI-generated image areas are annotated with a light yellow semi-transparent overlay, and a floating prompt box with the words "Low-confidence, possibly AI-generated" is displayed in the corner of the image. Multi-level visual encoding not only allows users to quickly identify AI-generated areas but also conveys the confidence level of the detection results through different visual styles, giving users a more intuitive understanding of the annotation information. For example, in a webpage containing multiple images and a large amount of text, this multi-level visual encoding allows users to quickly locate content that may be AI-generated and judge the credibility of the detection results based on color and prompt information.
[0088] While creating the visual annotation layer, the results aggregator performs multi-dimensional aggregation analysis on the output of the AI detection model. First, the aggregator categorizes content by type, dividing webpage content into different categories such as text, images, and videos, and counting the number and proportion of content judged as AI-generated in each category. Then, it performs confidence analysis on each judgment result, calculating the content distribution across different confidence intervals (e.g., high confidence, medium confidence, low confidence). Based on these analysis results, a detailed review report is generated. The review report clearly lists the specific details of AI generation in each type of content, for example, "In the text section, 5 paragraphs were judged as AI-generated, including 3 with high confidence and 2 with medium confidence; in the image section, 3 images were judged as AI-generated, all with low confidence." The review report also displays the confidence distribution in chart form, such as using bar charts to show the number of contents in different confidence intervals and pie charts to show the proportion of each type of content on the webpage, allowing users to more clearly understand the AI generation status of the webpage content.
[0089] After creating the visual annotation layer and generating the review report, the browser plugin integrates both into the user interface. When a user opens the target webpage, the browser plugin pops up a dedicated review results display panel at the top or side of the page. This panel displays the visual annotation layer on the webpage content in real time, allowing users to scroll or zoom to view annotation information at different locations. Simultaneously, the review report is presented in a clear and easy-to-read format, enabling users to quickly browse the AI generation status and confidence distribution of various types of content. For example, users can see text paragraphs highlighted in red on the webpage in the display panel, and simultaneously find corresponding detailed information in the review report below, understanding the confidence level at which the paragraph was judged as AI-generated and its proportion within the overall text content. This integrated display method eliminates the need for users to switch between webpages and different reports, providing a more convenient and comprehensive way to obtain review results.
[0090] The browser plugin's user interface also features interactive functionality, allowing users to interact with the visual annotation layer and review report, and dynamically adjust the displayed content and analytical perspective. Users can click on the annotation areas on the visual annotation layer to view detailed detection information for that area, such as the specific features generated by the AI and the judgment criteria of the detection model. Simultaneously, users can filter the review report by selecting different content categories or confidence ranges, viewing only content of specific types or confidence levels. For example, if a user is only interested in high-confidence AI-generated text, they can select the "High Confidence" and "Text" options in the review report. In this case, the visual annotation layer in the display panel will only show high-confidence AI-generated text annotations, and the review report will correspondingly only display detailed information for this part of the content. Furthermore, users can optimize their viewing experience by adjusting the size and position of the display panel. Through this dynamic adjustment function, users can flexibly view and analyze review results according to their needs and concerns, improving review efficiency and accuracy.
[0091] This embodiment can directly annotate AI-generated areas on the target webpage through a browser plugin based on the output of the AI detection model, and output detailed and intuitive review results, providing users with convenient and efficient webpage content review services.
[0092] In some embodiments, step S106 above, which involves dynamically adjusting model parameters based on website feature data collected in real time by browser plugins to perform adaptive detection of different website content features, specifically includes: By using the feature acquisition interface of the browser plugin, the URL features, content style and user interaction data of the current website can be obtained in real time to generate a unified multimodal feature vector. Based on multimodal feature vectors, the mapping relationship between website features and model parameters is analyzed through a parameter adjustment decision-maker to generate a parameter adjustment scheme for the current website. Based on the parameter adjustment plan, the model parameters are obtained from the local cache or remote server through the model dynamic loader to optimize the configuration; Based on the optimized configuration of model parameters, the detection threshold, feature weights, and inference strategy of the AI detection model are dynamically updated. The updated AI detection model performs adaptive detection of the current website content, resulting in optimized detection results.
[0093] In this embodiment, the browser plugin is equipped with a dedicated feature collection interface. During the analysis process of the AI detection model, this interface immediately initiates data collection. For URL features, it extracts key information from the website URL, such as the top-level domain, second-level domain, and keywords in the path. This information reflects the website's type and domain; for example, the ".edu" suffix typically indicates an educational institution website, while the ".gov" suffix represents a government agency website. Regarding content style, the feature collection interface analyzes the text content of the webpage, extracting features such as grammatical structure, vocabulary frequency, and topic category. For example, the text of a technology website may contain a large number of technical terms and new technology-related words, while the text of a literary website focuses more on rhetoric and emotional expression. Simultaneously, the interface also collects user interaction data, including user dwell time on the webpage, click behavior, and scroll depth. This data reflects the user's level of interest in the website content and their interaction methods. After collecting these different types of data, they are integrated into a unified multimodal feature vector through specific data preprocessing and fusion methods. This multimodal feature vector integrates a website's URL features, content style, and user interaction data, enabling a comprehensive and accurate description of the website's characteristics. For example, for an e-commerce website, the multimodal feature vector might include ".com" domain information, keyword features in product description text, and user interaction data such as frequent clicks on product images and adding items to the shopping cart.
[0094] After obtaining the multimodal feature vectors, they are input into the parameter adjustment decision-maker. The parameter adjustment decision-maker is a module built on machine learning algorithms, internally storing a large number of pre-trained mapping models between website features and model parameters. These mapping models are derived by analyzing data from a large number of different types of websites and their corresponding optimal model parameters. The parameter adjustment decision-maker analyzes the input multimodal feature vectors, matching and comparing them with the internally stored mapping models. By calculating similarity and relevance, it determines the optimal direction and magnitude of model parameter adjustment corresponding to the current website features. For example, if the multimodal feature vectors indicate that the current website is a news and information website, and user interaction data shows that users are more interested in in-depth reporting articles, the parameter adjustment decision-maker will generate a parameter adjustment scheme for the current website based on these features. This scheme may include increasing the weight of features such as long text and technical terms, adjusting the detection threshold to more accurately identify AI-generated parts in news content, etc.
[0095] Based on the parameter tuning plan generated by the parameter tuning decision-maker, the model dynamic loader begins its operation. The model dynamic loader first checks if a suitable optimized model parameter configuration is already stored in the local cache. This local cache, set up to improve data retrieval efficiency, stores recently used and frequently used model parameter configurations. If a matching optimized parameter configuration is found in the local cache, the model dynamic loader retrieves these parameters directly from the cache. If the required parameter configuration is not found in the local cache, the model dynamic loader establishes a connection with a remote server and downloads the corresponding optimized model parameter configuration from the remote server. The remote server stores a rich variety of model parameter configurations to meet the needs of different types of websites. For example, when the parameter tuning plan requires model parameters specifically optimized for art websites, if the local cache does not contain them, the model dynamic loader will download a model parameter configuration specifically optimized for art websites from the remote server, including adjusted feature weights, detection thresholds, etc.
[0096] After obtaining the optimized model parameters, the AI detection model is dynamically updated. First, the detection threshold is updated, as it's a crucial criterion for determining whether webpage content is AI-generated. Depending on the parameter adjustment plan, if the website's content is more rigorous and professional, the detection threshold might be increased to reduce false positives; conversely, if the content is more casual and colloquial, the threshold might be decreased to improve sensitivity. Next, feature weights are adjusted, as different website features have varying importance in determining AI-generated content. For example, for tech commentary websites, features like the frequency of technical terminology and logical structure might be more important, leading to increased weights for these features; while for entertainment news websites, features like emotional expression and trending words might be more critical, requiring adjusted weights. Finally, the inference strategy is optimized, determining how the model makes judgments and decisions based on the input features. Depending on the website's characteristics, different inference algorithms or adjustments to the inference order and method might be chosen to improve accuracy and efficiency. For instance, for websites with large datasets and complex content, a more efficient inference strategy might be employed to accelerate detection.
[0097] After dynamically updating the parameters of the AI detection model, the updated model is used to adaptively detect content on the current website. The updated model can better adapt to the characteristics of the current website and more accurately identify content that may be AI-generated. During the detection process, the model comprehensively analyzes various forms of content on the webpage, such as text, images, and videos, based on adjusted detection thresholds, feature weights, and inference strategies. For example, for a news article, the model will determine whether the article is AI-generated based on features such as the article's theme, language style, and vocabulary usage, combined with the adjusted parameters. After the detection is completed, optimized detection results are obtained. These results can more accurately reflect the AI-generated content on the current website, providing users with more valuable references. For example, the detection results may specify in detail which paragraphs or images are AI-generated and provide corresponding confidence levels, helping users better judge the authenticity and reliability of the website content.
[0098] This embodiment can dynamically adjust model parameters based on website feature data collected in real time by browser plugins, perform adaptive detection of different website content features, and improve the detection accuracy and adaptability of AI detection models in different website environments.
[0099] Reference Figure 2 An embodiment of the present invention provides an AI content review system 2 based on a browser plugin, wherein the system 2 specifically includes: The first review module 201 is used to install and start a browser plugin, and obtain the webpage content of the target webpage based on the browser plugin. The second review module 202 is used to select the review intensity configuration based on the web page content and the sensitivity level set by the user. The third review module 203 is used to configure review strategies based on review intensity configuration by matching the review strategy template library and according to the preset detection parameters of the industry or scenario. The fourth review module 204 is used to analyze whether web page content is generated by AI based on the review strategy, using an AI detection model that has been adversarially trained and model distilled. The fifth review module 205 is used to directly annotate the AI-generated areas on the target webpage based on the output results of the AI detection model, and output the review results. The sixth review module 206 is used to dynamically adjust the model parameters based on the website feature data collected in real time by the browser plugin during the analysis process of the AI detection model, and to perform adaptive detection of different website content features.
[0100] It is understandable that, such as Figure 1 The content shown in the examples of the AI content moderation method based on browser plugins is applicable to the AI content moderation system examples based on browser plugins. The specific functions implemented in the AI content moderation system examples based on browser plugins are the same as those shown in the examples. Figure 1 The illustrated embodiment of the AI content moderation method based on browser plugins is the same, and the beneficial effects achieved are the same as those shown. Figure 1 The beneficial effects achieved by the browser plugin-based AI content moderation method embodiment shown are the same.
[0101] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0102] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0103] Reference Figure 3 The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the AI content review method based on browser plugins as described in any of the above methods.
[0104] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0105] The processor 301 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0106] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.
[0107] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the AI content review method based on a browser plugin as described in any of the above methods.
[0108] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0109] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0110] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0111] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0112] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. < / h1>
Claims
1. A method of AI content review based on a browser plug-in, characterized by, The method specifically comprises: install and start the browser plug-in, and acquire the web page content of the target web page according to the browser plug-in; based on the web page content, select the review intensity configuration according to the sensitivity level set by the user; based on the review intensity configuration, configure the review strategy according to the detection parameters preset by the industry or scene by matching the review strategy template library; based on the review strategy, use the AI detection model trained by confrontation and model distillation to analyze whether the web page content is generated by AI; based on the output result of the AI detection model, directly mark the AI generated area on the target web page through the browser plug-in, and output the review result; in the analysis process of the AI detection model, dynamically adjust the model parameters based on the website feature data collected by the browser plug-in in real time, and perform adaptive detection on different website content features.
2. The method of claim 1, wherein, The method specifically comprises: through the intelligent injection script of the browser plug-in, dynamically select the optimal content script injection strategy according to the characteristic attributes of the target web page; based on the content script injection strategy, analyze the semantic features of the DOM structure of the target web page through the content extraction engine based on semantic perception, identify and preferentially extract the core content area in the target web page; in the content extraction process, use the DOM change listening interface to listen to and capture the dynamically loaded content in the target web page, and selectively capture the incremental content; according to the content priority, use the hierarchical content transmission controller to transmit and process the extracted core content area and incremental content to the background script of the browser plug-in in a hierarchical manner.
3. The method of claim 1, wherein, The method specifically comprises: through the user interface of the browser plug-in, receive the sensitivity level set by the user; through analyzing the historical review records, feedback data and operation mode of the user, build the user portrait containing professional level, content preference and risk tolerance; collect context factor features, automatically identify the key context factors that have the greatest impact on the review intensity configuration through the context importance evaluation algorithm, and the context factor features include access time, access website type, browsing content theme and device type; combined with the sensitivity level, user portrait and key context factors, call the pre-set review intensity configuration mapping table in the browser plug-in, query and acquire the corresponding review intensity configuration from the review intensity configuration mapping table.
4. The method of claim 1, wherein, The method specifically comprises: based on the review intensity configuration, analyze multiple characteristic dimensions of the target web page through the industry identifier, and automatically identify the industry type to which the target web page belongs; combined with the industry type and the review intensity configuration, perform multi-layer matching retrieval in the preset strategy template library, obtain a candidate template set with an adaptation score; according to the candidate template set, use the reinforcement learning fusioner to intelligently weight and fuse the technical parameters of the candidate templates, and generate an optimized parameter set adapted to the current scene; based on the optimized parameter set, build a complete review strategy containing industry characteristics, intensity requirements and scene adaptability.
5. The method of claim 1, wherein, The AI detection model based on the review strategy is used to analyze whether the webpage content is generated by AI, specifically including: Based on the latest AI generation technology of generative adversarial network simulation webpage content, dynamic generation of diversified adversarial samples, and mixing of adversarial samples and real content to form an enhanced training data set; According to the enhanced training data set, the initial AI detection model is iteratively adversarial trained to obtain a target AI detection model; The target AI detection model is used as a teacher model, and model distillation is performed on the teacher model to obtain a lightweight AI detection model; Through the lightweight student model, the webpage content of the target webpage is detected, and the probability score of the webpage content being AI generated is calculated; The probability score and the confidence threshold in the review strategy are combined to generate a content generation attribute judgment result.
6. The method of claim 1, wherein, The output result of the AI detection model is directly marked on the AI generated area of the target webpage through the browser plug-in, and the review result is output, specifically including: Based on the output result of the AI detection model, the intelligent marking engine creates a visual marking layer with multiple levels of visual coding on the content of the target webpage; Through the result aggregator, the output result is analyzed in multiple dimensions to generate a review report containing content classification and confidence distribution; The visual marking layer and the review report are combined and displayed in the user interface of the browser plug-in; Through the user interface, the user's interactive operation on the visual marking layer and the review report is received, and the display content and analysis perspective are dynamically adjusted.
7. The method according to any one of claims 1 to 6, characterized in that, The browser plug-in collects website feature data in real time to dynamically adjust model parameters and perform adaptive detection of different website content features, specifically including: Through the feature collection interface of the browser plug-in, the URL features, content style and user interaction data of the current website are obtained in real time to generate a unified multi-modal feature vector; Based on the multi-modal feature vector, the parameter adjustment decision maker analyzes the mapping relationship between the website features and the model parameters to generate a parameter adjustment scheme for the current website; According to the parameter adjustment scheme, the model dynamic loader obtains the model parameter optimization configuration from the local cache or the remote server; Based on the model parameter optimization configuration, the detection threshold, feature weight and inference strategy of the AI detection model are dynamically updated; Through the updated AI detection model, adaptive detection of the current website content is performed to obtain an optimized detection result. 8.A browser plug-in based AI content review system, characterized by, The system specifically includes: The first review module is used to install and start the browser plug-in, and obtain the webpage content of the target webpage according to the browser plug-in; The second review module is used to select the review intensity configuration according to the user's set sensitivity level based on the webpage content; The third review module is used to configure the review strategy according to the industry or scene preset detection parameters by matching the review strategy template library based on the review intensity configuration; The fourth review module is used to analyze whether the webpage content is generated by AI using the AI detection model based on the review strategy and adversarial training and model distillation; The fifth review module is used to directly mark the AI generated area on the target webpage through the browser plug-in based on the output result of the AI detection model, and output the review result. The sixth examination module is configured to dynamically adjust model parameters based on website feature data collected in real time by the browser plug-in during the analysis process of the AI detection model, and perform adaptive detection on different website content features.
9. A computer device, comprising: The application relates to a browser plug-in based AI content review method and device. The application relates to a browser plug-in based AI content review method and device.
10. A computer-readable storage medium, characterized in that, The application relates to a browser plug-in based AI content review method and device.