Abnormal webpage identification method and device, electronic equipment, storage medium and computer program product
By using an agent in a multimodal large language model to process webpage information step by step, the problem of low accuracy in identifying abnormal webpages in existing technologies is solved, and efficient and accurate identification of abnormal webpages is achieved.
Patent Information
- Application Number
- CN202411306308.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-20
AI Technical Summary
In existing technologies, clustering algorithms are sensitive to noise, have low accuracy in identifying abnormal web pages, and supervised learning and cluster diffusion methods are difficult to collect sample data, consume high resources, have poor interpretability, and are difficult to effectively identify abnormal web pages.
A multimodal large language model is used to construct a first agent, a second agent, and a third agent, which are used to screen information, collect information, and identify anomalies, respectively. The accuracy is improved through layer-by-layer processing.
It improves the accuracy and precision of abnormal webpage identification, reduces the possibility of false positives, enhances identification efficiency and quality, and reduces labor costs.
Smart Images

Figure CN121705928A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and includes, but is not limited to, an abnormal webpage identification method, apparatus, electronic device, storage medium, and computer program product. Background Technology
[0002] With the development of internet and terminal technologies, browsing web pages to obtain online information has become a daily activity for users. However, some suspicious web pages often require users to fill in personal information, leading to the leakage of users' personal information and poor security. Related technologies typically use clustering algorithms to identify suspicious web pages; however, clustering algorithms are very sensitive to noise, and the presence of noise in the data will reduce the accuracy of suspicious web page identification. Summary of the Invention
[0003] This application provides an abnormal webpage identification method, apparatus, electronic device, storage medium, and computer program product, which can be applied to at least the field of artificial intelligence. This application improves the accuracy of abnormal webpage identification by using an intelligent agent in a multimodal large language model to screen, collect, and identify webpage information.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] This application provides an abnormal webpage identification method, comprising: constructing a first agent, a second agent, and a third agent in a multimodal large language model; inputting webpage information of the webpage to be identified into the multimodal large language model; using the first agent to perform information screening on the webpage information to obtain screening information for abnormal webpage identification; using the second agent to collect information from the webpage to be identified based on the webpage information and the screening information to obtain collected information; using the third agent to perform abnormal webpage identification on the webpage to be identified based on the webpage information and the collected information; and outputting the abnormal webpage identification result through the multimodal large language model.
[0006] This application provides an abnormal webpage identification device, comprising: a construction module for constructing a first intelligent agent, a second intelligent agent, and a third intelligent agent in a multimodal large language model; an input module for inputting webpage information of a webpage to be identified into the multimodal large language model; a screening module for screening the webpage information through the first intelligent agent to obtain screening information for abnormal webpage identification; a collection module for collecting information from the webpage to be identified based on the webpage information and the screening information through the second intelligent agent to obtain collected information; and an identification module for identifying the webpage to be identified based on the webpage information and the collected information through the third intelligent agent, and outputting the abnormal webpage identification result through the multimodal large language model.
[0007] In the above scheme, the device further includes a construction module, used to acquire a first pre-trained model, and a first role prompt information, a first task prompt information, and a first input-output prompt information corresponding to the first pre-trained model; merge the first role prompt information, the first task prompt information, and the first input-output prompt information to obtain first prompt information; and construct the first agent in the multimodal large language model based on the first prompt information and the first pre-trained model.
[0008] In the above scheme, the screening module is further configured to: fill the web page information into the first input / output prompt information to obtain the updated first prompt information; and, through the first pre-trained model in the first agent, screen the web page information in the updated first prompt information based on the first role prompt information and the first task prompt information to obtain screening information for abnormal web page identification.
[0009] In the above scheme, the device further includes an output module for outputting the screening information in accordance with the output format specified in the updated first prompt information.
[0010] In the above scheme, the construction module is further configured to: obtain a second pre-trained model, and a second role prompting information, a second task prompting information, and a second input / output prompting information corresponding to the second pre-trained model; merge the second role prompting information, the second task prompting information, and the second input / output prompting information to obtain second prompting information; and construct the second agent in the multimodal large language model based on the second prompting information and the second pre-trained model.
[0011] In the above scheme, the collection module is further configured to: input the second prompt information, the webpage information, and the screening information into the second intelligent agent to obtain an output result; when the output result indicates that extended information needs to be collected based on the screening information, parse the tool identifier and input data from the output result; collect target extended information through the information collection tool corresponding to the tool identifier according to the input data, and concatenate the target extended information with the screening information to obtain collected information; when the output result indicates that extended information does not need to be collected, determine the screening information as the collected information.
[0012] In the above scheme, the device further includes a splicing module, which is used to extract features from the text data when the input data is text data to obtain a text embedding vector; to search for target text vectors with a similarity higher than a first preset threshold in a pre-built text vector library through the information collection tool; and to splice the information corresponding to the target text vector with the screening information to obtain the collected information.
[0013] In the above scheme, the splicing module is further configured to: when the input data is image data, perform feature extraction on the image data to obtain an image embedding vector; search for a target image vector with a similarity higher than a second preset threshold in a pre-built image vector library using the information collection tool; and splice the information corresponding to the target image vector with the screening information to obtain the collected information.
[0014] In the above scheme, the construction module is further configured to: obtain a third pre-trained model, and third role prompt information, third task prompt information, and third input / output prompt information corresponding to the third pre-trained model; merge the third role prompt information, the third task prompt information, and the third input / output prompt information to obtain third prompt information; and construct the third agent in the multimodal large language model based on the third prompt information and the third pre-trained model.
[0015] In the above scheme, the identification module is further configured to: fill the webpage information and the collected information into the third input / output prompt information to obtain the updated third prompt information; through the third pre-trained model in the third agent, based on the third role prompt information, the third task prompt information and the collected information in the updated third prompt information, perform anomaly identification on the webpage information in the updated third prompt information to obtain anomaly identification results; and determine the abnormal webpage identification result of the webpage to be identified based on the anomaly identification results.
[0016] In the above scheme, the output module is further configured to: output the abnormal webpage recognition result according to the output format specified by the updated third prompt information through the multimodal large language model.
[0017] This application provides an electronic device, including: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the above-described abnormal webpage identification method.
[0018] This application provides a computer program product that stores computer-executable instructions or a computer program, which is used to cause a processor to execute the computer-executable instructions or the computer program to implement the above-mentioned abnormal webpage identification method.
[0019] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs, which are used to cause a processor to execute the computer-executable instructions or computer programs to implement the above-described abnormal webpage identification method.
[0020] The above scheme has the following beneficial effects:
[0021] The abnormal webpage identification method of this application involves, firstly, constructing a first intelligent agent, a second intelligent agent, and a third intelligent agent within a multimodal large language model; then, inputting the webpage information of the webpage to be identified into the multimodal large language model; next, using the first intelligent agent, screening the webpage information to obtain screening information for abnormal webpage identification; then, using the second intelligent agent, collecting information about the webpage to be identified based on the webpage information and the screening information; finally, using the third intelligent agent, identifying the webpage to be identified as abnormal based on the webpage information and the collected information, and outputting the abnormal webpage identification result through the multimodal large language model. Thus, this application uses multiple intelligent agents to progressively process and analyze the webpage information of the webpage to be identified. The first intelligent agent performs preliminary screening, the second intelligent agent collects information, and the third intelligent agent performs anomaly identification. This layered processing improves the accuracy and precision of abnormal webpage identification, reduces the possibility of misjudgment, and ensures that each intelligent agent has a clear division of labor, allowing each agent to focus on a specific function, thereby improving the efficiency and quality of abnormal webpage identification. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating a method for identifying abnormal web pages based on supervised learning, provided by related technologies.
[0023] Figure 2This is a flowchart illustrating the abnormal webpage classification method based on clustering diffusion provided by related technologies;
[0024] Figure 3 This is a schematic diagram of an optional architecture of the abnormal webpage identification system provided in this application embodiment;
[0025] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0026] Figure 5 This is an optional flowchart illustrating the abnormal webpage identification method provided in this application embodiment;
[0027] Figure 6 This is another optional flowchart illustrating the abnormal webpage identification method provided in the embodiments of this application;
[0028] Figure 7 This is a schematic diagram of the process of screening web page information by a first intelligent agent according to an embodiment of this application;
[0029] Figure 8 This is a schematic diagram of the process of information collection through a second intelligent agent provided in an embodiment of this application;
[0030] Figure 9 This is a schematic diagram of the process for identifying abnormal web pages through a third-party intelligent agent, provided in an embodiment of this application.
[0031] Figure 10 This is a schematic diagram of a visual question-answering task provided in an embodiment of this application;
[0032] Figure 11 This is another optional flowchart illustrating the abnormal webpage identification method provided in the embodiments of this application;
[0033] Figure 12 This is a schematic diagram of the process of information screening performed by the first intelligent agent according to an embodiment of this application;
[0034] Figure 13 This is a schematic diagram of the process of information collection by the second intelligent agent provided in the embodiments of this application;
[0035] Figure 14 This is a schematic diagram of the process of a third intelligent agent identifying abnormal web pages provided in an embodiment of this application. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0037] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit the application.
[0038] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0039] Before describing the abnormal webpage identification method provided in the embodiments of this application, the technical terms involved in the embodiments of this application will be explained first:
[0040] 1) Multimodal Large Language Model (MLLM): It has multiple cross-modal capabilities such as image captioning (IC), visual question answering (VQA), and image-text matching (ITM), and is a large-scale deep learning model with more than one billion parameters.
[0041] 2) Agent: An entity with autonomous understanding, planning and decision-making capabilities, and information processing capabilities, built based on a multimodal large language model as the core controller.
[0042] 3) Abnormal web pages: Web pages that impersonate legitimate institutions or company websites to steal users' personal information, account passwords, etc.
[0043] 4) Supervised Learning: Supervised learning is a machine learning method that uses a set of samples of known categories (i.e., training data) to train a model, enabling the model to predict output results from input features. Training data includes input objects and corresponding expected output values (supervisory signals). The model learns the mapping relationship between inputs and outputs by analyzing the training data.
[0044] 5) Clustering: Clustering refers to the process of dividing a set of data into several categories according to a certain similarity measure.
[0045] To better understand the abnormal webpage identification method provided in the embodiments of this application, the abnormal webpage identification methods in related technologies will be described below.
[0046] 1. Anomaly webpage identification method based on supervised learning:
[0047] See Figure 1 , Figure 1 This is a flowchart illustrating a supervised learning-based method for identifying abnormal web pages, provided by related technologies. First, known web page image data is manually labeled to create a dataset 101 containing web page images and labels. Then, a supervised model is trained using dataset 101 through machine learning or deep learning methods to obtain a pre-trained model 102. Finally, the web page image 103 to be labeled is input into the pre-trained model 102 to obtain a recognition result 104. By analyzing the recognition result 104, it can be determined whether the web page image 103 to be labeled is an abnormal web page.
[0048] 2. Anomaly webpage classification method based on clustering diffusion:
[0049] See Figure 2 , Figure 2 This is a flowchart illustrating an abnormal webpage classification method based on cluster diffusion provided by related technologies. First, features are extracted from the webpage image 201 to be identified and the known abnormal webpage image 202 to obtain image features 203. Then, the image features 203 are clustered to obtain clusters 204. Finally, diffusion is performed based on indicators such as the number, proportion, and average cluster distance of abnormal webpage labels in clusters 204 to obtain the abnormal identification result 205 of the webpage image to be identified.
[0050] However, the above-mentioned methods for identifying abnormal web pages have at least the following problems:
[0051] For supervised learning-based methods for identifying abnormal web pages: 1) Supervised models require a large amount of abnormal web page sample data for supervised training. However, in actual business, collecting abnormal web page sample data is very difficult, consuming a lot of time and human resources, making it difficult to meet the model's requirements, resulting in underfitting and failing to achieve good recognition results; 2) Abnormal web pages and normal web pages have extremely high similarity, making it impossible to identify them based on content, thus greatly limiting the recognition effect of supervised learning; 3) Supervised models require a lot of resources and time for training, and the model update cost is high; 4) Machine learning models for text and images have poor interpretability. The model is a black box internally, and it cannot provide reliable evidence for the prediction results, which is not conducive to model operation and maintenance and practical application.
[0052] For cluster-diffusion-based abnormal webpage classification methods: 1) They are highly dependent on the initial seed label, and the choice of seed has a significant impact on the recognition results. If abnormal pages and normal pages have extremely high similarity, a large number of normal pages will be diffused and identified, resulting in a large number of misjudgments and seriously affecting the stability and effectiveness of the system; 2) Clustering algorithms are very sensitive to noise and outliers. If there is noise or outliers in the data, it may lead to the diffusion of incorrect categories; 3) For a small number of new webpage samples to be identified, it is still necessary to re-cluster and diffuse the whole, which will cause underfitting and fail to achieve good discrimination results.
[0053] Based on the problems existing in related technologies, this application provides an abnormal webpage identification method. This method involves multiple intelligent agents progressively processing and analyzing the webpage information of the webpage to be identified. The first intelligent agent performs preliminary screening, the second intelligent agent collects information, and the third intelligent agent identifies anomalies. This layered processing improves the accuracy and precision of abnormal webpage identification, reduces the possibility of misjudgment, and ensures that each intelligent agent has a clear division of labor, allowing each agent to focus on a specific function, thus improving the efficiency and quality of abnormal webpage identification.
[0054] Specifically, this application provides a method for identifying abnormal web pages. First, a first agent, a second agent, and a third agent are constructed in a multimodal large language model. Then, the web page information of the web page to be identified is input into the multimodal large language model. Next, the first agent performs information screening on the web page information to obtain screening information for abnormal web page identification. Then, the second agent collects information about the web page to be identified based on the web page information and the screening information to obtain collected information. Finally, the third agent performs abnormal web page identification on the web page to be identified based on the web page information and the collected information, and outputs the abnormal web page identification result through the multimodal large language model.
[0055] Here, we first describe an exemplary application of the abnormal webpage identification device according to the embodiments of this application. This abnormal webpage identification device is an electronic device used to implement the abnormal webpage identification method. In one implementation, the abnormal webpage identification device (i.e., electronic device) provided in the embodiments of this application can be implemented as a terminal or as a server. In one implementation, the electronic device provided in the embodiments of this application can be implemented as any terminal with abnormal webpage identification function, such as a laptop, tablet, desktop computer, or intelligent robot. In another implementation, the abnormal webpage identification device provided in the embodiments of this application can also be implemented as a server, wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of this application. Below, we will describe an exemplary application when the abnormal webpage identification device is implemented as a server.
[0056] See Figure 3 , Figure 3 This is an optional architecture diagram of the abnormal webpage identification system provided in this application embodiment. The abnormal webpage identification system 10 in this application embodiment includes at least a terminal 100, a network 200, and a server 300. An abnormal webpage identification application is deployed on the terminal 100, and the server 300 can be a backend server for the abnormal webpage identification application. The server 300 can constitute the abnormal webpage identification device of this application embodiment, that is, the abnormal webpage identification method of this application embodiment is implemented through the server 300. The terminal 100 is connected to the server 300 through the network 200, which can be a wide area network (WAN), a local area network (LAN), or a combination of both.
[0057] See Figure 3Terminal 100 receives a user's identification operation for an abnormal webpage. In response, terminal 100 generates an abnormal webpage identification request. Then, terminal 100 sends the abnormal webpage identification request to server 300 via network 200. Upon receiving the abnormal webpage identification request, server 300 obtains the webpage information of the webpage to be identified. Next, server 300 constructs a first agent, a second agent, and a third agent in a multimodal large language model. Then, server 300 inputs the webpage information of the webpage to be identified into the multimodal large language model. Finally, server... Server 300 uses a first intelligent agent to screen webpage information and obtain screening information for identifying abnormal webpages. Next, server 300 uses a second intelligent agent to collect information about the webpage to be identified based on the webpage information and the screening information. Then, server 300 uses a third intelligent agent to identify abnormal webpages based on the webpage information and the collected information, and outputs the abnormal webpage identification results through a multimodal large language model. Finally, server 300 sends the abnormal webpage identification results to terminal 100 through network 200, and displays the abnormal webpage identification results on the display interface of terminal 100.
[0058] In some embodiments, the above-described abnormal webpage identification method can also be executed by a terminal. That is, after receiving a user's identification operation for an abnormal webpage, the terminal 100 can obtain the webpage information of the webpage to be identified; then, the terminal 100 constructs a first intelligent agent, a second intelligent agent, and a third intelligent agent in a multimodal large language model; next, the terminal 100 inputs the webpage information of the webpage to be identified into the multimodal large language model; next, the terminal 100 uses the first intelligent agent to perform information screening on the webpage information to obtain screening information for abnormal webpage identification; next, the terminal 100 uses the second intelligent agent to collect information about the webpage to be identified based on the webpage information and the screening information to obtain collected information; next, the terminal 100 uses the third intelligent agent to perform abnormal webpage identification on the webpage to be identified based on the webpage information and the collected information, and outputs the abnormal webpage identification result through the multimodal large language model; finally, the terminal 100 displays the abnormal webpage identification result on the display interface.
[0059] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Figure 4The illustrated electronic device may be an abnormal webpage identification device, which includes at least one processor 410, a memory 450, at least one network interface 420, and a user interface 430. The various components in the abnormal webpage identification device are coupled together via a bus system 440. It is understood that the bus system 440 is used to implement communication between these components. In addition to a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 4 The general labeled all buses as Bus System 440.
[0060] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0061] User interface 430 includes one or more output devices 431 that enable the presentation of media content, and one or more input devices 432.
[0062] Memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. Memory 450 may optionally include one or more storage devices physically located remote from processor 410. Memory 450 may include volatile memory or non-volatile memory, or both. Non-volatile memory may be read-only memory (ROM), and volatile memory may be random access memory (RAM). The memory 450 described in this application embodiment is intended to include any suitable type of memory. In some embodiments, memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures, or subsets or supersets thereof, as illustrated below.
[0063] Operating system 451 includes system programs for handling various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks; network communication module 452 is used to reach other computing devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.; input processing module 453 is used to detect and translate one or more user inputs or interactions from one or more input devices 432.
[0064] In some embodiments, the apparatus provided in this application may be implemented in software. Figure 4 An abnormal webpage identification device 454 stored in memory 450 is shown. This device 454 can be an abnormal webpage identification device in an electronic device, and can be software in the form of programs and plug-ins, including the following software modules: a construction module 4541, an input module 4542, a screening module 4543, a collection module 4544, and an identification module 4545. These modules are logically connected and can therefore be arbitrarily combined or further divided according to their implemented functions. The functions of each module will be described below.
[0065] In some embodiments, the apparatus provided in this application can be implemented in hardware. For example, the apparatus provided in this application can be a processor in the form of a hardware decoding processor, which is programmed to execute the abnormal webpage identification method provided in this application. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0066] The abnormal webpage identification methods provided in the embodiments of this application can be executed by an electronic device, which can be a server or a terminal. That is, the abnormal webpage identification methods in the embodiments of this application can be executed by a server, by a terminal, or by interaction between a server and a terminal.
[0067] Figure 5 This is an optional flowchart illustrating the abnormal webpage identification method provided in this application embodiment. The following will be combined with... Figure 5 The steps shown are explained as follows: Figure 5 As shown, taking the server as the execution subject of the abnormal webpage identification method as an example, the method includes the following steps S101 to S105:
[0068] Step S101: Construct the first agent, the second agent, and the third agent in the multimodal large language model.
[0069] A multimodal large language model refers to a large-scale deep learning model with over a billion parameters, possessing multiple cross-modal capabilities such as image description, visual question answering, and image-text matching. An intelligent agent refers to an entity with autonomous understanding, planning and decision-making capabilities, and information processing capabilities, built based on a multimodal large language model as its core controller.
[0070] In some embodiments, the first intelligent agent includes a first pre-trained model and first prompt information. The first prompt information includes a first role prompt information, a first task prompt information, and a first input / output prompt information corresponding to the first pre-trained model. The first pre-trained model is the core part of the first intelligent agent. The first pre-trained model is a model pre-trained on a large amount of data and can be fine-tuned or used directly on a specific task. The first prompt information is an instruction that guides the first pre-trained model on how to work.
[0071] The first role prompt information specifies the role played by the first intelligent agent in the abnormal webpage identification process. It can help the first pre-trained model determine the identity or perspective from which to process the data. For example, if the role of the first intelligent agent specified in the first role prompt information is "violation content inspection expert", then the first pre-trained model can process webpage information from the perspective of "violation content inspection expert".
[0072] The first task prompt information is a prompt information indicating the type or goal of the task to be performed by the first pre-trained model. It is used to indicate the specific work that the first pre-trained model should currently complete or the goal to be achieved. For example, the first task prompt information specifies that the task goal of the first agent is: "to determine whether the webpage to be identified needs to be identified as an abnormal page". The first input-output prompt information is a prompt information indicating the input information and output format of the first pre-trained model. For example, the first input-output prompt information specifies that the input of the first pre-trained model includes: webpage URL and webpage title, and the output format is JSON format.
[0073] In some embodiments, the second agent includes a second pre-trained model and second prompt information. The second prompt information includes a second role prompt information, a second task prompt information, and a second input / output prompt information corresponding to the second pre-trained model. The second pre-trained model is the core part of the second agent. The second pre-trained model is a model pre-trained on a large amount of data and can be fine-tuned or used directly on a specific task. The second prompt information is an instruction that guides the second pre-trained model on how to work.
[0074] The second role prompt information specifies the role that the second agent plays in the process of identifying abnormal web pages. It can help the second pre-trained model determine the identity or perspective from which to process the data. For example, if the role of the second agent specified in the second role prompt information is "information gathering expert", then the second pre-trained model can process web page information and screening information from the perspective of "information gathering expert".
[0075] The second task prompt information is a prompt information indicating the type of task or goal to be performed by the second pre-trained model. It is used to indicate the specific work that the second pre-trained model should currently complete or the goal to be achieved. For example, the second task prompt information specifies that the task goal of the second agent is: "Use the information collection tool to collect information related to the input information". The second input-output prompt information is a prompt information indicating the input information and output format of the second pre-trained model. For example, the second input-output prompt information specifies that the input of the second pre-trained model includes: web page URL, web page title and screening information, and the output format is JSON format.
[0076] In some embodiments, the three agents include a third pre-trained model and third prompt information. The third prompt information includes third role prompt information, third task prompt information, and third input / output prompt information corresponding to the third pre-trained model. The third pre-trained model is the core part of the third agent. The third pre-trained model is a model pre-trained on a large amount of data and can be fine-tuned or used directly on a specific task. The third prompt information is an instruction that guides the third pre-trained model on how to work.
[0077] Among them, the third role prompt information is the prompt information that specifies the role played by the third intelligent agent in the abnormal webpage identification process. It can help the third pre-trained model determine what identity or perspective it should take to process data. For example, if the role of the third intelligent agent specified in the third role prompt information is "abnormal webpage identification expert", then the third pre-trained model can process webpage information and collect information from the perspective of "abnormal webpage identification expert" through the third role prompt information.
[0078] The third task prompt information is a prompt information indicating the type of task or goal to be performed by the third pre-trained model. It is used to indicate the specific work that the third pre-trained model should currently complete or the goal to be achieved. For example, the third task prompt information specifies that the task goal of the third agent is: "Determine whether the webpage to be identified is an abnormal webpage". The third input-output prompt information is a prompt information indicating the input information and output format of the third pre-trained model. For example, the third input-output prompt information specifies that the input of the third pre-trained model includes: webpage URL, webpage title and collected information, and the output format is JSON format.
[0079] Step S102: Input the webpage information of the webpage to be identified into the multimodal large language model.
[0080] Web pages to be identified refer to web pages that have not yet been processed or analyzed and require further operation. Web pages to be identified can be obtained by users manually adding Uniform Resource Locators (URLs) or by using web crawlers in batches. Web page information refers to all obtainable content on a web page, including but not limited to text, images, videos, audio, and metadata (such as titles, descriptions, keywords, etc.), web page screenshots, etc. Web page information can be structured information, such as tables and lists, facilitating data parsing and processing. Web page information can also be semi-structured or unstructured data, such as large amounts of natural language text or news articles on a web page, which require parsing using natural language processing technologies.
[0081] Step S103: The first intelligent agent performs information screening on the web page information to obtain screening information for abnormal web page identification.
[0082] The first intelligent agent refers to the intelligent agent used to screen web page information for identification. Information screening refers to extracting specific information from a large amount of web page information, usually information related to a certain goal or condition. Anomaly web page identification refers to identifying web pages that differ from normal web page behavior or content. Anomalies may include malicious web pages, phishing web pages, phishing websites, and content violations. Among them, malicious web pages are web pages specifically designed to perform malicious behavior. Malicious web pages usually contain malicious code or scripts that attempt to attack users' devices, data, or privacy without their knowledge. Phishing web pages are web pages that imitate the appearance and function of legitimate web pages, with the purpose of misleading users into believing that the phishing page is legitimate. Phishing websites are malicious websites that disguise themselves as legitimate websites to trick users into providing sensitive information (such as usernames, passwords, etc.). Content violations refer to content published on web pages that violates legal regulations, platform policies, or ethical standards. Screening information refers to the information obtained after screening by the first intelligent agent that needs to be used for anomaly web page identification.
[0083] In some embodiments, step S103 can be implemented by the following method: First, the web page information is filled into the input field of the first input-output information to update the first prompt information and obtain the updated first prompt information; then, through the first pre-trained model, based on the task objective determined by the first task prompt information, the web page information in the updated first prompt information is screened from the data processing perspective specified by the first role prompt information to determine the screening information that needs to be identified as abnormal web pages.
[0084] It should be noted that step S103 can automatically perform preliminary screening of a large amount of web page information through the first intelligent agent, reducing manual intervention and thus reducing labor costs; moreover, by guiding the first pre-trained model with the first prompt information, the first pre-trained model can pay more attention to the web page information that needs to be focused on, providing high-quality input for subsequent processing, thereby improving the recognition accuracy of abnormal web pages.
[0085] Step S104: The second intelligent agent collects information about the webpage to be identified based on webpage information and screening information, and obtains the collected information.
[0086] The second agent refers to the agent used to collect information based on web page information and screening information output by the first agent; information collection refers to extracting target information with a high degree of similarity to web page information or screening information from a pre-built database based on web page information and screening information; information collection refers to the set of target extended information extracted by the second agent from the pre-built database and screening information.
[0087] In some embodiments, step S104 can be implemented by the following method: First, inputting webpage information and screening information into the second input-output prompt information to update the second prompt information and obtain the updated second prompt information; then, inputting the second prompt information into the second pre-trained model, which will determine whether the amount of input webpage information and screening information is sufficient for abnormal webpage identification based on the role and task determined by the second prompt information. If sufficient, the screening information is directly identified as the collection information; if insufficient, the information collection tool searches for target vectors with a similarity higher than a preset threshold to the input data in a pre-built vector library, and concatenates the information corresponding to the target vectors with the screening information to obtain the collection information; wherein, the information collection tool refers to an algorithm module used to search for data that meets the conditions from a pre-built database.
[0088] It should be noted that step S104 automatically collects information based on webpage information and screening information through the second intelligent agent, reducing human intervention and minimizing errors that may be caused by human information collection. This can enrich the data information for abnormal webpage identification to the greatest extent, thereby further improving the accuracy of abnormal webpage identification and avoiding false detection or missed detection.
[0089] Step S105: Through a third intelligent agent, based on webpage information and collected information, abnormal webpage identification is performed on the webpage to be identified, and the abnormal webpage identification result is output through a multimodal large language model.
[0090] The third agent refers to an agent used to identify abnormal web pages based on web page information and the collected information output by the second agent.
[0091] In some embodiments, step S105 can be implemented by the following method: First, the web page information and collected information are filled into the input field of the first input-output information to update the first prompt information and obtain the updated first prompt information; then, through the third pre-trained model, based on the task objective determined by the third task prompt information, from the data processing perspective specified by the third role prompt information, anomaly identification is performed on the web page information and collected information in the updated third prompt information to determine the abnormal web page identification result of the web page to be identified.
[0092] The abnormal webpage identification method of this application embodiment processes and analyzes the webpage information of the webpage to be identified through multiple intelligent agents in a progressively deeper manner. First, the first intelligent agent can automatically perform preliminary screening of a large amount of webpage information, reducing manual intervention and thus lowering labor costs. Moreover, the first prompt information guides the first pre-trained model, enabling it to focus more on the webpage information that needs special attention, providing high-quality input for subsequent processing and thus improving the identification accuracy of abnormal webpages. Then, the second intelligent agent automatically collects information based on the webpage information and screening information, reducing manual intervention and minimizing the errors that may result from manual information collection. The process of identifying abnormal web pages can enrich the data information to the greatest extent, thereby further improving the accuracy of abnormal web page identification and avoiding false positives or false negatives. Finally, the anomaly identification is performed by a third intelligent agent. Through layer-by-layer processing, the accuracy and precision of abnormal web page identification can be improved, reducing the possibility of misjudgment. Moreover, each intelligent agent has a clear division of labor, allowing each agent to focus on a specific function, thereby improving the efficiency and quality of abnormal web page identification. In addition, through the screening and information collection of the first and second intelligent agents, the third intelligent agent already has sufficient data support when performing abnormal web page identification, which can effectively improve the accuracy of abnormal page identification.
[0093] The following examples illustrate the application scenarios of the abnormal webpage identification method provided in this application. This application embodiment can be applied to at least the following exemplary scenarios:
[0094] Scenario 1: Users are accessing web pages more and more frequently on internet platforms, application software, and community platforms. However, with this increased frequency, some malicious users or organizations exploit this trend to illegally obtain users' personal information by disguising abnormal web pages as legitimate ones. This phenomenon poses a significant security risk, especially when sensitive user information (such as usernames, passwords, bank accounts, etc.) is involved, making it easier for personal information to be leaked. To improve the security of user information, the abnormal web page identification method provided in this application embodiment can be adopted. First, the user inputs the web page information of the accessed web page into the terminal. The terminal generates an abnormal web page identification request based on the web page information and sends the request to the server. After receiving the request, the server parses the request to obtain the web page information. Then, a first intelligent agent performs information screening on the web page information to obtain screening information. Next, the screening information and the web page information are input into a second intelligent agent, which collects information to obtain collected information. Finally, the web page information and the collected information are input into a third intelligent agent to identify abnormal web pages accessed by the user and manage abnormal web pages based on the identification results.
[0095] Scenario 2: With the widespread use of the internet, fraudulent activities using abnormal web pages are becoming increasingly serious. To prevent infringement on users' privacy and property, relevant personnel need to identify abnormal web pages in a timely and accurate manner to warn potential victims. This can be achieved using the abnormal web page identification method provided in this application. First, relevant personnel input the web page information of the web page to be identified into a terminal. The terminal generates an abnormal web page identification request based on the web page information and sends the request to a server. After receiving the request, the server parses it to obtain the web page information. Then, a first intelligent agent screens the web page information to obtain screening information. Next, the screening information and the web page information are input into a second intelligent agent, which collects information to obtain collected information. Finally, the web page information and the collected information are input into a third intelligent agent to identify abnormal web pages. When an abnormal web page is identified, a warning message can be issued to the user to prevent infringement on their privacy.
[0096] The abnormal webpage identification method of this application embodiment will be described below using the above scenario as an example. Figure 6 This is another optional flowchart illustrating the abnormal webpage identification method provided in the embodiments of this application, such as... Figure 6 As shown, the method includes the following steps S201 to S211:
[0097] Step S201: The terminal receives an abnormal webpage identification operation for the webpage to be identified.
[0098] Here, an abnormal webpage identification application can run on the terminal, and the server constitutes the backend server for the abnormal webpage identification application. The abnormal webpage identification operation can be a selection operation or an input operation through the client of the abnormal webpage identification application running on the terminal. For example, the selection operation can select the webpage information to be identified, or the input operation can be that the user enters the webpage information to be identified on the client.
[0099] In some embodiments, the abnormal webpage identification application may provide an input interface or input box, allowing users to select or input webpage information of the webpage to be identified. The input interface may be in the form of a form, text box, or drop-down menu, etc., and the specific form is not limited in this application. Users can select webpage information of the webpage to be identified from predetermined options, or manually input the webpage information of the webpage to be identified.
[0100] In step S202, the terminal generates an abnormal webpage identification request in response to the abnormal webpage identification operation.
[0101] Here, the terminal can encapsulate the webpage information of the webpage selected by the user into the abnormal webpage identification request.
[0102] In some embodiments, to ensure the security of abnormal webpage identification requests, authentication parameters are added to the requests. These parameters verify the legitimacy of the request, ensuring that only authorized users can access the abnormal webpage identification service and preventing unauthorized access and abuse. A common authentication method is to use API keys or tokens. An API key is a unique string used to identify and verify a user's identity. A token is a credential similar to an access token or authentication token, containing the user's identity information and permissions.
[0103] In step S203, the terminal sends an abnormal webpage identification request to the server.
[0104] In some embodiments, the terminal sends the encapsulated abnormal webpage identification request to the server and requests the server to perform abnormal webpage identification operations. The abnormal webpage identification request is usually sent using protocols such as HTTP or WebSocket.
[0105] In step S204, the server responds to the abnormal webpage identification request and obtains the webpage information of the webpage to be identified.
[0106] Here, after receiving an abnormal webpage identification request, the server parses the request. For example, for an HTTP request, the server can parse the request header and request body. Parsing the request header retrieves relevant request information; parsing the request body retrieves the main data of the request, i.e., the webpage information of the webpage to be identified. Specific fields or parameters in the request body, which contain the webpage information, are parsed. A specific data format, such as JSON or XML, is extracted from the request body, and then this data format is parsed to obtain the webpage information of the webpage to be identified.
[0107] In step S205, the server constructs a first agent, a second agent, and a third agent in the multimodal large language model.
[0108] In some embodiments, the first intelligent agent can be constructed by the following method: First, a first pre-trained model is obtained, along with first role prompt information, first task prompt information, and first input / output prompt information corresponding to the first pre-trained model; then, the first role prompt information, first task prompt information, and first input / output prompt information are merged to obtain first prompt information; finally, the first intelligent agent is constructed based on the first prompt information and the first pre-trained model.
[0109] Here, the first role prompt information is a prompt information that specifies the role played by the first intelligent agent in the process of identifying abnormal web pages. The first role prompt information may include the role of the first intelligent agent, the language, and the rules that need to be followed; an example of the first role prompt information is as follows:
[0110] "You are an intelligent agent, adept at handling complex tasks based on your role and following the requirements."
[0111] Role: Expert in detecting prohibited content;
[0112] Language: Chinese;
[0113] Basis: You need to comply with the regulations of XXXX.
[0114] The specific content of the prompt information for the first role can be set according to actual needs, and this application does not impose any restrictions on it.
[0115] The first task prompt information refers to the prompt information that specifically instructs the first pre-trained model on the type of task or the goal it should perform. It indicates the specific work that the first pre-trained model should currently complete or the goal it should achieve. The first task prompt information may include the task, the goal, key information, judgment content, and decision-making execution steps. An example of first task prompt information is as follows:
[0116] "Task: Based on the webpage screenshot, webpage URL, and webpage title, complete the following tasks;"
[0117] Objective: To determine whether the webpage to be identified needs to be identified as an abnormal webpage, and to collect information related to the webpage to be identified;
[0118] Key information to focus on: sensitive personal information and organizational information;
[0119] Judgment content:
[0120] Whether the webpage to be identified requires the input of personal information, including mobile phone number, card number, password, etc.
[0121] The webpage to be identified determines whether it displays organizational information, such as name, address, and contact number.
[0122] Decision execution:
[0123] If the webpage to be identified requires the input of personal information, it is considered that anomaly identification is required, anomaly identification action is initiated, and the organizational information in the webpage to be identified is recorded.
[0124] If the webpage to be identified has no personal information input, it is considered that no abnormal identification is needed, and the page is discarded.
[0125] The specific content included in the first task prompt information can be set according to actual needs, and this application does not impose any restrictions here.
[0126] The first input / output prompt is a prompt indicating the input information and output format of the first pre-trained model. An example of the first input / output prompt is as follows:
[0127]
[0128]
[0129] Then, the first role prompt information, the first task prompt information, and the first input / output prompt information are merged to obtain the first prompt information; the first prompt information and the first pre-trained model are combined to obtain the first intelligent agent.
[0130] In some embodiments, the second agent can be constructed by the following method: First, a second pre-trained model is obtained, along with second role prompt information, second task prompt information, and second input-output prompt information corresponding to the second pre-trained model; then, the second role prompt information, second task prompt information, and second input-output prompt information are merged to obtain second prompt information; finally, the second agent is constructed based on the second prompt information and the second pre-trained model.
[0131] Here, the second role prompt information is a prompt information that specifies the role played by the second intelligent agent in the abnormal webpage identification process. The second role prompt information can include the role of the second intelligent agent and its language; an example of the second role prompt information is as follows:
[0132] "You are an intelligent agent, adept at handling complex tasks based on your role and following the requirements."
[0133] Role: Information gathering expert;
[0134] Language: Chinese.
[0135] The specific content of the second-role prompt can be set according to actual needs, and this application does not impose any restrictions on it.
[0136] Second task prompts refer to specific prompts indicating the type of task or goal that the second pre-trained model should perform. They instruct the model to complete a specific task or achieve a particular objective. Second task prompts may include the task, objective, available tools, and decision-making steps. An example of a second task prompt is shown below.
[0137] Task: Based on the information on the webpage, complete the following tasks.
[0138] Objective: To use information gathering tools to search for data related to the input information;
[0139] Available tools:
[0140] Text knowledge retrieval: Retrieves text information that is most similar to the input text. The input can be information such as organization name, address, and contact information, and the output will be the most similar search results.
[0141] Image knowledge retrieval: Retrieves image information that is most similar to the input image. The input is a webpage screenshot. The system searches the webpage screenshot and outputs the most similar search results.
[0142] Search engine: Search for keywords in a search engine, enter the keywords, and get the search results.
[0143] Decision-making and implementation steps:
[0144] Determine if the input information is sufficient for detecting abnormal web pages.
[0145] If there is enough, no action is needed.
[0146] If this is insufficient, provide the information gathering tools and input data required.
[0147] The specific content included in the second task prompt information can be set according to actual needs, and this application does not impose any restrictions here.
[0148] The second input / output prompt is a prompt indicating the input information and output format of the second pre-trained model. An example of the second input / output prompt is as follows:
[0149]
[0150] Then, the second role prompt information, the second task prompt information, and the second input / output prompt information are merged to obtain the second prompt information; the second prompt information and the second pre-trained model are combined to obtain the second agent.
[0151] In some embodiments, the third agent can be constructed by the following method: First, the server obtains a third pre-trained model, as well as third role prompt information, third task prompt information, and third input / output prompt information corresponding to the third pre-trained model; then, the server merges the third role prompt information, third task prompt information, and third input / output prompt information to obtain third prompt information; finally, based on the third prompt information and the third pre-trained model, the third agent is constructed.
[0152] Here, the third-party prompt information specifies the role played by the third agent in the abnormal webpage identification process. The third-party prompt information may include the third agent's role, language, and the rules to be followed. An example of third-party prompt information is as follows:
[0153] "You are an intelligent agent, adept at handling complex tasks based on your role and following the requirements."
[0154] Role: Abnormal Webpage Identification Expert;
[0155] Language: Chinese;
[0156] Basis: You need to comply with the regulations of XXXX.
[0157] The specific content of the third-party prompts can be set according to actual needs, and this application does not impose any restrictions on it.
[0158] Third-task prompts refer to specific prompts indicating the type of task or goal that the third pre-trained model should perform. They are used to indicate the specific task or goal that the third pre-trained model should currently complete. Third-task prompts may include the task, goal, reference information, and decision-making steps. An example of a third-task prompt is shown below:
[0159] Task: Based on the information on the webpage and the information collected, complete the following tasks.
[0160] Objective: To determine whether a webpage to be identified is an abnormal webpage.
[0161] Reference criteria: To determine whether a webpage to be identified is an abnormal webpage, the following criteria can be used as a reference:
[0162] The system checks whether the contact information, such as the address and telephone number, of the organization on the webpage to be identified matches the information in the target text vector retrieved from the text vector library.
[0163] The organization's name must match the organization name of the target image vector retrieved from the image vector library.
[0164] Does the search engine search result include information supporting the identification of the webpage as the official webpage of the corresponding organization?
[0165] Decision-making and implementation steps:
[0166] Based on the reference criteria, the webpage information and collected information are used to determine whether the webpage to be identified is an abnormal webpage.
[0167] The relevant web page information, collected information, and judgment criteria were compiled into a report.
[0168] The specific content included in the third task prompt information can be set according to actual needs, and this application does not impose any restrictions here.
[0169] The third input / output cue information is a cue information indicating the input information and output format of the third pre-trained model. An example of the third input / output cue information is as follows:
[0170]
[0171] Then, the third role prompt information, the third task prompt information, and the third input / output prompt information are merged to obtain the third prompt information; the third prompt information and the third pre-trained model are combined to obtain the third intelligent agent.
[0172] In step S206, the server inputs the webpage information of the webpage to be identified into the multimodal large language model.
[0173] In step S207, the server uses the first intelligent agent to screen the webpage information and obtain the screening information that needs to be used for abnormal webpage identification.
[0174] Here, the server uses a first intelligent agent to perform preliminary screening of web page information to determine whether the web page to be identified needs to be identified as abnormal. If it needs to be identified as abnormal, the server extracts the screening information for identifying abnormal web pages from the web page information.
[0175] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the process of screening web page information through a first intelligent agent, provided in an embodiment of this application. Figure 7 In step S207, the server, through the first intelligent agent, performs information screening on the webpage information to obtain screening information for identifying abnormal webpages. This can be achieved through steps S2071 to S2072.
[0176] In step S2071, the server fills the webpage information into the first input / output prompt information to obtain the updated first prompt information.
[0177] In some embodiments, the server populates the input fields of the first input / output prompt information with webpage information to update the first prompt information. For example, if the input fields of the first input / output prompt information are "URL: {url}; Title: {title}", and the webpage information contains URL1 and Title1, the server populates the corresponding fields of the first input / output prompt information with the URL and title from the webpage information, resulting in the updated input fields of the first prompt information being "URL: {URL1}; Title: {Title1}".
[0178] In step S2072, the server uses the first pre-trained model in the first intelligent agent to perform information screening on the web page information in the updated first prompt information based on the first role prompt information and the first task prompt information, to obtain the screening information that needs to be used for abnormal web page identification.
[0179] In some embodiments, the server uses a first pre-trained model to screen the webpage information in the updated first prompt information from the data processing perspective specified by the first role prompt information, based on the task objective determined by the first task prompt information, and determines the screening information that needs to be identified as abnormal webpages.
[0180] In some examples, the first role prompt information identifies the first agent's role as a "violation content inspection expert," and the first task prompt information identifies the first agent's task as "detecting whether the webpage to be identified needs to be anomaly identified." The server, through the first pre-trained model, performs information screening on the webpage information in the updated first prompt information based on the first role prompt information and the first task prompt information. The screening process is as follows: First, the server extracts features from the webpage information, extracting text content, links, etc.; then, based on the role of "violation content inspection expert," the server pays special attention to words and links related to finance, payment, and personal information in the webpage information. For example, if the server finds the phrase "Please log in to your bank account to verify your identity" and links related to financial institutions in the webpage information, the server can identify similar words and links as screening information.
[0181] In some embodiments, the server may output screening information in the output format specified in the updated first prompt information.
[0182] After determining the screening information, the server will output the screening information according to the output format specified in the first input / output prompt information in the updated first prompt information.
[0183] In step S208, the server, through the second intelligent agent, collects information about the webpage to be identified based on webpage information and screening information, and obtains the collected information.
[0184] Here, the server uses a second intelligent agent to determine whether the amount of web page information and screening information meets the requirements for web page anomaly detection. If the requirements are met, the screening information is identified as collected information. If the requirements are not met, information is collected based on the web page information and screening information. The collected information is then combined with the screening information to obtain the collected information.
[0185] In some embodiments, see Figure 8 , Figure 8 This is a schematic diagram of the process of information collection through a second intelligent agent provided in an embodiment of this application; Figure 8 In step S208, the server, through a second intelligent agent, collects information about the webpage to be identified based on webpage information and screening information. This collected information can be achieved through the following steps S2081 to S2084:
[0186] In step S2081, the server inputs the second prompt information, web page information, and screening information into the second intelligent agent and obtains the output result.
[0187] In some embodiments, the server fills the corresponding fields in the second prompt information with web page information and screening information to update the second prompt information. The server, through the second pre-trained model in the second agent, determines whether the number of web page information and screening information in the updated second prompt information meets the requirements from the data processing perspective specified by the second role prompt information in the updated second prompt information, based on the task objective determined by the second task prompt information in the updated second prompt information. For example, the server counts the number of web page information and screening information and compares it with a preset number threshold. When the number is less than the preset number threshold, the output result is that extended information needs to be collected based on the screening information, and the tool identifier and input data to be used are added to the output result. When the number is greater than the preset number threshold, the output result is that extended information does not need to be collected.
[0188] In step S2082, when the output result indicates that extended information needs to be collected based on the screening information, the server parses the tool identifier and input data from the output result.
[0189] Extended information refers to the information corresponding to data in the database that has a similarity to the input data that is higher than a preset threshold. Input data refers to the source data that corresponds to the extended information. For example, if the input data is data A, and data B that has a similarity to data A that is higher than a preset threshold is found in the database, then the information corresponding to data B is the extended information.
[0190] In some embodiments, the server parses the output of the second pre-trained model. When the output indicates that extended information needs to be collected based on the screening information, the server continues to parse the output to extract the usable tool identifiers and input data.
[0191] In step S2083, the server collects target extended information by using the information collection tool corresponding to the tool identifier based on the input data, and then concatenates the target extended information with the screening information to obtain the collected information.
[0192] Here, the server searches the database for target data with a similarity higher than a preset threshold to the input data based on the information collection tool corresponding to the tool identifier. Information corresponding to the target data is identified as target extended information. Then, the target extended information is concatenated with the screening information to obtain the collected information. For example, if the screening information is {screening information 1, screening information 2, screening information 3}, and the target extended information obtained through the information collection tool is {extended information 1, extended information 2}, concatenating the screening information and the target extended information yields the collected information as {screening information 1, screening information 2, screening information 3, extended information 1, extended information 2}.
[0193] In some embodiments, the input data may be text data. When the input data is text data, step S2083 can be implemented by the following method: First, when the input data is text data, the server performs feature extraction on the text data to obtain a text embedding vector; then, the server uses an information collection tool to search for target text vectors in a pre-built text vector library that have a similarity higher than a first preset threshold with the text embedding vector; finally, the server concatenates the information corresponding to the target text vector with the screening information to obtain the collected information.
[0194] Here, text embedding vectors are numerical representations of text data obtained after feature extraction; a pre-built text vector library refers to a database that has been created in advance and stores embedding vectors of a large amount of text data.
[0195] In some embodiments, the server first preprocesses the input text data, such as removing stop characters, punctuation marks, and segmenting words; then, it uses a neural network model to convert the preprocessed text into text embedding vectors; next, the server compares the similarity of the generated text embedding vectors with text vectors in a text vector library, where cosine similarity, Euclidean distance, or other similarity metrics can be used to evaluate the similarity between text vectors. The calculated similarity is compared with a first preset threshold. If it is higher than the first preset threshold, the corresponding vector is extracted as the target text vector, and further information corresponding to the target text vector is obtained. For example, if the first preset threshold is 0.8, and the similarity between text embedding vector A and text vector B in the text vector library is 0.85, then text vector B is used as the target text vector; finally, the information corresponding to the target text vector is concatenated with the screening information to obtain the collected information, where concatenation is simply information merging.
[0196] In some embodiments, the input data may be image data. When the input data is image data, step S2063 can be implemented by the following method: First, when the input data is image data, the server performs feature extraction on the image data to obtain an image embedding vector; then, the server uses an information collection tool to search for target image vectors in a pre-built image vector library that have a similarity higher than a second preset threshold with the image embedding vector; finally, the server concatenates the information corresponding to the target image vector with the screening information to obtain the collected information.
[0197] Here, image embedding vectors are numerical representations of image data obtained after feature extraction; a pre-built image vector library refers to a database that has been created in advance and stores embedding vectors of a large amount of image data.
[0198] In some embodiments, the server first extracts features from the input image data to obtain image embedding vectors. Then, the server compares the similarity of the generated image embedding vectors with image vectors in an image vector library. Cosine similarity, Euclidean distance, or other similarity metrics can be used to evaluate the similarity between image vectors. The calculated similarity is compared with a second preset threshold. If the similarity is higher than the second preset threshold, the corresponding vector is extracted as the target image vector, and information corresponding to the target image vector is further obtained. For example, if the second preset threshold is 0.8, and the similarity between image embedding vector A and image vector B in the image vector library is 0.85, then image vector B is used as the target image vector. Finally, the information corresponding to the target image vector is concatenated with screening information to obtain the collected information.
[0199] In step S2084, when the output result indicates that extended information does not need to be collected, the server determines the screened information as the information to be collected.
[0200] In some embodiments, the server parses the output of the second pre-trained model. When the output indicates that no extended information needs to be collected, the screening information output by the first agent is determined as the information to be collected.
[0201] In step S209, the server, through a third intelligent agent, performs abnormal webpage identification on the webpage to be identified based on webpage information and collected information, obtains the abnormal webpage identification result, and outputs the abnormal webpage identification result through a multimodal large language model.
[0202] Here, the server uses a pre-configured third-party intelligent agent to identify abnormal web pages based on web page information and collected information, and obtains the abnormal web page identification results.
[0203] In some embodiments, see Figure 9 , Figure 9 This is a schematic diagram of the process for identifying abnormal web pages through a third-party intelligent agent, provided in an embodiment of this application. Figure 9 In step S209, the server, through a third-party intelligent agent, identifies abnormal web pages based on web page information and collected information, and obtains the abnormal web page identification result. This can be achieved through the following steps S2091 to S2093:
[0204] In step S2091, the server fills the webpage information and collected information into the third input / output prompt information to obtain the updated third prompt information.
[0205] In some embodiments, the server populates the input fields of the third input / output prompt information with webpage information to update the third prompt information. For example, the input fields of the third input / output prompt information are "URL: {url}; Title: {title}; Info: {info}", and the webpage information contains URL1, Title1, and Info1. The server populates the corresponding fields of the third input / output prompt information with the URL and title from the webpage information, as well as the collected information, resulting in the updated input fields of the third prompt information being "URL: {URL1}; Title: {title1}; Info: {Info1}".
[0206] In step S2092, the server uses the third pre-trained model in the third intelligent agent to perform anomaly identification on the web page information in the updated third prompt information based on the third role prompt information, the third task prompt information, and the collected information, and obtains the anomaly identification result.
[0207] In some embodiments, the server uses a third pre-trained model to identify anomalies in the webpage information in the updated third prompt information from the data processing perspective specified by the third role prompt information, based on the task objective determined by the third task prompt information, and obtains the anomaly identification result.
[0208] In some examples, the third-role prompt information identifies the third agent's role as an "abnormal webpage identification expert," and the third-task prompt information identifies the third agent's task as "determining whether the webpage to be identified is an abnormal webpage." The server uses a third pre-trained model to perform anomaly identification on the webpage information in the updated third prompt information based on the third-role prompt information and the third-task prompt information. The anomaly identification process is as follows: First, the server obtains information such as the organization name, address, and phone number contained in the webpage information; then, it compares the obtained information with the information in the text vectors obtained from the text vector library to determine whether the obtained information is consistent with the information in the text vectors. If they are inconsistent, the webpage to be identified is judged as an abnormal webpage, and the prompt information "The webpage to be identified is an abnormal webpage" is output. Based on the judgment result, an identification basis report is generated.
[0209] Step S2093: The server determines the abnormal webpage identification result of the webpage to be identified based on the abnormal identification result.
[0210] Here, the server determines the abnormal webpage identification result based on the prompt information and identification basis report generated by the third-party intelligent agent.
[0211] In some embodiments, the server may output the abnormal webpage identification results in accordance with the output format specified in the updated third prompt information.
[0212] After determining the abnormal webpage identification result, the server will output the abnormal webpage identification result according to the output format specified in the third input / output prompt information in the updated third prompt information.
[0213] In step S210, the server sends the abnormal webpage identification result to the terminal.
[0214] Step S211: The terminal displays the abnormal webpage identification result on the current interface.
[0215] After determining the results of abnormal webpage identification, the webpages to be identified can be managed according to the results to prevent users from being harmed by abnormal webpages. At the same time, the identification report can also be submitted to regulatory agencies as clues to facilitate the identification of abnormal webpages by regulatory agencies.
[0216] The abnormal webpage identification method of this application embodiment processes and analyzes the webpage information of the webpage to be identified through multiple intelligent agents in a progressively deeper manner. First, the first intelligent agent can automatically perform preliminary screening of a large amount of webpage information, reducing manual intervention and thus lowering labor costs. Moreover, the first prompt information guides the first pre-trained model, enabling it to focus more on the webpage information that needs special attention, providing high-quality input for subsequent processing and thus improving the identification accuracy of abnormal webpages. Then, the second intelligent agent automatically collects information based on the webpage information and screening information, reducing manual intervention and minimizing the errors that may result from manual information collection. The process of identifying abnormal web pages can enrich the data information to the greatest extent, thereby further improving the accuracy of abnormal web page identification and avoiding false positives or false negatives. Finally, the anomaly identification is performed by a third intelligent agent. Through layer-by-layer processing, the accuracy and precision of abnormal web page identification can be improved, reducing the possibility of misjudgment. Moreover, each intelligent agent has a clear division of labor, allowing each agent to focus on a specific function, thereby improving the efficiency and quality of abnormal web page identification. In addition, through the screening and information collection of the first and second intelligent agents, the third intelligent agent already has sufficient data support when performing abnormal web page identification, which can effectively improve the accuracy of abnormal page identification.
[0217] The following will describe an exemplary application of the embodiments of this application in a real-world application scenario.
[0218] This application proposes an abnormal webpage identification method. This method uses a pre-trained model (which may be a multimodal large language model) to construct three interoperable intelligent agents: a screener (the first intelligent agent mentioned above), an information officer (the second intelligent agent mentioned above), and a discriminator (the third intelligent agent mentioned above). It establishes an abnormal webpage identification method that includes three processes: information screening, information collection, and abnormal identification, thereby achieving accurate and reliable identification of abnormal webpages.
[0219] A pre-trained model is a model with a large number of parameters that is pre-trained on a large number of samples and capable of supporting visual question answering tasks and instruction following. See also Figure 10 , Figure 10 This is a schematic diagram of a visual question answering task provided in an embodiment of this application. An image 1002 and an open-ended natural language question 1003 are input into a pre-trained model 1001. The pre-trained model can output a natural language answer 1004 based on the image and the question.
[0220] The abnormal webpage identification method provided in this application is based on the powerful capabilities of pre-trained models, establishing multiple intelligent agents to complete complex tasks such as information screening, information collection, and anomaly identification. See also Figure 11 , Figure 11This is another optional flowchart of the abnormal webpage identification method provided in this application embodiment. The abnormal webpage identification method is mainly divided into three stages: information screening stage 1101, information collection stage 1102, and abnormal identification stage 1103. The three stages will be described in detail below.
[0221] Phase 1: Information Screening Phase
[0222] See Figure 12 , Figure 12 This is a flowchart illustrating the information screening process of the first intelligent agent provided in this application embodiment. The input to the first intelligent agent 1201 is webpage information, including webpage URL, webpage title 1202, and webpage screenshot 1203. The webpage information and the first prompt information 1204 are input into the first pre-trained model 1205 to perform information screening on the webpage information. The output is whether anomaly identification is required. Here, the construction of the first prompt information is divided into three steps: 1) constructing the first role prompt information 1206; 2) constructing the first task prompt information 1207; 3) constructing the first input-output prompt information 1208. The construction process of the first role prompt information, the first task prompt information, and the first input-output prompt information is described in step S205, and will not be repeated here.
[0223] Then, the first role prompt, the first task prompt, and the first input / output prompt are merged into the first prompt. The information screening process begins, replacing `{url}` in the first input / output prompt with the URL of the webpage to be identified, and `{title}` with the title of the webpage to be identified. When the first agent's decision is "No need for anomaly identification" and its action is "Discard," the webpage to be identified is directly determined to be a normal webpage, without further processing. However, when the first agent's decision is "Anomaly identification is required" and its action is "Initiate anomaly identification," a screenshot of the webpage, the webpage URL, the webpage title, and the screening result are input to the second agent.
[0224] Phase Two: Information Gathering Phase
[0225] See Figure 13 , Figure 13This is a flowchart illustrating the information collection process of the second intelligent agent provided in this application embodiment. The input information of the second intelligent agent 1301 includes a webpage URL, a webpage title, screening information 1302, and a webpage screenshot 1303. The input information and the second prompt information 1304 are input into the second pre-trained model 1305. The model determines whether the amount of input information is sufficient for abnormal webpage detection and obtains an output result. If the output result indicates that no extended information needs to be collected, the screening information is determined as the collection information. If the output result indicates that extended information needs to be collected based on the screening information, the tool identifier and input data are parsed from the output result. Based on the input data, the target extended information is collected through the information collection tool 1306 corresponding to the tool identifier, and the target extended information is concatenated with the screening information to obtain the collection information 1307.
[0226] The information gathering phase includes three steps: the construction of information gathering tools, the construction of the second intelligent agent, and the construction of the process. Specifically:
[0227] Step 1: Building Information Gathering Tools
[0228] 1) Constructing a text vector library 1308:
[0229] First, information about various organizations, institutions, and companies is compiled into text information (1309 records). This text data includes organization name, address, contact number, official website, etc. If multiple branches or organizations exist, multiple records are used. Then, a neural network model is used to extract features from each text record, obtaining corresponding text vectors. These text vectors and their corresponding text information are stored in a database to generate a text vector library. Finally, when the input is text data, features are extracted to obtain text embedding vectors. Target text vectors with a similarity higher than a first preset threshold are searched in the text vector library, and the text information corresponding to these target text vectors is output.
[0230] 2) Constructing the image vector library 1310:
[0231] First, screenshots are taken from the known websites of various organizations to obtain page screenshot information 1311. Then, features are extracted from the page screenshot information to obtain corresponding image vectors. The image vectors and their corresponding page screenshot information are stored in a database to generate an image vector library. Finally, when the input is image data, features are extracted from the image data to obtain image embedding vectors. Target image vectors with a similarity higher than a second preset threshold are searched in the image vector library, and the page screenshot information corresponding to the target image vectors is used as the output.
[0232] 3) Building search engine tools 1312:
[0233] The search engine 1313 is packaged into an interface. The input of the search engine is the search keywords, and the output is the search results.
[0234] Step Two: Construction of the Second Intelligent Agent:
[0235] First, second role prompt information 1314, second task prompt information 1315, and second input / output prompt information 1316 are constructed. The construction process of the second role prompt information, second task prompt information, and second input / output prompt information is described in step S206 and will not be repeated here. The second role prompt information, second task prompt information, and second input / output prompt information are merged to form second prompt information. Then, the second prompt information and the second pre-trained model are combined to form the second intelligent agent.
[0236] Step 3: Process Construction:
[0237] After completing the construction of the information collection tools and the second intelligent agent, the entire information collection process is constructed, mainly in three steps:
[0238] 1) Input the webpage information and the second prompt information into the second pre-trained model, where {url} is the webpage URL, {title} is the webpage title, and {info} is the screening information in the second prompt information, and obtain the output result of the second agent.
[0239] 2) When the output result indicates that no extended information needs to be collected, the screening information is determined as the information to be collected; when the output result indicates that extended information needs to be collected based on the screening information, the tool identifier and input data are parsed from the output result, and the information collection tool corresponding to the tool identifier is called to collect the target extended information.
[0240] 3) Concatenate the target extension information with the screening information to obtain the collected information. Input the collected information back into the second pre-trained model and repeat steps 2) and 3) until the output of the second agent indicates that no extension information needs to be collected. Stop collecting information and input the final collected information into the third agent.
[0241] Phase 3: Anomaly Detection Phase
[0242] See Figure 14 , Figure 14This is a flowchart illustrating the process of an abnormal webpage identification performed by a third intelligent agent according to an embodiment of this application. The inputs to the third intelligent agent 1401 are the webpage URL, webpage title, collected information 1402, and webpage screenshot 1403. The webpage URL, webpage title, webpage screenshot, and first prompt information 1404 are input into the third pre-trained model 1405 to determine whether the webpage to be identified is an abnormal webpage. Here, the construction of the third prompt information is divided into three steps: 1) constructing third role prompt information 1406; 2) constructing third task prompt information 1407; 3) constructing third input-output prompt information 1408. The construction process of the third role prompt information, third task prompt information, and third input-output prompt information is described in step S207, and will not be repeated here.
[0243] Then, the {url} in the third input / output prompt information is replaced with the webpage URL, {title} with the webpage title, and {info} with the collected information. The third pre-trained model then determines whether the webpage to be identified is an abnormal webpage and outputs an identification report. Finally, based on the abnormal webpage identification results, the webpage to be identified is managed to prevent users from being harmed by abnormal webpages. Simultaneously, the identification report can be used as a clue to report to regulatory agencies, facilitating the identification of abnormal webpages.
[0244] The abnormal webpage identification method provided in this application embodiment eliminates the need for manual annotation and model training, reducing manual processing and review time by 90% and computational resource consumption by 30%, improving tag generation efficiency and effectively reducing costs. Furthermore, the abnormal webpage identification method provided in this application embodiment can accurately identify webpage images and provide a detailed identification basis report, demonstrating good identification performance in multiple application fields. In particular, in internet ecosystem maintenance, the abnormal webpage identification rate is improved by 14%, the accuracy rate is 99.99%, and the user appeal successful explanation rate is 100%.
[0245] It is understood that in the embodiments of this application, if data related to user information or enterprise information is involved, when the embodiments of this application are applied to specific products or technologies, it is necessary to obtain user permission or consent, or to obfuscate this information in order to eliminate the correspondence between this information and the user; and the collection and processing of related data should strictly comply with the requirements of relevant laws and regulations when applied in practice, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0246] The following continues to describe the exemplary structure of the abnormal webpage identification device 454 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 4As shown, the abnormal webpage identification device 454 includes: a construction module 4541, used to construct a first intelligent agent, a second intelligent agent, and a third intelligent agent in a multimodal large language model; an input module 4542, used to input the webpage information of the webpage to be identified into the multimodal large language model; a screening module 4543, used to screen the webpage information through the first intelligent agent to obtain screening information for abnormal webpage identification; a collection module 4544, used to collect information from the webpage to be identified based on the webpage information and the screening information through the second intelligent agent to obtain collected information; and an identification module 4545, used to identify the webpage to be identified as abnormal through the third intelligent agent based on the webpage information and the collected information, and output the abnormal identification result through the multimodal large language model.
[0247] In some embodiments, the apparatus further includes a construction module, configured to acquire a first pre-trained model, and a first role prompting information, a first task prompting information, and a first input / output prompting information corresponding to the first pre-trained model; merge the first role prompting information, the first task prompting information, and the first input / output prompting information to obtain first prompting information; and construct the first agent in the multimodal large language model based on the first prompting information and the first pre-trained model.
[0248] In some embodiments, the screening module is further configured to: fill the webpage information into the first input / output prompt information to obtain updated first prompt information; and, through the first pre-trained model in the first agent, perform information screening on the webpage information in the updated first prompt information based on the first role prompt information and the first task prompt information to obtain screening information for abnormal webpage identification.
[0249] In the above scheme, the device further includes an output module for outputting the screening information in accordance with the output format specified in the updated first prompt information.
[0250] In some embodiments, the construction module is further configured to: obtain a second pre-trained model, and second role prompt information, second task prompt information, and second input / output prompt information corresponding to the second pre-trained model; merge the second role prompt information, the second task prompt information, and the second input / output prompt information to obtain second prompt information; and construct the second agent in the multimodal large language model based on the second prompt information and the second pre-trained model.
[0251] In some embodiments, the collection module is further configured to: input the second prompt information, the webpage information, and the screening information into the second intelligent agent to obtain an output result; when the output result indicates that extended information needs to be collected based on the screening information, parse the tool identifier and input data from the output result; collect target extended information through the information collection tool corresponding to the tool identifier according to the input data, and concatenate the target extended information with the screening information to obtain collected information; when the output result indicates that extended information does not need to be collected, determine the screening information as the collected information.
[0252] In some embodiments, the device further includes a splicing module, used to extract features from the text data when the input data is text data to obtain a text embedding vector; to search for target text vectors with a similarity higher than a first preset threshold in a pre-built text vector library using the information collection tool; and to splice the information corresponding to the target text vectors with the screening information to obtain the collected information.
[0253] In some embodiments, the splicing module is further configured to: when the input data is image data, perform feature extraction on the image data to obtain an image embedding vector; search for a target image vector with a similarity higher than a second preset threshold in a pre-built image vector library using the information collection tool; and splice the information corresponding to the target image vector with the screening information to obtain the collected information.
[0254] In some embodiments, the construction module is further configured to: obtain a third pre-trained model, and third role prompt information, third task prompt information, and third input / output prompt information corresponding to the third pre-trained model; merge the third role prompt information, the third task prompt information, and the third input / output prompt information to obtain third prompt information; and construct the third agent in the multimodal large language model based on the third prompt information and the third pre-trained model.
[0255] In some embodiments, the identification module is further configured to: fill the webpage information and the collected information into the third input / output prompt information to obtain updated third prompt information; perform anomaly identification on the webpage information in the updated third prompt information based on the third role prompt information, the third task prompt information and the collected information in the updated third prompt information using the third pre-trained model in the third agent to obtain anomaly identification results; and determine the abnormal webpage identification result of the webpage to be identified based on the anomaly identification results.
[0256] In some embodiments, the output module is further configured to: output the abnormal webpage identification result according to the output format specified by the updated third prompt information through the multimodal large language model.
[0257] It should be noted that the description of the apparatus in this application embodiment is similar to the description of the method embodiment described above, and has similar beneficial effects as the method embodiment; therefore, it will not be repeated. For technical details not disclosed in this apparatus embodiment, please refer to the description of the method embodiment of this application for understanding.
[0258] This application provides an electronic device, including: a memory for storing computer-executable instructions or computer programs; and a processor for executing the computer-executable instructions or computer programs stored in the memory to implement the above-described abnormal webpage identification method.
[0259] This application provides a computer program product, which includes computer-executable instructions or a computer program, and the computer-executable instructions or computer program are stored in a computer-readable storage medium; wherein, when the processor of the electronic device reads the computer-executable instructions or computer program from the computer-readable storage medium and executes the computer-executable instructions or computer program, the above-mentioned abnormal webpage identification method is implemented.
[0260] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the aforementioned abnormal webpage identification method, for example, such as... Figure 5 The method shown.
[0261] In some embodiments, the storage medium may be a computer-readable storage medium, such as a ferromagnetic random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEROM), flash memory, magnetic surface memory, optical disk, or a compact disk-read-only memory (CD-ROM); or it may be a device that includes one or any combination of the above-mentioned memories.
[0262] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0263] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file containing other programs or data, for example, in one or more scripts within a Hypertext Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., files storing one or more modules, subroutines, or code sections). As an example, executable instructions may be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0264] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A method for identifying abnormal web pages, characterized in that, The method includes: Construct a first agent, a second agent, and a third agent in a multimodal large language model; The webpage information of the webpage to be identified is input into the multimodal large language model; The first intelligent agent performs information screening on the web page information to obtain screening information that needs to be identified as abnormal web pages. The second intelligent agent collects information about the webpage to be identified based on the webpage information and the screening information, and obtains the collected information. The third intelligent agent identifies abnormal web pages based on the web page information and the collected information, and outputs the abnormal web page identification result through the multimodal large language model.
2. The method according to claim 1, characterized in that, Constructing the first intelligent agent in a multimodal large language model includes: Obtain a first pre-trained model, and the first role prompt information, first task prompt information, and first input / output prompt information corresponding to the first pre-trained model; The first role prompt information, the first task prompt information, and the first input / output prompt information are combined to obtain the first prompt information; Based on the first prompt information and the first pre-trained model, the first intelligent agent is constructed in the multimodal large language model.
3. The method according to claim 2, characterized in that, The step of using the first intelligent agent to screen the webpage information to obtain screening information for identifying abnormal webpages includes: The webpage information is filled into the first input / output prompt information to obtain the updated first prompt information; Using the first pre-trained model in the first intelligent agent, based on the first role prompt information and the first task prompt information in the updated first prompt information, information screening is performed on the web page information in the updated first prompt information to obtain the screening information that needs to be identified as abnormal web pages. The method further includes: outputting the screening information according to the output format specified in the updated first prompt information.
4. The method according to claim 3, characterized in that, Constructing a second agent in a multimodal large language model includes: Obtain the second pre-trained model, as well as the second role prompt information, the second task prompt information, and the second input-output prompt information corresponding to the second pre-trained model; The second role prompt information, the second task prompt information, and the second input / output prompt information are combined to obtain the second prompt information; Based on the second prompt information and the second pre-trained model, the second agent is constructed in the multimodal large language model.
5. The method according to claim 4, characterized in that, The second intelligent agent collects information about the webpage to be identified based on the webpage information and the screening information to obtain collected information, including: The second prompt information, the webpage information, and the screening information are input into the second intelligent agent to obtain the output result; When the output result indicates that it is necessary to collect extended information based on the screening information, the tool identifier and input data are parsed from the output result. Based on the input data, target extended information is collected through the information collection tool corresponding to the tool identifier, and the target extended information is concatenated with the screening information to obtain the collected information; When the output result indicates that no extended information needs to be collected, the screening information is determined as the collected information.
6. The method according to claim 5, characterized in that, The step involves collecting target extended information using the information collection tool corresponding to the tool identifier based on the input data, and concatenating the target extended information with the screening information to obtain collected information, including: When the input data is text data, feature extraction is performed on the text data to obtain a text embedding vector; Using the information collection tool, target text vectors with a similarity higher than a first preset threshold to the text embedding vector are searched in a pre-built text vector library. The information corresponding to the target text vector is concatenated with the screening information to obtain the collected information.
7. The method according to claim 5, characterized in that, The step involves collecting target extended information using the information collection tool corresponding to the tool identifier based on the input data, and concatenating the target extended information with the screening information to obtain collected information, including: When the input data is image data, feature extraction is performed on the image data to obtain an image embedding vector; Using the information collection tool, target image vectors with a similarity higher than a second preset threshold to the image embedding vector are searched in a pre-built image vector library; The information corresponding to the target image vector is concatenated with the screening information to obtain the collected information.
8. The method according to claim 5, characterized in that, Constructing a third agent in a multimodal large language model includes: Obtain the third pre-trained model, as well as the third role prompt information, third task prompt information, and third input / output prompt information corresponding to the third pre-trained model; The third role prompt information, the third task prompt information, and the third input / output prompt information are combined to obtain the third prompt information; Based on the third prompt information and the third pre-trained model, the third intelligent agent is constructed in the multimodal large language model.
9. The method according to claim 8, characterized in that, The process of identifying abnormal web pages based on the web page information and the collected information by the third intelligent agent, and outputting the abnormal web page identification result through the multimodal large language model, includes: The webpage information and the collected information are filled into the third input / output prompt information to obtain the updated third prompt information; Using the third pre-trained model in the third intelligent agent, based on the third role prompt information, third task prompt information and collected information in the updated third prompt information, anomaly identification is performed on the web page information in the updated third prompt information to obtain anomaly identification results; Based on the anomaly identification results, the anomaly identification result of the webpage to be identified is determined; The method further includes: using the multimodal large language model, outputting the abnormal webpage recognition result according to the output format specified by the updated third prompt information.
10. An abnormal webpage identification device, characterized in that, The device includes: The building blocks are used to construct the first, second, and third agents in a multimodal large language model; The input module is used to input the webpage information of the webpage to be identified into the multimodal large language model; The screening module is used to screen the web page information through the first intelligent agent to obtain screening information that needs to be identified as abnormal web pages; The collection module is used to collect information from the webpage to be identified based on the webpage information and the screening information through the second intelligent agent, and obtain the collected information. The identification module is used by the third intelligent agent to identify abnormal web pages based on the web page information and the collected information, and outputs the abnormal identification result through the multimodal large language model.
11. An electronic device, characterized in that, include: Memory is used to store executable instructions or computer programs. A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the abnormal webpage identification method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions or computer programs for causing a processor to execute the computer-executable instructions or computer programs to implement the abnormal webpage identification method according to any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, the abnormal webpage identification method according to any one of claims 1 to 9 is implemented.