Data processing method, device, electronic device and storage medium based on large model
Through a data processing method based on a large model, combined with website feature parameters and keyword strategies, the problems of high false alarm rate and high resource consumption in identifying black market intelligence in traditional methods are solved, and efficient and accurate abnormal data identification and intelligence report generation are achieved.
Patent Information
- Application Number
- CN202411866008.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-17
AI Technical Summary
When identifying and managing black market intelligence, existing technologies such as traditional network monitoring and data analysis methods based on pre-trained models have problems such as high false alarm rate, high missed alarm rate, high maintenance cost, large computing resource requirements, and difficulty in model interpretation.
A data processing method based on a large model is used to determine the target website through website feature parameters and keyword strategies. The large model is used for data summary and report generation. In combination with positive and negative keyword strategies, the accuracy and relevance of data collection are improved, and the demand for computing resources is reduced.
Effectively identify high-value target websites, reduce false positive and false negative rates, save computing resources, improve the accuracy and efficiency of abnormal data mining, and support multi-language processing and efficient management of intelligence reports.
Smart Images

Figure CN119669908B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and big data technology, particularly internet security technology, and can be applied to scenarios where abnormal data is processed based on large models. More specifically, the present disclosure provides a data processing method, device, electronic device, and storage medium based on large models. Background Art
[0002] With the rapid development of the internet industry, more and more internet-related black industries ("black industries") have emerged, increasing network risks and posing a threat to internet security. Currently, the main solutions for mining and managing black industry intelligence are traditional network monitoring and filtering, as well as data analysis based on pre-trained models. Summary of the Invention
[0003] The present disclosure provides a data processing method, device, electronic device and storage medium based on a large model.
[0004] According to one aspect of the present disclosure, a data method based on a big model is provided, the method comprising: determining at least one target website from at least one website based on at least one first characteristic parameter and at least one second characteristic parameter of each website in at least one website, the first characteristic parameter being used to characterize the business matching degree of the website, and the second characteristic parameter being used to represent the activity degree of the website; collecting data from at least one target website based on at least one first keyword and at least one second keyword, the first keyword being used to collect data items including the first keyword in the data, and the second keyword being used to delete data items including the second keyword in the data; inputting the data into the big model, outputting summary information of the data; and generating a first report based on the summary information.
[0005] According to another aspect of the present disclosure, a data processing device based on a big model is provided, which includes: a first determination module, used to determine at least one target website from at least one website based on at least one of the first characteristic parameter and the second characteristic parameter of each website in the at least one website, the first characteristic parameter is used to characterize the business matching degree of the website, and the second characteristic parameter is used to represent the activity of the website; a collection module, used to collect data from at least one target website based on at least one first keyword and at least one second keyword, the first keyword is used to collect data items including the first keyword in the data, and the second keyword is used to delete data items including the second keyword in the data; an input and output module, used to input data into the big model and output summary information of the data; and a generation module, used to generate a first report based on the summary information.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided. The computer instructions are used to cause a computer to execute the method provided according to the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method provided according to the present disclosure when executed by a processor.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0011] Figure 1 is a schematic diagram of an exemplary system architecture to which a large model-based data processing method and apparatus can be applied according to an embodiment of the present disclosure;
[0012] Figure 2 is a flow chart of a data processing method based on a large model according to an embodiment of the present disclosure;
[0013] Figure 3 A schematic diagram schematically illustrates a keyword determination process according to an embodiment of the present disclosure;
[0014] Figure 4 The following schematically illustrates a method for collecting data based on the priority of keywords according to an embodiment of the present disclosure;
[0015] Figure 5 A schematic diagram of applying the large model-based data processing method according to an embodiment of the present disclosure to mine illegal data is shown;
[0016] Figure 6 is a block diagram of a data processing apparatus based on a large model according to an embodiment of the present disclosure;
[0017] Figure 7 Schematically shows a structural block diagram of a large model of artificial intelligence according to an embodiment of the present disclosure; and
[0018] Figure 8 A schematic block diagram of an example electronic device 800 is shown, which may be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0020] In the technical solutions disclosed herein, the information and data involved (including but not limited to data used for analysis, stored data, displayed data, etc.) are all authorized information and data, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for the choice of authorization or rejection.
[0021] Currently, the primary approaches for mining and managing illegal intelligence rely on traditional network monitoring and filtering, as well as data analysis based on pre-trained models. Network monitoring and filtering primarily rely on pre-defined rules and keywords to detect and prevent anomalous activity on the network. Pre-trained models analyze large amounts of data to automatically identify potential illegal behavior and threat patterns.
[0022] Network monitoring and filtering methods use simple keyword filtering, which is prone to false positives and false negatives, making it difficult to accurately identify complex illegal activities. Furthermore, this method requires constant updating of rules and keywords, resulting in high maintenance costs.
[0023] Data analysis methods based on pre-trained models require large amounts of high-quality data for training. Insufficient data can lead to model inaccuracies. Furthermore, some complex models are difficult to explain their decision-making processes, increasing the cost of trust. Furthermore, training and running complex models requires significant computing resources.
[0024] In view of this, an embodiment of the present disclosure provides a data processing method based on a big model, which determines at least one target website from at least one website based on at least one first characteristic parameter and at least one second characteristic parameter of each website in at least one website, the first characteristic parameter is used to characterize the business matching degree of the website, and the second characteristic parameter is used to represent the activity of the website; data from at least one target website is collected based on at least one first keyword and at least one second keyword, the first keyword is used to collect data items including the first keyword in the data, and the second keyword is used to delete data items including the second keyword in the data; the data is input into the big model, summary information of the data is output; and a first report is generated based on the summary information.
[0025] Figure 1 This is a schematic diagram of an exemplary system architecture to which a large-model-based data processing method and apparatus can be applied according to an embodiment of the present disclosure. It should be noted that: Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0026] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0027] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with display screens and support web browsing, including but not limited to smartphones, tablet computers, laptop computers, and desktop computers, etc.
[0028] Server 105 may be a server that provides various services, such as a background management server (for example only) that processes data for websites browsed by users using terminal devices 101, 102, and 103. The background management server may collect data from the websites, input the data into a large model, output data summaries, generate a first report based on the summary information, and save or feed the first report back to the terminal device.
[0029] It should be noted that the data processing method based on the big model provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the data processing device based on the big model provided in the embodiment of the present disclosure can generally be set in the server 105. The data processing method based on the big model provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the data processing device based on the big model provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.
[0030] It should be understood that Figure 1 The number and type of terminal devices, networks and servers in the embodiment are merely illustrative. Any number and type of terminal devices, networks and servers may be used as required.
[0031] Figure 2 is a flowchart of a data processing method based on a large model according to an embodiment of the present disclosure.
[0032] like Figure 2 As shown, the method 200 may include operations S210 to S230.
[0033] In operation S210, at least one target website is determined from the at least one website based on at least one of the first characteristic parameter and the second characteristic parameter of each website in the at least one website.
[0034] According to an embodiment of the present disclosure, the first characteristic parameter can be used to characterize the business compatibility of a website. The business compatibility of a website can also be referred to as the degree of fit between the website's business value and the core business or target market of the website, such as the website's content, design, functionality, and the services it provides.
[0035] According to an embodiment of the present disclosure, the second characteristic parameter can be used to represent the activity of a website. Activity can be the frequency and depth of visits and use of a website or its services within a certain period of time. It can include login frequency, number of pages viewed, dwell time, interactive behaviors (such as comments, likes, shares, etc.), and conversion behaviors such as purchases and registrations.
[0036] According to embodiments of the present disclosure, the at least one website can be understood as a collection of websites that have abnormal data (e.g., illegal) in historical abnormal data (e.g., illegal) statistics. The at least one target website can be understood as a collection of websites that need to be counted for abnormal data in the current application scenario, and can be determined from the at least one website. The website can also refer to a community website for mobile devices and software development.
[0037] In one possible implementation of the disclosed embodiment, each website may be prioritized and sorted based on its first characteristic parameter, with the highest-priority websites identified as target websites as the focus of abnormal data. For example, the top 50 or top 100 websites may be identified as target websites, although this disclosure does not limit this. A higher priority can be understood as indicating a greater degree of match between the website's content, design, functionality, and the services it provides and the website's core business or target market.
[0038] In another possible implementation of the disclosed embodiment, each website can be prioritized and sorted based on its second characteristic parameter, with websites with higher priorities being identified as target websites, serving as the focus of attention for abnormal data (e.g., illegal). For example, the top 50 or top 100 websites can be identified as target websites, although this disclosure does not limit this. A higher priority can be understood as indicating a more active website.
[0039] In another possible implementation of the embodiment of the present disclosure, each website can be prioritized and sorted based on its first and second characteristic parameters, with websites with higher priorities being identified as target websites as the focus of abnormal data (e.g., illegal). For example, the top 50 websites or the top 100 websites can be identified as target websites, although this disclosure does not limit this. A higher priority can be understood as a higher degree of match between the website's content, design, functionality, and the services it provides and the website's core business or target market, and a higher level of activity on the website.
[0040] In operation S220 , data from at least one target website is collected based on the at least one first keyword and the at least one second keyword.
[0041] According to the embodiments of the present disclosure, the data collected from websites can be public domain information, meaning it is accessible to the public within the scope permitted by law and does not pose a threat to any confidential information. When the collected website data involves private domain information, authorization or consent has been obtained from the user or the institution to which the data belongs, and the collection and processing of the data complies with relevant laws and regulations and does not violate public order and good morals.
[0042] According to an embodiment of the present disclosure, a first keyword can be used to collect data items that include the first keyword in the data. A second keyword can be used to delete data items that include the second keyword in the data. The first keyword can be called a positive keyword, that is, a keyword that has a positive effect on collecting abnormal data. The data item that includes the first keyword in the data is, for example, a data item that may be abnormal data. The second keyword can be called a negative keyword, that is, a keyword that has an interfering effect on collecting abnormal data, that is, a keyword that needs to be excluded. During the data collection process, the data item containing the second keyword is directly excluded and not collected to avoid capturing irrelevant or unnecessary information.
[0043] In operation S230, data is input into the large model, and summary information of the data is output.
[0044] According to the embodiments of the present disclosure, even if data from target websites is collected by designing positive and negative keywords, the amount of data collected may be large. To quickly identify abnormal data from the collected data, the interface of the understanding and summarization function of the large model can be called, and each collected data item can be input into the large model. The lengthy data content and related comments are summarized to obtain summary information for each data item.
[0045] In operation S240 , a first report is generated based on the summary information.
[0046] According to an embodiment of the present disclosure, based on summary data, it is possible to quickly determine which data is abnormal data, and then convert the abnormal data into a report for subsequent review.
[0047] By using the data processing method of the embodiment, the websites of focus are determined by at least one of the first characteristic parameter and the second characteristic parameter of each website, which can effectively identify and focus on high-value target websites, and is more targeted. Compared with simple keywords, the strategy of combining positive and negative keywords can effectively improve the accuracy and relevance of abnormal data mining, reduce the interference of irrelevant information, and thus reduce the false alarm rate and missed alarm rate. On this basis, since the object of collection is a high-value target website, the interference data captured is reduced, which can reduce the amount of data collection and reduce the amount of large model calculations, thereby saving computing resources and hardware resources.
[0048] The following will be combined Figure 3 The keyword determination process of the embodiment of the present disclosure is schematically described. Figure 3 The following is a schematic diagram schematically illustrating a keyword determination process according to an embodiment of the present disclosure.
[0049] like Figure 3 As shown, in some embodiments, the keyword determination process may be:
[0050] At least one first initial keyword and at least one second initial keyword of each target website are determined based on the first priori information.
[0051] Based on at least one of search traffic, search frequency, and keyword difficulty values of the at least one first initial keyword and the at least one second initial keyword, at least one first keyword is determined from the at least one first initial keyword, and at least one second keyword is determined from the at least one second initial keyword.
[0052] According to an embodiment of the present disclosure, the first prior information 311 includes keywords used to determine historical abnormal data, that is, existing or known knowledge, experience, or information in the process of mining historical abnormal data.
[0053] For example, during the historical anomaly data mining process, website data is collected based on pre-set keywords. If, after analysis, it is determined that the data mined based on these pre-set keywords is indeed anomaly data, these pre-set keywords can be stored as first prior information 311a for subsequent use in anomaly data mining. In this embodiment, the first initial keyword 312a can be determined based on the first prior information 311a.
[0054] For another example, during historical anomaly data mining, website data is collected based on pre-set keywords. After analysis, it is determined that some of the data mined based on these pre-set keywords is abnormal data and some is normal data. These pre-set keywords for the collected abnormal data can be stored as first prior information 311a, and these pre-set keywords for the collected normal data can be stored as first prior information 311b for subsequent use in anomaly data mining. In this embodiment, the second initial keywords 312b can be determined based on the first prior information 311b to eliminate interference information during the data collection process.
[0055] According to an embodiment of the present disclosure, after obtaining the first initial keyword 312a and the second initial keyword 312b, a search engine optimization strategy can be employed to analyze the keyword strategy of the target website and determine which keywords have high value within the target domain of the target website. Indicators of the search engine optimization strategy may include at least one of keyword search traffic, keyword search frequency, and keyword difficulty. Keyword search traffic is a metric that measures the number of times a keyword is searched on a search engine, reflecting user interest and demand for a particular keyword. Keyword search frequency is the total number of times users search for a specific keyword via a search engine within a certain period of time. This metric is an important indicator for measuring keyword popularity, user interest, and market demand. Keyword difficulty is an indicator that measures the difficulty of achieving a high ranking for a specific keyword on a search engine results page. For example, it can be a score or percentage ranging from 0 to 100 (or 0% to 100%). A larger value indicates a higher keyword difficulty, meaning it is more difficult to achieve a high ranking.
[0056] It should be understood that before obtaining the search traffic and frequency of keywords, the user's authorization or consent is obtained, and the collection and processing of data complies with relevant laws and regulations and does not violate public order and good customs.
[0057] Based on the indicators of the search engine optimization strategy, the initially obtained first initial keyword 312 a and second initial keyword 312 b may be optimized to determine the final first keyword 321 and second keyword 331 .
[0058] It should be noted that different target websites may have different corresponding search engine optimization strategies, and the first keyword 321 and the second keyword 331 obtained based on the analysis may also be different.
[0059] It should be noted that, in the keyword determination process, the preliminarily determined first initial keyword 312a and second initial keyword 312b can also be directly used as the first keyword 321 and the second keyword 331 without performing a search engine optimization strategy. They can be optimized according to specific application scenarios, and this disclosure does not limit them.
[0060] It should be noted that the first keyword and the second keyword can be determined according to the topic, that is, each topic can correspond to its own first keyword and second keyword. The topic can include, for example, technology-related topics such as xposed, hook, root, positioning tampering, etc.
[0061] The data processing method of this embodiment determines keywords based on prior information, ensuring the accuracy of keyword determination. Further optimization of the initially determined keywords based on search engine optimization strategy indicators can improve data collection efficiency, further reduce interference information, and effectively improve the accuracy of abnormal data mining.
[0062] The following will be combined Figure 4 The process of collecting target websites based on keywords in an embodiment of the present disclosure is schematically described. Figure 4 The diagram schematically shows a schematic diagram of data collection based on the priority of keywords according to an embodiment of the present disclosure.
[0063] In some embodiments, collecting data of the at least one target website based on the at least one first keyword and the at least one second keyword may include:
[0064] For each target website in the at least one target website, the priority of each keyword in the at least one first keyword is determined respectively, and the priority is used to indicate the collection order of the keywords.
[0065] Based on the respective priorities of the target websites, data of the target websites are collected in the keyword collection order indicated by the priorities.
[0066] Based on the respective priorities of the target websites, tags are added to the data items containing the first keywords.
[0067] According to an embodiment of the present disclosure, different types of target websites may have the same or different corresponding abnormal data, and the keywords contained in the corresponding abnormal data may be the same or different. In addition, there may be multiple pieces of abnormal data corresponding to a target website.
[0068] like Figure 4 As shown, for example, based on historical data, for target website 420a, target website 420a may contain abnormal data 421a, abnormal data 422b, and abnormal data 423c, and the corresponding first keywords may be first keyword 421, first keyword 422, and first keyword 423, respectively. The number of entries in abnormal data 421a, for example, is greater than the number of entries in abnormal data 422b, and the number of entries in abnormal data 421b, for example, is greater than the number of entries in abnormal data 422c. That is, for target website 420a, the probability of abnormal data 421a appearing is greater than the probability of abnormal data 422b appearing, and the probability of abnormal data 421b appearing is greater than the probability of abnormal data 422c appearing. Therefore, the priorities of first keywords 421, 422, and 423 may be, for example, first keyword 421 having a higher priority than first keyword 422, and first keyword 422 having a higher priority than first keyword 423.
[0069] For target website 420b, target website 420b may contain abnormal data 421a, abnormal data 422b, and abnormal data 424d, and the corresponding first keywords may be first keyword 421, first keyword 422, and first keyword 424, respectively. For example, the number of entries in abnormal data 422b is greater than the number of entries in abnormal data 421a, and the number of entries in abnormal data 421a is greater than the number of entries in abnormal data 422d. That is, for target website 420b, the probability of abnormal data 422b appearing is greater than the probability of abnormal data 421a appearing, and the probability of abnormal data 421a appearing is greater than the probability of abnormal data 424d appearing. Therefore, the priorities of first keywords 421, 422, and 424 may be, for example, that the priority of first keyword 422 is greater than the priority of first keyword 421, and the priority of first keyword 421 is greater than the priority of first keyword 424.
[0070] The data from the target website 420a and the target website 420b are collected respectively according to the priorities determined above.
[0071] The data processing method of this embodiment sets the search priority of the first keyword for different types of target websites, and collects data from the target websites according to their respective priorities. This allows the keyword strategy to better match the data characteristics of each website, further effectively improving the accuracy and relevance of abnormal data capture. Furthermore, adding tags after data collection based on keyword priority facilitates subsequent review of highly relevant data.
[0072] In some embodiments, determining the priority of each keyword in the at least one first keyword may include:
[0073] The priority of each keyword in the at least one first keyword is determined based on at least one of a long-tail keyword included in the at least one first keyword, a combination of uppercase and lowercase letters of the first keyword, homophones, typos, slang, and abbreviations.
[0074] For example, for a target website using a first language (e.g., Chinese), a first keyword containing homophones or misspellings may have a higher priority than a first keyword containing uppercase and lowercase combinations, slang, abbreviations, or long-tail keywords. For another example, for a target website using a second language (e.g., English), a first keyword containing uppercase and lowercase combinations, slang, abbreviations, or long-tail keywords may have a higher priority than a first keyword containing homophones or misspellings. The priority of the first keyword can be determined based on the type of the actual website, and the embodiments of the present disclosure are not limited thereto.
[0075] Through the data processing method of this embodiment, the priority of each first keyword is determined based on at least one of the long-tail keywords contained in at least one first keyword, the uppercase and lowercase combination of the first keyword, homophones, typos, slang, and abbreviations, so that the priority of keywords that are more suitable for each target website is determined, thereby making the keyword strategy more in line with the data characteristics of each website, and further effectively improving the accuracy and relevance of abnormal data capture.
[0076] In some embodiments, after determining the first keyword and the second keyword, a crawling technique may be used to collect data from the target website based on the first keyword and the second keyword. The crawling technique may be used to directly traverse each web page node of the target website to obtain data, such as comment posts.
[0077] The crawler strategy used by crawler technology can be a delay and time strategy. In one possible implementation, the delay and time strategy can be a fixed delay, which sends query requests to the server by setting a fixed delay time to obtain data from various target websites. In another possible implementation, the delay and time strategy can be a random delay, which generates random delay times so that the interval between requests conforms to natural access behavior. In yet another possible implementation, the delay and time strategy can be adaptive, dynamically adjusting the delay time of the next request based on the result or response time of the previous request. For example, if the response time of the previous request is long, the delay time of the next request can be appropriately increased; conversely, if the response time is short, the delay time can be appropriately reduced. Based on the delay and time strategy, sending a large number of requests in a short period of time to burden the server is avoided.
[0078] In some embodiments, the data collected from each target website can be stored according to tags. For example, the storage method can be "tag: website: link: title + time: content: comment + time". Subsequently, based on the tags, the set can be viewed to obtain the corresponding saved data according to business needs.
[0079] In some embodiments, generating a first report based on the summary information may include:
[0080] The summary information is filtered based on second priori information, and abnormal data is determined from the data based on the filtered summary information, where the second priori information includes historical abnormal data.
[0081] A first tag is added to the abnormal data, where the first tag is used to indicate at least one of the following: a derivative tool included in the target website where the abnormal data exists, a target website version, a model of a device used to log in to the target website, a vulnerability of the device, a tampered file, and a scenario where the abnormal data exists on the target website.
[0082] A first report is generated based on the first label and the abnormal data.
[0083] According to the embodiments of the present disclosure, since the data crawled based on the first and second keywords is not necessarily all abnormal data and may contain misjudgments, the currently acquired data can also be secondary screened based on historical abnormal data, deleting the misjudged data, determining the data that is truly abnormal, and extracting key information such as the technology, tools, and activities associated with the abnormal data. The remaining data can also be stored for subsequent learning without being discarded. After extracting this key information, more detailed labels can be added to the abnormal data based on this key information.
[0084] By performing a secondary screening of the acquired data based on the first prior information, the accuracy of the collected abnormal data can be further improved. Labeling the abnormal data based on the key information facilitates accurate review of the corresponding type of abnormal data in the later stage.
[0085] In some embodiments, generating the first report based on the summary information may include:
[0086] According to the stage type of data processing, a first report with corresponding timeliness is generated based on data, summary information and abnormal data. The stage types include data collection stage, reporting stage, risk strategy formulation stage and summary stage.
[0087] According to an embodiment of the present disclosure, during the data collection phase, a first report can be generated based on the collected original web page information of the target website. During the risk strategy formulation phase, the original web page information or summary information can be mined to obtain risk information, and a first report can be generated based on the risk information. During the reporting phase, a business threat report can be generated as the first report based on the risk information. During the summary phase, the reports produced in the previous phases can be summarized to generate a first report, such as a detailed security threat and risk monitoring report.
[0088] According to the embodiments of the present disclosure, first reports of different periods can be selectively uploaded according to business needs and timeliness to build a full life cycle reporting system.
[0089] Through the data processing method of this embodiment, relevant reports that meet corresponding timeliness are generated for different stage types, which can meet the timeliness requirements of abnormal data collection and analysis.
[0090] In some embodiments, generating the first report based on the summary information may further include:
[0091] A second tag is added to the first report, where the second tag includes at least an identity tag of the target website to which the first report belongs, a content index tag of the abnormal data, and life cycle tags of the first report of different stages and types.
[0092] The first report is stored and edited based on the first tag and the second tag.
[0093] According to an embodiment of the present disclosure, when the number of first reports exceeds a certain threshold, the first reports may be labeled to facilitate subsequent viewing.
[0094] The following will be combined Figure 5 The application of the large model-based data processing method according to the embodiment of the present disclosure to mining illegal data is schematically described. Figure 5 A schematic diagram of applying the large model-based data processing method according to an embodiment of the present disclosure to mine illegal data is shown schematically.
[0095] like Figure 5 As shown, based on the data processing method of the big model, the mining of illegal data can include four stages: preliminary work construction, data collection and identification, risk strategy formulation and report generation, and intelligence information construction.
[0096] For example, the preliminary work may include website account management, target website selection, and keyword library construction, that is, a positive keyword library containing the first keyword and a negative keyword library containing the second keyword.
[0097] For example, data collection and recognition can include information crawling, information storage, and target recognition.
[0098] For example, risk strategy development and report generation can include risk assessment, strategy formulation, and report generation. Risk assessment can involve secondary screening of summary information based on prior information to identify actual illegal data. Strategy development can involve mining illegal data, collecting illegal techniques, tools, and behaviors, conducting risk mining on detailed illegal tools, and developing relevant risk control strategies. Report generation can be based on "original webpage information," "summary summary," and "risk mining information," combined with actual cheating scenarios, relevant technical principles, detection methods, and strategies. This allows for the relatively easy production of a security threat and risk monitoring report. The report should include a report summary, information tags, source material information, illegal attack methods, illegal attack principles, illegal cheating scenarios, reproduction of illegal cheating methods, risk characteristics, safety factors, usage strategies, a brief description of relevant code, and the current status of strategy implementation. Detailed reporting is crucial for sharing illegal data.
[0099] For example, intelligence information construction can include intelligence database construction, information sharing and cooperation, and continuous optimization learning.
[0100] Intelligence database construction can include:
[0101] Based on previous work, we can generate different reports at different stages for the illegal techniques, tools, and actions involved in a specific forum: a "Risk Information Summary Report" based on a large model during information crawling, a "Business Threat Report" for reporting business value confirmation, a "Detection Strategy Report" for risk control strategy formulation, and a detailed "Security Threat and Risk Monitoring Report." Based on business needs and timeliness, we selectively upload relevant reports from different periods to build a full-lifecycle threat intelligence system.
[0102] When the number of reports reaches a certain level, it is necessary to carry out tagging. Based on previous work, we currently have multiple dimensions for report management, such as "topic tags", "content index tags", "secondary tags", and "lifecycle tags", which can be as follows:
[0103] 「Theme Tag」: The initially set theme, such as custom, xposed, hook, root, modified device, etc.
[0104] Content Index Tags: Based on crawler information: platform, link, time, etc. The value of this tag is that it can be filtered by time.
[0105] "Secondary Label": covers detailed illegal tools, system versions, device models, usage scenarios, Android vulnerabilities, modified files, etc.
[0106] "Life cycle label": information crawling period, business value confirmation period, risk control strategy formulation period, and final report of business review.
[0107] Given the large number of reports produced, the security team's diverse scope, distinct personalities, and meticulous focus, templated reports are necessary for efficient, aesthetically pleasing, and information-sharing purposes. Initially, we believe that each period and each detailed task requires a corresponding template.
[0108] It is of great help to other technical personnel in technical learning and business understanding; it increases the aesthetics of the report; facilitates information indexing; the prescribed modules are convenient and efficient in output; according to business needs, custom report writing modules are added to facilitate other personnel to understand the incremental information in the report.
[0109] If the technical staff feels that the report has too many modules, they can select the modules they need or use the "short report template". For example, some reports are just a table, but it is better to include an abstract or background and purpose.
[0110] Information sharing and collaboration can include:
[0111] Internal sharing: The granularity can be divided into business personnel, security SDK team members, and risk control team members.
[0112] Sharing on multiple platforms: In order to expand the team's influence, some reports can be published on major platforms after deleting some information.
[0113] After it has reached a certain scale, it can assist in the development of other businesses.
[0114] Resource sharing: involves technical resources, servers, network resources, etc. For example, when we build crawler IP resources, we may need to use the IP pool or account management resources of other units.
[0115] Data sharing: Share the latest threat information and attack patterns, IP information sources, and mobile security vulnerabilities with other friendly companies.
[0116] Strategy sharing: Communicate and share effective security defense strategies with other friendly companies.
[0117] Continuous learning and optimization can include:
[0118] Mid-term evaluation of system effectiveness: Based on the first phase of intelligence system construction, the overall construction environment will be optimized and feedback will be provided based on the alignment with business needs, the accuracy and effectiveness of intelligence results, and resource utilization. Optimization will focus on site target selection, crawler keywords, crawler strategies, model summary and translation, and other aspects.
[0119] Review of the development trend of illegal technologies: Based on existing reports, quarterly and annual report summaries can be prepared for a specific illegal topic.
[0120] It should be noted that the data collection involved in the above-mentioned illegal data mining are all authorized information and data, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data all comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for people to choose to authorize or refuse.
[0121] According to the embodiments of the present disclosure, the above-mentioned illegal data mining method, by setting target priorities and building a keyword library, can more effectively identify and focus on high-value targets, and is more targeted than competitors. The use of positive and negative keyword strategies, combined with SEO analysis, can improve the accuracy and relevance of information crawling and reduce the interference of irrelevant information. Utilize large model understanding and summarization technology to quickly screen and process large amounts of information, improve efficiency, and support multi-language processing capabilities. Through labeling management and standardized templates, intelligence reports can be efficiently organized and managed, facilitating information retrieval and sharing. Based on risk assessment and strategy formulation, it can respond and adjust strategies quickly to enhance the agility and adaptability of the system. Support multi-platform information sharing and cooperation with other manufacturers to expand influence and enhance the value of intelligence. The system design focuses on continuous learning and optimization, and can be dynamically adjusted according to business needs and technological trends to maintain competitive advantage.
[0122] Figure 6 is a block diagram of a data processing apparatus based on a large model according to an embodiment of the present disclosure.
[0123] like Figure 6 As shown, the large model-based data processing device 600 may include a first determination module 610 , a collection module 620 , an input and output module 630 and a generation module 640 .
[0124] The first determination module 610 is used to determine at least one target website from at least one website based on at least one of the first characteristic parameter and the second characteristic parameter of each website in the at least one website, where the first characteristic parameter is used to characterize the business matching degree of the website, and the second characteristic parameter is used to represent the activity degree of the website.
[0125] The collection module 620 is used to collect data from at least one target website based on at least one first keyword and at least one second keyword, the first keyword is used to collect data items including the first keyword in the data, and the second keyword is used to delete data items including the second keyword in the data.
[0126] The input and output module 630 is used to input data into the large model and output summary information of the data.
[0127] The generating module 640 is configured to generate a first report based on the summary information.
[0128] According to an embodiment of the present disclosure, the collection module collects data of at least one target website based on at least one first keyword and at least one second keyword, including:
[0129] For each target website in the at least one target website, the priority of each keyword in the at least one first keyword is determined respectively, and the priority is used to indicate the collection order of the keywords.
[0130] Based on the respective priorities of the target websites, data of the target websites are collected in the keyword collection order indicated by the priorities.
[0131] Based on the respective priorities of the target websites, tags are added to the data items containing the first keywords.
[0132] According to an embodiment of the present disclosure, the acquisition module determines the priority of each keyword in the at least one first keyword, including:
[0133] The priority of each keyword in the at least one first keyword is determined based on at least one of a long-tail keyword included in the at least one first keyword, a combination of uppercase and lowercase letters of the first keyword, homophones, typos, slang, and abbreviations.
[0134] According to an embodiment of the present disclosure, the data processing apparatus 600 based on the large model further includes:
[0135] The second determination module is configured to determine at least one first initial keyword and at least one second initial keyword for each target website based on first prior information, wherein the first prior information includes keywords used to determine historical abnormal data.
[0136] The analysis module is configured to determine at least one first keyword from the at least one first initial keyword and at least one second keyword from the at least one second initial keyword based on at least one of search traffic, search frequency, and keyword difficulty values of the at least one first initial keyword and the at least one second initial keyword.
[0137] According to an embodiment of the present disclosure, the generating module generates a first report based on the summary information, including:
[0138] The summary information is filtered based on second priori information, and abnormal data is determined from the public data based on the filtered summary information, where the second priori information includes historical abnormal data.
[0139] A first tag is added to the abnormal data, where the first tag is used to indicate at least one of the following: a derivative tool included in the target website where the abnormal data exists, a target website version, a model of a device used to log in to the target website, a vulnerability of the device, a tampered file, and a scenario where the abnormal data exists on the target website.
[0140] A first report is generated based on the first label and the anomaly data.
[0141] According to an embodiment of the present disclosure, the generating module generates a first report based on the summary information, further comprising:
[0142] According to the stage type of data processing, a first report with corresponding timeliness is generated based on data, summary information and abnormal data. The stage types include data collection stage, reporting stage, risk strategy formulation stage and summary stage.
[0143] According to an embodiment of the present disclosure, the generating module generates a first report based on the summary information, further comprising:
[0144] A second tag is added to the first report, where the second tag includes at least an identity tag of the target website to which the first report belongs, a content index tag of the abnormal data, and life cycle tags of the first report at different stages.
[0145] The first report is stored and edited based on the first tag and the second tag.
[0146] Figure 7 The structural block diagram of the large model of artificial intelligence according to the embodiment of the present disclosure is schematically shown.
[0147] In the embodiments of the present disclosure, Figure 7 As shown, the large model 700 may include five core modules: an input module 710 , a control module 720 , a storage module 730 , a calculation module 740 and an output module 750 .
[0148] Input module 710 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that can be understood and processed by large model 700. Input module 710 is the primary link for large model 700 to interact with the outside world. It enables large model 700 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0149] For example, let's take the example of applying a large model to data processing:
[0150] The input module 710 can input data collected from various target websites.
[0151] The control module 720 is the core support for the large model 700 to handle complex tasks. The control module 720 can perform summary information extraction based on the large model.
[0152] During operation, the control module 720 will continuously interact with the storage module 730, the computing module 740, and / or the output module 750. However, it should be noted that the control module 720 acts as a single initiator to initiate communications with the storage module 730, the computing module 740, and / or the output module 750, and there is no communication coupling between the storage module 730, the computing module 740, and the output module 750.
[0153] The performance of control module 720 may be closely related to the large model on which large model 700 is based. To fully utilize the capabilities of the large language model, the internal structure of control module 720 may be designed to be highly configurable and extensible to cope with various types of tasks and requirements in real scenarios.
[0154] The storage module 730 may be responsible for memorizing the information generated during the processing. The summary information as described above may be included in the storage module 730 .
[0155] After acquiring data from each target website, the large model 700 may use the summary extraction model to extract summary information of each data from the data from each target website. The summary information may be stored in the storage module 730. The large model 700 may then transmit the summary information to the output module 750.
[0156] The operation module 740 can be viewed as a predefined tool library.
[0157] When large model 700 needs to extract summaries from multiple data sets, it can call relevant tools from calculation module 740 and feed them back to control module 720. Control module 720 can then use the feedback tools to extract summary information and pass the extracted summary information to output module 750. It's understandable that while large language models possess excellent language understanding and generation capabilities, like humans, they can only perform limited tasks without the aid of tools. Once large model 700 is equipped with the ability to call tools, it can perform mathematical operations and data analysis, for example, using a calculator.
[0158] In an example, the output module 750 may output the summary information described above.
[0159] The large model 700 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0160] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0161] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0162] like Figure 8 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 can also store various programs and data required for the operation of electronic device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0163] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0164] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the agent publishing method. For example, in some embodiments, the agent publishing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the agent publishing method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the publishing method of the agent in any other appropriate manner (eg, by means of firmware).
[0165] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0166] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0167] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM) or flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0168] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) display or a liquid crystal display (LCD)) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0169] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0170] A computer system may include a terminal device and a server. The terminal device and the server are generally remote from each other and typically interact via a communication network. The relationship between the terminal device and the server is generated by computer programs running on the respective computers and having a terminal device-server relationship with each other.
[0171] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0172] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A data processing method based on a large model, comprising: determining at least one target website from the at least one website based on at least one of a first characteristic parameter and a second characteristic parameter of each website in the at least one website, wherein the first characteristic parameter is used to characterize the business matching degree of the website, and the second characteristic parameter is used to represent the activity degree of the website; collecting data from the at least one target website based on at least one first keyword and at least one second keyword, wherein the first keyword is used to collect data items including the first keyword in the data, and the second keyword is used to delete data items including the second keyword in the data; Input the data into a large model and output summary information of the data; as well as Generating a first report based on the summary information includes: filtering the summary information based on second prior information, and determining abnormal data from the data based on the filtered summary information, wherein the second prior information includes historical abnormal data; Adding a first tag to the abnormal data, the first tag being used to indicate at least one of a derivative tool included in a target website where the abnormal data exists, a version of the target website, a model of a device used to log in to the target website, a vulnerability of the device, a tampered file, and a scenario where the abnormal data exists on the target website; The first report is generated based on the first tag and the abnormal data.
2. The method according to claim 1, wherein The collecting data of the at least one target website based on the at least one first keyword and the at least one second keyword includes: For each target website in the at least one target website, determining a priority of each keyword in the at least one first keyword, wherein the priority is used to indicate a collection order of the keywords; Based on the priorities of the respective target websites, data of the target websites are collected according to the keyword collection order indicated by the priorities; Based on the respective priorities of the target websites, tags are added to the data items containing the first keywords in the data.
3. The method according to claim 2, wherein: The respectively determining the priority of each keyword in the at least one first keyword includes: The priority of each keyword in the at least one first keyword is determined based on at least one of a long-tail keyword included in the at least one first keyword, a combination of uppercase and lowercase letters, homophones, typos, slang, and abbreviations of the first keyword.
4. The method according to claim 1, further comprising: Determining at least one first initial keyword and at least one second initial keyword for each target website based on first prior information, wherein the first prior information includes keywords used to determine historical abnormal data; Based on at least one of search traffic, search frequency, and keyword difficulty values of the at least one first initial keyword and the at least one second initial keyword, at least one first keyword is determined from the at least one first initial keyword, and at least one second keyword is determined from the at least one second initial keyword.
5. The method according to claim 1, wherein The generating of the first report based on the summary information further includes: According to the stage type of data processing, a first report with corresponding timeliness is generated based on the data, the summary information and the abnormal data. The stage types include data collection stage, reporting stage, risk strategy formulation stage and summary stage.
6. The method according to claim 1 or 5, wherein: The generating of the first report based on the summary information further includes: Adding a second tag to the first report, the second tag including at least an identity tag of a target website to which the first report belongs, a content index tag of the abnormal data, and life cycle tags of the first report of different stages and types; The first report is stored and edited based on the first tag and the second tag.
7. A data processing device based on a large model, comprising: a first determining module configured to determine at least one target website from the at least one website based on at least one of a first characteristic parameter and a second characteristic parameter of each website in the at least one website, wherein the first characteristic parameter is used to represent the business matching degree of the website, and the second characteristic parameter is used to represent the activity degree of the website; a collection module, configured to collect data from the at least one target website based on at least one first keyword and at least one second keyword, wherein the first keyword is used to collect data items including the first keyword in the data, and the second keyword is used to delete data items including the second keyword in the data; An input and output module, used for inputting the data into the large model and outputting summary information of the data; as well as a generating module, configured to generate a first report based on the summary information, The generating module generates a first report based on the summary information, including: filtering the summary information based on second prior information, and determining abnormal data from the data based on the filtered summary information, wherein the second prior information includes historical abnormal data; Adding a first tag to the abnormal data, the first tag being used to indicate at least one of a derivative tool included in a target website where the abnormal data exists, a version of the target website, a model of a device used to log in to the target website, a vulnerability of the device, a tampered file, and a scenario where the abnormal data exists on the target website; The first report is generated based on the first tag and the abnormal data.
8. The apparatus according to claim 7, wherein the collecting module collects data of the at least one target website based on the at least one first keyword and the at least one second keyword, comprising: For each target website in the at least one target website, determining a priority of each keyword in the at least one first keyword, wherein the priority is used to indicate a collection order of the keywords; Based on the priorities of the respective target websites, data of the target websites are collected according to the keyword collection order indicated by the priorities; Based on the respective priorities of the target websites, tags are added to the data items containing the first keywords in the data.
9. The apparatus according to claim 8, wherein the acquisition module determines the priority of each keyword in the at least one first keyword, comprising: The priority of each keyword in the at least one first keyword is determined based on at least one of a long-tail keyword included in the at least one first keyword, a combination of uppercase and lowercase letters, homophones, typos, slang, and abbreviations of the first keyword.
10. The apparatus according to claim 7, further comprising: A second determination module is configured to determine at least one first initial keyword and at least one second initial keyword for each target website based on first prior information, wherein the first prior information includes keywords used to determine historical abnormal data; an analysis module for determining at least one first keyword from the at least one first initial keyword and at least one second keyword from the at least one second initial keyword based on at least one of search traffic, search frequency, and keyword difficulty values of the at least one first initial keyword and the at least one second initial keyword; 11. The apparatus according to claim 7, wherein the generating module generates a first report based on the summary information, further comprising: According to the stage type of data processing, a first report with corresponding timeliness is generated based on the data, the summary information and the abnormal data. The stage types include data collection stage, reporting stage, risk strategy formulation stage and summary stage.
12. The apparatus according to claim 7 or 11, wherein the generating module generates a first report based on the summary information, further comprising: Adding a second tag to the first report, the second tag including at least an identity tag of a target website to which the first report belongs, a content index tag of the abnormal data, and life cycle tags of the first report at different stages; The first report is stored and edited based on the first tag and the second tag.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Website defense method, device and equipment, and storage medium
CN110460620A
Data identification method and device, electronic equipment and storage medium
CN115879166A