Evaluation set construction method and device, equipment, storage medium and program product
By acquiring online public opinion data and performing semantic feature analysis and quantitative analysis, dynamic interception rules are generated, which solves the problems of lagging attack command mining and incomplete evaluation coverage in existing technologies, and realizes efficient security evaluation of generative large models.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-14
AI Technical Summary
Existing security assessment and defense systems suffer from problems such as delayed attack command mining, static defense strategies, and incomplete assessment coverage, making it difficult to effectively cope with complex attack methods of generative large models.
By acquiring online public opinion data, performing semantic feature analysis, mining target features, generating a set of target attack commands, and conducting quantitative analysis, interception rules are dynamically generated to achieve proactive mining and precise screening of attack commands.
It significantly shortens the time for discovering attack commands, improves the completeness and real-time performance of the evaluation set, covers various attack paths, dynamically adapts to the rapid iteration of attack methods, and provides a comprehensive and dynamically updated command library.
Smart Images

Figure CN121864407A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of information security and financial technology, and in particular to a method, apparatus, device, storage medium and program product for constructing an evaluation set. Background Technology
[0002] With the widespread application of generative large models in customer service, intelligent question answering, and automated report generation, the security threats they face are becoming increasingly complex. For example, in banking operations, attackers may construct specific instructions to induce the model to generate responses containing sensitive information, or exploit model vulnerabilities to conduct role-playing attacks. Such attacks may not only lead to data breaches and economic losses, but also trigger serious public opinion risks.
[0003] In existing technologies, security assessment systems struggle to detect new attack methods in a timely manner, allowing attackers to continuously exploit vulnerabilities before system fixes, creating a lag cycle of "attack-exposure-remediation." Furthermore, due to the stealth and diversity of attack commands, traditional manual review and static interception rules cannot cover all potential threats. Therefore, existing security assessment and defense systems suffer from problems such as delayed attack command discovery, static defense strategies, and incomplete assessment coverage. Summary of the Invention
[0004] This application provides a method, apparatus, device, storage medium, and program product for constructing an evaluation set, in order to solve the technical problems of existing security evaluation and defense systems, such as delayed attack instruction mining, static defense strategies, and incomplete evaluation coverage.
[0005] Firstly, this application provides a method for constructing an evaluation set, including:
[0006] Obtain online public opinion data;
[0007] Semantic feature analysis is performed on online public opinion data to mine target features, which are sentence features containing attack intent;
[0008] Based on target features, a set of target attack instructions is generated in batches through semantic expansion;
[0009] The set of target attack commands is quantitatively analyzed to obtain a score, and interception rules are generated based on the score.
[0010] Secondly, this application provides an evaluation set construction apparatus, comprising:
[0011] The acquisition module is used to acquire online public opinion data;
[0012] The mining module is used to perform semantic feature analysis on online public opinion data to mine target features, which are sentence features containing attack intent.
[0013] The generation module is used to generate a set of target attack instructions in batches based on target features and through semantic expansion;
[0014] The interception module is used to perform quantitative analysis on the set of target attack commands, obtain scoring results, and generate interception rules based on the scoring results.
[0015] Thirdly, this application provides an electronic device, including: a processor and a memory communicatively connected to the processor;
[0016] The memory stores the instructions that the computer executes;
[0017] The processor executes computer-executable instructions stored in memory to implement any of the methods of the first aspect.
[0018] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method of any one of the first aspects.
[0019] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method of any one of the first aspects.
[0020] The evaluation set construction method, apparatus, equipment, storage medium, and program products provided in this application acquire network public opinion data programmatically, replacing manual collection of attack cases and significantly shortening the time for discovering attack commands. Through joint processing of data mining and semantic expansion, the lag problem of existing technologies relying on manual extraction of attack features is solved. Data mining extracts statement features of attack intent through keyword matching and context analysis, avoiding misjudgments caused by single-keyword matching and overcoming the limitations of traditional keyword matching methods. The generated target commands cover various attack paths, ensuring that the evaluation set covers the attack patterns most likely to cause actual threats, significantly improving the completeness of the evaluation set. Risk levels are generated through quantitative analysis, dynamically adapting to the rapid iteration of attack methods, achieving proactive mining and precise screening of attack commands, providing a comprehensive and dynamically updated command library for subsequent evaluation and interception, significantly improving the real-time performance and coverage of large-scale model security evaluation. Attached Figure Description
[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0022] Figure 1 This is a diagram illustrating the attack methods.
[0023] Figure 2A flowchart illustrating a method for constructing an evaluation set as provided in an embodiment of this application;
[0024] Figure 3 A schematic diagram of the structure of an evaluation set construction device provided in an embodiment of this application;
[0025] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0026] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0027] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0029] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.
[0030] It should be noted that the evaluation set construction method, apparatus, equipment, storage medium and program product provided in this application can be used in the fields of information security and financial technology, as well as in any other field. The application fields of the evaluation set construction method, apparatus, equipment, storage medium and program product in this application are not limited.
[0031] The specific application scenario of this application is to address the security protection needs of generative large models in highly sensitive fields such as finance and healthcare. With the widespread application of generative large models in business scenarios, attack methods targeting these models are becoming increasingly complex. Figure 1 A diagram illustrating the attack methods, such as Figure 1 As shown, attack methods targeting large models can include, but are not limited to, the following three categories: 1) Target hijacking: inducing the model to output sensitive content; 2) Role-playing: impersonating a user to obtain permissions; 3) Reverse inducement: triggering the model to generate harmful information, etc. Attackers can exploit model vulnerabilities to cause public opinion crises, data leaks, or business interruptions. Taking the banking scenario as an example, large models are often used in intelligent customer service, document generation, and other processes. If attackers induce the model to output incorrect information through carefully designed attack commands, it may lead to damage to the bank's reputation and a significant increase in compliance risks. However, current large model attack methods are updated rapidly, and traditional manual monitoring and interception methods are difficult to cover new attack characteristics, resulting in a significant time lag between the occurrence of an attack and its repair. In addition, the output results of large models are highly dynamic, and attackers can bypass existing interception rules by modifying tiny commands, making the security protection system face the dual challenges of "lag" and "insufficient generalization".
[0032] Current technologies primarily rely on manual transmission and judgment of public opinion issues. For example, they involve manually analyzing abnormal events in social media, news reports, or internal logs to extract attack command characteristics and then manually adding them to the interception database. Therefore, the limitations of current technologies are: low automation (relying on manual discovery and processing, unable to respond to new types of attacks in real time); limited attack characteristic coverage (only able to intercept known attack patterns, unable to generalize to variant attacks); and rigid interception methods (fixed interception rules, unable to adapt to the rapid iteration of attack methods). Furthermore, current technologies lack mechanisms for the batch generation and evaluation of attack commands, making it difficult for the security protection system to form a closed loop of "discovery-verification-interception," and unable to effectively cope with the dynamic evolution of large-scale attack models.
[0033] The evaluation set construction method, apparatus, device, storage medium, and program product provided in this application acquire network public opinion data containing attack keywords, perform semantic feature analysis on the acquired text data, extract target features of attack commands, and expand and generate high-risk attack commands based on the extracted target features. The attack command set is then quantitatively analyzed to generate risk levels and dynamically generate interception rules, aiming to solve the aforementioned technical problems of existing technologies.
[0034] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0035] Figure 2 This is a flowchart illustrating a method for constructing an evaluation set, as provided in an embodiment of this application. Figure 2 As shown, the method includes:
[0036] S201. Obtain online public opinion data.
[0037] In one example, text containing keywords such as "command attack," "vulnerability," and "public opinion risk" can be used, for example, from social media, technology forums, and news reports.
[0038] S202. Perform semantic feature analysis on online public opinion data to mine target features.
[0039] In this application embodiment, the target feature is a statement feature containing an attack intent.
[0040] In one example, potential attack commands can be extracted from online public opinion data through keyword matching and semantic analysis. For instance, if the online public opinion data contains the phrase "command attack + vulnerability + bank A," the system extracts "vulnerability" and "bank A" as target features and combines this with click popularity (such as high-viewership pages) to determine its risk level. This can be achieved through text matching based on a pre-defined attack feature lexicon (such as "vulnerability" and "attack"). Natural language processing techniques are used to extract implicit semantic features from the online public opinion data. Click popularity can include the number of visits to the online public opinion data or user interaction data, used to measure the potential impact of the attack commands.
[0041] S203. Based on target features, generate a set of target attack instructions in batches through semantic expansion.
[0042] In one example, based on target features, a set of target attack instructions is generated in batches using semantic expansion methods (such as keyword replacement and sentence structure adjustment). For instance, if the initial instruction is "Induce the model to generate false results," the system replaces the keyword "false" with "simulate" and adjusts the sentence structure (such as step-by-step questioning) to generate a variant of the target attack instruction, "Please simulate the process and generate results." Exemplarily, keywords in the attack instructions are replaced with synonyms or near-synonyms (e.g., "vulnerability" → "defect"). Relevant attack instructions are generated based on contextual semantics. The expression of attack instructions is expanded through keyword replacement, utilizing semantic association to cover variant target attack instructions that attackers might employ.
[0043] In another example, a large model can be invoked to perform joint semantic encoding of the initial attack command and its context based on target features, extracting the association features between the attack intent and the target entity, and generating a set of target attack commands in batches. For example, the large model can learn the co-occurrence relationship between "vulnerability" and "bootstrapping" in the context through prompts to identify the attack intent.
[0044] S204. Perform quantitative analysis on the set of target attack commands to obtain a score result, and generate interception rules based on the score result.
[0045] In one example, the generated target attack commands are input into a large model to obtain the output results. The risk level of the output results is evaluated based on a scoring model, and interception rules are dynamically generated based on the evaluation results.
[0046] In another example, a large model can be invoked to filter the set of target attack commands through prompt word engineering and evaluate the risk level of the output results in order to generate corresponding interception rules based on the scoring results.
[0047] In one implementation scenario, a system for automatically constructing large-scale security evaluation sets can be used to build evaluation sets and generate interception rules. This system can include: 1. A data mining module: By summarizing existing problems, it identifies unified characteristics of command attack public opinion issues, such as combinations of terms like "command attack" + "victimized," "large-scale model" + "vulnerability" + organization name, or by combining information such as click popularity, to target and mine online public opinion data, obtaining data on online public opinion with potential risks.
[0048] 2. Command Regeneration Module: Based on the network public opinion data obtained through mining, it generates a batch of similar target attack commands by means of semantic expansion methods such as keyword replacement and command association.
[0049] 3. Batch Detection Module: Utilizes the generated set of target attack commands to input into a large model for batch detection, and obtains the output results of the large model.
[0050] 4. Evaluation Module: Based on the application scenario of the large model, a scoring model is constructed to score the output results and obtain the scoring results. Based on the scoring results, a set of target attack commands that meet the conditions is determined.
[0051] 5. Batch interception module: Generates batch interception rule combinations for a set of attack commands that meet the criteria.
[0052] 6. Manual Module: Based on the effectiveness of the combined blocking rules and the needs of the large model, selectively filter the combinations of blocking rules. The manual module is optional.
[0053] The evaluation set construction method provided in this embodiment acquires network public opinion data programmatically, replacing manual collection of attack cases and significantly shortening the time for discovering attack commands. Through joint processing of data mining and semantic expansion, it solves the problem of lag in existing technologies that rely on manual extraction of attack features. Data mining extracts sentence features of attack intent through keyword matching and context analysis, avoiding misjudgments caused by single-keyword matching and overcoming the limitations of traditional keyword matching methods. The generated target commands cover various attack paths, ensuring that the evaluation set covers the attack patterns most likely to pose actual threats, significantly improving the completeness of the evaluation set. By generating risk levels through quantitative analysis and dynamically adapting to the rapid iteration of attack methods, it achieves proactive mining and precise screening of attack commands, providing a comprehensive and dynamically updated command library for subsequent evaluation and interception, significantly improving the real-time performance and coverage of large-scale model security evaluation.
[0054] Optionally, semantic feature analysis can be performed on online public opinion data to mine target features, including: extracting keyword combinations containing attack intent from online public opinion data through keyword matching; determining the association between attack intent and target entity based on the contextual semantics of keyword combinations to obtain target features.
[0055] In one example, semantic analysis can be used to determine the correspondence between an attacker's intent (such as inducing the generation of false content) and the target (such as a specific organization name). Keyword matching extracts keyword combinations representing the attack intent from network data, and then contextual semantic analysis determines the association between the attack intent and the target entity. For example, if an attacker describes their attack using the phrase "vulnerability" + "Bank A" + "inducing the generation of false reports," extracting "vulnerability" and "Bank A" as keyword combinations, semantic analysis can determine that "inducing the generation of false reports" is an attack intent targeting Bank A.
[0056] By combining keyword matching with contextual semantic analysis, this approach solves the misjudgment problem caused by existing technologies relying solely on keyword matching. It significantly improves the accuracy of attack feature extraction, covers more covert attack patterns, reduces reliance on fixed keyword libraries, and enhances adaptability to new types of attacks.
[0057] Optionally, based on the contextual semantics of keyword combinations, the association between the attack intent and the target entity is determined, including: using a contextual embedding model to perform joint semantic encoding on the keyword combination and its contextual text to obtain the joint semantic encoding result; and extracting the association features between the attack intent and the target entity based on the joint semantic encoding result.
[0058] In one example, a context embedding model can be used to perform joint semantic encoding on attack instructions and their context contained in online public opinion data, extracting the association features between attack intent and target entities. For example, the model can identify the attacker's true purpose by learning the co-occurrence relationship between "vulnerability" and "generated content" in context. Specifically, the context embedding model is used to perform joint semantic encoding on text and its context through a pre-trained model. Joint semantic encoding includes a unified semantic vector representation of keyword combinations and their contextual text.
[0059] For example, jointly semantically encoding keyword combinations and their contextual text using a context embedding model can include: encoding "vulnerability" and "inducing the generation of false reports" into a unified vector using the context embedding model, and extracting the association features between the attack intent and the target entity through semantic similarity analysis. If the contextual text indicates that "vulnerability" and "Bank A" both point to "inducing the generation of false reports," the system can determine that the attack intent is directed at Bank A.
[0060] By employing joint semantic encoding through a contextual embedding model, this approach addresses the challenge of existing technologies in identifying metaphorical expressions or step-by-step induced attacks. It significantly improves the accuracy of attack feature extraction, covers more complex attack strategies, and enhances the generalization ability to novel attack patterns.
[0061] Optionally, based on target features, a set of target attack instructions is generated in batches through semantic expansion, including: based on target features, using a preset generator to generate target attack instructions in batches through semantic expansion; using a preset discriminator to verify the target attack instructions, and filtering to obtain a set of verified target attack instructions.
[0062] In one example, an adversarial generative network (GAN) can be introduced into the instruction regeneration module. A generator produces variants of attack instructions, and a discriminator verifies the aggressiveness of these variants. The generator generates variant target attack instructions based on the semantic features of the attack instructions (such as attack intent, target entity, etc.), and the discriminator verifies the aggressiveness of the variant target attack instructions using a pre-trained attack detection model. Through adversarial training, the generator and discriminator optimize the quality of the generated target attack instructions, ensuring that the variant target attack instructions are both aggressive and diverse.
[0063] By employing adversarial training between the generator and discriminator, the problem of insufficient variant coverage caused by traditional instruction re-engineering relying solely on keyword substitution is resolved. This significantly improves the completeness of the evaluation set, covering more potential attacker strategies and providing more comprehensive test samples for batch detection.
[0064] Optionally, a quantitative analysis is performed on the set of target attack commands to obtain a scoring result, including: multimodal feature fusion based on the attack intent, behavioral characteristics, and output features corresponding to the set of target attack commands to obtain multimodal features; and scoring of the multimodal features through ensemble learning to obtain a scoring result.
[0065] In one example, multimodal features include a comprehensive feature that combines attack intent, behavioral features (such as command input frequency), and output features (such as sensitive word density) corresponding to the set of target attack commands. By introducing multimodal feature fusion, combining attack intent, behavioral features (such as command input frequency and user history interactions), and output features (such as output length and the number of times sensitive words appear) corresponding to the set of target attack commands, a more comprehensive evaluation model can be constructed. For example, if a command frequently triggers the model to generate outputs containing keywords such as "vulnerability," the command is determined to be a high-risk command based on a combination of textual semantics and behavioral features.
[0066] For example, a multimodal fusion model is used to weight and fuse text features (including attack intent and semantic similarity), behavioral features (input frequency, user's historical attack records, etc.), and output features (sensitive word density, output length, etc.), and score them using an ensemble learning method. For instance, an attacker may frequently input low-risk commands (such as "query rule A"), but the model output may frequently contain sensitive words such as "vulnerability." In this case, the command is determined to be a high-risk command through multimodal feature fusion.
[0067] By employing multimodal fusion, textual, behavioral, and output features are combined, and risk scoring is performed through ensemble learning. This avoids false positives caused by single features and improves the detection capability for low-frequency but high-risk attacks. Furthermore, multimodal feature fusion can dynamically adapt to changes in attack patterns, significantly enhancing the robustness and generalization ability of the scoring model.
[0068] Optionally, generating interception rules based on the scoring results includes: generating regular expression rules based on the scoring results through online learning; and performing fuzzy matching based on the regular expression rules to generate interception rules.
[0069] In one example, a dynamic rule update mechanism is introduced to automatically adjust interception rules based on real-time attack data. For instance, when a certain type of attack command is detected to frequently appear in multiple scenarios, the keyword combination of that command is added to the rule base, and variant commands are intercepted using regular expression matching. An online learning algorithm is used to update the interception rule base in real time. By monitoring the distribution characteristics of attack commands (such as high-frequency keywords, attack intent clustering, etc.), regular expression rules are automatically generated, and the interception effect of the rules is verified through testing. These regular expression rules include rules that achieve fuzzy matching through pattern matching (such as wildcards).
[0070] By dynamically updating the interception rule base, the problem of reliance on manual rule base maintenance in existing technologies is solved. This significantly improves the real-time performance and adaptability of the interception rules while reducing the workload of manual rule base maintenance. Furthermore, by using fuzzy matching of regular expression rules, the problem of existing rule bases being unable to adapt to iterative attacker strategies is addressed.
[0071] Alternatively, the method may also include: optimizing the blocking rules based on manual screening.
[0072] In one example, a semi-automatic filtering mechanism is introduced. This mechanism performs initial screening of interception rules based on preset criteria (such as attack risk scoring thresholds and instruction complexity), and then optimizes the rules based on human feedback. For instance, the system marks instructions with scores higher than the threshold as high-risk, and humans only need to confirm the rule logic, rather than reviewing each instruction individually.
[0073] For example, a rule logic engine is used to automatically adjust the rule logic based on human feedback. For instance, if a human confirms that "A" is a variant of "B", the rule base for "A" will be automatically expanded to "A|B", and the target attack command of the variant will be matched using regular expressions.
[0074] By manually reviewing and optimizing the blocking rules, potential logical inconsistencies in dynamic rule bases were resolved. This significantly improved the interpretability of the blocking rules, making them easier for business personnel to understand and optimize.
[0075] Figure 3 A schematic diagram of an evaluation set construction device provided in an embodiment of this application is shown below. Figure 3 As shown, the evaluation set construction apparatus 30 provided in this embodiment includes:
[0076] Module 301 is used to acquire online public opinion data;
[0077] The mining module 302 is used to perform semantic feature analysis on online public opinion data and mine target features, which are sentence features containing attack intent.
[0078] The generation module 303 is used to generate a set of target attack instructions in batches based on target features and through semantic expansion;
[0079] The interception module 304 is used to perform quantitative analysis on the set of target attack commands, obtain a score result, and generate interception rules based on the score result.
[0080] In one possible implementation, the mining module 302 is specifically used to: extract keyword combinations containing attack intent from online public opinion data through keyword matching; and determine the association between the attack intent and the target entity based on the contextual semantics of the keyword combinations to obtain target features.
[0081] In one possible implementation, the mining module 302 is further specifically used to: employ a context embedding model to perform joint semantic encoding on the keyword combination and its context text to obtain the joint semantic encoding result; and extract the association features between the attack intent and the target entity based on the joint semantic encoding result.
[0082] In one possible implementation, the generation module 303 is specifically used to: generate target attack instructions in batches by semantic expansion based on target features using a preset generator; and verify the target attack instructions using a preset discriminator to obtain a set of verified target attack instructions.
[0083] In one possible implementation, the interception module 304 is specifically used to: perform multimodal feature fusion based on the output features corresponding to the set of attack intent, behavioral characteristics and target attack instructions to obtain multimodal features; and score the multimodal features through ensemble learning to obtain a scoring result.
[0084] In one possible implementation, the interception module 304 is further specifically used to: generate regular expression rules based on the scoring results through online learning; and perform fuzzy matching based on the regular expression rules to generate interception rules.
[0085] In one possible implementation, the evaluation set construction device is also specifically used to: optimize the interception rules based on manual screening.
[0086] The evaluation set construction apparatus provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0087] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 4 As shown, the electronic device 40 may include a memory 401 and a processor 402. Optionally, the electronic device may also include a transceiver 403, wherein the memory 401 and the processor 402 communicate with each other; for example, the memory 401, the processor 402 and the transceiver 403 may communicate via a communication bus 404, the memory 401 is used to store a computer program, and the processor 402 executes the computer program to implement the method of the above embodiments.
[0088] Optionally, the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps in the method embodiments disclosed in this application can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0089] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the methods in any of the above method embodiments.
[0090] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the methods in any of the above method embodiments.
[0091] All or part of the steps in the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a readable memory. When the program is executed, it performs the steps of the above method embodiments; and the aforementioned memory (storage medium) includes: read-only memory (ROM), RAM, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disk, and any combination thereof.
[0092] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0093] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0094] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0095] Obviously, those skilled in the art can make various modifications and variations to the embodiments of this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations to the embodiments of this application fall within the scope of the claims of this application and their equivalents, this application also intends to include these modifications and variations.
[0096] In this application, the term "comprising" and its variations can refer to non-limiting inclusion; the term "or" and its variations can refer to "and / or". The terms "first", "second", etc., in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this application, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0097] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0098] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0099] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.
[0100] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or in the form of software program modules.
[0101] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.
[0102] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0103] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0104] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0105] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A method for constructing an evaluation set, characterized in that, The method includes: Obtain online public opinion data; Semantic feature analysis is performed on the aforementioned online public opinion data to mine target features, which are sentence features containing attack intent; Based on the target characteristics, a set of target attack instructions is generated in batches through semantic expansion; The set of target attack commands is quantitatively analyzed to obtain a scoring result, and an interception rule is generated based on the scoring result.
2. The method according to claim 1, characterized in that, The semantic feature analysis of the online public opinion data to mine target features includes: By using keyword matching, keyword combinations containing attack intent are extracted from the online public opinion data; Based on the contextual semantics of the keyword combination, the association between the attack intent and the target entity is determined, and the target features are obtained.
3. The method according to claim 2, characterized in that, The determination of the association between the attack intent and the target entity based on the contextual semantics of the keyword combination includes: A context embedding model is used to perform joint semantic encoding on the keyword combination and its context text to obtain the joint semantic encoding result. Based on the joint semantic encoding results, the association features between the attack intent and the target entity are extracted.
4. The method according to claim 1, characterized in that, The process of generating a set of target attack instructions in batches based on the target features through semantic expansion includes: Based on the target characteristics, the target attack instructions are generated in batches using a preset generator through semantic expansion; The target attack commands are verified using a preset discriminator, and a set of verified target attack commands is obtained.
5. The method according to claim 1, characterized in that, The quantitative analysis of the set of target attack commands to obtain a scoring result includes: Multimodal features are obtained by performing multimodal feature fusion based on the output features corresponding to the set of attack intent, behavioral characteristics, and target attack instructions. The multimodal features are scored through ensemble learning to obtain the scoring results.
6. The method according to claim 1, characterized in that, The generation of interception rules based on the scoring results includes: Based on the scoring results, regular expression rules are generated through online learning; The interception rules are generated by performing fuzzy matching based on the regular expression rules.
7. The method according to claim 1, characterized in that, The method further includes: The interception rules are optimized based on manual screening.
8. An evaluation set construction apparatus, characterized in that, The device includes: The acquisition module is used to acquire online public opinion data; The mining module is used to perform semantic feature analysis on the online public opinion data and mine target features, wherein the target features are sentence features containing attack intent; A generation module is used to generate a set of target attack instructions in batches based on the target features through semantic expansion; The interception module is used to perform quantitative analysis on the set of target attack commands, obtain a scoring result, and generate interception rules based on the scoring result.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 7.
11. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 7.