Harmful cue word visual analysis system based on risk perception
By constructing a risk-aware visual analysis system for harmful prompt words, integrating security risk data and domain knowledge, the system enables security assessment and optimization of large language models. This solves the problems of lack of controllability and fine granularity in the generated results in existing technologies, and improves the quality and generation efficiency of harmful prompt word datasets.
Patent Information
- Application Number
- CN202510874491.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Existing domain-specific harmful keyword generation and visualization analysis methods are insufficient in terms of understandability, controllability, and feedback, making it difficult to comprehensively assess the security of large language models, and the generated results lack sufficient controllability and fine granularity.
Design a risk-aware-based visual analysis system for harmful warning words. Integrate security risk data and domain knowledge input, model-driven algorithm workflow, core knowledge graph operation process, and multi-model feedback. Through risk-aware visualization and expert intervention mechanisms, realize a closed loop of domain harmful warning word generation, evaluation, and optimization, and construct a high-quality domain-specific harmful warning word dataset.
It improves the efficiency and controllability of domain-specific harmful keyword synthesis, enhances the security of large language models, supports rapid iteration and refined improvement, and enables fine-grained expression and risk identification for diverse risk scenarios.
Smart Images

Figure CN120995040A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the cross field of artificial intelligence and visual analysis, and particularly relates to a harmful prompt word visual analysis system based on risk perception. BACKGROUND
[0002] With the rapid improvement of the capabilities of large language models such as ChatGPT and DeepSeek, their applications in various specific fields have also made remarkable progress. For example, BloombergGPT is a 500 billion parameter model that has been optimized through mixed field training and performs well and competes in financial field tasks. However, at the same time, these field-specific applications have also brought new security and ethical challenges: financial large language models may have user data leakage risks, and medical large language models may be misused to provide guidance on concealing medical errors or other illegal behaviors, which significantly restricts the widespread application of large language models. At present, the security challenges of large language models have received widespread attention, and many studies have been devoted to exploring effective solutions.
[0003] To alleviate this problem, researchers have proposed constructing harmful prompts to test the security of large language models. Constructing harmful prompts as an effective method to test the security of large language models can reveal the weaknesses of large language models and serve as a fine-tuning corpus. Although a variety of general harmful prompt datasets have been developed, they mainly focus on general security risks. However, field-specific large language models have unique vulnerabilities related to professional knowledge in specific fields: for example, in medical large language models, the risk of recommending excessive use of certain drugs is inherently dependent on pharmacological professional knowledge; similarly, in financial large language models, the risk of providing investment advice is closely related to the regulatory restrictions of specific jurisdictions.
[0004] In addition, in existing research, the generation of harmful prompts mainly relies on manual or semi-automatic processes, which are time-consuming and labor-intensive, making it difficult to support rapid iteration and limiting the scalability of the dataset. Moreover, existing datasets lack diversity and coverage, with some risk types being adequately represented while others are severely lacking, resulting in a lack of comprehensive evaluation of the security robustness of large language models in different application scenarios.
[0005] As a transformative tool for synthetic data generation, large language models can be used to create diverse and high-quality datasets, thereby addressing the data scarcity challenge and improving machine learning performance. However, a comprehensive evaluation of synthetic data quality and in-depth insights into effective generation strategies remain unresolved issues. Solving this problem requires building an efficient human-machine collaboration framework to facilitate data generation, evaluation, and iterative improvement. Two major challenges that still need to be addressed in this process are: (1) The implicit and fuzzy risk entity recognition in domain knowledge. The harmful entity in the domain is often implicitly or ambiguously embedded in the knowledge graph, and lacks clear risk indication. The definition and judgment standard of "harmful information" in different application fields are significantly different, and it is difficult to distinguish between normal entities and risk entities in semantics. Therefore, directly applying the general harmful prompt word dataset is not enough to cover the needs of each field, and a special method needs to be designed for different fields to mine and represent the implicitly or indirectly expressed risk entity, and to be included in the data synthesis process of high-risk concepts.
[0006] (2) Controllability and expert intervention optimization of harmful prompt word generation. The automatically generated domain harmful prompt word often lacks sufficient controllability, which is not conducive to comprehensive evaluation and target iteration. Although diversity, toxicity and semantic relevance and other automatic indicators can partially measure the generation quality, it is difficult to cover the domain details and corner cases, so the effective evaluation and fine improvement of the generation result highly depends on the domain experts. How to provide clear multi-level semantic insights to support experts to review the generated data, and efficiently integrate the expert feedback into the next round of data generation is a key and extremely challenging task to ensure the robustness and reliability of the harmful prompt word dataset.
[0007] In summary, the existing domain harmful prompt word and visualization analysis research still has significant deficiencies in "understandability, controllability, and feedback" three aspects, therefore, it is urgent to design a domain knowledge optimization prompt word generation task, integrate safety risk data and domain knowledge input, model-driven algorithm workflow, core knowledge graph operation process and multi-model feedback into a visual analysis system for safety enhanced data generation and knowledge graph management, realize the closed loop of domain harmful prompt word generation-evaluation-optimization through risk perception visualization and expert intervention mechanism. SUMMARY
[0008] In view of the above, the purpose of the present application is to provide a harmful prompt word visual analysis system based on risk perception, which constructs seven front-end function modules for integrating safety risk data and domain knowledge input, model-driven algorithm workflow, core knowledge graph operation process and knowledge structure visualization analysis, and constructs a back-end module for constructing a knowledge graph according to user configuration, vectorization embedding calculation and retrieval, realizes the identification of implicit risk concepts in domain knowledge and the controllability and iterative optimization of harmful prompt word generation, so as to construct a high-quality harmful prompt word dataset in a specific field.
[0009] To achieve the above invention purpose, the technical scheme provided by the present application is as follows: The harmful prompt word visual analysis system based on risk perception provided by the embodiment of the present application, the system comprises: A sidebar module provides security risk data and domain knowledge input, configuration generation engine import, and system settings and operations, and generates knowledge graph results through a topic distribution view and subsequent views to display the domain topic distribution; A knowledge graph exploration module includes a risk perception graph, a prompt word scatter plot, and a word flow view, which are used to display the knowledge graph structure of the domain topic and to cluster risky concepts across topics, and to filter out key risky concepts according to the importance of the knowledge graph structure and the toxicity index; A harmful prompt word view browsing module is used to present the harmful prompt word library constructed by the key risky concepts that the user focuses on, and to review the harmful prompt word library and iteratively optimize the harmful prompt word library; A strategy adjustment panel module provides a domain topic proportion pie chart and an index weight slider, which are used to adjust the sampling weight of the key risky concepts through the domain topic proportion pie chart, and to adjust the generation strategy of the harmful prompt words in combination with the index weight slider; An index trend view module includes a flow chart and a line chart, which are used to present the evaluation index of the selected harmful prompt words with the iteration effect of the synthetic data version, and to help the user focus on the performance fluctuation source of the harmful prompt words; An index monitoring view module includes a radar chart and a domain topic level column chart, which are used to quantitatively evaluate the iteration effect of the harmful prompt words by comparing the harmful prompt words of the current synthetic data version with the harmful prompt words of any historical data version; A version switching module provides switching of each version of data, supports switching to display the current, historical, and overall views, and updates each module.
[0010] Preferably, the security risk data and domain knowledge input in the sidebar module includes a harmful prompt word dataset, domain knowledge, and concepts related to the root node of the domain knowledge graph; The configuration generation engine in the sidebar module includes a synthesis model for generating prompt words based on security risk data and domain knowledge for data synthesis, a target model for security testing based on the generated prompt words, and an embedding model for converting security risk data and domain knowledge input into high-dimensional vector representation.
[0011] Preferably, the system settings in the sidebar module include providing domain settings, domain knowledge graph initialization, risk concept generation, data synthesis, and result export.
[0012] Further preferably, the node size of the domain knowledge graph is encoded as the importance of the concept in the entire knowledge graph; The node transparency is encoded as the coverage of the concept, which is used to evaluate the semantic coverage of the generated harmful prompt words to the risk concepts in the domain knowledge graph; The thickness and transparency of the edges encode the maximum coverage of the connected vertices, which is used to locate high-coverage risky concepts.
[0013] Preferably, the knowledge graph exploration module is displayed through a three-level hierarchical view, including a global overview view, an intermediate view, and a local view. The global overview view displays all nodes in an elliptical layout and arranges them from top to bottom according to the importance of the knowledge graph structure. The intermediate view highlights key nodes in the current selected domain theme to help users focus. The local view displays detailed information of the nodes and provides functions such as expansion, deletion, hiding, or switching to prompt word expansion deletion. The prompt word scatter plot encodes the toxicity of the prompt word through the horizontal coordinate axis and the importance of the corresponding node through the vertical coordinate axis. The domain theme grouping displays the prompt word in a vertical bar cluster. Clicking on the vertical bar cluster jumps to the corresponding theme view of the knowledge graph and supports the lasso tool to select the interested prompt word for display in the subsequent view and mark the interested prompt word as expansion data. The left side of the word flow view is used to display high-frequency word items sorted by the importance of the domain knowledge graph, and the right side is used to display high-frequency word items sorted by coverage. The connection lines between the two types of high-frequency word items and the difference in saturation of the connection lines highlight the overlap and difference to help users discover problems such as insufficient coverage of important concepts or excessive generation of secondary concepts. Meanwhile, the knowledge graph exploration module also provides a sub-risk graph button for switching the sub-risk graph. The sub-risk graph reveals concepts with common risks across themes in a radial layout and clusters risk concepts across themes into a sub-risk category.
[0014] Further preferably, the knowledge graph exploration module also includes a thumbnail and a coverage-based heat map mode. Switching between the thumbnail and the heat map mode helps users quickly locate the interested area.
[0015] Preferably, in the harmful prompt word view browsing module, reviewing the harmful prompt word library includes: The harmful prompt word library is displayed through a prompt word detail table. According to the number, prompt word text, domain theme, corresponding node name in the knowledge graph, importance and toxicity indicators of the corresponding node in the prompt word detail table, the harmful prompt word is sorted, keyword searched, edited and deleted, and high-quality harmful prompt words are marked as expansion data.
[0016] Preferably, in the strategy adjustment panel module, adjusting the sampling weight of the domain theme includes the weight of the prompt word toxicity, semantic relevance, and randomness.
[0017] Preferably, in the index trend view module, a flow chart is used to present the trend of each theme index as the data version evolves, a line chart is used to compare the absolute value difference of the index between different data versions, and the performance of different indexes can be observed by clicking the legend, so as to comprehensively observe the performance of data synthesis.
[0018] Preferably, in the index monitoring view module, a radar chart is used to compare the performance of harmful prompt words generated by the current data version with the performance of harmful prompt words generated by any historical data version through five-dimensional indexes, and linkage between the radar chart and the field theme level column chart is supported, so that the user can decide whether to terminate the iteration of harmful prompt words by observing the increase and decrease of each theme level in the field theme level column chart after clicking any dimensional index on the radar chart; wherein the five-dimensional indexes include toxicity, semantic relevance, diversity, attack success rate and theme coverage of the prompt words corresponding to each theme; the field theme level column chart uses solid columns to indicate the current version result, uses dashed lines to indicate the historical version result, uses a gray area to identify the decline range, and uses a shadow area to identify the improvement range.
[0019] Preferably, the system further comprises: a backend computing module, which comprises a field knowledge graph generation submodule and a harmful prompt word generation submodule, the field knowledge graph generation submodule is used to construct a field knowledge graph based on security risk data and field knowledge input, calculate the importance and toxicity score of the field knowledge graph structure for screening, and identify field theme risk entities by using a large language model and semantic embedding clustering to generate risk concepts; the harmful prompt word generation submodule generates field harmful prompt words based on a large language model, refines and refines the data version management and multi-dimensional index calculation, and interacts with the sidebar module, the knowledge graph exploration module, the harmful prompt word view browsing module, the strategy adjustment panel module, the index trend view module, the index monitoring view module and the version switching module for real-time updating.
[0020] Compared with the prior art, the present application has at least the following beneficial effects: (1) The present application is based on a general harmful prompt word data set and combined with knowledge graph guided field knowledge retrieval to realize fine-grained expression of diversified risk scenarios. In addition, an interactive graph visualization for risk perception is designed, which uses force-directed layout and theme clustering to display the semantic association of prompt words and entities, and supports field users to quickly identify corner cases and iterate and refine prompt words.
[0021] (2) By integrating automatic harmful prompt word synthesis and interactive analysis, the field harmful prompt word generation-evaluation-optimization closed loop is realized through risk perception visualization and expert intervention mechanism, the synthesis efficiency and controllability of field harmful prompt words are improved, and the safety of large language model is strengthened through safety testing. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0023] Figure 1 is an interface structure schematic diagram of a harmful prompt word visual analysis system based on risk perception provided by the embodiments of the present application.
[0024] Figure 2 is a side bar module schematic diagram provided by the embodiments.
[0025] Figure 3 is a knowledge graph exploration module schematic diagram provided by the embodiments.
[0026] Figure 4 is a harmful prompt word view browsing module schematic diagram provided by the embodiments.
[0027] Figure 5 is a strategy adjustment panel module schematic diagram provided by the embodiments.
[0028] Figure 6 is an index trend view module schematic diagram provided by the embodiments.
[0029] Figure 7 is an index monitoring view module schematic diagram provided by the embodiments.
[0030] Figure 8 is a version switching module schematic diagram provided by the embodiments. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the protection scope of the present application.
[0032] The embodiment of the application provides a harmful prompt word visual analysis system based on risk perception, which comprises a side bar module, a knowledge graph exploration module, a harmful prompt word view browsing module, a strategy adjustment panel module, an index trend view module, an index monitoring view module and a version switching module in the front end, and a back end calculation module. The system integrates security risk data and domain knowledge input, model-driven algorithm workflow, core knowledge graph operation process and knowledge structure visual analysis. User interaction operation can affect domain knowledge graph construction and optimization prompt word generation. Through front-end and back-end cooperation, a closed loop of harmful prompt word synthesis-evaluation-optimization in a specific field is realized, so that a high-quality harmful prompt word dataset in a specific field is constructed.
[0033] Figure 1 The embodiment of the application provides a harmful prompt word visual analysis system based on risk perception. As shown in Figure 1 The embodiment provides a harmful prompt word visual analysis system based on risk perception, which comprises a side bar module, a knowledge graph exploration module, a harmful prompt word view browsing module, a strategy adjustment panel module, an index trend view module, an index monitoring view module and a version switching module. In the interface of the visual analysis system, A is the side bar module (A1 is the data import submodule, A2 is the configuration generation engine submodule, A3 is the system setting submodule, and A4 is the theme distribution view), B is the knowledge graph exploration module (B1 is the knowledge graph overall view, B2 is the prompt word scatter plot, B3 is the word flow view), C is the harmful prompt word view browsing module, D is the strategy adjustment panel module (D1 is the field theme proportion pie chart, and D2 is the index weight slider), E is the index trend view module (E1 is the flow chart, and E2 is the line chart), F is the index monitoring view module (F1 is the radar chart, and F2 is the field theme level column chart), and G is the version switching module. The modules are described in detail below.
[0034] As shown in Figure 2 The data import submodule A1 in the side bar module A provides harmful prompt word datasets, domain knowledge and concepts related to the root node of the domain knowledge graph as data import. The configuration generation engine submodule A2 provides a synthesis model for synthesizing prompt words according to security risk data and domain knowledge, a target model for security testing according to generated prompt words, and an embedding model for converting security risk data and domain knowledge input into high-dimensional vector representation as model import system, and provides field setting, domain knowledge graph initialization, risk entity generation, data synthesis and result export in the system setting submodule A3. After the configuration set by the user is completed, the theme distribution view A4 displays the risk entity quantity distribution according to the field theme, guides the user to adjust the domain knowledge graph structure, and performs visual analysis through the subsequent view.
[0035] In Knowledge Graph Exploration Module B, such as Figure 3 As shown in the diagram, B1 is the overall overview of the knowledge graph, B2 is the sub-risk graph button, B3 is the scatter plot of the prompt words, and B4 is the word flow view. The overall knowledge graph overview B1 displays the global structure of the domain knowledge graph. The node size of the domain knowledge graph encodes the importance of a concept within the entire knowledge graph; the node transparency encodes the coverage of a concept, used to assess the semantic coverage of the generated harmful prompt words for risky concepts in the domain knowledge graph; and the edge thickness and transparency encode the maximum coverage of connected vertices, used to locate high-coverage risky concepts.
[0036] The knowledge graph exploration module is displayed through a three-level hierarchical view, including a global overview view, a middle view, and a local view. The global overview view displays all domain topics and nodes in an elliptical layout and is arranged from top to bottom according to the importance of the knowledge graph structure. The middle view is used to highlight key nodes in the currently selected domain topic to help users focus. The local view is used to display detailed node information and provides functions such as expanding, deleting, hiding, or switching to the prompt word expansion and deletion.
[0037] Clicking the sub-risk graph button B2 displays nodes related to risk concepts. The sub-risk graph reveals cross-topic risk concepts in a radial layout and clusters cross-topic risk concepts.
[0038] The cue word scatter plot B3 is used to encode the toxicity of cue words through the horizontal axis and the importance of the corresponding nodes through the vertical axis. Domain topic groups display cue words in vertical bar clusters. Clicking on a vertical bar cluster jumps to the topic view of the knowledge graph. It also supports the lasso tool to select cue words of interest for display in subsequent views and to mark cue words of interest as extended data.
[0039] In the word flow view B4, the left side displays high-frequency terms sorted by importance in the domain knowledge graph, while the right side displays high-frequency terms sorted by coverage. The overlap and differences are highlighted by the connecting lines between the two types of high-frequency terms and the difference in the saturation of the connecting lines, which helps users discover problems such as insufficient coverage of important concepts or excessive generation of secondary concepts.
[0040] In addition, the knowledge graph exploration module also includes thumbnail and coverage-based heatmap modes, which support cross-level navigation and coverage assessment by switching between thumbnail and heatmap modes, helping users quickly locate areas of interest.
[0041] In the harmful word warning view browsing module C, such as Figure 4As shown in FIG. 6, the focus of the user is presented with the harmful prompt word library of the risk concept construction, the harmful prompt word library is reviewed, the harmful prompt word library is iteratively optimized, and the selected prompt word in the prompt word scatter plot or the field knowledge graph is displayed through the prompt word detail table. According to the number, prompt word text, field theme, corresponding node name in the knowledge graph, importance and toxicity indicators of the corresponding node in the prompt word detail table, the harmful prompt word is sorted, keyword retrieval, edited and deleted, and high-quality harmful prompt words are marked as expansion data for subsequent synthesis iteration.
[0042] As shown in FIG. 7, the strategy adjustment panel module D provides a field theme proportion pie chart D1 and an index weight slider D2, which are used to adjust the sampling weight of different themes through the field theme proportion pie chart, and adjust the generation strategy of harmful prompt words in combination with the index weight slider; wherein the field theme proportion pie chart D1 adjusts the sampling proportion of each field theme, and the index weight slider D2 includes the correlation and toxicity index weight slider to jointly determine the weight of balancing correlation, toxicity and randomness when sampling, so as to support interactive strategy fine-tuning. Figure 5 As shown in FIG. 8, the index trend view module E includes a stream chart E1 and a line chart E2, the stream chart E1 is used to present the trend of the index of each field theme with the evolution of multi-version data in real time, and the line chart E2 is used to compare the absolute value difference of the index between different data versions. By clicking the legend to switch the index, the performance of different indexes can be observed, so as to comprehensively observe the performance of data synthesis, so as to help the user to identify the insufficient field theme and trace back the strategy adjustment.
[0043] Figure 6 As shown in FIG. 9, the index monitoring view module F includes a radar chart F1 and a field theme level column chart F2, which are used to quantitatively evaluate the iteration effect of harmful prompt words by comparing the harmful prompt words of the current synthesized data version with the harmful prompt words of any historical data version, wherein the radar chart F1 is used to compare the harmful prompt words of the current data version with the harmful prompt words of any historical data version in parallel through five-dimensional indexes: the toxicity, semantic correlation, diversity, attack success rate and theme coverage of the prompt words corresponding to each theme. And support radar chart and field theme level column chart linkage F2, by clicking any dimensional index on the radar chart, the increase and decrease of each field theme level in the field theme level column chart is viewed, the field theme level column chart is marked with solid columns to indicate the current version result, and the historical version result is marked with dashed lines. The gray area represents the decline amplitude, and the shadow area represents the improvement amplitude, which is used to assist the user to decide whether to terminate the iteration of harmful prompt words.
[0044] As shown in FIG. 10, the index monitoring view module F includes a radar chart F1 and a field theme level column chart F2, which are used to quantitatively evaluate the iteration effect of harmful prompt words by comparing the harmful prompt words of the current synthesized data version with the harmful prompt words of any historical data version, wherein the radar chart F1 is used to compare the harmful prompt words of the current data version with the harmful prompt words of any historical data version in parallel through five-dimensional indexes: the toxicity, semantic correlation, diversity, attack success rate and theme coverage of the prompt words corresponding to each theme. And support radar chart and field theme level column chart linkage F2, by clicking any dimensional index on the radar chart, the increase and decrease of each field theme level in the field theme level column chart is viewed, the field theme level column chart is marked with solid columns to indicate the current version result, and the historical version result is marked with dashed lines. The gray area represents the decline amplitude, and the shadow area represents the improvement amplitude, which is used to assist the user to decide whether to terminate the iteration of harmful prompt words. Figure 7 As shown in FIG. 11, the index monitoring view module F includes a radar chart F1 and a field theme level column chart F2, which are used to quantitatively evaluate the iteration effect of harmful prompt words by comparing the harmful prompt words of the current synthesized data version with the harmful prompt words of any historical data version, wherein the radar chart F1 is used to compare the harmful prompt words of the current data version with the harmful prompt words of any historical data version in parallel through five-dimensional indexes: the toxicity, semantic correlation, diversity, attack success rate and theme coverage of the prompt words corresponding to each theme. And support radar chart and field theme level column chart linkage F2, by clicking any dimensional index on the radar chart, the increase and decrease of each field theme level in the field theme level column chart is viewed, the field theme level column chart is marked with solid columns to indicate the current version result, and the historical version result is marked with dashed lines. The gray area represents the decline amplitude, and the shadow area represents the improvement amplitude, which is used to assist the user to decide whether to terminate the iteration of harmful prompt words.
[0045] Figure 8 As shown in the middle, the version switching module G is used for switching of each version data, and switching between the current view, the historical comparison view and the overview view, and real-time updating of the field knowledge graph coverage, the prompt word scatter plot and the word flow view and other module display contents.
[0046] In actual use, the user can start the field knowledge graph construction through the sidebar module, and after the configuration is completed, view the field knowledge graph construction result in the knowledge graph exploration module, display the risk concept coverage through the three-level view, present the harmful prompt word library in the harmful prompt word view browsing module through lasso selection of the risk concept, review the harmful prompt word library, and iteratively optimize the harmful prompt word library. The user can further check each index, adjust the index weight, and adjust the generation strategy of the harmful prompt word, to realize the field harmful prompt word generation-evaluation-optimization closed loop.
[0047] In addition, the system also includes a backend computing module including a field knowledge graph generation submodule and a harmful prompt word generation submodule, which constructs a knowledge graph based on a field root node and user configuration through SPARQL query and breadth-first strategy, calculates structural importance and toxicity score for screening, and uses a large language model and semantic embedding clustering to identify field theme risk entities to generate risk concepts. The harmful prompt word generation submodule constructs a field harmful prompt word vector database based on safety risk data and field knowledge and a retrieval enhancement generation system (RAG), combines cleaning and refining, data version management, retrieves general harmful prompt word context, and uses a fine-tuning synthesis model to generate a response, screens out low-quality prompt words after toxicity evaluation by the Perspective API and iteratively generates, and balances data performance through a dynamic seed selection strategy of relevance and toxicity weight.
[0048] The backend computing module performs specific field harmful prompt word generation based on a large language model during and after data synthesis, and performs cleaning and refining, data set version management and multi-dimensional index calculation, and interacts with the sidebar module, the knowledge graph exploration module, the harmful prompt word view browsing module, the strategy adjustment panel module, the index trend view module, the index monitoring view module and the version switching module to realize real-time updating and iterative control of the front-end interface.
[0049] Through the above module design and collaborative linkage, the system realizes the field safety enhancement and visual generation closed loop of “field harmful prompt word synthesis → risk identification → strategy regulation → iterative enhancement”, visualizes the implicit risk and establishes a real-time control link of an expert feedback direct generation engine for assisting the user in visual analysis and strategy fine-tuning of the entire process of harmful prompt word synthesis. The present application uses the above visual analysis system, which can effectively improve the synthesis efficiency and quality of the field-specific harmful prompt word data set.
[0050] The above detailed description of the specific embodiments of the present application has described the technical solutions and beneficial effects of the present application, and it should be understood that the above description is only the most preferred embodiment of the present application and is not intended to limit the present application. Any modifications, supplements and equivalent replacements made within the principle range of the present application shall be included in the protection range of the present application.
Claims
1. A risk perception based harmful cue word visual analysis system, characterized in that, The system comprises: a sidebar module providing security risk data and domain knowledge input, configuration generation engine import, and system settings and operation, generating knowledge graph results through topic distribution view and subsequent view to display domain topic distribution; a knowledge graph exploration module including risk perception graph, prompt word scatter plot, and word flow view, for displaying the knowledge graph structure of the domain topic and clustering risky concepts across topics, and screening key risky concepts according to the importance of the knowledge graph structure and the toxicity index; a harmful prompt word view browsing module for presenting the harmful prompt word library constructed by the key risky concepts focused by the user, and reviewing the harmful prompt word library to iteratively optimize the harmful prompt word library; a strategy adjustment panel module providing a domain topic proportion pie chart and an index weight slider, for adjusting the sampling weight of the key risky concepts through the domain topic proportion pie chart, and adjusting the generation strategy of the harmful prompt words in combination with the index weight slider; an index trend view module including a flow chart and a line chart, for presenting the evaluation index of the selected harmful prompt words with the iteration effect of the synthetic data version, and helping the user focus on the root cause of the performance fluctuation of the harmful prompt words; an index monitoring view module including a radar chart and a domain topic level column chart, for quantitatively evaluating the iteration effect of the harmful prompt words by comparing the harmful prompt words of the current synthetic data version with the harmful prompt words of any historical data version; a version switching module providing switching of each version of data, and supporting switching to display the current, historical, and overall views, and updating each module.
2. The risk perception based harmful cue word visual analysis system of claim 1, wherein, The security risk data and domain knowledge input in the sidebar module includes: harmful prompt word dataset, domain knowledge, and concepts related to the root node of the domain knowledge graph; The configuration generation engine in the sidebar module includes: a synthetic model for generating prompt words for data synthesis according to security risk data and domain knowledge, a target model for security testing according to the generated prompt words, and an embedding model for converting security risk data and domain knowledge input into high-dimensional vector representation.
3. The risk perception based harmful cue word visual analysis system of claim 2, wherein, The system settings in the sidebar module include: providing domain settings, domain knowledge graph initialization, risk concept generation, data synthesis, and result export.
4. The risk perception based harmful cue word visual analysis system of claim 3, wherein, The node size of the domain knowledge graph is encoded as the importance of the concept in the entire knowledge graph; The node transparency is encoded as the coverage of the concept, for evaluating the semantic coverage of the generated harmful prompt words to the risk concepts in the domain knowledge graph; The thickness and transparency of the edge are encoded as the maximum coverage of the connected vertex, for locating high-coverage risky concepts.
5. The risk perception based harmful cue word visual analysis system of claim 1, wherein, The knowledge graph exploration module displays through three-level hierarchical views, including global overview view, intermediate view, and local view, the global overview view displays all nodes in an elliptical layout and arranges them from top to bottom according to the importance of the knowledge graph structure, the intermediate view highlights the key nodes in the currently selected domain topic to help the user focus, and the local view displays detailed information of the nodes and provides functions of expansion, deletion, hiding, or switching to prompt word expansion and deletion; A prompt word scatter plot is used to encode the toxicity of the prompt word through the horizontal coordinate axis, and the importance of the corresponding node of the prompt word through the vertical coordinate axis. The domain topic group is displayed as a vertical bar cluster. Clicking on the vertical bar cluster jumps to the corresponding topic view of the knowledge graph. The lasso tool is used to select the interested prompt word in the subsequent view and mark the interested prompt word as extended data. The left side of the word flow view is used to display high-frequency word items sorted by the importance of the domain knowledge graph, and the right side is used to display high-frequency word items sorted by coverage. The connection lines between the two types of high-frequency word items and the difference in the saturation of the connection lines highlight the overlap and difference, so as to help users find problems such as insufficient coverage of important concepts or excessive generation of secondary concepts. Meanwhile, the knowledge graph exploration module also provides a sub-risk graph button. The sub-risk graph button is used to switch the sub-risk graph. The sub-risk graph discloses the concepts with common risks across themes in a radial layout and clusters the risk concepts across themes into a sub-risk category.
6. The risk perception based harmful cue word visual analysis system of claim 1, wherein, The knowledge graph exploration module also includes a thumbnail and a coverage-based heat map mode. By switching between the thumbnail and the heat map mode, the user can quickly locate the area of interest.
7. The risk perception based harmful cue word visual analysis system of claim 1, wherein, In the harmful prompt word view browsing module, reviewing the harmful prompt word library includes: The harmful prompt word library is displayed through a prompt word details table. According to the number, prompt word text, domain theme, corresponding node name in the knowledge graph, importance and toxicity indicators of the corresponding node in the prompt word details table, the harmful prompt words are sorted, keyword searched, edited and deleted. High-quality harmful prompt words are marked as extended data.
8. The risk perception based harmful cue word visual analysis system of claim 1, wherein, In the index trend view module, a stream graph is used to present the trend of each theme index as the data version evolves, and a line chart is used to compare the absolute value difference of the index between different data versions. By clicking on the legend, the user can observe the performance of different indicators, so as to comprehensively observe the performance of the data synthesis.
9. The risk perception based harmful cue word visual analysis system of claim 1, wherein, In the index monitoring view module, a radar chart is used to compare the performance of harmful prompt words generated by the current data version with the performance of harmful prompt words generated by any historical data version through five-dimensional indicators in parallel. The radar chart and the domain theme level column chart are linked. By clicking on any dimension indicator on the radar chart, the user can view the increase or decrease of each theme level in the domain theme level column chart, so as to help the user decide whether to terminate the iteration of harmful prompt words. The five-dimensional indicators include the toxicity, semantic relevance, diversity, attack success rate and theme coverage of the prompt words of each theme. The domain theme level column chart uses solid columns to represent the results of the current version and dashed lines to represent the results of the historical version. The gray area represents the decline range, and the shaded area represents the improvement range.
10. The risk perception based hazardous cue visual analysis system according to any one of claims 1 to 9, wherein, The system also includes: A backend computing module includes a domain knowledge graph generation submodule and a harmful prompt word generation submodule. The domain knowledge graph generation submodule is used to construct a domain knowledge graph based on safety risk data and domain knowledge input. The structure importance and toxicity score of the domain knowledge graph are calculated for screening. A large language model and semantic embedding clustering are used to identify domain theme risk entities to generate risk concepts. The harmful prompt word generation submodule generates domain harmful prompt words based on a large language model, combines cleaning and refining, data version management, and multi-dimensional index calculation, and interacts with and updates in real time with the sidebar module, the knowledge graph exploration module, the harmful prompt word view browsing module, the strategy adjustment panel module, the index trend view module, the index monitoring view module, and the version switching module.
Citation Information
Patent Citations
Multi-level visual analysis system for big language model red team drilling
CN119088951A
Large model target range construction method and system based on expert model and risk level-to-level management
CN119514505A
Harmful knowledge graph construction and harmful information identification method based on large language model
CN120179831A
System and method for providing an interactive visual learning environment for creation, presentation, sharing, organizing and analysis of knowledge on subject matter
US20210065569A1