Security event description automatic generation method based on semantic enhancement

By using a multi-dimensional quality scoring method and a gradient prompt word template design, combined with a large language model to generate security event description text, the problem of uncontrollable text quality in existing technologies is solved, the controllability and compliance of security event descriptions are achieved, and the efficiency of compliance verification and the accuracy of event attribution are improved.

CN122065829APending Publication Date: 2026-05-19INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
Filing Date
2025-12-17
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

The quality of security incident description texts generated by existing technologies is uncontrollable, and they lack the ability to understand the semantics of regulatory contexts, resulting in low efficiency of compliance verification and difficulty in incident attribution.

Method used

A multi-dimensional quality scoring method is adopted, which generates security incident description text by combining multiple prompt word templates with a large language model, and uses machine and human scoring systems to ensure text quality, including evaluation of the dimensions of factual completeness, logical consistency, terminology professionalism, conciseness and fluency.

Benefits of technology

It achieves controllability and compliance in the quality of security incident description text, ensuring that the generated text can directly serve the security compliance audit needs of the digital service field, and improves the efficiency of compliance verification and the accuracy of incident attribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122065829A_ABST
    Figure CN122065829A_ABST
Patent Text Reader

Abstract

The invention provides a semantic enhancement-based security event description automatic generation method, which belongs to the technical field of artificial intelligence, and comprises the following steps of: respectively inputting a combination of each cue word template in a plurality of cue word templates and application behavior data into a plurality of large language models to obtain security event description texts generated by each large language model; performing quality scoring on each security event description text; and outputting the security event description text with the highest quality score. According to the method, the security event description text with the best quality is selected from the security event description texts generated by the combination of the various prompt word templates and the large language model through the quality score, so that the quality of the security event description text can be ensured to be always good, and the controllability of the quality of the security event description text is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for automatically generating security event descriptions based on semantic enhancement. Background Technology

[0002] The massive growth of digital services has generated a large number of security incidents related to application behavior, and has also driven the implementation of relevant regulations. Security incidents are usually presented in the form of structured logs containing technical descriptions, while regulatory rules are requirements expressed in natural language using legal terminology. Therefore, it is necessary to parse the security logs that record the original application behavior and generate semantic textual descriptions in order to establish a direct and interpretable link between specific security incidents and abstract regulatory requirements.

[0003] Raw security incidents are typically structured logs or system alerts. This data needs to be converted into semantically rich, natural language text descriptions. These descriptions must fully encompass the incident context, attributes, and compliance-related characteristics, while avoiding information loss. Because Large Language Models (LLMs) possess powerful understanding and logical reasoning capabilities, they can be used to generate security incident description texts that retain the original information and have rich contextual semantics.

[0004] However, the large language model and prompt words used to generate security incident description text are often randomly selected by the user, resulting in uncontrollable quality of the generated security incident description text. Summary of the Invention

[0005] This invention provides an automatic generation method for security event descriptions based on semantic enhancement, which solves the technical problem of uncontrollable quality of security event description text generated by existing technologies.

[0006] This invention provides a method for automatically generating security event descriptions based on semantic enhancement, comprising: Each of the multiple prompt word templates and the combination of application behavior data is input into multiple large language models to obtain the security event description text generated by each of the large language models. Quality scores were assigned to the description texts of each security incident. Output the description text of the security event with the highest quality score.

[0007] According to the present invention, an automatic generation method for security event descriptions based on semantic enhancement is provided, wherein the quality scoring of each security event description text includes: The security incident description text is scored from multiple dimensions to obtain scores for each dimension; these multiple dimensions include several from the following dimensions: factual completeness, logical consistency, terminology professionalism, conciseness, and fluency. The quality score is obtained by weighted summation of the scores for each dimension.

[0008] According to the present invention, a method for automatically generating security event descriptions based on semantic enhancement is provided, wherein the multiple dimensions include the factual integrity dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Determine the total number of fields contained in the application behavior data and the number of behavior fields contained in the security event description text, respectively; The score for the fact completeness dimension is determined based on the number of behavioral fields and the total number of fields.

[0009] According to the present invention, a method for automatically generating security event descriptions based on semantic enhancement is provided, wherein the multiple dimensions include the logical consistency dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Obtain multiple expert scores for the security event description text in the logical consistency dimension; If the kappa coefficient of the multiple expert scores is greater than a preset threshold, then the multiple expert scores will be used as the score for the logical consistency dimension.

[0010] According to the present invention, a method for automatically generating security event descriptions based on semantic enhancement is provided, wherein the multiple dimensions include the terminology professionalism dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Obtain multiple expert scores for the security incident description text in the terminology expertise dimension; If the kappa coefficient of the multiple expert scores is greater than a preset threshold, then the multiple expert scores will be used as the score for the terminology expertise dimension.

[0011] According to the present invention, a method for automatically generating security event descriptions based on semantic enhancement is provided, wherein the multiple dimensions include the simplicity dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Determine the total number of valid characters, the total number of repeated words, and the total number of valid information words contained in the security event description text, respectively; The length compliance score of the security event description text is determined based on the total number of valid characters. The redundancy score of the security event description text is determined based on the total number of valid characters, the total number of repeated words, and the total number of valid information words. The score for the simplicity dimension is obtained by weighted summation of the length compliance score and the redundancy score.

[0012] The present invention also provides an automatic generation device for security event descriptions based on semantic enhancement, comprising: The input module is used to input the combination of each of the multiple prompt word templates and application behavior data into multiple large language models to obtain the security event description text generated by each of the large language models. The scoring module is used to score the quality of description texts for each security incident. The output module is used to output the description text of the security event with the highest quality score.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the automatic generation method for semantically enhanced security event descriptions as described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the semantically enhanced security event description automatic generation method as described above.

[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the semantically enhanced security event description automatic generation method as described above.

[0016] This invention provides an automatic generation method for security event descriptions based on semantic enhancement. By using quality scoring to select the highest quality security event description text from a combination of various prompt word templates and a large language model, the method can ensure that the quality of the security event description text is consistently good and achieve controllable quality of the security event description text. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating the automatic generation method for security event descriptions based on semantic enhancement provided by the present invention.

[0019] Figure 2 This is one of the schematic diagrams of the prompt word template provided by the present invention.

[0020] Figure 3 This is the second schematic diagram of the prompt word template provided by the present invention.

[0021] Figure 4 This is the third illustration of the prompt word template provided by the present invention.

[0022] Figure 5 This is the fourth illustration of the prompt word template provided by the present invention.

[0023] Figure 6 This is an example diagram illustrating the application behavior data and security event description text provided by the present invention.

[0024] Figure 7 This is a schematic diagram of the structure of the automatic generation device for security event descriptions based on semantic enhancement provided by the present invention.

[0025] Figure 8 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0027] First, a brief explanation of the technologies related to the automatic generation of security event descriptions based on semantic enhancement will be given.

[0028] Security incidents related to application behavior are typically logged as structured or semi-structured logs. Early log analysis relied on rule-based or clustering methods to extract templates, while recent research has leveraged classifiers for semantic understanding. Beyond parsing, neural text generation techniques based on structured data have made significant progress. Early research focused on general domains, such as generating biographies from Wikipedia infographics and weather forecasts from meteorological records. This demonstrates that encoder-decoder structures can generate fluent and information-rich text, driving the development of increasingly complex data-to-text conversion tasks. In recent years, this paradigm has been applied to system log analysis. Large language models are now directly used for log understanding and summarization.

[0029] One related technology improves the efficiency and accuracy of log analysis and report generation through AI-automated processing. Users upload log files through a data interface. The system first performs data cleaning, transformation, and normalization preprocessing on the logs, and then uses the GPT-4 model to perform in-depth analysis on the preprocessed data. Based on the training data and algorithm, the model extracts meaningful patterns, trends, and outliers, and then automatically generates text descriptions based on the analysis results.

[0030] While related technology one achieves network security log analysis through automated processing, its limitation lies in focusing solely on security event analysis and report output, lacking the ability to understand the semantics of regulatory context. This deficiency stems from design limitations in each stage of its technical solution: the data preprocessing stage only cleans, transforms, and normalizes the log data itself, without incorporating regulatory corpora or annotating the logs with regulatory-related fields, thus lacking the foundational support for regulatory semantic understanding. The GPT-4 model is only used to extract key information about security events (such as attack source IP and event type) and identify security trends. It has not undergone fine-tuning related to regulatory semantic mapping, nor has it constructed semantic association rules between "security event terms" and "regulatory terms." It can only parse the technical attributes of security events and cannot transform them into regulatory contextual expressions.

[0031] Related technology two leverages the collaboration of knowledge graphs and large language models to achieve intelligent log analysis, multimodal data integration, and personalized report generation. The technical solution involves: collecting log data, preprocessing it through cleaning and denoising, format standardization, and time-series generation; constructing and dynamically updating a domain knowledge graph; integrating external knowledge bases; and combining the knowledge graph with a large language model to parse log semantics, infer causal relationships, detect anomalies, and fuse multi-source information. Finally, a prompt is customized according to user needs to generate multi-format personalized reports, while simultaneously integrating multimodal data to generate a comprehensive report including charts and timelines.

[0032] The large language model used in related technology 2 relies solely on the System Prompt definition to generate rules, without domain-specific fine-tuning. This makes it difficult to accurately parse technical terms and fault-related logic, resulting in insufficient semantic parsing depth, omissions in key event extraction, and analytical biases. The multimodal data integration module lacks conflict judgment and weight allocation rules, making it unable to determine validity when faced with data contradictions (such as conflicts between normal sensor readings and fault log alarms), leading to ambiguous output results.

[0033] The following is combined Figures 1-8 This invention describes an automatic generation method for security event descriptions based on semantic enhancement.

[0034] Figure 1 This is a flowchart illustrating the automatic generation method for security event descriptions based on semantic enhancement provided by the present invention, as shown below. Figure 1 As shown, steps S1, S2 and S3 are included but are not limited to.

[0035] Step S1: Input the combination of each prompt word template and application behavior data from the multiple prompt word templates into multiple large language models to obtain the security event description text generated by each large language model.

[0036] Considering the varying impacts of different prompt word templates on LLM semantic enhancement, this invention progressively designs four prompt word templates from four dimensions: instruction explicitness, constraint refinement, example guidance, and thought chain decomposition, such as... Figures 2-5 As shown. Prompt1 is the simplest basic instruction, including only system prompts, input data, and output requirements. Prompt2 is an enhanced version with finer-grained constraints, including system prompts, input data, semantic enhancement requirements (such as factual retention, word limits, etc.), and output requirements. Prompt3 adds some few-sample examples to Prompt2, providing intuitive reference for LLM through "example-output". Prompt4 introduces the Chain of Thought (CoT) model, breaking down the enhancement process step by step into steps such as element extraction, logical organization, and language organization, explicitly guiding LLM to complete the semantic enhancement task. In addition, all Prompts use HTML. Labeling modules (original events, constraint rules, examples, thought processes, enhanced outputs) can ensure the parsability of the output and strengthen LLM's understanding of module boundaries.

[0037] Application behavior data can include Figure 6 The security log shown. Combining the application behavior data with the four prompt word templates yields four different input data combinations.

[0038] Considering that raw application behavior data (logs, alerts) is fragmented and difficult to analyze in context, and given that LLMs possess powerful understanding and logical reasoning capabilities, they can be used to generate security event description text that retains the original information and has rich contextual semantics. This invention primarily selects open-source text generation LLMs, including Qwen2.5, Llama3, Deepseek-R1, etc., with parameter sizes ranging from 7b to 13b. Guided by prompting engineering, the LLM can achieve semantic enhancement of the original behavior and parse the output HTML structured tags to effectively extract natural language description text that retains the original behavioral facts, i.e., security event description text, such as... Figure 6 As shown.

[0039] Leveraging the semantic enhancement capabilities of LLM, structured data recording original application behavior can be semantically parsed to generate textual descriptions of application behavior that conform to natural language logic. Ultimately, this enables a direct and interpretable association between specific security events and abstract regulatory requirements. It can be used for security compliance auditing in the digital service field, solving problems such as low compliance verification efficiency and difficulty in attributing security events caused by the semantic gap between the technical expression of security logs and the legal language of regulatory rules.

[0040] For the same LLM, each combination of input data generates a security event description text, so each LLM can generate 4 security event description texts.

[0041] Step S2: Quality score the description text of each security event.

[0042] The quality score of the security incident description text represents the quality of the security incident description text, and also represents the ability of the combination of the corresponding prompt word template and LLM to generate security incident description text.

[0043] Step S3: Output the security event description text with the highest quality score.

[0044] The highest quality score for the security incident description text indicates the best quality, and the stronger the ability of the combination of the prompt word template and LLM to generate security incident description text.

[0045] The specific process of steps S1-S3 is as follows: (1) Initialization ; (2) Construct a prompt word template library; (3) Prepare application behavior data (logs, alarms, etc.) for input and select prompt word templates in order from the prompt word template library; (4) Select LLM; (5) Input the application behavior data and prompt word template into the LLM to generate the security event description text T; (6) Calculate the quality score of T. ; (7) Determine the current quality score Is it greater than If so, then update. Otherwise proceed to the next step (8); (8) Determine whether the security event description text has been generated for all prompt word templates and all LLMs. If yes, proceed to the next step (9); otherwise, proceed to (3) to perform the next selection, generation and calculation. (9) Complete all calculations and record them. Corresponding security event description text Select The corresponding prompt word templates and LLMs are solidified into the optimal enhancement template; (10) Output security event description text .

[0046] As can be seen from the above, this invention selects the best quality security event description text from the security event description text generated by the combination of multiple prompt word templates and large language models through quality scoring. This ensures that the quality of the security event description text is always good and achieves controllable quality of the security event description text.

[0047] In one embodiment, step S2 of the present invention may further include: The security incident description text is scored from multiple dimensions, resulting in multiple scores for each dimension. These dimensions include several from the following dimensions: factual completeness, logical consistency, terminology professionalism, conciseness, and fluency. The scores for each dimension are weighted and summed to obtain the quality score.

[0048] The scoring of the factual completeness dimension, logical consistency dimension, terminology professionalism dimension, conciseness dimension, and fluency dimension can be completed by machines or humans.

[0049] For the scoring of the factual completeness dimension, since the input application behavior data structure is flexible, it is not possible to directly extract the structured information in the security event description text. Instead, the information in each field of the original application behavior data can be matched and compared with the information in the security event description text to assess whether the security event description text completely contains the content of the original application behavior data.

[0050] For the scoring of the logical consistency dimension, since application behavior data is mostly structured data and there is no direct contextual relationship between the data of each dimension, it is difficult for machines to directly judge whether the logic between application behavior data and security event description text is consistent. Moreover, it involves some legal and privacy security professional terms. Therefore, the scoring of the logical consistency dimension can be carried out manually.

[0051] The terminology professionalism dimension score is used to measure whether the use of domain terminology in security incident description text is standardized, to address whether the security incident description text conforms to industry expression habits and colloquial language (such as "invoking permissions" instead of "getting permissions"), and to ensure the professionalism and readability of the security incident description text.

[0052] The conciseness dimension score measures whether the security incident description text is free of redundancy and meets the length requirements, addressing the issue of whether the security incident description text is too long or too short. Tasks typically require the security incident description text to be concise (e.g., ≤100 words), avoiding repetitive and redundant expressions.

[0053] The fluency dimension score addresses the readability issue of security incident description text, meaning the text must be fluent and conform to human reading habits. Currently, LLM, after being pre-trained on massive amounts of data, possesses strong understanding and generation capabilities, resulting in mostly very fluent security incident description text. Therefore, fluency evaluation can be disregarded for now.

[0054] Scoring security incident description texts from multiple dimensions in this way yields objective and reliable quality scores, which helps ensure the highest quality output of security incident description texts.

[0055] In one embodiment, the multiple dimensions of the present invention may include a factual integrity dimension; The security incident description text is scored from multiple dimensions to obtain scores for each dimension, which may further include: Determine the total number of fields contained in the application behavior data and the number of behavior fields contained in the security event description text, respectively; The score for the factual completeness dimension is determined based on the number of behavioral fields and the total number of fields.

[0056] First, the application behavior data can be parsed to obtain all behavior fields. Then, each behavior field in the application behavior data is traversed, and direct matching and semantic matching are used to determine whether each behavior field in the application behavior data is covered by the security event description text. The number of behavior fields contained in the security event description text is then counted. Direct matching means using the behavior fields in the application behavior data to search for keywords in the fields in the security event description text. If no keyword is found, semantic matching is used, which involves segmenting the security event description text into sentences / phrases / words. A pre-trained multilingual semantic model (sentence-transformers / paraphrase-multilingual-MiniLM-L12-v2) is used to calculate the semantic similarity between the behavior fields in the application behavior data and the security event description text. If the semantic similarity is greater than or equal to a set threshold, the security event description text is determined to contain that behavior field.

[0057] The score for the factual integrity dimension can be determined based on the ratio of the number of behavioral fields to the total number of fields. For example, the score for the factual integrity dimension = (number of behavioral fields in the security event description text / total number of fields in the application behavioral data) * 5.

[0058] This allows the security event description text to be scored based on the total number of fields contained in the application behavior data and the number of behavior fields contained in the security event description text, from the perspective of factual completeness.

[0059] In one embodiment, the multiple dimensions of the present invention may include a logical consistency dimension; The security incident description text is scored from multiple dimensions to obtain scores for each dimension, which may further include: Obtain multiple expert scores for the logical consistency dimension of the security incident description text; If the kappa coefficient of multiple expert ratings is greater than a preset threshold, then the multiple expert ratings will be used as the score for the logical consistency dimension.

[0060] The security incident description text can be scored by multiple experts with over two years of experience in the security / privacy field, based on logical consistency. This assesses the logical consistency between the security incident description text and application behavior data, and then obtains the scores from each expert. Expert scores can range from 1 to 5 points, where 1-2 points indicate significant logical problems or incorrect interpretations in the security incident description text; 3 points indicate the security incident description text is basically acceptable; 4 points indicate the security incident description text is slightly lacking but generally correct; and 5 points indicate the security incident description text has no problems.

[0061] Next, the kappa coefficient of each expert's score is calculated. The kappa coefficient is a statistical indicator used to assess classification consistency. The larger the kappa coefficient, the more consistent the scores of each expert, and the more reliable the expert scores are. The preset threshold can be 0.7. A kappa coefficient greater than 0.7 can be considered as a reliable expert score and can be used as a score for the logical consistency dimension of the security event description text.

[0062] This allows us to determine the reliability of human scoring for the logical consistency dimension using the kappa coefficient, which helps us obtain reliable scores for the logical consistency dimension and thus improves the accuracy of quality scoring.

[0063] In one embodiment, the multiple dimensions of the present invention may include a terminology dimension; The security incident description text is scored from multiple dimensions to obtain scores for each dimension, which may further include: Obtain multiple expert scores for the terminology expertise dimension of the security incident description text; If the kappa coefficient of multiple expert ratings is greater than a preset threshold, then the multiple expert ratings will be used as the score for the terminology expertise dimension.

[0064] The scoring of the terminology expertise dimension requires the construction of a domain-specific glossary of standardized or incorrect terms, which is currently difficult to achieve through machine learning. Therefore, a manual evaluation method can be used. Multiple experts with over two years of experience in the security / privacy field can score the security incident description text on the terminology expertise dimension, and then obtain the scores from each expert. Expert scores can be assigned from 1 to 5 points, where 1-2 points indicate serious colloquialism issues in the security incident description text, 3 points indicate the text is basically acceptable, 4 points indicate the terminology is slightly lacking but the overall expression is good, and 5 points indicate no terminology issues.

[0065] Next, the kappa coefficient of each expert score is calculated. A kappa coefficient greater than 0.7 indicates that the expert score is reliable and can be used as the score of the security incident description text in the terminology professionalism dimension.

[0066] This allows us to use the kappa coefficient to determine the reliability of the manual scoring of the terminology specialization dimension, which helps to obtain reliable scores for the terminology specialization dimension and thus improves the accuracy of the quality scoring.

[0067] In one embodiment, the multiple dimensions of the present invention may include a simplicity dimension; The security incident description text is scored from multiple dimensions to obtain scores for each dimension, which may further include: Determine the total number of valid characters, the total number of repeated words, and the total number of valid information words in the security event description text; The length compliance score of the security event description text is determined based on the total number of valid characters. The redundancy score of the security event description text is determined based on the total number of valid characters, the total number of repeated words, and the total number of valid information words. The score for simplicity is obtained by weighting and summing the length compliance score and the redundancy score.

[0068] The simplicity score of this invention mainly includes two parts: length compliance score and redundancy score.

[0069] Regarding length compliance scores, it's understood that the fewer the total number of valid characters (character length after removing spaces) in the security incident description text, the more concise the description. Therefore, determining the length compliance score of the security incident description text based on the total number of valid characters can include: determining the length compliance score based on the principle that the length compliance score of the security incident description text is negatively correlated with the total number of valid characters.

[0070] This invention analyzes collected application behavior events and statistically finds that the description of each event can be clearly stated within 100 characters. Therefore, length compliance can be calculated based on the ≤100-character length constraint. ; in, The total number of valid characters contained in the security incident description text. The length penalty coefficient is set to 0.01, meaning 0.01 points are deducted for each extra character, and 0 points are awarded for exceeding 100 characters, in order to avoid extremely long texts. The range of values ​​is 1 indicates that the length is fully compliant, while 0 indicates that the length exceeds 100 characters, is not compliant, and has too much redundancy.

[0071] The redundancy score is determined based on the total number of valid characters, the total number of repeated words, and the total number of valid information words. This can be further explained by: determining the repeated word density based on the total number of repeated words and the total number of valid characters, and determining the information density based on the total number of valid information words and the total number of valid characters; and then weighting and summing the repeated word density and the information density to obtain the redundancy score.

[0072] The repetition density is the proportion of meaningless repeated words (such as "user user" or "application application") in the security event description text. The repetition density is calculated as: total number of repeated words / total number of valid characters. It can be used to match adjacent, frequently repeated domain-specific words using the regular expression `(user|application|data|permission)+\s*(\1)`.

[0073] Information density is the proportion of effective information words (including domain terminology and key dimension values) in the security incident description text. Information density = total number of effective information words / total number of effective characters. Effective information words are words extracted from the key values ​​of the original application behavior data (such as social apps and access permissions).

[0074] The formula for calculating redundancy score can be: ; If the density of repeated words is explicit redundancy, the weight is 60%; if the information density is implicit redundancy, the weight is 40%. The value range is [0,1], where 1 indicates no repetition and dense information, and 0 indicates all repeated words and no effective information.

[0075] Finally, the formula for calculating the score in the simplicity dimension can be: Considering that length compliance and redundancy are equally important, each is weighted 50%. The score range for the simplicity dimension is... 1 indicates that the length is compliant and there is no redundancy, while 0 indicates that it is extremely long or completely redundant.

[0076] This allows for a scoring of security incident description text based on its conciseness, considering the total number of valid characters, the total number of repeated words, and the total number of valid information words.

[0077] Finally, the scores for each dimension are weighted and summed to obtain the quality score. The weights for the factual completeness, logical consistency, terminology, conciseness, and fluency dimensions can be 0.3, 0.3, 0.15, 0.15, and 0.1, respectively.

[0078] This invention will collect application behavior data (including but not limited to various classic application behavior scenarios such as application permission calls and data leaks), and use methods such as... Figures 2-5 The four prompt word templates shown are used to generate security event description text on multiple LLMs. A self-designed quality scoring method is used to score the text, and the scores of each dimension are weighted to obtain the quality score of the security event description text. The prompt word template with the highest quality score is selected and combined with the LLM, and then solidified as the optimal enhanced template to ensure the stability of the quality of subsequent generation.

[0079] This invention innovatively proposes four graded prompt word templates from basic instructions to thought chain decomposition, covering four dimensions: instruction clarity, constraint refinement, example guidance, and thought chain decomposition. At the same time, it divides functional modules through HTML tags, which not only strengthens LLM's understanding of task boundaries but also ensures the parsability of output results, effectively solving the core problem of "how to fully utilize LLM's semantic enhancement capabilities".

[0080] To address the pain points of evaluation without datasets or reference texts, this invention constructs a five-dimensional evaluation system that combines machine and human expertise. Factual completeness is achieved through a two-layer matching strategy (direct matching + semantic similarity matching) for automatic machine evaluation. Logical consistency and terminology professionalism are evaluated by domain experts and their effectiveness is verified through the kappa coefficient. Simplicity is achieved through length compliance and redundancy quantification calculations, comprehensively ensuring the quality and credibility of the generated text.

[0081] This invention integrates multiple open-source LLMs and gradient prompt word templates. Through batch data testing, the combination of "LLM + prompt word template" with the best overall performance is selected and solidified into a template. This not only solves the problem of unstable generation quality of single model / single prompt word template, but also realizes efficient reuse of subsequent application behavior description generation, taking into account both flexibility and stability.

[0082] This invention overcomes the shortcomings of existing technologies in lacking semantic understanding of regulatory context. The generated text not only fully preserves the core field information of the original security log, but also strengthens the compliance association characteristics through semantic expansion, realizing the semantic connection between technical logs and legal regulatory rules, and can directly serve the security compliance audit needs of the digital service field.

[0083] This invention focuses on the core objective of "preserving original information and enriching contextual semantics," specifically addressing two key issues: insufficient design of prompt word templates leading to underutilization of LLM capabilities, and difficulty in quality assessment without datasets / reference text. It achieves greater accuracy and richness in generating security event description texts, significantly outperforming existing technologies. Specific effects are as follows: (1) Preserving the original information while achieving deep semantic enhancement of the context, fully covering all dimensions of event features. Existing background technologies have two major defects: first, they only focus on the analysis of the technical attributes of security events, losing the event background or compliance association; second, the degree of semantic enhancement is shallow, resulting in the text being unable to support subsequent compliance audits and event attribution. This invention achieves precise context expansion while preserving the original information by combining the three constraints of "forced element extraction + logical sorting + compliance association embedding" in the prompt word template design with the semantic reasoning capabilities of LLM. (2) The gradient prompt word template system solves the problem of LLM capability utilization and achieves accurate context expansion and improved semantic interpretability. This invention designs four types of prompt word templates from simple to complex and in a hierarchical manner (Prompt1 + basic instructions - Prompt2 + fine-grained constraints - Prompt3 + few-sample guidance - Prompt4 + thought chain decomposition), and encapsulates the module boundaries through HTML tags to accurately guide LLM to focus on core tasks.

[0084] (3) Overcoming the limitations of no dataset / no reference text, constructing a closed-loop quality assessment method to ensure the reliability and usability of the generated text. Existing background technologies generally lack effective quality assessment mechanisms, relying entirely on subjective human judgment or labeled datasets or reference texts. This completely fails in real-world scenarios without datasets or reference texts, resulting in uncontrollable quality of the generated text. This invention addresses the pain point of no reference by constructing a multi-dimensional assessment system of "automatic machine assessment + domain expert calibration," enabling accurate quality assessment even without reference text.

[0085] like Figure 7 As shown, the automatic generation device for security event descriptions based on semantic enhancement provided by the present invention includes, but is not limited to: The input module is used to input the combination of each prompt word template and application behavior data from multiple prompt word templates into multiple large language models to obtain the security event description text generated by each large language model. The scoring module is used to score the quality of description texts for each security incident. The output module is used to output the description text of the security event with the highest quality score.

[0086] It should be noted that the automatic generation device for security event descriptions based on semantic enhancement provided by the present invention can execute the automatic generation method for security event descriptions based on semantic enhancement described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0087] Figure 8 This is a schematic diagram of the electronic device provided by the present invention. The electronic device may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. The processor can call logical instructions in the memory to execute a semantically enhanced automatic generation method for security event descriptions. This method includes: inputting the combination of each prompt word template and application behavior data from multiple prompt word templates into multiple large language models to obtain security event description texts generated by each large language model; scoring the quality of each security event description text; and outputting the security event description text with the highest quality score.

[0088] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0089] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the automatic generation method for security event descriptions based on semantic enhancement provided in the above embodiments, the method including: inputting the combination of each prompt word template and application behavior data from multiple prompt word templates into multiple large language models respectively to obtain security event description texts generated by each large language model; performing quality scoring on each security event description text; and outputting the security event description text with the highest quality score.

[0090] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the semantically enhanced security event description automatic generation method provided in the above embodiments. The method includes: inputting the combination of each prompt word template and application behavior data from a plurality of prompt word templates into a plurality of large language models to obtain security event description texts generated by each large language model; performing quality scoring on each security event description text; and outputting the security event description text with the highest quality score.

[0091] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0092] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for automatically generating security event descriptions based on semantic enhancement, characterized in that, include: Each of the multiple prompt word templates and the combination of application behavior data is input into multiple large language models to obtain the security event description text generated by each of the large language models. Quality scores were assigned to the description texts of each security incident. Output the description text of the security event with the highest quality score.

2. The automatic generation method for security event descriptions based on semantic enhancement according to claim 1, characterized in that, The quality scoring of each security incident description text includes: The security incident description text is scored from multiple dimensions to obtain scores for each dimension; these multiple dimensions include several from the following dimensions: factual completeness, logical consistency, terminology professionalism, conciseness, and fluency. The quality score is obtained by weighted summation of the scores for each dimension.

3. The automatic generation method for security event descriptions based on semantic enhancement according to claim 2, characterized in that, The multiple dimensions include the factual integrity dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Determine the total number of fields contained in the application behavior data and the number of behavior fields contained in the security event description text, respectively; The score for the fact completeness dimension is determined based on the number of behavioral fields and the total number of fields.

4. The method for automatically generating security event descriptions based on semantic enhancement according to claim 2, characterized in that, The multiple dimensions include the logical consistency dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Obtain multiple expert scores for the security event description text in the logical consistency dimension; If the kappa coefficient of the multiple expert scores is greater than a preset threshold, then the multiple expert scores will be used as the score for the logical consistency dimension.

5. The method for automatically generating security event descriptions based on semantic enhancement according to claim 2, characterized in that, The multiple dimensions include the terminology professionalism dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Obtain multiple expert scores for the security incident description text in the terminology expertise dimension; If the kappa coefficient of the multiple expert scores is greater than a preset threshold, then the multiple expert scores will be used as the score for the terminology expertise dimension.

6. The method for automatically generating security event descriptions based on semantic enhancement according to claim 2, characterized in that, The multiple dimensions include the simplicity dimension; The security event description text is scored from multiple dimensions to obtain scores across these dimensions, including: Determine the total number of valid characters, the total number of repeated words, and the total number of valid information words contained in the security event description text, respectively; The length compliance score of the security event description text is determined based on the total number of valid characters. The redundancy score of the security event description text is determined based on the total number of valid characters, the total number of repeated words, and the total number of valid information words. The score for the simplicity dimension is obtained by weighted summation of the length compliance score and the redundancy score.

7. An automatic generation device for security event descriptions based on semantic enhancement, characterized in that, include: The input module is used to input the combination of each of the multiple prompt word templates and application behavior data into multiple large language models to obtain the security event description text generated by each of the large language models. The scoring module is used to score the quality of description texts for each security incident. The output module is used to output the description text of the security event with the highest quality score.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the automatic generation method for security event descriptions based on semantic enhancement as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the automatic generation method for security event descriptions based on semantic enhancement as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the automatic generation method for security event descriptions based on semantic enhancement as described in any one of claims 1 to 6.