Data security risk assessment method and system based on large model

By employing a large-scale model-based data security risk assessment method, and utilizing a pre-trained language model and a vertical large-scale model for data security risk assessment, an assessment item prompt is constructed for multi-level reasoning. This solves the problem of low efficiency in manual verification and achieves efficient, accurate, and traceable data security risk assessment.

CN121256285AActive Publication Date: 2026-01-02NAT IND INFORMATION SECURITY DEV RES CENT

Patent Information

Application Number
CN202511831601.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-01-02
Estimated Expiration
2045-12-08

AI Technical Summary

Technical Problem

Existing data security risk assessments rely on manual verification, which is inefficient, subjective, difficult to standardize, lacks traceability, and is prone to missing information.

Method used

A data security risk assessment method based on a large model is adopted. By using a pre-trained language model and a vertical large model for data security risk assessment, compliance and security risk analysis and assessment prompts are constructed. The assessment results are generated through multi-level reasoning, and an assessment report is generated.

Benefits of technology

It has achieved automated and intelligent data security risk assessment, improved assessment efficiency, reduced human error, and has efficient, accurate, and traceable assessment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121256285A_ABST
    Figure CN121256285A_ABST
Patent Text Reader

Abstract

The invention discloses a data security risk assessment method and system based on a large model, and relates to the field of artificial intelligence large models. The method comprises the following steps: firstly, preprocessing a benchmarking and checking material, and then encoding the preprocessed benchmarking and checking material and a standard specification into a unified semantic space by utilizing a pre-training language model to form a benchmarking and checking vector database; then, on the basis of standard specifications, a data security risk assessment vertical large model is utilized to construct a compliance and security risk analysis assessment item prompt, and a structured assessment problem library is formed after manual verification; and then respectively calling each evaluation item prompt in the structured evaluation problem library, carrying out semantic retrieval in the benchmarking check vector database to generate an enhanced prompt, and then carrying out multi-stage reasoning to obtain a data security risk evaluation result. Through the pre-training language model and the data security risk assessment vertical large model, automatic and intelligent data security risk assessment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence large models, in particular to a data security risk assessment method and system based on a large model. BACKGROUND

[0002] With the release of laws, regulations, policy documents, standard specifications, etc., enterprises are encouraged to carry out data security risk assessment to review the degree of data security risk and compliance gap to meet the requirements of data security protection and supervision. The data security risk assessment specification requires that the data objects be checked one by one in terms of basic security assessment, data life cycle security assessment, etc., and relevant supporting materials need to be provided as the basis. The existing data security risk assessment workflow usually relies on personnel interviews, data checking, manual verification, tool testing, etc. When carrying out data security risk assessment for each data item, it needs to be checked one by one against numerous sub-clauses. The assessment personnel need to accurately locate the corresponding content for evaluation from a large amount of data security risk assessment benchmarking check materials and give the corresponding evaluation results. Artificial screening of materials to be checked as supporting materials occupies a large part of the entire evaluation work, and information is easily missed. In addition, the manual evaluation process has low efficiency, strong subjectivity, is difficult to standardize, lacks traceability and transparency, etc. An intelligent evaluation method and system with explainable rule chain and automatic quantitative scoring capability are needed. SUMMARY

[0003] The purpose of the present application is to provide a data security risk assessment method and system based on a large model to realize automatic and intelligent data security risk assessment.

[0004] To achieve the above-mentioned purpose, the present application provides the following solutions.

[0005] In a first aspect, the present application provides a data security risk assessment method based on a large model. The data security risk assessment method based on a large model applies a data security risk assessment vertical large model. The data security risk assessment method based on a large model includes: Obtaining benchmarking check materials and standard specifications for data security risk assessment, and preprocessing the benchmarking check materials to obtain preprocessed benchmarking check materials; Encoding the preprocessed benchmarking check materials and the standard specifications into a unified semantic space using a pre-trained language model to form a benchmarking check vector database, and based on a data security risk assessment standard, using a data security risk assessment vertical large model to construct a compliance and security risk analysis evaluation item prompt, and after artificial verification, forming a structured evaluation question library; The compliance evaluation item prompt in the structured evaluation question library is called to perform semantic retrieval in the benchmark checking vector database to generate an enhanced prompt for compliance evaluation; The security risk analysis evaluation item prompt in the structured evaluation question library is called to perform semantic retrieval in the benchmark checking vector database to generate an enhanced prompt for security risk analysis; According to the enhanced prompt for compliance evaluation and the enhanced prompt for security risk analysis, multi-level reasoning is performed by using the data security risk assessment vertical large model to obtain a compliance evaluation record and result, a risk level determination process and conclusion as a data security risk assessment result; the multi-level reasoning includes compliance determination, risk source identification, hazard degree and occurrence probability determination; According to the data security risk assessment result, an evaluation report is generated by using an embedded template.

[0006] In a second aspect, the application provides a data security risk assessment system based on a large model, which applies the data security risk assessment method based on a large model described above, and includes a data loading and preprocessing module, a semantic coding module, a compliance evaluation engine module, a security risk analysis engine module, a multi-level reasoning module and a report generation module. The data loading and preprocessing module is configured to obtain benchmark checking materials and standard specifications for data security risk assessment, and preprocess the benchmark checking materials to obtain preprocessed benchmark checking materials. The semantic coding module is configured to code the preprocessed benchmark checking materials and the standard specifications to a unified semantic space by using a pre-trained language model to form a benchmark checking vector database, and construct compliance and security risk analysis evaluation item prompts by using a data security risk assessment vertical large model based on a data security risk assessment standard, and form a structured evaluation question library after artificial checking. The compliance evaluation engine module is configured to call a compliance evaluation item prompt in the structured evaluation question library to perform semantic retrieval in the benchmark checking vector database to generate an enhanced prompt for compliance evaluation. The security risk analysis engine module is configured to call a security risk analysis evaluation item prompt in the structured evaluation question library to perform semantic retrieval in the benchmark checking vector database to generate an enhanced prompt for security risk analysis. A multi-level reasoning module is configured to perform multi-level reasoning on the data security risk assessment vertical large model according to the enhanced prompt for compliance evaluation and the enhanced prompt for security risk analysis, and obtain a compliance evaluation record and result, a risk level determination process and conclusion as a data security risk assessment result; the multi-level reasoning includes compliance determination, risk source identification, hazard degree and possibility determination; A report generation module is configured to generate an evaluation report according to the data security risk assessment result by using an embedded template.

[0007] According to the specific embodiments provided in the application, the application has the following technical effects.

[0008] The application provides a data security risk assessment method and system based on a large model. First, the benchmark checking materials are preprocessed, and then the preprocessed benchmark checking materials and standard specifications are coded into a unified semantic space by using a pre-training language model to form a benchmark checking vector database. Then, based on the data security risk assessment standard, a data security risk assessment vertical large model is used to construct a compliance and security risk analysis evaluation item prompt, and after artificial verification, a structured evaluation question library is formed. The compliance and security risk analysis evaluation item prompt in the structured evaluation question library is called to perform semantic retrieval in the benchmark checking vector database to generate an enhanced prompt for compliance evaluation and an enhanced prompt for security risk analysis, and a data security risk assessment vertical large model is used for reasoning to obtain a data security risk assessment result. According to the data security risk assessment result, an evaluation report is generated by using an embedded template. The application realizes automatic and intelligent data security risk assessment by using a pre-training language model and a data security risk assessment vertical large model. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.

[0010] Figure 1 A flowchart of a data security risk assessment method based on a large model according to an embodiment of the application.

[0011] Figure 2 A principle diagram of a data security risk assessment method based on a large model according to an embodiment of the application.

[0012] Figure 3An implementation process schematic diagram of a data security risk assessment method based on a large model provided by an embodiment of the present application.

[0013] Figure 4 A structural schematic diagram of a data security risk assessment system based on a large model provided by an embodiment of the present application. DETAILED DESCRIPTION

[0014] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0015] The above purposes, features and advantages of the present application can be more apparent and easy to understand. The present application will be described in further detail below with reference to the drawings and specific embodiments.

[0016] With the rapid rise of artificial intelligence technology, especially large models, it has deeply penetrated into key fields such as industry, telecommunications, medical treatment, government affairs, finance, and bears core tasks such as data processing, intelligent decision-making, content generation. Under this background, the embodiments of the present application provide a data security risk assessment method and system based on a large model, which can deeply analyze the structured information and unstructured information in the data security risk assessment benchmarking materials (including system configuration screenshots, access log records, management system texts, etc.) by virtue of the powerful multi-modal semantic understanding and knowledge association ability, realize the automatic and intelligent identification of the corresponding relationship between the data security risk assessment benchmarking materials and the standard clauses, and accurately extract key elements as evidence. This data security risk assessment method and system based on a large model can automatically analyze the benchmarking materials of the enterprise to be evaluated, output the compliance judgment conclusion, and automatically evaluate the possibility level and security impact level of the data security event in combination with the standard, generate a data security risk assessment report. The data security risk assessment method that combines expert knowledge, rule-driven and large model reasoning advantage reduces human errors through standardized reasoning logic, and realizes data security compliance risk judgment and evaluation in an efficient, accurate and traceable manner, providing technical support for enterprises to meet compliance requirements in a timely manner.

[0017] In an exemplary embodiment, a large model-based data security risk assessment method is provided, which applies a data security risk assessment vertical large model. The large model-based data security risk assessment method of the present application is based on a data security risk assessment vertical large model of a target field (an exemplary target field is the industry and communications field, and hereinafter the industry and communications field is taken as an example for illustration). The data security risk assessment method of the present application takes the data security risk assessment benchmarking materials of an enterprise as input data, uses natural language processing technology to cut the text after data cleaning into multiple text segments according to semantic units, and encodes them into a unified semantic space through contrastive learning. For each item of industry standard clause, a structured template is designed, including clause description, logical judgment standard, and generation format specification, forming a compliance judgment chain. According to the semantic space and the compliance judgment chain, multi-level reasoning is performed to learn the corresponding relationship between the industry standard clauses and the data security risk assessment benchmarking files, and the learned corresponding relationship is converted into a natural language description. The compliance degree of each standard clause is judged, and the corresponding evaluation result is given. Finally, a data security risk assessment report is generated according to the data security risk assessment report template.

[0018] As shown in Figure 1 and Figure 2 The large model-based data security risk assessment method includes the following steps 101-106.

[0019] Step 101, obtaining benchmarking materials and standard specifications for data security risk assessment, and preprocessing the benchmarking materials to obtain preprocessed benchmarking materials.

[0020] Step 102, using a pre-trained language model to encode the preprocessed benchmarking materials and the standard specifications into a unified semantic space to form a benchmarking vector database, and based on the data security risk assessment standard, using a data security risk assessment vertical large model to construct a compliance and security risk analysis and evaluation item prompt, and after artificial verification, forming a structured evaluation question library.

[0021] Step 103, calling the compliance evaluation item prompt in the structured evaluation question library, and performing semantic retrieval in the benchmarking vector database to generate an enhanced prompt for compliance evaluation.

[0022] Step 104, calling the security risk analysis and evaluation item prompt in the structured evaluation question library, and performing semantic retrieval in the benchmarking vector database to generate an enhanced prompt for security risk analysis.

[0023] Step 105, according to the enhanced prompt for compliance assessment and the enhanced prompt for security risk analysis, using the data security risk assessment vertical large model to perform multi-level reasoning, obtaining compliance assessment records and results, risk level determination process and conclusion as the data security risk assessment results; the multi-level reasoning includes: compliance determination, risk source identification, hazard degree and occurrence probability determination.

[0024] Step 106, according to the data security risk assessment results, using the built-in template to generate an evaluation report.

[0025] In another exemplary embodiment, the data security risk assessment vertical large model of the present application is trained by the following steps 201-203.

[0026] Step 201, obtaining evaluation records and corresponding supporting materials in historical data security risk assessment projects, and performing standardization processing to construct a training set.

[0027] 2.1, data acquisition.

[0028] Comprehensively sort out the evaluation records and corresponding supporting materials in past data security risk assessment projects (covering data of enterprises in different industries and of different sizes), and convert the evaluation items, evaluation results, and evaluation records into a form that can be used for training. At the same time, convert the clauses and indicators in the data security risk assessment specification standard files, such as the "Industrial Field Data Security Risk Assessment Specification" and the "Telecommunications Field Data Security Risk Assessment Specification", into a form that can be used for training.

[0029] 2.2, data preprocessing.

[0030] Clean the obtained data, remove duplicate, incorrect, and incomplete data records, unify the data format, and ensure the accuracy, consistency, and usability of the data.

[0031] 2.3, data annotation.

[0032] According to the professional knowledge and standards of data security risk assessment, carefully annotate the cleaned data. Adopt entity annotation method to annotate key entities in standard clauses (such as "important data, general data, personal information, etc."); adopt relationship annotation method to annotate the correspondence between material content and standard clauses; adopt evidence annotation method to mark key evidence fragments in supporting materials. After data annotation is completed, the annotation content is summarized to form a model training data set. For the compliance assessment link, the annotation content includes evaluation category, evaluation subclass, evaluation item, evaluation record, and evaluation result, and for the security risk analysis link, the annotation content includes determination factor, factor content, determination item, result record, and probability level, providing clear supervision basis for model training.

[0033] During the annotation process, multi-person cross-annotation and auditing mechanisms are introduced to control the quality of the annotation results. Cohen's Kappa coefficient is used to calculate the consistency between annotators. For controversial annotation content, uncertain sampling is used to automatically screen controversial samples, and experts are organized to discuss and determine.

[0034] Step 2.4, data augmentation.

[0035] To expand the scale of the training data set and improve the generalization ability of the model, after annotation, data augmentation techniques in natural language processing are used to replace synonyms, transform sentence structures, add noise, etc. to the text data, generating more samples with similar semantics but different forms to the original data. At the same time, using domain ontology knowledge, the data is annotated for semantic expansion, enriching the semantic information of the data.

[0036] Finally, the processed data set is stored in a special data storage system, and data index is established to facilitate fast data reading during model training.

[0037] Step 202, pre-training the selected large model using unlabeled target domain data, so that the large model learns the general language patterns, semantic features, and potential knowledge related to data security in the target domain, and obtains the pre-trained large model.

[0038] Step 203, fine-tuning the pre-trained large model using the training set, so that the large model can understand the industry standards of the target domain and the semantic and semantic relationships in the text of the benchmark checking material, and obtain the trained large model as the data security risk assessment vertical large model.

[0039] In another exemplary embodiment, steps 202-203 above are the process of training and fine-tuning the data security risk assessment vertical large model, and the specific process is as follows.

[0040] First, pre-train using a large amount of unlabeled industry and information technology field data, so that the model learns the general language patterns, semantic features, and potential knowledge related to data security in the industry and information technology field.

[0041] Next, based on the training data set with clear evaluation items and supporting materials, fine-tuning is performed in a supervised learning manner to guide the model to accurately learn the semantic relationship between the data security risk assessment specification standard and the enterprise benchmarking materials. During fine-tuning, the learning rate, batch size and other hyperparameters are dynamically adjusted to ensure the performance balance of the model on the training set and the validation set, preventing overfitting or underfitting. Through fine-tuning training, the model can deeply understand the semantics of the data security risk assessment related texts in the industry field, including accurately interpreting each requirement in the evaluation standard and the actual situation expressed by the enterprise-provided verification materials. The model learns the complex semantic relationships between the evaluation standard and the verification materials, such as causal relationship, matching relationship, and inclusion relationship. For example, when given an evaluation standard for data security risk assessment, the model can accurately determine whether the corresponding institutional documents provided by the enterprise meet the standard, and identify the gaps and potential risks.

[0042] According to the established fine-tuning strategy, the training data set is divided into training set, validation set and test set in a certain proportion. During training, the training set is input into the selected general large model, and the model parameters are continuously adjusted through the backpropagation algorithm to make the model's prediction results as close as possible to the labeled data. After each training cycle, the model is evaluated using the validation set to monitor the performance indicators (accuracy, recall, F1 value) of the model. Based on the evaluation results, the fine-tuning strategy is dynamically adjusted, such as adjusting the learning rate, replacing the optimizer, etc., to ensure the continuous improvement of the model's performance on the validation set. When the model's performance on the validation set reaches a stable and satisfactory state, the test set is used for final testing of the model to evaluate the model's generalization ability.

[0043] Incremental fine-tuning is used to freeze the bottom layer of the general semantic layer of the large model and only adjust the upper layer network related to the evaluation task. Through contrastive learning, the model learns the semantic mapping relationship between industry standards and verification materials, and designs clause identification and matching special training tasks to improve the model's related capabilities.

[0044] During training, the model performance is monitored in real time, and the fine-tuning effect is optimized by adjusting the learning rate and increasing difficult example samples, and finally a special model adapted to the data security risk assessment scene is output.

[0045] The comparative test includes comparison of different fine-tuning strategies, influence of different training data models, and effect comparison of different learning rates and batch sizes. Training is stopped when the performance of the validation set no longer improves, and the parameter fine-tuning method is used to reduce the training parameters.

[0046] A series of training tasks targeting semantic understanding and relationship learning are designed, such as text matching tasks, semantic judgment tasks, risk assessment case analysis tasks, etc. During the training process, a large number of diversified training samples are provided to the model, including positive samples (cases that meet the evaluation criteria) and negative samples (cases that do not meet the evaluation criteria), so that the model can understand the semantic matching relationship and difference characteristics between the evaluation criteria and the verification materials through comparative learning. Using attention mechanisms, graph neural networks, and other technologies, the model can better capture key information in the text and the semantic associations between information, improving the model's understanding and processing capabilities for complex semantic relationships. From known risk case data, common data security risk features are manually extracted and summarized, and these features are converted into feature vectors or patterns that can be used for model training. Through supervised learning, the model learns the mapping relationship between these risk features and risk types, risk levels. Regularly update the risk feature library and include the features in the newly emerging data security risk cases in the training, so that the model can adapt to the changing data security risk situation in a timely manner.

[0047] In another exemplary embodiment, after training the above-mentioned data security risk assessment vertical large model, the process of using the vertical large model for data security assessment is as shown in Figure 3

[0048] In another exemplary embodiment, as shown in Figure 3 The above-mentioned preprocessing process of the aforementioned benchmark verification materials in step 101 can be replaced by the following steps 301-306.

[0049] Step 301, converting different format files in the benchmark verification materials into machine-processable text to obtain the converted benchmark verification materials.

[0050] Step 302, using a sentence splitter to cut the machine-processable text in the converted benchmark verification materials into multiple text segments according to semantic units.

[0051] Step 303, using natural language processing technology to segment each text segment, and using a word vector model to vectorize the words in each text segment to obtain a vectorized representation of each text segment.

[0052] Step 304, extracting the chapter titles and paragraph themes of each text segment to generate a tree structure representation of each text segment.

[0053] Step 305, extracting key information from each text segment; the key information includes: mandatory statements, prohibited statements, dates, risk levels, and evaluation indicators.

[0054] ​Step 306, the vectorized representation, tree structure representation, and key information of each text segment constitute the preprocessed benchmarking materials.

[0055] In another exemplary embodiment, the above-mentioned step 301 can be implemented by the following process.

[0056] 3.1. Convert the benchmarking materials for data security risk assessment (including system configuration screenshots, access log records, management system texts, etc.) in different formats into machine-processable text.

[0057] 3.2. For database tables, JSON format files, remove irrelevant symbols, white spaces, redundant information, and special format processing, and perform data cleaning.

[0058] In another exemplary embodiment, in the above-mentioned step 302, for text, picture, etc. files, high-precision OCR technology is used for text content extraction. Semi-structured or unstructured information is extracted and organized into a more structured form, and the processed data is cut into multiple text segments by semantic units using a sentence segmenter, improving retrieval efficiency and accuracy.

[0059] For example, in this embodiment, by batch loading the benchmarking materials for enterprise data security risk assessment, including system configuration screenshots, access log records, management system texts, etc. Data cleaning techniques are used to remove duplicate, invalid or format error data. The processed data is cut into multiple text segments by semantic units using a sentence segmenter, and the structured data is uniformly formatted (such as converting different formats of dates and numbers to standard formats). The processed data is classified and stored in the corresponding data pool, providing high-quality basic data for subsequent modules.

[0060] A1. Data loading strategy Load the benchmarking materials for enterprise data security risk assessment. Support structured data (CSV, Excel, SQL database export file), semi-structured data (XML, JSON, Markdown format file), and unstructured data (support PDF, Word, PPT, etc. Office document format).

[0061] Based on file path, metadata label, preliminary classification, and construction of document indexing system, support fast retrieval based on file name, creation time, etc. Metadata.

[0062] A2. Text content extraction and cleaning Support multiple document format content extraction, text content extraction and cleaning to extract text from original documents and remove noise, remove irrelevant information such as headers and footers, page numbers, watermarks, etc.; process nested structures, which require to be sorted out; identify and extract text information in tables, lists, and charts.

[0063] Data cleaning process, remove HTML tags, special characters, garbled characters and other noise data; process abbreviations, acronyms; standardize units of measurement, date formats, etc.

[0064] Text segmentation and reorganization, based on punctuation marks, chapter titles, etc. for text segmentation; merge the contents of paragraphs interrupted by pagination; identify and extract key information in the document.

[0065] A3, natural language processing technology Natural language processing technology is responsible for basic processing of text to prepare for subsequent analysis. Use jieba, spaCy and other tools for Chinese and English word segmentation, remove meaningless words, and unify different words expressing the same concept.

[0066] Text vectorization uses pre-trained models to convert text into vector representation, uses domain-adapted word vector models for industry domains, and builds a text similarity calculation mechanism to support subsequent semantic retrieval.

[0067] A4, document structure analysis and feature extraction Document structure analysis and feature extraction are responsible for analyzing the structure information of the document and extracting the key information. Adopt hierarchical structure to identify the chapter structure and clause relationship of the document; build a tree structure representation of the document to extract chapter titles, paragraph topics, etc.

[0068] Key information extraction, based on template matching to extract standard specifications, date and other data information; extract structured information such as risk level and evaluation index; identify required statements and prohibited statements in the standard specification.

[0069] In another exemplary embodiment, as shown in Figure 3 The process of using a pre-trained language model to encode the preprocessed benchmark checking materials and the standard specification into a unified semantic space in step 102 above to form a benchmark checking vector database can be replaced by steps 401-403 as follows.

[0070] Step 401, use the data security risk assessment vertical large model to locate the terms in the preprocessed benchmark checking materials and the standard specification, and generate context-aware vector representation of each term.

[0071] Step 402, analyze the context-aware vector representation of each term using the attention mechanism to obtain the correlation strength between different terms.

[0072] Step 403, according to the correlation strength between different terms, encode the preprocessed benchmark checking materials and the standard specification into a unified semantic space, and construct a benchmark checking vector database.

[0073] In another exemplary embodiment, in steps 401-403 described above, a knowledge base is constructed by screening and labeling after training data set, including standard clauses, compliance descriptions corresponding to clauses, supporting materials corresponding to clauses, etc. Based on the knowledge base, a pre-trained language model is used for learning, and industry standards and benchmark checking materials are respectively input into two vector spaces. Through comparative learning, the model masters the semantic correlation of the two, and then encodes them into a unified semantic space. And use attention mechanism to analyze the difference characteristics of terms in the text, generate term mapping rule library containing corresponding relationship and logic, and establish vector index library to store text semantic vector. First, by accurately positioning the relevant terms in the text, the exact position of each term in the text sequence is recorded, and the complete text is input into the data security risk assessment vertical large model to generate the context-aware vector representation of each token. The model automatically generates a multi-layer, multi-head attention weight matrix during calculation, capturing the correlation strength between words. Second, extract the attention distribution pattern of terms at different levels of the data security risk assessment vertical large model, calculate the attention distribution difference of different terms at each layer, identify the attention receiving range of the term, and quantify the degree of dependence of the term on the context. The difference characteristics are extracted by deep convolution. Finally, set up a feedback loop to trigger the attention mechanism to retrain. In the retrieval link, after the query question is encoded, the semantic similarity is obtained by calculating the included angle cosine value of the query vector and the text semantic vector in the vector index library, the most relevant result is returned, and new data is introduced periodically for adversarial training to optimize the semantic matching effect.

[0074] When specific knowledge is needed later, the semantic similarity search can be performed in the knowledge base according to the query to accurately find the most relevant text segment, rather than just keyword matching.

[0075] In this embodiment, based on the data security risk assessment vertical large model and the pre-processed benchmarking data, the terminology gap between industry standards and data security risk assessment benchmarking materials is focused on solving. By building a dynamic knowledge base and using a pre-trained language model for contrastive learning, the industry standard clauses and enterprise data security risk assessment benchmarking materials are encoded into a unified semantic space. The attention mechanism is used to dynamically identify the differences in terminology and generate an interpretable terminology mapping rule base. By calculating the semantic similarity between the query question and the text segments in the vector index library, the most relevant text content is accurately fed back for subsequent structured template construction. Combined with adversarial training to reduce the difference in field distribution, high-precision semantic matching between industry standards and enterprise practices is achieved.

[0076] B1. Pre-training language model design The pre-training language model is mainly used to unify the semantic space of industry standards and enterprise data security risk assessment benchmarking materials. The model architecture design is as follows: the query tower mainly processes the risk assessment questions in the industry standard; the document tower processes the knowledge base content such as risk assessment cases; the shared basic model parameters are only differentiated in the output layer.

[0077] Contrastive learning training: the positive sample pair is extracted from related paragraphs in the same document; the negative sample pair is randomly selected for negative sample pair generation training; the InfoNCE loss function is used to maximize the similarity of the positive sample pair and minimize the similarity of the negative sample pair.

[0078] B2. Knowledge base construction and indexing The knowledge base construction and indexing are responsible for constructing the data security risk assessment knowledge base and creating an efficient indexing structure. The knowledge source processing extracts the clause requirements and protection measures from industry standards such as “Industrial Domain Data Security Risk Assessment Specification” and “Telecommunications Domain Data Security Risk Assessment Specification”; extracts information from enterprise data security risk assessment benchmarking materials to provide logical ideas for subsequent benchmarking materials.

[0079] Knowledge representation and storage use a graph database to store entity relationships, and store the converted text paragraphs as vectors in a vector database. Combined with inverted indexes and vectors, it supports hybrid semantic and keyword retrieval.

[0080] Knowledge update mechanism only processes the changed part when new documents are added, records the update history of the knowledge base, supports rollback and comparison, and ensures the accuracy of new knowledge through expert review or automatic verification mechanism.

[0081] B3. Attention mechanism enhanced retrieval Attention mechanism improves the accuracy of semantic retrieval. Hierarchical attention is used to focus on keywords and important terms; identify key sentences and central arguments; understand document structure and contextual relationships. Context-aware retrieval dynamically adjusts the weights of different parts during the retrieval process.

[0082] B4. Semantic similarity calculation and matching Implement multiple semantic similarity calculation methods. Similarity calculation method, calculate the cosine angle of two vectors; calculate the Euclidean distance in vector space; calculate the Manhattan distance in vector space.

[0083] Similarity optimization, L2 normalization of embedding vectors, makes similarity calculation more stable; use PCA dimension reduction, reduce vector dimension, improve calculation efficiency; adjust similarity distribution through contrast learning, improve discriminability.

[0084] Semantic matching strategy, match the same expression, identify similar expressions through similarity threshold, consider word-level and sentence-level matching.

[0085] B5. Re-ranking and fusion of retrieval results Re-rank the initial retrieval results to improve the quality of the results. Multi-stage retrieval uses vector similarity to quickly screen the candidate set, combined with the results of keyword retrieval and semantic retrieval.

[0086] Result fusion strategy, assign different weights to the results of different retrieval methods. Group by category first, then sort within the group to ensure that the retrieval results cover different angles and sources.

[0087] In another exemplary embodiment, as shown in Figure 3 the data security risk assessment standard in step 102 above is based on a data security risk assessment vertical large model, a compliance and security risk analysis evaluation item prompt is constructed, and after artificial verification, a structured evaluation question library is formed. The following steps 501-504 can be used instead.

[0088] Step 501, use the data security risk assessment vertical large model to sort the data security risk assessment standard, obtain structured compliance judgment items, and generate an evaluation question library. This step combines data security laws and regulations, industry standards, and fine-tunes clear, independent, and structured compliance judgment items to form a unified structured evaluation question library.

[0089] Step 502, determine the dependency relationship between different evaluation questions in the evaluation question library.

[0090] Step 503, generate an evaluation process template according to the dependency relationship between different evaluation questions in the evaluation question library.

[0091] Step 504, using a directed acyclic graph to represent the evaluation process template, constructing a compliance and security risk analysis evaluation item prompt, and forming a structured evaluation question library after manual verification.

[0092] Steps 502-504 above design a structured template for each question in the structured evaluation question library, including clause description, logical judgment criteria, and generation format specification, forming a compliance judgment chain.

[0093] This process has dynamic management functions, allowing the addition, modification or deletion of questions in the question library and corresponding templates according to actual needs, ensuring the flexibility of the system.

[0094] C1, Construction and management of structured evaluation question library The structured evaluation question library is constructed and maintained to support rule classification, retrieval and version control. The structured evaluation question library is structured to add tags to each question and establish dependencies between questions, such as question B must be evaluated after question A is passed.

[0095] Problem templating defines the compliance judgment criteria for each question and associates relevant regulations or industry best practices.

[0096] Dynamic update mechanism, record the history version of question library, support rollback and change tracking, support quick update when new regulation requirement or enterprise policy change, after update automatically verify the integrity and consistency of judgment chain.

[0097] C2, Design and implementation of structured template Define standardized evaluation process templates to support automatic execution of judgment chains. Template structure design, use directed acyclic graph to represent evaluation process, node is evaluation question, edge is flow condition. Define branch logic, such as condition A is met, execute path B, otherwise execute path C. Define the conditions for completing the evaluation, all mandatory questions have been answered.

[0098] Template engine supports variables and expressions in templates, supports logical operations and comparison operations, and maintains state data during the process.

[0099] C3, Dynamic management mechanism design Dynamic management mechanism design supports dynamic configuration and execution of judgment chain, adapts to different evaluation scenarios. Dynamic rule loading supports loading evaluation rules in different fields on demand, dynamically selects and combines judgment chains according to evaluation objectives; automatically adjust rule weights according to properties of evaluation objects.

[0100] Real-time feedback and correction, real-time verification of the rationality of intermediate results during the evaluation process, detect unreasonable results and prompt manual review, automatically adjust rule parameters and weights based on historical evaluation data.

[0101] C4, compliance judgment chain execution engine The engine architecture converts natural language rules into executable logical expressions, executes nodes in the judgment chain in order, handles branching logic, aggregates evaluation results, and generates compliance scores.

[0102] The execution flow loads the evaluation template and the question library, creates an execution context, executes the evaluation nodes in the order defined by the template, determines the subsequent path based on the node execution results and flow conditions, collects the evaluation results of all nodes, and calculates the final compliance score.

[0103] C5, integration with other functions Integration with semantic encoding process, obtain compliance and security risk analysis evaluation items prompt from structured evaluation question library, perform semantic search in benchmark checking vector database to generate enhanced prompt.

[0104] Integration with multi-level reasoning process, call multi-level reasoning module for deep analysis of complex problems, and use reasoning results as input of judgment chain.

[0105] Integration with report generation process, provide detailed evaluation results and compliance evidence, and support the generation of compliance reports that meet industry standards.

[0106] In another exemplary embodiment, the above steps 103-105 first call compliance evaluation items prompt in the structured evaluation question library, perform semantic search in the benchmark checking vector database to generate enhanced prompt for compliance evaluation; call security risk analysis evaluation items prompt in the structured evaluation question library, perform semantic search in the benchmark checking vector database to generate enhanced prompt for security risk analysis. Then, according to the enhanced prompt for compliance evaluation, use the data security risk assessment vertical large model to perform compliance reasoning, obtain the evaluation conclusion and related basis of the compliance evaluation items for manual judgment, form the compliance evaluation record and result; according to the enhanced prompt for security risk analysis, use the data security risk assessment vertical large model to perform risk comprehensive reasoning, identify risk sources, hazard degree and occurrence probability, and obtain risk level determination process and conclusion.

[0107] Among them, step 105 uses the data security risk assessment vertical large model to perform multi-level reasoning based on the relevant text content provided by the structured template and semantic similarity comparison. The multi-level reasoning capability realizes the accurate diagnosis from the data security risk assessment benchmark checking materials to the data security risk assessment. In this embodiment, a three-level progressive reasoning architecture is used as an example.

[0108] D1, first level reasoning Primary reasoning conducts in-depth analysis on a single assessment item. Based on the compliance judgment, combined with the detailed information in the benchmark checking materials and the relevant knowledge in the knowledge base, the causes of the violation behavior are further reasoned, including: Through analyzing the hierarchical structure and semantic relationship of the document, key information is extracted. The document structure analysis identifies the chapter structure of the document based on the title format and indentation, analyzes the reference, dependency and hierarchical relationship between clauses, and divides the document into paragraphs or sentences with independent semantics.

[0109] Key information extraction identifies entities such as industry standards, data types, risk levels, etc.; extracts the relationship between entities, judges the applicable data types of the standard; extracts the attributes of entities, and the risk levels of different data items.

[0110] D2, secondary reasoning Secondary reasoning starts from the correlation between multiple related assessment items, analyzes the diffusion and correlation of risks, including: Starting from the correlation between multiple related assessment items, a graph attention network is constructed and applied to capture the complex semantic relationship between entities. Graph construction and representation take entities as graph nodes and construct edges based on the relationship between entities.

[0111] Among them, the graph attention mechanism uses multiple attention heads to capture different types of relationships, automatically learns the importance weight of different neighbor nodes, and updates the node representation through message passing between nodes.

[0112] D3, tertiary reasoning Tertiary reasoning rises to the level of the enterprise's overall data security system, and comprehensively analyzes the overall situation and weak links of enterprise data security risks based on the reasoning results of all assessment items.

[0113] This stage uses pointer network positioning to reason, that is, it uses a pointer network to accurately locate the key information segment in the document. The specific process is as follows: use a pre-trained language model to encode the document content, learn the pointer to a specific location in the document through attention mechanism, and maximize the probability of predicting the pointer position.

[0114] Accurate positioning of the specific clauses cited in the document extracts specific text segments from the document to support risk assessment, identifies specific descriptions of data flow paths and processing operations, and analyzes the overall situation and weak links of enterprise data security risks based on the reasoning results of all assessment items.

[0115] In another exemplary embodiment, the Monte Carlo Dropout quantifies the confidence in the multi-level reasoning process, and low-confidence results trigger human review to ensure zero-miss of critical evidence. Based on the reasoning process and the corresponding evidence materials, the corresponding evaluation results are given, and the risk source identification and safety impact analysis are carried out. The model output strictly follows the requirements of the structured template, including clear compliance conclusions, cited evidence, problem explanations, and rectification measures, etc.

[0116] The uncertainty of the reasoning result is quantified by multiple random forward propagation. The confidence quantification method keeps Dropout open during the reasoning phase and introduces randomness. Multiple forward propagations are performed on the same input to obtain multiple prediction results, and the confidence score and uncertainty interval are calculated based on the multiple prediction results.

[0117] The credibility of the risk assessment result is quantified, the uncertain edge cases of the model are identified, and human review is triggered. In retrieval and recommendation, the results are reordered based on the confidence.

[0118] In another exemplary embodiment, the technical solution of the present application also includes the process of generating logical explanation, that is, after the above step 105, the process of tracing the key nodes in the reasoning process of multi-level reasoning, extracting the input data, reasoning logic and intermediate results of each level in multi-level reasoning, and generating visual reasoning chain diagram.

[0119] This process provides a clear, traceable, and logical explanation for the risk assessment result, enhancing the credibility of the report. It clearly explains which standard clause, which part of the enterprise material, and through which reasoning steps leads to a certain evaluation conclusion. The judgment basis is linked to form a complete evidence chain. The technical judgment basis and logical explanation process are converted into easy-to-understand, professional natural language descriptions.

[0120] After the model generates the evaluation conclusion, the key nodes in the reasoning process are automatically traced, and the input data, reasoning logic and intermediate results of each level in multi-level reasoning are extracted.

[0121] Natural language generation technology is used to convert the traced logic into easy-to-understand text descriptions and associate them with the corresponding standard provisions and knowledge base content.

[0122] A visual reasoning chain diagram is generated to intuitively display the reasoning path from the original data to the evaluation conclusion, supporting users to click to view detailed information at each link.

[0123] E1, Explainable Chain Construction A complete reasoning chain from input to output is constructed to achieve traceability of explanation. The explanation unit is designed to correspond to a specific text fragment in the original document, recording the model's reasoning process from evidence to conclusion, as well as the stage conclusions in the reasoning process, and generating the final result of the risk assessment.

[0124] The chain-building process involves locating key text from the document that supports the conclusion and analyzing the logical relationships between the evidence. Explanatory units are then organized in logical order to form a complete chain, and a credibility weight is assigned to each explanatory unit.

[0125] E2, Structured Logic Representation Design standardized logical representation methods to support the interpretation of complex reasoning. Logical structure types include those representing conditional logical relationships; relationships between cause and effect; relationships between whole and part, general and specific; and comparative relationships between different entities or viewpoints.

[0126] Formal representation uses predicates, variables, and quantifiers to represent logical relationships; nodes and edges to represent concepts and relationships; and conditions-based decision-making processes.

[0127] E3. Explanation, Generation, and Evaluation The explanation generation process obtains evidence supporting the conclusion from multi-level reasoning modules, and organizes the evidence and reasoning steps using structured logic representations. It constructs a complete explanation chain from evidence to conclusion, converting the structured explanation into understandable text.

[0128] Explain the evaluation metrics; whether they accurately reflect the model's decision-making process; whether human users can understand and trust the explanations; whether the explanations include all necessary reasoning steps; and whether the explanations provide details for specific cases.

[0129] In another exemplary embodiment, the built-in template in step 106 above is a data security risk assessment report template. This step is based on a structured knowledge graph (related assessment items, evidence, risk levels and rectification suggestions), and the content is intelligently filled in by a dynamic template engine, which can realize the automated output of data security risk assessment conclusions.

[0130] First, there are pre-set templates that cover the assessment overview, compliance assessment, security risk analysis, assessment conclusions, security risks and rectification suggestions, etc. The templates allow users to customize and adjust the chapters and format.

[0131] Then, after obtaining the data security risk assessment results, key information is automatically extracted and filled into the corresponding positions according to the template structure.

[0132] Then, by integrating related texts and reasoning explanations, a complete evaluation report is generated.

[0133] It offers online editing capabilities, allowing users to modify report content, add manual annotations, and export to PDF, DOCX, and other formats. The report is automatically updated and saved after manual corrections for easy review and traceability. Regarding compliance assessment results, it calculates average normalized arithmetic scores for basic security assessment and data lifecycle security assessment indicators to determine whether the compliance assessment is passed. For risk source identification in security impact analysis, it determines the probability level of data processing activities that may trigger data security incidents. For security impact analysis, it analyzes and assesses the security impact level after a data security incident occurs, referring to a security impact level judgment table. It supports Word / PDF dual-format output and has a built-in Jinja2 template parser to dynamically insert risk scores and visual charts, generating a complete data security risk assessment report.

[0134] F1, Manual Correction The collaborative design allows users to adjust report content and format in real time via an interface, and to review system-generated reports and make necessary revisions. User feedback is processed by identifying feedback types, such as content modifications and format adjustments. The system automatically adjusts report generation strategies based on feedback, learning from user feedback to improve the quality of future reports.

[0135] F2, Scoring Calculation After completing the assessment of all evaluation items, the compliance assessment results are scored, risk sources are identified and determined, and security impacts are analyzed and assessed to finally determine the security risk level of the data.

[0136] F3, Adaptive Template The system automatically selects or generates suitable report templates based on different scenarios and needs, and supports user-defined template structure and content requirements. It automatically selects appropriate templates based on assessment objectives and industry, and combines portions of content from multiple templates to generate customized reports. The template structure is dynamically adjusted based on the characteristics of the assessment results.

[0137] F4, Report Generation Process Integration Understand users' report generation needs, select or create appropriate templates based on those needs, support user-system interaction and collaboration, and generate data security risk assessment reports that meet the requirements.

[0138] Based on the same inventive concept, this application also provides a large-model-based data security risk assessment system for implementing the aforementioned large-model-based data security risk assessment method. The solution provided by this device is similar to the implementation scheme described in the above method; therefore, the specific limitations of one or more large-model-based data security risk assessment system embodiments provided below can be found in the limitations of the large-model-based data security risk assessment method described above, and will not be repeated here.

[0139] In one exemplary embodiment, a data security risk assessment system based on a large model is provided, such as... Figure 4 As shown, it includes: a data loading and preprocessing module, a semantic encoding module, a compliance assessment engine module, a security risk analysis engine module, a multi-level inference module, and a report generation module; The data loading and preprocessing module is used to acquire benchmarking verification materials and standard specifications for data security risk assessment, and to preprocess the benchmarking verification materials to obtain preprocessed benchmarking verification materials. The semantic encoding module is used to encode the pre-processed benchmarking and verification materials and the standard specifications into a unified semantic space using a pre-trained language model, forming a benchmarking and verification vector database. Based on the data security risk assessment standard, it uses a vertical large model for data security risk assessment to construct compliance and security risk analysis and assessment items prompts, which are then manually verified to form a structured assessment question library. The compliance assessment engine module is used to call the compliance assessment prompts in the structured assessment question library and perform semantic retrieval in the benchmarking and verification vector database to generate enhanced prompts for compliance assessment. The security risk analysis engine module is used to call the security risk analysis assessment item prompts in the structured assessment question library and perform semantic retrieval in the benchmarking and verification vector database to generate enhanced prompts for security risk analysis. The multi-level reasoning module is used to perform multi-level reasoning based on the enhanced prompts for compliance assessment and security risk analysis using the vertical large model of data security risk assessment, to obtain compliance assessment records and results, risk level determination process and conclusions, as the data security risk assessment results; the multi-level reasoning includes: compliance determination, risk source identification, and determination of the degree of harm and probability of occurrence. The report generation module is used to generate an assessment report based on the data security risk assessment results using a built-in template.

[0140] The structural diagram of the data security risk assessment system based on a large model in this application is referenced. Figure 4 Based on the vertical large model of data security risk assessment, the benchmarking and verification materials used for data security risk assessment are input. After data preprocessing, semantic encoding, compliance assessment, security risk analysis, multi-level reasoning and report generation, the benchmarking and verification materials and reasons on which the risk occurs are output in combination with the assessment results and the probability of risk occurrence. After the output results are manually corrected, a score is calculated and the risk is comprehensively assessed to obtain the assessment conclusion, and finally a data security risk assessment report is generated.

[0141] The data security risk assessment method and system based on a large model proposed in this application address the systematic compliance pressure faced by enterprises after the implementation of industry standards such as the "Data Security Risk Assessment Specification for Industrial Sector" and the "Data Security Risk Assessment Specification for Telecommunications Sector," as well as the pain points of existing assessment work, such as high reliance on manual labor, low efficiency, strong subjectivity, and difficulty in traceability. It constructs an automated assessment scheme that integrates large-scale model semantic understanding, knowledge association, and rule-driven reasoning. Based on a vertical large-scale model for data security risk assessment, this method and system, through the collaborative work of data loading and preprocessing, semantic encoding, compliance assessment engine, security risk analysis engine, multi-level reasoning, and report generation, can deeply analyze the structured and unstructured information in the enterprise's data security risk assessment benchmarking materials. It automatically locates the corresponding content of standard clauses, extracts key elements as supporting evidence, and realizes a line-by-line review of subdivided assessment items such as legitimacy and necessity, basic security, and data lifecycle security. It then generates an assessment report containing compliance conclusions, risk levels, and supporting evidence. Its core advantage lies in empowering the entire assessment process with large-scale model technology, integrating expert knowledge and standardized reasoning logic, improving assessment efficiency and accuracy through automated processing, and addressing the subjectivity and lack of transparency inherent in manual assessments through a traceable reasoning chain. This application provides technical support for enterprises to meet industry-specific compliance requirements, helping them efficiently complete data security risk assessments and promptly meet compliance requirements, demonstrating significant practical value and application prospects.

[0142] The data security risk assessment method and system based on large models proposed in this application have the following technical effects and advantages: ① Based on professional knowledge and standards for data security risk assessment, we reviewed the assessment items and corresponding supporting materials from past data security risk assessment projects, and meticulously annotated them after data cleaning. For the compliance assessment stage, the annotations included assessment category, assessment subcategory, assessment item, assessment record, and assessment result. For the security risk analysis stage, the annotations included judgment factors, factor content, judgment item, result record, and probability level. This resulted in a data security risk assessment training dataset, providing clear supervisory basis for training a large-scale vertical model for data security risk assessment.

[0143] ② Effectively bridges the terminology gap between industry standards and enterprise practices, achieving high-precision semantic matching, ensuring semantic consistency between industry standards and data security risk assessment benchmarking materials, improving terminology matching accuracy, and resolving the issue of missed detections caused by "different expressions of the same concept." It generates an interpretable terminology mapping rule base, supporting manual review and adjustment, and adversarial training significantly reduces the impact of domain differences, ensuring semantic matching accuracy in cross-industry scenarios.

[0144] ③ By employing a pre-trained language model for comparative learning and attention mechanisms, the limitations of traditional keyword matching are overcome, capturing implicit semantic relationships. The terminology mapping rule base supports dynamic updates; when adding new industry standards or data security risk assessment benchmarking materials, only minor model adjustments are needed for rapid adaptation, reducing secondary development costs. A unified semantic space provides a consistent cognitive standard for subsequent reasoning modules, avoiding evaluation errors caused by misunderstandings of terminology.

[0145] ④ By meticulously reviewing compliance items and using structured templates, we ensured consistent evaluation standards, transparent processes, and highly comparable results, avoiding subjective arbitrariness. Mandatory templates ensured uniform formatting and completeness of model-generated results, greatly improving readability, auditability, and interpretability, and facilitating manual review and traceability.

[0146] ⑤ Supports association of multi-format documents (PDF / Word / charts) with complex evidence. The three-level progressive reasoning architecture improves the efficiency of evaluation item matching, and the graph attention network effectively improves the modeling accuracy of evaluation item dependencies. The pointer network enables character-level localization of supporting evidence fragments, and combined with Monte Carlo Dropout to quantify confidence, the rate of missed detection of key evidence decreases.

[0147] ⑥ Automate the entire data security risk assessment process, creating a closed loop from data input to report output, improving assessment efficiency and reducing human error. Supports both Word and PDF formats and visual charts to meet report preparation standards.

[0148] The data security risk assessment method and system based on large-scale models has achieved a revolutionary leap from technical capabilities to industrial value. Breaking down domain barriers with deep semantic cognition, the system uses a pre-trained language model for comparative learning and dynamic attention mechanisms to establish a high-precision semantic bridge between industry standards and data security risk assessment benchmarking materials, essentially building a cross-domain compliance semantic infrastructure. When new standards or document types are added, the fine-tuning mechanism reduces the adaptation cycle from months to days, significantly lowering the cost for enterprises to cope with regulatory iterations.

[0149] A multi-level inference engine reshapes decision-making accuracy. A three-tiered progressive architecture integrates graph attention networks and pointer networks, achieving character-level evidence localization while ensuring global logical coherence. Monte Carlo Dropout confidence mechanisms and manual feedback enable continuous system iteration, reducing the false negative rate. This capability is particularly prominent in fragmented documents, providing enterprises with the data security risk assessment capabilities of a professional team.

[0150] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0151] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A data security risk assessment method based on a large model, characterized in that, The data security risk assessment method based on a large model applies a vertical large model for data security risk assessment, and the data security risk assessment method based on a large model includes: Obtain benchmarking verification materials and standard specifications for data security risk assessment, and preprocess the benchmarking verification materials to obtain preprocessed benchmarking verification materials; The pre-trained language model is used to encode the pre-processed benchmarking and verification materials and the standard specifications into a unified semantic space to form a benchmarking and verification vector database. Based on the data security risk assessment standard, a compliance and security risk analysis and assessment prompt is constructed using a vertical large model for data security risk assessment. After manual verification, a structured assessment question library is formed. The compliance assessment prompt from the structured assessment question library is called, and semantic retrieval is performed in the benchmarking and verification vector database to generate an enhanced prompt for compliance assessment. The system calls upon the security risk analysis and assessment prompt from the structured assessment question library and performs semantic retrieval in the benchmarking and verification vector database to generate an enhanced prompt for security risk analysis. Based on the enhanced prompts used for compliance assessment and security risk analysis, the data security risk assessment vertical big model is used to perform multi-level reasoning to obtain compliance assessment records and results, risk level determination process and conclusions, which serve as the data security risk assessment results. The multi-level reasoning includes: compliance determination, risk source identification, and determination of the degree of harm and probability of occurrence. Based on the data security risk assessment results, an assessment report is generated using a built-in template.

2. The data security risk assessment method based on a large model according to claim 1, characterized in that, The aforementioned vertical large-scale model for data security risk assessment was trained using the following steps: Obtain assessment records and corresponding supporting materials from historical data security risk assessment projects, standardize them, and construct a training set; By using unlabeled target domain data to pre-train a selected large model, the large model learns the common language patterns, semantic features, and potential knowledge related to data security in the target domain, thus obtaining a pre-trained large model. The pre-trained large model is fine-tuned using the training set to enable it to understand the textual semantics and semantic relationships in industry standards and benchmarking verification materials in the target domain, thus obtaining the trained large model, which serves as the vertical large model for data security risk assessment.

3. The data security risk assessment method based on a large model according to claim 2, characterized in that, Obtain assessment records and corresponding supporting materials from historical data security risk assessment projects, standardize them, and construct a training set, specifically including: Obtain assessment records and supporting materials from historical data security risk assessment projects to construct the original dataset; The original dataset is cleaned to obtain a cleaned dataset; Entity annotation and relation annotation methods are used, and industry standards in the target domain are used to annotate the cleaned dataset to obtain the annotated dataset; The labeled dataset is augmented using natural language processing techniques to obtain the training set; the data augmentation includes synonym replacement, sentence structure transformation, and noise addition.

4. The data security risk assessment method based on a large model according to claim 3, characterized in that, The method of fine-tuning the pre-trained large model using the training set is called incremental fine-tuning.

5. The data security risk assessment method based on a large model according to claim 1, characterized in that, The benchmarking verification materials are preprocessed to obtain preprocessed benchmarking verification materials, specifically including: Convert the different format files in the benchmarking and verification materials into machine-processable text to obtain the converted benchmarking and verification materials; A sentence segmenter was used to cut the machine-processable text in the converted benchmarking materials into multiple text fragments according to semantic units; Natural language processing techniques are used to segment each text segment into words, and word vector models are used to vectorize the words in each text segment to obtain the vectorized representation of each text segment. Extract the chapter titles and paragraph topics of each text segment to generate a tree structure representation of each text segment; Extract key information from each text segment; the key information includes: requirement statements, prohibition statements, dates, risk levels, and assessment indicators; The vectorized representation, tree structure representation, and key information of each text fragment constitute the preprocessed benchmarking and verification materials.

6. The data security risk assessment method based on a large model according to claim 1, characterized in that, The pre-trained language model is used to encode the pre-processed benchmarking verification materials and the aforementioned standard specifications into a unified semantic space, forming a benchmarking verification vector database, specifically including: By using the vertical large-scale model of data security risk assessment to locate the terminology in the preprocessed benchmarking and verification materials and standard specifications, a context-aware vector representation of each term is generated; The attention mechanism is used to analyze the context-aware vector representations of each term to obtain the correlation strength between different terms; Based on the correlation strength between different terms, the preprocessed benchmarking verification materials and the aforementioned standard specifications are encoded into a unified semantic space to construct a benchmarking verification vector database.

7. The data security risk assessment method based on a large model according to claim 1, characterized in that, Based on the enhanced prompts used for compliance assessment and security risk analysis, multi-level reasoning is performed using the aforementioned vertical big data security risk assessment model to obtain compliance assessment records and results, risk level determination processes and conclusions, specifically including: Based on the enhanced prompt used for compliance assessment, compliance reasoning is performed using a large vertical model for data security risk assessment to obtain assessment conclusions and related evidence for compliance assessment items, and to form compliance assessment records and results. Based on the enhanced prompt used for security risk analysis, a comprehensive risk reasoning process is conducted using a vertical large-scale model for data security risk assessment to identify risk sources, severity levels, and likelihood of occurrence, and to arrive at the risk level determination process and conclusions.

8. The data security risk assessment method based on a large model according to claim 1, characterized in that, Based on data security risk assessment standards, and utilizing a vertical data security risk assessment model, a compliance and security risk analysis and assessment prompt is constructed. After manual verification, a structured assessment question library is formed, specifically including: Using a vertical big data security risk assessment model, the data security risk assessment standards are sorted out to obtain structured compliance judgment items and generate an assessment question library; Determine the dependencies between different evaluation questions in the evaluation question bank; Based on the dependencies between different assessment questions in the assessment question library, an assessment process template is generated; The assessment process template is represented by a directed acyclic graph, and a compliance and security risk analysis and assessment prompt is constructed. After manual verification, a structured assessment question library is formed.

9. The data security risk assessment method based on a large model according to claim 1, characterized in that, Based on the enhanced prompts used for compliance assessment and security risk analysis, multi-level reasoning is performed using the aforementioned vertical big data security risk assessment model to obtain compliance assessment records and results, risk level determination processes and conclusions, which serve as the data security risk assessment results. This also includes: By tracing the key nodes in the reasoning process of multi-level reasoning, the input data, reasoning logic and intermediate results of each level in multi-level reasoning are extracted to generate a visual reasoning chain diagram.

10. A data security risk assessment system based on a large model, characterized in that, The large-model-based data security risk assessment system applies the large-model-based data security risk assessment method according to any one of claims 1-9. The large-model-based data security risk assessment system includes: a data loading and preprocessing module, a semantic encoding module, a compliance assessment engine module, a security risk analysis engine module, a multi-level inference module, and a report generation module. The data loading and preprocessing module is used to acquire benchmarking verification materials and standard specifications for data security risk assessment, and to preprocess the benchmarking verification materials to obtain preprocessed benchmarking verification materials. The semantic encoding module is used to encode the pre-processed benchmarking and verification materials and the standard specifications into a unified semantic space using a pre-trained language model, forming a benchmarking and verification vector database. Based on the data security risk assessment standard, it uses a vertical large model for data security risk assessment to construct compliance and security risk analysis and assessment items prompts, which are then manually verified to form a structured assessment question library. The compliance assessment engine module is used to call the compliance assessment prompts in the structured assessment question library and perform semantic retrieval in the benchmarking and verification vector database to generate enhanced prompts for compliance assessment. The security risk analysis engine module is used to call the security risk analysis assessment item prompts in the structured assessment question library and perform semantic retrieval in the benchmarking and verification vector database to generate enhanced prompts for security risk analysis. The multi-level reasoning module is used to perform multi-level reasoning based on the enhanced prompts for compliance assessment and security risk analysis using the vertical large model of data security risk assessment, to obtain compliance assessment records and results, risk level determination process and conclusions, as the data security risk assessment results; the multi-level reasoning includes: compliance determination, risk source identification, and determination of the degree of harm and probability of occurrence. The report generation module is used to generate an assessment report based on the data security risk assessment results using a built-in template.

Citation Information

Patent Citations

  • Vertical domain document question and answer method and system based on knowledge graph enhanced large model

    CN119646026A

  • Contract risk assessment and compliance check method, system and equipment based on large language model, and medium

    CN119863119A

  • Intelligent hospital guide method and device based on large language model and multi-source knowledge dynamic enhancement

    CN120581161A

  • Large model application construction method based on configurable workflow and domain knowledge base

    CN120631869A

  • Artificial-intelligence-based system and method for questionnaire / security policy cross-correlation and compliance level estimation for cyber risk assessments

    US20240273214A1

Cited By

  • AI model security alignment method, device and system based on vertical domain detection engine

    CN122021974A

  • Private data compliance auxiliary evaluation method and system based on large language model

    CN122065342A

  • Framework mining-based port entry and exit article risk assessment method and system

    CN122311890A