Data security management evaluation method and related equipment
By combining multimodal large models and vector knowledge bases, the data security management system is automatically sliced, retrieved, and evaluated, solving the problems of low verification efficiency and poor accuracy in existing technologies, and realizing efficient and objective data security management evaluation.
Patent Information
- Application Number
- CN202511236266.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2026-01-16
AI Technical Summary
The existing data security management system suffers from problems such as inefficiency, low accuracy, strong subjectivity, insufficient coverage, and untimely knowledge updates in its verification methods.
The document to be checked is transformed by a multimodal large model, and information is sliced according to the logical structure of the text document. The pre-built vector knowledge base is used for retrieval and matching. The text slices and candidate vectors are then input into the large language model to output the data security management evaluation results.
It has enabled automated and intelligent analysis of data management system documents, improving the efficiency, accuracy, and comprehensiveness of verification, and reducing omissions and subjective biases.
Smart Images

Figure CN121349971A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security technology, and in particular to a data security management and evaluation method and related equipment. Background Technology
[0002] Currently, the verification of data security management systems mainly relies on manual methods, where data security experts, legal or auditing personnel read through the policy documents line by line and compare them with the requirements of laws, regulations, industry standards, and best practices. This method suffers from problems such as inefficiency, low accuracy, strong subjectivity, insufficient coverage, and untimely knowledge updates. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a data security management evaluation method and related equipment to solve the problems of low efficiency, low accuracy, strong subjectivity, insufficient coverage and untimely knowledge updates in the current data security management system verification methods.
[0004] To achieve the above objectives, the first aspect of this application provides a data security management evaluation method, comprising:
[0005] The document to be checked is converted using a multimodal large model to obtain a text file;
[0006] The text file is sliced according to its logical structure information to obtain multiple text slices;
[0007] Based on multiple text slices, a search and matching process is performed in a pre-built vector knowledge base to determine multiple candidate vectors;
[0008] Multiple text slices and multiple candidate vectors are input into a large language model, and the large language model outputs data security management evaluation results.
[0009] Optionally, the step of slicing the text file according to its logical structure information to obtain multiple text slices includes:
[0010] The text file is sliced according to its logical structure information to obtain multiple initial slices;
[0011] Determine whether the slice lengths of multiple initial slices meet preset conditions;
[0012] In response to the fact that the slice lengths of multiple initial slices meet the preset conditions, a semantic integrity check is performed on the multiple initial slices;
[0013] In response to multiple initial slices passing semantic integrity checks, the multiple initial slices are standardized to generate multiple text slices.
[0014] Optional, also includes:
[0015] In response to the fact that the slice length of multiple initial slices does not meet the preset conditions and there is a slice length greater than the preset threshold, the text file is semantically subdivided so that the slice length of each initial slice meets the preset conditions.
[0016] In response to the fact that the slice lengths of multiple initial slices do not meet the preset conditions and there are slices with lengths less than the preset threshold, adjacent slices are merged so that the slice lengths of multiple initial slices meet the preset conditions.
[0017] Optionally, the step of searching and matching in a pre-built vector knowledge base based on multiple text slices to determine multiple candidate vectors includes:
[0018] Multiple text slices are converted into query vectors, and vector similarity is calculated to determine multiple first initial candidate vectors in the vector knowledge base.
[0019] Multiple initial candidate vectors are filtered based on a similarity threshold to obtain multiple second initial candidate vectors;
[0020] Multiple candidate vectors are scored in multiple dimensions to determine multiple candidate vectors.
[0021] Optional methods for constructing vector knowledge bases include:
[0022] Collect data source files and perform data cleaning and formatting on the data source files;
[0023] The data source file, after data cleaning and formatting, is sliced to obtain multiple data source slices;
[0024] Vectorize multiple data source slices, converting them into multiple data source vectors;
[0025] The vector knowledge base is constructed based on multiple data source vectors.
[0026] Optional, also includes:
[0027] Construct a vector knowledge base retrieval interface;
[0028] The vector knowledge base retrieval interface is used to perform retrieval tests, and the vector knowledge base is deployed in response to passing the retrieval test.
[0029] Optionally, the step of inputting multiple text slices and multiple candidate vectors into a large language model, and outputting data security management evaluation results through the large language model, includes:
[0030] Construct cue words based on multiple text slices and multiple candidate vectors;
[0031] Multiple text slices, multiple candidate vectors, and prompt words are input into the large language model to generate analysis results;
[0032] The analysis results are format-validated. In response to passing the format validation, a multi-dimensional evaluation is performed based on the analysis results, and the data security management evaluation results are output through the large language model.
[0033] Based on the same inventive concept, a second aspect of this application also provides a data security management and evaluation device, comprising:
[0034] The conversion module is configured to convert the document to be checked using a multimodal large model to obtain a text file.
[0035] The slicing module is configured to slice the text file according to the logical structure information of the text file to obtain multiple text slices;
[0036] The retrieval module is configured to perform retrieval and matching in a pre-built vector knowledge base based on multiple text slices to determine multiple candidate vectors;
[0037] The evaluation module is configured to input multiple text slices and multiple candidate vectors into a large language model, and output data security management evaluation results through the large language model.
[0038] Based on the same inventive concept, a third aspect of this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.
[0039] Based on the same inventive concept, a fourth aspect of this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described above.
[0040] As described above, the data security management evaluation method and related equipment provided in this application include the following steps: converting the document to be verified using a multimodal large model to obtain a text file; slicing the text file according to its logical structure to obtain multiple text slices; performing retrieval and matching based on the multiple text slices in a pre-built vector knowledge base to determine multiple candidate vectors; inputting the multiple text slices and candidate vectors into a large language model, and outputting the data security management evaluation result through the large language model. The method provided in this application enables automated and intelligent analysis of data management system documents, improving the efficiency, accuracy, and comprehensiveness of verification. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 This is a flowchart illustrating the data security management and evaluation method according to an embodiment of this application.
[0043] Figure 2 This is a schematic diagram of the data security management and evaluation device according to an embodiment of this application;
[0044] Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0046] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0047] Currently, the verification of data security management systems primarily relies on manual methods. Teams composed of data security experts, legal counsel, and auditors hired internally or externally manually read the policy documents, using their professional knowledge and experience, along with legal provisions and standards, to compare, judge, and record each item. The initial stage of automated management system verification involves searching and matching policy documents using preset keywords (such as "data breach," "access control," and "single consent"). Alternatively, rule engines are used, with experts pre-writing numerous "IF-THEN" logical rules. This approach has the following significant drawbacks:
[0048] 1. Inefficient and costly: Manual review is time-consuming and labor-intensive, requiring a large investment of expert resources. For large organizations with extensive and complex systems, the review cycle is long and the cost is high.
[0049] 2. Low accuracy: Search and rule engines often lack semantic understanding. For example, content in management systems that is semantically correct may be judged as "not covered" because precise keywords are not used.
[0050] 3. Highly subjective and inconsistent standards: Audit results heavily rely on the professional level, experience, and understanding of the auditors. Different people may reach different conclusions, lacking objectivity and consistency.
[0051] 4. Insufficient coverage and prone to omissions: Among the massive amount of laws and regulations and standards, manual comparison is bound to result in omissions, especially when dealing with frequently updated laws and regulations, it is difficult to achieve real-time and comprehensive coverage.
[0052] 5. Delayed knowledge updates: Laws, regulations, and safety standards are constantly evolving, while the update speed of human knowledge bases is slow, which may lead to the review process being based on outdated requirements.
[0053] In view of this, this application proposes a data security management evaluation method, aiming to solve the problems of low efficiency, strong subjectivity, and easy omissions in the verification of data security management systems in existing technologies. It provides an automated verification method for data security management systems based on a large model and knowledge base. This method can achieve automated and intelligent analysis of system documents, improving the efficiency, accuracy, and comprehensiveness of verification.
[0054] The embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0055] This application proposes a data security management evaluation method, which is referenced in the embodiments of this application. Figure 1 This includes the following steps:
[0056] Step 101: Convert the document to be checked using a multimodal large model to obtain a text file.
[0057] Specifically, leveraging the document understanding capabilities of a multimodal large model, the document to be reviewed is processed, automatically identifying elements such as text, charts, and titles, and accurately converting them into plain text format. The document to be reviewed can be a management policy document.
[0058] Step 102: Slice the text file according to its logical structure information to obtain multiple text slices.
[0059] Specifically, the extracted text files are automatically sliced according to their inherent logical structure (such as chapters, clauses, and paragraphs). Each slice is treated as a unit of text to be analyzed, ensuring granular analysis and contextual integrity.
[0060] Step 103: Based on multiple text slices, perform retrieval and matching in a pre-built vector knowledge base to determine multiple candidate vectors.
[0061] Specifically, for each text slice, a text embedding model is used to vectorize it. The vectorized slice vectors are then used to perform similarity retrieval in a pre-built vector knowledge base. The system retrieves several knowledge base vectors (i.e., candidate vectors) that are semantically most similar to the current text slice, and uses these knowledge base vectors as the most relevant references for the current text slice.
[0062] Step 104: Input multiple text slices and multiple candidate vectors into the large language model, and output the data security management evaluation results through the large language model.
[0063] Specifically, multiple text slices and multiple candidate vectors are used as input to a large language model. These are pushed to the large language model via an API interface. The large language model then compares and analyzes the multiple text slices and candidate vectors to determine whether the text slices adequately and clearly cover the requirements of the multiple candidate vectors, and outputs a data security management evaluation result. This evaluation result may include an overall compliance score, a coverage score, and risk and gap analysis.
[0064] Based on steps 101 to 104 above, the data security management evaluation method provided in this embodiment includes: converting the document to be verified using a multimodal large model to obtain a text file; slicing the text file according to its logical structure information to obtain multiple text slices; performing retrieval and matching based on the multiple text slices in a pre-built vector knowledge base to determine multiple candidate vectors; inputting the multiple text slices and multiple candidate vectors into a large language model, and outputting the data security management evaluation result through the large language model. This application automatically parses documents of various formats using a multimodal large model, and the system automatically completes slicing, retrieval, comparison, analysis, and result output. It compresses work that might have previously taken weeks into less than an hour, making high-frequency, full-volume policy verification possible and providing a technical foundation for agile compliance for enterprises. This application utilizes a vector knowledge base to store massive amounts of laws, regulations, industry standards, and best practices, far exceeding the knowledge reserves of any individual expert. The automated process ensures that every paragraph of the policy document is systematically compared with the entire knowledge base, eliminating omissions. Simultaneously, the analysis based on algorithms and models eliminates subjective human factors, making the verification results highly objective. The method provided in this application enables automated and intelligent analysis of data management system documents, improving the efficiency, accuracy, and comprehensiveness of verification.
[0065] In some embodiments, slicing the text file according to its logical structure information to obtain multiple text slices includes:
[0066] The text file is sliced according to its logical structure information to obtain multiple initial slices; it is determined whether the slice length of the multiple initial slices meets a preset condition; in response to the multiple initial slices meeting the preset condition, a semantic integrity check is performed on the multiple initial slices; in response to the multiple initial slices passing the semantic integrity check, the multiple initial slices are standardized to generate multiple text slices.
[0067] Specifically, the text file is sliced according to its logical structure information, which may include chapters, clauses, and paragraphs, resulting in multiple initial slices. The length of each initial slice is calculated to determine if it meets preset conditions. The initial slice length cannot be too long or too short. If the initial slice length is equal to or within a preset threshold, it is considered to meet the preset conditions. For example, the preset threshold can be 32KB, and the preset threshold range can be 30KB to 32KB. When multiple initial slices are determined to meet the preset conditions, a semantic integrity check is performed on each initial slice to determine if it is semantically complete, thus improving the accuracy of subsequent vector knowledge base matching. The semantic integrity check can be implemented using a model with semantic understanding capabilities. If multiple initial slices pass the semantic integrity check, each initial slice is standardized. Standardization includes adding context markers to each initial slice and generating slice metadata. After standardization, multiple text slices are generated. If multiple initial slices fail the semantic integrity check, or if there are semantically incomplete initial slices, the slice boundaries are adjusted, such as by re-dividing slice boundaries based on punctuation marks or paragraph centers. By performing text slicing, we can ensure the granularity and contextual integrity of the analysis.
[0068] In some embodiments, it also includes:
[0069] In response to the fact that the slice length of multiple initial slices does not meet the preset conditions and there is a slice length greater than the preset threshold, the text file is semantically subdivided so that the slice length of each initial slice meets the preset conditions.
[0070] In response to the fact that the slice lengths of multiple initial slices do not meet the preset conditions and there are slices with lengths less than the preset threshold, adjacent slices are merged so that the slice lengths of multiple initial slices meet the preset conditions.
[0071] Specifically, if the lengths of multiple initial slices do not meet the preset conditions and some slices exceed the preset threshold, it indicates that the slice length is too long and needs to be shortened. This shortening is achieved by semantically subdividing the text file. Semantic boundaries refer to naturally existing nodes in the text that represent semantic unit transitions. Cutting at these nodes ensures that each slice possesses independent and complete semantics to the greatest extent possible. Common semantic boundaries include: paragraph boundaries (a paragraph typically expresses a complete sub-point); chapter titles (marking transitions to major themes); sentence boundaries (a sentence is the basic unit of complete semantics); punctuation boundaries (indicating pauses in tone or small semantic units); and word boundaries (but, however, on the other hand, in short) (these words imply logical transitions). Natural language processing tools (such as NLP libraries and regular expressions) are used to locate all the above semantic boundary points and increase the cutting positions to achieve semantic boundary subdivision.
[0072] If the slice lengths of multiple initial slices do not meet the preset conditions and some slices are shorter than the preset threshold, it indicates that the slice length is too short and needs to be increased. In practice, adjacent slices can be merged to increase the slice length. Through this embodiment, the slice length can be adaptively adjusted to a reasonable length, ensuring the granularity and contextual integrity of the analysis.
[0073] In some embodiments, the step of retrieving and matching multiple candidate vectors based on multiple text slices in a pre-built vector knowledge base includes:
[0074] Multiple text slices are converted into query vectors, and vector similarity is calculated to determine multiple first initial candidate vectors in the vector knowledge base. Multiple initial candidate vectors are filtered based on a similarity threshold to obtain multiple second initial candidate vectors. Multiple second initial candidate vectors are scored in multiple dimensions to determine multiple candidate vectors.
[0075] Specifically, multiple text slices are vectorized using a text embedding model to obtain query vectors. The query vectors are then compared with vectors in a pre-built vector knowledge base, and the first initial candidate vectors are sorted in descending order of similarity. Next, the top-K first initial candidate vectors are selected to obtain multiple second initial candidate vectors. The value of K can be chosen according to the actual situation. Multiple second initial candidate vectors are scored across multiple dimensions, including vector similarity, keyword matching, and semantic relevance. The scores for vector similarity, keyword matching, and semantic relevance can be calculated separately, allowing for a comprehensive evaluation of each second initial candidate vector to obtain the vectors from the vector knowledge base that have the highest relevance to the text slice vectors. Furthermore, the weights of each dimension can be adjusted according to actual needs, ensuring that the selected candidate vectors better meet user requirements. In this embodiment, combining vectorized semantic understanding with simple keyword matching allows for a deeper understanding of the inherent meaning of institutional clauses and regulatory requirements, avoiding subjective biases in manual review and ensuring the objectivity and consistency of the verification results.
[0076] In some embodiments, the method for constructing a vector knowledge base includes:
[0077] Collect data source files and perform data cleaning and formatting on the data source files;
[0078] The data source file, after data cleaning and formatting, is sliced to obtain multiple data source slices;
[0079] Vectorize multiple data source slices, converting them into multiple data source vectors;
[0080] The vector knowledge base is constructed based on multiple data source vectors.
[0081] Specifically, data sources include, but are not limited to, various types of data such as laws and regulations, national / industry standards, enterprise standards, and best practice cases. The collected data source files are cleaned, formatted, and structured to extract core requirements. Each clause or requirement is treated as an independent knowledge unit. A pre-trained text embedding model is used to vectorize each knowledge unit, converting it into a high-dimensional vector that captures the deep semantic information of the text. The vector representations of all knowledge units are stored in a dedicated vector knowledge base and indexed. For example, each vector in the vector knowledge base can represent a specific data security management requirement. By constructing a vector knowledge base from massive amounts of laws, regulations, standards, and best practices, and performing systematic slice comparisons, omissions can be minimized, ensuring comprehensive verification. When new laws, regulations, or standards are introduced, they can be vectorized and incrementally updated in the vector knowledge base to quickly acquire the ability to verify new requirements and adapt to the ever-changing compliance environment.
[0082] In some embodiments, it also includes:
[0083] Construct a vector knowledge base retrieval interface;
[0084] The vector knowledge base retrieval interface is used to perform retrieval tests, and the vector knowledge base is deployed in response to passing the retrieval test.
[0085] Specifically, the retrieval interface is an application programming interface (API) that receives user queries (text) and returns the most relevant answers (text). After constructing the vector knowledge base retrieval interface, searches and matches can be performed within the vector knowledge base through this interface. Before the vector knowledge base is officially used, a retrieval test is required to ensure its normal operation. That is, vector searches and matches can be performed through the constructed retrieval interface. If the retrieval interface returns the correct matching answer, the vector knowledge base has passed the retrieval test. If the retrieval interface fails to return the correct matching answer, the vector knowledge base has failed the retrieval test and further optimization and adjustments are needed until it passes the retrieval test. After passing the retrieval test, the vector knowledge base is deployed on the system or platform. The method in this embodiment ensures the normal operation of the vector knowledge base and improves the efficiency of subsequent retrieval and matching.
[0086] In some embodiments, the step of inputting multiple text slices and multiple candidate vectors into a large language model, and outputting data security management evaluation results through the large language model, includes:
[0087] Construct cue words based on multiple text slices and multiple candidate vectors;
[0088] Multiple text slices, multiple candidate vectors, and prompt words are input into the large language model to generate analysis results;
[0089] The analysis results are format-validated. In response to passing the format validation, a multi-dimensional evaluation is performed based on the analysis results, and the data security management evaluation results are output through the large language model.
[0090] Specifically, cue words can be constructed based on multiple text slices and multiple candidate vectors. These cue words act as a "language" for communicating with the large language model, guiding and constraining the model to generate high-quality, expected output through instructions, context, and format. In this embodiment, cue words guide the large model to perform comparative analysis on multiple text slices and multiple candidate vectors. Multiple text slices, multiple candidate vectors, and cue words are input into the large language model. The large language model compares the content of the text slices with the content of the candidate vectors to determine whether the text slices sufficiently cover the requirements of the candidate vectors and whether they meet the standards specified by the candidate vectors. The large language model then generates analysis results. The analysis results include an evaluation of each text slice and the degree to which the text slices cover the candidate vectors.
[0091] The output format of the large language model is pre-defined to ensure that it can output data security management evaluation results that meet user needs and format requirements. Therefore, after the large language model generates the analysis results, it is necessary to verify the format of the results to determine if they conform to the pre-defined output format. If they do, a multi-dimensional evaluation is performed based on the analysis results, including coverage, compliance, quantitative indicators, and problem identification. The coverage dimension determines whether the text slice completely covers the content of the candidate vectors. The compliance dimension determines whether the text slice belongs to a compliant document. The quantitative indicator dimension determines whether the large language model outputs deterministic quantitative analysis results, allowing users to clearly understand the comparison results. The problem identification dimension determines which planning clauses or regulations the text slice involves. In this way, the above multi-dimensional analysis provides users with a relatively comprehensive data security management evaluation result. For the coverage dimension, the large language model can output corresponding labels, such as "full coverage," "partial coverage," "no coverage," and "irrelevant," so that users clearly understand the analysis and comparison results. Simultaneously, the large language model can also provide analysis reasons and explanations. In addition, the data security management evaluation results can be presented to users in a visual format, allowing them to understand the results more intuitively. Users can also quickly understand whether the documents to be checked comply with the data security management rules, and if not, what the specific reasons are, etc.
[0092] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0093] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0094] Based on the same inventive concept, and corresponding to any of the above embodiments, this application also provides a data security management and evaluation device.
[0095] refer to Figure 2 The data security management and evaluation device includes:
[0096] The conversion module 201 is configured to convert the document to be checked using a multimodal large model to obtain a text file.
[0097] Slicing module 202 is configured to slice the text file according to the logical structure information of the text file to obtain multiple text slices;
[0098] The retrieval module 203 is configured to perform retrieval and matching in a pre-built vector knowledge base based on multiple text slices to determine multiple candidate vectors;
[0099] Evaluation module 204 is configured to input multiple text slices and multiple candidate vectors into a large language model, and output data security management evaluation results through the large language model.
[0100] In some embodiments, the slicing module 202 is further configured to slice the text file according to the logical structure information of the text file to obtain multiple initial slices;
[0101] Determine whether the slice lengths of multiple initial slices meet preset conditions;
[0102] In response to the fact that the slice lengths of multiple initial slices meet the preset conditions, a semantic integrity check is performed on the multiple initial slices;
[0103] In response to multiple initial slices passing semantic integrity checks, the multiple initial slices are standardized to generate multiple text slices.
[0104] In some embodiments, the slicing module 202 is further configured to perform semantic boundary subdivision on the text file in response to the fact that the slice length of multiple initial slices does not meet the preset conditions and there is a slice length greater than a preset threshold, so that the slice length of each initial slice meets the preset conditions.
[0105] In response to the fact that the slice lengths of multiple initial slices do not meet the preset conditions and there are slices with lengths less than the preset threshold, adjacent slices are merged so that the slice lengths of multiple initial slices meet the preset conditions.
[0106] In some embodiments, the retrieval module 203 is further configured to convert multiple text slices into query vectors, and use vector similarity calculation to determine multiple first initial candidate vectors in the vector knowledge base;
[0107] Multiple initial candidate vectors are filtered based on a similarity threshold to obtain multiple second initial candidate vectors;
[0108] Multiple candidate vectors are scored in multiple dimensions to determine multiple candidate vectors.
[0109] In some embodiments, a construction module is also included, configured to collect data source files and perform data cleaning and formatting on the data source files;
[0110] The data source file, after data cleaning and formatting, is sliced to obtain multiple data source slices;
[0111] Vectorize multiple data source slices, converting them into multiple data source vectors;
[0112] The vector knowledge base is constructed based on multiple data source vectors.
[0113] In some embodiments, the building module is also configured to build a vector knowledge base retrieval interface;
[0114] The vector knowledge base retrieval interface is used to perform retrieval tests, and the vector knowledge base is deployed in response to passing the retrieval test.
[0115] In some embodiments, the evaluation module 204 is further configured to construct cue words based on multiple text slices and multiple candidate vectors;
[0116] Multiple text slices, multiple candidate vectors, and prompt words are input into the large language model to generate analysis results;
[0117] The analysis results are format-validated. In response to passing the format validation, a multi-dimensional evaluation is performed based on the analysis results, and the data security management evaluation results are output through the large language model.
[0118] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0119] The apparatus described above is used to implement the corresponding data security management and evaluation method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0120] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the data security management and evaluation method described in any of the above embodiments.
[0121] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0122] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0123] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0124] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0125] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0126] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0127] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0128] The electronic devices described above are used to implement the corresponding data security management and evaluation methods in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0129] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the data security management and evaluation method as described in any of the above embodiments.
[0130] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0131] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the data security management and evaluation method as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0132] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0133] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0134] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0135] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A data security management evaluation method characterized by comprising: The method comprises the following steps: Converting the to-be-checked file by a multi-modal large model to obtain a text file; Slicing the text file according to logical structure information of the text file to obtain a plurality of text slices; Based on the plurality of text slices, searching and matching in a pre-constructed vector knowledge base to determine a plurality of candidate vectors; Inputting the plurality of text slices and the plurality of candidate vectors into a large language model to output a data security management evaluation result by the large language model.
2. The method of claim 1, wherein, The slicing the text file according to the logical structure information of the text file to obtain a plurality of text slices comprises the following steps: Slicing the text file according to the logical structure information of the text file to obtain a plurality of initial slices; Determining whether the slice length of the plurality of initial slices meets a preset condition; In response to the slice length of the plurality of initial slices meeting the preset condition, performing semantic integrity checking on the plurality of initial slices; In response to the plurality of initial slices passing the semantic integrity checking, performing standardization processing on the plurality of initial slices to generate a plurality of text slices.
3. The method of claim 2, wherein, Further comprising: In response to the slice length of the plurality of initial slices not meeting the preset condition and there being an initial slice with a slice length greater than a preset threshold, performing semantic boundary subdivision on the text file to make the slice length of each initial slice meet the preset condition; In response to the slice length of the plurality of initial slices not meeting the preset condition and there being an initial slice with a slice length less than a preset threshold, merging adjacent slices to make the slice length of the plurality of initial slices meet the preset condition.
4. The method of claim 1, wherein, The searching and matching in the pre-constructed vector knowledge base based on the plurality of text slices to determine a plurality of candidate vectors comprises the following steps: Converting the plurality of text slices into query vectors, and determining a plurality of first initial candidate vectors in the vector knowledge base by vector similarity calculation; Filtering the plurality of initial candidate vectors based on a similarity threshold to obtain a plurality of second initial candidate vectors; Performing multi-dimensional scoring on the plurality of second initial candidate vectors to determine a plurality of candidate vectors.
5. The method of claim 1, wherein, A method for constructing a vector knowledge base comprises the following steps: Collecting a data source file and performing data cleaning and format processing on the data source file; Performing document slicing on the data source file after data cleaning and format processing to obtain a plurality of data source slices; Performing vectorization processing on the plurality of data source slices to convert them into a plurality of data source vectors; Constructing the vector knowledge base based on the plurality of data source vectors.
6. The method of claim 5, wherein, Further comprising: Constructing a vector knowledge base retrieval interface; Performing retrieval testing through the vector knowledge base retrieval interface, and in response to passing the retrieval testing, deploying the vector knowledge base.
7. The method of claim 1, wherein, The inputting the plurality of text slices and the plurality of candidate vectors into a large language model to output a data security management evaluation result by the large language model comprises the following steps: Constructing prompt words according to the plurality of text slices and the plurality of candidate vectors; Inputting the plurality of text slices, the plurality of candidate vectors and the prompt words into the large language model to generate an analysis result; Performing format checking on the analysis result, and in response to passing the format checking, performing multi-dimensional evaluation according to the analysis result, and outputting the data security management evaluation result by the large language model.
8. A data security management evaluation device, characterized by comprising: The method comprises the following steps: The conversion module is configured to convert the file to be checked by a multi-modal large model to obtain a text file; The slicing module is configured to slice the text file according to the logical structure information of the text file to obtain a plurality of text slices; The retrieval module is configured to perform retrieval matching in a pre-constructed vector knowledge base based on the plurality of text slices to determine a plurality of candidate vectors; The evaluation module is configured to input the plurality of text slices and the plurality of candidate vectors into a large language model, and output a data security management evaluation result by the large language model.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method of any one of claims 1 to 7 when executing the program.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to make the computer execute the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic auditing method for security management system of information security level protection evaluation
CN117952531A
Medical question and answer method and system based on large language model and storage medium
CN118152533A
PortRAG-based intelligent question-answering system for port safety production management knowledge
CN118861236A
Digital-intelligent power regulation and control question answering method and system based on large language model
CN119003742A
System comparison system and method based on multi-granularity large language model
CN119513621A