Large model fine-grained access control method and device for data classification and classification
By introducing permission annotation model and access policy engine in the big model question and answer system, permission annotation and compliance detection of user problems and documents is solved, and the problem of intricate access control in the existing system is achieved, and fine-grained access control and security improvement of the large language model is achieved.
Patent Information
- Application Number
- CN202510252414.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing big model question and answer system is difficult to effectively implement fine-grained permission access control, resulting in users with different permissions being able to easily obtain mismatched information, and the data hierarchical and classification tag system is complex, and the access decision-making process is cumbersome.
By obtaining user information and user problems, using the pre-built permission annotation model for text permission annotation, obtaining the permission hierarchical classification annotation results of user problems, and conducting compliance detection through the access policy engine to obtain access policies based on user problems. For compliance issues, relevance document search and large language model access are carried out to generate a holistic reply; for non-compliance issues, directly output rejection answers.
The fine-grained access control of the large language model is realized, ensuring that users with different permissions can only obtain information matching the permissions, improve the accuracy and efficiency of access, and enhance the security of the system.
Smart Images

Figure CN119761525B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large-model question-answering technology, and in particular to a large-model fine-grained access control method and device for data grading and classification. Background Art
[0002] The traditional retrieval enhancement generates a large model question-answering mode. The large model directly answers user questions based on the references retrieved from the external document library. However, different documents in the external document library usually have different permission levels and types, and users with different permission types can easily obtain information that does not match their permissions by asking questions.
[0003] In the existing large-model basic security algorithms, the method of building system prompt words and fine-tuning the alignment model is usually adopted to prevent the model from outputting toxic, harmful, illegal, and biased responses. However, the needs of users with different permissions who expect the model to give answers that match their permissions have not yet been properly met. That is, not all sensitive information should be treated as harmful information and blocked, but answers should be generated conditionally within a controllable range. At the same time, due to the complex labeling system of data classification and the cumbersome access decision-making process, it is difficult for traditional retrieval enhancement generation models to achieve this goal end-to-end. Therefore, it is particularly necessary to study the construction of a large-model permission fine-grained access control question-and-answer model. Summary of the invention
[0004] Based on this, it is necessary to provide a large-model fine-grained access control method and device for data classification and grading to address the above technical issues.
[0005] A large-model fine-grained access control method for data classification and categorization, the method comprising:
[0006] Obtain user information and corresponding user questions, input the user questions into a pre-built permission annotation model for text permission annotation, obtain the hierarchical and classified annotation results of the user question permissions, and input the hierarchical and classified annotation results of the user question permissions and the user access rights in the user information into the access policy engine for compliance detection, determine whether the user question is compliant and obtain the access policy based on the user question;
[0007] For compliance questions or partial compliance questions, they are first input into the two-step retrieval model for related document retrieval to obtain the reference document set corresponding to the user question, and then the user question, the access policy based on the user question and the reference document set are input into the large language model for access to obtain the overall answer to the user question; for non-compliance questions, after inputting into the large language model for access, a response of refusing to answer is directly output;
[0008] Input each reference document in the reference document collection into the permission annotation model for text permission annotation, obtain the hierarchical classification annotation results of each reference document permission, and input the hierarchical classification annotation results of each reference document permission and the user access rights into the access policy engine for compliance detection, determine whether each reference document is compliant and obtain the access policy based on the reference document;
[0009] Each reference document and the corresponding access strategy based on the reference document are input into the large language model for access, and the answer to each reference document is obtained. By sorting out the user questions and the answers to each reference document, the final answer to the user question output by the large language model is obtained.
[0010] In one embodiment, taking user questions as an example, the process of the permission annotation model for text permission annotation includes: building a hierarchical classification tag library And the corresponding classification sample library ,and ;in, Represents the mapping relationship from samples to hierarchical classification labels, Indicates the first m tags, Represents the first n documents, each label corresponds to several documents;
[0011] For user input questions First, based on the vector retrieval model, in the hierarchical classification sample library The vector search method is used to obtain the A similar document set consisting of the two most similar documents , and get the labels corresponding to the two most similar documents ;in, and express The two most similar documents are Represents a vector retrieval model;
[0012] Then, the K nearest neighbor algorithm is used to search In the classification label library , and obtain the documents corresponding to the neighboring tags, expand the number of documents in the similar document set, and obtain the reference example ;
[0013] Reference examples , User issues And the prompt template Input the large language model for permission prediction and entity extraction to obtain user questions The hierarchical classification and labeling results of permissions are expressed as
[0014] ;
[0015] in, Represents the predicted user question The label category of the permission, Indicates the extracted user questions The entity set of represents the annotation model based on the large language model, Represents the annotation results, namely, label categories and entity sets.
[0016] In one embodiment, the hierarchical classification and labeling results of the user's problem permissions and the user's access rights in the user information are input into the access policy engine for compliance detection, to determine whether the user's problem is compliant and obtain an access policy based on the user's problem, including:
[0017] User Questions Permissions label categories, user issues The entity collection and user access rights are input into the access policy engine for compliance detection, and the user problem is judged to be compliant, partially compliant, or non-compliant, and the access policy based on the user problem is obtained. ;in, Indicates user information. Represents the user's access rights matrix, including the user's organization, business, field, level, position, and tasks. Indicates the user's basic personal information, including the user's name and gender, subscript k Indicates k users.
[0018] In one embodiment, for a compliance question or part of a compliance question, the question is first input into a two-step search model for related document search to obtain a reference document set corresponding to the user question, and then the user question, the access policy based on the user question, and the reference document set are input into a large language model for access to obtain an overall answer to the user question, including:
[0019] If the user has any questions Compliant or partially compliant, user issues Input the two-step retrieval model. The two-step retrieval model first uses a vector retrieval model with smaller parameters to perform a preliminary search and obtain the user's question The 200 most relevant documents are then re-ranked using a vector retrieval model with larger parameters to obtain the user's question The 10 most relevant documents make up the user's question A collection of reference documents , expressed as
[0020] ;
[0021] in, represents the vector retrieval model for re-ranking, represents the vector retrieval model used for preliminary retrieval, Represents the document set obtained by the initial retrieval. Represents the collection of all documents in the document library;
[0022] User Questions , access policy based on user questions , Reference Document Collection And the prompt template Input the large language model to access and get the user's question Overall response , expressed as
[0023] ;
[0024] in, Represents a large language model.
[0025] In one embodiment, each reference document in the reference document set is input into a permission annotation model for text permission annotation, and a hierarchical and classified annotation result of the permission of each reference document is obtained. The hierarchical and classified annotation result of the permission of each reference document and the user access permission are input into an access policy engine for compliance detection, and whether each reference document is compliant is determined and an access policy based on the reference document is obtained, including:
[0026] Reference document collection Each reference document in the document is input into the permission annotation model for text permission annotation, and the label category of each reference document permission and the corresponding entity set are obtained;
[0027] Input the label category of each reference document permission, the entity set corresponding to each reference document, and the user access rights into the access policy engine for compliance detection, determine whether each reference document is compliant or non-compliant, and obtain the access policy based on the reference document ;in, Indicates user information. Represents the user's access rights matrix, including the user's organization, business, field, level, position, and tasks. Indicates the user's basic personal information, including the user's name and gender, subscript k Indicates k Users; Indicates j Reference documents.
[0028] In one embodiment, each reference document and a corresponding access policy based on the reference document are input into a large language model for access, and a response to each reference document is obtained, including:
[0029] If you refer to the document Compliance, user issues , Reference Documents , access strategy based on reference documents And the prompt template Enter the large language model to access and obtain reference documents Reply , expressed as ;
[0030] in, Represents a large language model;
[0031] If you refer to the document If it is not compliant, after entering the large language model for access, it directly outputs a response that refuses to answer.
[0032] In one embodiment, by collating the user question and the answer of each reference document, a final answer to the user question output by the large language model is obtained, including:
[0033] Organize user questions and each reference document Reply , and use the prompt template , get the user question output by the large language model The final answer is expressed as
[0034] .
[0035] In one embodiment, the method further comprises:
[0036] The large language model is fine-tuned based on the LoRA algorithm. The original parameters of the large language model are: , these parameters are fixed during training, and the trainable parameters of the large language model are weights , where the matrix ,matrix ; At initialization, the matrix Initialized by Gaussian function, the matrix It is initialized to zero, so that the bypass does not affect the original language model before training begins, that is, the parameter change is 0; for this weight Input For example, the output is as follows:
[0037] ;
[0038] in, represents the set of real numbers, , and Represent the parameter dimensions respectively.
[0039] In one embodiment, the loss function of the large language model is a negative log-likelihood loss, expressed as ;
[0040] in, represents the value of negative log-likelihood loss, Indicates the length of the input text. is the current word, Represents input text, Represents the text before the current word. represents the model parameters that are fine-tuned, Indicates that the current word is predicted based on the text and model parameters before the current word probability.
[0041] A large-model fine-grained access control device for data classification and classification, the device comprising:
[0042] A user question compliance detection module is used to obtain user information and corresponding user questions, input the user questions into a pre-built permission annotation model for text permission annotation, obtain hierarchical and classified annotation results of user question permissions, and input the hierarchical and classified annotation results of user question permissions and user access rights in the user information into an access policy engine for compliance detection, determine whether the user questions are compliant and obtain access policies based on the user questions;
[0043] The reference document retrieval and overall response module is used to first input compliance issues or partial compliance issues into the two-step retrieval model for related document retrieval to obtain a reference document set corresponding to the user's question, and then input the user's question, the access policy based on the user's question, and the reference document set into the large language model for access to obtain the overall response to the user's question; for non-compliant issues, after inputting the large language model for access, directly output a response of refusing to answer;
[0044] A reference document compliance detection module is used to input each reference document in the reference document set into the permission annotation model for text permission annotation, obtain the hierarchical classification annotation results of each reference document permission, and input the hierarchical classification annotation results of each reference document permission and user access rights into the access policy engine for compliance detection, determine whether each reference document is compliant and obtain an access policy based on the reference document;
[0045] The answer output module is used to input each reference document and the corresponding access strategy based on the reference document into the large language model for access, obtain the answer to each reference document, and obtain the final answer to the user question output by the large language model by sorting out the user questions and the answers to each reference document.
[0046] Compared with the prior art, the large-model fine-grained access control method and device for data classification and classification have the following beneficial effects:
[0047] 1. Before accessing the large language model, the user questions and reference documents are graded and classified based on the permission annotation model. This allows the system to retrieve documents with matching permissions for different questions of users with different permissions, avoiding information redundancy or mismatch caused by ambiguous permissions. This improves the accuracy and efficiency of user information acquisition when accessing the large language model, and implements fine-grained access control of the large language model.
[0048] 2. Compliance checks are performed on user questions and reference documents based on the access policy engine, which can prevent the large language model from outputting illegal and harmful responses, thereby improving the security of access to the large language model.
[0049] 3. The large language model not only gives an overall response to the user's question, but also responds to each reference document in the reference document set corresponding to the user's question, thus avoiding confusion and erroneous output that may occur in a one-time response, and improving the accuracy of the model's question and answer. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 It is a flowchart of a large-model fine-grained access control method for data classification and classification in one embodiment;
[0051] Figure 2 A schematic diagram of a framework of a large-model fine-grained access control method for data classification and classification in one embodiment;
[0052] Figure 3 A schematic diagram of a framework for text permission annotation in a permission annotation model in an embodiment. DETAILED DESCRIPTION
[0053] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] In one embodiment, Figure 1 and Figure 2 As shown, a large-model fine-grained access control method for data classification and classification is provided, comprising the following steps:
[0055] Step 1: Obtain user information and corresponding user questions, input user questions into a pre-built permission annotation model for text permission annotation, obtain hierarchical and classified annotation results of user question permissions, and input the hierarchical and classified annotation results of user question permissions and user access rights in user information into the access policy engine for compliance detection to determine whether the user question is compliant and obtain access policies based on the user question.
[0056] Wherein, step 1 specifically includes the following steps:
[0057] First, get the user information And the corresponding user questions ;in, Represents the user's access rights matrix, including the user's organization, business, field, level, position, and tasks. Indicates the basic personal information of the user, including the user's name and gender, etc. k Indicates k users.
[0058] Then, the user question Input the pre-built permission annotation model to perform text permission annotation. Figure 3 As shown in the figure, the permission labeling model takes into account the characteristics of data classification labels, which are often large in number and fine in granularity. In particular, for different target units, there are often different classification systems. At the same time, there is a large uncertainty in the number of samples that can be used for training and the degree of balance. Specifically, the generative language model is used in conjunction with the small sample prompt learning method to label the permissions of text data such as user questions or reference documents. When the number of labels is , each class label sampling If there are Considering that too many sample examples will greatly reduce the computational efficiency of the permission labeling model, the KNN (K-Nearest Neighbors) algorithm is specifically used to select the examples most relevant to the text data to be labeled to provide permission labeling model prompts. Taking user questions as an example, the process of text permission labeling by the permission labeling model specifically includes:
[0059] Build a hierarchical classification tag library And the corresponding classification sample library ,and ;in, Represents the mapping relationship from samples to hierarchical classification labels, Indicates the first m tags, Represents the firstn documents, and each label corresponds to several documents.
[0060] For user input questions First, based on the vector retrieval model, in the hierarchical classification sample library The vector search method is used to obtain the A similar document set consisting of the two most similar documents , and get the labels corresponding to the two most similar documents ;in, and express The two most similar documents are Represents a vector retrieval model;
[0061] Then, the K nearest neighbor algorithm is used to search In the classification label library , and obtain the documents corresponding to the neighboring tags, expand the number of documents in the similar document set, and obtain the reference example ;
[0062] Reference examples , User issues And the prompt template Input the large language model for permission prediction and entity extraction to obtain user questions The hierarchical classification and labeling results of permissions are expressed as
[0063] ;
[0064] in, Represents the predicted user question The label category of the permission, Indicates the extracted user questions The entity set of represents the annotation model based on the large language model, Represents the annotation results, namely, label categories and entity sets.
[0065] Getting user question permissions After the hierarchical classification and labeling results, the user questions are further Permissions label categories, user issues The entity collection and user access rights are input into the access policy engine for compliance detection, and the user problem is judged to be compliant, partially compliant, or non-compliant, and the access policy based on the user problem is obtained. .
[0066] Furthermore, when the number of labeled samples is large and the quality is high, this method also fine-tunes the large language model based on the LoRA (low-rank adaptation of LLM) algorithm. The original parameters of the large language model are , these parameters are fixed during training, and the trainable parameters of the large language model are weights , where the matrix ,matrix ; At initialization, the matrix Initialized by Gaussian function, the matrix It is initialized to zero, so that the bypass does not affect the original language model before training begins, that is, the parameter change is 0; for this weight Input For example, the output is as follows: ;
[0067] in, represents the set of real numbers, , and Represent the parameter dimensions respectively.
[0068] The loss function of the large language model is the negative log-likelihood loss, expressed as
[0069] ;
[0070] in, represents the value of negative log-likelihood loss, Indicates the length of the input text. is the current word, Represents input text, Represents the text before the current word. represents the model parameters that are fine-tuned, Indicates that the current word is predicted based on the text and model parameters before the current word probability.
[0071] Step 2: For compliance issues or partial compliance issues, first input them into the two-step retrieval model for relevant document retrieval to obtain the reference document set corresponding to the user question, and then input the user question, the access policy based on the user question and the reference document set into the large language model for access to obtain the overall answer to the user question; for non-compliance issues, after inputting into the large language model for access, directly output a response of refusing to answer.
[0072] Wherein, step 2 specifically includes the following steps:
[0073] If the user has any questions Compliant or partially compliant, user issues Input the two-step retrieval model. In order to balance the efficiency and accuracy of the retrieval, the two-step retrieval model first uses a vector retrieval model with smaller parameters to perform a preliminary retrieval to obtain the user's question The 200 most relevant documents are then re-ranked using a vector retrieval model with larger parameters to obtain the user's question The 10 most relevant documents make up the user's question A collection of reference documents , expressed as
[0074] ;
[0075] in, represents the vector retrieval model for re-ranking, represents the vector retrieval model used for preliminary retrieval, Represents the document set obtained by the initial retrieval. Represents the collection of all documents in the document library;
[0076] User Questions , access policy based on user questions , Reference Document Collection And the prompt template Input the large language model to access and get the user's question Overall response , expressed as ;
[0077] in, Represents a large language model.
[0078] If the user has any questions Non-compliant, directly access the large language model and output a response that refuses to answer.
[0079] Step 3: Input each reference document in the reference document collection into the permission annotation model for text permission annotation, obtain the hierarchical and classified annotation results of each reference document's permissions, and input the hierarchical and classified annotation results of each reference document's permissions and user access rights into the access policy engine for compliance detection to determine whether each reference document is compliant and obtain an access policy based on the reference document.
[0080] Wherein, step 3 specifically includes the following steps:
[0081] First, collect the reference documents Each reference document in the document is input into the permission annotation model for text permission annotation, and the label category of each reference document permission and the corresponding entity set are obtained.
[0082] Then, the label category of each reference document permission, the entity set corresponding to each reference document, and the user access rights are input into the access policy engine for compliance detection to determine whether each reference document is compliant or non-compliant, and obtain the access policy based on the reference document. ; Indicates j Reference documents.
[0083] It can be understood that in the above steps 1 to 3, by using the permission annotation model to grade and classify the text permissions of user questions and reference documents, it is possible to retrieve documents with permissions matching different questions of users with different permissions, avoid information redundancy or mismatch caused by permission ambiguity, improve the accuracy and efficiency of user information acquisition when accessing the large language model, and realize fine-grained access control of the large language model. The access policy engine is used to perform compliance checks on user questions and reference documents respectively, and filter and process illegal questions or documents, which improves the security of access to the large language model.
[0084] Step 4: Input each reference document and the corresponding access strategy based on the reference document into the large language model for access, obtain the answer to each reference document, and obtain the final answer to the user question output by the large language model by sorting out the user questions and the answers to each reference document.
[0085] Wherein, step 4 specifically includes the following steps:
[0086] If you refer to the document Compliance, user issues , Reference Documents , access strategy based on reference documents And the prompt template Enter the large language model to access and obtain reference documents Reply , expressed as
[0087] ;
[0088] in, Represents a large language model.
[0089] If you refer to the document If it is not compliant, after entering the large language model for access, it directly outputs a response that refuses to answer.
[0090] It is understandable that since the access policies for each reference document are not necessarily exactly the same, if the large language model generates a response all at once during the response process, it is easy to cause confusion and erroneous output. Therefore, responding to each reference document and its corresponding access policy can improve the accuracy of the large language model question answering.
[0091] Finally, organize user questions and each reference document Reply , and use the prompt template , get the user question output by the large language model The final answer is expressed as .
[0092] Furthermore, this application generates user permission access data based on the heuristic induction expansion of the big model, and verifies the effectiveness of the above-mentioned big model fine-grained access control method for data classification and classification by simulating user access situations, which specifically includes the following steps:
[0093] For different fields and institutions, the basic model of data classification is similar, but the specific system settings and management models are different. At the same time, since most of the data involve the secrets of the unit or institution and may involve the personal privacy of personnel, it is difficult to obtain comprehensive and systematic data to simulate the user's access behavior and optimize and verify the algorithm proposed in this application. Therefore, it is planned to use the method of synthetic data generation and use the generative large model to construct the user permission access data set. .
[0094] ① Design the core business of the target unit:
[0095] The core business of the target unit needs to be able to summarize the industry field, main business work and related unit characteristics of the unit, as a general inspiration to guide the whole process of data generation. Let the large language model randomly select a field and generate a text describing the core business and the name of the target unit. .
[0096] ②Structural unit structure:
[0097] The unit structure is usually represented as a tree structure , the second level refines the first level , expand different business working groups, and for different departments, it can be refined to the third level. Indicates the first level of the unit structure n Department.
[0098] ③ Set up staff composition:
[0099] Considering that the target unit has many departments and a complex structure, a loop traversal method is used to simulate and generate personnel information for each department at each level in turn. , including the person’s <name, position, gender, and business responsibilities>. A collection of personnel information.
[0100] ④Generate text data and set access policy:
[0101] For data types of different dimensions and levels in the data classification system, different types of texts are generated under different business departments to obtain text collections. , and mark the user's access rights for the text data;
[0102] ;
[0103] in, Represents the combination of different business units, Indicates the number of combined departments. In this application, The maximum value of is set to 2. The probability is 0.8, The probability is 0.2, that is, most of the data is generated by a single part, and a small part of the data is generated by two units together; Represents a combination of different tag types, where tags must include text type, business field, etc. represents the number of combined labels, represents the model that generates simulated data, Represents a combination of multiple departments. Represents a combination of multiple tags.
[0104] According to common sense, set each text data Access policy , the access strategy is a mapping of text data to personnel groups ,in .
[0105] ⑤Simulate user access:
[0106] For each piece of text data , generate 3 questions at different levels and angles based on the text content According to the access policy, the users who can legally access this question are , its complement For each question, , randomly select 1 user Composition Compliance Visit , randomly select 1 user Composition of non-compliant access , which eventually constitutes the user access dataset .
[0107] Since the dataset The number of two labels is balanced, so when verifying the algorithm, accuracy, precision, and recall can be used for calculation.
[0108] In one embodiment, a large-model fine-grained access control device for data classification and classification is provided, comprising:
[0109] A user question compliance detection module is used to obtain user information and corresponding user questions, input the user questions into a pre-built permission annotation model for text permission annotation, obtain hierarchical and classified annotation results of user question permissions, and input the hierarchical and classified annotation results of user question permissions and user access rights in the user information into an access policy engine for compliance detection, determine whether the user questions are compliant and obtain access policies based on the user questions;
[0110] The reference document retrieval and overall response module is used to first input compliance issues or partial compliance issues into the two-step retrieval model for related document retrieval to obtain a reference document set corresponding to the user's question, and then input the user's question, the access policy based on the user's question, and the reference document set into the large language model for access to obtain the overall response to the user's question; for non-compliant issues, after inputting the large language model for access, directly output a response of refusing to answer;
[0111] A reference document compliance detection module is used to input each reference document in the reference document set into the permission annotation model for text permission annotation, obtain the hierarchical classification annotation results of each reference document permission, and input the hierarchical classification annotation results of each reference document permission and user access rights into the access policy engine for compliance detection, determine whether each reference document is compliant and obtain an access policy based on the reference document;
[0112] The answer output module is used to input each reference document and the corresponding access strategy based on the reference document into the large language model for access, obtain the answer to each reference document, and obtain the final answer to the user question output by the large language model by sorting out the user questions and the answers to each reference document.
[0113] For the specific definition of the large-model fine-grained access control device for data classification and classification, please refer to the definition of the large-model fine-grained access control method for data classification and classification mentioned above, which will not be repeated here. Each module in the above-mentioned large-model fine-grained access control device for data classification and classification can be implemented in whole or in part through software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0114] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0115] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A large-model fine-grained access control method for data classification, characterized in that: The method comprises: Obtain user information and corresponding user questions, input the user questions into a pre-built permission annotation model for text permission annotation, obtain hierarchical and classified annotation results of user question permissions, and input the hierarchical and classified annotation results of user question permissions and user access rights in the user information into an access policy engine for compliance detection, determine whether the user questions are compliant and obtain an access policy based on the user questions; For compliance questions or partial compliance questions, they are first input into the two-step retrieval model for related document retrieval to obtain a reference document set corresponding to the user question, and then the user question, the access policy based on the user question and the reference document set are input into the large language model for access to obtain the overall answer to the user question; for non-compliance questions, after inputting into the large language model for access, a reply of refusing to answer is directly output; Input each reference document in the reference document set into the permission annotation model for text permission annotation, obtain the hierarchical classification annotation result of the permission of each reference document, and input the hierarchical classification annotation result of the permission of each reference document and the user access permission into the access policy engine for compliance detection, determine whether each reference document is compliant and obtain the access policy based on the reference document; Each reference document and the corresponding access strategy based on the reference document are input into the large language model for access, and the answer to each reference document is obtained. By sorting out the user questions and the answers to each reference document, the final answer to the user question output by the large language model is obtained.
2. The method according to claim 1, characterized in that Taking user questions as an example, the process of text permission annotation by the permission annotation model includes: Build a hierarchical classification tag library And the corresponding classification sample library ,and ;in, Represents the mapping relationship from samples to hierarchical classification labels, Indicates the first m tags, Represents the first n documents, each label corresponds to several documents; For user input questions First, based on the vector retrieval model, in the hierarchical classification sample library The vector search method is used to obtain the A similar document set consisting of the two most similar documents , and get the labels corresponding to the two most similar documents ;in, and express The two most similar documents are Represents a vector retrieval model; Then, the K nearest neighbor algorithm is used to search In the classification label library , and obtain the documents corresponding to the neighboring tags, expand the number of documents in the similar document set, and obtain a reference example ; The reference example , User issues And the prompt template Input the large language model for permission prediction and entity extraction to obtain user questions The hierarchical classification and labeling results of permissions are expressed as ; in, Represents the predicted user question The label category of the permission, Indicates the extracted user questions The entity set of represents the annotation model based on the large language model, Represents the annotation results, namely, label categories and entity sets.
3. The method according to claim 2, characterized in that By inputting the hierarchical classification and annotation results of user problem permissions and the user access rights in the user information into the access policy engine for compliance detection, it is determined whether the user problem is compliant and an access policy based on the user problem is obtained, including: User Questions Permissions label categories, user issues The entity collection and user access rights are input into the access policy engine for compliance detection, and the user's problem is judged to be compliant, partially compliant, or non-compliant, and the access policy based on the user's problem is obtained. ;in, Indicates user information. Represents the user's access rights matrix, including the user's organization, business, field, level, position, and tasks. Indicates the user's basic personal information, including the user's name and gender, subscript k Indicates k users.
4. The method according to claim 3, characterized in that For compliance questions or partial compliance questions, they are first input into the two-step retrieval model for related document retrieval to obtain a reference document set corresponding to the user question. Then, the user question, the access policy based on the user question, and the reference document set are input into the large language model for access to obtain the overall answer to the user question, including: If the user has any questions Compliant or partially compliant, user issues Input a two-step search model, which first uses a vector search model with smaller parameters to perform a preliminary search to obtain the user's question The 200 most relevant documents are then re-ranked using a vector retrieval model with larger parameters to obtain the user's question The 10 most relevant documents make up the user's question A collection of reference documents , expressed as ; in, represents the vector retrieval model for re-ranking, represents the vector retrieval model used for preliminary retrieval, Represents the document set obtained by the initial retrieval. Represents the collection of all documents in the document library; User Questions , access policy based on user questions , Reference Document Collection And the prompt template Input the large language model to access and get the user's question Overall response , expressed as ; in, Represents a large language model.
5. The method according to claim 4, characterized in that Input each reference document in the reference document set into the permission annotation model for text permission annotation, obtain the hierarchical classification annotation result of the permission of each reference document, and input the hierarchical classification annotation result of the permission of each reference document and the user access permission into the access policy engine for compliance detection, determine whether each reference document is compliant and obtain the access policy based on the reference document, including: Reference document collection Each reference document in the document is input into the permission annotation model for text permission annotation, and the label category of each reference document permission and the corresponding entity set are obtained; Input the label category of each reference document permission, the entity set corresponding to each reference document, and the user access rights into the access policy engine for compliance detection, determine whether each reference document is compliant or non-compliant, and obtain the access policy based on the reference document ;in, Indicates user information. Represents the user's access rights matrix, including the user's organization, business, field, level, position, and tasks. Indicates the user's basic personal information, including the user's name and gender, subscript k Indicates k Users; Indicates j Reference documents.
6. The method according to claim 5, characterized in that Each reference document and the corresponding access policy based on the reference document are input into the large language model for access, and the response of each reference document is obtained, including: If you refer to the document Compliance, user issues , Reference Documents , access strategy based on reference documents And the prompt template Enter the large language model to access and obtain reference documents Reply , expressed as ; in, Represents a large language model; If you refer to the document If it is not compliant, after inputting the large language model for access, it directly outputs a response that refuses to answer.
7. The method according to claim 6, characterized in that By collating the user questions and the answers to each reference document, the final answers to the user questions output by the large language model are obtained, including: Organize user questions and each reference document Reply , and use the prompt template , get the user question output by the large language model The final answer is expressed as 。 8. The method according to claim 7, characterized in that The method further includes: fine-tuning the large language model based on the LoRA algorithm, where the original parameters of the large language model are , these parameters are fixed during training, and the trainable parameters of the large language model are weights , where the matrix ,matrix ; At initialization, the matrix Initialized by Gaussian function, the matrix Initialized to zero, so that the bypass does not affect the original language model before training begins, that is, the parameter change is 0; for this weight Input For example, the output is as follows: ; in, represents the set of real numbers, , and Represent the parameter dimensions respectively.
9. The method according to claim 8, characterized in that The loss function of the large language model is the negative log-likelihood loss, expressed as ; in, represents the value of negative log-likelihood loss, Indicates the length of the input text. is the current word, Represents input text, Represents the text before the current word. represents the model parameters that are fine-tuned, Indicates that the current word is predicted based on the text and model parameters before the current word probability.
10. A large-model fine-grained access control device for data classification, characterized in that: The device comprises: A user question compliance detection module is used to obtain user information and corresponding user questions, input the user questions into a pre-built permission annotation model for text permission annotation, obtain hierarchical and classified annotation results of user question permissions, and input the hierarchical and classified annotation results of user question permissions and user access rights in the user information into an access policy engine for compliance detection, determine whether the user questions are compliant and obtain access policies based on the user questions; The reference document retrieval and overall response module is used to first input compliance issues or partial compliance issues into the two-step retrieval model for related document retrieval to obtain a reference document set corresponding to the user's question, and then input the user's question, the access policy based on the user's question, and the reference document set into the large language model for access to obtain the overall response to the user's question; for non-compliant issues, after inputting the large language model for access, directly output a response of refusing to answer; A reference document compliance detection module is used to input each reference document in the reference document set into the permission annotation model for text permission annotation, obtain the hierarchical classification annotation results of each reference document permission, and input the hierarchical classification annotation results of each reference document permission and user access rights into the access policy engine for compliance detection, determine whether each reference document is compliant and obtain an access policy based on the reference document; The answer output module is used to input each reference document and the corresponding access strategy based on the reference document into the large language model for access, obtain the answer to each reference document, and obtain the final answer to the user question output by the large language model by sorting out the user questions and the answers to each reference document.
Citation Information
Patent Citations
Access control strategy generation method and system based on large language model
CN119337358A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1