Knowledge base construction and use method and device of trusted data space, equipment and medium

By building distributed clusters in trusted data spaces, based on sensitivity hierarchy and vectorization processing, the problem of limited coverage of a single institution knowledge base is solved, and comprehensive and accurate answers to cross-organization knowledge sharing and large models are achieved.

CN120338085AActive Publication Date: 2025-07-18HANGZHOU DBAPPSECURITY CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510799442.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The existing knowledge base is mainly limited to a single enterprise or institution, with limited knowledge coverage, and large models cannot effectively utilize undisclosed sensitive data, resulting in incomplete and inaccurate answers, and it is difficult to achieve cross-institutional knowledge sharing.

Method used

A distributed cluster is built through trusted data space, based on sensitivity hierarchy and vectorization processing, the knowledge bases of multiple subjects are interconnected, cross-institutional knowledge sharing and collaborative reasoning are realized, and different types of large models are used for secure isolation and computing resource optimization.

Benefits of technology

It realizes cross-organization knowledge sharing, improves the comprehensiveness and accuracy of large-scale model answers, and optimizes the allocation of computing resources while meeting compliance requirements, reducing the risk of sensitive information leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338085A_ABST
    Figure CN120338085A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge base construction and use method and device for a trusted data space, equipment and a medium, and relates to the field of artificial intelligence, and the method comprises the steps: carrying out the sensitive grading of a target knowledge block, obtaining a corresponding first grading result, carrying out the vectorization of the target knowledge block through employing different types of preset large models according to the first grading result, and obtaining a second grading result; obtaining target vector data, and constructing a target knowledge base based on the target vector data; vectorizing the knowledge query request to obtain a query request vector, and determining target data corresponding to the query request vector from each constructed target knowledge base; performing sensitivity grading on the target data to obtain a corresponding second grading result, and inputting the target data into a preset large model according to the second grading result to obtain a corresponding reasoning result; and determining a request result corresponding to the knowledge query request by using the reasoning result, and returning the request result to the knowledge user. Different knowledge bases are connected together, and the knowledge retrieval range is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and particularly to a method, apparatus, device, and medium for constructing and using a knowledge base in a trusted data space. Background Art

[0002] A trusted data space is a data circulation and utilization infrastructure that connects multiple parties based on consensus rules to achieve shared use of data resources. It is an application ecosystem for co-creating the value of data elements and an important carrier for supporting the construction of a national integrated data market. Through the trusted data space, knowledge bases between different organizations can be connected, and data sharing among all parties can be ensured under the consensus rules, thereby releasing data value.

[0003] Currently, the construction of knowledge bases based on large models is mainly limited to within a single enterprise or institution, with a limited knowledge coverage. The reasoning of large models relies on publicly available data on the Internet, but a large amount of unpublicized sensitive data (such as enterprise internal business data and classified materials) cannot circulate on the Internet, resulting in the large models being unable to accurately answer questions related to such data. Although combining knowledge bases can partially solve the problem, existing knowledge bases are only used within enterprises and cannot achieve cross-institutional knowledge sharing, restricting the comprehensiveness and accuracy of large model answers. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a method, apparatus, device, and medium for constructing and using a knowledge base in a trusted data space, which can interconnect the local knowledge bases of multiple entities into a distributed cluster based on the trusted data space, support global retrieval and collaborative reasoning, and achieve cross-institutional knowledge sharing. The specific solutions are as follows:

[0005] In a first aspect, this application provides a method for constructing and using a knowledge base in a trusted data space, including:

[0006] Obtain a target document uploaded by a data provider, and perform a chunking operation on the target document to obtain target knowledge chunks;

[0007] Based on a preset sensitivity detection method, perform sensitivity grading on the target knowledge chunks to obtain a first grading result corresponding to each target knowledge chunk, and use different types of preset large models to perform vectorization processing on the target knowledge chunks according to the first grading result to obtain target vector data, and construct a target knowledge base based on the target vector data;

[0008] Obtain a knowledge query request sent by a knowledge user, perform the vectorization processing on the knowledge query request to obtain a query request vector, and determine target data corresponding to the query request vector from the constructed target knowledge bases;

[0009] Perform the sensitivity grading on the target data based on the preset sensitivity detection method to obtain a second grading result corresponding to each piece of target data, and input the target data into different types of preset large models according to the second grading result to obtain corresponding inference results;

[0010] Use the inference results to determine the request result corresponding to the knowledge query request, and return the request result to the knowledge user.

[0011] Optionally, the operation of dividing the target document into blocks to obtain target knowledge blocks includes:

[0012] Detect the target document to determine each complete paragraph included in the target document, and determine any complete paragraph in the target document as the first target knowledge block;

[0013] Based on a preset block size, divide the document data in the target document except the first target knowledge block to obtain second target knowledge blocks;

[0014] Determine the target knowledge blocks corresponding to the target document according to the first target knowledge block and the second target knowledge blocks, and set indexes for each of the target knowledge blocks.

[0015] Optionally, the operation of performing sensitivity grading on the target knowledge blocks based on a preset sensitivity detection method to obtain a first grading result corresponding to each of the target knowledge blocks includes:

[0016] Perform sensitivity detection on the target knowledge blocks based on a preset sensitivity detection method;

[0017] If the current target knowledge block meets the preset high-sensitivity standard, determine the current target knowledge block as a high-sensitivity knowledge block; if the current target knowledge block meets the preset medium-sensitivity standard, determine the current target knowledge block as a medium-sensitivity knowledge block; if the current target knowledge block meets the preset low-sensitivity standard, determine the current target knowledge block as a low-sensitivity knowledge block.

[0018] Optionally, the operation of using different types of preset large models to perform vectorization processing on the target knowledge blocks according to the first grading result to obtain target vector data, and constructing a target knowledge base based on the target vector data includes:

[0019] For the high-sensitivity knowledge blocks, perform vectorization processing on the high-sensitivity knowledge blocks through a preset large model constructed based on a trusted execution environment cluster to obtain first vector data;

[0020] For the medium-sensitivity knowledge chunks, vectorize the medium-sensitivity knowledge chunks through a pre-trained large model deployed in the local infrastructure platform corresponding to the data provider to obtain second vector data;

[0021] For the low-sensitivity knowledge chunks, vectorize the low-sensitivity knowledge chunks through a pre-trained large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain third vector data;

[0022] Determine the target vector data corresponding to the target knowledge chunk based on the first vector data, the second vector data, and the third vector data;

[0023] Construct a target knowledge base based on the target knowledge chunk, the index corresponding to each target knowledge chunk, and the target vector data.

[0024] Optionally, determining the target data corresponding to the query request vector from the constructed target knowledge bases includes:

[0025] Determine the vector similarity between the query request vector and the target vector data in the constructed target knowledge bases based on a preset search method;

[0026] Sort the target vector data in descending order of the vector similarity, and determine a preset number of vector data from the sorted target vector data as the target vectors corresponding to the query request vector, and use the target knowledge chunks corresponding to the target vectors as the target data corresponding to the knowledge query request.

[0027] Optionally, inputting the target data into the pre-trained large models of different types according to the second classification result to obtain corresponding inference results includes:

[0028] Input the first target data representing high sensitivity in the second classification result into a pre-trained large model constructed based on a trusted execution environment cluster to obtain a corresponding first inference result;

[0029] Input the second target data representing medium sensitivity in the second classification result into a pre-trained large model deployed in the local infrastructure platform corresponding to the data provider to obtain a corresponding second inference result;

[0030] Input the third target data representing low sensitivity in the second classification result into a pre-trained large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain a corresponding third inference result.

[0031] Optionally, before inputting the target data into the preset large models of different types according to the second classification result to obtain corresponding inference results, the method further includes:

[0032] For the first target data representing high sensitivity in the second classification result, determine the first target credibility corresponding to the first target data respectively based on a preset credibility standard, construct a first target query request related to the content of the first target data according to the content of the first target data and the corresponding first target credibility, and input the first target query request and the first target data into the corresponding preset large model to obtain corresponding inference results;

[0033] For the second target data representing medium sensitivity in the second classification result, determine the second target credibility corresponding to the second target data respectively based on the preset credibility standard, construct a second target query request related to the content of the second target data according to the content of the second target data and the corresponding second target credibility, and input the second target query request and the second target data into the corresponding preset large model to obtain corresponding inference results;

[0034] For the third target data representing low sensitivity in the second classification result, determine the third target credibility corresponding to the third target data respectively based on the preset credibility standard, construct a third target query request related to the content of the third target data according to the content of the third target data and the corresponding third target credibility, and input the third target query request and the third target data into the corresponding preset large model to obtain corresponding inference results.

[0035] In a second aspect, the present application provides a device for constructing and using a knowledge base of a trusted data space, including:

[0036] A data acquisition module, configured to acquire a target document uploaded by a data provider and perform a chunking operation on the target document to obtain target knowledge chunks;

[0037] A knowledge base construction module, configured to perform sensitivity classification on the target knowledge chunks based on a preset sensitivity detection method to obtain a first classification result corresponding to each target knowledge chunk, perform vectorization processing on the target knowledge chunks by using preset large models of different types according to the first classification result to obtain target vector data, and construct a target knowledge base based on the target vector data;

[0038] A data determination module, configured to acquire a knowledge query request sent by a knowledge user, perform vectorization processing on the knowledge query request to obtain a query request vector, and determine target data corresponding to the query request vector from the constructed target knowledge bases;

[0039] A result acquisition module, configured to perform the sensitivity grading on the target data based on the preset sensitivity detection method to obtain a second grading result corresponding to each target data, and input the target data into different types of preset large models according to the second grading result to obtain corresponding inference results;

[0040] A request response module, configured to use the inference results to determine a request result corresponding to the knowledge query request and return the request result to the knowledge user.

[0041] In a third aspect, the present application provides an electronic device, including:

[0042] A memory, configured to store a computer program;

[0043] A processor, configured to execute the computer program to implement the foregoing method for constructing and using a knowledge base in a trusted data space.

[0044] In a fourth aspect, the present application provides a computer-readable storage medium, configured to store a computer program, wherein when the computer program is executed by a processor, the foregoing method for constructing and using a knowledge base in a trusted data space is implemented.

[0045] This application obtains the target document uploaded by the data provider, and performs a chunking operation on the target document to obtain target knowledge chunks; based on a preset sensitivity detection method, the target knowledge chunks are classified by sensitivity to obtain the first classification results corresponding to the target knowledge chunks. According to the first classification results, different types of preset large models are used to perform vectorization processing on the target knowledge chunks to obtain target vector data, and a target knowledge base is constructed based on the target vector data; a knowledge query request sent by the knowledge user is obtained, and the knowledge query request is vectorized to obtain a query request vector, and the target data corresponding to the query request vector is determined from the constructed target knowledge bases; based on the preset sensitivity detection method, the target data is classified by sensitivity to obtain the second classification results corresponding to the target data, and according to the second classification results, the target data is input into the different types of preset large models to obtain corresponding inference results; the inference results are used to determine the request result corresponding to the knowledge query request, and the request result is returned to the knowledge user. As can be seen from the above, this application classifies target knowledge chunks based on a preset sensitivity detection method. Through classification processing, knowledge chunks with different sensitivity levels can achieve secure isolation during the vectorization stage, meeting compliance requirements while optimizing the allocation of computing resources. At the same time, each data provider constructs a target knowledge base based on the vectorization results of local knowledge chunks, and all target knowledge bases form a distributed cluster through a trusted data space, supporting global retrieval and collaborative reasoning across data sources, improving the comprehensiveness and accuracy of large model answers. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.

[0047] Figure 1 It is a flowchart of a method for constructing and using a knowledge base in a trusted data space disclosed in the present application;

[0048] Figure 2 It is a schematic diagram of the process for a data provider to establish a local knowledge base disclosed in the present application;

[0049] Figure 3 It is a schematic diagram of the vectorization operation of knowledge chunks based on a large model deployed in a TEE cluster disclosed in the present application;

[0050] Figure 4 It is a schematic diagram of a knowledge base cluster disclosed in the present application;

[0051] Figure 5 A schematic diagram of a user query knowledge process disclosed in this application;

[0052] Figure 6 A schematic diagram of the structure of a knowledge base construction and usage device for a trusted data space disclosed in this application;

[0053] Figure 7 A schematic diagram of the structure of an electronic device disclosed in this application. Specific implementation manners

[0054] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0055] Currently, the construction of knowledge bases based on large models is mainly limited to within a single enterprise or institution, with a limited knowledge coverage. The inference of large models relies on publicly available Internet data, but a large amount of unpublicized sensitive data cannot circulate on the Internet, resulting in the large models being unable to accurately answer questions related to such data. Although combining with a knowledge base can partially solve the problem, the existing knowledge bases are only used within the enterprise and cannot achieve cross-institutional knowledge sharing, restricting the comprehensiveness and accuracy of the large model's answers. Therefore, this application provides a method for constructing and using a knowledge base for a trusted data space, which can interconnect the local knowledge bases of multiple entities to form a distributed cluster, support global retrieval and collaborative inference, and achieve cross-institutional knowledge sharing.

[0056] See Figure 1 As shown, the embodiments of this application disclose a method for constructing and using a knowledge base for a trusted data space, including:

[0057] Step S11: Obtain a target document uploaded by a data provider, and perform a chunking operation on the target document to obtain target knowledge chunks.

[0058] In this embodiment, performing a chunking operation on the target document to obtain target knowledge chunks may include: First, detect the target document to determine each complete paragraph included in the target document, and determine any complete paragraph in the target document as the first target knowledge chunk; then, based on a preset chunk size, split the document data in the target document excluding the first target knowledge chunk to obtain second target knowledge chunks; finally, determine the target knowledge chunks corresponding to the target document according to the first target knowledge chunk and the second target knowledge chunks, and set indexes for each target knowledge chunk.

[0059] For example Figure 2As shown, after the data provider uploads the knowledge document locally to the access-connected machine, the access connector performs a chunking operation on the uploaded knowledge document based on the document splitting module. When chunking, it first checks whether the part to be chunked contains a complete paragraph. If it does, this paragraph is used as a knowledge chunk. Otherwise, the data is split according to a preset chunk size to obtain the segmented knowledge chunks, and an index is set for the last obtained knowledge chunk.

[0060] In this way, through the combined strategy of preferentially retaining complete paragraphs and splitting the remaining content according to the preset size, a reasonable segmentation of the knowledge chunks can be achieved.

[0061] Step S12: Based on a preset sensitivity detection method, perform sensitivity grading on the target knowledge chunks to obtain the first grading results corresponding to the target knowledge chunks. According to the first grading results, use different types of preset large models to perform vectorization processing on the target knowledge chunks to obtain target vector data, and construct a target knowledge base based on the target vector data.

[0062] In this embodiment, first, the sensitivity grading of the target knowledge chunks can be performed based on a preset sensitivity detection method to obtain the first grading results corresponding to the target knowledge chunks, which specifically may include: first, performing sensitivity detection on the target knowledge chunks based on the preset sensitivity detection method; if the current target knowledge chunk meets the preset high-sensitivity standard, the current target knowledge chunk is determined as a high-sensitivity knowledge chunk; if the current target knowledge chunk meets the preset medium-sensitivity standard, the current target knowledge chunk is determined as a medium-sensitivity knowledge chunk; if the current target knowledge chunk meets the preset low-sensitivity standard, the current target knowledge chunk is determined as a low-sensitivity knowledge chunk.

[0063] For example Figure 2 As shown, the sensitivity recognition module is locally called in the access connector. The sensitivity recognition module performs sensitivity detection on each knowledge chunk based on the preset sensitivity detection method, and the detection results are divided into three levels, namely high-sensitivity knowledge chunks, medium-sensitivity knowledge chunks, and low-sensitivity knowledge chunks.

[0064] Furthermore, for highly sensitive knowledge chunks, a pre-trained large model built based on a trusted execution environment cluster can be used to vectorize the highly sensitive knowledge chunks to obtain first vector data; for moderately sensitive knowledge chunks, a pre-trained large model deployed in the local infrastructure platform corresponding to the data provider (i.e., the large model mentioned later) can be used to vectorize the moderately sensitive knowledge chunks to obtain second vector data; for low-sensitivity knowledge chunks, a pre-trained large model provided by a platform other than the local infrastructure platform corresponding to the data provider can be used to vectorize the low-sensitivity knowledge chunks to obtain third vector data; then, based on the first vector data, the second vector data, and the third vector data, the target vector data corresponding to the target knowledge chunk is determined; finally, a target knowledge base is constructed based on the target knowledge chunk, the index corresponding to each target knowledge chunk, and the target vector data. Among them, when vectorizing the knowledge chunk, a pre-trained embedding model can be selected, such as the BERT model (Bidirectional Encoder Representations from Transformers, a pre-trained language model based on the Transformer architecture).

[0065] It should be noted that, based on technical means such as keyword matching, data label recognition, and content compliance analysis, the sensitivity of the target knowledge chunk can be graded to obtain highly sensitive knowledge chunks containing classified information, personal privacy, business secrets, etc., moderately sensitive knowledge chunks containing internal business data, industry-sensitive information, etc., and low-sensitivity knowledge chunks containing publicly available information, etc.

[0066] After grading the sensitivity of the target knowledge chunk, different types of pre-trained large models can be used to vectorize knowledge chunks with different sensitivities based on the differentiated secure computing environment provided by the trusted data space and according to the grading results corresponding to the target knowledge chunk. Specifically, for highly sensitive knowledge chunks, the large model deployed in the TEE (Trusted Execution Environment) cluster can be used to vectorize the knowledge chunks to ensure that the highly sensitive data is vectorized in an encrypted isolation environment, such as Figure 2 the "AI deployed based on TEE" module in Figure 3As shown in the figure, a large model constructed based on the TEE cluster method will be deployed on the infrastructure support platform, and the credibility of the large model is ensured by the hardware TEE. Before the user requests the large model, the user's query message will first be encrypted by the local user-side proxy program. The encrypted information will then be sent to the proxy program within the TEE through an encrypted channel. The proxy program within the TEE will send the decrypted message to the large model for inference, and after the large model completes the inference, it will encrypt the result and return it to the user-side proxy program. Through the above process, the inference task based on the large model can be completed while ensuring that the user's request data is not leaked. For medium-sensitivity knowledge blocks, the large model deployed on the local infrastructure platform of the data provider can be used for processing to complete the calculation within the trusted local area network, balancing security and efficiency, such as Figure 2 the "Localized Deployment of AI" module in Figure 2 For low-sensitivity knowledge blocks, vectorization processing can be directly performed by requesting an external large model API (Application Programming Interface), such as

[0067] the "Open AI API" module in

[0068] In this way, the above process constructs a "security fortress" for sensitive data through technologies such as the TEE cluster to prevent unauthorized access. At the same time, it dynamically schedules computing resources according to the sensitivity level to avoid resource waste of "processing low-sensitivity data with high security configurations", providing a secure and compliant underlying support for subsequent scenarios such as cross-institutional knowledge retrieval, and promoting the efficient circulation of knowledge elements in the trusted data space. Figure 2 As shown in the figure, after the vectorization processing of the knowledge block is completed, the knowledge block, knowledge block index, and knowledge block vector data can be stored in the local vectorized database, that is, the target knowledge base.

[0069] Step S13: Obtain the knowledge query request sent by the knowledge user, perform the vectorization processing on the knowledge query request to obtain a query request vector, and determine the target data corresponding to the query request vector from each of the constructed target knowledge bases.

[0070] In this embodiment, the target knowledge bases of all knowledge providers are used as a knowledge base cluster so that when the user queries, the optimal knowledge can be retrieved from the knowledge base cluster. As Figure 4 shown in the figure is a schematic diagram of the knowledge base cluster.

[0071] As Figure 5 shown in the figure, the user can first input a knowledge query request locally at the access connector to ask a question, and then the access connector will process the knowledge query request based on the vectorization processing process in step S12 to obtain a query request vector, and then synchronize the query request vector to each of the constructed target knowledge bases.

[0072] Further, the vector similarity between the query request vector and the target vector data in each of the constructed target knowledge bases can be determined based on a preset search method; then, the target vector data is sorted in descending order of vector similarity, and a preset number of vector data are determined from the sorted target vector data as the target vectors corresponding to the query request vector, and the target knowledge blocks corresponding to the target vectors are used as the target data corresponding to the knowledge query request.

[0073] For example Figure 5 As shown, after each node in the knowledge base cluster receives the query request vector, search methods such as Exact Nearest Neighbor (ENN) and Approximate Nearest Neighbor (ANN) can be used to recall knowledge. After the recall is completed, the knowledge is re-sorted in descending order according to the relevance, and the top N pieces of knowledge in the sorting result are used as the retrieved target data.

[0074] Step S14: Perform the sensitivity grading on the target data based on the preset sensitivity detection method to obtain the second grading results corresponding to the target data, and input the target data into the preset large models of different types according to the second grading results to obtain the corresponding inference results.

[0075] In this embodiment, the sensitivity recognition module in step S12 can be used to grade the retrieved target data to obtain the grading results corresponding to the target data. Then, the first target data indicating high sensitivity in the second grading results can be input into the preset large model constructed based on the trusted execution environment cluster to obtain the corresponding first inference result; the second target data indicating medium sensitivity in the second grading results can be input into the preset large model deployed in the local infrastructure platform corresponding to the data provider to obtain the corresponding second inference result; the third target data indicating low sensitivity in the second grading results can be input into the preset large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain the corresponding third inference result.

[0076] For example Figure 5 As shown, the target data with the grading results is sent to the AI agent service module deployed on the infrastructure support platform to use the AI agent service, and large models with different deployment methods are used for inference according to different sensitivity gradings. The specific process can refer to step S12.

[0077] It should be noted that adding a trusted AI inference module (including an AI agent service module, a prompt word sorting module, and a large model deployment module) to the trusted data space provides the basic ability to combine the trusted data space with AI. Among them, the large model deployment module can be divided into large models with three deployment methods according to different levels of trust, namely:

[0078] 1. For a high level of trust: a large model deployed based on a TEE cluster;

[0079] 2. For a medium level of trust: a large model deployed based on a local cluster;

[0080] 3. For a low level of trust: an API provided by an external large model service provider.

[0081] Among them, the "level of trust" refers to the reliability level of different large model deployment methods in terms of data privacy, security, and controllability. The TEE cluster deployment with a high level of trust ensures the security of the entire data processing link through hardware-level encryption technology (such as a trusted execution environment); the local cluster deployment with a medium level of trust relies on the institution's own infrastructure, which has certain controllability but weaker security than the TEE cluster deployment; the external API with a low level of trust has risks such as data outbound and protocol constraints due to relying on a third-party service provider, and has the lowest controllability.

[0082] Before inputting the target data into different types of preset large models according to the second classification result to obtain the corresponding inference results, it may further include: for the first target data representing high sensitivity in the second classification result, determining the first target credibility corresponding to the first target data respectively based on a preset credibility standard, constructing a first target query request related to the content of the first target data according to the content of the first target data and the corresponding first target credibility, and inputting the first target query request and the first target data into the corresponding preset large model to obtain the corresponding inference result; for the second target data representing medium sensitivity in the second classification result, determining the second target credibility corresponding to the second target data respectively based on a preset credibility standard, constructing a second target query request related to the content of the second target data according to the content of the second target data and the corresponding second target credibility, and inputting the second target query request and the second target data into the corresponding preset large model to obtain the corresponding inference result; for the third target data representing low sensitivity in the second classification result, determining the third target credibility corresponding to the third target data respectively based on a preset credibility standard, constructing a third target query request related to the content of the third target data according to the content of the third target data and the corresponding third target credibility, and inputting the third target query request and the third target data into the corresponding preset large model to obtain the corresponding inference result.

[0083] It should be noted that the quality of target data from different sources varies. For example, the credibility of enterprise internal documents is usually higher than that of publicly available online information. Therefore, based on the prompt word sorting module, by presetting credibility criteria such as source authority, content timeliness, and data integrity, the credibility corresponding to different contents in the target data can be evaluated. Then, the contents in the target data are sorted according to the credibility, and a complete question, that is, the target query request, is generated based on the sorted result and combined with the knowledge query request sent by the user. The target query request is used as the sorted prompt word, and the sorted prompt word is sent to different large models for reasoning.

[0084] Step S15: Use the inference result to determine the request result corresponding to the knowledge query request, and return the request result to the knowledge user.

[0085] In this embodiment, the AI agent forms a complete answer by integrating the inference results of different large models, and returns the answer as the request result corresponding to the knowledge query request to the user.

[0086] It should be noted that in this embodiment, on the basis of AI basic capabilities, the scenario of combining the trusted data space with AI is explored, that is, the knowledge base construction and usage method of the trusted data space described above. By constructing a distributed knowledge base cluster that can be dynamically scaled, and deeply integrating various capabilities in the trusted data space, such as data trading and data usage control, during the usage process of the knowledge base cluster, cross-institutional knowledge sharing is realized. Among them, data trading, data usage control, etc. make knowledge holders willing to let knowledge circulate.

[0087] As can be seen from the above, this embodiment provides a method for constructing and using a knowledge base in a trusted data space, which is mainly divided into two parts. The first part is to establish a local knowledge base based on the knowledge documents provided by each data provider, and the second part is that the data user uses the knowledge base in the trusted data space to retrieve relevant knowledge by asking questions. In this way, through the above two parts, the creation and usage of the knowledge base cluster are basically realized, allowing the knowledge in the knowledge base to circulate, making the answer results of the large model more accurate and more in line with the user's questions. At the same time, by performing sensitivity grading processing on knowledge blocks and requesting large models with different protection capabilities according to different sensitivities, more fine-grained protection of knowledge is achieved, improving the security of large model reasoning and reducing the risk of sensitive information leakage.

[0088] See Figure 6 As shown, the embodiment of the present application also discloses a device for constructing and using a knowledge base in a trusted data space, including:

[0089] A data acquisition module 11, configured to acquire the target document uploaded by the data provider, and perform a chunking operation on the target document to obtain target knowledge chunks;

[0090] The knowledge base construction module 12 is configured to perform sensitivity grading on the target knowledge chunks based on a preset sensitivity detection method to obtain first grading results corresponding to the target knowledge chunks, perform vectorization processing on the target knowledge chunks using different types of preset large models according to the first grading results to obtain target vector data, and construct a target knowledge base based on the target vector data;

[0091] The data determination module 13 is configured to obtain a knowledge query request sent by a knowledge user, perform the vectorization processing on the knowledge query request to obtain a query request vector, and determine target data corresponding to the query request vector from the constructed target knowledge bases;

[0092] The result acquisition module 14 is configured to perform the sensitivity grading on the target data based on the preset sensitivity detection method to obtain second grading results corresponding to the target data, and input the target data into the different types of preset large models according to the second grading results to obtain corresponding inference results;

[0093] The request response module 15 is configured to determine a request result corresponding to the knowledge query request using the inference result and return the request result to the knowledge user.

[0094] As can be seen from the above, this application grades target knowledge chunks based on a preset sensitivity detection method. Through the grading process, knowledge chunks with different sensitivity levels can be safely isolated during the vectorization stage, meeting compliance requirements while optimizing the allocation of computing resources. At the same time, each data provider constructs a target knowledge base based on the vectorization results of local knowledge chunks, and all target knowledge bases form a distributed cluster through a trusted data space, supporting global retrieval and collaborative reasoning across data sources, improving the comprehensiveness and accuracy of the large model's answers.

[0095] In some specific embodiments, the data acquisition module 11 includes:

[0096] The first knowledge chunk determination unit is configured to detect the target document to determine each complete paragraph included in the target document, and determine any complete paragraph in the target document as the first target knowledge chunk;

[0097] The second knowledge chunk determination unit is configured to segment the document data in the target document except the first target knowledge chunk based on a preset chunk size to obtain second target knowledge chunks;

[0098] The target knowledge chunk determination unit is configured to determine the target knowledge chunks corresponding to the target document according to the first target knowledge chunk and the second target knowledge chunks, and set indexes for the target knowledge chunks.

[0099] In some specific embodiments, the knowledge base construction module 12 includes:

[0100] A detection unit for performing sensitivity detection on the target knowledge block based on a preset sensitivity detection method;

[0101] A sensitivity classification unit for, if the current target knowledge block meets the preset high-sensitivity standard, determining the current target knowledge block as a high-sensitivity knowledge block, if the current target knowledge block meets the preset medium-sensitivity standard, determining the current target knowledge block as a medium-sensitivity knowledge block, and if the current target knowledge block meets the preset low-sensitivity standard, determining the current target knowledge block as a low-sensitivity knowledge block.

[0102] In some specific embodiments, the knowledge base construction module 12 includes:

[0103] A first data determination unit for, for the high-sensitivity knowledge block, performing vectorization processing on the high-sensitivity knowledge block through a preset large model constructed based on a trusted execution environment cluster to obtain first vector data;

[0104] A second data determination unit for, for the medium-sensitivity knowledge block, performing vectorization processing on the medium-sensitivity knowledge block through a preset large model deployed in the local infrastructure platform corresponding to the data provider to obtain second vector data;

[0105] A third data determination unit for, for the low-sensitivity knowledge block, performing vectorization processing on the low-sensitivity knowledge block through a preset large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain third vector data;

[0106] A target data determination unit for determining the target vector data corresponding to the target knowledge block based on the first vector data, the second vector data, and the third vector data;

[0107] A knowledge base construction unit for constructing a target knowledge base based on the target knowledge block, the index corresponding to each target knowledge block, and the target vector data.

[0108] In some specific embodiments, the data determination module 13 includes:

[0109] A similarity determination unit for determining the vector similarity between the query request vector and the target vector data in each of the already constructed target knowledge bases based on a preset search method;

[0110] A fourth data determination unit, configured to sort the target vector data in descending order of the vector similarity, and determine a preset number of vector data from the sorted target vector data as target vectors corresponding to the query request vector, and use the target knowledge blocks corresponding to the target vectors as target data corresponding to the knowledge query request.

[0111] In some specific embodiments, the result acquisition module 14 includes:

[0112] A first result determination unit, configured to input first target data representing high sensitivity in the second classification result into a preset large model constructed based on a trusted execution environment cluster to obtain a corresponding first inference result;

[0113] A second result determination unit, configured to input second target data representing medium sensitivity in the second classification result into a preset large model deployed in a local infrastructure platform corresponding to a data provider to obtain a corresponding second inference result;

[0114] A third result determination unit, configured to input third target data representing low sensitivity in the second classification result into a preset large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain a corresponding third inference result.

[0115] In some specific embodiments, the result acquisition module 14 includes:

[0116] A first request determination unit, configured to, for first target data representing high sensitivity in the second classification result, determine first target credibility corresponding to the first target data based on a preset credibility standard, construct a first target query request related to the content of the first target data according to the content of the first target data and the corresponding first target credibility, and input the first target query request and the first target data into a corresponding preset large model to obtain a corresponding inference result;

[0117] A second request determination unit, configured to, for second target data representing medium sensitivity in the second classification result, determine second target credibility corresponding to the second target data based on the preset credibility standard, construct a second target query request related to the content of the second target data according to the content of the second target data and the corresponding second target credibility, and input the second target query request and the second target data into a corresponding preset large model to obtain a corresponding inference result;

[0118] A third request determination unit, configured to, for third target data representing low sensitivity in the second classification result, determine third target credibility corresponding to the third target data respectively based on the preset credibility standard, construct a third target query request related to the content of the third target data according to the content of the third target data and the corresponding third target credibility, and input the third target query request and the third target data into a corresponding preset large model to obtain a corresponding inference result.

[0119] Further, an embodiment of the present application also discloses an electronic device. Figure 7 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure should not be regarded as any limitation on the scope of use of the present application.

[0120] Figure 7 It is a structural schematic diagram of an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the method for constructing and using the knowledge base of the trusted data space disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0121] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0122] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc., and the resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.

[0123] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the method for constructing and using the knowledge base of the trusted data space executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0124] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the method for constructing and using the knowledge base of the trusted data space disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0125] In this specification, the various embodiments are described in a progressive manner. The focus of each embodiment is on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for the relevant parts.

[0126] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0127] The steps of the method or algorithm described in combination with the embodiments disclosed in this article can be directly implemented by hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0128] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0129] The technical solutions provided in this application have been introduced in detail above. Specific examples are used in this text to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for constructing and using a knowledge base of a trusted data space, characterized in that Including: Obtain a target document uploaded by a data provider, and perform a chunking operation on the target document to obtain target knowledge chunks; Based on a preset sensitivity detection method, perform sensitivity grading on the target knowledge chunks to obtain a first grading result corresponding to each target knowledge chunk. According to the first grading result, use different types of preset large models to perform vectorization processing on the target knowledge chunks to obtain target vector data, and construct a target knowledge base based on the target vector data; Obtain a knowledge query request sent by a knowledge user, perform the vectorization processing on the knowledge query request to obtain a query request vector, and determine target data corresponding to the query request vector from each constructed target knowledge base; Based on the preset sensitivity detection method, perform the sensitivity grading on the target data to obtain a second grading result corresponding to each target data. According to the second grading result, input the target data into the different types of preset large models to obtain corresponding inference results; Use the inference results to determine a request result corresponding to the knowledge query request, and return the request result to the knowledge user.

2. The method for constructing and using a knowledge base of a trusted data space according to claim 1, wherein The performing a chunking operation on the target document to obtain target knowledge chunks includes: Detect the target document to determine each complete paragraph included in the target document, and determine any complete paragraph in the target document as a first target knowledge chunk; Based on a preset chunk size, segment the document data in the target document after removing the first target knowledge chunk to obtain second target knowledge chunks; Determine target knowledge chunks corresponding to the target document according to the first target knowledge chunk and the second target knowledge chunks, and set indexes for each target knowledge chunk.

3. The method for constructing and using a knowledge base of a trusted data space according to claim 1, characterized in that The performing sensitivity grading on the target knowledge chunks based on a preset sensitivity detection method to obtain a first grading result corresponding to each target knowledge chunk includes: Perform sensitivity detection on the target knowledge chunks based on a preset sensitivity detection method; If the current target knowledge chunk meets a preset high-sensitivity standard, determine the current target knowledge chunk as a high-sensitivity knowledge chunk. If the current target knowledge chunk meets a preset medium-sensitivity standard, determine the current target knowledge chunk as a medium-sensitivity knowledge chunk. If the current target knowledge chunk meets a preset low-sensitivity standard, determine the current target knowledge chunk as a low-sensitivity knowledge chunk.

4. The method for constructing and using a knowledge base of a trusted data space according to claim 3, characterized in that, The using different types of preset large models to perform vectorization processing on the target knowledge chunks according to the first grading result to obtain target vector data and constructing a target knowledge base based on the target vector data includes: For the high-sensitivity knowledge chunks, perform vectorization processing on the high-sensitivity knowledge chunks through a preset large model constructed based on a trusted execution environment cluster to obtain first vector data; For the medium-sensitivity knowledge chunks, perform vectorization processing on the medium-sensitivity knowledge chunks through a preset large model deployed in the local infrastructure platform corresponding to the data provider to obtain second vector data; For the low-sensitivity knowledge chunks, vectorize the low-sensitivity knowledge chunks through a preset large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain third vector data; Determine the target vector data corresponding to the target knowledge chunks based on the first vector data, the second vector data, and the third vector data; Construct a target knowledge base based on the target knowledge chunks, the indexes corresponding to the target knowledge chunks, and the target vector data.

5. The method for constructing and using a knowledge base of a trusted data space according to claim 4, characterized in that, The determining the target data corresponding to the query request vector from the constructed target knowledge bases includes: Determine the vector similarity between the query request vector and the target vector data in the constructed target knowledge bases based on a preset search method; Sort the target vector data in descending order of the vector similarity, and determine a preset number of vector data as the target vectors corresponding to the query request vector from the sorted target vector data, and use the target knowledge chunks corresponding to the target vectors as the target data corresponding to the knowledge query request.

6. The method for constructing and using the knowledge base of the trusted data space according to any one of claims 1 to 5, characterized in that, The inputting the target data into the preset large models of different types according to the second classification result to obtain corresponding inference results includes: Input the first target data representing high sensitivity in the second classification result into a preset large model constructed based on a trusted execution environment cluster to obtain a corresponding first inference result; Input the second target data representing medium sensitivity in the second classification result into a preset large model deployed in the local infrastructure platform corresponding to the data provider to obtain a corresponding second inference result; Input the third target data representing low sensitivity in the second classification result into a preset large model provided by a platform other than the local infrastructure platform corresponding to the data provider to obtain a corresponding third inference result.

7. The method for constructing and using a knowledge base of a trusted data space according to claim 1, characterized in that, Before inputting the target data into the preset large models of different types according to the second classification result to obtain corresponding inference results, it further includes: For the first target data representing high sensitivity in the second classification result, determine the first target credibility corresponding to the first target data based on a preset credibility standard, construct a first target query request related to the content of the first target data according to the content of the first target data and the corresponding first target credibility, and input the first target query request and the first target data into the corresponding preset large model to obtain a corresponding inference result; For the second target data representing medium sensitivity in the second classification result, determine the second target credibility corresponding to the second target data based on the preset credibility standard, construct a second target query request related to the content of the second target data according to the content of the second target data and the corresponding second target credibility, and input the second target query request and the second target data into the corresponding preset large model to obtain a corresponding inference result; For the third target data representing low sensitivity in the second classification result, determine the corresponding third target credibility of the third target data based on the preset credibility standard, and construct a third target query request related to the content of the third target data according to the content of the third target data and the corresponding third target credibility, so as to input the third target query request and the third target data into the corresponding preset large model to obtain the corresponding inference result.

8. An apparatus for constructing and using a knowledge base of a trusted data space, characterized in that Including: A data acquisition module, configured to acquire a target document uploaded by a data provider, and perform a chunking operation on the target document to obtain target knowledge chunks; A knowledge base construction module, configured to perform sensitivity classification on the target knowledge chunks based on a preset sensitivity detection method to obtain a first classification result corresponding to each target knowledge chunk, perform vectorization processing on the target knowledge chunks using different types of preset large models according to the first classification result to obtain target vector data, and construct a target knowledge base based on the target vector data; A data determination module, configured to acquire a knowledge query request sent by a knowledge user, perform the vectorization processing on the knowledge query request to obtain a query request vector, and determine target data corresponding to the query request vector from each constructed target knowledge base; A result acquisition module, configured to perform the sensitivity classification on the target data based on the preset sensitivity detection method to obtain a second classification result corresponding to each target data, and input the target data into the different types of preset large models according to the second classification result to obtain the corresponding inference result; A request response module, configured to determine a request result corresponding to the knowledge query request by using the inference result, and return the request result to the knowledge user.

9. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the method for constructing and using a knowledge base of a trusted data space according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program, when executed by a processor, implements the method for constructing and using a knowledge base of a trusted data space according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data security control method for establishing intelligent assistant based on AIGC + enterprise internal knowledge base

    CN117573819A

  • Multi-party joint vector knowledge base retrieval method and system for privacy protection

    CN117708263A

  • Multi-party participation privacy security knowledge base construction method

    CN118245565A

  • Retrieval enhancement method and device, equipment and storage medium

    CN118394793A

  • Data query method and device, electronic equipment and computer program product

    CN119149579A