Method for secure sharing of data in a shared network for knowledge exchange
By constructing an agent knowledge graph and analyzing node characteristics, filtering individual nodes, and calculating adaptive privacy budget parameters, the problem of the inability of fixed privacy budget parameters to dynamically balance privacy and accuracy is solved, thereby improving the data sharing effect.
Patent Information
- Application Number
- CN202511465207.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-10-14
AI Technical Summary
In the process of knowledge data sharing, existing technologies cannot dynamically balance privacy and accuracy by using fixed privacy budget parameters, resulting in poor data sharing effects and failing to adapt to the differences in data value and privacy sensitivity among different intelligent agents.
Construct a knowledge graph for intelligent agents, analyze the characteristics of the number of references, the trend of reference frequency, and the topology of the graph nodes, screen individual nodes, and calculate adaptive privacy budget parameters for privacy protection based on the distribution of shared data in the graph.
By adaptively adjusting privacy budget parameters, the privacy and accuracy in the data sharing process are dynamically balanced, improving the effectiveness of data sharing and avoiding the inadequacy of static strategies.
Smart Images

Figure CN120951388B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data protection, in particular to a data security sharing method in a shared network for knowledge exchange. BACKGROUND
[0002] An AI agent is a high-level artificial intelligence system that can autonomously perceive, think and act in a specific environment. It can understand, learn and reason to perform complex tasks and make decisions. It is commonly applied in various knowledge fields through interactive dialogue. In order to expand the knowledge coverage of the agent, a shared network platform can be introduced, and agents in different fields can share data in the shared network platform to realize knowledge exchange.
[0003] In the prior art, a static privacy protection strategy is usually used in the process of knowledge data sharing, such as a fixed differential privacy budget. The privacy budget parameter is essentially noise caused by the data sharing process, which affects the accuracy and privacy of data sharing. However, since the value and privacy sensitivity of the data shared by different agents are different, the accuracy and privacy requirements of the data in the sharing process are also different. Therefore, relying only on a preset fixed privacy budget parameter cannot dynamically balance the privacy and accuracy, which seriously affects the data sharing effect. SUMMARY
[0004] In order to solve the technical problem that since the value and privacy sensitivity of the data shared by different agents are different, the accuracy and privacy requirements of the data in the sharing process are also different, and therefore relying only on a preset fixed privacy budget parameter cannot dynamically balance the privacy and accuracy, which seriously affects the data sharing effect, the purpose of the present application is to provide a data security sharing method in a shared network for knowledge exchange, and the technical solution adopted is as follows:
[0005] Obtain the knowledge graph corresponding to each agent in the shared network, wherein each graph node in the knowledge graph has label information and contains a plurality of references, and each graph node corresponds to an extended graph;
[0006] In the graph node, analyze the number of references, the number of references and the change trend of the number of references, and combine them with the topology structure of the extended graph of the graph node, so as to filter out individual nodes in the knowledge graph of each agent;
[0007] In the current knowledge sharing process, analyze the distribution between the corresponding graph nodes and individual nodes of the shared data in the knowledge graph of the sharing agent and the receiving agent, and the general situation of the graph nodes corresponding to the shared data in the knowledge graph of all agents, and determine the sharing benefit performance of the sharing party and the receiving party.
[0008] Based on the data receiving situation of the receiver in the historical knowledge sharing process and the privacy budget parameter, and in combination with the sharing benefit performance degree, an adaptive privacy budget parameter is calculated for the privacy protection of the shared data.
[0009] Further, the acquisition method of the individual node comprises:
[0010] In each graph node corresponding to each agent, based on the change trend of the reference times of each reference address within a preset period and the numerical characteristics of the reference times, a reference confidence factor of each reference address is determined;
[0011] Based on the number of reference addresses in the graph nodes of all agents and the reference confidence factors, the availability of each graph node is obtained;
[0012] In the knowledge graph corresponding to all agents, the differences between the topological structures of the extended graphs of the graph nodes with the same label information are analyzed, so as to obtain the content coverage comprehensive degree of each graph node in the knowledge graph of each agent;
[0013] In the knowledge graph of each agent, the product of the content coverage comprehensive degree and the availability of each graph node after normalization is taken as the individual attribute reflection degree of each graph node;
[0014] In the knowledge graph of each agent, the graph node with an individual attribute reflection degree greater than a preset individual threshold is taken as an individual node.
[0015] Further, the acquisition method of the reference confidence factor comprises:
[0016] In each graph node corresponding to each agent, the reference times of each reference address within a preset period are sorted in time sequence, and then linear fitting is performed based on the least square method, and the slope value of the fitted straight line is obtained. The value after normalization of the slope value is taken as a reference growth characteristic value;
[0017] The product of the total number of reference times of each reference address within a preset period and the reference growth characteristic value after normalization is taken as the reference confidence factor of each reference address in each graph node.
[0018] Further, the acquisition method of the availability comprises:
[0019] The number of reference addresses in each graph node is taken as a quantity factor of each graph node;
[0020] The ratio of the quantity factor of each graph node to the maximum value of the quantity factors of the graph nodes of all agents in the sharing network is taken as an importance factor of each graph node;
[0021] The importance factor of each graph node is multiplied by the average of the reference confidence factors of all references in each graph node, and the normalized value of the resulting product is the availability of each graph node.
[0022] Further, the content coverage comprehensiveness acquisition method comprises:
[0023] In the extended graph of each graph node, the number of all extended nodes is the diffusion degree of each graph node.
[0024] In the knowledge graph of all agents in the shared network, the average of the diffusion degrees of the graph nodes with the same label information is the average diffusion value.
[0025] In the knowledge graph of each agent, the normalized value of the difference between the diffusion degree of each graph node and the corresponding average diffusion value is the content coverage comprehensiveness of each graph node.
[0026] Further, the shared benefit performance degree acquisition method comprises:
[0027] In the current knowledge sharing process, based on the positional relationship between the corresponding graph node of the shared data and the individual node in the knowledge graph of the sharing agent, the authority coefficient of the shared data in the sharing agent is determined.
[0028] In the current knowledge sharing process, based on the positional relationship between the corresponding graph node of the shared data and the individual node in the knowledge graph of the receiving agent, and the number characteristics of the individual nodes in the corresponding graph node of the shared data in the knowledge graph of the receiving agent, the deficiency coefficient of the shared data in the receiving agent is determined.
[0029] In the current knowledge sharing process, the ratio of the corresponding graph node of the shared data in all agents to the total number of graph nodes in all agents is the universal coefficient of the shared data.
[0030] In the current knowledge sharing process, the value of the universal coefficient after negative correlation mapping is normalized with the product of the authority coefficient, the deficiency coefficient and the universal coefficient to obtain the shared benefit performance degree of the sharing agent and the receiving agent.
[0031] Further, the authority coefficient acquisition method comprises:
[0032] In the knowledge graph of the sharing agent, the graph node with the same information label as the shared data is taken as the corresponding graph node of the shared data, denoted as the target node.
[0033] In the knowledge graph of the sharing agent, the shortest path between each target node and the nearest individual node is obtained, and the number of graph nodes passed on the shortest path is taken as the distance factor;
[0034] The sum of the distance factors corresponding to all target nodes is negatively correlated and normalized, and the value after normalization is taken as the authority coefficient of the shared data in the sharing agent.
[0035] Further, the method for obtaining the deficiency coefficient comprises:
[0036] In the knowledge graph of the receiving agent, the graph node with the same information label as the shared data is taken as the graph node corresponding to the shared data, and is recorded as a reference node;
[0037] The number of individual nodes in the reference node is negatively correlated and mapped, and the value after mapping is taken as the deficiency factor;
[0038] In the knowledge graph of the receiving agent, the shortest path between each reference node and the nearest individual node is obtained, and the number of graph nodes passed on the shortest path is taken as the distance length;
[0039] The product of the distance length of all reference nodes and the deficiency factor is normalized, and the value after normalization is taken as the deficiency coefficient of the shared data in the receiving agent; wherein, if there is no reference node, the deficiency coefficient is a preset value.
[0040] Further, the method for obtaining the adaptive privacy budget parameter comprises:
[0041] The data receiving situation and the privacy budget parameter of the receiving agent in the historical knowledge sharing process are analyzed to determine the preset privacy budget parameter in the current knowledge sharing process;
[0042] The product of the difference between the sharing benefit performance degree and the preset parameter and the preset maximum adjustment amplitude is taken as the adjustment value, and the sum of the adjustment value and the preset privacy budget parameter is taken as the adaptive privacy budget parameter in the current knowledge sharing process.
[0043] Further, the method for obtaining the preset privacy budget parameter comprises:
[0044] In each historical knowledge sharing process of the receiving agent, the ratio of the number of graph nodes successfully imported by the shared data to the number of graph nodes corresponding to the shared data in the sharing agent is taken as the confidence degree;
[0045] The privacy budget parameter is weighted using the confidence degree corresponding to the historical knowledge sharing process of the receiving agent, and the obtained weighted result is taken as the preset privacy budget parameter in the current knowledge sharing process.
[0046] The present application has the following beneficial effects:
[0047] Firstly, in the shared network, the knowledge graph corresponding to each agent is constructed, and the knowledge graph includes graph nodes, each graph node has label information and includes a plurality of references in the graph node, and each graph node also corresponds to an extension graph, which can be regarded as a subgraph, and is used for detailed explanation of the graph node. Different agents focus on different knowledge fields in the shared network, that is, the individual attributes are different, and the individual attributes of the agents affect the privacy parameter demand, so the number characteristics of the references, the change trend of the reference times and the topological structure of the extension graph can be analyzed in the graph node, and the individual nodes are screened out in the knowledge graph of each agent to represent the unique attributes of each agent. Further, the value of the shared data to the shared agent and the receiving agent can be quantified, that is, whether the shared data is the "key knowledge" that is scarce to the receiving agent, so in the current knowledge sharing process, the distribution relationship between the graph nodes and the individual nodes corresponding to the shared data in the knowledge graphs of the two parties is analyzed, and the general situation of the shared data in the knowledge graphs of all agents is analyzed, and the sharing benefit performance of the shared data to the two parties is determined. Finally, based on the data receiving situation of the receiving party in the historical knowledge sharing process, the privacy budget parameter and the benefit characteristics (sharing benefit performance) in the current knowledge sharing process, an adaptive budget privacy parameter is obtained, at this time, the index can better maintain the dynamic balance of the accuracy and privacy of the data in the current sharing process, avoid the inadaptability of the static strategy, and effectively improve the data sharing effect. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art and the advantages thereof, a brief introduction will be given to the drawings needed in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.
[0049] Figure 1 A method flowchart of a data security sharing method in a shared network for knowledge exchange provided by an embodiment of the present application;
[0050] Figure 2 A knowledge graph of an agent provided by an embodiment of the present application;
[0051] Figure 3 A method flowchart of an individual node acquisition method provided by an embodiment of the present application;
[0052] Figure 4A method flow chart of a shared benefit performance degree acquisition method provided by one embodiment of the present application;
[0053] Reference signs: 1 - graph node, 2 - indexing. DETAILED DESCRIPTION
[0054] In order to further clarify the technical means and effects taken by the present application to achieve the predetermined object of the application, the following describes in detail the specific implementation, structure, features and effects of a data security sharing method in a shared network for knowledge exchange according to the present application in combination with the preferred embodiments and the drawings. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.
[0056] The following specifically describes the specific scheme of the data security sharing method in a shared network for knowledge exchange provided by the present application in combination with the drawings.
[0057] Please refer to Figure 1 , which shows a method flow chart of the data security sharing method in a shared network for knowledge exchange provided by one embodiment of the present application. The method includes the following steps:
[0058] Step S1: acquiring the knowledge graph corresponding to each agent in the shared network, wherein each graph node in the knowledge graph has label information and contains several indexings, and each graph node corresponds to an extended graph.
[0059] The agents in different fields can realize knowledge exchange through data sharing in the shared network platform. When the knowledge exchange is performed, the shared agent receives a sharing request for a certain range of knowledge from the shared agent (the receiving agent), and the shared agent shares the part of knowledge data to the shared agent after receiving the request. The shared agent trains and imports the part of shared data into the local database, so as to perform the dialogue and exchange between the agents. In the process of importing and training the shared knowledge by the shared agent, the differential privacy algorithm is usually used to protect the privacy in the sharing process. The privacy budget parameter is an important parameter of the differential privacy algorithm. The parameter is mainly used to control the difference degree of the output distribution of adjacent data sets. The essence is the noise amount introduced in the data sharing process. The smaller the value (the larger the noise), the more similar the output distribution, and the strongest the privacy protection, but the data distortion is serious, and the receiving agent cannot effectively use the information. The larger the value (the smaller the noise), the higher the risk of privacy leakage, although the data accuracy is ensured. Therefore, the embodiment of the present application adaptively modifies the preset privacy budget parameter according to the individual attributes and data sharing performance of the agents in the actual shared network scene, so as to improve the data security sharing efficiency of the knowledge exchange parties.
[0060] Firstly, each agent in the shared network platform constructs a corresponding knowledge graph according to the local database. In the embodiment of the present application, the object of constructing the knowledge graph is mainly for knowledge points, concept explanations, reference papers, reference websites and the like. There are a plurality of topological structure connected graph nodes in the knowledge graph. Each graph node has label information (which can describe the general explanation of the graph node). There are a plurality of addresses (such as addresses of reference websites or reference papers) in each graph node. Please refer to Figure 2 which shows the knowledge graph of an agent in one embodiment of the present application; and in the knowledge graph, all parts (subgraphs) diffused outward from each graph node are regarded as the extended graph of the graph node.
[0061] It should be noted that the construction of the knowledge graph can be generated by multi-source data extraction and NLP technology, which is a known technology, and the specific process is not described here.
[0062] Step S2: In the graph node, the number characteristics of the address, the reference times of the address and the change trend of the reference times are analyzed, and the topological structure of the extended graph of the graph node is combined, so as to screen out the individual node in the knowledge graph of each agent.
[0063] The knowledge fields corresponding to different agents in the shared network platform have individual attribute differences, and the individual attributes of the agents affect the privacy parameter requirements. For example, the more the shared data tends to be the individual attribute knowledge field (core content) of the agent on the sharing side, the less the privacy budget parameter can be adjusted to ensure the sharing accuracy of the shared data. Therefore, in the embodiment of the present application, the individual nodes in the knowledge graph of each agent, i.e., the most representative graph nodes, need to be analyzed.
[0064] In the graph nodes, since the indexing is linked with other knowledge or external resources, the reference times and reference conditions of the indexing can reflect whether the indexing is frequently accessed and whether it is a commonly used tool, and the quantity characteristics of the indexing can reflect whether there is extensive association with other fields or resources. Therefore, the various characteristics described above are used as reference data to describe the individual attribute conditions of the graph nodes. At the same time, since the extended graph of each graph node is its subgraph, it also represents more detailed interpretation information of the graph node, so the topology of the extended graph can be used to reflect the knowledge coverage of the graph node and can be used as a reference basis to characterize the individual attribute conditions. Therefore, the various data characteristics described above are fused to filter out the individual nodes in the knowledge graph of each agent.
[0065] Preferably, the method for obtaining the individual nodes in the embodiment of the present application includes:
[0066] Please refer to Figure 3 which shows a method flowchart of the method for obtaining the individual nodes in the embodiment of the present application. The method includes the following steps:
[0067] Step S201: In each graph node corresponding to each agent, based on the change trend of the reference times of each indexing in a preset period and the numerical characteristics of the reference times, a reference confidence factor of each indexing is determined.
[0068] In each graph node corresponding to each agent, after the reference times of each indexing in a preset period are sorted in time sequence, a straight line fitting is performed based on the least square method, and the slope value of the fitted straight line is obtained. The greater the slope value is, the higher the hot performance of the indexing is, and the greater the demand is. Therefore, the value of the slope value after normalization is used as a reference growth characteristic value to represent the active trend. The greater the growth characteristic value is, the greater the reference demand degree of the indexing is, and the higher the credibility is. Since the slope value here can be positive or negative, the normalization method uses function.
[0069] Meanwhile, since the total number of citation times can represent the long-term frequency of referencing, the greater the value, the greater the demand for the sound quality of the reference, and the higher the reliability, therefore, the product of the total number of citation times of each reference in the preset period and the reference growth feature value after normalization is taken as the reference confidence factor of each reference in each graph node. The greater the reference confidence factor, the more frequently the reference in the graph node is called, the higher the reliability, and the more likely it represents the individual attributes of the graph node. The normalization is a well-known technical means for those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, and the specific normalization method is not limited here.
[0070] It should be noted that the least squares method is a well-known technique, and the specific process is not repeated here; the preset period is in units of days, which can be set to the last week, and the specific duration can be adjusted according to the implementation scenario, which is not limited here.
[0071] Step S202: Based on the number of references in the graph node of all agents and the reference confidence factor, the availability of each graph node is obtained.
[0072] The reference is the connection between knowledge or resources, and the more references, the more extensive the association of the graph node with other knowledge fields or resources, so the number of references in each graph node is taken as the quantity factor of each graph node, and the ratio of the quantity factor of each graph node to the maximum value of the quantity factor of all graph nodes of the agents in the sharing network is taken as the importance factor of each graph node. The greater the importance factor, the more prominent the association ability of the references in the graph node, and thus the graph node is more likely to reflect the individual characteristics of the agent.
[0073] Finally, the product of the importance factor of each graph node and the average of the reference confidence factors of all references in each graph node is multiplied, and the normalized value of the resulting product is taken as the availability of each graph node. Based on the foregoing analysis, the availability integrates the importance and stability of the graph node, so the greater the value, the higher the demand degree and the higher the reliability, and thus the graph node can better represent the individual attribute characteristics of the agent knowledge graph.
[0074] Step S203: In the knowledge graph corresponding to all agents, the differences between the topological structures of the extended graphs of the graph nodes with the same label information are analyzed, thereby obtaining the content coverage degree of each graph node in the knowledge graph of each agent.
[0075] The extended graph of each graph node is a subgraph that spreads outward from the graph node, and the more complex the topological structure of the spread subgraph, the more subnodes the graph node is associated with, and the more extensive the coverage content, and thus the more likely it represents the unique attributes of the knowledge graph of the agent.
[0076] Therefore, in the extended knowledge graph of each knowledge graph node, the number of all extended nodes is taken as the diffusion degree of each knowledge graph node, and in the knowledge graph of all agents in the shared network, the average of the diffusion degrees of the knowledge graph nodes with the same label information is taken as the average diffusion value.
[0077] Then, in the knowledge graph of each agent, the difference between the diffusion degree of each knowledge graph node and the corresponding average diffusion value is calculated. The difference is positive, and the greater the difference, the stronger the knowledge coverage ability of the knowledge graph node, and the stronger the individual attribute. Therefore, the value of the difference after normalization is taken as the content coverage degree of each knowledge graph node. The greater the content coverage degree, the wider the content interpretation coverage of the knowledge graph node, and the better the individual attribute characteristics of the knowledge graph of the agent to which the knowledge graph node belongs. Since the difference here can be positive or negative, the normalization method adopts function.
[0078] Step S204: In the knowledge graph of each agent, the content coverage degree and the availability of each knowledge graph node are fused to obtain the individual attribute reflection degree of each knowledge graph node, and the individual nodes are screened in the knowledge graph node based on the individual attribute reflection degree.
[0079] Based on the analysis in the foregoing steps, in the knowledge graph of each agent, the content coverage degree and the availability of each knowledge graph node are positively correlated with the ability of the knowledge graph node to represent the individual attribute characteristics of the knowledge graph. Therefore, the product of the content coverage degree and the availability of each knowledge graph node after normalization is taken as the individual attribute reflection degree of each knowledge graph node. The greater the individual attribute reflection degree, the better the individual attribute of the knowledge graph node to represent the individual attribute of the knowledge domain of the agent to which the knowledge graph node belongs. The normalization is a technical means familiar to those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, and the specific normalization method is not limited herein.
[0080] Therefore, in the knowledge graph of each agent, the knowledge graph node with an individual attribute reflection degree greater than a preset individual threshold is taken as an individual node.
[0081] It should be noted that the preset individual threshold is 0.85, and the specific value can be adjusted according to the implementation scenario, which is not limited herein.
[0082] Step S3: In the current knowledge sharing process, the distribution of the shared data in the knowledge graph of the sharing agent and the receiving agent is analyzed, and the distribution of the shared data in the knowledge graph of the sharing agent and the receiving agent is analyzed. The distribution of the shared data between the corresponding knowledge graph nodes and the individual nodes, and the general situation of the shared data corresponding to the knowledge graph nodes in the knowledge graph of all agents is determined, and the sharing benefit performance of the sharing agent and the receiving agent is determined.
[0083] Based on the analysis process in step S2, the individual node in the knowledge graph of each agent in the sharing network platform can be obtained. Further, for the sharing parties in the current knowledge sharing process, the value of the shared data to the receiving agent can be analyzed, that is, whether the data is lacking for the receiving agent. Therefore, in this step, the distribution relationship between the graph nodes corresponding to the shared data in the knowledge graph of the two parties and the individual nodes is analyzed, and the general situation of the shared data in the knowledge graph of all agents is determined, so as to determine the sharing benefit performance degree of the shared data to the two parties. The quantitative result of the sharing benefit performance degree directly affects the adjustment of the privacy budget, and provides decision support for the subsequent dynamic adjustment of the privacy budget parameter.
[0084] Preferably, in an embodiment of the present application, the method for obtaining the sharing benefit performance degree comprises:
[0085] Referring to Figure 4 , a method flowchart of the method for obtaining the sharing benefit performance degree in an embodiment of the present application is shown, and the method comprises the following steps:
[0086] Step S301: In the current knowledge sharing process, based on the positional relationship between the graph nodes corresponding to the shared data in the knowledge graph of the sharing agent and the individual nodes, the authority coefficient of the shared data in the sharing agent is determined.
[0087] In the knowledge graph of the sharing agent, the graph node with the same information tag as the shared data is taken as the graph node corresponding to the shared data, and is denoted as a target node.
[0088] Since the individual node in the knowledge graph can better represent the individual attribute of the knowledge field of the agent, it can be regarded as a core node. Therefore, in the knowledge graph of the sharing agent, the shortest path between each target node and the nearest individual node is obtained (the shortest path can be obtained by breadth-first search, which is a known technology, and the process is not described in detail), and the number of graph nodes passing through the shortest path is taken as a distance factor. The smaller the distance factor is, the closer the graph node is to the individual node, and the closer the association between the core knowledge is. Therefore, the higher the authority is. Conversely, if the distance factor is larger, it is more likely to be edge knowledge, and the authority is lower. Therefore, the sum of the distance factors corresponding to all target nodes is negatively correlated and normalized to correct the logical relationship, so as to obtain the authority coefficient of the shared data in the sharing agent. The larger the authority coefficient is, the closer the distance between the shared data and the core content in the knowledge graph is, and the more likely the core data is. Therefore, the value in the sharing process is higher, and the authority is larger. The negative correlation mapping and normalization processing can be performed by using the formula , wherein, represents an exponential function with a natural constant e as the base, and x represents the independent variable.
[0089] Step S302: In the current knowledge sharing process, based on the positional relationship between the corresponding graph nodes and individual nodes in the knowledge graph of the receiving agent and the quantity feature of the individual nodes in the corresponding graph nodes of the shared data in the knowledge graph of the receiving agent, the deficiency coefficient of the shared data in the receiving agent is determined.
[0090] If the shared data does not exist or is relatively deficient in the knowledge graph of the receiving agent, the sharing value of the shared data will be higher, so in the knowledge graph of the receiving agent, the graph nodes with the same information label as the shared data are taken as the corresponding graph nodes of the shared data, which are referred to as reference nodes.
[0091] Then the number of individual nodes in the reference nodes is counted, and the smaller the number is, the more deficient the distribution of the shared data in the knowledge graph of the receiving agent is, so the number is negatively correlated and mapped to obtain the deficiency factor. The larger the deficiency factor is, the more important the shared data is to the receiving agent. The negative correlation mapping and normalization processing can be performed by the formula , wherein represents an exponential function with natural constant e as the base, and x represents the independent variable.
[0092] In the knowledge graph of the receiving agent, the shortest path between each reference node and the nearest individual node is obtained (the shortest path can be obtained by breadth-first search, which is a known technology, and the process is not described in detail), and the number of graph nodes passed on the shortest path is taken as the distance length. The larger the distance length is, the farther the distance between the shared data and the core data in the knowledge graph of the receiving agent is, and the more steps of reasoning are needed to obtain it, so the importance to the receiving agent will be higher, the deficiency will be higher, and the sharing demand will be greater.
[0093] Finally, the product of the distance length of all reference nodes and the deficiency factor after normalization is taken as the deficiency coefficient of the shared data in the receiving agent. Based on the foregoing analysis, the larger the deficiency coefficient is, the greater the demand of the receiving agent for the shared data, and the higher the value of the shared data. If there is no reference node, it means that there is no shared data in the receiving agent, so the deficiency is greater, and the deficiency coefficient is a preset value and the preset value is 1. The normalization is a known technology to those skilled in the art, and the selection of the normalization function can be linear normalization or standard normalization, and the specific normalization method is not limited herein.
[0094] Step S303: In the current knowledge sharing process, the general situation of the shared data in the corresponding knowledge graph of all agents is analyzed to determine the general coefficient of the shared data.
[0095] The popularity of the shared data can be quantified by counting the proportion of the graph nodes corresponding to the shared data in all agent knowledge graphs, which can represent whether it is general knowledge. Therefore, in the current knowledge sharing process, the ratio of the graph nodes corresponding to the shared data (the graph nodes with the same label information) to the total number of graph nodes in all agents is used as the general coefficient of the shared data. The larger the general coefficient is, the more the shared data tends to be general knowledge of all agents in the network platform. Therefore, in this data sharing process, the importance is weak, and the sharing value will be reduced.
[0096] Step S304: In the current knowledge sharing process, the authority coefficient, the scarcity coefficient and the general coefficient of the shared data are comprehensively analyzed to obtain the sharing benefit performance between the sharing party and the receiving party.
[0097] Based on the foregoing analysis, in the current knowledge sharing process, the authority coefficient and the scarcity coefficient are positively correlated with the sharing value of the shared data, and the general coefficient is negatively correlated with the sharing value of the shared data. Therefore, the general coefficient can be negatively correlated and mapped, and after the logical relationship is corrected, it is multiplied by the authority coefficient and the scarcity coefficient. The value of the product after normalization is used as the sharing benefit performance of the sharing party and the receiving party. The greater the sharing benefit performance is, the higher the importance of the shared data to the receiving party in the current knowledge sharing process is. Therefore, it should be ensured that the data can be accurately transmitted. The negative correlation mapping and normalization processing can be performed by the formula wherein, represents the exponential function with the natural constant e as the base, and x represents the independent variable.
[0098] Step S4: Based on the data receiving situation of the receiving party in the historical knowledge sharing process and the privacy budget parameter, and in combination with the sharing benefit performance, an adaptive privacy budget parameter is calculated for the privacy protection of the shared data.
[0099] In the foregoing steps, the sharing benefit performance of the sharing party and the receiving party in the current knowledge sharing process can be obtained, which reflects the sharing value of the shared data. In this step, the index can be combined with the data receiving situation of the receiving party in the historical knowledge sharing process and the privacy budget parameter in the historical knowledge sharing process, so as to obtain the adaptive privacy budget parameter most suitable for the current knowledge sharing process, which is used for the privacy protection of the shared data in the current knowledge sharing process, and effectively balances the privacy degree and the accuracy. The historical knowledge sharing process can be obtained through the log file in the database.
[0100] Preferably, in an embodiment of the present application, the method for obtaining the adaptive privacy budget parameter comprises:
[0101] The data receiving condition of the receiving party in the historical knowledge sharing process and the privacy budget parameter are analyzed to determine the preset privacy budget parameter in the current knowledge sharing process: in each historical knowledge sharing process of the receiving party, the ratio of the number of successfully imported graph nodes of the shared data to the number of corresponding graph nodes of the shared data in the sharing party agent is calculated. The larger the ratio is, the larger the volume of the shared data successfully received by the receiving party in the historical knowledge sharing process is, and the greater the degree of appropriateness of the budget privacy parameter in the historical knowledge sharing is. Therefore, the ratio is used as the confidence of the privacy budget parameter in the historical knowledge sharing process. Then, the privacy budget parameters in all historical knowledge sharing processes are weighted using the confidence of the privacy budget parameter, and the obtained weighted result is used as the preset privacy budget parameter in the current knowledge sharing process. If the receiving party has no historical knowledge sharing process, the preset privacy budget parameter can be set, for example, to 10.
[0102] Since the greater the sharing benefit performance degree is, the higher the sharing value of the shared data in the current knowledge sharing process is, the preset privacy budget parameter should be adjusted to be larger, so as to ensure that the data can be accurately shared. Conversely, when the sharing benefit performance degree is smaller, the sharing value of the shared data is lower, and the preset privacy budget parameter can be appropriately adjusted to be smaller to reduce the accuracy and ensure privacy. Therefore, the difference between the sharing benefit performance degree and the preset parameter can be calculated. The difference is positive and larger, and the accuracy should be focused on, and then the preset privacy budget parameter is adjusted to be larger. Conversely, the difference is negative and smaller, and the privacy degree should be focused on, and then the preset privacy budget parameter is adjusted to be smaller. Therefore, the product of the difference and the preset maximum adjustment range can be used as an adjustment value, and finally the adjustment value is added to the preset privacy budget parameter, so as to realize the adjustment of the preset privacy budget parameter to obtain the adaptive privacy budget parameter.
[0103] It should be noted that in this embodiment of the application, the preset parameter is 0.5, and the preset maximum adjustment range is one half of the preset privacy budget parameter, which is rounded up. The specific values can be adjusted according to the implementation scene, which is not limited herein.
[0104] Finally, the differential privacy algorithm can be used to realize the privacy protection of the shared data in the current knowledge sharing process based on the adaptive privacy budget parameter.
[0105] It should be noted that the differential privacy algorithm is a known technology, and the specific process is not described herein.
[0106] To sum up, first, in the shared network, the knowledge graph corresponding to each agent is constructed, and the knowledge graph includes graph nodes, each graph node has label information and includes a number of references in the graph node, and each graph node also corresponds to an extension graph, which can be regarded as a subgraph, used to explain the graph node in detail. Different agents focus on different knowledge fields in the shared network, that is, different individual attributes, and the individual attributes of the agents affect the privacy parameter requirements, so the number of references, the trend of the number of references and the topology structure of the extension graph can be analyzed in the graph node, and the individual nodes in the knowledge graph of each agent are screened out to represent the unique attributes of each agent. Further, the value of the shared data to the shared agent and the receiving agent can be quantified, that is, whether the shared data is a "key knowledge" that is scarce to the receiving agent, so in the current knowledge sharing process, the distribution relationship between the graph nodes and the individual nodes corresponding to the shared data in the knowledge graphs of both parties is analyzed, and the general situation of the shared data in the knowledge graphs of all agents is analyzed, and the sharing benefit performance of the shared data to both parties is determined. Finally, based on the data receiving situation of the receiving party in the historical knowledge sharing process, the privacy budget parameter and the benefit characteristics (sharing benefit performance) in the current knowledge sharing process, an adaptive budget privacy parameter is obtained, at this time, the index can better maintain the dynamic balance between the accuracy and privacy of the data in the current sharing process, avoiding the inadaptability of the static strategy, and effectively improving the data sharing effect.
[0107] It should be noted that the above-mentioned sequence of embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0108] Each embodiment in the specification is described in a progressive manner, and the same and similar parts between each embodiment can be referred to each other, and each embodiment mainly describes the difference from other embodiments.
Claims
1. A method for secure data sharing in a shared network for knowledge exchange, characterized in that, The method includes: Obtain the knowledge graph corresponding to each agent in the shared network. Each graph node in the knowledge graph has label information and contains several references. Each graph node corresponds to an extended graph. In the knowledge graph nodes, the quantitative characteristics of the addresses, the number of references, and the trend of the number of references are analyzed, and combined with the topological structure of the extended graph of the knowledge graph nodes, so as to select individual nodes in the knowledge graph of each agent. In the current knowledge sharing process, we analyze the distribution of graph nodes and individual nodes of the shared data in the knowledge graphs of the sharing party and the receiving party, as well as the commonality of the graph nodes corresponding to the shared data in the knowledge graphs of all parties, to determine the sharing benefit performance of the sharing party and the receiving party. Based on the recipient’s data reception status and privacy budget parameters during the historical knowledge sharing process, and combined with the sharing benefit performance, an adaptive privacy budget parameter is calculated for privacy protection of shared data. The methods for obtaining the personalized nodes include: In each graph node corresponding to each agent, the reference confidence factor of each address is determined based on the changing trend of the number of references within a preset time period and the numerical characteristics of the number of references. The availability of each graph node is obtained based on the number of references in the graph nodes of all agents and the reference confidence factor. In the knowledge graphs corresponding to all agents, we analyze the differences in the topological structure of the extended graphs of graph nodes with the same label information, thereby obtaining the content coverage of each graph node in the knowledge graph of each agent. In the knowledge graph of each agent, the normalized value of the product of the comprehensiveness of the content coverage and the availability of each graph node is used as the individual attribute reflectivity of each graph node. In the knowledge graph of each agent, graph nodes whose individual attribute reflectance is greater than a preset individuality threshold are designated as individuality nodes.
2. The method for secure data sharing in a shared network for knowledge exchange according to claim 1, characterized in that, The method for obtaining the reference confidence factor includes: In each graph node corresponding to each agent, the number of references of each address within a preset time period is sorted according to time sequence, and then a straight line is fitted based on the least squares method. The slope value of the fitted straight line is obtained, and the normalized value of the slope value is used as the reference growth feature value. The normalized value of the product of the total number of references to each address within a preset time period and the reference growth characteristic value is used as the reference confidence factor for each address in each graph node.
3. The method for secure data sharing in a shared network for knowledge exchange according to claim 1, characterized in that, The methods for obtaining availability include: The number of references in each graph node is used as the quantity factor for each graph node; The ratio of the quantity factor of each graph node to the maximum quantity factor of all graph nodes of all agents in the shared network is used as the important factor of each graph node. The importance factor of each graph node is multiplied by the mean of the reference confidence factors of all addresses in each graph node, and the normalized value of the resulting product is used as the availability of each graph node.
4. A method for secure data sharing in a shared network for knowledge exchange according to claim 1, characterized in that, The methods for obtaining the comprehensiveness of the content coverage include: In the extended graph of each graph node, the number of all extended nodes is taken as the diffusion degree of each graph node. In the knowledge graph of all agents in the shared network, the average diffusion value is taken as the mean diffusion value of graph nodes with the same label information. In the knowledge graph of each agent, the normalized value of the difference between the diffusion degree of each graph node and the corresponding average diffusion value is used as the content coverage of each graph node.
5. A method for secure data sharing in a shared network for knowledge exchange according to claim 1, characterized in that, The method for obtaining the shared benefit performance includes: In the current knowledge sharing process, the authority coefficient of the shared data in the sharing agent is determined based on the positional relationship between the graph nodes and individual nodes corresponding to the shared data in the knowledge graph of the sharing agent. In the current knowledge sharing process, based on the positional relationship between the graph nodes and individual nodes corresponding to the shared data in the knowledge graph of the receiving agent, and the quantitative characteristics of the individual nodes in the graph nodes corresponding to the shared data in the knowledge graph of the receiving agent, the scarcity coefficient of the shared data in the receiving agent is determined. In the current knowledge sharing process, the ratio of the graph nodes corresponding to the shared data in all agents to the total number of graph nodes in all agents is used as the universal coefficient of the shared data. In the current knowledge sharing process, the value of the general coefficient after negative correlation mapping is normalized by the product of the authority coefficient and the scarcity coefficient to obtain the sharing benefit performance of the sharing party and the receiving party.
6. A method for secure data sharing in a shared network for knowledge exchange according to claim 5, characterized in that, The methods for obtaining the authority coefficient include: In the knowledge graph of the sharing agent, the graph node with the same information label as the shared data is regarded as the graph node corresponding to the shared data and is denoted as the target node; In the knowledge graph of the shared agent, the shortest path from each target node to the nearest individual node is obtained, and the number of graph nodes traversed on the shortest path is used as the distance factor. The sum of the distance factors corresponding to all target nodes is negatively correlated and normalized, and the resulting value is used as the authority coefficient of the shared data in the shared agent.
7. A method for secure data sharing in a shared network for knowledge exchange according to claim 5, characterized in that, The method for obtaining the scarcity coefficient includes: In the knowledge graph of the receiving agent, the graph node with the same information label as the shared data is regarded as the graph node corresponding to the shared data and is recorded as the reference node. The value obtained by negatively mapping the number of individual nodes in the reference nodes is used as the scarcity factor; In the knowledge graph of the receiving agent, the shortest path from each reference node to the nearest individual node is obtained, and the number of graph nodes traversed on the shortest path is used as the distance length. The normalized value of the product of the distance length of all reference nodes and the scarcity factor is used as the scarcity coefficient of the shared data in the receiving agent; where, if there are no reference nodes, the scarcity coefficient is a preset value.
8. A method for secure data sharing in a shared network for knowledge exchange according to claim 1, characterized in that, The method for obtaining the adaptive privacy budget parameter includes: Analyze the recipient's data reception and privacy budget parameters in the historical knowledge sharing process to determine the preset privacy budget parameters in the current knowledge sharing process; The product of the difference between the shared benefit performance and the preset parameter and the preset maximum adjustment range is used as the adjustment value, and the sum of the adjustment value and the preset privacy budget parameter is used as the adaptive privacy budget parameter in the current knowledge sharing process.
9. A method for secure data sharing in a shared network for knowledge exchange according to claim 8, characterized in that, The method for obtaining the preset privacy budget parameter includes: In each historical knowledge sharing process of the receiver, the ratio of the number of graph nodes successfully imported by the shared data to the number of graph nodes corresponding to the shared data in the sharing agent is used as the confidence level. The privacy budget parameters are weighted using the confidence levels corresponding to the recipient's historical knowledge sharing processes, and the resulting weighted values are used as the preset privacy budget parameters for the current knowledge sharing process.
Citation Information
Patent Citations
Large-model-oriented multi-agent information processing and auxiliary analysis system
CN120181125A
Method, server and storage medium for data distribution
US20190386962A1