Multi-source heterogeneous encrypted data sharing method, equipment and medium

By building an encryption syntax tree and time-sensitive key information extraction, the problem of inefficient positioning data nodes in the encrypted state is solved, and efficient and secure encrypted data sharing is achieved.

CN120474742APending Publication Date: 2025-08-12INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510530571.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Traditional query algorithms are difficult to efficiently and accurately locate the target data nodes in an encrypted state, resulting in inefficient data sharing.

Method used

The node level is determined through the nested relationship in the data directory information, and the encryption syntax tree is built layer by layer. Key information is extracted based on the time area of shared requests and the historical request interval, and local matching is performed and global matching is performed to determine the reference encrypted data node to ensure the security and privacy of data sharing.

Benefits of technology

Improve the accuracy and efficiency of encrypted data queries, ensure security and privacy in the data sharing process, and prevent data leakage and illegal access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474742A_ABST
    Figure CN120474742A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a multi-source heterogeneous encrypted data sharing method and device and a medium, belongs to the technical field of data communication, and solves the problem of low data sharing efficiency caused by the fact that a traditional query algorithm is difficult to efficiently and accurately position a target data node in an encrypted state. Determining a data structure type and data directory information corresponding to the encrypted data; determining a node hierarchy through a nesting relationship in the data directory information, so as to hierarchically construct an encrypted syntax tree based on the data structure type and the node hierarchy; receiving a sharing request sent by a user, and performing key information extraction on the sharing request based on the time region corresponding to the sharing request and the time interval between the sharing request and the historical sharing request; based on the extracted key information, performing local matching and global matching in the encrypted syntax tree to determine a reference encrypted data node based on a matching result; and sending the encrypted data corresponding to the reference encrypted data node to the user to realize encrypted data sharing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data communication technology, and in particular to a method, device and medium for sharing multi-source heterogeneous encrypted data. Background Art

[0002] With the rapid development of information technology, data has become a core asset for businesses and society. The demand for data sharing and circulation is growing across a wide range of sectors, including healthcare, finance, and government. However, data sharing faces severe security and privacy challenges, making encryption a key means of ensuring data security.

[0003] At the data analysis level, encrypted data, due to its ciphertext form, is difficult to process directly with conventional data parsing and analysis tools. Traditional data analysis relies on feature extraction and pattern recognition of plaintext data, but encrypted data streams lose their original semantic and structural characteristics, making it difficult to identify data types, content, and relationships through conventional methods.

[0004] Traditional query algorithms are usually based on the indexing and comparison of plaintext data, which makes it difficult to efficiently and accurately locate the target data nodes in an encrypted state, resulting in inefficient data sharing. Summary of the Invention

[0005] The embodiments of the present application provide a multi-source heterogeneous encrypted data sharing method, device and medium for solving the following technical problems: traditional query algorithms are usually based on the indexing and comparison of plaintext data, which makes it difficult to efficiently and accurately locate the target data node in the encrypted state, resulting in low data sharing efficiency.

[0006] The embodiments of this application adopt the following technical solutions:

[0007] The embodiment of the present application provides a multi-source heterogeneous encrypted data sharing method. The method includes performing identifier detection on the acquired encrypted data to determine the data structure type and data directory information corresponding to the encrypted data; determining the node hierarchy through the nested relationship in the data directory information, and hierarchically constructing an encryption syntax tree based on the data structure type and node hierarchy; receiving a sharing request sent by a user, extracting key information from the sharing request based on the time zone corresponding to the sharing request and the time interval between historical sharing requests; performing local matching and global matching in the encryption syntax tree based on the extracted key information, and determining a reference encrypted data node based on the matching results; and sending the encrypted data corresponding to the reference encrypted data node to the user to achieve encrypted data sharing.

[0008] The embodiment of the present application determines the node hierarchy through the nested relationship in the data directory information, and constructs an encrypted syntax tree based on the data structure type and the node hierarchy, clearly presenting the hierarchical structure of the encrypted data, and providing an intuitive and effective model for data management and operation. By extracting key information through the time zone corresponding to the sharing request and the time interval between the historical sharing request, it is possible to grasp user needs more comprehensively and accurately, make the extracted key information more in line with the actual application scenario, and improve the efficiency and accuracy of information extraction. Secondly, local matching can carefully find the parts directly related to the key information, and global matching can grasp the relevance of the data as a whole, avoid missing important information, thereby improving the accuracy of data query and ensuring that the encrypted data node that best meets the user's needs is found. The embodiment of the present application operates on the basis of encrypted data, from data detection, syntax tree construction to processing sharing requests and final data sharing, the data is kept encrypted throughout the process, and the corresponding encrypted data is sent to the user only after the user authority verification is passed, fully guaranteeing the security and privacy of the data during the sharing process, and preventing data leakage and illegal access.

[0009] In one implementation of the present application, an encrypted syntax tree is constructed hierarchically based on the data structure type and the node level, specifically including: determining the encrypted data corresponding to each node based on the mapping relationship between the data directory information and the encrypted data; performing hash calculations on each node based on the directory information to obtain a basic hash value, and extracting reference features on each node to generate a corresponding reference hash value based on the reference features; generating an encrypted syntax tree based on the node level, the basic hash value, the reference hash value, and the encrypted data corresponding to each node; determining the sensitivity level of each node level based on the data structure type, and setting corresponding access control policies based on the sensitivity levels to perform access control on the encrypted syntax tree.

[0010] In one implementation of the present application, a corresponding access control policy is set according to the sensitivity level to perform access control on the encrypted syntax tree, specifically including: constructing a knowledge graph, analyzing the knowledge graph based on the sensitivity level, node hierarchy, data structure type and user permissions, and constructing a semantic association network; determining the association relationship between each node based on the semantic association network; performing real-time sensitivity detection on each node, and when fluctuations in sensitivity are detected, adjusting the sensitivity level of each node in the semantic association network based on the association relationship between each node; and dynamically adjusting the corresponding access control policy based on the adjusted sensitivity level to perform access control on the encrypted syntax tree.

[0011] In one implementation of the present application, a sharing request sent by a user is received, and key information of the sharing request is extracted based on the time zone corresponding to the sharing request and the time interval between the sharing request and the historical sharing request, specifically including: obtaining the historical sharing request sent by the user; filtering the historical sharing requests based on the time interval between the current sharing request and the historical sharing request to construct an associated request set; weighting the associated request set based on the time zones corresponding to the current sharing request and the historical sharing request respectively; filling in key information of the current sharing request based on the weighted associated request set; and extracting key information of the filled current sharing request through a preset key information extraction model.

[0012] In one implementation of the present application, local matching and global matching are performed in the encrypted syntax tree based on the extracted key information to determine a reference encrypted data node based on the matching results, specifically including: matching the key information with the directory information to determine the first directory information in the directory information whose matching degree is greater than a preset matching threshold; mapping the encrypted key information to a feature space that is the same as the encryption feature of the encrypted syntax tree; dividing the encrypted key information into multiple first sub-vectors, and dividing the encryption feature of the encrypted syntax tree into multiple second sub-vectors; determining the local feature similarity based on the multiple first sub-vectors and the multiple second sub-vectors, and determining the overall similarity between the encrypted mapping feature of the key information and the encryption feature of the encrypted syntax tree; determining the second directory information based on the confidence levels corresponding to the local similarity and the overall similarity; and locating the reference encrypted data node based on the first directory information and the second directory information.

[0013] In one implementation of the present application, determining the second directory information based on the confidence levels corresponding to the local similarity and the overall similarity specifically includes: according to the similarity score function:

[0014]

[0015] Determine the similarity between the key information and the encrypted syntax tree, and filter the data directory information based on the similarity to obtain the second directory information; where S is the similarity score function; β t is the dynamic attenuation coefficient that changes with time t; α is the resistance factor; c global is the overall confidence; s global is the overall similarity; h ij is the level weight; r ij is the correlation between nodes; is the local similarity; is the local confidence; i is the element serial number corresponding to the key information; j is the element serial number corresponding to the syntax tree.

[0016] In one implementation of the present application, the encrypted data corresponding to the reference encryption data node is sent to the user, specifically including: determining the data volume of the encrypted data corresponding to the reference encryption data node; when the data volume is greater than a preset data volume threshold, scanning the encrypted data corresponding to the reference encryption data node through a sliding window to determine the data statistical characteristics corresponding to the encrypted data in the window; determining multiple data block boundaries based on the data statistical characteristics corresponding to different windows; dividing the encrypted data corresponding to the reference encryption data node into multiple data blocks based on the multiple data block boundaries; and assigning a corresponding transmission path to each data block to send the encrypted data corresponding to the reference encryption data node to the user through multiple transmission paths.

[0017] In one implementation of the present application, each data block is assigned a corresponding transmission path, specifically including: determining performance data corresponding to each transmission path in the network; wherein the performance data includes at least one of the bandwidth, delay time, and packet loss rate corresponding to each transmission path; constructing a network status model based on the performance data and the network topology; determining a set of paths for transmitting each data block to a user end in the network status model using a preset shortest path algorithm; sorting each path in the path set based on the path distance to the user end to generate a path sequence; dynamically adjusting the path sequence based on the current load status of each path, and assigning a corresponding transmission path to each data block based on the dynamically adjusted path sequence.

[0018] An embodiment of the present application provides a multi-source heterogeneous encrypted data sharing device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so as to enable the at least one processor to: perform identifier detection on the acquired encrypted data to determine the data structure type and data directory information corresponding to the encrypted data; determine the node hierarchy through the nested relationship in the data directory information, and construct an encrypted syntax tree in layers based on the data structure type and the node hierarchy; receive a sharing request sent by a user, and extract key information of the sharing request based on the time zone corresponding to the sharing request and the time interval between the sharing request and the historical sharing request; perform local matching and global matching in the encrypted syntax tree based on the extracted key information, and determine a reference encrypted data node based on the matching result; and send the encrypted data corresponding to the reference encrypted data node to the user to realize encrypted data sharing.

[0019] A non-volatile computer storage medium provided by an embodiment of the present application stores computer-executable instructions, which are configured to: perform identifier detection on the acquired encrypted data to determine the data structure type and data directory information corresponding to the encrypted data; determine the node hierarchy through the nested relationship in the data directory information, and construct a hierarchical encryption syntax tree based on the data structure type and the node hierarchy; receive a sharing request sent by a user, and extract key information of the sharing request based on the time zone corresponding to the sharing request and the time interval between the sharing request and the historical sharing request; perform local matching and global matching in the encryption syntax tree based on the extracted key information, and determine a reference encrypted data node based on the matching result; and send the encrypted data corresponding to the reference encrypted data node to the user to realize encrypted data sharing.

[0020] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: The embodiments of the present application determine the node hierarchy through the nested relationship in the data directory information, and construct an encrypted syntax tree based on the data structure type and the node hierarchy, clearly presenting the hierarchical structure of the encrypted data, and providing an intuitive and effective model for data management and operation. By extracting key information through the time zone corresponding to the sharing request and the time interval between historical sharing requests, it is possible to more comprehensively and accurately grasp user needs, make the extracted key information more in line with actual application scenarios, and improve the efficiency and accuracy of information extraction. Secondly, local matching can carefully find the parts directly related to the key information, while global matching can grasp the relevance of the data as a whole, avoid missing important information, thereby improving the accuracy of data query and ensuring that the encrypted data node that best meets the user's needs is found. The embodiments of the present application operate on the basis of encrypted data, from data detection, syntax tree construction to processing sharing requests and final data sharing, the data is kept encrypted throughout the process. The corresponding encrypted data will only be sent to the user after the user's permission verification is passed, fully ensuring the security and privacy of the data during the sharing process, preventing data leakage and illegal access. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments described in the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:

[0022] Figure 1 A flowchart of a multi-source heterogeneous encrypted data sharing method provided in an embodiment of the present application;

[0023] Figure 2A schematic diagram of the structure of a multi-source heterogeneous encrypted data sharing device provided in an embodiment of the present application.

[0024] Reference numerals:

[0025] 200: Multi-source heterogeneous encrypted data sharing device, 201: Processor, 202: Memory. DETAILED DESCRIPTION

[0026] Embodiments of the present application provide a method, device, and medium for sharing multi-source heterogeneous encrypted data.

[0027] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0028] The technical solutions proposed in the embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0029] Figure 1 A flow chart of a multi-source heterogeneous encrypted data sharing method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the multi-source heterogeneous encrypted data sharing method includes the following steps:

[0030] S101: Perform identifier detection on the obtained encrypted data to determine the data structure type and data directory information corresponding to the encrypted data.

[0031] In one implementation of the present application, the encrypted data in the embodiment of the present application is information processed by an encryption algorithm, the purpose of which is to protect data privacy and security, and prevent unauthorized access and data leakage. An identifier is a special mark or information fragment in the data used to identify data characteristics, structure or content. In different data types, the expression of identifiers varies. For example, in structured database table data, table names, column names, specific constraint keywords, etc. are all important identifiers; in semi-structured XML or JSON data, tag names, attribute names, etc. can be used as identifiers. Perform identifier detection on it, and judge the data structure type of the encrypted data based on the characteristics and combination of the identifiers. If an identifier in the format of an SQL statement, as well as elements such as table names and column names, is detected, it can be judged as structured database table data; when an identifier in the tag format of XML or JSON is found, such as <root>, {"key":"value"}, etc., can be determined as semi-structured data.

[0032] Furthermore, data directory information describes the organization and content of the data, documenting its various components and their interrelationships. During the process of determining the data structure type, relevant identifier information is simultaneously extracted to construct the data directory. For structured data, data directory information includes table names, column names, data types, and constraints. Directory information for semi-structured data covers things like the label hierarchy and attribute definitions. Directory information for unstructured data may include basic information such as the file's name, size, creation time, and storage path, as well as descriptive information obtained through metadata extraction techniques. This data directory information acts as an "index" for the data, helping to quickly locate and understand its content.

[0033] S102. Determine the node level through the nested relationship in the data directory information, and construct an encrypted syntax tree hierarchically based on the data structure type and the node level.

[0034] In one implementation of the present application, the encrypted data corresponding to each node is determined based on the mapping relationship between data directory information and encrypted data. A hash calculation is performed on each node based on the directory information to obtain a basic hash value, and a reference feature is extracted for each node to generate a corresponding reference hash value based on the reference feature. An encrypted syntax tree is generated based on the node hierarchy, the basic hash value, the reference hash value, and the encrypted data corresponding to each node. Based on the data structure type, the sensitivity level of each node hierarchy is determined, and the corresponding access control policy is set according to the sensitivity level to control access to the encrypted syntax tree.

[0035] Specifically, data directory information describes the organizational structure and hierarchical relationships of the data, including multiple levels of directories and subdirectories, forming a nested structure. By analyzing this nested relationship, the hierarchical position of each data element in the entire data system can be determined. For example, in a file system's data directory, the root directory may have multiple first-level subdirectories, each of which may contain multiple second-level subdirectories, and so on. By identifying these parent-child relationships, a node hierarchy can be assigned to each directory or file.

[0036] Furthermore, data directory information not only describes the data's hierarchical structure but also maps to specific encrypted data. This mapping clearly identifies which encrypted data each directory or node corresponds to. For example, in an enterprise's data management system, a department's data directory might contain different types of encrypted data, such as employee information and project data. This mapping identifies which specific encrypted data files or records correspond to each department's directory node. This way, each node is associated with specific encrypted data, providing a basis for subsequent management and manipulation of encrypted data.

[0037] Furthermore, identical directory information will generate identical hash values, while different directory information will almost never generate identical hash values. By calculating the base hash value, nodes can be easily identified and verified. Reference features can include the node's data type, data size, creation time, modification time, and so on. After extracting these features, they are also processed using a hash algorithm to generate a reference hash value. The reference hash value provides more information about the node, facilitating more detailed node differentiation and management. For example, when searching for a specific type of node, you can quickly filter out eligible nodes by comparing the reference hash value.

[0038] Furthermore, the encryption syntax tree in the embodiment of the present application is based on the node hierarchy, and each node is organized according to its hierarchical relationship. Each node contains its basic hash value, reference hash value and corresponding encrypted data. The encrypted data is stored in the corresponding node, ensuring the security of the data. When it is necessary to access or process encrypted data, the required nodes and data can be located by traversing the encryption syntax tree. Different types of data structures have different levels of sensitivity. For example, data containing personal privacy information usually has a higher sensitivity level, while some public statistical data may have a lower sensitivity level. According to the data structure type, a corresponding sensitivity level can be determined for each node level. For example, for the node level containing user personal information, its sensitivity level can be set to high; for the node level of some system configuration information, its sensitivity level can be set to medium; for the node level of some public documents, its sensitivity level can be set to low.

[0039] Furthermore, access control policies are used to restrict access rights to nodes and data in the encrypted syntax tree. Different access control policies can be set based on the sensitivity level of each node level.

[0040] In one implementation of the present application, the sensitivity level corresponding to each node is determined based on the data structure type. A knowledge graph is constructed, and based on the sensitivity level, node hierarchy, data structure type, and user permissions, the knowledge graph is analyzed to construct a semantic association network. Based on the semantic association network, the association relationship between each node is determined. The sensitivity of each node is detected in real time. When fluctuations in sensitivity are detected, the sensitivity level of the nodes in the semantic association network is adjusted based on the association relationship between each node. Based on the adjusted sensitivity level, the corresponding access control policy is dynamically adjusted to perform access control on the encrypted syntax tree.

[0041] Specifically, identifier detection is used to determine the data structure type corresponding to the encrypted data, such as structured database tables, semi-structured XML / JSON data, unstructured documents or multimedia data. Based on this, combined with the business attributes and content characteristics of the data, an initial sensitivity level is assigned to each node. Taking each node as an entity, the node's sensitivity level, node hierarchy, data structure type, and user permissions are used as attributes of the entity. At the same time, the relationship between entities is defined to construct a knowledge graph. Graph algorithms and natural language processing technologies are used to conduct in-depth analysis of the knowledge graph. Potential semantic connections are mined through node attributes and relationships. For example, it is found that nodes with the same sensitivity level and data structure type may be associated in business logic; nodes at the same level and accessed by the same user permissions are also closely connected. These associations are further abstracted and integrated to construct a semantic association network, which intuitively presents the association between each node at the semantic level.

[0042] Furthermore, in the semantic association network, path analysis, cluster analysis and other methods are used to clarify the specific association relationship between each node. Based on the node attributes and association characteristics, similar nodes are clustered into one category, and nodes of the same category have a strong association relationship. The data changes of each node are monitored in real time. For example, the sensitivity of the node is judged by indicators such as data access frequency, data update operation, and data flow direction. If a node originally had a low access frequency and suddenly a large number of abnormal access requests appear, it may mean that its sensitivity has increased. When a fluctuation in the sensitivity of a node is detected, the sensitivity level of other related nodes in the semantic association network is adjusted based on the association relationship between the nodes previously determined. If the sensitivity level of node A is increased, and there is a strong association between node A and node B, the sensitivity level of node B will also be increased accordingly.

[0043] Furthermore, based on the adjusted sensitivity level, the access control policy is re-established or modified to ensure that only users with the appropriate permissions can access nodes with the corresponding sensitivity level. The new access control policy is applied to the encrypted syntax tree. When a user initiates an access request to a node, the permission is verified according to the latest policy. If the user's permissions do not meet the policy requirements, access is denied. If they do, the user is allowed to obtain the corresponding encrypted data, thus achieving dynamic and secure access control over the encrypted syntax tree.

[0044] S103: Receive a sharing request sent by a user, and extract key information from the sharing request based on a time zone corresponding to the sharing request and a time interval between the sharing request and historical sharing requests.

[0045] In one implementation of the present application, historical sharing requests sent by a user are obtained. Based on the time interval between the current sharing request and the historical sharing requests, the historical sharing requests are filtered to construct a set of associated requests. The associated request set is weighted based on the time zones corresponding to the current sharing request and the historical sharing requests, respectively. Based on the weighted associated request set, key information is populated into the current sharing request. Key information is extracted from the populated current sharing request using a preset key information extraction model.

[0046] Specifically, the embodiment of the present application is provided with a special request storage module, which records the sharing request information sent by the user each time in chronological order, including detailed data such as request content, initiation time, and request results. When a new sharing request needs to be processed, all historical sharing request data is retrieved from the storage module. A time interval threshold is set, and the initiation time of the current sharing request is compared with the initiation time of each historical sharing request to calculate the time interval. If the time interval is greater than the time interval threshold, it means that the historical request is too long from the current time and has a weak correlation with the current request. If the time interval is not greater than the time interval threshold, the historical sharing request is included in the associated request set. Furthermore, the time is divided into different areas, such as "short term", "medium term" and "long term", and each area corresponds to a different weight coefficient. According to the initiation time of the historical sharing request, its time area is judged and the corresponding weight is assigned.

[0047] Furthermore, the associated request set is traversed to extract the key information in each historical request, and for each historical request, its key information is merged with the current shared request according to the corresponding weight. If the current request already has some key information, it is merged using strategies such as weighted average or priority coverage. The filled current shared request is input into the preset key information extraction model, and the model automatically identifies and extracts the key information therein, and outputs a structured key information set. Among them, the preset key information extraction model in the embodiment of the present application is obtained by training the neural network with historical data. For example, for the filled shared request "Get the sales data report for September 2024 (supplemented based on historical requests)", the model extracts key information such as the data time range "September 2024", the data type "sales data", and the data presentation form "report", providing an accurate basis for subsequent matching and retrieval in the encrypted syntax tree.

[0048] S104: Based on the extracted key information, perform local matching and global matching in the encryption syntax tree to determine a reference encryption data node based on the matching results.

[0049] In one implementation of the present application, key information is matched with directory information to determine first directory information in the directory information whose matching degree is greater than a preset matching threshold. The encrypted key information is mapped to a feature space that is the same as the encryption feature of the encrypted syntax tree. The encrypted key information is divided into multiple first sub-vectors, and the encryption feature of the encrypted syntax tree is divided into multiple second sub-vectors. Based on the multiple first sub-vectors and the multiple second sub-vectors, the local feature similarity is determined, and the overall similarity between the encrypted mapping feature of the key information and the encryption feature of the encrypted syntax tree is determined. The second directory information is determined based on the confidence levels corresponding to the local similarity and the overall similarity. Based on the first directory information and the second directory information, the reference encrypted data node is located.

[0050] Specifically, the data directory information is structured and parsed to extract key content such as core identifiers and attributes. For example, for the directory information of a database table, the table name, column name, data type, etc. are extracted; for the directory information of a file system, the file name, file type, storage path, etc. are extracted. A variety of matching algorithms are used to calculate the matching degree between key information and directory information. For example, for text information, the edit distance algorithm can be used to calculate the similarity between character strings, or the word vector model can be used to calculate the semantic similarity; for numerical or structured information, the matching degree is calculated by comparing attribute values, data ranges, etc. A preset matching threshold is set, and only directory information with a matching degree higher than the threshold is determined as the first directory information.

[0051] Furthermore, the feature space used by the encryption feature of the encrypted syntax tree is clarified, and the feature space is composed of multiple attributes and feature dimensions of the data. The encrypted key information is input into the trained mapping model, and the model output is the result of mapping to the same feature space as the encryption feature of the encrypted syntax tree. For example, for a section of encrypted key information text describing product sales data, after being processed by the mapping model, it is converted into a feature vector containing dimensions such as the sales amount range and product category distribution, which is consistent with the feature space form of the corresponding sales data node in the encrypted syntax tree. Among them, the mapping model in the embodiment of the present application is based on deep learning or machine learning algorithms. The model is trained with a large amount of training data to learn the mapping relationship between key information and the features of the encrypted syntax tree. During the training process, the encryption features of the encrypted syntax tree are used as the target, and the parameters of the mapping model are adjusted so that the encrypted key information can be as close as possible to the representation of the feature space after being processed by the model.

[0052] Furthermore, the sub-vector division method and length are determined based on the characteristics of the data and computational requirements. For structured data, the division can be performed by columns or data blocks; for text data, the division can be performed by sentences, paragraphs, or fixed-length character sequences. The encrypted key information is divided into multiple first sub-vectors according to the determined strategy, and the encrypted features of the encrypted syntax tree are similarly divided into multiple second sub-vectors. For example, for a structured data encryption feature containing multiple columns, the data in each column is treated as a sub-vector; for a longer text key information, the division is performed based on a sub-vector consisting of every 100 characters.

[0053] Furthermore, for each set of corresponding first sub-vectors and second sub-vectors, the local feature similarity between them is determined. Based on all first sub-vectors and second sub-vectors, the similarity between the encrypted mapping features of the key information and the encrypted features of the encrypted syntax tree is comprehensively evaluated as a whole.

[0054] Furthermore, confidence levels are calculated for the local similarity and the overall similarity, respectively. Based on the local similarity, the overall similarity and their corresponding confidence levels, a comprehensive decision rule is formulated to determine the second directory information.

[0055] Specifically, according to the similarity score function:

[0056]

[0057] Determining the similarity between the key information and the encrypted syntax tree, and filtering the data directory information based on the similarity to obtain second directory information;

[0058] Among them, S is the similarity score function; β t is the dynamic attenuation coefficient that changes with time t; α is the resistance factor; c global is the overall confidence; s global is the overall similarity; h ij is the level weight; r ij is the correlation between nodes; is the local similarity; is the local confidence; i is the element serial number corresponding to the key information; j is the element serial number corresponding to the syntax tree.

[0059] Furthermore, the first directory information and the second directory information are integrated to further filter out the node range most likely to contain the target data. In the encrypted syntax tree, based on the filtered directory information, the reference encrypted data node is gradually located according to the tree structure and node identification.

[0060] S105: Send the encrypted data corresponding to the reference encrypted data node to the user to achieve encrypted data sharing.

[0061] In one implementation of the present application, the amount of encrypted data corresponding to a reference encryption data node is determined. When the data amount is greater than a preset data amount threshold, the encrypted data corresponding to the reference encryption data node is scanned using a sliding window to determine statistical characteristics of the encrypted data within the window. Based on the statistical characteristics of the data corresponding to different windows, multiple data block boundaries are determined. Based on the multiple data block boundaries, the encrypted data corresponding to the reference encryption data node is divided into multiple data blocks. A corresponding transmission path is assigned to each data block, so that the encrypted data corresponding to the reference encryption data node is sent to the user via the multiple transmission paths.

[0062] Specifically, after locating a reference encrypted data node, the system obtains the size of the encrypted data stored there. The system pre-sets a data volume threshold and compares the volume of the encrypted data corresponding to the reference encrypted data node with this threshold. If the data volume is less than or equal to the threshold, conventional transmission methods are used. If the data volume is greater than the threshold, subsequent data processing is triggered to optimize the transmission process.

[0063] Furthermore, when the data volume exceeds a preset threshold, a sliding window technique is used to scan the encrypted data. The sliding window size and sliding step size are set, and the window size should be appropriately set based on the data characteristics and computing resources. Starting from the starting position of the encrypted data, the window is moved according to the set step size, and statistical features of the encrypted data within each window are calculated. For example, the statistical features in the embodiments of the present application can be mean, variance, and data distribution frequency.

[0064] Furthermore, the statistical characteristics of the data corresponding to different windows are analyzed, and the data block boundaries are determined based on the locations where the statistical characteristics of the data change significantly. For example, when the mean or variance of adjacent windows fluctuates greatly, or the data distribution frequency changes significantly, it is considered that this may be the dividing point of different data blocks. Based on the determined data block boundaries, the encrypted data corresponding to the reference encrypted data node is divided into multiple data blocks. Each data block contains the complete data content from one boundary to the next boundary, ensuring the integrity and continuity of the data. Based on the size and importance of each data block and the evaluation results of the network transmission path, a corresponding transmission path is assigned to each data block. For larger data blocks, paths with high bandwidth and good stability are given priority; for data blocks with high importance, paths with strong reliability and low latency are selected.

[0065] In one implementation of the present application, performance data corresponding to each transmission path in the network is determined; wherein the performance data includes at least one of the bandwidth, delay time, and packet loss rate corresponding to each transmission path. Based on the performance data and the network topology, a network status model is constructed. By presetting the shortest path algorithm, a set of paths for transmitting each data block to the user end is determined in the network status model. Based on the path distance from each path in the path set to the user end, each path is sorted to generate a path sequence. Based on the current load status of each path, the path sequence is dynamically adjusted, and a corresponding transmission path is assigned to each data block based on the dynamically adjusted path sequence.

[0066] Specifically, we use the collected performance data to construct a model that accurately reflects the current state of the network. We abstract the nodes in the network as vertices in a graph, and the transmission paths as edges. Each edge is accompanied by corresponding performance data, such as bandwidth, latency, and packet loss rate. The entire network can be represented as a weighted directed graph. This model clearly demonstrates the status of each path in the network and their relationships.

[0067] Furthermore, within the constructed network status model, a preset shortest path algorithm is used to find possible paths for transmitting each data block to the user end. After obtaining the path set, the paths are sorted according to their distance to the user end to form a path sequence. Because the network load situation changes dynamically, even if a path performs well in the evaluation, if it is currently under high load, it is not suitable for allocating data blocks for transmission. Therefore, it is necessary to monitor the load situation of each path in real time, for example, by checking the port utilization, traffic statistics, and other information of the router. The path sequence is adjusted according to the load situation, and paths with excessive load are placed at the back, while paths with lower load and better performance are placed at the front, to ensure that data blocks can be allocated to the most appropriate paths.

[0068] After dynamic adjustment, a transmission path is assigned to each data block according to the new path sequence, so that network resources can be fully utilized, the efficiency and stability of data transmission can be improved, and problems such as transmission delay and packet loss caused by improper path selection can be avoided.

[0069] Figure 2 This is a schematic diagram of the structure of a multi-source heterogeneous encrypted data sharing device provided in an embodiment of the present application. Figure 2 As shown, a multi-source heterogeneous encrypted data sharing device 200 includes: at least one processor 201; and a memory 202 communicatively connected to the at least one processor 201; wherein the memory 202 stores instructions that can be executed by the at least one processor 201, and the instructions are executed by the at least one processor 201 so that the at least one processor 201 can: perform identifier detection on the acquired encrypted data to determine the data structure type and data directory information corresponding to the encrypted data; determine the node hierarchy through the nested relationship in the data directory information, and construct an encrypted syntax tree in layers based on the data structure type and the node hierarchy; receive a sharing request sent by a user, and extract key information of the sharing request based on the time zone corresponding to the sharing request and the time interval between the sharing request and the historical sharing request; perform local matching and global matching in the encrypted syntax tree based on the extracted key information to determine a reference encrypted data node based on the matching result; send the encrypted data corresponding to the reference encrypted data node to the user to realize encrypted data sharing.

[0070] A non-volatile computer storage medium provided by an embodiment of the present application stores computer-executable instructions, which are configured to: perform identifier detection on the acquired encrypted data to determine the data structure type and data directory information corresponding to the encrypted data; determine the node hierarchy through the nested relationship in the data directory information, and construct a hierarchical encryption syntax tree based on the data structure type and the node hierarchy; receive a sharing request sent by a user, and extract key information of the sharing request based on the time zone corresponding to the sharing request and the time interval between the sharing request and the historical sharing request; perform local matching and global matching in the encryption syntax tree based on the extracted key information, and determine a reference encrypted data node based on the matching result; and send the encrypted data corresponding to the reference encrypted data node to the user to realize encrypted data sharing.

[0071] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, apparatus, and non-volatile computer storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.

[0072] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. It will be apparent to those skilled in the art that various modifications and variations may be made to the embodiments of the present application. However, such modifications or substitutions do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.< / root>

Claims

1. A multi-source heterogeneous encrypted data sharing method, characterized in that: The method comprises: Performing identifier detection on the obtained encrypted data to determine the data structure type and data directory information corresponding to the encrypted data; Determining the node hierarchy through the nested relationship in the data directory information, and constructing the encryption syntax tree hierarchically based on the data structure type and the node hierarchy; receiving a sharing request sent by a user, and extracting key information from the sharing request based on a time zone corresponding to the sharing request and a time interval between the sharing request and historical sharing requests; Based on the extracted key information, performing local matching and global matching in the encryption syntax tree to determine a reference encryption data node based on the matching results; The encrypted data corresponding to the reference encrypted data node is sent to the user to achieve encrypted data sharing.

2. A multi-source heterogeneous encrypted data sharing method according to claim 1, characterized in that: The hierarchical construction of the encryption syntax tree based on the data structure type and the node level specifically includes: Determining the encrypted data corresponding to each node according to the mapping relationship between the data directory information and the encrypted data; Performing hash calculations on each node based on the directory information to obtain a basic hash value, and extracting reference features from each node to generate a corresponding reference hash value based on the reference features; generating the encrypted syntax tree based on the node hierarchy, the basic hash value, the reference hash value, and the encrypted data corresponding to each node; According to the data structure type, the sensitivity level of each node level is determined, and a corresponding access control policy is set according to the sensitivity level to perform access control on the encrypted syntax tree.

3. A multi-source heterogeneous encrypted data sharing method according to claim 2, characterized in that: The step of setting a corresponding access control policy according to the sensitivity level to perform access control on the encrypted syntax tree specifically includes: Constructing a knowledge graph, analyzing the knowledge graph based on the sensitivity level, the node hierarchy, the data structure type, and user permissions, and constructing a semantic association network; Determining the association relationship between nodes based on the semantic association network; Performing sensitivity detection on each node in real time, and adjusting the sensitivity level of each node in the semantic association network based on the association relationship between the nodes when fluctuations in sensitivity are detected; Based on the adjusted sensitivity level, the corresponding access control policy is dynamically adjusted to perform access control on the encrypted syntax tree.

4. A multi-source heterogeneous encrypted data sharing method according to claim 1, characterized in that: The receiving of the sharing request sent by the user and extracting key information from the sharing request based on a time zone corresponding to the sharing request and a time interval between the sharing request and historical sharing requests may specifically include: Obtaining historical sharing requests sent by the user; filtering the historical sharing requests based on a time interval between the current sharing request and the historical sharing requests to construct an associated request set; weighting the associated request set based on time zones corresponding to the current sharing request and the historical sharing request; Filling key information of the current sharing request based on the weighted association request set; By presetting a key information extraction model, key information is extracted from the filled current sharing request.

5. A multi-source heterogeneous encrypted data sharing method according to claim 1, characterized in that: The step of performing local matching and global matching in the encrypted syntax tree based on the extracted key information to determine a reference encrypted data node based on the matching results specifically includes: Matching the key information with the directory information to determine first directory information in the directory information whose matching degree is greater than a preset matching threshold; Mapping the encrypted key information to a feature space identical to the encryption feature of the encrypted syntax tree; Dividing the encrypted key information into a plurality of first sub-vectors, and dividing the encrypted features of the encrypted syntax tree into a plurality of second sub-vectors; Determining local feature similarities based on the plurality of first sub-vectors and the plurality of second sub-vectors, and determining overall similarities between the encrypted mapping features of the key information and the encrypted features of the encrypted syntax tree; Determining second directory information based on the confidence levels respectively corresponding to the local similarity and the overall similarity; Based on the first directory information and the second directory information, the reference encrypted data node is located.

6. A multi-source heterogeneous encrypted data sharing method according to claim 5, characterized in that: The determining of the second directory information based on the confidences respectively corresponding to the local similarity and the overall similarity specifically includes: According to the similarity score function: Determining a similarity between the key information and the encrypted syntax tree, and filtering the data directory information based on the similarity to obtain the second directory information; Among them, S is the similarity score function; βt is the dynamic attenuation coefficient that changes with time t; α is the confrontation factor; c global is the overall confidence; s global is the overall similarity; h ij is the level weight; r ij is the correlation between nodes; is the local similarity; is the local confidence; i is the element serial number corresponding to the key information; j is the element serial number corresponding to the syntax tree.

7. A multi-source heterogeneous encrypted data sharing method according to claim 1, characterized in that: The sending the encrypted data corresponding to the reference encrypted data node to the user specifically includes: Determining the amount of encrypted data corresponding to the reference encrypted data node; When the data volume is greater than a preset data volume threshold, the encrypted data corresponding to the reference encrypted data node is scanned through a sliding window to determine the data statistical features corresponding to the encrypted data in the window; Determining a plurality of data block boundaries based on the data statistical features corresponding to different windows; Based on the plurality of data block boundaries, the encrypted data corresponding to the reference encrypted data node is divided into a plurality of data blocks; A corresponding transmission path is allocated to each of the data blocks, so that the encrypted data corresponding to the reference encrypted data node is sent to the user through the multiple transmission paths.

8. A multi-source heterogeneous encrypted data sharing method according to claim 7, characterized in that: The step of allocating a corresponding transmission path to each of the data blocks specifically includes: Determining performance data corresponding to each transmission path in the network; wherein the performance data includes at least one of bandwidth, delay time, and packet loss rate corresponding to each transmission path; Building a network status model based on the performance data and the network topology; Determining, in the network state model, a set of paths for transmitting each of the data blocks to the user end by using a preset shortest path algorithm; sorting the paths based on the path distances from each path in the path set to the user end to generate a path sequence; Based on the current load conditions of each path, the path sequence is dynamically adjusted, and a corresponding transmission path is allocated to each data block based on the dynamically adjusted path sequence.

9. A multi-source heterogeneous encrypted data sharing device, characterized in that: The device comprises a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions can execute the method according to any one of claims 1 to 8.