Data management method and related equipment
By building an isolated operating environment and data management system, using data reading policies and permission control, the problems of abuse and leakage during data use are solved, and the transparency and security management of data use are achieved.
Patent Information
- Application Number
- CN202410532731.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-11
- Filing Date
- 2024-04-29
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to effectively manage the data usage process provided by data providers, resulting in the data being abused or leaked, and it is impossible to ensure the reasonable and safe use of data users.
Build an isolated operation environment, and use data reading policies and permission policies to control the access and use behavior of data users through execution nodes and storage nodes in the data management system, including data permission policies, data reading policies, access credential management and derived data tracking to ensure that data is used in the isolated environment.
It realizes transparent management of the data usage process, reduces the risk of data leakage, ensures the safe and reasonable use of data providers, and improves the transparency and controllability of data usage.
Smart Images

Figure CN120296722A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the Chinese application number 202410044120.5 and the invention title "A Data Management Method, System, Device, and Storage Medium", which was filed with the National Intellectual Property Administration on January 11, 2024. The entire content thereof is incorporated herein by reference. Technical Field
[0002] This application relates to the field of data processing technologies, and particularly to a data management method and related devices. Background Art
[0003] With the continuous development of information technology, data has a large number of application requirements in various business fields, and the purchase of data also appears in various application scenarios. For example, the performance improvement of artificial intelligence (AI) models fully depends on high-quality training data. Currently, public data and private data are difficult to meet the training needs of AI models. Therefore, AI developers (i.e., data users) need to purchase a large amount of high-quality data from third parties (i.e., data providers).
[0004] Currently, data providers usually clarify the data usage rights through forms such as written agreements. However, the data usage process by data users is usually not transparent enough to data providers, resulting in the data provided by data providers may be misused or even leaked.
[0005] Therefore, there is an urgent need for a method to manage the usage process of the data provided by data providers to ensure that the data provided by data providers is used safely and reasonably by data users. Summary of the Invention
[0006] The embodiments of this application provide a data management method, which can manage the usage process of the data provided by data providers to ensure that the data provided by data providers is used safely and reasonably by data users. This application also provides corresponding devices, equipment, computer-readable storage media, computer program products, etc.
[0007] The first aspect of the present application provides a data management method. This method is applied to a data management system, which is set up in at least one data center and located in an isolated operating environment. The isolated operating environment is connected to a data provider and a data user respectively through an application programming interface (API). The data management system includes an execution node and a storage node, and the storage node stores third-party data of the data provider. The method includes: the execution node runs a business program, and the business program is used to execute the business of the data user; the execution node sends a data reading request to the storage node, and the data reading request indicates that the running business program requests to read first data from the storage node, and the first data belongs to the third-party data; the storage node generates feedback information according to the data reading request and the data reading policy regarding the third-party data, and sends the feedback information to the execution node.
[0008] In the first aspect, since the data management system is located in an isolated operating environment, it can make it difficult for data users to directly transfer the third-party data provided by data providers outside the isolated operating environment, effectively restricting the use of the third-party data within the isolated operating environment and avoiding the leakage of the third-party data provided by data providers.
[0009] Moreover, the third-party data is stored in the storage node of the data management system. During the process that the execution node of the data management system runs a business program to execute the business of the data user and needs to use the third-party data, the execution node needs to send a data reading request to the storage node. In this way, the storage node can control the reading of the third-party data by the business program based on the data reading policy, enabling the use of the third-party data by the data user through the business program to be monitored by the storage node. Thus, the storage node manages the usage process of the third-party data provided by the data provider to ensure that the third-party data is used safely and reasonably by the data user and reduce the risk of data leakage.
[0010] In a possible implementation manner of the first aspect, the method further includes: monitoring the access behavior of the data user to the data management system through the API, and monitoring the data output behavior of the data management system through the API.
[0011] In this possible implementation manner, the API can be managed through an API gateway or other forms of API management tools to monitor the data interaction related to the API through the API gateway. In this way, through the API gateway, etc., the access behavior of external nodes (data users) to internal nodes (data management system) through the API can be monitored, and the data output behavior of the data management system through the API can be monitored.
[0012] Among them, access behavior can be monitored by one or more of monitoring access frequency, corresponding response size, response content, etc., so that abnormal access behavior can be identified. And the monitoring data output behavior can be to detect the similarity between the data to be output by this data output behavior and the third-party data provided by the data provider, so as to avoid the leakage of third-party data.
[0013] In a possible implementation manner of the first aspect, the data reading policy is determined by the data provider. The data reading policy includes policy items and operations determined based on the policy items. The policy items include one or more of the following items: data policy items, usage policy items, quantity policy items, time policy items, and subject policy items; among them, the data policy items are used to describe the data range in the third-party data that can be read; the usage policy items are used to describe the tasks in the business program that can use the third-party data; the quantity policy items are used to describe the usage quantity threshold of the third-party data; the time policy items are used to describe the usage time range of the third-party data; the subject policy items are used to describe the subject types to which the data reading policy can be applied.
[0014] In this possible implementation manner, the data reading policy can be flexibly configured from one or more aspects through one or more policy items. The configuration method is clear and easy to implement, which can improve the configuration efficiency, facilitate the understanding of data users and data providers, and thus facilitate the negotiation and unification between data users and data providers.
[0015] In a possible implementation manner of the first aspect, the storage node includes one or more data reading policies regarding the third-party data; the storage node generates feedback information according to the data reading request and the data reading policy regarding the third-party data, and sends the feedback information to the execution node, including: the storage node matches the data reading request with the policy items of at least one data reading policy among the one or more data reading policies; if there is a data reading policy that matches the data reading request, the storage node generates feedback information according to the operation in the data reading policy that matches the data reading request and sends the feedback information to the execution node.
[0016] In this possible implementation manner, the number of data reading policies to be matched can be one or more. When the number of data reading policies to be matched is multiple, the multiple data reading policies can be matched in sequence. Specifically, the matching can be performed according to the policy items in the data reading policy. Match the multiple data reading policies in sequence until there is a matching data reading policy, and generate feedback information according to the operation in the data reading policy, or until it is determined that all data reading policies do not match, then feedback information can be generated according to the default operation, or feedback information indicating the rejection of data reading can be generated.
[0017] In a possible implementation of the first aspect, the storage node further includes a data permission policy, which is used to determine the read and write permissions for data according to the source of the data. Among them, the data permission policy indicates that the data user has read permission for third-party data from the data provider; before the storage node generates feedback information based on the first data and the data reading policy for the third-party data and sends the feedback information to the execution node, it further includes: the storage node determines, according to the data permission policy, that the business program has read permission for the first data based on the first data belonging to third-party data.
[0018] In this possible implementation, the data permission policy can be a system-level policy of the data management system, so that the read and write permissions of the data in the storage node can be basically managed through this system-level policy. The data permission policy may include indication information indicating that the data user has read permission for third-party data from the data provider. In this way, after it is determined that the first data has read permission, data reading can be performed according to the data reading policy.
[0019] It can be seen that in this possible implementation, the reading operation of the data in the storage node can be more perfectly managed through at least two levels of policies.
[0020] In a possible implementation of the first aspect, the data management system includes a management node; the method further includes: after the management node obtains the running instruction for the business program, it creates an access credential for the storage node; the management node instructs the execution node to run the business program and transmits the information of the access credential to the execution node; the execution node sends a data reading request to the storage node, including: the execution node sends a data reading request to the storage node based on the information of the access credential.
[0021] In this possible implementation, the management node can manage the access credential of the execution node to the storage node, that is to say, the management node can manage the permission for the execution node to interact with the storage node. In this way, the execution node cannot obtain the access permission to the storage node by itself, avoiding the execution node from reading and using the third-party data in the storage node without permission, thus avoiding the leakage and unreasonable use of the third-party data by the execution node.
[0022] In a possible implementation of the first aspect, the method further includes: after the execution node stops running the business program, deleting the access credential.
[0023] In this possible implementation, deleting the access credential after the execution node stops running the business program can ensure the validity of the access credential.
[0024] In this way, each time the execution node starts the operation of the business program, it needs to obtain the permission to interact with the storage node from the management node again. In this way, the execution node cannot obtain the access permission to the storage node by itself, nor can it abuse the obtained access permission, effectively ensuring the controllability of each access of the execution node to the storage node, and avoiding the execution node from reading and using the third-party data of the storage node without permission, thus avoiding the leakage and unreasonable use of the third-party data by the execution node.
[0025] In a possible implementation manner of the first aspect, after the storage node generates feedback information according to the first data and the data reading policy regarding the third-party data and sends the feedback information to the execution node, it further includes: the storage node records the data access behavior related to the data reading request.
[0026] In this possible implementation manner, after the storage node stores the information of the data access behavior, the data provider can conveniently and flexibly view the usage situation of the third-party data stored in the storage node by it.
[0027] In a possible implementation manner of the first aspect, the method further includes: the storage node obtains a query request input by the data provider, where the query request is used to request to query the access records related to the third-party data; the storage node outputs the access records, and the access records include the data access behavior related to the data reading request.
[0028] In this possible implementation manner, the data provider can conveniently query the access and usage situation of the third-party data by the data user through the storage node, making the usage of the third-party data transparent. In addition, the data provider can also timely discover the abnormal usage of the third-party data and timely adjust the data reading policy. The data provider can also initiate a policy negotiation process with the data user through the data management system, using the data management system as an intermediate platform to negotiate a new data reading policy with the data user to meet the requirements of both the data provider and the data user.
[0029] In a possible implementation manner of the first aspect, the method further includes: during the process of running the business program, the execution node sends the data to be stored to the storage node; after receiving the data to be stored, the storage node determines and stores the derivative link of the data to be stored according to the third-party data read by the execution node from the storage node within the specified time period, where the specified time period includes the time period when the execution node runs the business program this time, and the derivative link is used to indicate the data in the storage node for generating the data to be stored.
[0030] In this possible implementation, during the process of the execution node running the business program, after reading the third-party data from the storage node, derivative data may be generated during the application of the third-party data. To ensure data security, in this possible implementation, the derivative data can be tracked and managed through a derivative link to avoid leakage of the derivative data.
[0031] In a possible implementation of the first aspect, the method further includes: after the execution node stops running the business program, deleting the running data related to the business program in the execution node.
[0032] In this possible implementation, it can effectively ensure that the usage traces of the third-party data used during the process of running the business program are deleted, avoiding the execution node from unreasonably retaining the third-party data used during the process of running the business program, thereby ensuring that the third-party data is not leaked or misused.
[0033] The second aspect of the present application provides a data management system, and the system has the function of implementing the method of the first aspect or any possible implementation of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions, such as a management node, an execution node, and a storage node.
[0034] The third aspect of the present application provides a computing device cluster, which includes at least one computing device. The at least one computing device includes a processor and a memory. The memory of the at least one computing device stores computer execution instructions that can run on the processor. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation of the first aspect as described above.
[0035] The fourth aspect of the present application provides a computer-readable storage medium storing one or more computer execution instructions. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation of the first aspect as described above.
[0036] The fifth aspect of the present application provides a computer program product storing one or more computer execution instructions. When the computer execution instructions are executed by the processor, the processor executes the method of the first aspect or any possible implementation of the first aspect as described above.
[0037] The sixth aspect of this application provides a chip system, which includes a processor for supporting the processor to implement the functions involved in the above-mentioned first aspect or any possible implementation manner of the first aspect. In a possible design, the chip system may further include a memory for storing necessary program instructions and data. The chip system may be composed of chips or may include chips and other discrete devices.
[0038] Among them, for the technical effects brought by the second aspect to the sixth aspect or any possible implementation manner thereof, reference may be made to the technical effects brought by the first aspect or the relevant possible implementation manners of the first aspect, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is an exemplary flowchart provided by an embodiment of this application;
[0040] Figure 2 is an exemplary system framework diagram of the AI platform provided by an embodiment of this application;
[0041] Figure 3a is an exemplary diagram of the development space in the logical multi-tenancy scenario provided by an embodiment of this application;
[0042] Figure 3b is an exemplary diagram of the development space in the physical multi-tenancy scenario provided by an embodiment of this application;
[0043] Figure 4 is an exemplary diagram of an isolated operating environment provided by an embodiment of this application;
[0044] Figure 5 is an exemplary diagram of a management node, an execution node, and a storage node in the training scenario of an AI model provided by an embodiment of this application;
[0045] Figure 6 is an exemplary diagram of a policy item of a data reading policy provided by an embodiment of this application;
[0046] Figure 7 is an exemplary diagram of a data management method provided by an embodiment of this application;
[0047] Figure 8 is an exemplary flowchart provided by an embodiment of this application;
[0048] Figure 9 is an exemplary flowchart provided by an embodiment of this application;
[0049] Figure 10 is an exemplary diagram of a derived link provided by an embodiment of this application;
[0050] Figure 11 is an exemplary schematic diagram of a derivative link provided by an embodiment of the present application;
[0051] Figure 12 is an exemplary schematic diagram of a data management system provided by an embodiment of the present application;
[0052] Figure 13 is a schematic structural diagram of a computing device provided by an embodiment of the present application;
[0053] Figure 14 is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application;
[0054] Figure 15 is a schematic structural diagram of a computing device cluster provided by an embodiment of the present application. Detailed implementation manners
[0055] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation manners part of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.
[0056] As known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0057] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0058] In many application scenarios, the business of purchasing third-party data is involved.
[0059] For example, in the scenario of AI model training, AI developers (i.e., data users) need to purchase a large amount of high-quality data from a third party (i.e., data providers).
[0060] Currently, one way to restrict the usage rights of the data purchased by data users is to clarify the usage rights of the purchased data by data users through a written contract. However, since it is difficult to monitor the data usage process of data users, that is to say, the data usage process of data users is usually not transparent enough to data providers, it may lead to the abuse or even leakage of the data provided by data providers (for the sake of description, hereinafter referred to as third-party data).
[0061] Currently, a commonly used data usage control method is to adopt a user-role-permission authorization model. Data providers assign the minimum permissions to data users to restrict the usage rights of data. However, this control method is relatively simple. For example, it cannot control the number of times data users use third-party data, nor can it record the usage of third-party data by data users, making it impossible for data providers to perceive the usage of third-party data. In addition, data users can still copy and output third-party data and its derivative data during the process of using third-party data through applications, resulting in data leakage.
[0062] Another data usage control method is that data providers attach the usage rights of third-party data to the third-party data through a permission management service and encrypt it. For example, for third-party data of the document type, data providers can set usage rights and encrypt the third-party data. Then, data users can obtain the encrypted third-party data and open it through the application that applies the document. The application can decrypt it to obtain the decrypted document and control the usage of the document by data users according to the usage rights attached to the document.
[0063] However, in this data usage control method, it is necessary to encrypt and decrypt the data, which takes a long time when there is a large amount of data. Moreover, data users can obtain the plaintext data of third-party data through this application, so the third-party data is easily copied and forwarded by data users through this application, resulting in leakage, and making it difficult for data providers to manage the usage and deletion of third-party data and its derivative data in the application.
[0064] It can be seen that traditional data usage control methods are difficult to ensure that the third-party data provided by data providers is used safely and reasonably by data users.
[0065] Based on this, the embodiments of the present application provide a data management method, which can manage the usage process of third-party data provided by data providers to ensure that the third-party data provided by data providers is used safely and reasonably by data users.
[0066] As Figure 1 shown, the data management method of the embodiments of the present application mainly involves one or more of the following aspects:
[0067] Building an isolated operating environment, configuring policies, using policies to read data, managing derivative data, and stopping the running business program.
[0068] The following provides an exemplary introduction to each aspect.
[0069] I. Building an isolated operating environment.
[0070] This data management method can be applied to a data management system, which is located in an isolated operating environment.
[0071] This isolated operating environment can be implemented through a computing device cluster, where the computing device cluster can include one or more computing devices.
[0072] Among them, the type of any computing device is not limited here. Exemplarily, any computing device can be a terminal device, or a server, a server cluster, a container, or a virtual machine, etc. When the computing device cluster includes multiple computing devices, the types of different computing devices can be the same or different.
[0073] The one or more computing devices can be included in one or more data centers, and the data management system can be a system built based on the resources in the data center.
[0074] The specific form of the isolated operating environment where the data management system is located can be various.
[0075] For example, the isolated operating environment can be located in a cloud platform or an AI platform implemented through a computing device cluster, or can be deployed on a server and provide data computing capabilities such as cloud services and applications through a network or an application programming interface (API), such as function cloud services, online document editors, network disks, etc.
[0076] The following takes the AI platform as an example for an exemplary introduction. It should be noted that the isolated operating environment is not limited to the AI platform.
[0077] In the training scenario of an AI model, the business operations that can be performed by the data management system located on the AI platform include AI model training.
[0078] As Figure 2 shown, it is a schematic diagram of an exemplary system framework of the AI platform.
[0079] As Figure 2 shown in the example, the AI platform can provide access interfaces to users such as data consumers and data providers.
[0080] For example, by providing an interface through a console, users can access the console of the AI platform through a browser or the like, thereby accessing the AI platform; alternatively, an application in the user's client device can access the APIs provided by the AI platform through an API gateway, thereby accessing the AI platform.
[0081] The AI platform can also provide one or more of data management services, annotation services, training services, inference services, and storage services, and can provide access interfaces to users such as data consumers and data providers.
[0082] The annotation service, training service, and inference service can provide their respective services through corresponding application programs based on corresponding algorithms, etc., and can access the storage service through the storage software development kit (SDK) provided by the storage service.
[0083] The storage service can store data. For example, it can store the proprietary data of data consumers, third-party data of data providers, and derivative data generated based on other data (such as proprietary data and third-party data, etc.).
[0084] The storage service can also store metadata. The metadata can include information on proprietary data, third-party data, and derivative data, and can also include relevant policies in subsequent steps for controlling the use of data such as third-party data. In addition, the storage service can also store access records of data, etc., to facilitate data providers, etc., to understand the usage of third-party data.
[0085] In this way, data consumers can access the AI platform to utilize the resources of the AI platform for AI model training in the AI platform.
[0086] Moreover, data providers can store and sell the third-party data of the data providers through the AI platform. In this way, data consumers can purchase the third-party data provided by data providers through the AI platform for AI model training.
[0087] Among them, in the embodiments of the present application, the data user can be understood as an individual, an enterprise or an organization, and can also be understood as a subject represented by a device such as a client, an account, a network address or other forms of identifiers. The identifier and specific form of the data user can be various; similarly, the data provider can be understood as an individual, an enterprise or an organization, and can also be understood as a subject represented by a device such as a client, an account, a network address or other forms of identifiers. The identifier and specific form of the data provider can be various, and the embodiments of the application do not limit this.
[0088] In this example, an isolated operating environment can be built in the AI platform, and a data management system can be implemented in this isolated operating environment.
[0089] Among them, as Figure 3a shown, when the development space provided by the AI platform to the tenant (which can be a data user) is logically multi-tenant, the computing device cluster of the AI platform can process requests from multiple tenants. That is to say, multiple tenants (such as Figure 3a the development space 1 of tenant 1 and the development space 2 of tenant 2 shown) share the resources of the AI platform. Then, the isolated operating environment can be the AI platform. Among them, the isolated operating environment can include the network resources, computing resources and storage resources of the AI platform.
[0090] However, as Figure 3b shown, when the development space provided by the AI platform to the tenant (which can be a data user) is physically multi-tenant, the resources of the development space leased by the tenant from the AI platform are exclusive to this tenant, and the resources among the tenants in the AI platform are fully isolated. That is to say, in the example as Figure 3b shown, the network resources, computing resources and storage resources in the development space 1 of tenant 1 are the isolated operating environment 1, and the network resources, computing resources and storage resources in the development space 2 of tenant 2 are the isolated operating environment 2. The isolated operating environment 1 and the isolated operating environment 2 are isolated from each other.
[0091] It can be seen that in the embodiments of the present application, the isolated operating environment can be implemented from one or more aspects of network, computing and storage.
[0092] The following will be introduced separately.
[0093] 1. Network isolation
[0094] As Figure 4 shown in the example, the isolated operating environment is connected to other nodes outside the isolated operating environment through the API. For example, the isolated operating environment is connected to the data provider and the data user respectively through the API.
[0095] In this way, internal nodes (such as execution nodes and storage nodes in a data management system) in the isolated operating environment cannot communicate directly with external nodes (such as client devices of data consumers) outside the isolated operating environment.
[0096] In this way, the information transmission between internal nodes and external nodes can be monitored through a monitoring API.
[0097] Specifically, in some embodiments, the access behavior of data consumers to the data management system through the API can be monitored, and the data output behavior of the data management system through the API can be monitored.
[0098] Exemplarily, the API can be managed through an API gateway or other forms of API management tools. The data management system can register the API with the API gateway to monitor the data interaction related to the API through the API gateway.
[0099] The API gateway can access internal nodes in the isolated operating environment internally. The API gateway can be deployed in the isolated operating environment. For example, the API gateway can be an elastic load balance (ELB) service in a cloud platform, or the API gateway can also be located outside the isolated operating environment and communicate with the isolated operating environment through technologies such as virtual private cloud (VPC) peer connection, source network address translation (SNAT), and elastic Internet Protocol (EIP).
[0100] In addition, the API gateway can be accessed by external nodes externally. For example, it can communicate with the external network where the external nodes are located through the ELB service in the cloud platform.
[0101] Moreover, the network flow of the isolated operating environment can also be controlled through the API gateway, etc. For example, the network flow can be unidirectional, from the external node to the API gateway, and then from the API gateway to the internal node. That is to say, only the external node can send a request to the internal node through the API, and after the internal node processes the request, it can send a response to the external node through the API. The internal node cannot actively send data to the external node through the API without receiving a request from the external node.
[0102] In this way, through the API gateway, etc., the access behavior of external nodes (data consumers) to internal nodes (data management system) through the API can be monitored, and the data output behavior of the data management system through the API can be monitored.
[0103] Among them, access behavior can be monitored by one or more of monitoring access frequency, corresponding response size, response content, etc., so that abnormal access behavior can be identified. The monitoring data output behavior can be to detect the similarity between the data to be output by the data output behavior and the third-party data provided by the data provider, so as to avoid the leakage of third-party data.
[0104] The specific form of this isolated operating environment is not limited here.
[0105] Exemplarily, in a cloud platform, the isolated operating environment can be to create an independent network environment using a virtual private cloud (VPC). In a non-cloud platform, the isolated operating environment can be a small network environment or an independent network area isolated by a firewall.
[0106] 2. Computing Isolation
[0107] In some examples, the data management system in the isolated operating environment includes a management node, an execution node, and a storage node. The management node can be used to execute the system operations of the data management system. The execution node is used to run business programs, and the business programs are used to execute the business of data users. The storage node can be used to provide storage services. For example, it stores the third-party data provided by the data provider.
[0108] Among them, the specific forms of nodes such as the management node, the execution node, and the storage node, the forms of the above-mentioned various nodes can be the same or different, and are not limited here.
[0109] Exemplarily, any node can be a software module or plugin in the data management system; or, it can be a virtual machine or a container; or, it can also be the hardware in the computing device cluster that implements the data management system. For example, the execution node is a server or a server cluster in the computing device cluster, and the storage node is a storage device in the computing device cluster.
[0110] In the embodiments of the present application, the business programs that execute the business of data users (such as performing AI model training) are executed by the execution nodes, so that the data management system can isolate and manage the operation and data interaction of the business programs, and avoid the leakage of third-party data and related derivative data, etc. during the execution of the business programs.
[0111] Moreover, the management node can also monitor the business programs running on the execution nodes. For example, it can monitor the business programs with high usage frequency to timely discover the anomalies of the business programs.
[0112] In addition, the data management system can also limit users' viewing rights to the execution results, execution logs and other execution data of business programs. For example, the scale and frequency of users' viewing of execution data can be limited. For example, users can view up to 1,000 execution results within a specified time period, or can view up to 10MB of execution logs.
[0113] 3. Storage Isolation
[0114] The storage nodes of the data management system in the data management system can provide storage services.
[0115] Among them, data users can transfer, download, and delete their own data to storage nodes.
[0116] The data provider may transmit the third-party data provided by the data provider to the storage node, and the storage node may store the third-party data.
[0117] In this way, data users can apply to storage nodes for permission to use third-party data, for example, they can obtain permission to use third-party data by purchasing. In this way, the use and management of third-party data of data providers can be achieved through storage nodes.
[0118] In actual scenarios, third-party data and related derivative data stored in storage nodes cannot be downloaded directly. If data users or other users need to download third-party data from storage nodes, they can initiate a download request for the third-party data to the data provider through the data management system. Only when the data provider of the third-party data approves the download request through the data management system will the data user or other users be allowed to download the corresponding third-party data from the storage node.
[0119] like Figure 5 As shown, taking the data management system as an AI platform as an example, the relevant functions of the management node, execution node, and storage node in the training scenario of the AI model are introduced exemplarily.
[0120] exist Figure 5 In the example shown, the management node of the AI platform can provide platform-level AI services.
[0121] AI services can trigger execution nodes to perform business services based on instructions from data users.
[0122] Specifically, when executing a business program, the AI service first initializes the execution node, and then assigns the business program to the execution node for execution. The business program can use the infrastructure of the AI platform (such as a graphics processing unit (GPU) cluster, etc.) to train and reason about the AI model.
[0123] During the execution of the business process, the execution node can write the temporary data generated during the execution process to the local storage. After the business process is executed, the local storage is cleared and released. The execution node can write the logs generated during the execution process to the log service, or write them to the local storage and have the log agent collect them to the log service. The execution node can also write the persistent data such as AI models to the storage service. Users need to view the logs and persistent data through the AI service and configure policies in the AI service to control the use of the data, etc.
[0124] II. Configure policies.
[0125] In the embodiments of this application, the data management system may include one or more of the following policies:
[0126] Data permission policy, data reading policy.
[0127] Through the above policies, the data management system can manage the read operations of the data in the storage node.
[0128] The above policies are introduced exemplarily below.
[0129] 1. Data permission policy
[0130] In this example, the data permission policy is used to determine the read and write permissions for the data according to the source of the data.
[0131] In some examples, the data permission policy can be a system-level policy, that is to say, the data permission policy can be applied to the data management system and be pre-configured by the developer of the data management system.
[0132] Among them, the source of the data can include one or more of the following: self-owned data, third-party data, and derivative data. Among them, self-owned data is the data provided by the data user, third-party data is the data provided by the data provider, and derivative data is the data obtained based on data such as third-party data. In some examples, the storage node can consider the data to be stored in the storage node when the execution node executes the business process (also called the data to be stored) as derivative data.
[0133] In some examples, the data permission policy can also be used to determine the read and write permissions for the data according to the content of the data.
[0134] Among them, the content of the data can have various classification methods. For example, the content of the data can include ordinary data and model data, and all data other than model data can be considered ordinary data.
[0135] In addition, the data permission policy can also be used to determine the read and write permissions for data according to the use of the data.
[0136] The following uses a specific example to illustrate an exemplary content of this data permission policy.
[0137] Table 1: Data Permission Policy
[0138]
[0139] It can be seen that in the example shown in Table 1, the data permission policy can indicate the read permission for third-party data from a data provider, and thus can determine the third-party data that a business program can read from a storage node according to the data reading policy subsequently.
[0140] 2. Data Reading Policy
[0141] In the embodiments of the present application, the data reading policy can be a reading policy for the corresponding third-party data. The specific type of this third-party data is not limited herein. For example, the third-party data can be a file or a group of files (such as a directory, a compressed data packet, etc.). The third-party data can include attributes such as type, name, path, number of lines, and storage size.
[0142] In this example, the data reading policy corresponding to a certain third-party data can be determined by the data provider of this third-party data. Specifically, it can be that the data management system provides a configuration interface, and the data provider of the third-party data inputs configuration information through this configuration interface to configure the data reading policy; or, it can also be that the data management system provides candidate data reading policies (which can be obtained from historical data reading policies or default data reading policies) to the data provider of the third-party data, and the data provider determines the candidate data reading policy as the data reading policy for this third-party data.
[0143] The data reading policy includes policy items and operations determined based on the policy items. The number and content of the policy items can have various situations. Among them, the object of any policy item can be the data itself or a file in the data, and the specific type of the object of each policy item can be pre-configured. For example, it is defined through PolicyType. Specifically, PolicyType being DATA indicates that the object of the policy item is the data itself, and PolicyType being FILE indicates that the object of the policy item is a file. The operation determined based on the policy item is the action indicated after the policy item is matched. For example, it can include allowing reading, denying reading, etc.
[0144] As Figure 6 shown, in some examples, the policy item includes one or more of the following items:
[0145] Data policy item, usage policy item, quantity policy item, time policy item, subject policy item.
[0146] Next, various exemplary policy items will be introduced exemplarily.
[0147] 1) Data Policy Item
[0148] The data policy item is used to describe the data range in the third-party data that can be read. This data policy item can split the data and determine the data that can be read from the split data.
[0149] Exemplarily, this data policy item may include path filtering, selected columns, and row filtering.
[0150] Path filtering means screening out some files from multiple files included in the third-party data, and the default setting is to select all files. Specifically, filtering can be performed using file extensions, wildcards, regular expressions, etc. Among them, wildcards can use relatively simple and easy-to-use path matching expressions such as the Ant style.
[0151] Selected columns mean selecting one or more columns, and the default setting is to select all columns. For third-party data of data types such as comma-separated values (CSV), JavaScript object notation (JSON), and Excel, these third-party data can define columns when being transmitted to the data management system. If a certain third-party data does not define columns, the content of selected columns is not applicable. For example, third-party data of binary file types such as pictures and videos does not apply to the content of selected columns.
[0152] Row filtering means filtering the corresponding third-party data by rows, and only the rows that meet the conditions can be read, such as "name in ('beijing', 'nanjing') AND age > 30", where name and age are columns of the data. For third-party data of data types such as CSV, JSON, and Excel, these third-party data can define columns when being transmitted to the data management system. If a certain third-party data does not define rows, the content of row filtering is not applicable. For third-party data of binary file types such as pictures and videos, the corresponding metadata can be filtered, such as "imageSize < 10MB AND imageWidth < 1024", where imageSize describes the size of the picture and imageWidth describes the width of the picture.
[0153] Exemplarily, the data policy item can be configured through the following code:
[0154]
[0155] 2) Usage Policy Item
[0156] The usage policy item is used to describe the tasks in the business process that can use third - party data, that is, it is used to describe which functions can use the third - party data.
[0157] For example, in the application scenario of an AI model, the functions involved may include one or more of the following:
[0158] a) Labeling: Annotation;
[0159] b) Feature: Feature engineering;
[0160] c) Train: Training;
[0161] d) Reasoning: Inference.
[0162] Exemplarily, the usage policy item can be configured through the following code:
[0163]
[0164] 3) Count Policy Item
[0165] The count policy item is used to describe the usage quantity threshold of third - party data, that is to say, it is used to describe the quantity of third - party data that can be used. There are various situations for this usage quantity. When the object of the data policy item is data or a file, the specific content of the usage quantity can be different.
[0166] In one example, PolicyType is DATA, indicating that the object of the policy item is the data itself. At this time, the usage quantity of the data policy item can be the number of uses. Specifically, this data policy item is used to describe the usage times threshold (maxUseTimes) of third - party data.
[0167] This usage times threshold can be the maximum number of times the third - party data can be used. Among them, the calculation rule of this usage times can be defined by the user or the data management platform, and the calculation rule includes but is not limited to:
[0168] a) Count only once: When the algorithm program reads data one or more times during a single run, useTimes is incremented only once. Multiple runs will be cumulatively counted.
[0169] b) Count only once for multiple uses within a period: When the algorithm program reads data one or more times within a period, useTimes is incremented only once.
[0170] c) Count once for each use: Each time the algorithm program uses data, useTimes is incremented once.
[0171] In another example, PolicyType is FILE, indicating that the object of the policy item is a file. In this case, the usage threshold can include one or more of the following:
[0172] a) The maximum number of rows (maxRowCount) refers to the maximum number of rows that can be used for a single file. For unstructured third-party data such as pictures and videos, this usage threshold is not applicable.
[0173] b) The maximum storage size (maxStorageSize) refers to the maximum amount of data that can be used for a single file.
[0174] Exemplarily, the Count Policy Item can be configured through the following code:
[0175]
[0176] 4) Time Policy Item
[0177] The Time Policy Item is used to describe the usage time range of third-party data.
[0178] For example, the Time Policy Item can specify a specific deadline. For example, the RFC 3339 date format can be adopted, and the time zone can be included. For example, the media time can be set to 2023-12-01T12:21:32Z.
[0179] Alternatively, the Time Policy Item can specify the usage duration. For example, the starting time can be the time when the third-party data is transferred to the data management system. The usage duration can be expressed in various ways. For example, it can be expressed using Duration in RFC 3339. For example, the usage duration is expressed as P2DT3H4M, indicating a usage duration of "2 days, 3 hours, and 4 minutes".
[0180] Exemplarily, the Time Policy Item can be configured through the following code:
[0181]
[0182] 5) Subject Policy Item
[0183] The Subject Policy Item is used to describe the types of subjects to which the data reading policy can be applied, so as to clarify the types of subject objects to which the corresponding data reading policy applies.
[0184] Considering that the data management system may be used by multiple users, different data reading policies can be specified for different users through the Subject Policy Item.
[0185] There can be multiple specific situations for this subject type, which can be determined according to the specific application scenario.
[0186] Exemplarily, the subject type may include one or more of the following:
[0187] a) User: The accounts involved in the data management system, such as the users in the cloud management platform or AI platform where the data management system is located, with independent identity identifiers.
[0188] b) Group: A group is a collection of users. The data reading policy applicable to the group can be determined through the Subject Policy Item, thus simplifying the authorization operation of relevant reading permissions.
[0189] c) Role: A role is a collection of a set of permissions, defined according to the user's job functions.
[0190] d) OrganizationalUnit: An organizational unit can be mapped to departments, subsidiaries, project families, etc. of an enterprise or organization. A user can belong to an organizational unit.
[0191] Exemplarily, the Subject Policy Item can be configured through the following code:
[0192]
[0193] Based on any of the above examples, one or more data reading policies can be configured for third-party data.
[0194] Among them, there can be multiple ways to express the data reading policy.
[0195] For example, when applied to data management systems and business programs therein, the data reading policy can be expressed in ways such as open digital rights language (ODRL), extensible markup language (XML), JSON, YAML, etc.
[0196] When facing users (such as data providers), it can be expressed in text for easy understanding by users.
[0197] For example, when expressing the data reading policy in text, the format of "
Allow / Deny
Subject
Quantity
Purpose
Time
[0198] It can be seen that the data reading policy can be flexibly configured from one or more aspects through one or more policy items, with a clear configuration method and easy implementation, which can improve the configuration efficiency, facilitate the understanding of data users and data providers, and thus facilitate the negotiation and unification between data users and data providers.
[0199] Based on the above policy, the metadata information corresponding to each data can be recorded in the storage node.
[0200] Exemplarily, the specific content of the metadata information of each data can be as shown in Table 2.
[0201] Table 2: Metadata Information of Data Stored in the Storage Node
[0202]
[0203] In this way, the use of data can be managed through the metadata information of the data in the storage node. The specific management method can refer to related embodiments such as the processing stage of data reading using the usage policy, which will not be elaborated here.
[0204] III. Reading Data Using the Usage Policy.
[0205] In the embodiments of the present application, in actual application scenarios, the reading operations of data users on third-party data in the storage node can be managed according to the data reading policy and / or data permission policy.
[0206] Specifically, as Figure 7 shown, in some embodiments, the method includes steps 701 - 703.
[0207] Step 701, the execution node runs the business program.
[0208] The business program is used to execute the business of the data user.
[0209] In the embodiments of the present application, the execution node is used to run the business program to execute the business of the data user.
[0210] The specific functions of this business program are not limited herein.
[0211] For example, in the scenario of AI model training, this business program is used to train the AI model. For example, it can call resources such as the GPU cluster of the AI platform, as well as the training data stored in the storage node to train the AI model.
[0212] In some examples, the execution node can be created according to the instructions of the management node in the data management system and start running the business program.
[0213] For example, specifically, the management node can receive instructions from the data user to indicate the execution of a specified business.
[0214] After the management node receives this instruction, it can initialize the execution node and instruct the execution node to run the business program to execute the business of the data user.
[0215] After the execution node completes the current execution of this business program, it can delete the relevant running data of this execution stored in the execution node, such as temporary data, etc., but the execution node can be retained to execute other businesses of this data user, or to run this business program again later; or, after the execution node completes the current execution of this business program, it can destroy this execution node.
[0216] Step 702, the execution node sends a data reading request to the storage node.
[0217] The data reading request indicates that the running business program requests to read the first data from the storage node, and the first data belongs to third-party data.
[0218] In the embodiments of the present application, during the process of running the business program, when the execution node needs to use the first data belonging to third-party data, it can generate and send a data reading request to the storage node.
[0219] It can be seen that the use of third-party data by the execution node is controlled by the storage node. That is to say, the execution node cannot obtain third-party data without being monitored. Compared with the traditional data control method that controls the use of third-party data through the business program using third-party data, the solution of the embodiments of the present application can better monitor the business program using third-party data, so as to better ensure that third-party data is not misused and leaked.
[0220] Step 703: The storage node generates feedback information according to the data reading request and the data reading policy for third-party data, and sends the feedback information to the execution node.
[0221] In this storage node, one or more data reading policies may be stored. In addition, in some examples, data permission policies may also be stored.
[0222] In the embodiments of the present application, the data reading policy for third-party data can be matched according to the data reading request, and the data permission policy for third-party data can also be matched. Depending on different matching results, the specific content of the feedback information may be different.
[0223] For example, in the case where it is determined that data reading is allowed according to the matching result, the feedback information may include the data allowed to be read; while in the case where it is determined that data reading is not allowed according to the matching result, the feedback information may indicate the data reading operation that rejects the data reading request.
[0224] The following provides an exemplary introduction to the matching process related to the data permission policy and / or the data reading policy.
[0225] 1. Perform data permission policy matching.
[0226] Specifically, in some embodiments, the storage node further includes a data permission policy for determining the read and write permissions for data according to the data source. Among them, the data permission policy indicates that the data user has read permission for third-party data from the data provider;
[0227] Before the storage node generates feedback information according to the first data and the data reading policy for third-party data and sends the feedback information to the execution node, it further includes:
[0228] The storage node determines, based on the fact that the first data belongs to third-party data according to the data permission policy, that the business program has read permission for the first data.
[0229] In the embodiments of the present application, the relevant content of this data permission policy can refer to the relevant content of the above configuration policy section, which will not be elaborated here.
[0230] This data permission policy can be a system-level policy of the data management system, so that the read and write permissions of the data in the storage node can be basically managed through this system-level policy.
[0231] In the embodiments of the present application, the data permission policy may include indication information indicating that the data user has read permission for third-party data from the data provider. In this way, after determining that the first data has read permission, the data reading can be performed according to the data reading policy.
[0232] It can be seen that in the embodiments of the present application, the reading operation of data in the storage node can be more perfectly managed through at least two levels of policies.
[0233] 2. Perform data reading policy matching.
[0234] Specifically, in some embodiments, the storage node includes one or more data reading policies regarding third-party data;
[0235] The storage node generates feedback information according to the data reading request and the data reading policy regarding third-party data, and sends the feedback information to the execution node, including:
[0236] The storage node matches the data reading request with the policy items of at least one data reading policy among one or more data reading policies;
[0237] If there is a data reading policy that matches the data reading request, feedback information is generated according to the operations in the data reading policy that matches the data reading request, and the feedback information is sent to the execution node.
[0238] In the embodiments of the present application, the data reading policy includes policy items and operations determined based on the policy items.
[0239] The number of data reading policies to be matched can be one or more. When the number of data reading policies to be matched is multiple, the multiple data reading policies can be matched in sequence. Specifically, the matching can be performed according to the policy items in the data reading policy. The multiple data reading policies are matched in sequence until there is a matching data reading policy, and feedback information is generated according to the operations in this data reading policy, or until it is determined that all data reading policies do not match, then feedback information can be generated according to the default operation, or feedback information indicating the rejection of the data reading can be generated.
[0240] In addition, in some examples, in order to better manage the data reading operation of the execution node, the execution node itself can be made unable to obtain the permission to interact with the storage node by itself, but the data management system manages the permission for the execution node to interact with the storage node. After the data management system authorizes the execution node to obtain the permission to interact with the storage node, the execution node can send a data reading request to the storage node.
[0241] Specifically, in some embodiments, the data management system includes a management node;
[0242] After the management node obtains the operation instruction for the service program, it creates an access credential for the storage node;
[0243] The management node instructs the execution node to run the service program and transfers the information of the access credential to the execution node;
[0244] The execution node sends a data reading request to the storage node, including:
[0245] Based on the information of the access credential, the execution node sends a data reading request to the storage node.
[0246] The management node can be regarded as a system-level node of the data management system. This management node can manage the permissions for the execution node to interact with the storage node. The data management system, through the management node, conducts system-level management of the access credentials of the storage node, rather than having the execution node apply for the access credentials of the storage node by itself, thus preventing data users from accessing third-party data improperly through the execution node.
[0247] It can be seen that in the embodiments of the present application, the management node can manage the access credentials of the execution node to the storage node, that is to say, this management node can manage the permissions for the execution node to interact with the storage node. In this way, the execution node cannot obtain the access permission to the storage node by itself, avoiding the execution node from reading and using the third-party data of the storage node without permission, thus avoiding the leakage and unreasonable use of the third-party data by the execution node.
[0248] Among them, the specific form of the access credential can have various situations and is not limited here.
[0249] Exemplarily, the access credential can be a storage session. Specifically, the access credential can be uniquely represented by the identifier of the storage session. Or the access credential can also be a specified secret key or other forms of credentials.
[0250] Such as Figure 8 shown, it is an exemplary process schematic diagram for information interaction among the management node, the execution node, and the storage node.
[0251] In Figure 8 the example, it can include the following steps:
[0252] 1) The data user sends a start instruction for starting a task to the management node to instruct the execution of a specified service through the business program.
[0253] 2) After receiving the start instruction, the management node can send indication information to the storage node to instruct the storage node to create a storage session.
[0254] The indication information can include the ID of the isolated operating environment where the management node is located (such as the ID of the development space where it is located), the subject corresponding to the storage session, and the purpose corresponding to the storage session, to request that the subject can obtain the permission to read data for this purpose from the storage node through this storage session.
[0255] 3) After receiving the indication information, the storage node can return the ID of the storage session to the management node. In this way, the subject of the corresponding subject type can obtain the permission to read the data for this purpose from the storage node through the ID of the storage session.
[0256] 4) After receiving the ID of the storage session, the management node instructs the execution node to run the business program and passes the ID of the storage session to the business program. For example, it can be passed through parameters, environment variables, etc.
[0257] 5) When running the business program, the execution node requests the storage node to list the files included in the data of the storage node according to the ID of the storage session.
[0258] 6) Without verifying the policy, the storage node returns the file list.
[0259] 7) When running the business program, the execution node sends a data reading request to the storage node according to the ID of the storage session. The data reading request instructs the running business program to request to read the first data from the storage node, and the first data belongs to third-party data.
[0260] 8) The storage node reads the metadata of the first data according to the data reading request.
[0261] 9) The storage node verifies the policy according to the metadata of the first data, the subject corresponding to the storage session, the purpose corresponding to the storage session, etc. For example, it verifies the data permission policy first and then the data reading policy.
[0262] 10) According to the matching result, feedback information is returned.
[0263] Among them, the matching result may be as follows:
[0264] A. If a certain data reading policy does not match, the next data reading policy is judged.
[0265] B. If a certain data reading policy matches, then:
[0266] (a) If the operation in the data reading policy is ALLOW, it is determined that the data reading policy is the matching policy, and the file content allowed to be read is fed back according to the data reading policy, and the subsequent policy matching operation is stopped.
[0267] (b) If the operation in the data reading policy is DENY, access to the first data is denied, and the subsequent policy matching operation is stopped.
[0268] C. If all do not match, the storage node can adopt the default operation, and the default operation can be specified by the data management system.
[0269] Specifically, to determine whether a certain data reading policy matches, it can be achieved from the policy items of the data reading policy.
[0270] For example, a certain data reading policy includes a data policy item, a usage policy item, a quantity policy item, a time policy item, and a subject policy item.
[0271] Among them, the exemplary matching methods for each policy item are as follows:
[0272] A. Data policy item:
[0273] (a) Path filtering: The file path specified in the data reading request needs to match the specified conditions;
[0274] (b) Select columns: It does not participate in the matching, but is applied when reading the content;
[0275] (c) Row filtering: It does not participate in the matching, but is applied when reading the content.
[0276] B. Usage policy item:
[0277] The usage specified in the data reading request needs to be in the usage list defined in the usage policy item.
[0278] C. Quantity policy item:
[0279] The control object of the data reading policy can be the data itself (DATA) or the file in the data (FILE), which is defined by PolicyType. The PolicyType in the metadata of the first data can be DATA or FILE.
[0280] (a) When PolicyType in the metadata is DATA:
[0281] Usage times: Calculate the usage times of the first data according to the calculation rule specified by the user in the quantity policy item, and determine whether it exceeds the range. Among them, reading a certain data multiple times within one storage session can only be counted as one usage.
[0282] (b) When PolicyType in the metadata is FILE:
[0283] ① Number of rows: It does not participate in the matching, but is applied when reading the content;
[0284] ② Storage size: It does not participate in the matching, but is applied when reading the content.
[0285] D. Time policy item:
[0286] The current time needs to be within the usage time range.
[0287] E. Subject policy item:
[0288] The subject corresponding to the data reading request can match the definition of the subject policy item. For example, the userId of the subject is within the userId defined in the subject policy item.
[0289] If the storage node determines to allow data reading based on the matching result, it returns the data allowed to be read.
[0290] Among them, when the storage node returns the data allowed to be read, it needs to control the content read based on the finally matched data reading policy, mainly one or more of the following control methods:
[0291] Select columns: Exclude unselected columns:
[0292] Row filtering: Filter the content:
[0293] Number of rows: Limit the maximum number of rows that can be read;
[0294] Storage size: Limit the maximum size that can be read.
[0295] In this way, the storage node can complete the feedback on the data reading request.
[0296] In some embodiments, to ensure the validity of the access credential, after the execution node stops running the business program, the access credential is deleted.
[0297] In this way, each time the execution node starts running the business program, it needs to obtain the permission to interact with the storage node again through the management node. In this way, the execution node cannot obtain the access permission to the storage node by itself, nor can it abuse the obtained access permission, effectively ensuring the controllability of each access of the execution node to the storage node, avoiding the execution node from reading and using the third-party data of the storage node without permission, and thus avoiding the leakage and unreasonable use of the third-party data by the execution node.
[0298] In addition, in some embodiments, the storage node can record the data access behavior of the third-party data for the data provider to view.
[0299] Specifically, in some embodiments, after the above step 703, the method further includes:
[0300] The storage node records the data access behavior related to the data reading request.
[0301] The data access behavior includes the access behavior that successfully realizes data reading, and can also include the access behavior of data reading failure.
[0302] In the embodiments of the present application, the data access behavior may include one or more of a data read request, a user associated with the data read request, and a response corresponding to the data read request (such as feedback information, etc.).
[0303] Exemplarily, the data access behavior may include one or more of the following information:
[0304] Access time, user (including data user and / or data provider), purpose, operation (such as read operation or write operation), data, file, number of lines, storage size.
[0305] For example, the access record may be as shown in Table 3.
[0306] Table 3: Access Record
[0307] Field Name Data Type Remarks ID of Development Space String Stored Session ID String Time DateTime Read / Write String Read (R) / Write (W) Data ID String
[0308] After the storage node stores the information of the above data access behavior, the data provider can conveniently and flexibly view the usage situation of the third-party data stored in the storage node by it.
[0309] Specifically, in some embodiments, the method further includes:
[0310] The storage node obtains a query request input by the data provider, and the query request is used to request to query the access record related to the third-party data;
[0311] The storage node outputs the access record, and the access record includes the data access behavior related to the data read request.
[0312] In the embodiments of the present application, the data provider may input a query request to the data management system through a client, etc., based on the API, to request to query the access record related to the third-party data provided by the data provider. After receiving the query request, the storage node may, according to the identifier of the data provider and / or the identifier of the third-party data, query the data access behavior information related to the third-party data from the stored data access behavior information, so as to obtain and output the access record of the third-party data to the data provider and / or other devices. At this time, the access record may include the data access behavior related to the data read request.
[0313] In this way, the data provider can conveniently query the situation of the third-party data being accessed and used by the data user through the storage node, making the usage of the third-party data transparent.
[0314] In addition, the data provider can also promptly detect the abnormal use of third-party data and adjust the data reading strategy in a timely manner. The data provider can also initiate a policy negotiation process with the data user through the data management system, using the data management system as an intermediate platform to negotiate a new data reading strategy with the data user to meet the requirements of both the data provider and the data user.
[0315] IV. Derivative data management.
[0316] In the embodiments of the present application, during the process of the execution node running the business program, after reading the third-party data from the storage node, derivative data may be generated during the application of the third-party data.
[0317] To ensure data security, in the embodiments of the present application, the derivative data can be tracked and managed to avoid the leakage of derivative data.
[0318] Specifically, as Figure 7 shown, in some embodiments, the method further includes:
[0319] Step 704, during the process of running the business program, the execution node sends the data to be stored to the storage node.
[0320] Step 705, after receiving the data to be stored, the storage node determines and stores the derivative link of the data to be stored according to the third-party data read by the execution node from the storage node within the specified time period.
[0321] The specified time period includes the time period when the execution node runs the business program this time, and the derivative link is used to indicate the data in the storage node that is used to generate the data to be stored.
[0322] In the embodiments of the present application, there can be various situations for the data to be stored.
[0323] For example, in one example, during the process of running the business program, the execution node can detect the derivative data generated according to the third-party data read from the storage node and use the derivative data as the data to be stored.
[0324] In another example, during the process of the execution node executing the business program, temporary data and persistent data that need to be persistently stored may be generated. Among them, the temporary data can be stored in the execution node, for example, through the cache, memory, etc. of the execution node. And the persistent data needs to be stored through the storage node. Therefore, the execution node can use the persistent data as the data to be stored and send it to the storage node. In this example, it does not actually detect whether the data to be stored is truly generated according to the third-party data.
[0325] After receiving the data to be stored, the storage node may consider that the data to be stored is likely to be generated based on the third-party data read by the execution node from the storage node within a specified time period. Therefore, according to the third-party data read by the execution node from the storage node within the specified time period, determine and store the derivation link of the data to be stored.
[0326] Among them, the specified time period may include the time period between the time point when the execution node starts running the business program this time and the time point when the storage node receives the data to be stored.
[0327] For example, as Figure 9 shown in the process schematic diagram, after the management node instructs the execution node to run the business program and passes the ID of the storage session to the execution node as an access credential, the following steps may be included:
[0328] The execution node reads data A from the storage node, and data A is third-party data;
[0329] The storage node records the access behavior related to data A. For example, it records the ID of the storage session, the information of data A, the time, etc.;
[0330] The execution node reads data B from the storage node, and data B is third-party data;
[0331] The storage node records the access behavior related to data B. For example, it records the ID of the storage session, the information of data B, the time, etc.;
[0332] The execution node writes data C to the storage node, and this data C can be considered as the data to be stored;
[0333] The storage node can, according to the recorded access behavior, obtain the third-party data read by the execution node from the storage node within the specified time period from the time of obtaining the ID of the storage session created this time to before this write operation;
[0334] According to the information of this third-party data, update the information in the "derived from" field in the metadata of data C in the storage node to indicate that data C is derived from data A and data B;
[0335] The storage node determines and stores the derivation link of data C;
[0336] The storage node writes the content of data C and returns information to the execution node to indicate that data C has been written;
[0337] The execution node reads data D from the storage node, and data D is third-party data. The relevant read operations will not be elaborated;
[0338] The execution node writes data E to the storage node. The relevant read operations will not be elaborated.
[0339] After performing the above steps, a derived link as shown in Figure 10 can be generated and stored in the storage node, where data C is derived from data A and data B, and data E is derived from data A, data B, data C, and data D.
[0340] Exemplarily, the derived link can be recorded through the fields shown in Table 4.
[0341] Table 4: Fields of the Derived Link
[0342] Field Name Data Type Remarks Space ID String Source Data ID String Target Data ID String
[0343] Among them, in the derived link of data C, data C can be regarded as the target data, and data A and data B are the source data of data C.
[0344] In addition, after the business program in the execution node is executed multiple times within a development space, the relevant derived links can also be automatically merged and updated, as shown in Figure 11 which can merge the derived link related to storage session 1 and the derived link related to storage session 2.
[0345] When the data provider or data user views the derived link of data A, according to the record of the derived link, the subsequent derived link of data A and the related derived data can be found.
[0346] In addition, the data provider can also request to delete the third-party data provided by the data provider, and can also request to delete the derived data of the third-party data.
[0347] Specifically, the data provider can initiate a deletion request to the storage node. At the same time, the storage node can give a certain time as a buffer. For example, the data user can receive a deletion request notification, and then can view all the derived links corresponding to the third-party data to be deleted in time to adjust the business program. After the buffer time expires, the data management system can perform the deletion operation of the third-party data and the related derived data according to the above-derived link and notify the data user. If the data user cannot complete it within the buffer time, an extension can be requested, and after the data provider agrees, it will be subject to the new time.
[0348] V. Stop running the business program.
[0349] Specifically, in some embodiments, the method further includes:
[0350] After the execution node stops running the business program, delete the running data related to the business program in the execution node.
[0351] After the execution node completes the current execution of the service program, relevant running data of the current execution stored in the execution node, such as temporary data, etc., can be deleted, but the execution node can be retained to execute other services of the data user; or execute the service program again later; or, after the execution node completes the current execution of the service program, the execution node can be destroyed.
[0352] In this way, it can effectively ensure that the usage traces of third-party data used in the process of running the service program are deleted, and avoid the execution node unreasonably retaining the third-party data used in the process of running the service program, thereby ensuring that the third-party data is not leaked and misused.
[0353] The above introduced the data management method provided by the embodiments of the present application from multiple aspects. Next, in combination with the accompanying drawings, the data management system provided by the embodiments of the present application will be introduced.
[0354] As Figure 12 shown, the embodiments of the present application provide a data management system 120. The data management system 120 is set in at least one data center. The data management system 120 is located in an isolated operating environment. The isolated operating environment is connected to the data provider and the data user respectively through an application programming interface API. The data management system 120 includes an execution node 1201 and a storage node 1202. The storage node 1202 stores third-party data of the data provider.
[0355] The execution node 1201 is used for:
[0356] Running a service program, and the service program is used to execute the service of the data user;
[0357] Sending a data reading request to the storage node, and the data reading request instructs the running service program to request to read the first data from the storage node, and the first data belongs to the third-party data;
[0358] The storage node 1202 is used for: generating feedback information according to the data reading request and the data reading policy regarding the third-party data, and sending the feedback information to the execution node.
[0359] Optionally, the system 120 further includes a management node 1203;
[0360] The management node 1203 is used for monitoring the access behavior of the data user to the data management system through the API, and monitoring the data output behavior of the data management system through the API.
[0361] Optionally, the data reading policy is determined by the data provider. The data reading policy includes policy items and operations determined based on the policy items. The policy items include one or more of the following items:
[0362] Data policy item, usage policy item, quantity policy item, time policy item, subject policy item;
[0363] Among them, the data policy item is used to describe the data range in the third-party data that can be read;
[0364] The usage policy item is used to describe the tasks in the business program that can use the third-party data;
[0365] The quantity policy item is used to describe the usage quantity threshold of the third-party data;
[0366] The time policy item is used to describe the usage time range of the third-party data;
[0367] The subject policy item is used to describe the subject types to which the data reading policy can be applied.
[0368] Optionally, the storage node 1202 includes one or more data reading policies regarding the third-party data;
[0369] The storage node 1202 is used for:
[0370] Matching the data reading request with the policy items of at least one data reading policy among the one or more data reading policies;
[0371] If there is a data reading policy that matches the data reading request, generate feedback information according to the operations in the data reading policy that matches the data reading request and send the feedback information to the execution node.
[0372] Optionally, the storage node 1202 further includes a data permission policy, and the data permission policy is used to determine the read and write permissions for the data according to the source of the data. Among them, the data permission policy indicates that the data user has the read permission for the third-party data from the data provider;
[0373] The storage node 1202 is further used for:
[0374] According to the data permission policy, based on the fact that the first data belongs to the third-party data, determine that the business program has the read permission for the first data.
[0375] Optionally, the management node 1203 is used for:
[0376] After obtaining the operation instruction for the business program, create an access credential for the storage node;
[0377] Instruct the execution node to run the business program and transfer the information of the access credential to the execution node;
[0378] The execution node 1201 is used for:
[0379] Based on the information of the access credential, send a data reading request to the storage node.
[0380] Optionally, the execution node 1201 is further configured to: delete the access credentials after the execution node stops running the service program.
[0381] Optionally, the storage node 1202 is further configured to: record the data access behaviors related to the data read request.
[0382] Optionally, the storage node 1202 is further configured to:
[0383] Obtain a query request input by the data provider, where the query request is used to request to query the access records related to the third-party data;
[0384] Output the access records, where the access records include the data access behaviors related to the data read request.
[0385] Optionally, the execution node 1201 is further configured to: send the data to be stored to the storage node during the process of running the service program.
[0386] The storage node 1202 is further configured to: after receiving the data to be stored, determine and store the derivative link of the data to be stored according to the third-party data read by the execution node from the storage node within the specified time period. The specified time period includes the time period when the execution node runs the service program this time. The derivative link is used to indicate the data in the storage node that is used to generate the data to be stored.
[0387] Optionally, the execution node 1201 is further configured to: delete the running data related to the service program in the execution node after the execution node stops running the service program.
[0388] In the embodiments of the present application, the execution node, the storage node, and the management node can all be implemented by software or can be implemented by hardware. Exemplarily, next, taking the execution node as an example, the implementation manner of the execution node will be introduced. Similarly, the implementation manners of the storage node and the management node can refer to the implementation manner of the execution node.
[0389] As an example of a software functional unit, an execution node may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, an execution node may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.
[0390] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set up within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC, and the interconnection between VPCs is achieved through the communication gateway.
[0391] As an example of a hardware functional unit, an execution node may include at least one computing device, such as a server. Alternatively, an execution node may also be a device implemented using a central processing unit (CPU), or may be implemented using an application-specific integrated circuit (ASIC), or a programmable logic device (PLD). The above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), a data processing unit (DPU), a neural network processing unit (NPU), a system on chip (SoC), an offload card, an acceleration card, or any combination thereof.
[0392] The multiple computing devices included in the execution node can be distributed in the same region or in different regions. The multiple computing devices included in the execution node can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the execution node can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), generic array logic (GALs), data processing units (DPUs), neural processing units (NPUs), system-on-chips (SoCs), offloading cards, and acceleration cards.
[0393] It should be noted that in other embodiments, the execution node can be used to execute any step in the data management method, the storage node can be used to execute any step in the data management method, and the management node can be used to execute any step in the data management method. The steps to be implemented by the execution node, the storage node, and the management node can be specified as needed. The full functions of the data management system can be realized by implementing different steps in the data management method through the execution node, the storage node, and the management node respectively.
[0394] The embodiment of the present application also provides a computing device 130. As Figure 13 shown, the computing device 130 includes: a bus 132, a processor 134, a memory 136, and a communication interface 138. The processor 134, the memory 136, and the communication interface 138 communicate with each other through the bus 132. The computing device 130 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 130.
[0395] The bus 132 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 13 only one line is shown here, but it does not mean that there is only one bus or one type of bus. The bus 134 can include a path for transmitting information between various components (such as the memory 136, the processor 134, and the communication interface 138) of the computing device 130.
[0396] The processor 134 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0397] The memory 136 may include a volatile memory, such as a random access memory (RAM). The memory 136 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0398] The executable program code is stored in the memory 136, and the processor 134 executes the executable program code to respectively implement the functions of the foregoing execution node and storage node, thereby implementing the data management method applied to the computing device cluster in the above embodiments. That is, the memory 136 stores instructions for executing the data management method applied to the computing device cluster in the above embodiments.
[0399] The communication interface 138 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 130 and other devices or a communication network.
[0400] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server or an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0401] As Figure 14 shown, the computing device cluster includes at least one computing device 130. The same instructions for executing the data management method may be stored in the memory 136 of one or more computing devices 130 in the computing device cluster.
[0402] In some possible implementations, the memories 136 of one or more computing devices 130 in the computing device cluster may also store some instructions for executing the data management method respectively. In other words, the combination of one or more computing devices 130 can jointly execute the instructions for executing the data management method.
[0403] It should be noted that the memories 136 in different computing devices 130 in the computing device cluster may store different instructions, respectively for executing some functions of the data management method. That is to say, the instructions stored in the memories 136 of different computing devices 130 can implement the functions of one or more modules in the execution node and the storage node.
[0404] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 15 A possible implementation is shown. As Figure 15 shown, two computing devices 130A and 130B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 136 in the computing device 130A may store instructions for executing the functions of the execution node. At the same time, the memory 136 in the computing device 130B may store instructions for executing the functions of the storage node. Or, the memory 136 in the computing device 130A may store instructions for executing some functions of the storage node. At the same time, the memory 136 in the computing device 130B may store instructions for executing another part of the functions of the storage node.
[0405] It should be understood that Figure 15 the functions of the computing device 130A shown in
[0406] can also be completed by multiple computing devices 130. Similarly, the functions of the computing device 130B can also be completed by multiple computing devices 130. Figure 14 and Figure 15 the connection manner of the computing device cluster. The difference is that the memories 136 of one or more computing devices 130 in this computing device cluster may store the same instructions for executing the data management method.
[0407] In some possible implementations, the memories 136 of one or more computing devices 130 in the computing device cluster may also store some instructions for executing the data management method respectively. In other words, the combination of one or more computing devices 130 can jointly execute the instructions for executing the data management method.
[0408] It should be noted that the memories 136 in different computing devices 130 in the computing device cluster may store different instructions for performing partial functions of the data management method. That is, the instructions stored in the memories 136 in different computing devices 130 can implement the functions of one or more modules in the execution node and the storage node.
[0409] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the data management method.
[0410] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the data management method.
[0411] The embodiments of the present application also provide a chip system, which includes a processor for implementing the steps executed by the above-mentioned computing device cluster. In a possible design, the chip system may further include a memory for storing necessary program instructions and data. The chip system can be composed of chips or can include chips and other discrete devices.
[0412] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0413] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0414] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0415] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may be physically present individually for each unit, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0416] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
Claims
1. A data management method, characterized in that, Applied to a data management system, the data management system is set up in at least one data center, the data management system is located in an isolated operating environment, the isolated operating environment is connected to a data provider and a data user respectively through an application programming interface (API), the data management system includes an execution node and a storage node, the storage node stores third-party data of the data provider, and the method includes: The execution node runs a business program, and the business program is used to execute the business of the data user; The execution node sends a data reading request to the storage node, and the data reading request indicates that the running business program requests to read first data from the storage node, and the first data belongs to the third-party data; The storage node generates feedback information according to the data reading request and a data reading strategy for the third-party data, and sends the feedback information to the execution node.
2. The method according to claim 1, characterized in that The method further includes: Monitoring the access behavior of the data user to the data management system through the API, and monitoring the data output behavior of the data management system through the API.
3. The method according to claim 1 or 2, characterized in that The data reading strategy is determined by the data provider, and the data reading strategy includes a policy item and an operation determined based on the policy item. The policy item includes one or more of the following items: Data policy item, usage policy item, quantity policy item, time policy item, subject policy item; Among them, the data policy item is used to describe the data range in the third-party data that can be read; The usage policy item is used to describe the tasks in the business program that can use the third-party data; The quantity policy item is used to describe the usage quantity threshold of the third-party data; The time policy item is used to describe the usage time range of the third-party data; The subject policy item is used to describe the subject type to which the data reading strategy can be applied.
4. The method according to claim 3, characterized in that The storage node includes one or more data reading strategies for the third-party data; The storage node generates feedback information according to the data reading request and a data reading strategy for the third-party data, and sends the feedback information to the execution node, including: The storage node matches the data reading request with the policy items of at least one data reading strategy among the one or more data reading strategies; If there is a data reading strategy that matches the data reading request, the storage node generates the feedback information according to the operation in the data reading strategy that matches the data reading request, and sends the feedback information to the execution node.
5. The method according to any one of claims 1-4, characterized in that The storage node further includes a data permission policy, and the data permission policy is used to determine the read and write permissions for the data according to the data source. Among them, the data permission policy indicates that the data user has the read permission for the third-party data from the data provider; Before the storage node generates feedback information according to the first data and a data reading strategy for the third-party data, and sends the feedback information to the execution node, it further includes: The storage node determines that the service program has read permission for the first data based on the data permission policy and that the first data belongs to the third-party data.
6. The method according to any one of claims 1-5, characterized in that, The data management system includes a management node, and the method further includes: After the management node obtains a running instruction for the service program, it creates an access credential for the storage node; The management node instructs the execution node to run the service program and transmits information about the access credential to the execution node; The execution node sends a data read request to the storage node, including: The execution node sends the data read request to the storage node based on the information of the access credential.
7. The method according to claim 6, characterized in that, The method further includes: After the execution node stops running the service program, the access credential is deleted.
8. The method according to any one of claims 1 to 7, characterized in that, After the storage node generates feedback information according to the first data and the data read policy for the third-party data and sends the feedback information to the execution node, it further includes: The storage node records the data access behavior related to the data read request.
9. The method according to claim 8, wherein The method further includes: The storage node obtains a query request input by the data provider, and the query request is used to request a query of the access record related to the third-party data; The storage node outputs the access record, and the access record includes the data access behavior related to the data read request.
10. The method according to any one of claims 1-9, characterized in that The method further includes: During the running of the service program, the execution node sends data to be stored to the storage node; After receiving the data to be stored, the storage node determines and stores the derivative link of the data to be stored according to the third-party data read by the execution node from the storage node within a specified time period. The specified time period includes the time period when the execution node runs the service program this time, and the derivative link is used to indicate the data in the storage node for generating the data to be stored.
11. The method according to any one of claims 1 to 10, characterized in that The method further includes: After the execution node stops running the service program, the running data related to the service program in the execution node is deleted.
12. A data management system, characterized in that, The data management system is set up in at least one data center. The data management system is located in an isolated running environment, and the isolated running environment is connected to the data provider and the data user respectively through an application programming interface (API). The data management system includes an execution node and a storage node, and the storage node stores the third-party data of the data provider; The execution node is used for: Running a service program, and the service program is used to execute the service of the data user; Sending a data read request to the storage node, and the data read request indicates that the running service program requests to read first data from the storage node, and the first data belongs to the third-party data; The storage node is used for: generating feedback information according to the data read request and the data read policy for the third-party data and sending the feedback information to the execution node.
13. The system according to claim 12, wherein The system further includes a management node; The management node is used to monitor the access behavior of the data consumer to the data management system through the API, and monitor the data output behavior of the data management system through the API.
14. The system according to claim 12 or 13, wherein The data reading policy is determined by the data provider, and the data reading policy includes policy items and operations determined based on the policy items. The policy items include one or more of the following items: Data policy item, usage policy item, quantity policy item, time policy item, subject policy item; Among them, the data policy item is used to describe the data range in the third-party data that can be read; The usage policy item is used to describe the tasks in the business program that can use the third-party data; The quantity policy item is used to describe the usage quantity threshold of the third-party data; The time policy item is used to describe the usage time range of the third-party data; The subject policy item is used to describe the type of subject to which the data reading policy can be applied.
15. The system according to claim 14, wherein The storage node includes one or more data reading policies regarding the third-party data; The storage node is used for: Matching the data reading request with the policy items of at least one data reading policy among the one or more data reading policies; If there is a data reading policy that matches the data reading request, generating the feedback information according to the operation in the data reading policy that matches the data reading request and sending the feedback information to the execution node.
16. The system according to any one of claims 12 - 15, characterized in that, The storage node further includes a data permission policy, and the data permission policy is used to determine the read and write permissions for the data according to the data source. Among them, the data permission policy indicates that the data consumer has read permission for the third-party data from the data provider; The storage node is further used for: According to the data permission policy, based on the fact that the first data belongs to the third-party data, determining that the business program has read permission for the first data.
17. The system according to any one of claims 12-16, characterized in that, The system further includes a management node; The management node is used for: After obtaining the operation instruction for the business program, creating an access credential for the storage node; Instructing the execution node to run the business program and transmitting the information of the access credential to the execution node; The execution node is used for: Based on the information of the access credential, sending the data reading request to the storage node.
18. The system according to claim 17, wherein The execution node is further used for: after the execution node stops running the business program, deleting the access credential.
19. The system according to any one of claims 12 - 18, wherein The storage node is further used for: recording the data access behavior related to the data reading request.
20. The system according to claim 19, wherein The storage node is further used for: Obtaining the query request input by the data provider, where the query request is used to request to query the access records related to the third-party data; Outputting the access records, where the access records include the data access behavior related to the data reading request.
21. The system according to any one of claims 12-20, wherein: The execution node is further configured to: during the running of the service program, send data to be stored to the storage node; The storage node is further configured to: after receiving the data to be stored, determine and store the derivative link of the data to be stored according to the third-party data read by the execution node from the storage node within a specified time period, where the specified time period includes the time period when the execution node runs the service program this time, and the derivative link is used to indicate the data in the storage node for generating the data to be stored.
22. The system according to any one of claims 12-21, wherein: The execution node is further configured to: after the execution node stops running the service program, delete the running data related to the service program in the execution node.
23. A cluster of computing devices, characterized in that, Comprising at least one computing device, the at least one computing device comprising a processor and a memory; The processor is configured to execute instructions stored in the memory to cause the computing device cluster to execute the method according to any one of claims 1-11.
24. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program runs on the processor, the processor is caused to execute the method according to any one of claims 1-11.
25. A computer program product containing instructions, characterized in that, When the instructions are executed by the processor, the method according to any one of claims 1-11 is implemented.
Citation Information
Cited By
Data management method and related device
WO2025148456A1