Data management system and method

By implementing access policy management and operation/data logging in federated learning systems, transparent data usage is tracked, maintaining access rights and ensuring compliance, which supports continued data provider participation and personalized medicine.

JP7854956B2Active Publication Date: 2026-05-07HITACHI LTD
View PDF 13 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HITACHI LTD
Filing Date
2023-03-23
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The management of privacy data in federated learning systems is challenging due to the difficulty in tracking data usage and implementing transparent access policies, leading to potential decreases in data provider participation and hindering personalized medicine.

Method used

A storage device manages access policies for entities, logging operations and data usage, and maintains entity lists to ensure transparent data management by tracking and controlling access rights and usage history.

Benefits of technology

This approach enables transparent data management, maintaining data access rights and ensuring compliance with user preferences, thereby preserving data provider participation and supporting personalized medicine.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007854956000001
    Figure 0007854956000001
  • Figure 0007854956000002
    Figure 0007854956000002
  • Figure 0007854956000003
    Figure 0007854956000003
Patent Text Reader

Abstract

To provide data management that is transparent while maintaining access authority for data.SOLUTION: For each entity, there is an access policy including access authority of each (n)th-order data for use by an application or model. For each operation, there is a data log which is a log of data on the source or target of the operation and associated with an operation log. With respect to the operation log and / or data log, there is an entity list based upon the access policy. When one or more entity lists in which an entity specified based upon a request is recorded are found among a plurality of entity lists, a processor specifies a use state at the request based upon one or more operation logs and one or more data logs specified by using the one or more entity lists, and returns data representing the specified use state to the request source.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to data management.

Background Art

[0002] As an example of data to be managed, there is privacy data such as medical data of patients. In order to provide personalized medicine, it is necessary to segment patient characteristics at a fine granularity, and a large amount of privacy data is required. Federated learning is known as model learning using privacy data.

[0003] Patent Document 1 discloses a technique related to enhancing privacy. Patent Document 2 discloses a technique related to federated learning. Patent Document 3 discloses a technique related to machine learning.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0005] With the advent of a data circulation society, due to secondary and tertiary data utilization, data providers may suffer unexpected disadvantages. Although it is possible to request the suspension, deletion, or disclosure of privacy data by an opt-out application, it is difficult to implement an appropriate opt-out application if it is not known how the data is being used. Also, if the usage history is opaque (difficult to trace), there is a risk that the number of data providers will decrease and it will become difficult to provide personalized medicine.

[0006] These points highlight the need for careful management of private data in federative learning, and the desire for transparent federative learning services. Furthermore, similar challenges may arise in data management outside of federative learning.

[0007] Patent documents 1-3 do not disclose or suggest any such problems or solutions to them. [Means for solving the problem]

[0008] The storage device stores access policies for each entity. For each entity, the access policy includes, for each of one or more operation attributes, access rights per nth data for use in the application or model for that entity (where n is two or more integers such that n ≥ 0). For each operation, the storage device stores an operation log, which is the log of that operation; a data log corresponding to the source data, which is the data for that operation, and associated with the operation log; and / or a data log corresponding to the target data, which is the data resulting from that operation, and associated with the operation log. The storage device stores entity lists based on access policies for the operation logs and / or data logs. Each entity list includes an inclusion list, which is a list of entities whose data is permitted for use in the operation or data corresponding to the log to which the entity list is associated; and / or an exclusion list, which is a list of entities whose data is prohibited from being used in the operation or data corresponding to the log to which the entity list is associated. In response to a request, if the processor finds one or more entity lists containing the entities identified based on the request among multiple entity lists, it identifies usage based on one or more operation logs and one or more data logs identified using those entity lists, and returns data representing the identified usage to the requester. [Effects of the Invention]

[0009] According to the present invention, it is possible to provide transparent data management while maintaining data access rights. [Brief explanation of the drawing]

[0010] [Figure 1] This shows the overall system configuration for one embodiment of the present invention. [Figure 2] This shows the configuration of a client-server system as a data management system. [Figure 3] This diagram schematically illustrates the processing flow up to feature storage on the client computer. [Figure 4] This diagram schematically illustrates the processing flow for learning on the client computer and associative learning on the server computer. [Figure 5] Multiple operation logs are shown. [Figure 6] Multiple data logs are shown. [Figure 7] Shows multiple user lists. [Figure 8] Show multiple operation lists. [Figure 9] Show multiple access policies. [Figure 10A] This shows the access policy before the access permissions were changed. [Figure 10B] This shows the access policy after the access permissions have been changed. [Figure 11A] This shows the usage data before the access permission change. [Figure 11B] This shows the usage data after the access permissions were changed. [Figure 12A] This indicates the lineage represented by the lineage data before the access permission change. [Figure 12B] This indicates the lineage represented by the lineage data after the access permission has been changed. [Figure 13A] An example of contribution data before the access permissions were changed is shown. [Figure 13B] An example of contribution data after a change in access permissions is shown. [Figure 14] Shows the functions of the client computer and the server computer. [Figure 15] Shows the flow of usage management processing. [Figure 16] Shows the flow of lineage management processing. [Figure 17] Shows the flow of data processing. [Figure 18] Shows the flow of learning processing. [Figure 19] Shows the flow of inference processing. [Figure 20] Shows the flow of contribution management processing. [Figure 21] Shows the flow of the first access management processing. [Figure 22] Shows the flow of the second access management processing.

Mode for Carrying Out the Invention

[0011] In the following description, the "interface device" may be one or more interface devices. The one or more interface devices may be at least one of the following. · One or more I / O (Input / Output) interface devices. The I / O (Input / Output) interface device is an interface device for at least one of an I / O device and a remote display computer. The I / O interface device for the display computer may be a communication interface device. At least one I / O device may be any of a user interface device, for example, an input device such as a keyboard and a pointing device, and an output device such as a display device. · One or more communication interface devices. The one or more communication interface devices may be one or more of the same type of communication interface devices (for example, one or more NICs (Network Interface Cards)) or two or more different types of communication interface devices (for example, a NIC and an HBA (Host Bus Adapter)).

[0012] Furthermore, in the following explanation, "memory" refers to one or more memory devices, which are examples of one or more storage devices, and may typically be main memory devices. At least one memory device in memory may be a volatile memory device or a non-volatile memory device.

[0013] Furthermore, in the following explanation, "persistent storage device" may refer to one or more persistent storage devices, which are examples of one or more storage devices. Persistent storage devices are typically non-volatile storage devices (e.g., auxiliary storage devices), and specifically may be, for example, HDDs (Hard Disk Drives), SSDs (Solid State Drives), NVME (Non-Volatile Memory Express) drives, or SCMs (Storage Class Memory).

[0014] Furthermore, in the following explanation, "storage device" may refer to at least memory, including both memory and persistent storage.

[0015] Furthermore, in the following explanation, "processor" may refer to one or more processor devices. At least one processor device may typically be a microprocessor device such as a CPU (Central Processing Unit), but may also be other types of processor devices such as a GPU (Graphics Processing Unit). At least one processor device may be single-core or multi-core. At least one processor device may be a processor core. At least one processor device may be a broad processor device such as a circuit that is a collection of gate arrays according to a hardware description language that performs some or all of the processing (e.g., FPGA (Field-Programmable Gate Array), CPLD (Complex Programmable Logic Device), or ASIC (Application Specific Integrated Circuit)).

[0016] Furthermore, in the following explanation, functions may be described using the expression "yyy section," but a function may be realized by the execution of one or more computer programs by a processor, by one or more hardware circuits (e.g., FPGA or ASIC), or by a combination thereof. When a function is realized by the execution of a program by a processor, the defined processing is carried out using memory and / or interface devices as appropriate, so the function may be at least a part of the processor. Processing described with a function as the subject may be processing performed by the processor or a device having that processor. Programs may be installed from program source. Program source may be, for example, a program distribution computer or a storage medium that the computer can read (e.g., a non-temporary storage medium). The description of each function is an example, and multiple functions may be combined into one function, or one function may be divided into multiple functions.

[0017] Furthermore, in the following explanation, the expression "xxxDS" may be used to describe information from which an output is obtained for a given input. This information can be data of any structure (for example, structured data or unstructured data), or a learning model such as a neural network, genetic algorithm, or random forest that generates an output for a given input. Therefore, "xxxDS" can be referred to as "xxx information." Also, in the following explanation, the structure of each table is just an example; one table may be divided into two or more tables, or all or part of two or more tables may be combined into one table. Note that "DS" is an abbreviation for dataset, and can be, for example, a database.

[0018] Furthermore, any information (for example, at least one of "ID," "Name," and "Number") may be used as information to identify an element (identification information, identifier).

[0019] Furthermore, in the following explanation, when describing similar elements without distinction, the common reference code will be used, and when describing similar elements with distinction, the reference code will be used.

[0020] An embodiment of the present invention will be described below with reference to the drawings. However, the present invention is not limited by the following description.

[0021] Figure 1 shows the overall system configuration of one embodiment of the present invention.

[0022] A computer system including multiple computers 40 constitutes the data management system according to this embodiment. The data managed by the data management system according to this embodiment is medical data relating to the user's medical care (for example, data representing information such as date of birth, gender, and medical history). However, the present invention can also manage data other than medical data, for example, data relating to insurance, finance, or security (for example, various types of privacy data).

[0023] At least one computer 40 can communicate with a user terminal 101. The user terminal 101 may be an information processing terminal such as a personal computer or a smartphone. Each computer 40 and each user terminal 101 is connected to a communication network 35. Communication between computers 40 and communication between computers 40 and user terminals 101 are conducted via the communication network 35. The communication network for communication between computers 40 and the communication network for communication between computers 40 and user terminals 101 may be different.

[0024] Computer 40 may be a general-purpose computer or a dedicated computer. Computer 40 may be a physical computer or a logical computer (for example, a virtual machine or cloud computing service).

[0025] The computer 40 illustrated in Figure 1 is a physical computer. The computer 40 has an interface device 170, a storage device 120, and a processor 160 connected thereto. Furthermore, the computer 40 can act as either a client or a server, or both. The data and functions of the computer 40 differ depending on which role it plays. The computer 40 illustrated in Figure 1 is capable of acting as both a client and a server.

[0026] The interface device 170 communicates via the communication network 35. The storage device 120 stores data and programs. The processor 160 executes programs.

[0027] The data stored in the storage device 120 includes, for example, user list DS132, access policy DS133, usage DS134, provision DS135, application DS136, lineage DS137, feature DS138, operation list DS139, contribution DS140, model DS141, log DS142, and promotion DS143. User list DS132 is a set of user lists. Access policy DS133 is a set of access policies for each user. Usage DS134 is a set of usage data (for example, data representing the usage history of applications or models). Provision DS135 is a set of provision data (data available for data processing). Application DS136 is a set of application data (for example, data including data of the application program itself and metadata of the application program). Lineage DS137 is a set of lineage data (for example, lineage data representing how an application used the provision data). Feature DS138 is a set of feature data for the provision data. Feature data may include features for each user represented by the provided data, or it may include features for the entire provided data. Operation list DS139 is a set of operation lists. Contribution DS140 is a set of contribution data (e.g., data representing application or model contributions). Model DS141 is a set of model data (e.g., data including data for the model itself or model metadata). Log DS142 is a set of logs. Promotion DS143 is a set of promotion data (e.g., data including data representing the conditions for promotions to users). Each DS may consist of one data, but typically contains multiple data.

[0028] The processor 160 executes the program in the storage device 120, thereby realizing functions such as the federated learning control unit 121, federated learning unit 122, machine learning control unit 123, data loading unit 124, learning unit 125, data preprocessing unit 126, inference unit 127, usage management unit 128, lineage management unit 129, access management unit 130, and contribution management unit 131. These functions 121 to 131 will be described later. These functions 121 to 131 are realized on the OS (Operating System) 150, but at least some of these functions may be incorporated into the OS 150.

[0029] Figure 2 shows the configuration of a client-server system as a data management system.

[0030] The data management system includes multiple client computers 100 and a server computer 110. Each of the multiple client computers 100 may be a physical computer located in a hospital or other location, or a logical computer in the cloud, etc. The server computer 110 may also be a physical computer or a logical computer. Federative learning is performed in this data management system.

[0031] The client computer 100 is a computer that acts as a client. The server computer 110 is a computer that acts as a server. Either or both of the client computer 100 and the server computer 110 are capable of receiving requests (queries) from the user terminal 101. In Figure 2, for the same elements present in both the client computer 100 and the server computer 110, the reference code of the element in the client computer 100 ends with "C", and the reference code of the element in the server computer 110 ends with "S". Requests may be sent from the client computer 100 to the server computer 110, or requests may be sent from the server computer 110 to the client computer 100.

[0032] The client computer 100 includes a machine learning control unit 123, a privacy DS201, a data loading unit 124, a provision DS135, a data preprocessing unit 126, a feature DS138, a learning unit 125, a model DS141C, an inference unit 127, a user list DS132, an access policy DS133C, an application DS136C, a log DS142, a usage DS134C, a lineage DS137C, a contribution DS140C, and a promotion DS143C.

[0033] The server computer 110 includes a federated learning control unit 121, a federated learning unit 122, a model DS141S, an operation list DS139, an access policy DS133S, an application DS136S, a log DS142S, a usage DS134S, a lineage DS137S, a contribution DS140S, and a promotion DS143S.

[0034] This data management system performs the following processes, for example:

[0035] The privacy data DS201 is stored in the storage device 120 of the client computer 100. The privacy data DS201 is often, for example, a hospital database and contains privacy data. The privacy data is so-called raw data and includes medical data for each of multiple users. The privacy data may exist for each period, for example. For example, there may be privacy data for January, privacy data for February, and so on.

[0036] On the client computer 100, the data loading unit 124, in response to a request from the machine learning control unit 123, retrieves privacy data from the privacy DS201, converts the privacy data into provided data (data usable for data processing), and includes the provided data in the provided DS135. The data preprocessing unit 126, in response to a request from the machine learning control unit 123, retrieves provided data from the provided DS135, generates feature data for that provided data, and includes that feature data in the feature DS138. The learning unit 125, in response to a request from the machine learning control unit 123, creates a model using the feature data in the feature DS138 (or retrains the model in the model DS141C), and includes the model data in the model DS141C. The inference unit 127 performs inference using the model in the model DS141C.

[0037] On the client computer 100, data logs and operation logs are generated as appropriate and stored in log DS142C. Additionally, a user list is generated as appropriate and stored in user list DS132. Furthermore, the machine learning control unit 123 refers to or updates user list DS132, access policy DS133C, application DS136C, log DS142, usage DS134C, lineage DS137C, contribution DS140C, and promotion DS143C as appropriate.

[0038] In this embodiment, the applications are as follows. Application DS136 includes data representing each application (e.g., AID and other data). The AID and other data relating to the application in this embodiment can be identified from Application DS136. • Data loading process (This process involves retrieving data from the privacy DS201, performing filtering and anonymization, and storing the processing results in the provision DS135). • Data preprocessing (This involves obtaining data from the provided DS135, generating features for use in training, and storing these features in the feature array DS138). • Learning process (This process involves obtaining features from feature array DS138, generating a model based on these features, and storing the model in model DS141C). • Associative learning process (a process that works in conjunction with the above learning process, exchanging model information (also called parameters) between the server computer 110 and the client computer 100, generating an integrated model from multiple models on the server computer 110 side, and storing the integrated model in model DS141S). • Inference processing (the process of obtaining the model from model DS141C, obtaining the data to be inferred from privacy DS201, provision DS135, and feature DS138, and generating inference results using the model and data).

[0039] These processes (applications) may be instructed from the user terminal 101 to the client computer 100 or the server computer 110 as appropriate (or all at once). In addition, in order to prevent the leakage of data relating to individual users (for example, privacy data or related data), the federated learning process described above may be executed on the server computer 110, while the processes other than the federated learning process described above may be executed on the client computer 100.

[0040] In client computer 100, the machine learning control unit 123 acquires model data for model DS141C and transmits that model data to server computer 110. Server computer 110 receives model data from multiple client computers 100, and federated learning control unit 121 incorporates the received model data into model DS141S. In server computer 110, the federated learning unit 122, in response to a request from the federated learning control unit 121, acquires model data from multiple client computers 100 from model DS141S, generates model data based on that model data (for example, generates merged model data), and transmits the generated model data to each of the multiple client computers 100. In client computer 100 that has received model data from server computer 110, the machine learning control unit 123 incorporates that model data into model DS141C. Thereafter, the model represented by the latest model data received from server computer 110 is used for learning by the learning unit 125 or for inference by the inference unit 127.

[0041] In the server computer 110, the federated learning control unit 121 appropriately refers to or updates the operation list DS139, access policy DS133S, application DS136S, log DS142S, usage DS134S, lineage DS137S, contribution DS140S, and promotion DS143S. For example, DS133S, 136S, 142S, 134S, 137S, 140S, and 143S in the server computer 110 may include DS133C, 136C, 142C, 134C, 137C, 140C, and 143C in each client computer 100.

[0042] Figure 3 schematically shows the processing flow up to feature storage on the client computer 100.

[0043] The data log 31-1 of the privacy data in the privacy DS201 is stored in the log DS142C. The data load unit 124 retrieves the privacy data corresponding to the data log 31-1, generates provided data based on that privacy data, and includes the provided data in the provided DS135. In this series of operations (processing), the data load unit 124 generates an operation log 32-1 based on the data log 31-1, and generates a data log 31-2 corresponding to the stored provided data based on the operation log 32-1. The data load unit 124 also generates a user list 33-1 associated with the operation log 32-1, and a user list 33-3 associated with the data log 31-2.

[0044] The data preprocessing unit 126 acquires the provided data, generates three types of feature data based on that data, and includes these three types of feature data in the feature DS138. In this series of operations (processing), the data preprocessing unit 126 generates an operation log 32-3 based on the data log 31-2, and generates data logs 31-a to 31-c corresponding to the stored three types of feature data based on the operation log 32-3. The data loading unit 124 also generates a user list 33-3 associated with the operation log 32-3, and generates user lists 33-a to 33-c associated with the data logs 31-a to 31-c.

[0045] The machine learning control unit 123 includes data logs 31-1, 31-2, and 31-a to 33-c, and operation logs 32-1 and 32-3 in log DS142C. The machine learning control unit 123 also includes user lists 33-1 to 33-3 and 33-a to 33-c in user list DS132.

[0046] As shown in the example in Figure 3, three types of feature data are generated from a single set of data (feature data used for model training, feature data used for model validation, and feature data used for model evaluation). However, the number of types of feature data generated may be more or less than three.

[0047] Additionally, a user list 33 is generated for each data log and each operation log. The user list 33 corresponding to the data includes a list of users whose data is included in the corresponding data (data corresponding to the data log), and / or a list of users whose data is not included in the corresponding data. The user list 33 corresponding to the operation log includes a list of users whose data is included in the target data (data as the target of the operation corresponding to the operation log), and / or a list of users whose data is not included in the target data.

[0048] Figure 4 schematically shows the processing flow related to learning on the client computer 100 and associative learning on the server computer 110.

[0049] In this process as explained with reference to Figure 3, a user list 33 is generated for each data log and each operation log.

[0050] On the client computer 100A, the learning unit 125A acquires feature data and performs model training, validation, or evaluation. Taking training as an example, the learning unit 125A acquires training feature data, trains the model, and sends the trained model data to the server computer 110. In this series of operations (processes), the learning unit 125A generates an operation log 32-4 based on the data log 31-a1 corresponding to the training feature data. The user list 33-4 is associated with the operation log 32-4.

[0051] On client computer 100A, the learning unit 125A includes the data of the trained model in model DS141C. The learning unit 125A generates a data log 31-6 corresponding to this data. A user list 33-6 is associated with the data log 31-6.

[0052] On the client computer 100A, the inference unit 127A performs inference using the trained model. The inference unit 127 generates an operation log 32-7 corresponding to this inference. A user list 33 may be associated with this operation log 32-7.

[0053] On the client computer 100B, the learning unit 125B acquires feature data for training, trains the model, and sends the trained model data to the server computer 110. In this series of operations (processes), the learning unit 125B generates an operation log 32-5 based on the data log 31-a2 corresponding to the feature data for training. The user list 33-5 is associated with the operation log 32-5.

[0054] On the server computer 110, the federated learning unit 122 generates model data by merging model data from client computers 100A and 100B, and includes the merged model data in the model DS141S. The federated learning unit 122 generates an operation log 32-8 corresponding to this federated learning, and also generates a data log 31-9 corresponding to the merged model data. The federated learning unit 122 generates an operation list 34-8 associated with the operation log 32-8, and also generates an operation list 34-9 associated with the data log 31-9. The federated learning control unit 121 includes logs 32-8 and 31-9 in the log DS142S, and includes operation lists 34-8 and 34-9 in the operation list DS139.

[0055] The transmission of model data from multiple client computers 100 to a server computer 110, and the merging (federated learning) of that model data on the server computer 110, may be repeated. The server computer 110 may send the merged model data with the highest accuracy to each client computer 100. That is, the transmission of the merged model data to each client computer 100 may be performed each time a merged model is obtained, but it may also be performed after repeated model merging.

[0056] Figure 5 shows multiple operation logs 32.

[0057] Figure 5 shows multiple rows, but each row represents one operation log 32. An operation log 32 contains information such as Log ID 501, Location 502, Application ID 503, Date 504, Type 505, Sub Type 506, Source 507, Target 508, List ID 509, Usage 510, and Reproduction 511. We will explain these pieces of information 501-511 using one operation log 32 as an example.

[0058] Log ID 501 represents the OID (Operation ID) of the corresponding operation (the operation corresponding to Operation Log 32). Location 502 represents the location where the corresponding operation was performed.

[0059] Application ID 503 represents the Application ID (AID) of the application to which the corresponding operation belongs. In the AID, "DP" stands for Data Preprocessing, "LT" stands for Learning (Training), "LV" stands for Learning (Validation), "LE" stands for Learning (Evaluation), "I" stands for Inference, and "FL" stands for Federated Learning.

[0060] Date504 represents the date and time the corresponding operation was performed (the unit of time can be year, month, day, hour, minute, second, or it can be coarser or finer). Type505 represents the type of function on which the corresponding application was performed. Sub Type506 represents the details of the corresponding application (specifically, whether it is Training, Validation, or Evaluation) when the function on which the corresponding application was performed is the learning unit 125.

[0061] Source507 represents the ID of the data used as input in the corresponding operation. Target508 represents the ID of the data used as output in the corresponding operation. "PDID" means the ID of the provided data. "FSID" means the ID of the feature data. "MID" means the ID of the model.

[0062] List ID 509 represents the ID of either User List 33 or Operation List 34 associated with the corresponding operation. "ULID" means the ID of User List 33. "OLID" means the ID of Operation List 34.

[0063] Usage 510 indicates whether or not the application to which the corresponding operation belongs is permitted to be used. "Approve" means that use is permitted. "Disapprove" means that use is prohibited. If the application provider does not prohibit the use of the application, "Approve" (usable) may follow.

[0064] Reproduction 511 indicates whether the output data (data identified from Target 508) can be reproduced by the corresponding operation. "Possible" means that reproduction is possible. "Impossible" means that reproduction is not possible.

[0065] According to the multiple operation logs 32 shown in Figure 5, for example, the following can be observed: The operation corresponding to Location502 "US1" was performed on one client computer 100, and the operation corresponding to Location502 "US2" was performed on another client computer 100. The operation corresponding to Location502 "Cloud-1" was performed on server computer 110. According to Source 507 and Target 508 of Operation Log 32 for Log ID 501 "OID-001" to "OID-003", the data was input and output in the following flow: Privacy DS201 → Provision DS135 → Feature DS138 → Model DS141C. According to the operation log 32 for Log ID 501 “OID-005” and “OID-006”, data was obtained from feature vector DS138, inference was performed using the model in model DS141C, and the data output by the inference was provided to the user (user terminal 101).

[0066] Figure 6 shows multiple data logs 31.

[0067] Figure 6 shows multiple rows, but each row represents one data log 31. A data log 31 contains information such as Log ID 601, Location 602, Date 603, Type 604, Sub Type 605, Data 606, Version 607, List ID 608, Usage 609, and Reproduction 610. We will explain these pieces of information 601-610 using one data log 31 as an example.

[0068] Log ID 601 represents the DID (Data ID) of the corresponding data. Here, "corresponding data" refers to the data corresponding to the operation, and the DID is the log ID for that data. Location 602 represents the location where the corresponding data exists. Date 603 represents the date and time the corresponding data was obtained.

[0069] Type 604 represents the type of data it corresponds to. Sub Type 605 represents the process or application that uses the corresponding data.

[0070] Data606 represents the ID recognized by the operation as the ID of the corresponding data (specifically, the ID linked to Source507 or Target508 in Figure 5). Version607 represents the version of the corresponding data.

[0071] List ID 608 represents the ID of the user list 33 or operation list 34 associated with the corresponding data.

[0072] Usage 609 indicates whether the application that processes the corresponding data is permitted to use it. "Approve" means that use is permitted. "Disapprove" means that use is prohibited.

[0073] Reproduction 610 indicates whether the corresponding data can be reproduced. "Possible" means that reproduction is possible. "Impossible" means that reproduction is not possible.

[0074] An operation log 32 is generated (output) using data log 31 as input, and a data log 31 is generated (output) using operation log 32 as input. As shown in Figures 5 and 6, for example, the following occurs: That is, using data log 31 of Log ID 501 "DID-001" as input, an operation log 32 of Log ID 501 "OID-001" is generated (for example, Data 606 "Raw-001" is passed on to Source 507). Next, using the operation log 32 of Log ID 501 "OID-001" as input, a data log 31 of Log ID 501 "DID-002" is generated (for example, "PDID-001" in Target 508 is passed on to Data 606 and reflected in Type 604).

[0075] Based on the examples shown in Figures 5 and 6, the following can be said (details will be provided later). • A user list 33 or operation list 34 is managed for each operation and data. This makes it possible to track which user's data was used / not used for each operation and data. By tracing the input / output relationship between the operation log 32 and the data log 31, a data lineage can be constructed. This makes it possible to understand when and how user data was used. • Availability and reproducibility are managed. This allows for the management of processes and data whose reproducibility is lost due to changes in access permissions.

[0076] Figure 7 shows multiple user lists 33.

[0077] Figure 7 shows multiple rows, but each row represents one user list 33. A user list 33 contains information such as List ID 701, Log ID 702, Location 703, Version 704, Exclusion Users 705, and Inclusion Users 706. We will explain these pieces of information 701-706 using one user list 33 as an example.

[0078] List ID 701 represents the ID of user list 33. Log ID 702 represents the ID of the corresponding log (operation log 32 or data log 31 associated with user list 33). Location 703 is the same information as Location 502 or 601 in the corresponding log. Version 704 represents the version of user list 33.

[0079] Exclusion Users705 is a list of users (e.g., a list of usernames) who are excluded (not included) in the corresponding log for the corresponding data or operation. Inclusion Users706 is a list of users who are included in the corresponding log for the corresponding data or operation. As shown in the example in Figure 7, this is as follows: • Alice: A user who refused to provide data for machine learning / federal learning (see the row with “OID-001” in Figures 7 and 5). • Bob: A user who refused to provide data to the application (AID-123) (see the row with “OID-002” in Figures 7 and 5). • Ellen: A user who prohibited the use of data for training feature data (see the row with “DID-013” in Figures 7 and 6). • Frank: A user who has prohibited the use of data for training and validation features (see the rows with “DID-013” and “DID-014” in Figures 7 and 6). • Carol: A user who visits multiple hospitals (in Figure 7, "Carol" is listed in both the row with Location703 "US-1" and "US-2").

[0080] Figure 8 shows multiple operation lists 34.

[0081] Figure 8 shows multiple rows, but each row represents one operation list 34. An operation list 34 contains information such as List ID 801, Log ID 802, Location 803, Version 804, Exclusion Operations 805, and Inclusion Operations 806. We will explain these pieces of information 801-806 using one operation list 34 as an example.

[0082] List ID 801 represents the ID of operation list 34. Log ID 802 represents the ID of the corresponding log (operation log 32 or data log 31 to which operation list 34 is associated). Location 803 is the same information as Location 502 or 601 in the corresponding log. Version 804 represents the version of operation list 34.

[0083] Exclusion Operations 805 is a list of operations or data (e.g., a list of OIDs or DIDs) that are excluded (not included) in the corresponding log for the corresponding data or operation. Inclusion Operations 806 is a list of operations or data that are included in the corresponding log for the corresponding data or operation.

[0084] Furthermore, by having an operation list 34 instead of a user list 33, the server computer 110 can manage operation logs 32 and data logs 31 for each client computer 100 when there are multiple client computers 100, and it is also possible to prevent the server computer 110 from being informed of the user's name.

[0085] Furthermore, operation list 34 is used in the server-side processing shown in Figure 22. For example, if the reproducibility of a certain preprocessing is lost, the learning / associative learning process that uses the features of that preprocessing will also lose its reproducibility. If the reproducibility of "OID-005" in Figure 5 is lost, it is necessary to propagate the inability to reproduce to "OID-027" in Figure 5 and "MID-027" in Figure 6. Operation list 34 in Figure 8 is used to achieve this propagation.

[0086] Figure 9 shows multiple access policies.

[0087] The access policy DS133 has an access policy 910 for each user. Each access policy 910 has a column 901 for each operation attribute. Each column 901 represents access rights (e.g., "allowed" or "denied") for each data type. Each user's access policy 910 is reflected in the user list 33 and the operation list 34.

[0088] For example, for Alice, there is only one operation attribute: "Retention". Therefore, Alice's access policy 910-1 consists of a single column 901-1. According to access policy 910-1, Alice's data may be included in the privacy data, but it is prohibited for Alice's data to be used as provided data, feature data, and model data. In other words, Alice is a user who has refused to provide data for machine learning / federal learning, and for example, "Alice" is set as Exclusion Users 705 in user list 33 of "ULID-001".

[0089] Furthermore, for example, Ellen has three operation attributes: "Reference," "Create," and "Retain." Therefore, Ellen's access policy 910-2 consists of three columns: 901-2a, 901-2b, and 901-2c. According to access policy 910-2, Ellen's data is prohibited from being used as training feature data (and training model data). In other words, Ellen is a user whose data is prohibited from being used as training feature data, and for example, "Ellen" is set as Exclusion Users 705 in user list 33 of "ULID-013."

[0090] Furthermore, for Frank, for example, the operation attributes are "reference," "create," and "retain." Therefore, Frank's access policy 910-3 consists of three columns: 901-3a, 901-3b, and 901-3c. According to column 901-3a, use of training and validation feature data is prohibited, but use of training and validation model data is permitted. This means that models previously generated using feature data containing Frank's data are permitted to be used.

[0091] The following examples, using Ellen as an example, show access policy 910-2, usage data, lineage data, and contribution data before and after the change in access permissions.

[0092] Figure 10A shows access policy 910-2 before the access rights were changed. Figure 11A shows usage data before the access rights were changed. Figure 12A shows the lineage represented by the lineage data before the access rights were changed. Figure 13A shows an example of contribution data before the access rights were changed. The data shown in Figures 10A, 11A, 12A, and 13A corresponds to Ellen.

[0093] The example usage data in Figure 11A shows some of the usage data in the DS134C. The usage data includes information such as ID1101, Date1102, Application / Model ID1103, Start Date1104, End Date1105, Available Models (Total1107, Using Training Data1108, Using Validation Data1109, and Using Evaluation Data1110), Unavailable Models (Total1111, Using Training Data1112, Using Validation Data1113, and Using Evaluation Data1114), and Usage Counts1115.

[0094] ID1101 represents the ID of the data being used. Date1102 represents the date and time the data was acquired. Application / Model ID1103 represents the ID of the application or model corresponding to the data being used. Start Date1104 represents the start date and time of the application or model corresponding to the data being used. End Date1105 represents the end date and time of the application or model corresponding to the data being used.

[0095] Regarding Available Models, Total 1107 represents the total number of available models, while Information 1108-1110 shows the breakdown. Specifically, Using Training Data 1108 represents the number of available training models, Using Validation Data 1109 represents the number of available validation models, and Using Evaluation Data 1110 represents the number of available evaluation models. An application may have multiple models (versions), and when used for inference, the model with the highest accuracy is used. Specifically, for example, if the application is "cancer risk prediction," multiple models may be provided for a single application, such as a prediction model prepared in one month and a prediction model prepared in the following month.

[0096] Regarding Unavailable Models, Total1111 represents the total number of unavailable models, while Information1112-1114 shows the breakdown. Specifically, Using Training Data1112 represents the number of unavailable training models, Using Validation Data1113 represents the number of unavailable validation models, and Using Evaluation Data1114 represents the number of unavailable evaluation models.

[0097] Usage Counts1115 represents the number of times a model corresponding to the usage data is used, or the number of times a model belonging to the application corresponding to the usage data is used.

[0098] The lineage illustrated in Figure 12A has a DAG structure with DID (Data Log 31) or OID (Operation Log 32) as nodes. For example, by tracing the operation log 32 and data log 31 from the data log 31 with a DID identified by AID as the key, it is possible to identify the model used among the models belonging to the application corresponding to that AID. Note that the AID, DID, OID, and MID shown in Figure 12A (and Figure 12B) correspond to the AID, DID, OID, and MID in Figures 5 and 6. For example, the flow from AID(123) to MID(027) is a flow related to federative learning. Also, the flow from AID(123) to MID(101) means that the validation feature data corresponding to "DID-014" was reused. In this way, it is possible to see how the application corresponding to AID(123) used the data (data such as Ellen's data).

[0099] The contribution data illustrated in Figure 13A is a portion of the contribution data in contribution DS140C. The contribution data includes information such as ID1301, AID1302, MID1303, contribution level (information amount 1304 and general information amount), and metadata (number of available data 1306 and total number of data 1307).

[0100] ID1301 represents the ID of the contributed data. AID1302 represents the ID of the application corresponding to the contributed data. MID1303 represents the ID of the model corresponding to the contributed data.

[0101] The information quantity of 1304 represents the information quantity of the data involved in the model corresponding to the contributing data, and is based on the number of available data points (1306) and the total number of data points (1307).

[0102] The number of available data points, 1306, represents the number of data points for the target user (e.g., "Ellen") of the application corresponding to the contributing data. For example, "100" means that there are 100 data points for the target user (100 data points out of the data provided to the application). In federated learning processing, "number of available data points" refers to the number of target user data points included in the privacy DS201 of all 100 client computers.

[0103] The total number of data points, 1307, represents the total number of data points in Privacy DS201. For example, "100000" means that Privacy DS201 contains 100,000 data points. In federated learning processing, "total number of data points" refers to the number of data points contained in Privacy DS201 across all 100 client computers.

[0104] The information quantity of 1304 is, for example, -log2(number of available data 1306 / total number of data 1307). Therefore, in the case illustrated in Figure 13A, the information quantity of 1304 is "9.97".

[0105] The general-purpose information quantity 1305 is the amount of information that can be used generally for data related to the model corresponding to the contributed data, and is based on the information quantity 1304 and the setting of access permissions (e.g., the number of "permissions") for the operation attribute "create" related to the model corresponding to the contributed data. Specifically, for example, the general-purpose information quantity 1305 is the product of the information quantity 1304 and the number of "permissions" for Validation and Evaluation for the operation attribute "create" related to the model corresponding to the contributed data. Therefore, if the number of "permissions" is 2, the general-purpose information quantity 1305 is "19.94", as illustrated in Figure 13A.

[0106] Thus, the degree of contribution represented by the contribution data is based on the amount of information in the data related to the model corresponding to the contribution data and the access rights related to the model.

[0107] Figure 10B shows the access policy 910-2 after the access permission change. Figure 11B shows the usage data after the access permission change. Figure 12B shows the lineage data after the access permission change. Figure 13B shows an example of contribution data after the access permission change.

[0108] As shown in Figure 10B, assume that, for each of the operation attributes "Read" and "Create" in the access rights represented by access policy 910-2, the use of Ellen's data for validation feature data has been changed from "Allowed" to "Forbidden," and the use of Ellen's data for validation models has been changed from "Allowed" to "Forbidden."

[0109] As a result of this access permission change, as shown in Figure 11B, both "AID-123" and "MID-001" will change from available models to unavailable models. Specifically, for "AID-123", the Using Validation Data 1109 in Available Models will change from "3" to "0", and the Using Validation Data 1113 in Unavailable Models will change from "0" to "3". Similarly, for "MID-001", the Using Validation Data 1109 in Available Models will change from "1" to "0", and the Using Validation Data 1113 in Unavailable Models will change from "0" to "1".

[0110] As a result of this access permission change, as shown in Figure 12B, the data corresponding to “DID-014” (feature data for validation corresponding to the validation model that was deemed unavailable), the data corresponding to “DID-116” (feature data for validation corresponding to the validation model that was deemed unavailable), “DID-016” (data for the validation model that was deemed unavailable), and “DID-037” (data for the model in which the validation models that were deemed unavailable were merged through federative learning) will become unavailable. Specifically, for example, the usage 609 in data log 31 corresponding to this data will be changed to “Disapprove”.

[0111] Furthermore, as a result of this access permission change, the access permission for Validation becomes "prohibited" and the access permission for Evaluation becomes "permitted," so the number of "permitted" items becomes 1. Therefore, the general information quantity 1305 is the same as information quantity 1304, "9.97".

[0112] Figure 14 shows the functions of the client computer 100 and the server computer 110.

[0113] The User Management Unit 128, Lineage Management Unit 129, Access Management Unit 130, and Contribution Management Unit 131 are present in both the client computer 100 and the server computer 110. Hereinafter, for each of these functions 129 to 131, the function in the client computer 100 will be referred to as "client," and the function in the server computer 110 will be referred to as "server."

[0114] The client usage management unit 128C and the server usage management unit 128S communicate with each other. The client lineage management unit 129C and the server lineage management unit 129S communicate with each other. The client access management unit 130C and the server access management unit 130S communicate with each other. The client contribution management unit 131C and the server contribution management unit 131S communicate with each other.

[0115] The user terminal 101 can send a request to either the server computer 110 or the client computer 100. For example, if the target is a specific location, the user terminal 101 may send a request to the client computer 100 for that location. On the other hand, if the target is all locations, the user terminal 101 may send a request to the server computer 110.

[0116] The following describes some of the processes performed in this embodiment.

[0117] Figure 15 shows the flow of the user management process.

[0118] The client usage management unit 128C receives a usage management request from the user terminal 101, either via the server usage management unit 128S or otherwise, and in response to the request, performs the processing from S1501 onwards.

[0119] In S1501, the client usage management unit 128C searches for a target from the user list DS132 and obtains a target ID, which is the OID and / or DID of the found target. In this paragraph, "target" may be any element that satisfies the conditions specified in the request from the user terminal 101. For example, if "Ellen" is found as the "target", the target ID is "OID-004" (see Figure 7). The scope of the target search is Inclusion Users 705, but Exclusion Users 706 may be used instead of Inclusion Users 705 as the scope of the target search. This may also be the case in S1601 and / or S2001, instead of or in addition to S1501.

[0120] In S1502, the client usage management unit 128C uses the target log to determine whether a model has been generated. In this paragraph, the "target log" refers to the log (operation log 32 and / or data log 31) corresponding to the target ID obtained in S1501. In S1502, for example, by searching multiple operation logs 32 of log DS142C using "OID-004" as the key, it is possible to determine that a model for "MID-001" has been generated (see Figure 5).

[0121] In S1503, the client usage management unit 128C determines whether the model is available using the target log (based on usage 510 or 609 of the target log). The client usage management unit 128C also identifies how the provided data was processed to generate the model, and determines whether the model is available for each data processing step (for example, for Training, Validation, and Evaluation). For example, the client usage management unit 128C identifies "OID-003" to "OID-005" corresponding to "MID-001", and determines whether they are available based on usage 510 of the operation log 32 for each of "OID-003" to "OID-005".

[0122] In S1504, the client usage management unit 128C uses the target log to determine the number of times the model has been used. For example, the client usage management unit 128C may determine the number of operation logs 32 with Type 505 "Inference" corresponding to "MID-001" as the number of inferences, or determine the number of times the model has been provided, and determine the number of inferences, the number of provision, or the sum thereof as the number of uses.

[0123] When S1504 is completed, the usage data exemplified in Figure 11A is complete for the model (or application to which that model belongs) identified in S1502. Specifically, Date 1102 may be the date and time this usage management process was performed, Application / Model ID 1103 may be the ID of that model (or application), and Start Date 1104 and End Date 1105 may be set from Date 504 or 603 of the log for that model (or application) for the target found in S1501. Information 1107 to 1114 may be set in S1503. Usage Counts 1115 may be the number of uses identified in S1504.

[0124] In S1505, the client usage management unit 128C returns the result (completed usage data) to the user terminal 101 that requested the usage management, either via or without the server usage management unit 128S. The client usage management unit 128C also includes the result (completed usage data) in the usage DS 134C.

[0125] Figure 16 shows the flow of the lineage management process.

[0126] The client lineage management unit 129C receives a lineage management request from the user terminal 101, either via the server lineage management unit 129S or otherwise, and in response to the request, performs the processing from S1601 onwards.

[0127] In S1601, the client lineage management unit 129C searches for the target from the user list DS132 and obtains the target ID, which is the OID and / or DID of the found target. In this paragraph, "target" may be an element that satisfies the conditions specified in the request from the user terminal 101 (for example, "Ellen"). Note that since the target ID of a target may exist in multiple user lists 33, multiple target IDs may be obtained for a target. In this case, S1602 and subsequent steps are performed for each of the multiple target IDs.

[0128] In S1602, the client lineage management unit 129C uses the target log to obtain a first path (usage history) from the target ID found in S1601 to the generation of the model, and a second path from the privacy data to the generation of the target ID. In other words, a bidirectional search (pathfinding) is performed starting from the target ID. The first path is the path from the target ID to the downstream side, and the second path is the path from the upstream side to the target ID. In this paragraph, "target log" refers to the log (operation log 32 and / or data log 31) corresponding to the target ID obtained in S1601. For example, in Figure 12A, if the target ID is "OID-002", the first path from the target ID to "MID-101", "MID-001", and "MID-027", and the second path from the privacy data ("AID-123") to "OID-002" are identified. The combination of the first and second paths constitutes the lineage.

[0129] Identifying the first route is done, for example, as follows: (a1) The client lineage management unit 129C identifies the DID of the data log 31 that has the ID in Target 508 of the operation log 32 which has an OID as Data 606. (If the target ID is an OID, (a1) is the start of route identification.) (a2) The client lineage management unit 129C determines whether the Type 604 of the data log 31 with a DID is "Model" (an example of the first predetermined value). (If the target ID is a DID, (a2) is the start of the identification of the first route. If (a2) is a process after (a1), the DID is the DID identified in (a1).) If the result of (a3)(a2) is false, the client lineage management unit 129C identifies the OID of the operation log 32 that has Source 507 containing the DID in (a2). Then the process returns to (a1).

[0130] Furthermore, the identification of the second route can be done, for example, as follows: (b1) The client lineage management unit 129C identifies the DID of the data log 31 whose ID in Source 507 of the operation log 32, which has an OID, is set as Data 606. (If the target ID is an OID, (b1) is the start of identifying the second route.) (b2) The client lineage management unit 129C determines whether the Type 604 of the data log 31 with a DID is "Privacy" (an example of a second predetermined value). (If the target ID is a DID, (b2) is the start of the identification of the second route. If (b2) is a process that follows (b1), the DID is the DID identified in (b1).) If the result of (b3)(b2) is false, the client lineage management unit 129C identifies the OID of the operation log 32 that has Source 507 containing the DID in (b2). Then the process returns to (b1).

[0131] In S1603, the client lineage management unit 129C returns the result (lineage data representing the lineage identified in S1602) to the user terminal 101 that requested lineage management, either via or without the server lineage management unit 129S. The client lineage management unit 129C also includes the result (lineage data) in the lineage DS 137C.

[0132] Figure 17 shows the data processing flow.

[0133] In S1701, the data preprocessing unit 126 references the application DS136C, the access policy DS133C, and the provision DS135 through the machine learning control unit 123.

[0134] In S1702, the data preprocessing unit 126 allocates user data in the provided data corresponding to the application identified from application DS136C, in accordance with the access policy DS133C. Specifically, for example, it is as follows: (S1702-1) The data preprocessing unit 126 obtains user data from the provided data. (S1702-2) The data preprocessing unit 126 refers to the access policy related to the user corresponding to the data acquired in S1702-1 and identifies the access rights corresponding to “Feature”. (S1702-3) The data preprocessing unit 126 obtains feature types (e.g., Training, Validation, or Evaluation) for which all of the operation attributes "reference", "create", and "retain" are set to "allow". In other words, if even one of the operation attributes "reference", "create", and "retain" is set to "prohibited", the user's data cannot be used to create a model, so the data preprocessing unit 126 performs the following S1702-4. (S1702-4) If there is no corresponding feature type, the data cannot be used as feature data, so the data preprocessing unit 126 sets the name of that user in the Exclude Users 705 of the user list 33 associated with the operation log 32. (S1702-5) If there is a corresponding feature type, the data preprocessing unit 126 assigns the data to one of the corresponding feature types and sets the name of the user in the Inclusion Users 706 of the user list 33 associated with the assigned data log 31 (the data log 31 corresponding to the feature data of that feature type).

[0135] In S1702, steps S1702-1 to S1702-5 are performed on the data of all users in the provided data.

[0136] In S1703, the data preprocessing unit 126 generates feature data based on the data of users identified from the Inclusion Users 706 in the user list 33 of the assigned data log 31, and includes this feature data in the feature DS138.

[0137] Figure 18 shows the flow of the learning process.

[0138] In S1801, the learning unit 125 accesses the application DS136C through the machine learning control unit 123.

[0139] In S1802, the learning unit 125 refers to multiple data logs 31.

[0140] In S1803, the learning unit 125 determines whether the target data log 31 exists among the multiple data logs 31. The target data log 31 corresponds to the application identified from application DS136C, and is a data log 31 where Type 604 is "Feature" and Usage 510 is "Approved". The determination in S1803 is to determine whether usable feature data exists.

[0141] If the result of S1803 is true (S1803:Yes), then in S1804, the learning unit 125 performs model training (Training, Validation, or Evaluation) using the feature data.

[0142] If the result of the S1803 determination is true (S1803: No), the data preprocessing unit 126 performs the processing shown in Figure 17. After that, the process returns to S1802.

[0143] Figure 19 shows the flow of the inference process.

[0144] In S1901, the inference unit 127 references the application DS136C through the machine learning control unit 123.

[0145] In S1902, the inference unit 127 refers to multiple data logs 31.

[0146] In S1903, the inference unit 127 determines whether the target data log 31 exists among the multiple data logs 31. The target data log 31 corresponds to the application identified from application DS136C, and is a data log 31 whose Type 604 is "Model" and whose Usage 510 is "Approve". The determination in S1903 is to determine whether or not an available model exists.

[0147] If the result of the determination in S1903 is true (S1903:Yes), then in S1904, the inference unit 127 performs inference using that model.

[0148] If the result of the judgment in S1903 is true (S1903: No), in S1905, the learning unit 125 performs the process shown in Figure 18. After that, the process returns to S1902.

[0149] Figure 20 shows the flow of the contribution management process.

[0150] The client contribution management unit 131C receives a contribution management request from the user terminal 101, either via the server contribution management unit 131S or otherwise, and in response to the request, performs processing from S2001 onwards.

[0151] In S2001, the client contribution management unit 131C searches for the target from the user list DS132 and obtains the target ID, which is the OID and / or DID of the found target. In this paragraph, "target" may be any element that satisfies the conditions specified in the request from the user terminal 101.

[0152] In S2002, the client contribution management unit 131C uses the target log to identify the AID and MID corresponding to the target ID. The "target log" refers to the log (operation log 32 and / or data log 31) corresponding to the target ID obtained in S2001.

[0153] In S2003, the Client Contribution Management Unit 131C calculates the number of data points used to create models corresponding to the AID and MID identified in S2002. The calculated number of data points is the sum of the number of users represented by Inclusion Users 706 in the user list 33 corresponding to the data log 31 having Type 604 "Feature" and the number of users represented by Inclusion Users 706 in the user list 33 corresponding to the operation log 32 having Type 505 "Learning".

[0154] In S2004, the Client Contribution Management Unit 131C identifies the total number of data points related to the provided data (how many user data points are included in the provided data) using the target ID as the key.

[0155] In S2005, the client contribution management unit 131C calculates the amount of information (= -log2(number of available data / total number of data)). The "number of available data" is the number of data calculated in S2003. The "total number of data" is the number of data identified in S2004.

[0156] In S2006, the client contribution management unit 131C calculates the general information amount based on the access policy (for example, the access policy of the user identified in S2001) and the amount of information calculated in S2005.

[0157] When S2006 is completed, the example contribution data shown in Figure 13A is complete. That is, AID1302 and MID1303 can be the AID and MID identified in S2002, information quantity 1304 can be the information quantity calculated in S2005, general information quantity 1305 can be the general information quantity calculated in S2006, number of available data 1306 can be the number of data calculated in S2003, and total number of data 1307 can be the number of data identified in S2004.

[0158] In S2007, the client contribution management unit 131C returns the results (completed contribution data) to the user terminal 101 that requested contribution management, either via or without the server contribution management unit 131S. The client contribution management unit 131C also includes the results (contribution data) in the contribution DS 140C.

[0159] Figure 21 shows the flow of the first access management process.

[0160] The client access management unit 130C receives a first access management request from the user terminal 101, either via or without the server access management unit 130S, and performs processing from S2101 onwards in response to the request. The first access management request is a request for an impact investigation due to a change in access rights (an impact investigation assuming that the access rights change specified in this request has been made). In the first access management request, for example, the user corresponding to the access policy 910 whose access rights will be changed, and how the access rights corresponding to each data type and operation type will be changed (for example, the access policy 910 after the access rights change) may be specified. In the explanation of Figure 21 (and Figure 22), "before change" means before the access rights in the access policy have been changed, and "after change" means after the access rights in the access policy have been changed.

[0161] In S2101, the client access management unit 130C obtains the usage data before the change from the usage DS 134C (or client usage management unit 128C). The client access management unit 130C also obtains the lineage data before the change from the lineage DS 137C (or client lineage management unit 129C). The client access management unit 130C also obtains the contribution data before the change from the contribution DS 140C (or client contribution management unit 131C). The data before the change obtained in S2101 may be the data in DS 134C, 137C, or 140C at the start of the first access management process.

[0162] In S2102, the client access management unit 130C generates modified usage data based on the modified usage data obtained in S2101. Specifically, for example, the client access management unit 130C generates modified usage data based on the data type ("Privacy", "Provided", "Feature (Training)", "Feature (Validation)", "Feature (Evaluation)", "Model (Training)", "Model (Validation)", or "Model (Evaluation)") and operation type ("Reference", "Create", or "Retain") of the access policy 910 of the authorized user (the user represented by the user list 33 (Inclusion Users 706) associated with log 32 and / or 31 corresponding to the Application / Model ID 1103 of the usage data), which has been changed from "Allowed" to "Denyed", and the operation type ("Reference", "Create", or "Retain"), and then generates the Available Models The value to subtract from the information inside 1107~1110, and Unvailable ModelsThe values ​​to be added to the information 1111-1114 inside are determined. Also, for example, the client access management unit 130C determines the access policy 910 of a prohibited user (a user represented by the user list 33 (Inclusion Users 706) associated with log 32 and / or 31 corresponding to Application / Model ID 1103 of the usage data) based on the data type and operation type of the access rights that have been changed from "prohibited" to "allowed". Models The value to add to the information inside 1107~1110, and Unvailable Models Determine the value to subtract from the information items 1111-1114. Note that both the permitted and prohibited users can be a single user.

[0163] Furthermore, the client access management unit 130C generates modified lineage data based on the pre-change lineage data obtained by S2101. Specifically, for example, for the access policy 910 of an authorized user, the DID corresponding to the permitted data is set to the DID corresponding to the prohibited data, based on the data type and operation type of the access rights that have been changed from "allowed" to "forbidden". Also, for example, for the access policy 910 of a prohibited user, the client access management unit 130C sets to the DID corresponding to the prohibited data to the DID corresponding to the permitted data, based on the data type and operation type of the access rights that have been changed from "forbidden" to "allowed".

[0164] Furthermore, the client access management unit 130C generates post-change contribution data based on the pre-change contribution data acquired by S2101. Specifically, for example, for the access policy 910 of permitted users, it modifies the general information quantity 1305 based on the data type and operation type of the access rights that have been changed from "allowed" to "forbidden". Also, for example, the client access management unit 130C modifies the general information quantity 1305 for the access policy 910 of prohibited users based on the data type and operation type of the access rights that have been changed from "forbidden" to "allowed".

[0165] In S2103, the client access management unit 130C returns the results (at least the latter of the usage data, lineage data, and contribution data before the change and the usage data, lineage data, and contribution data after the change) to the user terminal 101 that made the first access management request, either via or without the server access management unit 130S.

[0166] Figure 22 shows the flow of the second access management process.

[0167] The client access management unit 130C receives a second access management request from the user terminal 101, either via or without the server access management unit 130S, and performs the processing from S2201 onwards in response to the request. The second access management request is a request to change access rights. The second access management request may specify, for example, the user corresponding to the access policy 910 whose access rights will be changed, and how the access rights corresponding to each data type and operation type will be changed (for example, the access policy 910 after the access rights change).

[0168] In S2201, the client access management unit 130C performs S2101 and S2102 in Figure 21, and returns the results and a query (a query on whether or not to implement the access permission change as requested) to the second access management requesting user terminal 101, either via or without the server access management unit 130S.

[0169] In S2202, the client access management unit 130C determines whether the response to the inquiry (the response from the user) is a response indicating a change in access rights. Note that in S2201, there does not need to be an inquiry to the user terminal 101. In that case, the determination in S2202 may be a determination of whether the result obtained in S2201 satisfies predetermined conditions. These conditions may differ for each user, and the predetermined conditions compared with the result in S2202 may be conditions corresponding to the user specified in the second access management request. Furthermore, the predetermined conditions may include at least one of the following conditions: conditions relating to information among usage data, lineage data, and contribution data that may be affected by the change in access rights, for example, conditions relating to the information of Available Models 1107~1110, conditions relating to the information of Unavailable Models 1111~1114, conditions relating to nodes in lineage (e.g., DID) (e.g., conditions relating to the decrease or increase in the number of DIDs), and conditions relating to the amount of general-purpose information (e.g., conditions relating to the increase or decrease in the amount of general-purpose information). Furthermore, the specified conditions may include conditions regarding the reproducibility of operations and data (e.g., feature data and model data).

[0170] If the result of S2202 is false (S2202: No), in S2208, the client access management unit 130C returns the result (that access rights will not be changed) to the user terminal 101 that requested the second access management, either via or without the server access management unit 130S.

[0171] If the result of S2202 is true (S2202:Yes), then in S2203, the client access management unit 130C determines whether the impact of the change in access rights falls under the promotion conditions. Specifically, for example, the promotion DS143C includes promotion condition data for each element such as user and application. The promotion condition data represents the element's ID, the content of the promotion, and the promotion trigger (the conditions under which the promotion is presented to the user (in other words, the conditions related to the impact of the change in access rights)). According to the promotion condition data exemplified in Figure 22, if the number of data points required to provide the application "AID-123" falls below 80, it means that a promotion consisting of the product of general contribution points and $1 will be presented to the user whose access rights are changed. In S2203, for example, the client access management unit 130C identifies promotion condition data from the promotion DS 143C using "AID-123" (or the ID of the user whose access rights are being changed (e.g., name)) corresponding to the first node (DID) of the graph represented by the modified lineage data as the key. The client access management unit 130C determines whether the number of data identified in S2201 (the number of data required to provide the application for "AID-123") corresponds to a promotion trigger. If the result of the determination in S2203 is false (S2203: No), the process proceeds to S2206.

[0172] If the result of the S2203 judgment is true (S2203: Yes), in S2204, the client access management unit 130C presents a promotion to the user (user terminal 101). The promotion is displayed on the user terminal 101. In the presentation in S2204, the client access management unit 130C informs the user that they can receive the presented promotion on the condition that they either relax their access rights (modify the desired access rights change so that the amount of data available to the user increases) or cancel the desired access rights change.

[0173] In S2205, the client access management unit 130C determines whether the response to the promotion offer (the user's response) is a change in access rights as desired (as requested). If the result of the determination in S2205 is false (S2205: No), in S2208, the client access management unit 130C returns the result (the result that access rights will not be changed) to the user terminal 101 that made the second access management request, either via or without the server access management unit 130S. In the case of relaxation of access rights, a second access management request for a change in access rights representing the relaxed access rights is sent from the user terminal 101, and the processing shown in Figure 22 may be performed on that second access management request.

[0174] If the result of the S2205 judgment is true (S2205:Yes), in S2206, the client access management unit 130C updates the use 510 and / or reproduction 511 of the operation log 32, and / or the use 609 and / or reproduction 610 of the data log 31, in accordance with the change in access rights. Specifically, for example, the use 609 "Disapprove" and reproduction 610 "Impossible" of the data log 31 depend on the access rights "prohibited" for the operation attribute "read", "create", or "keep". Also, the reproduction 511 "Possible" of the operation log 32 may be the case where the use 609 "Approve" of the data log 31 corresponding to the Source 507 of the operation log 32 is matched with the reproduction 610 "Possible" of the data log 31 corresponding to the Target 508 of the operation log 32. Operation log 32, usage 510 "Disapprove," indicates that the application has become unavailable, and access permissions do not affect whether the application is available or not.

[0175] In S2207, the client access management unit 130C updates at least one DS from among the provided DS135, feature DS138, and model DS141C, etc., in response to a change in access rights. For example, if the access rights change is to "prohibit" the operation attribute "retention", the client access management unit 130C deletes the data of the user to whom such access rights have been changed, or data that uses that data, from at least one DS such as the provided DS135, feature DS138, and model DS141C. If the user data can be identified, the client access management unit 130C creates or duplicates a data log 31 corresponding to the data containing the user data to be deleted, records the name of the user whose access rights have been changed in the Exclusion Users 705 of the user list 33 associated with the data log 31, and updates the Version 704. If the data that uses the user data to be deleted is feature data, the client access management unit 130C deletes the feature data itself because it is difficult to identify the user data.

[0176] In S2208, following S2207, the client access management unit 130C returns the result to the user terminal 101 that requested the second access management, either via or without the server access management unit 130S.

[0177] In the flow shown in Figure 22, for example, steps S2203 to S2205 may be omitted.

[0178] According to the embodiment described above, when data becomes unavailable due to changes in access permissions (changes from "allowed" to "prohibited") such as opt-out requests or withdrawal from federated learning, it is possible to identify and display models that can or cannot be reproduced or used. It is also possible to identify and display lineage as the usage path of private data. Furthermore, in order to provide appropriate changes in access permissions such as opt-out requests, it is possible to set granular usage scopes for private data (for example, the method of distributing between Training, Validation, and Evaluation).

[0179] For example, in S2201, the client access management unit 130C may check whether the values ​​of reproduction 511 and 610 change based on the pre-change data and the post-change data, and present the results of that check to the user. For example, the user can decide whether or not to exercise the access rights, such as "If reproduction is not possible, do not change the access rights" or "If the range in which reproduction is not possible is within the range presented, change the access rights." This decision result may be the response (answer to the inquiry) that the client access management unit 130C receives.

[0180] Although one embodiment has been described above, this is merely an example for the purpose of explaining the present invention and is not intended to limit the scope of the present invention to this embodiment alone. For example, anonymized privacy data may be used instead of privacy data (for example, a correspondence table or processing (such as a hash function) between the original privacy data and the anonymized data may be used).

[0181] The above explanation can be summarized, for example, as follows. This summary may include supplementary explanations and descriptions of variations of the above explanation.

[0182] A storage device (e.g., storage device 120) and a processor (e.g., processor 160) are provided. The storage device stores an access policy for each entity (e.g., access policy DS133 (e.g., 133C)). For each entity, the access policy includes, for each of one or more operational attributes, access rights for each nth-order data for use in the application or model for that entity (where n is a non-negative integer). In the above embodiment, the model is a model generated by machine learning, but it may also be a model generated by a method other than machine learning (e.g., statistical information generated by a statistical method). Here, "statistical information" may be a model as a program or data. Also, "nth-order data" may be called a data type, and the value of "n" may be incremented each time the data is used (e.g., including processing), such as 0th-order data (e.g., raw data such as privacy data), 1st-order data (e.g., provided data), 2nd-order data (e.g., feature data), 3rd-order data (e.g., model data), etc.

[0183] The storage device stores, for each operation, an operation log (e.g., operation log 32) which is a log of the operation, a data log corresponding to the source data which is the data for the operation and is associated with the operation log (e.g., data log 31 corresponding to Source 507), and / or a data log corresponding to the target data which is the data resulting from the operation and is associated with the operation log (e.g., data log 31 corresponding to Target 508).

[0184] The storage device stores entity lists (e.g., user list 33) based on access policies for operation logs and / or data logs. Each entity list includes an inclusion list (e.g., Inclusion Users 706), which is a list of entities whose data is permitted to be used for operations or data corresponding to the log to which the entity list is associated (e.g., log 32 or 31), and / or an exclusion list (e.g., Exclusion Users 705), which is a list of entities whose data is prohibited from being used for operations or data corresponding to the log to which the entity list is associated (e.g., log 32 or 31). Here, “entity” is a user in the embodiments described above, but may also be an entity other than a user (e.g., rights such as copyright). The entity’s data may be data that is undesirable to be used without the permission of a specific person (e.g., a user, rights holder, agent), such as privacy data or security data, and the entity may be that specific person. Furthermore, the entity is not limited to a human being, but may be any tangible or intangible object.

[0185] In response to a request, if the processor finds one or more entity lists containing the entities identified based on the request among multiple entity lists, it identifies usage based on one or more operation logs and one or more data logs identified using those entity lists, and returns data representing the identified usage to the requester.

[0186] This allows for transparent data management while maintaining data access permissions.

[0187] Furthermore, operation logs, data logs, and entity lists may all be generated and stored by the processor. For example, each time an operation involving data input / output is performed, the processor may identify the data log corresponding to the input data (source data), generate an operation log for that operation based on the identified data log, and generate a data log corresponding to the output data (target data) (a data log associated with the operation log), and store these logs in a storage device. The processor may also perform operations based on the access policies of each entity, and for this purpose, it may input (use) data from entities that have permission to access the input data (e.g., m-th order data), and output data as a result of the operation. The processor may generate an entity list for operation logs and output data logs, including entities whose data was used in an inclusion list (or entities whose data was not used in an exclusion list), and associate the generated entity list with the operation log and data log. Entity lists may exist for each operation log and for each data log.

[0188] "Usage status" can be a lineage identified based on one or more identified operation logs and one or more data logs. The lineage can be a Directed Acyclic Graph (DAG) where data or operations are nodes and models correspond to intermediate nodes or leaf nodes. This allows anyone who receives usage status information from a data management system to understand how the data was used.

[0189] "Usage status" may include the contribution of a specified entity. "Contribution status" may include a value calculated based on the total number of data points and the number of available data points (e.g., information quantity 1304 and / or general information quantity 1305). The total number of data points (e.g., total number of data points 1307) may be the number of data points in the (mk)th order data that can be used to generate the model as mth-order data (both m and k are integers less than or equal to the maximum value of n, and m is greater than k). The number of available data points (e.g., number of available data points 1306) may be the number of data points corresponding to entities for which access rights up to the mth-order data are permitted in the access policy, out of the total number of data points. This allows anyone who receives usage status information from the data management system to understand how much an entity's data contributes to the model (or the application to which the model belongs). Furthermore, the access policy may include access permissions for each of the multiple types of models as m-th order data, and the calculated value (for example, the general information quantity of 1305 included in the contribution score) may be based on the total number of data points, the number of available data points, and the number of predetermined types of models for which access permissions are granted.

[0190] The access policy may include access permissions for each of multiple types of models as m-th data, and the "usage status" may include the number of models for each type of model generated using the data of a specified entity. This allows anyone receiving usage status from the data management system to understand what types of models the entity's data was used for.

[0191] "Usage status" may include usage status before and after changes to access permissions in the access policy of the identified entity. This allows anyone receiving usage status from the data management system to understand the scope of impact of changes to access permissions.

[0192] At least one of the operation log and the data log associated with the operation log may include reproducibility information, which indicates whether the data provided as a result of the operation can be reproduced. "Usage status" may include the reproducibility information expressed by the reproducibility information after the access permission change, with respect to the reproducibility information held by one or more identified operation logs and one or more data logs. This allows a person who receives usage status from the data management system to understand whether the data can be reproduced after the access permission change.

[0193] If the usage after the change in access permissions meets certain conditions (e.g., a Promotion Trigger), the processor may offer the requester a promotion to ease or cancel the access permission change. This provides the requester with an incentive to ease or cancel the access permission change, thereby promoting data usage.

[0194] The data management system may include a plurality of client computers, each having a processor and memory, and a server computer that communicates with the plurality of client computers. The server computer may generate a model through federative learning using machine learning models from the plurality of client computers and transmit the generated model to the plurality of client computers. The server computer may also store operation logs and data logs regarding the operations performed by the server computer. An operation list (e.g., operation list 34) may be associated with the operation logs and data logs in the server computer instead of an entity list. The operation list may include an inclusion list, which is a list of operations or data corresponding to entities for which data use is permitted for the operation or data corresponding to the log to which the operation list is associated, and / or an exclusion list, which is a list of operations or data corresponding to entities for which data use is prohibited for the operation or data corresponding to the log to which the operation list is associated. In this way, the server computer stores an operation list that does not contain information that can identify entities, instead of an entity list. [Explanation of symbols]

[0195] 40...Calculator

Claims

1. Equipped with a memory device and a processor, The aforementioned storage device stores access policies for each entity, For each entity, the access policy includes, for each of one or more operational attributes, access permissions per nth data point for use in the application or model for that entity (where n is a non-negative integer). The aforementioned storage device, for each operation, The operation log, which is the log of the operation in question, A data log corresponding to the source data, which is the data for the operation, and associated with the operation log, and / or a data log corresponding to the target data, which is the data resulting from the operation, and associated with the operation log. Remember this, The storage device stores an entity list based on the access policy for the operation log and / or data log. Each entity list includes an inclusion list, which is a list of entities whose data is permitted for use in the operation or data corresponding to the log to which the entity list is associated, and / or an exclusion list, which is a list of entities whose data is prohibited from being used in the operation or data corresponding to the log to which the entity list is associated. The aforementioned processor, If, in response to a request, one or more entity lists containing the entities identified based on the request are found among multiple entity lists, the usage status is determined based on one or more operation logs and one or more data logs identified using those one or more entity lists. The data representing the identified usage status is returned to the requester of the aforementioned request. Data management system.

2. The aforementioned usage status is a lineage identified based on one or more operation logs and one or more data logs identified above. The aforementioned lineage is a Directed Acyclonic Graph (DAG) in which data or operations are nodes and models correspond to intermediate nodes or leaf nodes. The data management system according to claim 1.

3. The aforementioned usage includes the contribution of the identified entity, The aforementioned contribution includes a value calculated based on the total number of data points and the number of available data points. The total number of data points mentioned above is the number of data points in the (m-k)th order data that can be used to generate a model as the mth order data (both m and k are integers less than or equal to the maximum value of n, and m is greater than k), The number of available data items is the number of data items corresponding to entities for which access rights are permitted up to the mth level of data in the access policy, out of the total number of data items. The data management system according to claim 1.

4. The aforementioned access policy includes access rights for each of several types of models as m-th order data. The calculated value is based on the total number of data points and the number of available data points, and the number of a predetermined type of model for which access rights are permitted. The data management system according to claim 3.

5. The aforementioned access policy includes access rights for each of several types of models as m-th order data. The aforementioned usage includes the number of models for each type of model generated using the data of the identified entity. The data management system according to claim 1.

6. The aforementioned usage includes the usage before and after the change in the access permissions in the access policy of the identified entity. The data management system according to claim 1.

7. At least one of the operation log and the data log associated with the operation log includes reproducibility information, which is information indicating whether the data can be reproduced. The aforementioned usage status includes, with respect to the reproducibility information contained in one or more operation logs and one or more data logs identified, the reproducibility information after the change in access rights. The data management system according to claim 1.

8. If the usage after the change in access rights meets the specified conditions, the processor will offer the requester a promotion to relax or cancel the change in access rights. The data management system according to claim 1.

9. Entity lists exist for each operation log and each data log. The data management system according to claim 1.

10. A plurality of client computers, including a client computer having the processor and the storage device, A server computer that communicates with the aforementioned multiple client computers and Equipped with, The server computer generates a model through federated learning using machine learning models from the multiple client computers and transmits the generated model to the multiple client computers. In the aforementioned server computer, operation logs and data logs are stored regarding the operations performed by the server computer. In the aforementioned server computer, the operation log and data log are associated with an operation list instead of an entity list. An operation list includes an inclusion list, which is a list of operations or data corresponding to entities for which data use is permitted for the operations or data corresponding to the log to which the operation list is associated, and / or an exclusion list, which is a list of operations or data corresponding to entities for which data use is prohibited for the operations or data corresponding to the log to which the operation list is associated. The data management system according to claim 1.

11. If, in response to a request, the computer finds one or more entity lists in which entities identified based on the request are recorded, it will determine the usage based on one or more operation logs and one or more data logs identified using those one or more entity lists. For each entity, the access policy includes, for each of one or more operational attributes, access permissions per nth data point for use in the application or model for that entity (where n is a non-negative integer). For each operation, The operation log, which is the log of the operation in question, A data log corresponding to the source data, which is the data for the operation, and associated with the operation log, and / or a data log corresponding to the target data, which is the data resulting from the operation, and associated with the operation log. There is, Regarding the operation log and / or data log, there is an entity list based on the aforementioned access policy. Each entity list includes an inclusion list, which is a list of entities whose data is permitted for use in the operation or data corresponding to the log to which the entity list is associated, and / or an exclusion list, which is a list of entities whose data is prohibited from being used in the operation or data corresponding to the log to which the entity list is associated. The computer returns data representing the identified usage to the requester of the request. Data management methods.

Citation Information

Patent Citations

  • Method and device for managing data on literary work

    JP1996137686A

  • Copyrighted matter provision system using network

    JP1998055390A

  • Log managing method / Device

    JP2000029751A

  • Information processing device, information processing method, and information processing system

    JP2014029587A

  • Information providing method, information management system and control method for terminal equipment

    JP2016006553A