Computing system, computing method, computing program, and computing apparatus
The computing system addresses the challenge of combining personal data from multiple sources by converting and securely sharing non-identifying information, enabling comprehensive data analysis while preserving privacy through federated learning techniques.
Patent Information
- Application Number
- JP2023222868
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
The challenge lies in safely and efficiently utilizing personal information owned by multiple companies while adhering to data privacy regulations, as individual companies may lack comprehensive data but face difficulties in sharing personal information due to privacy concerns.
A computing system that converts personal information into non-identifying processed information and performs federated learning across multiple terminal devices to generate an inference model, allowing secure and easy access to combined data for analysis, while ensuring privacy through methods like bucketing or MPC.
Enables the safe and easy use of personal information across multiple owners, facilitating more comprehensive data analysis without directly sharing raw personal data, thus enhancing data utility while maintaining privacy.
Smart Images

Figure 2025104795000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computing system, a computing method, a computing program, and a computing device.
Background Art
[0002] In recent years, many companies have been using the personal information of their customers for marketing purposes. For example, analyzing the purchasing trends for each customer segment and using them for sales promotions have been carried out.
[0003] As a technique for using customer personal information, for example, in Cited Document 1, there is disclosed a customer information collection means for collecting sales information from customer POS data and creating customer aggregation data in which the collected sales information of the customer is associated with the personal information of the customer, and based on the customer aggregation data, a segmentation analysis means for clustering the customer into segments according to the lifestyle of the customer by non-hierarchical clustering and hierarchical clustering.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Since the personal information owned by one company is limited, there is a need to use not only the personal information owned by the company itself but also the personal information owned by other companies. On the other hand, from the viewpoint of personal information protection, it is difficult to use the personal information owned by other companies.
[0006] The present invention has been made in view of the above problems, and an object thereof is to enable the safe and easy use of personal information owned by a plurality of owners.
Means for Solving the Problem
[0007] A computing system according to an embodiment is a computing system including a plurality of terminal devices and a user terminal. Each of the plurality of terminal devices stores personal information including identification information and attribute information, converts the personal information into non-identifying processed information, sorts the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information, and generates a part of an inference model for inferring a target variable by collaborative learning. One of the plurality of terminal devices generates the inference model from a plurality of parts of the inference models generated by the plurality of terminal devices, searches for a set of explanatory variables for which the target variable satisfies a constraint condition from the inference model, and the user terminal receives the target variable, the explanatory variables, and the constraint condition set by the user, displays the set of explanatory variables, and the target variable and the explanatory variables are selected by the user from among the attribute information.
Advantages of the Invention
[0008] According to an embodiment, personal information owned by a plurality of owners can be used safely and easily.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Mode for Carrying Out the Invention
[0010] Hereinafter, each embodiment of the present invention will be described with reference to the accompanying drawings. In addition, regarding the description in the specification and drawings according to each embodiment, for components having substantially the same functional configuration, the same reference numerals are attached and duplicate descriptions are omitted.
[0011] <System Configuration> First, the outline of the computing system 1000 according to the present embodiment will be described. The computing system 1000 is an information processing system that makes the entire personal information owned by a plurality of owners legally available by converting the personal information owned by a plurality of owners into non-identifying processed information and then performing federated learning to generate an inference model.
[0012] Personal information is information about living individuals, which can identify a specific individual by the name, date of birth, and other descriptions included in the information. Personal information includes identifying information and attribute information.
[0013] Identification information refers to information that identifies an individual. Identification information includes, for example, but is not limited to, name, phone number, email address, or face image. Personal information may include multiple pieces of identification information.
[0014] Attribute information refers to information that indicates an individual's attributes. Attribute information includes, for example, but is not limited to, age, gender, address, occupation, or purchased goods. Personal information may include multiple pieces of attribute information.
[0015] Non-identifying processed information refers to information obtained by de-identifying personal information, that is, information processed so that a specific individual cannot be identified unless it is matched with at least other information. Non-identifying processed information can be personal information or non-personal information. Non-identifying processed information that is personal information is, for example, pseudonymized information. Non-identifying processed information that is non-personal information is, for example, pseudonymized information, anonymized information, personally related information, or other non-personal information.
[0016] Pseudonymized information refers to information related to an individual obtained by processing personal information so that a specific individual cannot be identified unless it is matched with at least other information.
[0017] Anonymized information refers to information related to an individual obtained by processing personal information so that a specific individual cannot be identified and the original personal information cannot be restored.
[0018] Personally related information refers to information related to a living individual that does not fall under any of personal information, pseudonymized information, or anonymized information.
[0019] FIG. 1 is a diagram showing an example of the configuration of a computing system 1000. As shown in FIG. 1, the computing system 1000 includes a plurality of terminal devices 1 and a user terminal 2 that are communicably connected to each other via a network N. The network N is, for example, a wired LAN (Local Area Network), a wireless LAN, the Internet, a public switched telephone network, a mobile data communication network, or a combination thereof. In the example of FIG. 1, the computing system 1000 includes two terminal devices 1A and 1B, but may include three or more terminal devices 1. Also, in the example of FIG. 1, the computing system 1000 includes one user terminal 2, but may include two or more user terminals 2.
[0020] The terminal device 1 is an information processing device that converts personal information into non-identifying processed information and performs federated learning. The terminal device 1 is, for example, a PC (Personal Computer), a smartphone, a tablet terminal, a server device, or a microcomputer, but is not limited thereto.
[0021] The user terminal 2 is an information processing device used by a user of the computing system 1000. The user may or may not be the owner of the personal information. The user terminal 2 is, for example, a PC, a smartphone, or a tablet terminal, but is not limited thereto. The user terminal 2 may function as the terminal device 1.
[0022] <Hardware Configuration of Terminal Device 1> Next, the hardware configuration of the terminal device 1 will be described. FIG. 2 is a diagram showing an example of the hardware configuration of the terminal device 1. As shown in FIG. 2, the terminal device 1 includes a processor 101, a memory 102, a storage 103, a communication I / F 104, an input device 105, an output device 106, and a drive device 107 that are interconnected via a bus B1.
[0023] The processor 101 controls each component of the terminal device 1 and realizes the functions of the terminal device 1 by expanding and executing various programs including the OS (Operating System) and calculation programs stored in the storage 103 in the memory 102. The processor 101 is, for example, a CPU, MPU (Micro Processing Unit), GPU (Graphics Processing Unit), ASIC (Application Specific Integrated Circuit), or DSP (Digital Signal Processor), but is not limited thereto.
[0024] The memory 102 is, for example, a ROM (Read Only Memory), RAM, or a combination thereof. The ROM is, for example, a PROM (Programmable ROM), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), or a combination thereof. The RAM is, for example, a DRAM (Dynamic RAM) or SRAM (Static RAM), but is not limited thereto.
[0025] The storage 103 stores various programs and data including the OS and calculation programs. The storage 103 is, for example, a flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or SCM (Storage Class Memories), but is not limited thereto.
[0026] The communication I / F 104 is an interface for connecting the terminal device 1 to an external device via the network N and controlling communication. The communication I / F 104 is, for example, Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), or Ethernet (registered trademark), but is not limited thereto.
[0027] The input device 105 is a device for inputting information into the terminal device 1. The input device 105 is, for example, a mouse, a keyboard, a touch panel, a microphone, a scanner, a photographing device (camera), various sensors, or operation buttons, but is not limited thereto.
[0028] The output device 106 is a device for outputting information from the terminal device 1. The output device 106 is, for example, a display device (display), a projector, a printer, a speaker, or a vibrator, but is not limited thereto.
[0029] The drive device 107 is a device for reading and writing data on the recording medium 108. The drive device 107 is, for example, a magnetic disk drive, an optical disk drive, a magneto-optical disk drive, or an SD card reader, but is not limited thereto. The recording medium 108 is, for example, a CD (Compact Disc), a DVD (Digital Versatile Disc), an FD (Floppy Disk), an MO (Magneto-Optical disk), a BD (Blu-ray (registered trademark) Disc), a USB (registered trademark) memory, or an SD card, but is not limited thereto.
[0030] Note that, in the present embodiment, the calculation program may be written in the memory 102 or the storage 103 at the manufacturing stage of the terminal device 1, may be provided to the terminal device 1 via the network N, or may be provided to the terminal device 1 via a non-transitory computer-readable recording medium such as the recording medium 108.
[0031] <Hardware Configuration of User Terminal 2> Next, the hardware configuration of the user terminal 2 will be described. FIG. 3 is a diagram showing an example of the hardware configuration of the user terminal 2. As shown in FIG. 3, the user terminal 2 includes a processor 201, a memory 202, a storage 203, a communication I / F 204, an input device 205, an output device 206, and a drive device 207, which are interconnected via a bus B2.
[0032] The processor 201 controls each component of the user terminal 2 and realizes the functions of the user terminal 2 by expanding and executing various programs including the OS and calculation programs stored in the storage 203 in the memory 202. The processor 201 is, for example, a CPU, MPU, GPU, ASIC, or DSP, but is not limited thereto.
[0033] The memory 202 is, for example, a ROM, RAM, or a combination thereof. The ROM is, for example, a PROM, EPROM, EEPROM, or a combination thereof. The RAM is, for example, a DRAM or SRAM, but is not limited thereto.
[0034] The storage 203 stores various programs and data including the OS and calculation programs. The storage 203 is, for example, a flash memory, HDD, SSD, or SCM, but is not limited thereto.
[0035] The communication I / F 204 is an interface for connecting the user terminal 2 to an external device via the network N and controlling communication. The communication I / F 204 is, for example, Bluetooth (registered trademark), Wi-Fi (registered trademark), ZigBee (registered trademark), or Ethernet (registered trademark), but is not limited thereto.
[0036] The input device 205 is a device for inputting information to the user terminal 2. The input device 205 is, for example, a mouse, keyboard, touch panel, microphone, scanner, imaging device (camera), various sensors, or operation buttons, but is not limited thereto.
[0037] The output device 206 is a device for outputting information from the user terminal 2. The output device 206 is, for example, a display device (display), projector, printer, speaker, or vibrator, but is not limited thereto. The user terminal 2 includes a display device 206D as the output device 206.
[0038] The drive device 207 is a device for reading and writing data on the recording medium 208. The drive device 207 is, for example, a magnetic disk drive, an optical disk drive, a magneto-optical disk drive, or an SD card reader, but is not limited thereto. The recording medium 208 is, for example, a CD, a DVD, an FD, an MO, a BD (Blu-ray (registered trademark) Disc), a USB (registered trademark) memory, or an SD card, but is not limited thereto.
[0039] Note that, in the present embodiment, the calculation program may be written in the memory 202 or the storage 203 at the manufacturing stage of the user terminal 2, may be provided to the user terminal 2 via the network N, or may be provided to the user terminal 2 via a non-transitory computer-readable recording medium such as the recording medium 208.
[0040] <Functional Configuration of Terminal Device 1> Next, the functional configuration of the terminal device 1 will be described. FIG. 4 is a diagram showing an example of the functional configuration of the terminal device 1. As shown in FIG. 4, the terminal device 1 includes a communication unit 11, a storage unit 12, and a control unit 13.
[0041] The communication unit 11 is realized by the communication I / F 104. The communication unit 11 transmits and receives information to and from the user terminal 2 and other terminal devices 1 via the network N.
[0042] The storage unit 12 is realized by the memory 102 and the storage 103. The storage unit 12 stores personal information 121, non-identifying processed information 122, search condition information 123, correspondence information 124, learning data 125, inference model information 126, and search result information 127.
[0043] The personal information 121 includes identification information and attribute information. Each terminal device 1 stores personal information 121 owned by different owners.
[0044] FIG. 5 is a diagram showing an example of the personal information 121A and 121B respectively stored in the terminal devices 1A and 1B.
[0045] The personal information 121A in FIG. 5 includes "No", "Name", "Age", "Gender", and "Product" as information items. "No" is the record number. "Name" is the information indicating the personal name. "Age" is the information indicating the personal age. "Gender" is the information indicating the personal gender. "Product" is the information indicating the purchased product of the individual. "Name" corresponds to the identification information. "Age", "Gender", and "Product" correspond to the attribute information.
[0046] The personal information 121B in FIG. 5 includes "No", "Name", "Occupation", and "Address" as information items. "No" is the record number. "Name" is the information indicating the personal name. "Occupation" is the information indicating the personal occupation. "Address" is the information indicating the personal address. "Name" corresponds to the identification information. "Occupation" and "Address" correspond to the attribute information.
[0047] Since the personal information 121A and 121B are personal information owned by different owners, as shown in FIG. 5, they may include the personal information of the same person (aaa, ccc), or they may include the personal information of different persons (bbb, ddd, fff). Similarly, the personal information 121A and 121B may include the same attribute information, or they may include different attribute information (age, gender, product, occupation, address).
[0048] Note that the personal information 121 is not limited to the example in FIG. 5. The personal information 121 may not include some of the above information items, or may include information items other than the above.
[0049] The non-identifying processed information 122 is non-identifying processed information obtained by converting the personal information 121. The identifying information of the non-identifying processed information 122 is de-identified. The method for de-identifying the identifying information of the non-identifying processed information 122 is, for example, hashing, adding noise, or salted hashing, but is not limited thereto. Also, the attribute information of the non-identifying processed information 122 may be de-identified. The method for de-identifying the attribute information of the non-identifying processed information 122 is, for example, deleting some of the attribute information, abstracting the attribute information, upper concept generalization, quantification, changing the granularity, top coding, bottom coding, data exchange, k-anonymization, or adding noise, but is not limited thereto.
[0050] FIG. 6 is a diagram showing an example of the non-identifying processed information 122A and 122B respectively stored in the terminal devices 1A and 1B.
[0051] The non-identifying processed information 122A in FIG. 6 is information obtained by de-identifying the personal information 121A, and includes "No", "Name", "Age", "Gender", and "Product" as information items. "No" is a record number. Personal information. "Name" corresponds to identifying information and is salted hashed. "Age", "Gender", and "Product" correspond to attribute information, "Age" has its granularity changed, and "Gender" and "Product" are quantified.
[0052] The non-identifying processed information 122B in FIG. 6 is information obtained by de-identifying the personal information 121B, and includes "No", "Name", "Occupation", and "Address" as information items. "No" is a record number. "Name" corresponds to identifying information and is salted hashed. "Occupation" and "Address" are k-anonymized.
[0053] Since the personal information 121A and 121B are de-identified by the same method, the identifying information (aaa, ccc) and the values of the attribute information (bc944ff0, 48aa00cb) common to the personal information 121A and 121B are also common values in the non-identifying processed information 122A and 122B.
[0054] The exploration condition information 123 is information indicating exploration conditions for exploring a set of explanatory variables for which the target variable satisfies the constraint conditions from the inference model m. The exploration conditions include a target variable, explanatory variables, and constraint conditions. Details of the exploration conditions will be described later.
[0055] The correspondence information 124 is information indicating the correspondence relationship between personal information stored in each of the plurality of terminal devices 1. In other words, the correspondence information 124 is information indicating which of the plurality of personal information stored in each of the plurality of terminal devices 1 is personal information of the same individual. Personal information of the same individual refers to personal information having the same identification information.
[0056] FIG. 7 is a diagram showing an example of correspondence information indicating the correspondence relationship between personal information 121A and 121B. The example of FIG. 7 shows that the personal information in personal information 121A where "No" is "001" and "003" corresponds to the personal information in personal information 121B where "No" is "001" and "002", respectively (personal information of the same individual).
[0057] The learning data 125 is information obtained by sorting the non-identified processed information 122 using the correspondence information 124 and is used for joint learning.
[0058] The inference model information 126 is information regarding the inference model m. The inference model m is a learned machine learning model that outputs the ratio at which the target variable is a certain value when an explanatory variable is input. Any machine learning model can be used as the inference model m, but it is preferable that the inference model m is a single decision tree. When the inference model m is a single decision tree, the exploration of the set of explanatory variables described later can be easily executed.
[0059] The exploration result information 127 is information indicating the exploration result of exploring a set of explanatory variables that satisfy the exploration conditions from the inference model m. The exploration result includes a set of explanatory variables for which the target variable satisfies a predetermined condition. Details of the exploration result will be described later.
[0060] The control unit 13 is realized by the processor 101 reading and executing a program from the memory 102 and collaborating with other hardware components. The control unit 13 controls the overall operation of the terminal device 1. The control unit 13 includes an acquisition unit 131, a conversion unit 132, a sorting unit 133, a learning unit 134, a generation unit 135, and a search unit 136.
[0061] The acquisition unit 131 acquires search condition information 221 from the user terminal 2 and stores it in the storage unit 12 as search condition information 123.
[0062] The conversion unit 132 converts the personal information 121 into non-identifying processed information 122 and stores it in the storage unit 12.
[0063] The sorting unit 133 uses the correspondence information 124 to sort the non-identifying processed information 122 to generate learning data 125 and stores it in the storage unit 12.
[0064] The learning unit 134 executes federated learning using the learning data 125 to generate a part of the inference model m, and stores information indicating the generated part of the inference model m in the storage unit 12 as inference model information 126. The learning units 134 of multiple terminal devices 1 cooperate with each other to execute vertical federated learning. As the method of vertical federated learning, any method can be selected, such as Feder Boost, Secure Boost, Fed GBF, PIVODL, Secure GBM, Fed XGB, Pivot, Pri VDT, VF CART, MP Fed XGB, and methods modified based on decision trees for these algorithms.
[0065] The generation unit 135 generates the inference model m from parts of multiple inference models m stored in the storage units 12 of multiple terminal devices 1, and stores information indicating the inference model m in the storage unit 12 as inference model information 126.
[0066] The search unit 136 searches for a set of explanatory variables that satisfy the search conditions from the inference model m, and stores information indicating the search results in the storage unit 12 as search result information 127.
[0067] Note that the functional configuration of the terminal device 1 is not limited to the above example. For example, the terminal device 1 may have a functional configuration other than the above, or some of the functional configurations may be provided in other devices. Also, each functional configuration of the terminal device 1 may be implemented by software as described above, or may be implemented by hardware such as an IC chip, SoC, LSI, or microcomputer.
[0068] <Functional Configuration of User Terminal 2> Next, the functional configuration of the user terminal 2 will be described. FIG. 8 is a diagram showing an example of the functional configuration of the user terminal 2. As shown in FIG. 8, the user terminal 2 includes a communication unit 21, a storage unit 22, and a control unit 23.
[0069] The communication unit 21 is realized by the communication I / F 204. The communication unit 21 transmits and receives information to and from the terminal device 1 via the network N.
[0070] The storage unit 22 is realized by the memory 202 and the storage 203. The storage unit 22 stores search condition information 221 and search result information 222.
[0071] The search condition information 221 is information indicating search conditions for searching for a set of explanatory variables from the inference model m. The search conditions include a target variable, explanatory variables, and constraint conditions. The search conditions will be described in detail later.
[0072] The search result information 222 is information indicating the search results of searching for a set of explanatory variables that satisfy the search conditions from the inference model m. The search results include a set of explanatory variables for which the target variable satisfies a predetermined condition. The search results will be described in detail later.
[0073] The control unit 23 is realized by the processor 201 reading and executing a program from the memory 202 and cooperating with other hardware configurations. The control unit 23 controls the overall operation of the user terminal 2. The control unit 23 includes an acquisition unit 231, a display unit 232, and a reception unit 233.
[0074] The acquisition unit 231 acquires the search result information 127 from the terminal device 1 and stores it in the storage unit 22 as the search result information 222.
[0075] The display unit 232 controls the screen displayed on the display device 206D. The display unit 232 displays the search result information 222 on the display device 206D according to the user's operation.
[0076] The reception unit 233 receives the search conditions set by the user and stores the information indicating the received search conditions in the storage unit 22 as the search condition information 221.
[0077] Note that the functional configuration of the user terminal 2 is not limited to the above example. For example, the user terminal 2 may have a functional configuration other than the above, or some functional configurations may be provided in other devices. Also, each functional configuration of the user terminal 2 may be realized by software as described above, or may be realized by hardware such as an IC chip, SoC, LSI, microcomputer, etc.
[0078] <Processes executed by the computing system 1000> Next, the processes executed by the computing system 1000 will be described. FIG. 10 is a flowchart showing an example of the processes executed by the computing system 1000. Hereinafter, the case where the user uses the personal information 121A, 121B stored in the terminal devices 1A, 1B will be described as an example, but the user can also use the personal information 121 stored in three or more terminal devices 1. Hereinafter, the configuration of the terminal device 1A will be described with an A appended to the end of the reference numeral, and the configuration of the terminal device 1B will be described with a B appended to the end of the reference numeral.
[0079] (Step S101) First, the display unit 232 of the user terminal 2 displays the personal information selection screen sc1 on the display device 206D according to the user's operation (Step S101). The personal information selection screen sc1 is a screen for the user to select the personal information to be used.
[0080] FIG. 10 is a diagram showing an example of the personal information selection screen sc1. The personal information selection screen sc1 in FIG. 10 includes a personal information selection section sc11, a selection button sc12, and a return button sc13.
[0081] In the personal information selection section sc11, a list of personal information available to the user is displayed. In the example of FIG. 10, three pieces of personal information are displayed, and as descriptions of each piece of personal information, "Title", "Owner", "Registration Date", "Number of Records", and "Information Items" are described. "Title" is the name given to the personal information. "Owner" is the name of the owner of the personal information. "Registration Date" is the date on which the personal information was registered in the terminal device 1. "Number of Records" is the number of records included in the personal information. "Information Items" is the type of information items included in the personal information.
[0082] The user refers to the personal information selection section sc11 and selects the personal information to be used. In the example of FIG. 10, personal information 121A and 121B are temporarily selected by the user.
[0083] The selection button sc12 is a button for the user to select the personal information to be used. When the user presses the selection button sc12, the temporarily selected personal information is selected as the personal information to be used by the user.
[0084] The return button sc13 is a button for returning to the previous screen of the personal information selection screen sc1.
[0085] Note that the personal information selection screen sc1 is not limited to the example of FIG. 10. The personal information selection screen sc1 can be any screen on which the user can select the personal information to be used.
[0086] (Step S102) When the user selects personal information, the display unit 232 displays the search condition input screen sc2 on the display device 206D (Step S102). The search condition input screen sc2 is a screen for the user to input search conditions.
[0087] FIG. 11 is a diagram showing an example of a search condition input screen sc2. The search condition input screen sc2 in FIG. 11 includes an explanatory variable selection section sc21, an objective variable selection section sc22, an objective variable value selection section sc23, a ratio selection section sc24, an execution button sc25, and a return button sc26.
[0088] The explanatory variable selection section sc21 is a part for the user to select an explanatory variable from among the information items of the attribute information included in the personal information selected by the user. In the example of FIG. 11, in the explanatory variable selection section sc21, "age", "gender", "product", "occupation", and "address", which are the attribute information included in the personal information 121A and 121B, are displayed so as to be selectable by check boxes, and "age", "gender", "occupation", and "address" have been selected as explanatory variables by the user.
[0089] The objective variable selection section sc22 is a part for the user to select an objective variable from among the information items of the attribute information included in the personal information selected by the user. In the example of FIG. 11, in the objective variable selection section sc22, "age", "gender", "product", "occupation", and "address", which are the attribute information included in the personal information 121A and 121B, are displayed so as to be selectable by check boxes, and "product" has been selected as the objective variable by the user.
[0090] The objective variable value selection section sc23 is a part for the user to select the value of the objective variable selected by the user. The value of the objective variable selected here is an example of a constraint condition. In the example of FIG. 11, in the objective variable value selection section sc23, the values (A, B, C, and D) of the objective variable "product" selected in the objective variable selection section sc22 are displayed so as to be selectable by check boxes, and "A" has been selected as the value of the objective variable by the user.
[0091] The ratio selection unit sc24 is a part for the user to select the ratio of the value of the target variable selected by the target variable value selection unit sc23. The ratio selected here is an example of a constraint condition. In the example of FIG. 11, a text box capable of inputting numerical values is displayed in the ratio selection unit sc24, and a ratio of "70"% or more has been selected (input) by the user. Note that the method of selecting the ratio is not limited to text input and is arbitrary.
[0092] The execution button sc25 is a button for requesting the terminal device 1 to execute a search according to the search conditions.
[0093] The back button sc26 is a button for returning to the previous screen (personal information selection screen sc1) of the search condition input screen sc2.
[0094] Note that the search condition input screen sc2 is not limited to the example of FIG. 11. The search condition input screen sc2 can be any screen that allows the user to select a target variable, constraint conditions, and explanatory variables for the search. Also, the constraint conditions are not limited to the value and ratio of the target variable. For example, the constraint condition may be the number of records.
[0095] (Step S103) When the user inputs search conditions on the search condition input screen sc2 and presses the execution button sc25, the reception unit 233 receives the search conditions selected by the user and stores information indicating the received search conditions in the storage unit 22 as search condition information 221 (Step S103).
[0096] (Step S104) The communication unit 21 of the user terminal 2 transmits the search condition information 221 to the terminal devices 1A and 1B that store the personal information 121A and 121B selected by the user (Step S104).
[0097] (Step S105) The acquisition unit 131A of the terminal device 1A acquires the search condition information 221 received from the user terminal 2 and stores it in the storage unit 12A as the search condition information 123A (step S105A). Similarly, the acquisition unit 131B of the terminal device 1B acquires the search condition information 221 received from the user terminal 2 and stores it in the storage unit 12B as the search condition information 123B (step S105B).
[0098] (Step S106) When the conversion unit 132A of the terminal device 1A acquires the search condition information 123A, it converts the personal information 121A into non-identifying processed information 122A and stores it in the storage unit 22A (step S106A). Similarly, when the conversion unit 132B of the terminal device 1B acquires the search condition information 123B, it converts the personal information 121B into non-identifying processed information 122B and stores it in the storage unit 22B (step S106B).
[0099] Here, the conversion units 132A and 132B anonymize the personal information 121A and 121B in the same way. For example, when salt-hashing the identification information of the personal information 121A and 121B, the conversion units 132A and 132B share the same salt and hash function in advance, use these to salt-hash the identification information, and then discard the salt.
[0100] (Step S107) The sorting unit 133A of the terminal device 1A acquires the identification information of the non-identifying processed information 122B from the terminal device 1B, compares it with the identification information of the non-identifying processed information 122A, generates corresponding information 124A, and stores it in the storage unit 12A (step S107). Since the identification information of the non-identifying processed information 122A and 122B is anonymized in the same way, the identification information of the non-identifying processed information 122A and 122B corresponding to the same individual matches. Therefore, the sorting unit 133A may determine the non-identifying processed information 122A and 122B with matching identification information as the corresponding non-identifying processed information 122A and 122B.
[0101] (Step S108) The terminal device 1A transmits the correspondence information 124A to the terminal device 1B (step S108). The terminal device 1B stores the received correspondence information 124A in the storage unit 12B as the correspondence information 124B.
[0102] Note that the correspondence information 124 may be generated by the terminal device 1B and transmitted to the terminal device 1A. Further, the correspondence information 124 may be generated by an information processing device other than the terminal devices 1A and 1B (for example, the terminal device 1 other than the terminal devices 1A and 1B) and transmitted to the terminal devices 1A and 1B.
[0103] (Step S109) The sorting unit 133A sorts the non-identified processed information 122A using the correspondence information 124A and stores it in the storage unit 12A as learning data 125A (step S109A). Similarly, the sorting unit 133B sorts the non-identified processed information 122B using the correspondence information 124B and stores it in the storage unit 12B as learning data 125B (step S109B).
[0104] Specifically, the sorting unit 133 extracts from the non-identified processed information 122 those that show a correspondence relationship with the correspondence information 124, and arranges the extracted ones in the order indicated by the correspondence information 124. In other words, the sorting unit 133 removes those from the non-identified processed information 122 that do not show a correspondence relationship with the correspondence information 124, and arranges the remaining ones in the order indicated by the correspondence information 124.
[0105] FIG. 12 is a diagram showing an example of the learning data 125A and 125B. The learning data 125A and 125B in FIG. 12 are generated from the non-identified processed information 122A and 122B in FIG. 6 using the correspondence information 124 in FIG. 7. As shown in FIG. 12, the learning data 125A is obtained by extracting from the non-identified processed information 122A the non-identified processed information where "No" is "001" and "003" and arranging them in this order, and the learning data 125B is obtained by extracting from the non-identified processed information 122B the non-identified processed information where "No" is "001" and "002" and arranging them in this order.
[0106] In this way, by sorting the non-identifying processed information 122 using the corresponding information 124, consistent learning data 125A and 125B in which the non-identifying processed information of the same individual is arranged in the same order can be generated between different terminal devices 1A and 1B.
[0107] (Step S110) The learning units 134A and 134B execute federated learning using the learning data 125A and 125B, respectively generate a part of the inference model m, and store information indicating a part of the inference model m in the storage units 12A and 12B as inference model information 126A and 126B (Step S110). The method of federated learning will be described in detail later.
[0108] (Step S111) Thereafter, the terminal device 1B transmits the inference model information 126B including the information indicating a part of the inference model m to the terminal device 1A (Step S111).
[0109] (Step S112) When the generation unit 135A of the terminal device 1A receives the inference model information 126B, it generates the inference model m from the information indicating a part of the inference model m included in the inference model information 126A and 126B, respectively, and stores the information indicating the inference model m in the storage unit 12A as the inference model information 126A (Step S112). Thereby, the inference model m for inferring the target variable based on the explanatory variable is generated. The method for generating the inference model m will be described in detail later.
[0110] (Step S113) The exploration unit 136A refers to the exploration condition information 123A, explores a set of explanatory variables that satisfy the exploration conditions from the inference model m, and stores the exploration result as exploration result information 127A in the storage unit 12A (step S113). The exploration result includes the set of explanatory variables discovered by the exploration and the number of learning data 125A (non-identification processed information 122A) extracted by the explanatory variables. More specifically, the exploration unit 136A explores a set of explanatory variables for which the target variable satisfies the constraint conditions. In the example of FIG. 11, the exploration unit 136A explores a set (combination of values) of "age", "gender", "occupation", and "address" for which the learning data 125A (records) with "product" being "A" is 70% or more.
[0111] FIG. 13 is a diagram showing an example of the exploration method by the exploration unit 136A. In the example of FIG. 13, the inference model m is a decision tree, and the exploration method is a decision tree algorithm. In FIG. 13, X:Y represents the number of learning data 125A for which X satisfies the constraint conditions and the number of learning data 125A for which Y does not satisfy the constraint conditions. That is, X / (X + Y) corresponds to the ratio. Here, it is assumed that there are 100 pieces (100 records) of learning data 125A, the target variable is "product", the explanatory variables are "age", "prefecture", "occupation", and "gender", the constraint conditions are "product = A" and "ratio ≥ 70%", and the end condition of the exploration is X ≤ 5. In FIG. 13, "address" and "gender" are digitized.
[0112] (Step S201) First, the exploration unit 136A tabulates the learning data 125A with "product = A" (40 pieces) and the learning data 125A without "product = A" (60 pieces) among all the learning data 125A, and calculates the ratio of the learning data 125A with "product = A" (step S201). Since the ratio here is 40%, the constraint conditions ("product = A" and "ratio ≥ 70%") are not satisfied.
[0113] (Step S202) Therefore, the exploration unit 136A sets a threshold value (50 years old) for "age" (step S202).
[0114] (Step S203) The search unit 136A aggregates the learning data 125A of "product = A" (10 pieces) and the learning data 125A that is not "product = A" (50 pieces) among the learning data 125A of "age > 50", and calculates the ratio of the learning data 125A of "product = A" (step S203). Since the ratio here is about 17%, the constraint conditions ("product = A" and "ratio ≥ 70%") are not satisfied.
[0115] (Step S204) Therefore, the search unit 136A sets a threshold value (8) for "address (pref)" (step S204).
[0116] (Step S205) The search unit 136A aggregates the learning data 125A of "product = A" (2 pieces) and the learning data 125A that is not "product = A" (30 pieces) among the learning data 125A of "age > 50" and "address > 8", and calculates the ratio of the learning data 125A of "product = A" (step S205). Since the ratio here is about 6%, the constraint conditions ("product = A" and "ratio ≥ 70%") are not satisfied. Since X ≤ 5 for the search unit 136A, the search for this branch ends.
[0117] (Step S206) In addition, the search unit 136A aggregates the learning data 125A of "product = A" (8 pieces) and the learning data 125A that is not "product = A" (20 pieces) among the learning data 125A of "age > 50" and "address ≤ 8", and calculates the ratio of the learning data 125A of "product = A" (step S206). Since the ratio here is about 29%, the constraint conditions ("product = A" and "ratio ≥ 70%") are not satisfied.
[0118] (Step S207) Therefore, the search unit 136A sets a threshold value (0) for "gender (sex)" (step S207).
[0119] (Step S208) Among the learning data 125A of "age > 50", "address ≤ 8" and "gender > 0", the exploration unit 136A aggregates the learning data 125A of "product = A" (7 pieces) and the learning data 125A that is not "product = A" (1 piece), and calculates the ratio of the learning data 125A of "product = A" (step S208). Since the ratio here is about 88%, the constraint conditions ("product = A" and "ratio ≥ 70%") are satisfied.
[0120] (Step S209) As a result, "age > 50", "address ≤ 8" and "gender > 0" are discovered as a set of explanatory variables for which the target variable satisfies the constraint conditions (step S209). Since the exploration unit 136A has discovered a set of explanatory variables, the exploration of this branch ends.
[0121] (Step S210) Among the learning data 125A of "age > 50", "address ≤ 8" and "gender ≤ 0", the exploration unit 136A aggregates the learning data 125A of "product = A" (1 piece) and the learning data 125A that is not "product = A" (19 pieces), and calculates the ratio of the learning data 125A of "product = A" (step S210). Since the ratio here is 5%, the constraint conditions ("product = A" and "ratio ≥ 70%") are not satisfied. Since X ≤ 5 for the exploration unit 136A, the exploration of this branch ends.
[0122] (Step S211) Among the learning data 125A of "age ≤ 50", the exploration unit 136A aggregates the learning data 125A of "product = A" (30 pieces) and the learning data 125A that is not "product = A" (10 pieces), and calculates the ratio of the learning data 125A of "product = A" (step S211). Since the ratio here is 75%, the constraint conditions ("product = A" and "ratio ≥ 70%") are satisfied.
[0123] (Step S212) As a result, "age ≤ 50" is discovered as a set of explanatory variables for which the target variable satisfies the constraint conditions (step S212). Since the search unit 136A has discovered a set of explanatory variables, it ends the search for this branch.
[0124] By the above method, as search results, "age > 50", "address ≤ 8" and "gender > 0", and "age ≤ 50" are discovered.
[0125] The search unit 136A can search for a set of explanatory variables for which the target variable satisfies the constraint conditions by repeatedly performing such a decision tree algorithm while changing the threshold values of each attribute information. Note that the order of the attribute information for setting the threshold values is arbitrary. The order of the attribute information for setting the threshold values may be common throughout all searches, may be randomly changed for each search, or may be changed every predetermined number of searches. Also, the value of the threshold set for each search is arbitrary. The value of the threshold may be randomly set for each search, or may be changed at a predetermined interval for each search. Further, the search unit 136A may perform the search using a decision tree algorithm other than the above, or may perform the search by a method other than the decision tree algorithm (such as a heuristic).
[0126] Note that the terminal device 1B may generate the inference model m and execute the search. Also, an information processing device other than the terminal devices 1A and 1B (for example, a terminal device 1 other than the terminal devices 1A and 1B) may generate the inference model m and execute the search.
[0127] Here, return to the description of FIG. 9.
[0128] (Step S114) The terminal device 1A transmits the search result information 127A to the user terminal 2 (step S114). The terminal device 1A may transmit the search result information 127A as it is, or may transmit it using an encrypted communication path.
[0129] (Step S115) The acquisition unit 231 of the user terminal 2 stores the search result information 127A received from the terminal device 1A in the storage unit 22 as the search result information 222 (step S115).
[0130] (Step S116) The display unit 232 refers to the search result information 222 according to the user's operation and displays the search result on the display device 206D (step S116).
[0131] FIG. 14 is a diagram showing an example of the search result display screen sc3. The search result display screen sc3 is a screen for displaying search results. As shown in FIG. 14, on the search result display screen sc3, for each set of explanatory variables, the set of explanatory variables sc31 and the number (volume) sc32 of the learning data 125A extracted by the explanatory variables are displayed. Thereby, the user can easily grasp how many people (segments) having attribute information that satisfies the constraint conditions exist without specifying an individual. Further, in the example of FIG. 14, the sets of explanatory variables sc31 are displayed in descending order of the number (volume) sc32 of the learning data 125A extracted by the explanatory variables. Thereby, the user can easily grasp a segment with a large scale.
[0132] <Method of collaborative learning> Here, the method of collaborative learning will be described in detail. This corresponds to the internal process of step S110. Hereinafter, it is assumed that the inference model m is a single decision tree, and the terminal device 1 storing the personal information 121 including the target variable set by the user is referred to as the terminal device AP (Active Party), and the other terminal devices 1 are referred to as the terminal device PP (Passive Party). Hereinafter, the configuration of the terminal device AP will be described with an A at the end of the reference numeral, and the configuration of the terminal device PP will be described with a P at the end of the reference numeral.
[0133] (1) Example 1 According to this embodiment, the computing system 1000 generates the inference model m using Feder Boost. More specifically, the learning unit 134 of each terminal device 1 buckets the learning data 125 and generates the inference model m using the bucketed learning data 125'. Bucketing is a process of replacing the attribute value, which is the value of the explanatory variable (attribute information), with a bucket value.
[0134] FIG. 15 is a flowchart showing an example of the federated learning method. FIG. 16 is a diagram for explaining the bucketing of the learning data 125. In FIG. 16, No is the record number of the learning data 125, and F is the explanatory variable (attribute information).
[0135] (Step S301) First, the terminal device AP sorts each explanatory variable (attribute information) included in the learning data 125A by the attribute value (step S301A). Similarly, the terminal device PP sorts each explanatory variable (attribute information) included in the learning data 125P by the attribute value (step S301P).
[0136] (Step S302) Next, the terminal device AP sets buckets for the sorted learning data 125A (step S302A). Similarly, the terminal device PP sets buckets for the sorted learning data 125P (step S302P).
[0137] Specifically, as shown in FIG. 16, the terminal device 1 groups a plurality of learning data 125 sorted by the attribute value of the explanatory variable F into one bucket in that order, and sets a bucket value B for each bucket. For example, in the example of FIG. 16, the learning data 125 with "No" being "000", "002", and "012" arranged in ascending order of the attribute value of the explanatory variable F are grouped as one bucket B0. The method of setting the bucket value is arbitrary.
[0138] (Step S303) The terminal device AP replaces the attribute value of the explanatory variable F with the bucket value B (step S303A). Similarly, the terminal device PP replaces the attribute value of the explanatory variable F with the bucket value B (step S303P). As a result, the value of the explanatory variable F is replaced with the bucket value B that has no direct relationship with the attribute value.
[0139] (Step S304) Also, the terminal device AP generates threshold information thA (step S304A). Similarly, the terminal device PP generates threshold information thP (step S304P).
[0140] The threshold information th is information indicating the threshold Fth of the attribute value between buckets. The threshold Fth is part of the inference model m and is calculated by an arbitrary method based on the attribute values included in the buckets. For example, in the example of FIG. 16, the threshold Fth of bucket B0 is 0.5. This indicates that the attribute value that is the threshold between buckets B0 and B1 is 0.5. This means that when the attribute value is less than 0.5, it is classified into bucket B0, and when the attribute value is 0.5 or more, it is classified into bucket B1.
[0141] (Step S305) Thereafter, the terminal device AP sorts each explanatory variable included in the learning data 125A by the bucket value (step S305A). Similarly, the terminal device PP sorts each explanatory variable included in the learning data 125P by the bucket value (step S305P). As a result, as shown in FIG. 16, the learning data 125' in which the value of the explanatory variable (attribute information) is converted from the attribute value to the bucket value is generated.
[0142] (Step S306) The terminal device PP transmits the learning data 125P' to the terminal device AP (step S306). As described above, since the learning data 125P' is data including bucket values that have no direct relationship with the attribute values, the security of the personal information 121P is ensured even when it is transmitted from the terminal device PP to the terminal device AP.
[0143] (Step S307) The terminal device AP learns the learning data 125A’ and 125P’, generates an inference model m’, and stores the information indicating the inference model m’ in the storage unit 12A as inference model information 126A (Step S307). The inference model m’ is a part of the inference model m and is a decision tree in which the threshold value of each node is represented by a bucket value instead of an attribute value. The information indicating the inference model m’ includes information indicating the structure of the inference model m’ and the threshold value of each node.
[0144] (Step S308) The terminal device AP transmits the inference model information 126A to the terminal device PP (Step S308). The inference model information 126A includes information indicating the structure of the inference model m’ and the threshold value of each node, and threshold value information thA.
[0145] (Step S309) The terminal device PP replaces the threshold value of the inference model m’ with an attribute value based on the inference model information 126A and the threshold value information thP (Step S309). Thereby, an inference model m, which is a decision tree in which the threshold value of each node is represented by an attribute value, is generated. The terminal device PP stores the information indicating the inference model m in the storage unit 12P as inference model information 126P.
[0146] In addition, when there are a plurality of terminal devices PP, any one of the plurality of terminal devices PP may receive the inference model information 126P (threshold value information thP) from another terminal device PP, receive the inference model information 126A from the terminal device AP, and generate the inference model m.
[0147] Alternatively, instead of the terminal device PP, the terminal device AP may receive the inference model information 126P (threshold value information thP) from the terminal device PP and generate the inference model m. However, from the viewpoint of the security of the personal information 121, it is preferable that the terminal device PP generates the inference model m.
[0148] As described above, according to this embodiment, since the learning data 125 (non-identifying and processing information 122) is not directly exchanged between the terminal devices AP and PP, the inference model m can be generated while ensuring the security of the personal information 121.
[0149] (2) Embodiment 2 According to this embodiment, the computing system 1000 uses MP Fed XGB to generate the inference model m. More specifically, the learning unit 134 of each terminal device 1 divides the learning data 125 into shares and generates the inference model m by MPC (Multi-Party Computation). A share is individual data obtained by dividing the original data.
[0150] FIG. 17 is a flowchart showing an example of the method of federated learning.
[0151] (Step S401) First, the terminal device AP divides the learning data 125A by the number of information processing devices participating in the MPC (step S401A). Similarly, the terminal device PP divides the learning data 125P by the number of information processing devices participating in the MPC (step S401P). Here, since there are two information processing devices, the terminal devices AP and PP, participating in the MPC, the learning data 125A and 125B are each divided into two shares.
[0152] (Step S402) Next, the terminal device AP transmits one of the shares obtained by dividing the learning data 125A to the terminal device PP (step S402A). Similarly, the terminal device PP transmits one of the shares obtained by dividing the learning data 125P to the terminal device AP (step S402P). As a result, the terminal devices AP and PP each hold a share of the learning data 125A and a share of the learning data 125P.
[0153] (Step S403) Subsequently, the terminal devices AP and PP use the shares to execute federated learning by MPC and generate an inference model m". The inference model m" is a part of the inference model m and is a decision tree in which data is composed of shares. The terminal devices AP and PP respectively store information indicating the inference model m" in the storage units 12A and 12P as inference model information 126A and 126P (step S403). The information indicating the inference model m" includes information indicating the structure of the inference model m", the ratio of the records in which the target variable is the value set by the user and is shared, the values (weights) of the shared leaf nodes (lowest-level nodes), and the threshold value of the plaintext.
[0154] Specifically, first, the terminal devices AP and PP share the information of the records existing in the node and the information of the ratio of the records in which the target variable is the value set by the user. Note that all records are included in the root node (topmost node).
[0155] Next, the terminal devices AP and PP calculate, by MPC, the value of the loss function when dividing the current node with each value of the feature amounts respectively possessed by the terminal devices AP and PP.
[0156] The terminal devices AP and PP compare, by MPC, the best value of the loss function when dividing the current node with each value of the feature amount and the value of the loss function when dividing the current node. If the latter is better than the former (the loss function is not improved even if the node is further divided), the division of the current node is terminated. In this case, the current node becomes a leaf node.
[0157] On the other hand, when the former is better than the latter (the loss function is improved when the node is divided), the terminal devices AP and PP restore the value of the feature amount with the best loss function value (thereby obtaining the threshold value of the plaintext). Assuming that the terminal device 1 having the restored feature amount has all the records, the terminal device 1 divides the records with the value of the feature amount with the best loss function value, shares the divided records and the ratio information, multiplies them by the share of the current node, and creates the share of the next node.
[0158] The terminal devices AP and PP repeat the above processes and end the processes when there are no nodes in the entire decision tree that can improve the loss function. As a result, the inference model m” is generated.
[0159] (Step S404) Based on the inference model information 126A and the share that the terminal device PP has, the terminal device AP restores the ratio of the records in the shared inference model m” where the target variable is the value set by the user, and the values of the leaf nodes (Step S404). As a result, the inference model m, which is a decision tree composed of plain text, is generated. The terminal device AP stores the information indicating the inference model m in the storage unit 12A as the inference model information 126A.
[0160] Alternatively, instead of the terminal device AP, the terminal device PP may restore the ratio of the records in the shared inference model m” where the target variable is the value set by the user and the values of the leaf nodes based on the inference model information 126P and the share that the terminal device AP has, generate the inference model A, and store the information indicating the inference model m in the storage unit 12P as the inference model information 126P.
[0161] As described above, according to this embodiment, since the learning data 125 (non-identifying processed information 122) is not directly exchanged between the terminal devices AP and PP, the inference model m can be generated while ensuring the security of the personal information 121.
[0162] Furthermore, in this embodiment, the terminal device AP may use the inference model m” generated in Step S403 to search for a set of explanatory variables (Step S113) by MPC in cooperation with the terminal device PP.
[0163] <Summary> As described above, according to this embodiment, a computing system 1000 including a plurality of terminal devices 1 and a user terminal 2 is realized. Each of the plurality of terminal devices 1 stores personal information 121 including identification information and attribute information, converts the personal information 121 into non-identifying processed information 122 (step S106), sorts the non-identifying processed information 122 using correspondence information 124 indicating the correspondence relationship between the personal information 121 (step S109), and generates a part of an inference model m for inferring a target variable by collaborative learning (step S110). One of the plurality of terminal devices 1 generates an inference model m from parts of the plurality of inference models m generated by the plurality of terminal devices 1 (step S112), and searches for a set of explanatory variables for which the target variable satisfies the constraint conditions from the inference model m (step S113). The user terminal 2 receives a target variable, explanatory variables, and constraint conditions set by the user (step S103), and displays a set of explanatory variables (step S116). The target variable and the explanatory variables are selected by the user from among the attribute information (step S102).
[0164] For example, when a user who owns personal information 121A analyzes the customer segment that purchases product A only from the personal information 121A owned by the company, since the item of "address" does not exist in the personal information 121A, an analysis considering "address" cannot be performed. By using the computing system 1000, this user can use not only the personal information 121A owned by the company but also the personal information 121B owned by other companies, so that the customer segment that purchases product A can be analyzed considering "address". As a result, it becomes possible to discover a customer segment (for example, segment 2 in FIG. 14) that could not be discovered only from the personal information 121A owned by the company. Also, during the analysis, since the user does not touch the personal information 121B itself owned by other companies, the personal information 121B can be used safely. In addition, since the user can easily use the personal information just by selecting the search conditions and the desired results are displayed, the computing system 1000 enables the user to safely and easily use the personal information owned by multiple owners.
[0165] <Supplementary Note> This embodiment includes the following disclosure.
[0166] (Appendix 1) A computing system comprising a plurality of terminal devices and a user terminal, each of the plurality of terminal devices, stores personal information including identification information and attribute information, converts the personal information into non-identifying processed information, sorts the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information, generates a part of an inference model for inferring a target variable by federated learning, one of the plurality of terminal devices, generates the inference model from a plurality of parts of the inference model generated by the plurality of terminal devices, searches for a set of explanatory variables for which the target variable satisfies the constraint conditions from the inference model, the user terminal, receives the target variable, the explanatory variables, and the constraint conditions set by the user, displays the set of explanatory variables, the target variable and the explanatory variables are selected by the user from among the attribute information Computing system.
[0167] (Appendix 2) The plurality of terminal devices execute the federated learning by bucketing the attribute information The computing system according to Appendix 1.
[0168] (Appendix 3) A part of the inference model includes a bucket threshold The computing system according to Appendix 1.
[0169] (Appendix 4) One of the plurality of terminal devices is a terminal device that does not store the attribute information selected as the target variable. The computing system according to Appendix 1.
[0170] (Appendix 5) The plurality of terminal devices execute federated learning by MPC The computing system according to Appendix 1
[0171] (Appendix 6) The plurality of terminal devices use a part of the inference model to search for a set of explanatory variables for which the target variable satisfies the constraint conditions by MPC The computing system according to Appendix 1
[0172] (Appendix 7) The inference model is a decision tree The computing system according to Appendix 1
[0173] (Appendix 8) The user terminal displays the number of pieces of the non-identified processed information extracted by the set of explanatory variables The computing system according to Appendix 1
[0174] (Appendix 9) The user terminal displays the set of explanatory variables in descending order of the number of pieces of the non-identified processed information extracted by the set of explanatory variables The computing system according to Appendix 1
[0175] (Appendix 10) Each of the plurality of terminal devices hashes the identification information, adds noise, or performs salted hashing The computing system according to Appendix 1
[0176] (Appendix 11) The non-identified processed information is personal information, pseudonymized information, anonymized information, personally related information, or non-personal information The computing system according to Appendix 1
[0177] (Appendix 12) A computing method executed by a computing system including a plurality of terminal devices and a user terminal, Each of the plurality of terminal devices Store personal information including identification information and attribute information, Convert the personal information into non-identifying processed information, Sort the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information, Generate part of an inference model for inferring a target variable through federated learning, One of the plurality of terminal devices, Generate the inference model from parts of the plurality of inference models generated by the plurality of terminal devices, Search for a set of explanatory variables for which the target variable satisfies the constraint conditions from the inference model, The user terminal, Receive the target variable, the explanatory variables, and the constraint conditions set by the user, Display the set of explanatory variables, The target variable and the explanatory variables are selected by the user from among the attribute information Calculation method.
[0178] (Appendix 13) In a computing system including a plurality of terminal devices and a user terminal, Each of the plurality of terminal devices, Store personal information including identification information and attribute information, Convert the personal information into non-identifying processed information, Sort the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information, Generate part of an inference model for inferring a target variable through federated learning, One of the plurality of terminal devices, Generate the inference model from parts of the plurality of inference models generated by the plurality of terminal devices, Search for a set of explanatory variables for which the target variable satisfies the constraint conditions from the inference model, The user terminal, Receive the target variable, the explanatory variables, and the constraint conditions set by the user, Display the set of explanatory variables, The target variable and the explanatory variable are selected by the user from among the attribute information. A program for executing a calculation method.
[0179] The embodiments disclosed this time should be considered as illustrative in all respects and not restrictive. The scope of the present invention is shown not by the above description but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims be included. Also, the present invention is not limited to the above-described embodiments, and various modifications are possible within the scope of the claims, and embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.
Explanation of Reference Numerals
[0180] 1: Terminal device 2: User terminal 11: Communication unit 12: Storage unit 13: Control unit 121: Personal information 122: Non-identifying processed information 123: Search condition information 124: Correspondence information 125: Learning data 126: Inference model information 127: Search result information 131: Acquisition unit 132: Conversion unit 133: Sorting unit 134: Learning unit 135: Generation unit 136: Search unit
Claims
1. A computing system comprising a plurality of terminal devices and a user terminal, wherein: Each of the plurality of terminal devices: Stores personal information including identification information and attribute information; Converts the personal information into non-identifying processed information; Sorts the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information; Generates a part of an inference model for inferring a target variable by federated learning; One of the plurality of terminal devices: Generates the inference model from a plurality of parts of the inference model generated by the plurality of terminal devices; Searches for a set of explanatory variables for which the target variable satisfies the constraint conditions from the inference model; The user terminal: Receives the target variable, the explanatory variables, and the constraint conditions set by the user; Displays the set of explanatory variables; The target variable and the explanatory variables are selected by the user from among the attribute information Computing system.
2. The plurality of terminal devices perform the federated learning by bucketing the attribute information The computing system according to claim 1.
3. A part of the inference model includes a bucket threshold The computing system according to claim 1.
4. One of the plurality of terminal devices is a terminal device that does not store the attribute information selected as the target variable. The computing system according to claim 1.
5. The plurality of terminal devices perform the federated learning by MPC The computing system according to claim 1.
6. The plurality of terminal devices use a part of the inference model to search for a set of explanatory variables for which the target variable satisfies the constraint conditions by MPC The computing system according to claim 1.
7. The inference model is a decision tree The computing system according to claim 1.
8. The user terminal displays the number of non-identifying processed information extracted by the set of explanatory variables The computing system according to claim 1.
9. The user terminal displays the set of explanatory variables in descending order of the number of non-identifying processed information extracted by the set of explanatory variables The computing system according to claim 1.
10. Each of the plurality of terminal devices hashes the identification information, adds noise, or performs salted hashing The computing system according to claim 1.
11. The non-identifying processed information is personal information, pseudonymized information, anonymized information, personally related information, or non-personal information The computing system according to claim 1.
12. A calculation method executed by a calculation system including a plurality of terminal devices and a user terminal, wherein each of the plurality of terminal devices stores personal information including identification information and attribute information, converts the personal information into non-identifying processed information, sorts the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information, generates a part of an inference model for inferring a target variable by federated learning, one of the plurality of terminal devices generates the inference model from a plurality of parts of the inference models generated by the plurality of terminal devices, searches for a set of explanatory variables for which the target variable satisfies a constraint condition from the inference model, wherein the user terminal receives the target variable, the explanatory variables, and the constraint condition set by the user, displays the set of explanatory variables, wherein the target variable and the explanatory variables are selected by the user from among the attribute information Calculation method.
13. In a calculation system including a plurality of terminal devices and a user terminal, wherein each of the plurality of terminal devices stores personal information including identification information and attribute information, converts the personal information into non-identifying processed information, sorts the non-identifying processed information using correspondence information indicating the correspondence relationship between the personal information, generates a part of an inference model for inferring a target variable by federated learning, one of the plurality of terminal devices generates the inference model from a plurality of parts of the inference models generated by the plurality of terminal devices, searches for a set of explanatory variables for which the target variable satisfies a constraint condition from the inference model, wherein the user terminal receives the target variable, the explanatory variables, and the constraint condition set by the user, displays the set of explanatory variables, wherein the target variable and the explanatory variables are selected by the user from among the attribute information A program for executing the calculation method.
Citation Information
Patent Citations
Marketing device, marketing method, program, and recording medium
JP2014048780A