Method of producing ai agent set for supporting business processing and support service providing system
The encryption process with key updates in blockchain systems addresses the challenge of ensuring information authenticity while allowing controlled deletion, enhancing user privacy and compliance with data erasure rights.
Patent Information
- Application Number
- JP2025187140
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-01-06
- Filing Date
- 2025-11-06
- Publication Date
- 2026-02-18
AI Technical Summary
Blockchain-based information recording systems face a dilemma between ensuring the authenticity of recorded information and allowing users to delete their personal information, as erasure is difficult once recorded, violating the right to be forgotten.
Implement an encryption process using a first and second key for recording and decrypting information, with the second key being updated to a different key to render the information undecryptable upon request, and allowing decryption with a first key distribution to authorized parties.
Resolves the trade-off between information authenticity and user deletion rights by enabling secure and controlled access to personal information.
Smart Images

Figure 2026027361000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a processing system and a program for an information recording method that is difficult to tamper with or erase, such as a blockchain. Regarding. [Background technology]
[0002] Blockchain has been widely known as an information recording method that is difficult to tamper with. For example, Patent Document 1 discloses an example of using blockchain to record various information related to cargo transportation. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2018-128723 Summary of the Invention [Problem to be solved by the invention]
[0004] However, such information recorded using blockchain is not only difficult to tamper with, but also difficult to erase (hereinafter referred to as "inerasability"). As a result, once personal information is recorded using blockchain, even if the owner of the personal information wants to erase it, it cannot be erased, which is a drawback in that the right to erase personal information (the so-called right to be forgotten) is violated.
[0005] In other words, there is a drawback in that a dilemma arises between guaranteeing the authenticity of recorded information and guaranteeing the right to delete that information, which are contradictory to each other.
[0006] The present invention was conceived in light of the above circumstances, and its purpose is to resolve the trade-off between guaranteeing the authenticity of recorded information and guaranteeing the right to delete that information. [Means for solving the problem]
[0007] The present invention includes: an encryption unit that performs an encryption process for encrypting information to be recorded; a recording means for recording the information after the encryption process; a decryption means for decrypting the information recorded by the recording means using a first key and a second key to generate plaintext information; a decryption-disabling means for making the information recorded by the recording means into a decryption-disabled state in which the information cannot be decrypted, the decryption means includes second key secret storage means for storing the second key in secret, The decryption disablement means updates the second key held by the second key secret holding means to a different key, thereby making the second key decryption disabled.
[0008] Preferably, the decryption means further includes first key distribution means for distributing the first key to a person who wishes to view the information. Preferably, the information storage device further comprises a search means for searching the information recorded by the recording means without converting the information into plain text.
[0009] More preferably, the information recorded by the recording means includes personal information, The decryption-disabling means makes the personal information of the personal information owner in the decryption-disabled state in response to a request from the personal information owner.
[0010] Another aspect of the present invention is a method for encrypting information to be recorded, comprising: a decryption step of decrypting the information recorded by the recording means for recording the encrypted information using a first key and a second key to generate plaintext information; a step of rendering the information recorded by the recording means in an undecodable state so that the information cannot be decoded; Let the computer run the decrypting step includes a step of keeping the second key secret; The step of making the second key undecryptable is performed by updating the second key held in the holding step to another key. [Effects of the Invention]
[0011] According to the present invention, it is possible to resolve as much as possible the dilemma of the trade-off between guaranteeing the authenticity of recorded information and guaranteeing the right to delete that information. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a system diagram showing the overall configuration of a processing system. [Figure 2] (A) is a diagram explaining the information stored on the HDD of a user terminal that constitutes a node in the blockchain, and (B) is a diagram explaining the information stored in the personal information database of a certified business operator. [Figure 3] (A) is a flowchart showing the main routine program of the public chain user terminal, and (B) is a flowchart showing the subroutine program of personal information recording processing and the flowchart of the certified business operator's server. [Figure 4] 10 is a flowchart showing a subroutine program of smart contract processing executed on a user terminal of a public chain. [Figure 5] 10A is a flowchart showing a main routine program of a user terminal of a private chain, and FIG. 10B is a flowchart showing a subroutine program of personal information search processing. [Figure 6] A flowchart showing a subroutine program of smart contract processing executed on a user terminal of a private chain. [Figure 7] (A) is a continuation of the flowchart showing the subroutine program for smart contract processing executed on a user terminal of a private chain, and (B) is a flowchart showing the subroutine program for machine learning processing executed on a user terminal of a private chain. [Figure 8]This is a flowchart showing a subroutine program of the AI smart contract generation process executed on a user terminal of a private chain. [Figure 9] (A) is a flowchart showing a subroutine program of the simulation learning process executed on a user terminal of a private chain, and (B) is a flowchart showing a subroutine program of the AI smart contract group generation process executed on a user terminal of a private chain. [Figure 10] 10 is a flowchart showing a subroutine program of the smart contract trust undertaking process executed on a user terminal of a private chain. [Figure 11] (A) is a flowchart showing a subroutine program of the reinforcement learning process of the personalized AI smart contract trained model, and (B) is a flowchart showing the main routine program executed by the user terminal of the consortium chain. [Figure 12] (A) is a flowchart showing a subroutine program for IoT sensor data aggregation processing executed by a user terminal of the consortium chain, and (B) is a flowchart showing a subroutine program for simulation processing executed by a user terminal of the consortium chain. [Figure 13] (A) is a flowchart showing a subroutine program for the AI smart contract group generation process executed by a user terminal of the consortium chain, and (B) is a flowchart showing a subroutine program for the smart contract processing executed by a user terminal of the consortium chain. [Figure 14] This is an explanatory diagram of how recorded information can be made inaccessible using blockchain. (A) shows the normal state where it can be viewed, and (B) shows the state where it has been made indecipherable and cannot be viewed. [Figure 15] This is a diagram explaining the information stored in the HDD of a user terminal that constitutes a node of the blockchain. [Figure 16]10 is a flowchart showing the main routine program of a user terminal of a public chain and a user terminal of a private chain. [Figure 17] (A) is a flowchart showing a subroutine program for recording personal information to a blockchain, and (B) is a flowchart showing a subroutine program for making records unreadable. [Figure 18] 10 is a flowchart showing a subroutine program for personal information acquisition processing and personal information provision processing. [Figure 19] (A) is a diagram explaining the information stored on the HDD of a user terminal that constitutes a node in the blockchain, and (B) is a diagram explaining the information stored in the personal information database of a certified business operator. [Figure 20] 10 is a flowchart showing the main routine program of the public chain user terminal, the certified business server, and the private chain user terminal. [Figure 21] 10 is a flowchart showing a subroutine program for recording personal information and hash values in a blockchain. [Figure 22] 10 is a flowchart showing a subroutine program for processing to make a recording unreadable; [Figure 23] 10 is a flowchart showing subroutine programs of personal information provision processing, personal information acquisition processing, and ciphertext transmission processing. [Figure 24] FIG. 1 is a system diagram showing the overall configuration of a processing system. [Figure 25] 10 is a flowchart showing the main routine program of the user terminal of the public chain, the server of the key registration center, and the user terminal of the private chain. [Figure 26] (A) is a flowchart showing the subroutine program for processing personal information recording and key registration in the blockchain, and (B) is a flowchart showing the subroutine program for processing a request to make records unreadable and processing records unreadable. [Figure 27]10 is a flowchart showing subroutine programs of a counterpart common key providing process, a data obtaining process, and a data decrypting process. [Figure 28] FIG. 1 is an explanatory diagram of a mirror world as a simulation environment. [Figure 29] FIG. 1 is a diagram illustrating a specific example of a city digital twin in the mirror world. [Figure 30] 10 is a flowchart of the main routine of the mirror world server and the user terminal. [Figure 31] 10 is a flowchart showing a subroutine program of a personal AI generation and sales process. [Figure 32] 10 is a flowchart showing a subroutine program of a simulation preparation process and a simulation preparation response process. [Figure 33] 10 is a flowchart showing a subroutine program of a simulation process. [Figure 34] (A) is a schematic diagram of the multi-service DAO construction system, and (B) is a flowchart of the main routine between the mirror world server and the user terminal. [Figure 35] 10 is a flowchart showing a subroutine program of a simulation reinforcement learning preparation process and a simulation reinforcement learning preparation response process. [Figure 36] This is an explanatory diagram for registering a multi-service DAO digital twin in the mirror world as a simulation target. [Figure 37] This is a schematic system diagram of simulation reinforcement learning for multi-service DAO. [Figure 38] 10 is a flowchart showing a subroutine program of a DAO agent reinforcement learning process. [Figure 39] (A) is a diagram showing a reward table stored as knowledge by the DAO agent, and (B) is a flowchart showing a subroutine program of the persona agent reinforcement learning process. [Figure 40]10A is a flowchart showing a subroutine program for idea proposal service execution processing, and FIG. 10B is a flowchart showing a subroutine program for improvement proposal service execution processing. [Figure 41] 10A is a flowchart showing a subroutine program for the commercialization service execution process, and FIG. 10B is a flowchart showing a subroutine program for the infringement countermeasure service execution process. [Figure 42] 10A is a flowchart showing a subroutine program for the token purchase service execution process, and FIG. 10B is a diagram explaining the price fluctuations at the fluctuating market price of tokens due to the purchase of tokens by persona agents. [Figure 43] A diagram showing an element integration DAO construction system. [Figure 44] (A) is a flowchart of the main routine between the mirror world server and the user terminal, and (B) is a flowchart showing the subroutine programs of the simulation reinforcement learning preparation processing and the simulation reinforcement learning preparation response processing. [Figure 45] This is an explanatory diagram for registering an element-integrated DAO digital twin in the mirror world as a simulation target. [Figure 46] This is a diagram showing a schematic system illustrating simulation reinforcement learning of element-integrated DAO digital twins. [Figure 47] (A) is a diagram showing the calculation algorithm for performance and distribution rate stored as knowledge by the material procurement element agent, (B) is a diagram showing the calculation algorithm for performance and distribution rate stored as knowledge by the assembly element agent, and (C) is a diagram showing the calculation algorithm for performance and distribution rate stored as knowledge by the advertising element agent. [Figure 48] 10A is a diagram showing a calculation algorithm for performance and distribution rate stored as knowledge by a sales element agent, and FIG. 10B is a diagram showing a remuneration table stored as knowledge by a control agent. [Figure 49]10 is a flowchart showing a subroutine program of a simulation reinforcement learning process. [Figure 50] 10 is a flowchart showing a subroutine program of a control agent reinforcement learning process. [Figure 51] 10 is a flowchart showing a subroutine program of a material procurement element agent reinforcement learning process. [Figure 52] 10A is a flowchart showing a subroutine program for information collection processing by a crawler, and FIG. 10B is a diagram showing various data stored in a material procurement DB. [Figure 53] 10 is a flowchart showing a subroutine program of assembly element agent reinforcement learning processing. [Figure 54] 10 is a flowchart showing a subroutine program of an advertising element agent reinforcement learning process. [Figure 55] 10A is a flowchart showing a subroutine program of information collection processing by a crawler, and FIG. 10B is a diagram showing various data stored in an advertisement DB. [Figure 56] 10 is a flowchart showing a subroutine program of a sales element agent reinforcement learning process. [Figure 57] 10A is a flowchart showing a subroutine program of an information collection process, and FIG. 10B is a diagram showing various data stored in a sales DB. [Figure 58] 10A is a flowchart showing a subroutine program of reinforcement learning processing for a personal AI in charge of material procurement, and FIG. 10B is a flowchart showing a subroutine program of reinforcement learning processing for a personal AI in charge of assembly. [Figure 59] 10A is a flowchart showing a subroutine program of the promotional staff personal AI reinforcement learning process, and FIG. 10B is a flowchart showing a subroutine program of the sales staff personal AI reinforcement learning process. [Figure 60] FIG. 10 is an explanatory diagram showing a program installation method. DETAILED DESCRIPTION OF THE INVENTION
[0013] [First embodiment] A first embodiment of the present invention will be described with reference to Figures 1 to 13. First, referring to the overall system in Figure 1, three types of blockchain networks—private chain 2, consortium chain 3, and public chain 4—are connected to a centralized oracle 21. The public chain 4 is a completely open system, allowing any individual or organization to transact on it. Transactions can be effectively verified by the blockchain. Mining (competition for bookkeeping rights) is also free and open to anyone. The consortium chain 3 is a blockchain that can only be used by partners belonging to an association or union. Individuals (each node) within the consortium chain are designated bookkeepers. Block generation is also predetermined, and other individuals (nodes) can transact but do not have bookkeeping rights. The private chain 2 only records transactions using blockchain technology, and bookkeeping rights are not open but are monopolized by individuals or companies, and only internal transactions are recorded. Polkadot is used to connect blockchains and exchange tokens and data between them. Polkadot is a blockchain designed to connect different blockchains. Blockchains developed using Substrate can be connected to Polkadot, allowing them to exchange tokens and data with other blockchains connected to Polkadot.
[0014] Centralized Oracle 21 is a system that acts as a bridge between blockchain and Internet 1. It is connected to Internet 1, collects various information scattered across the internet, and provides this information to the blockchain's smart contracts.
[0015] Each node 19 of the private chain 2, consortium chain 3, and public chain 4 is composed of a user terminal such as a personal computer (hereinafter referred to as "PC") 16. This PC (hereinafter also referred to as "user terminal") 16 is connected to the Internet 1. The Internet 1 is further connected to a server 20 of a social networking service (SNS) 40 and a server 18 of a blockchain certified business operator 17. The server 18 of the certified business operator 17 may participate in the blockchain as a node 19. In addition, the server of a certification authority that issues digital certificates in a public key infrastructure (PKI) may be connected to the Internet 1.
[0016] The certified business operator 17 receives personal information, issues an electronic ID based on the personal information, and records the hash value of the personal information in the blockchain. The received personal information is stored in a personal information database (hereinafter referred to as "personal information DB") 29. The certified business operator 17 may also participate in the blockchain as a node 19.
[0017] The PC 16 is composed of a CPU (Central Processing Unit) 10 as a control center, a RAM (Random Access Memory) 9 that functions as a work area for the CPU 10, a ROM (Read Only Memory) 11 that stores data and programs, a storage unit such as an HDD (Hard Disk Drive) 12, an input operation unit 7 such as a display and keyboard, a communication unit 5, a display unit 6, an interface 8, a bus 13, and various other hardware. Various servers such as server 20 and server 18 are also composed of hardware similar to that of the PC 16, and therefore will not be illustrated or described again here. A solid state drive (SDD) may be used as the storage unit in addition to or instead of the HDD.
[0018] An IoT (Internet of Things) device 14 and a wireless sensor network 15 are connected to a node 19 of the consortium chain 3. Sensor signals from the IoT device 14 and the wireless sensor network 15 are input to the node 19, and a drive signal for the IoT device 14 is output from the node 19. The IoT device 14 is a variety of sensors, actuators, etc. for IoT.
[0019] A wireless sensor network 15 is a wireless network in which multiple sensor-equipped wireless terminals are dispersed throughout space, enabling them to collaborate and collect information on the environment and physical conditions. For example, sensor devices can be created using energy harvesting, M2M, or batteries. Pressure and gauge sensors can constantly monitor, for example, metal fatigue degradation, and report any changes. Wireless sensor networks are primarily installed on structures such as bridges and tunnels. They typically include multiple sensor nodes and a gateway sensor node. These nodes typically consist of one or more sensors, a wireless chip, a microprocessor, and a power source (e.g., a battery). Wireless sensor networks typically have ad hoc functionality and a routing algorithm for sending data from each node to a central node. In other words, they have the ability to autonomously reconstruct an alternative communication path if a communication failure occurs between nodes. Because nodes work together as a group, they also have elements of distributed processing. Additionally, they can operate for long periods without external power supply, providing power-saving or self-powering capabilities.
[0020] In this embodiment, the IoT device 14 and the wireless sensor network 15 are connected to the consortium chain 3 via the node 19, but one or both of the IoT device 14 and the wireless sensor network 15 may themselves be part of the node 19 of the consortium chain 3 without going through the node 19.
[0021] Next, with reference to Figure 2(A), the information stored in the HDD 12 of the PC 16 will be described. The HDD 12 stores the user's private key SK, public key PK, common key K1, trapdoor common key K2, the user's address in the blockchain, smart contracts, tokens, artificial intelligence (also called "AI (Artificial Intelligence)"), blockchain data, etc. Note that the term "user" is a broad concept that includes not only natural persons but also corporations.
[0022] The private key SK and the public key PK are a key pair used in PKI (Public Key Infrastructure), and data encrypted with the public key PK is decrypted with the private key SK. The private key SK is also used for electronic signatures. The common key K1 is a key used for common key encryption such as DES (Data Encryption Standard) or AES (Advanced Encryption Standard). Data encrypted with the common key K1 is decrypted using the same common key K1. In this embodiment, a different common key is used for each piece of personal information to be encrypted. In the first embodiment, the encrypted personal information E K1 An index for keyword search is provided for the encrypted text (personal information). The index is encrypted with a common key K2. To perform a keyword search, an encrypted search query (called a "trapdoor") is used, in which the keyword (search query) used for the search is encrypted with the common key K2. This common key K2 is stored in the HDD 12 as the trapdoor common key K2.
[0023] A user's address on the blockchain is generated through the following process: 1 Generate a public key using ECDSA from the private key. 2. Pass the public key through the hash function SHA-256 to obtain a hash value. 3 The hash value is then passed through the RIPEMD-160 hash function to obtain a new hash value. 4. Add a prefix of 00 to the beginning of the hash value. 5. Pass it through the hash function SHA-256. 6. Pass it through the SHA-256 hash function again. 7 Add a 4-byte checksum to the end. 8 Encode in Base58 format.
[0024] A smart contract is a computer protocol intended to smoothly verify, check conditions, enforce, execute, and negotiate contracts. A token is a unique currency issued on a blockchain by a company or individual.
[0025] Next, we will explain blockchain data. The data in each block of the blockchain includes the hash value of the previous block, a nonce, and data on multiple transactions (also called transactions). Although not shown in the figure, a timestamp is also embedded in the blockchain. Such a blockchain is generated and added as a new blockchain when each node 19 performs blockchain processing (see S3, S19, S30, S51, S117, S122, S153, etc., described below). Blockchain processing mainly consists of three phases: transaction, propagation, and recording.
[0026] The transaction phase is what is generally called a transaction, and refers to legal acts such as buying and selling, transferring, lending, etc. More specifically, this transaction phase can be divided into three phases: generation → signing → propagation.
[0027] The generation phase involves generating a transaction. For example, Person A decides to lend Person B his or her idle PC resources (computing resources) for 39,005 seconds in exchange for 25.78 tokens, and digitally signs the transaction. This digital signature is created by passing the transaction data through a predetermined hash function to generate a hash value, which is then encrypted using the private key SK of the parties to the transaction (Person A and Person B). A digital public key certificate may also be issued by a certification authority. Figure 2 shows an example of lending PC resources (computing resources), but the lending target is not limited to this. Other possible values include electricity generated by a home or business, a user's specialized knowledge, experience, skills, personal connections (including online personal networks), and credibility.
[0028] The propagation phase involves having other nodes confirm that the transaction has been created and signed correctly. If it is determined that the transaction was not created and signed correctly, the transaction is discarded.
[0029] In the recording phase, if it is confirmed that the transaction has been correctly generated and signed, the miner performs mining to record the transaction. Once it is confirmed that the transaction has been correctly generated and signed, it moves to a place called a mining pool. The miner then selects a transaction to record from the mining pool and mines it.
[0030] Mining is the process of calculating a nonce. A nonce is a value used to adjust the hash value so that a very small hash value with many leading zeros is generated when the block data is passed through a hash function. If a nonce can be calculated that results in a hash value equal to or less than the target value, a new block is generated.
[0031] The transaction data is stored as E, which is the user's personal information encrypted with key K1, as shown in transaction I on the right side of Figure 2. K1The hash value of (personal information), its electronic ID, the index of personal information, and the compensation for providing personal information (provided as 2.4 tokens in Figure 2) are encrypted with key K2. K2 (Provided at index + 2.4 tokens) Specific examples of personal information include vital information such as the user's heart rate, blood pressure, body temperature, and brain waves, behavioral history information such as purchase history and website browsing history, user location information such as GPS, race, creed, social status, medical history, electronic medical record data, ID (identification), and information posted to social media sites.
[0032] Information posted to SNS etc. is information that has already been posted to SNS 25 and stored in server 20 and that has been transferred from server 20 to personal information DB 29 and the blockchain. Specifically, the user encrypts all of their past posted information and stores it in the certified business's personal information DB 29, while also recording its hash value in the blockchain. Thereafter, rather than posting to SNS 25, the user encrypts and stores the posted content in the certified business's personal information DB 29, while also recording its hash value in the blockchain. This enables the user to retrieve their personal information from the SNS etc. business and keep it under their own control.
[0033] The compensation for providing personal information (provided in 2.4 tokens in Figure 2) may be recorded in plain text on the blockchain without encryption. In this case, other users can search the blockchain to find out the compensation without obtaining the encryption key K2. Furthermore, transaction terms such as compensation for providing personal information and compensation for lending PC resources (computing resources) (in Figure 2, PC resources (computing resources) are lent for 39,005 seconds in exchange for receiving 25.78 tokens) may be coded as a smart contract, and transactions (legal acts) may be automated using the smart contract.
[0034] The encrypted personal information itself, which is the subject of the hash value recorded as transaction I, is stored in the personal information DB 29 of the certified business operator 17. Specifically, as shown in FIG. 2(B),K1 (Personal information) is associated with the electronic ID issued and encrypted personal information E K1 (Personal information) is stored in the personal information DB 29.
[0035] The index is the encrypted personal information K1 This is an index for keyword search of (personal information). In this embodiment, a common key encryption method is used in which the index is encrypted with a common key K2, so in order to perform a keyword search, an encrypted search query (called a "trapdoor") is used in which the keyword (search query) used for the search is encrypted with the common key K2. The common key K2 is a different key for each user, but the same key is used for encrypted indexes of the same user. Therefore, for example, if user A performs a transaction to distribute common key K2 to user B using a smart contract described later, user B can use E K2 (search query) can be used to search all of User A's encrypted indexes on the blockchain.
[0036] In addition, searchable encryption such as homomorphic encryption or fully homomorphic encryption, which allows searching ciphertext while it is encrypted, may be used. In this case, personal information may be encrypted using homomorphic encryption or fully homomorphic encryption, and the encrypted personal information may be recorded directly on the blockchain. In addition, the encrypted personal information E K1 (Personal information) may be recorded directly on the blockchain.
[0037] Next, referring to Figure 3(A), we will explain the flowchart of the main routine program of the user terminal of the public chain 19. Step S (hereinafter simply referred to as "S") 1 performs personal information recording processing, S2 performs smart contract processing, and S3 performs blockchain processing.
[0038] Personal information recording processing is a process in which the personal information owner encrypts personal information, registers it with a certified business operator 17, and records the hash value of the encrypted personal information on the blockchain. Smart contract processing is a process in which legal acts such as the conclusion and execution of a contract are automatically performed in accordance with predetermined rules. The specific content of blockchain processing is as described above with reference to Figure 2(A).
[0039] The personal information recording process will be explained with reference to Figure 3 (B). In S5, it is determined whether or not a personal information registration operation has been performed on a user terminal constituting a node 19 of the public chain 19. If not, this personal information recording process returns and moves to smart contract processing in S2. If it is determined that a personal information registration operation has been performed, in S6, the personal information stored in the memory (HDD 12, etc.) of the user terminal is encrypted with key K1 and then digitally signed with private key SK, and the index is also encrypted with key K2 and sent to server 18 of certified business operator 17.
[0040] The server 18 of the certified business 17 that receives the information in S7 issues an electronic ID and stores the received encrypted personal information, E K1 The certified business operator 17 generates a hash value of the electronic ID (personal information). Next, the issued electronic ID is returned to the user terminal (S9). The user terminal that receives the electronic ID stores the electronic ID in memory (HDD 12, etc.). In S10, the certified business operator 17's server 18 performs processing to record the electronic ID, hash value, and encrypted index in the blockchain.
[0041] A flowchart of the subroutine program for smart contract processing shown in S2 is described below. Referring to FIG. 4, in S13, it is determined whether a distribution contract for the common key K2 has been concluded. If not, in S15, it is determined whether a rental contract for the PC resources (computing resources) of the user terminal has been concluded. If not, it is determined whether a contract for the provision of personal information has been concluded. If not, it is determined whether an ordering contract for ordering custom-made products, etc. has been concluded. If not, it is determined whether a sales contract for products, etc. has been concluded in S22. If not, it returns. These determinations are made by the smart contract. For example, if the aforementioned consideration, etc., is coded as a smart contract, the smart contract determines whether the terms of consideration, etc., of both parties are met. If it is determined that they are met, the contract is automatically concluded and executed.
[0042] If it is determined that a distribution contract for the common key K2 has been concluded, control proceeds to S14, and after the common key K2 is sent to the distribution destination, control proceeds to S19. In S19, processing is performed to record the concluded contract as a transaction in the blockchain. If it is determined that a rental contract for PC resources (computing resources) has been concluded, control proceeds to S18, and rental processing for the PC resources (computing resources) is performed. If it is determined that a contract for the provision of personal information has been concluded, control proceeds to S20, and the electronic ID of the personal information to be provided and a signature agreeing to the provision of personal information are returned to the provider, and the common key K1 used to encrypt the personal information to be provided is encrypted with the provider's public key and returned to the provider.
[0043] If it is determined that an order contract has been concluded, control proceeds to S21, where an order processing is performed, and then control proceeds to S19. If it is determined that a sales contract has been concluded, control proceeds to S23, where processing to acquire the purchase item is performed, and then control proceeds to S19.
[0044] Next, a flowchart of the main routine program of a user terminal constituting node 19 of private chain 2 will be described with reference to Figure 5(A). Personal information search processing is performed in S28, smart contract processing is performed in S29, blockchain processing is performed in S30, machine learning processing is performed in S31, AI smart contract generation processing is performed in S32, and smart contract trust undertaking processing is performed in S33. AI smart contracts are a concept that includes both "integrated type" and "linked type." The "integrated type" is a type in which AI and smart contracts are integrated, machine learning is performed based on contract (legal act) data, and the smart contract itself is AI-based. The "linked type" is a type in which AI that has performed machine learning based on contract (legal act) data is linked to a smart contract. In the linked type, the trained AI (hereinafter referred to as the "linked AI") adds, modifies, and updates the smart contract depending on the situation.
[0045] Personal information search processing is a process that searches encrypted indexes recorded in a blockchain using a trapdoor (encrypted search query). Smart contract processing is a process that automatically performs legal actions such as concluding and enforcing contracts according to predetermined rules. The specific content of blockchain processing is as described above with reference to Figure 2(A). Machine learning processing is a process that generates a trained artificial intelligence model by machine learning the personal information of a large number of users as training data. More specifically, it is a process that generates a general trained artificial intelligence model by machine learning a large amount of personal information in a form that does not identify the personal information owner as training data, and then generates a personalized trained model for each personal information owner (each blockchain address) using personal information classified by data that can identify the personal information owner (e.g., an address in the blockchain).
[0046] AI smart contract generation processing is a process that uses machine learning to generate a trained model of a smart contract using artificial intelligence by using personal information related to contracts (legal acts) as learning data. More specifically, it is a process that uses machine learning to generate a general trained model of a smart contract using artificial intelligence by using personal information related to a huge amount of contracts (legal acts) in a form in which the subject of the personal information cannot be identified as learning data, and then generates a personalized AI smart contract trained model that is personalized for each subject of the personal information (for example, each address in the blockchain) using personal information related to the contracts (legal acts) classified by data that can identify the subject of the personal information (for example, addresses in the blockchain).
[0047] Smart contract trust contract processing is a service that is contracted by a principal to automatically perform legal acts such as contract conclusion and execution on behalf of the principal. More specifically, a personalized AI smart contract trained model is generated for the trustor, and the personalized AI smart contract trained model is used to perform legal acts on behalf of the trustor. A reward for the AI is determined based on the results of the execution, and the personalized AI smart contract trained model is further reinforced with the reward.
[0048] Next, a flowchart of the subroutine program for the personal information search process shown in S28 will be described with reference to FIG. 5(B). In S37, it is determined whether the common key K2 is stored, and if not, the process returns. If K2 is stored in S45 (described below), it is determined in S37 that the common key K2 is stored, and control proceeds to S38. In S38, a process is performed to search the encrypted index on the blockchain using a search query (trapdoor) encrypted with K2. As a result of this search, it is determined in S39 whether or not there is any personal information desired to be obtained. If it is determined that there is no personal information desired to be obtained, the process returns. However, if it is determined that there is personal information desired to be obtained, the electronic ID of the desired personal information is stored in S40.
[0049] Next, a flowchart of the subroutine program for smart contract processing shown in S29 will be described with reference to FIGS. 6 and 7A. In S42, it is determined whether there is any personal information desired to be searched. This determination is made, for example, by sequentially negotiating with the smart contract of a personal information owner who has not yet been searched, and determining that the personal information of that personal information owner is the personal information desired to be searched if the conditions are met. If it is determined in S42 that there is no personal information desired to be searched, it is determined in S46 whether there is any memory of personal information desired to be obtained. If it is determined that there is no memory, control proceeds to S55 in FIG. 7A, where it is determined whether a PC resource (computing resource) rental contract has been established. If it is determined that there is no rental contract, it is determined in S56 whether an ordering contract has been established. If it is determined that there is no rental contract, it is determined in S57 whether a sales contract has been established. If it is determined that there is no rental contract, control returns.
[0050] If S42 determines that there is personal information desired to be searched, control proceeds to S43, where the common key K2 is requested from the personal information owner. Specifically, the common key K2 is requested by sending its own address and attribute certificate to the address on the blockchain of the personal information owner of the personal information desired to be searched. S44 determines whether K2 has been returned and waits until it is received. The personal information owner or the personal information owner's smart contract checks the sent attribute certificate and determines whether K2 can be returned, and returns K2 if it is determined that it is allowed to be returned. When K2 is returned from the personal information owner, control proceeds to S45, where the returned K2 is stored, and then control proceeds to S54. In S54, the established contract is stored in the blockchain as a transaction. In this case, a contract is stored in the blockchain to the effect that the common key K2 used to encrypt the index has been distributed from the address of the personal information owner who sent the reply to the address of the user who received the reply.
[0051] If it is determined in S46 that the desired personal information is stored, control proceeds to S47, where a process is performed to request the desired personal information from the personal information owner. Specifically, the user's own address and attribute certificate are sent to the blockchain address of the personal information owner of the desired personal information to request the desired personal information. In the user terminal on the public chain that receives this, the smart contract checks the attribute certificate and determines whether the conditions for providing the personal information are met. If it is determined that the personal information can be provided (YES in S16), the smart contract replies with the electronic ID of the personal information to be provided and a signature agreeing to the provision of the personal information, and also encrypts the common key K1 used to encrypt the personal information to be provided with the public key of the recipient and replies (see S20).
[0052] If a reply is received, it is determined in step S48 that a reply has been received from the personal information owner, and control proceeds to step S49, where the returned encrypted common key, E PK The operation of decrypting (K1) with one's own private key SK, i.e., D SK (E PK Next, in S50, the electronic ID and signature returned from the personal information owner are transmitted to the server 18 of the certified business 17. The server 18 of the certified business 17 receives the electronic ID and signature, checks the transmitted signature, searches the personal information DB 29 (see FIG. 2(B)) based on the received electronic ID, and retrieves the encrypted personal information E stored in association with the received electronic ID. K1 Read (personal information) and reply.
[0053] If the reply is received, the answer is determined as YES in S51, and the control proceeds to S52, where the returned encrypted personal information, E K1 (personal information) by K1 calculated in S49, that is, D K1 (E K1 (Personal information)) to obtain plaintext personal information. The personal information is stored in S53. Then, the process proceeds to S54, where processing is performed to record the contract for providing the personal information as a transaction on the blockchain.
[0054] Next, if a PC resource (computing resource) loan contract is concluded with the public chain user terminal (YES in S15), S55 determines YES, control proceeds to S58, and PC resource (computing resource) borrowing processing is performed. Control then proceeds to S54, and processing is performed to record the loan contract as a transaction in the blockchain. If an order contract is concluded with the public chain user terminal (YES in S17), S56 determines that the order contract has been concluded, control proceeds to S59, and processing is performed to store an order receipt for that order. Control then proceeds to S54, and processing is performed to record the order receipt contract as a transaction in the blockchain.
[0055] If a sales contract is concluded with the public chain user terminal (YES in S15), S57 determines that a PC resource (computing resource) rental contract has been concluded, and control proceeds to S60, where processing to provide the item for sale is performed. Control then proceeds to S54, where processing to record the sales contract as a transaction on the blockchain is performed. Note that in the smart contract processing described above, actions (e.g., S42, S47, S58, S56, S60) may be executed with the consent of the owner of the smart contract rather than being executed solely based on the judgment of the private chain user terminal 16 based on the smart contract (e.g., S42, S46, S55, S56, S57). This owner's consent may also be obtained when executing a contract using a smart contract, as described below.
[0056] Next, a flowchart of the subroutine program for the machine learning process shown in S31 will be described with reference to Figure 7(B). In S63, a process is performed to convert a huge amount of personal information stored in a form that does not identify the owner of the personal information into learning data. Various learning algorithms are available for this machine learning, such as regression and discrimination as supervised learning, model estimation and data mining as unsupervised learning, and reinforcement learning and deep learning as intermediate methods.
[0057] Next, S64 performs machine learning using the training data using the borrowed PC resources (computational resources). For example, in the case of regression as supervised learning, a large data set consisting of input information (vector x) and correct answer information y is used as training data (learning data). In the case of supervised learning, the function ci(x) (ci:x→y) that maps input x to correct answer y is learned, so the trained model includes the function ci. Note that the machine learning performed by the machine learning means 34 is not limited to supervised learning, and may be any type of unsupervised learning such as model estimation or pattern mining (data mining), or semi-supervised learning, reinforcement learning, or deep learning, which are intermediate methods between supervised learning and unsupervised learning.
[0058] A general trained model is generated by machine learning in S64 and stored (S65). This general trained model is a model that uses as training data a huge amount of personal information of many users that is stored in a form that does not allow the personal information owner to be identified, and is an average trained model that can be widely applied to many users.
[0059] Next, S66 determines whether there is an order memory. If not, the process returns. If there is, S67 determines whether the order is from an AI. If not, the process returns. If it is, S68 requests personal information from the orderer's address. If the orderer replies with personal information, S69 determines YES and proceeds to S70. In S70, the general trained model is personalized based on the returned personal information to generate a personalized trained model. This process of personalizing a general trained model to generate a personalized trained model is described in Japanese Patent No. 6432859. The personal information required for personalization is collected using a blockchain, which ensures the anonymity of the blockchain, thereby minimizing privacy issues. Next, S71 sends the personalized trained model to the orderer's address.
[0060] Next, a flowchart of the subroutine program for the AI smart contract generation process shown in S32 will be described with reference to Figure 8. In S80, personal information related to the contract (legal act) is extracted from the vast amount of personal information stored in a form that does not identify the owner of the personal information, and this information is used as training data. Next, in S81, machine learning is performed using the training data using borrowed PC resources (computing resources). The processes of S80 and S81 are similar to those described in S63 and S64 above, and therefore will not be repeated here.
[0061] Next, S82 generates and stores a trained model of the general AI smart contract. Next, S83 executes the simulation learning process. This simulation learning process involves having a large number of AI smart contracts virtually execute (simulate) legal actions, such as contract verification, condition confirmation, execution, implementation, and negotiation, which are normally performed by multiple people. Each AI smart contract is then rewarded based on its performance, allowing reinforcement learning. Because reinforcement learning is performed through computer simulation rather than real-world reinforcement learning, it has the advantage of enabling massive reinforcement learning in a short period of time. Reinforcement learning is a mechanism in which an agent placed in a certain environment acquires a strategy that maximizes the cumulative reward from the initial state to the goal based on the rewards given when selecting an action. In reinforcement learning, learning progresses through interactions between a software agent (hereinafter referred to as "agent"), a type of AI, and the environment. An agent is a type of AI software that communicates with users and other software, behaves autonomously, has a certain degree of judgment ability, and operates continuously. When an agent performs an action a on the environment, the state s of the environment changes and a certain goal state is reached, and a reward r is given to the agent. The agent learns a function that takes the state s as input and outputs an action a with the goal of maximizing this reward r.
[0062] Reinforcement learning progresses over time by repeating the following simple steps: 1 The agent receives an observation o from the environment (or directly the state s of the environment) and returns an action a to the environment based on a policy π. 2 Based on the action a received from the agent and the current state s, the environment changes to the next state s', and based on that transition, it returns to the agent the next observation o' and a single number (scalar quantity) called the reward r, which indicates the quality of the previous action. 3 Time progression: t←t+1 Here, ← represents an assignment operation. For example, an alpha-zero reinforcement learning algorithm may be used as reinforcement learning. Unlike algorithms such as DQN (Deep Q-Network), this alpha-zero reinforcement learning algorithm uses Monte Carlo tree search (MCTS) for search. Values and policies are all predicted by a neural network, and predictions are revised solely based on experience gained from self-play using tree search. Compared to conventional AlphaGo, the value network that predicts values and the policy network that predicts policies are integrated into a single neural network, improving prediction accuracy through multitask learning. Furthermore, improved neural network performance eliminates the need for processor layout (extending the search tree until a reward is received) in tree search, enabling faster search. Furthermore, evolutionary computation, genetic algorithms, and generative adversarial networks may also be used.
[0063] Next, S84 determines whether an order for an AI smart contract has been stored, and if not, returns. If an order has been stored in S59 described above, S84 determines YES and control proceeds to S85, where it determines whether the stored order is for a simulation-trained AI smart contract. If the order is not for a simulation-trained AI smart contract, it is for a personalized AI smart contract trained model, in which case control proceeds to S86, where a process is performed to request personal information related to the contract (legal act) from the purchaser's address. When the purchaser replies with the personal information, S87 determines YES and control proceeds to S88.
[0064] In S88, a process is performed to personalize the trained model of the general AI smart contract based on the personal information related to the returned contract (legal act) to generate a personalized AI smart contract trained model. This process of personalizing the trained model of the general AI smart contract to generate a personalized AI smart contract trained model is described in Japanese Patent No. 6432859. The personalized trained model is then sent to the orderer's address in S89.
[0065] On the other hand, if an order is placed for a simulation-trained AI smart contract, S85 will determine YES, control will proceed to S90, and the simulation-trained AI smart contract will be sent to the orderer's address.
[0066] Next, a flowchart of the simulation learning process subroutine shown in S83 will be described with reference to FIG. 9(A). In S334, it is determined whether a simulation has been input. If not, the program returns. This simulation is input from a user terminal on the private chain 2. For example, it may be a trading simulation in an investment market such as stock trading or futures trading, a company management simulation, or a consumer behavior simulation, assuming that a policy or law that the government is planning to adopt (e.g., a reduced tax rate following a consumption tax increase, a revised immigration law, the UK's withdrawal from the EU (European Union), partial or full adoption of basic income, or an amendment to Article 9 of the Japanese Constitution) is adopted. It may also be a simulation of the promotion of new products (including financial products and life insurance) or new services through various media. If it is determined in S334 that a simulation has been input, control proceeds to S335, where the AI smart contract group generation process is executed.
[0067] The flowchart of the subroutine program for this AI smart contract group generation process is explained based on Figure 9(B). In step S344, borrowed PC resources (computational resources) are used to set personas that match the input simulation. A persona is generally defined as a fictitious person who is a typical target for a company, product, or service. In this embodiment, a persona is defined as a fictitious person who is a typical target for the simulation content. For example, in the case of a consumption behavior simulation under a reduced tax rate following the aforementioned consumption tax increase, personas corresponding to general consumers are set as personas for each group grouped by gender, age, region, annual income, etc.
[0068] The number of personas to be set should be proportional to the number of users belonging to the group. For example, if the age distribution of general consumers is 5% in their teens, 5% in their twenties, 10% in their thirties, 10% in their forties, 20% in their fifties, 20% in their sixties, 20% in their seventies, 5% in their eighties, and 5% in their nineties, the number of personas representing teens should be set as 1, the number of personas representing twenties, 2, the number of personas representing thirties, 2, the number of personas representing forties, 4, the number of personas representing fifties, 4, the number of personas representing sixties, 4, the number of personas representing seventies, 1, the number of personas representing eighties, and 1, the number of personas representing nineties.
[0069] Next, in S345, a process is performed to select a user group belonging to each persona using the borrowed PC resources (computing resources). Next, in S346, a process is performed to group the users belonging to each persona and collect transaction data for each group from the blockchain. For example, in the case of a consumption behavior simulation under a reduced tax rate following the aforementioned consumption tax increase, the users are grouped by gender, age, region, annual income, etc., and transaction data for each group is collected from the blockchain. For example, the selection of the user group and the collection of transaction data for the user group in S345 and S346 can be effectively performed using a database of survey monitor members held by an internet survey research company. The internet survey research company stores the contact information (email address, etc.) of survey monitor members in a database associated with attributes such as gender, age, place of residence, single / married status, occupation, and annual household income, and monitor member data by these attributes is used. Similarly, in steps S145 and S146, S586 and S587, S622 and S623, etc., which will be described later, it is useful to use a database of survey response monitor members held by an internet survey research company. Next, in step S347, machine learning is performed using the borrowed PC resources (computing resources) to learn from the transaction data, generating a trained AI smart contract for each persona. The number of AI smart contracts generated is equal to the number of corresponding personas. This prepares an environment for running the simulation, and the simulation is carried out within that environment.
[0070] Returning to Figure 9(A), each AI smart contract generated as described above executes a contract (legal act) in accordance with action a (S336). This "action a" is the result of reinforcement learning in S338. Next, in S337, the contract concluded between each AI smart contract is recorded on the blockchain.
[0071] Next, in S338, the borrowed PC resources (computing resources) are used to calculate the reward r based on the established contract, and the optimal policy π * A process is performed to determine an action a according to the above. For example, in the case of a consumption behavior simulation under a reduced tax rate following the aforementioned consumption tax increase, a higher reward r is awarded for smaller values of (expenses before the tax increase - expenses after the tax increase). Then, S339 determines whether the simulation has ended. If not, control returns to S336, and reinforcement learning progresses by repeatedly cycling through S337 → S338 → S339 → S336. When the simulation ends, S339 determines YES, and control proceeds to S340, where the AI smart contract that earned the highest reward r is stored and then returned. Note that this is not limited to the AI smart contract that earned the highest reward r; for example, the top 5% of AI smart contracts may be stored. Recording to the blockchain by S337 is not necessarily required. In that case, in the aforementioned collaboration type, "AI smart contract" in S335, S336, S340, and S347 is changed to "collaboration AI." In other words, in the case of a computer simulation, since it does not involve the execution of a contract (legal act) in the real world, if it is not recorded on the blockchain, there is no need to bother using a smart contract, as it is sufficient for each collaboration AI to perform action a and perform reinforcement learning. Once reinforcement learning is completed and at the actual quoting stage, the learned collaboration AI can work with the smart contract to execute the contract and record it on the blockchain.
[0072] Next, a flowchart of the subroutine program for the smart contract trust contract processing shown in S33 will be described with reference to Figure 10. In S94, it is determined whether or not there is a memory of an order for a smart contract trust, and if it is determined that there is no memory of an order, the program returns. The contents of the memory of the order stored in S59 described above are checked, and if it is an order for a smart contract trust, S94 judges YES and control proceeds to S95. In S95, personal information related to the contract (legal act) is requested from the orderer's address, and when the personal information is returned from the orderer, S96 judges YES and control proceeds to S97.
[0073] In S97, a process is performed to personalize the trained model of the general AI smart contract based on the personal information related to the returned contract (legal act) to generate a personalized AI smart contract trained model. This process of personalizing the trained model of the general AI smart contract to generate a personalized AI smart contract trained model is described in Japanese Patent No. 6432859.
[0074] Next, in S98, the orderer's address and the personalized AI smart contract trained model are associated and stored. This stored personalized AI smart contract trained model is used to perform trust contract processing for the orderer (S99). Next, in S100, reinforcement learning processing of the personalized AI smart contract trained model is executed.
[0075] Based on Figure 11(A), we will explain the flowchart of the reinforcement learning processing subroutine program of the personalized AI smart contract trained model shown in S100. In reinforcement learning, when the value of performing action at in state st is defined as Q(st,at), a model-based method can be used to estimate the Q value if knowledge of the environment model is given, i.e., if the state transition probability and probability distribution of the reward are given, but if the environment model is unknown, TD (Temporal Difference) learning is used. First, since the environment must be explored, the ε-greedy method is used. In the early stages of exploration, various actions are tried, and as the exploration settles, the concept of temperature is introduced so that the most optimal actions are selected. With temperature as T, actions are selected according to the probability expressed by the following formula.
[0076] P(a|s)={exp(Q(s,a) / T)} / {Σexp(Q(s,b) / T} (Note that b∈A is written under Σ, but is omitted in the above formula) Here, a is an action, Q(s, a) is the value of performing action a in state s,
[0077] T is called the annealing temperature; a high value results in actions being selected with near-equal probability, while a low value biases the selection toward the optimal one. As learning progresses, decreasing the value of T stabilizes the learning results. This method of estimating the Q value can be applied to all reinforcement learning methods, including those discussed below. While a typical von Neumann-type computer is used to perform machine learning such as reinforcement learning, a neural net processor (NNP) can also be used. The NNP chip is equipped with numerous "artificial neurons" modeled after real neurons, and each neuron connects to each other in a network. A quantum computer employing the "quantum annealing" method can also be used. In particular, using a quantum computer employing the "quantum annealing" method can significantly reduce the time required for optimization calculations in machine learning.
[0078] In S105, an evaluation of the results of the trust contract processing using the personalized AI smart contract trained model is received from the trustor. Next, in S106, a process of calculating a reward r based on the received evaluation is performed. Next, in S107, the borrowed PC resources (computational resources) are used to calculate the optimal policy π by TD learning. * Next, in S108, the contract according to the action a is executed on behalf of the trustor. An evaluation of the result is received from the trustor (S105), and the processes of S106 to S108 are executed.
[0079] Next, a flowchart of the main routine program of a user terminal constituting a node of the consortium chain 3 will be described with reference to Figure 11(B). IoT sensor data aggregation processing is executed in S114, smart contract processing is executed in S115, simulation processing is executed in S116, and blockchain processing is executed in S117. The specific content of the blockchain processing is as described above with reference to Figure 2(A).
[0080] Next, a flowchart of the subroutine program for the IoT sensor data aggregation process shown in S114 will be described with reference to FIG. 12(A). In S120, the IoT sensor data is classified and grouped by type, period, region, etc. This process is performed for not only IoT sensor data but also data from wireless sensor networks. Next, in S121, a process is performed to determine the value of each grouped data. Depending on this determined value, the compensation for providing the data (amount of tokens) is coded as a smart contract for each corresponding data. Next, in S122, a process is performed to record each grouped data in the blockchain.
[0081] Next, a flowchart of the subroutine program for smart contract processing shown in S115 will be described with reference to FIG. 13(B). S150 determines whether a sales contract with a person wishing to acquire data has been concluded. This determination is automatically made as YES in S150 when the conditions for the compensation (amount of tokens) for the provision of data coded as a smart contract are met. If S150 determines NO, control proceeds to S151, where it is determined whether a PC resource rental contract has been concluded, and if not, the process returns.
[0082] If S150 returns YES, control proceeds to S152, where data is sent to the address of the person requesting the acquisition and tokens are obtained in exchange. The established contract is recorded in the blockchain as a transaction in S153. On the other hand, if a PC resource (computing resource) lending contract is established, control proceeds to S154, where the PC resource (computing resource) is borrowed and the contract is recorded in the blockchain as a transaction in S153.
[0083] Next, based on Figure 12(B), we will explain the flowchart of the simulation processing subroutine program shown in S116. This simulation processing involves having a large number of AI smart contracts take over legal actions such as contract verification, condition confirmation, execution, implementation, and negotiation, which are normally performed by multiple people, and virtually executing (simulating) them in a computer to verify the simulation results under certain conditions. Specific examples of the above "certain conditions" include policies and laws that the government is considering adopting (e.g., reduced tax rates following a consumption tax increase, revised immigration law, the UK's withdrawal from the EU (European Union), partial or full adoption of basic income, and amendments to Article 9 of the Japanese Constitution), marketing-related conditions (e.g., pricing and compensation for new products (including financial products and life insurance) and new services, promotional effects through various media, etc.), and investment market-related conditions (e.g., weather conditions in futures trading, monetary tightening policies in the stock market, etc.).
[0084] In S134, it is determined whether a simulation request has been made, and if not, the process returns. If it is determined in S134 that a simulation request has been made, control proceeds to S135, where the AI smart contract group generation process is executed.
[0085] A flowchart of the subroutine program for this AI smart contract group generation process is explained based on Figure 13(A). In S144, borrowed PC resources (computational resources) are used to set a persona group that matches the requested simulation. A persona is generally defined as a fictitious person who is a typical target of a company, product, or service. In this embodiment, a persona is defined as a fictitious person who is a typical target of the simulation content. For example, in the case of the simulation of the reduced tax rate following the consumption tax increase mentioned above, a persona corresponding to an average consumer is set as a persona for each group grouped by gender, age, region, annual income, etc.
[0086] The number of personas to be set should be proportional to the number of users belonging to the group. For example, if the age distribution of the general public is 5% in their teens, 5% in their twenties, 10% in their thirties, 10% in their forties, 20% in their fifties, 20% in their sixties, 20% in their seventies, 5% in their eighties, and 5% in their nineties, the number of personas representing teens should be set as 1, the number of personas representing twenties, 2, the number of personas representing thirties, 2, the number of personas representing forties, 4, the number of personas representing fifties, 4, the number of personas representing sixties, 4, the number of personas representing seventies, 1, the number of personas representing eighties, and 1, the number of personas representing nineties.
[0087] Next, in S145, the borrowed PC resources (computing resources) are used to select a user group belonging to each persona. Next, in S146, the user groups belonging to each persona are grouped and transaction data of the user groups for each group is collected from the blockchain. For example, in the case of a simulation of a reduced tax rate due to the aforementioned consumption tax increase, the users are grouped by gender, age, region, annual income, etc., and transaction data of the user groups for each group is collected from the blockchain. Next, in S147, the borrowed PC resources (computing resources) are used to perform machine learning using the transaction data as learning data to generate a trained AI smart contract for each persona. The same number of AI smart contracts as the set number of corresponding personas are generated. This prepares an environment for executing the simulation, and the simulation is performed within that environment.
[0088] Returning to Figure 12(B), each AI smart contract generated as described above executes a contract (legal act) in accordance with action a (S136). This "action a" is the result of reinforcement learning in S138. Next, in S137, the contract concluded between each AI smart contract is recorded on the blockchain. Furthermore, the changes in the situation as the simulation progresses are recorded on the blockchain. For example, in the case of the simulation of a reduced tax rate following the aforementioned consumption tax increase, how domestic demand and the economy changed as the simulation progressed is recorded on the blockchain.
[0089] Next, in S138, the borrowed PC resources (computational resources) are used to calculate the reward r based on the established contract, and the optimal policy π * Then, a process is performed to obtain an action a according to the above. Then, in S139, it is determined whether the simulation has ended, and if it has not yet ended, control returns to S136, and reinforcement learning progresses by repeatedly circulating S137 → S138 → S139 → S136. When the simulation has ended, a YES determination is made in S139, and control proceeds to S140, where a process is performed to derive the simulation results, and then the process returns.
[0090] As a specific example of a process for deriving simulation results, for example, in a simulation of economic fluctuations following the adoption of a reduced tax rate in response to a consumption tax hike, the process derives how each item in the economic activity index fluctuated as a result of the simulation. In a simulation of a monetary tightening policy in the stock market, the process derives how the stock market fluctuated as a result of the simulation. Alternatively, a simulation optimization method may be employed, in which the specific form of the reduced tax rate following a consumption tax hike (e.g., which items are subject to the reduced tax rate and the reduced tax rate for each subject item) is varied in multiple ways to determine the optimal form of the reduced tax rate. In this case, the optimal form of the reduced tax rate is defined as (tax revenue increase rate (%) + diffusion index (DI) as an economic activity index) / 50, and the simulation result that maximizes this expected value E is obtained. The control parameter in the simulation (the form of the reduced tax rate) is defined as θ, the simulation result is defined as Y(θ), and θ is found at maxE[Y(θ)]. Specific methods for optimizing such simulations include, for example, particle swarm optimization (PSO) as a metaheuristic algorithm, and derivative-free optimization (DFO), which finds the optimal solution when analytical expression of the objective function is difficult or when information on the objective function's derivative is unavailable. Recording on the blockchain according to S137 is not necessarily required. In that case, in the aforementioned collaboration type, "AI smart contract" in S135, S136, and S147 is changed to "collaboration AI." In other words, in the case of a computer simulation, since it does not involve the execution of a contract (legal act) in the real world, if recording on the blockchain is not required, there is no need to bother using a smart contract; instead, each collaboration AI simply executes action a to perform reinforcement learning. [Variations]
[0091] (1) The certified business operator 17 shall store the encrypted personal information of the user E K1(personal information), but instead stores encrypted personal information, E K1 (Personal information) may be recorded directly on the blockchain.
[0092] (2) Transaction data other than personal information, such as transactions C and F in Figure 2, may also be encrypted with key K1, etc., in the same way as personal information, and recorded on the blockchain.
[0093] (3) In the above explanation, the operational processing of the node 19 of each blockchain network 2, 3, and 4 was shown, but the operational processing shown for the node 19 of the private chain 2 may be performed by the node 19 of the other blockchain networks 3 and 4, the operational processing shown for the node 19 of the consortium chain 3 may be performed by the node 19 of the other blockchain networks 2 and 4, and the operational processing shown for the node 19 of the public chain 4 may be performed by the node 19 of the other blockchain networks 2 and 3. This variation may also be applied to the embodiments described below.
[0094] (4) Borrowed PC resources (computing resources) may be used to perform mining (competition for bookkeeping rights) in the blockchain. In this case, the PC resources (computing resources) may be lent out at an hourly rate as in the first embodiment, but a percentage of the profits (tokens, etc.) obtained by miners who succeed in mining (competition for bookkeeping rights) may be distributed (dividends) to the lender of the PC resources (computing resources). The dividend rate (dividend amount) is controlled so as to be proportional to the amount of PC resources (computing resources) lent out (number of PCs lent out x lending time, etc.).
[0095] (5) Some project may be carried out by borrowing (using) the aforementioned PC resources (computing resources) or self-generated electricity to be loaned (provided). Specific examples of projects include research and development (e.g., artificial intelligence development, machine learning, human genome analysis, new product development, new drug development, etc.), exploration and excavation of rare metals, oil, natural gas, marine resources, etc., and space development. In this case, as in the first embodiment, the lender (provider) may receive compensation (tokens, etc.) according to the amount of the loaned (provided) object. Alternatively, a percentage of the profits obtained by the project executor (individual, corporation, or organization) upon the success of the project may be distributed (dividend) to the resource lender (resource provider). The dividend rate (dividend amount) is controlled to be proportional to the amount of the resource loaned (provided).
[0096] Furthermore, resource lenders (resource providers) may acquire the right to receive dividends (hereinafter referred to as "dividend rights") rather than receiving the dividends themselves. These dividend rights may be controlled so that they are acquired by the lender (provider), for example, in the form of tokens issued by the project executor. The lender (provider) may then control the system so that the acquired dividend rights (tokens) can be transferred to others at a price (token) according to the market price at the time. By configuring the system in this way, dividend rights (tokens) can be managed just like stock trading in the secondary market of a stock market.
[0097] (6) In S71, the generated personalized trained model is sent to the address of the client and delivered. In addition to or instead of this, the generated personalized trained model may be utilized to provide personalized services to the client.
[0098] (7) In the above description, smart contracts automate contract verification, condition confirmation, execution, implementation, and negotiation. However, the system may also be controlled to request the user's consent before concluding and executing a contract (a legal act such as a transaction). Furthermore, instead of requesting the user's consent for all contracts (a legal act such as a transaction), the system may determine whether a contract (a legal act such as a transaction) is of predetermined importance, and if it is determined to be an important contract (a legal act such as a transaction), the system may request the user's consent. Furthermore, the system may determine whether a contract (a legal act such as a transaction) is one for which conclusion and execution are urgent, and if it is determined to be a contract (a legal act such as a transaction), the system may control the conclusion and execution of the contract (a legal act such as a transaction) without the user's consent, and then report the result to the user. This variation may also be applied to the embodiments described below.
[0099] (8) The above-mentioned programs that run on the user terminals 16, etc. and various servers that constitute the nodes 19 of each blockchain may be downloaded and installed from a predetermined website, etc., or may be recorded on a recording medium (non-transitory recording medium) such as a CD-ROM 99 and distributed, and those who purchase the CD-ROM 99, etc., may install the programs on the user terminals 16 and various servers (see Figure 60).
[0100] (9) While the above explanation uses a centralized oracle, a decentralized oracle managed across the entire network can also be used. The information collected by multiple oracles distributed across the network is aggregated to extract average information, which is then considered correct and incorporated into the blockchain for use in smart contracts. This is based on the theory proposed by James Surowiecki in his book, "The Common Sense of the Common Sense," which states that "by aggregating information from a group, the group's conclusions can be better than any individual within the group." Oracles that collect information close to the average provide incentives for operating decentralized oracles by offering rewards such as tokens.
[0101] In summary, the system comprises an extraction means for aggregating information collected by a plurality of oracles distributed on a network and extracting average information, an adoption means for adopting the average information extracted by the extraction means, and a reward granting means for granting rewards to the oracles, wherein the plurality of oracles include a first oracle and a second oracle, and the reward granting means grants a larger reward to the second oracle that has collected information closer to the average information than the first oracle. Note that the reward granted to the first oracle may be 0 or may be negative.
[0102] (10) A function may be provided to determine whether or not to establish one or both of the K2 distribution contract in S13 and the personal information provision contract in S16 based on information that identifies the intention of the personal information owner (hereinafter referred to as "intention-specifying information"). Specifically, when the personal information owner wants to receive recommendations for products or services that suit them at a physical store or online shopping mall, the owner reads the specific information on their mobile device (smartphone, IC card, etc.) and enters a PIN indicating that they are willing to establish a contract, thereby notifying the smart contract of the personal information owner of the intention-specifying information consisting of the specific information and the PIN, and the smart contract makes a decision based on the intention-specifying information.
[0103] By doing this, users can use the personal information they have retrieved from SNS or other service providers and placed under their own management for their own purposes, according to their own will. [Second embodiment]
[0104] Next, a second embodiment will be described. This second embodiment responds to the need for a right to delete personal information (the so-called right to be forgotten) by applying encryption technology to personal information recorded using a blockchain. A feature of blockchain is that once recorded, information cannot be tampered with or is extremely difficult to tamper with. As a result, once recorded, it is impossible or extremely difficult to delete the information (hereinafter referred to as "impossibility of deletion"). Meanwhile, the General Data Protection Regulation (GDPR) in Europe requires that a right to delete personal information (the so-called right to be forgotten) be guaranteed, which allows the owner of the personal information to delete the personal information that has been recorded. The GDPR's requirement for the right to delete personal information and the impossibility of deletion in blockchain are in direct conflict, creating a dilemma of contradiction. In other words, this second embodiment solves the dilemma of the requirement to guarantee the right to delete information that one wants to delete and the impossibility of deletion. An overview will be given with reference to FIG. 14.
[0105] FIG. 14(A) shows the normal state where the right to delete is not exercised, and FIG. 14(B) shows the state where the right to delete is exercised and the information is rendered unreadable. Referring to FIG. 14(A), the information holder (also called the information owner) 40 double-encrypts the information using two symmetric keys KA and KB. Expressed as a formula, E KA (E KB (information)). Next, the encrypted information E KA (E KB (information)) is recorded in a blockchain or the like. Note that the other symmetric key KA is stored in a secret state in the user terminal or the like of the information holder 40.
[0106] In this state, when the information requester 41 requests information from the information owner 40, the information owner 40 sends the recorded encrypted information, E KA (E KB (information)) is decrypted with key KA. Expressed as a formula, D KA (E KA (E KB (information))=E KB (information). And this E KBThe information owner 40 transmits (information) and the other common key KB to the information requester 41.
[0107] The information requester 41 who receives them uses the received E KB (Information) is decrypted using the received shared key KB. Expressed as a formula, D KB (E KB (information))=information As a result, the information requester 41 can obtain the plaintext information.
[0108] Next, the state in which information has been rendered unreadable by exercising the right to delete will be described with reference to Fig. 14(B). The information owner 40 updates one of the symmetric keys KA and KB used to encrypt the information to be rendered unreadable, KA, to a random number R (≠KA). Next, when an information requester 41 who has already stored the symmetric key KB requests information from the information owner 40, the information owner 40 will update the encrypted information E KA (E KB (information)) is decrypted with key R (random number). Expressed as a formula, D R (E KA (E KB (information)). And this D R (E KA (E KB The information provider 40 transmits (information) to the information requester 41.
[0109] The information requester 41 who receives it uses the other half of the common key KB that he has already memorized to send it to D R (E KA (E KB (information)) is decrypted. In formula form, D KB (D R (E KA (E KB(Information)))) ≠ Information. In this way, in the decryption-decryption-defective state, even an information requester 41 who already has the other half of the common key KB stored cannot obtain the plaintext information, and the dilemma of the contradiction between the requirement to guarantee the right to delete the information desired and the impossibility of deletion can be resolved. Note that if copy protection is implemented to prevent copying and pasting of information such as personal information recorded in a blockchain, the guarantee of the right to delete information can be made more complete. Note that it is not necessary to be limited to double encryption using two keys KA and KB, and multiple encryption using three or more keys (n keys) can also be used. In this case, the decryption-defective state can be achieved by replacing at least one of the n keys with a random number R.
[0110] The outline of the second embodiment explained above will be explained in more detail. The points in common with the first embodiment will not be explained repeatedly, and differences will be mainly explained. Figure 15 corresponds to Figure 2 in the first embodiment. Referring to a transaction I in the blockchain, in this second embodiment, encrypted personal information E KA (E KB (personal information)) is recorded directly in the block. Therefore, in the second embodiment, the certified business operator 17 is not required. Here, KA and KB are symmetric common keys.
[0111] Next, referring to Figure 16, a flowchart of the main routine of a user terminal constituting node 19 of private chain 2 and a user terminal 16 constituting a node of public chain 4 will be described. This main routine omits the flowchart of the operational processing shown in the first embodiment, and shows only a flowchart of the operational processing that is added to or changed from the operational processing shown in the first embodiment. In the user terminal 16 constituting a node of public chain 4, personal information recording processing to the blockchain is performed in S160, recording indecipherable processing is performed in S161, and personal information provision processing is performed in S162. In the user terminal constituting node 19 of private chain 2, personal information acquisition processing is performed in S170.
[0112] The process of recording personal information to the blockchain is the process of recording personal information to the blockchain. The process of making the record unreadable is the process of exercising the right of deletion to make the information unreadable. The process of providing personal information is the process of a user terminal 16 on the public chain 4 providing personal information to a user terminal on the private chain 2. The process of obtaining personal information is the process of a user terminal on the private chain 2 obtaining personal information from a user terminal 16 on the public chain.
[0113] A flowchart of a subroutine program for processing personal information recording in a blockchain will be described with reference to Figure 17(A). In step S174, processing is performed to generate two random numbers. For example, in the case of DES, two 56-bit random numbers are generated, and these 56-bit random numbers are used as the other symmetric keys KA and KB. In the case of ADS, two 128-bit random numbers are generated, and these 128-bit random numbers are used as the other symmetric keys KA and KB.
[0114] Next, in S177, E KA (E KB (Personal Information)) and E K2 The process of recording (index + compensation for personal information provision) and the ciphertext identifier on the blockchain is performed. This ciphertext identifier is stored in the encrypted personal information E KA (E KB It is an identifier for identifying the user (personal information) and corresponds to the electronic ID in the first embodiment.
[0115] Next, in S183, KA and KB are stored in association with the ciphertext identifier in the HDD 12 of the user terminal 16.
[0116] Next, a flowchart of a subroutine program for the record undeciphering process will be described with reference to FIG. 17(B). In step S190, the ciphertext (for example, E KA (E KBIt is determined whether there is any ciphertext that needs to be made undecipherable among the ciphertexts (personal information), etc. If there is no ciphertext that needs to be made undecipherable, the process returns. However, if there is any ciphertext that needs to be made undecipherable, the process proceeds to S191, where the HDD 12 is searched for the other symmetric key KA that is stored in association with the ciphertext identifier of that ciphertext.
[0117] Next, in S192, a random number R is generated. For example, in the case of DES, a 56-bit random number is generated. In the case of ADS, a 128-bit random number is generated. Next, in S193, it is determined whether the generated random number R=KA. If the generated random number R is the same as the other symmetric key KA stored in HDD 12, control returns to S192, and a random number is generated again. If the determination in S193 is NO, control proceeds to S194, and a process is performed to update the other symmetric key KA stored in HDD 12 to R.
[0118] Next, a flowchart of the subroutine program for the personal information provision processing and personal information acquisition processing will be described with reference to Figure 18. In the user terminal of the private chain 2, it is determined in S198 whether the other half of the common key KB for the personal information desired to be obtained has already been stored. If the other half of the common key KB for the personal information desired to be obtained has already been distributed from the user terminal 16 of the public chain 4 to the user terminal of the private chain 2, it is determined in S198 that it is stored and control proceeds to S203, but if it has not yet been stored, control proceeds to S199.
[0119] In S199, the ciphertext identifier of the personal information desired to be obtained is sent to the user terminal 16 of the public chain 4 to request the other half of the common key KB. The user terminal 16 of the public chain 4 receives this in S200 and determines by a smart contract whether or not to carry out a transaction to provide the personal information specified by the ciphertext identifier (see S16), and if the transaction to provide the personal information is to be carried out, a signature agreeing to the provision of the personal information and the other half of the common key KB corresponding to the ciphertext identifier are returned in S201.
[0120] In S202, the user terminal of the private chain 2 receives the signature and the ciphertext identifier of the personal information desired to be obtained, and in S203, transmits the signature and the ciphertext identifier of the personal information to the user terminal 16 of the public chain 4. In S206, the user terminal 16 of the public chain 4 receives the signature and the ciphertext identifier and determines by a smart contract whether or not to carry out a transaction to provide the personal information specified by the ciphertext identifier (see S16). If a transaction to provide the personal information is to be carried out, in S207, D KA (Encrypted Personal Information) or D R Specifically, if the other symmetric key KA stored in the HDD 12 of the user terminal 16 has already been updated to the random number R, the process returns D R (encrypted personal information) is calculated and returned, but if it has not yet been updated to the random number R, KA The (encrypted personal information) is calculated and returned.
[0121] In the user terminal of the private chain 2 that receives the reply from the user terminal 16 of the public chain 4 in S208, in S209, KB (D KA (Encrypted personal information) = plain text or, D KB (D R (Encrypted personal information) ≠ plaintext. KA If you receive (encrypted personal information), KB (D KA (encrypted personal information))=D KB (D KA (E KA (E KB (Personal information)))) = plaintext is calculated to obtain the plaintext personal information. R (Encrypted Personal Information)) is received, KB (D R (encrypted personal information))=D KB (D R (E KA (E KBThe calculation results in (personal information)))) ≠ plaintext, and the plaintext personal information cannot be obtained. This resolves the conflicting dilemma between the requirement to guarantee the right to delete information that one wants to delete and the impossibility of deleting it. [Variations]
[0122] (1) In the above explanation, encrypted personal information E KA (E KB If personal information (personal information) is recorded directly on the blockchain, and this large amount of encrypted personal information were stored on each node (all nodes in the case of Public Chain 4), each node (user terminal) would be required to have a large amount of storage capacity, which is an inconvenience. To solve this problem, secret sharing technology is applied, which stores divided data on multiple computers. Data is divided into fragments, and each fragmented data is distributed and stored on multiple nodes. The data stored on each node is also stored redundantly (in duplicate). By providing sufficient redundancy, even if some of the fragmented data is lost, it can be restored without any problems, and the blockchain can be made tamper-resistant. Furthermore, the amount of storage space used to store fragmented data can be controlled so that each node can decide for itself, and compensation can be awarded to each node in the form of tokens or other compensation according to the amount of storage space used.
[0123] (2) In the above explanation, encrypted personal information E KA (E KB In the previous version, the encrypted personal information E was recorded directly on the blockchain. However, in the modified examples shown in FIGS. 19 to 23, the encrypted personal information E KA (E KB (Personal information)) is stored in the personal information DB 29 of the certified business operator 17, and the hash value of the encrypted personal information is recorded in the blockchain. KA (E KB Hash value of (personal information) + E K2 (Index + 2.4 token) + ciphertext identifier and digital signature are recorded. In addition, the personal information DB 29 of the certified business 17 shown in FIG. 19(B) stores the encrypted personal information E in association with the ciphertext identifier. KA (E KB(Personal information)) is stored.
[0124] 20, a flowchart of the main routine of the user terminal constituting node 19 of private chain 2 in this modified example, the user terminal 16 constituting node 19 of public chain 4, and the server 18 of the certified business operator 17 will be described. The server 18 of the certified business operator 17 participates in the blockchain as node 19. S215 executes a personal information recording process to the blockchain, S216 executes a recording indecipherable process, S217 executes a personal information provision process, S220 executes a hash value recording process, S221 executes a ciphertext transmission process, and S224 executes a personal information acquisition process.
[0125] The process of recording personal information on the blockchain is carried out by a user terminal 16 constituting a node 19 of the public chain 4 transmitting encrypted personal information E to a server 18 of a certified business operator 17. KA (E KB The hash value recording process is a process of transmitting the encrypted personal information E KA (E KB The server 18 of the certified business operator 17 that receives the encrypted personal information E (personal information) stores it and generates a hash value for it, which is then recorded in the blockchain. KA (E KB The personal information provision process is a process executed by a user terminal 16 constituting a node 19 of the public chain 4 to provide personal information to a user terminal constituting a node 19 of the private chain 2. The personal information acquisition process is a process executed by a user terminal constituting a node 19 of the private chain 2 to acquire personal information. The ciphertext transmission process is a process executed by a server 18 of the certified business operator 17 to transmit encrypted personal information E to a user terminal constituting a node 19 of the private chain 2. KA (E KB This is a process for transmitting encrypted text such as (personal information).
[0126] Details of each process will be explained below based on the flowcharts of each subroutine program, but differences from the second embodiment will be mainly explained.
[0127] Based on FIG. 21, a flowchart of the subroutine program for recording personal information and hash values in the blockchain will be described. KA (E KB (Personal Information)) and E K2 (index + compensation for providing personal information) is transmitted to the server 18 of the certified business 17. The server 18 receives it in S240 and then, in S241, KA (E KB Next, in step S242, a process is performed to generate a hash value of the E KA (E KB (Personal information)) hash value and E K2 A process is performed to record (index + compensation for providing personal information) and the ciphertext identifier on the blockchain.
[0128] Next, in S243, the ciphertext identifier is transmitted to the user terminal 16 of the public chain 4. In the user terminal 16 of the public chain 4 that received it in S232, in S233, the other keys KA and KB are associated with the received ciphertext identifier and stored in the HDD 12. In S244, the server 18 of the certified business operator 17 transmits the E KA (E KB The personal information (personal information) is stored in the personal information DB 29 in association with the ciphertext identifier.
[0129] The record unreadable process shown in FIG. 22 is the same as that already explained in FIG. 17(B) of the second embodiment, and therefore will not be explained again.
[0130] Next, a flowchart of the subroutine programs for the personal information provision processing, personal information acquisition processing, and ciphertext transmission processing will be described with reference to Figure 23. In this modification, when the user terminal of the private chain 2 receives a signature agreeing to the provision of personal information and the other symmetric key KB corresponding to the ciphertext identifier from the user terminal 16 of the public chain 4 (S264), in S265 it transmits the received signature and the ciphertext identifier of the personal information desired to be obtained to the server 18 of the certified business operator 17. The server 18 of the certified business operator 17 receives this in S266, and after verifying the signature, searches the personal information DB 29 for the ciphertext corresponding to the ciphertext identifier and transmits it to the user terminal 16 of the public chain 4 in S267.
[0131] The user terminal 16 of the public chain 4 receives the token in S268 and sends it to the D KA (Encrypted Personal Information) or D R (encrypted personal information)) and returns it to the user terminal of the private chain 2. Specifically, if the other symmetric key KA stored in the HDD 12 of the user terminal 16 has already been updated to R, R (encrypted personal information) is calculated and returned, but if it has not yet been updated to R, KA The (encrypted personal information) is calculated and returned.
[0132] In the user terminal of the private chain 2 that receives the reply from the user terminal 16 of the public chain 4 in S270, in S271, KB (D KA (Encrypted personal information) = plain text or, D KB (D R (Encrypted personal information) ≠ plaintext. KA If you receive (encrypted personal information), KB (D KA (encrypted personal information))=D KB (D KA (E KA (E KB (Personal information)))) = plaintext is calculated to obtain the plaintext personal information.R (Encrypted Personal Information)) is received, KB (D R (encrypted personal information))=D KB (D R (E KA (E KB The calculation results in (personal information)))) ≠ plaintext, and the plaintext personal information cannot be obtained. This resolves the conflicting dilemma between the requirement to guarantee the right to delete information that one wants to delete and the impossibility of deleting it.
[0133] In addition, the server 18 of the certified business operator 17 may be connected to the user terminal of the private chain 2 and the user terminal of the public chain 4 via the Internet 1 without participating in the blockchain as a node 19.
[0134] (3) In the above explanation, the other symmetric key KA is held by the personal information owner (stored in the HDD 12 of the user terminal 16), but instead, the other symmetric key KA may be registered in the key DB 32 of the key registration center 30, which is an example of a predetermined organization (third-party organization). The other symmetric key KA is stored in a secret state in the key DB 32. This modification will be explained with reference to Figs. 24 to 27.
[0135] 24, a server 31 of a key registration center 30 is connected to the Internet 1. A key DB 32 connected to the server 31 stores a ciphertext identifier and a counterpart symmetric key KA associated with each other for each user address, which is each node 19 of the public chain 4. When a user requests record decryption, the counterpart symmetric key KA stored in association with the ciphertext identifier corresponding to the requested record is updated to a random number R. In FIG. 24, the counterpart symmetric key stored in association with the ciphertext identifier 307cd4 at address 0x6079dd has been updated to the random number 1R2, the counterpart symmetric key stored in association with the ciphertext identifier 4arb56 at address 0x6080dd has been updated to the random number 2Rn, and the counterpart symmetric key stored in association with the ciphertext identifier e2c87r at address 0x6978dd has been updated to the random number mR1.
[0136] 25, a flowchart of the main routine between the user terminal 16 of the public chain 4, the server 31 of the key registration center 30, and the user terminal of the private chain 2 will be described. The points in common with the second embodiment will not be repeated, and differences will be mainly described.
[0137] In the user terminal 16 of the public chain 4, a personal information recording process to the blockchain is executed in S468, a recording undecryption request process is executed in S469, and a decryption key provision process is executed in S470. In the server 31 of the key registration center 30, a key registration process is executed in S463, a recording undecryption process is executed in S464, and a data decryption process is executed in S465. In the user terminal of the private chain 2, a data acquisition process is executed in S460.
[0138] Next, a flowchart of the subroutine program for the personal information recording process and key registration process in the blockchain will be described with reference to FIG. 26(A). In the user terminal 16 of the public chain 4, in S479, KA (E KB (Personal Information)) and E K2 The (index + compensation for providing personal information) and the ciphertext identifier are recorded in the blockchain, and in S480, the other symmetric key KA and the ciphertext identifier are transmitted to the key registration center 30.
[0139] The server 31 of the key registration center 30 receives the key in S474 and stores the received other symmetric key KA and the ciphertext identifier in the key DB 32 in association with each other in S475.
[0140] Next, a flowchart of the subroutine program for the record decryption indecryption request process and the record decryption indecryption process will be described with reference to Figure 26 (B). In the user terminal 16 of the public chain 4, in S494, it is determined whether there is a ciphertext to be made indecryptable, and if there is not, the process returns. However, if there is, in S495, a process is performed to transmit the ciphertext identifier for which indecryption is requested to be made indecryptable to the server 31 of the key registration center 30.
[0141] The server 31 of the key registration center 30 receives the ciphertext identifier in S485 and searches the key DB 32 for the other symmetric key KA stored in association with the received ciphertext identifier. Next, a random number R is generated in S487, and it is determined in S488 whether the random number R=KA. If R=KA, a new random number R is generated in S487, and when R≠KA holds, the other symmetric key KA is updated to R in S489.
[0142] Next, a flowchart of the subroutine programs for the counterpart symmetric key provision process, data acquisition process, and data decryption process will be described with reference to Figure 27. In the user terminal of the private chain 2, a process is performed in S503 to transmit the ciphertext identifier of the data to be acquired to the server 31 of the key registration center 30. The server 31 of the key registration center 30 receives this in S504 and performs a process in S505 to search the blockchain for a ciphertext (encrypted personal information, etc.) corresponding to the ciphertext identifier. Next, a process is performed in S506 to search for the counterpart symmetric key KA or R corresponding to the ciphertext identifier.
[0143] Next, by S507, D KA (Encrypted Personal Information) or D R The encrypted personal information is sent back to the user terminal of the private chain 2. Specifically, if the other symmetric key KA stored in the key DB 32 of the key registration center 30 has already been updated to R, the process returns D R (encrypted personal information) is calculated and returned, but if it has not yet been updated to R, KA The (encrypted personal information) is calculated and returned.
[0144] The user terminal of the private chain 2 that receives it in S509 receives the D KA (D KB (Encrypted personal information) = plain text or, D KA (D R (Encrypted personal information) ≠ plaintext. KA If you receive (encrypted personal information), KB (D KA (encrypted personal information))=D KB (D KA (E KA (E KB (Personal information)))) = plaintext is calculated to obtain the plaintext personal information. R (Encrypted Personal Information)) is received, KB (D R (encrypted personal information))=D KB (D R (E KA (E KB The calculation (personal information))))≠plaintext is performed, and the plaintext personal information cannot be obtained. This solves the dilemma of the trade-off between the requirement to guarantee the right to delete information and the impossibility of deletion. Furthermore, since the process of updating the other symmetric key KA to R is performed at the key registration center 30, it is easy to ensure the reliability of the guarantee of the right to delete by updating KA to R. For example, there is an advantage in that the update of KA to R can be easily performed under audit by a designated organization.
[0145] (4) As another way to resolve the conflicting dilemma between the request to guarantee the right to delete information that one wants to delete and the impossibility of deletion, the encrypted personal information stored in the personal information database 29 of the certified business operator 17, E KA (E KB(Personal information) may be deleted at the request of the owner of the personal information. In this case, a contradictory state occurs in which the hash value of the personal information is recorded on the blockchain but the corresponding personal information is not stored in the personal information DB 29. However, if this contradiction can be tolerated, deleting the personal information is also an effective measure.
[0146] (5) The information that guarantees the right to delete is not limited to personal information, but may be any information, such as information posted to social media or blogs (including data on posted photos and videos), notarized documents such as wills or voluntary guardianship contracts, private documents or articles of incorporation of a company, or other information that requires a fixed date. In addition, in the second embodiment, the information that guarantees the right to delete is recorded using a blockchain, but the blockchain is merely an example, and other information may be used for recording.
[0147] (6) The above-mentioned programs that run on the user terminals 16, etc. and various servers that constitute the nodes 19 of each blockchain may be downloaded and installed from a predetermined website, etc., or may be recorded on a recording medium (non-transitory recording medium) such as a CD-ROM 99 and distributed, and those who purchase the CD-ROM 99, etc., may install the programs on the user terminals 16 and various servers (see Figure 60).
[0148] (7) In the above explanation, plaintext is obtained by encrypting once with the symmetric key KA and then encrypting it again with the symmetric key KB twice, and then decrypting once with the symmetric key KB and then decrypting it again with the symmetric key KA twice. However, this is not limited to this, and encryption using the symmetric key KA or KB may be performed multiple times, and decryption using the symmetric key KA or KB may be performed multiple times. Furthermore, the number of symmetric keys KA and KB is not limited to two, and three or more symmetric keys may be used.
[0149] Furthermore, the exclusive OR of the two symmetric keys KA and KB is calculated to generate a single key K (KA(+)KB=K), and personal information is encrypted with this key K (E K The personal information owner receives a request from the personal information requester and sends the encrypted personal information to the server 31 of the key registration center 30, and the personal information requester sends the distributed shared key KB to the server 31 of the key registration center 30. The server 31 of the key registration center 30 performs an exclusive OR operation between the received shared key KB and the shared key KA registered in the key DB 32 to generate one key K (KA(+)KB=K), and sends the received encrypted personal information (E K (Personal information)) is decrypted with the key K to plain text (D K (E K (Personal information) = plain text, and the plain text personal information may be sent to the requester of the personal information. The (+) above is a symbolic representation of exclusive OR.
[0150] Note that exclusive OR is merely an example, and any algorithm may be used as long as it generates one key K from the other symmetric keys KA and KB.
[0151] Furthermore, the above method of generating key K using an additive group such as exclusive OR has the advantage of maintaining security by periodically updating the symmetric keys KA and KB. For example, if one symmetric key KA is updated to KC, the other symmetric key KB becomes K(+)KC, which can be calculated by calculation. Updating the symmetric keys KA and KB in this manner not only prevents leakage of the symmetric key but also prevents a personal information requester who has once distributed the symmetric key KB from decrypting the encrypted personal information on the blockchain again, thereby preventing access. By updating the symmetric key KA in this way every time the symmetric key KB is distributed, even if the recipient of the symmetric key KB sells it to someone else, it becomes impossible for the recipient to decrypt the encrypted personal information on the blockchain. In other words, the distributed symmetric key KB can be a one-time key that can be used only once.
[0152] Furthermore, the encryption method is not limited to a common key, and a public key encryption method such as RSA or elliptic curve cryptography may also be used.
[0153] In addition, in order to realize the above key update, an encryption algorithm that satisfies the following conditions may be adopted. Let M be the plaintext, C be the ciphertext, and KA, KB, KC, and KD be the encryption keys. E KA (E KB (M))=E KC (E KD (M))=C An algorithm in which the formula holds.
[0154] If such an algorithm is a symmetric key encryption algorithm, when the symmetric keys KA and KB are updated to KC and KD, the plaintext M can be obtained by decrypting the ciphertext C recorded in the blockchain with the symmetric keys KC and KD. On the other hand, in the case of a public key encryption algorithm, when the private keys KA and KB are updated to KC and KD, the plaintext M can be obtained by decrypting the ciphertext C recorded in the blockchain with the public keys PKC and PKD that correspond to the private keys KC and KD. (8) In the above explanation, the information holder provides the information requester with encrypted personal information (D KA (Encrypted Personal Information) or D R The information requester transmits the encrypted personal information (D) and the other half of the common key KB (S201, S207), and the information requester himself decrypts the encrypted personal information using the other half of the common key KB to make it plain text (S209). However, the decryption using the other half of the common key KB may be performed by a third party (a designated service organization). In this case, the information holder transmits the encrypted personal information (D KA (Encrypted Personal Information) or D R The encrypted personal information and the other half of the common key KB are sent to a third party (a designated service provider), which then decrypts the information and sends it to the information requester. [Features of the disclosure]
[0155] Next, the features of the disclosure of the above-described embodiment will be listed below. (Feature 1) [Technical field]
[0156] Feature 1 relates to a processing system and program for an information recording method that is difficult to tamper with or erase, such as a blockchain. [Background technology]
[0157] Blockchain has long been known as an information recording method that is difficult to tamper with or erase. For example, JP 2018-128723 A discloses a system that uses blockchain to record various information related to cargo transportation. [Summary of Feature 1] [Problem that Feature 1 aims to solve]
[0158] However, such information recorded using blockchain is not only difficult to tamper with, but also difficult to erase (hereinafter referred to as "inerasability"). As a result, once personal information is recorded using blockchain, even if the owner of the personal information wants to erase it, it cannot be erased, which is a drawback in that the right to erase personal information (the so-called right to be forgotten) is violated.
[0159] In other words, there is a drawback in that a dilemma arises between guaranteeing the authenticity of recorded information and guaranteeing the right to delete that information, which are contradictory to each other.
[0160] Feature 1 was devised in light of this situation, and its purpose is to resolve the trade-off between ensuring the authenticity of recorded information and ensuring the right to delete that information. [Means for solving the problem]
[0161] The subject of Feature 1 can be expressed as, for example, the following items: (Item 1) Encryption means (for example, S174, S177, S228, S231, or S478, S479) for encrypting information to be recorded (for example, personal information); A recording means for recording information after the encryption process (for example, S177 and a block chain, or S231, S240, S242, S244, a block chain and personal information DB29, or S479 and a block chain); a decryption means (for example, S201, S202, S207 to S209, S263 to S271, or S500 to S510) that decrypts the information recorded by the recording means using a first key and a second key to generate plaintext information; a decryption-disabling means (for example, S191 to S194, or S250 to S254, or S494, S495, S485 to S489) for making the information recorded by the recording means undecryptable, the decryption means includes a second key secret storage means (e.g., S194, S233, or S475) that stores the second key (e.g., the other symmetric key KA) in secret, The decryption-disabling means updates the second key held by the second key secret holding means to another key (e.g., a random number R) to make the second key decryption-disabled (e.g., S190 to S194, or S250 to S254, or S494, S495, S485 to S489), processing system.
[0162] (Item 2) The processing system described in item 1, wherein the decryption means further includes a first key distribution means (e.g., S200, S201, or S2562, S263, or S500, S501) that distributes the first key (e.g., the other half of the common key KB) to a person who wishes to view the information. (Item 3) 3. The processing system according to item 1 or 2, further comprising a search means (for example, S37 to S40, S42 to S45) for searching the information recorded by the recording means without converting it into plain text.
[0163] (Item 4) the information recorded by the recording means includes personal information; The processing system described in any of items 1 to 3, wherein the decryption-disabling means renders the personal information of the personal information owner in the decryption-disabled state in response to a request from the personal information owner (for example, S190 to S194, or S250 to S254, or S494, S495, S485 to S489).
[0164] (Item 5) a step of performing an encryption process for encrypting information to be recorded (for example, personal information) (for example, S174, S177, S228, S231, or S478, S479); A decryption step (e.g., S201, S202, S207 to S209, S263 to S271, or S500 to S510) of decrypting the information recorded by a recording means (e.g., S177 and the block chain, or S231, S240, S242, S244, the block chain and personal information DB29, or S479 and the block chain) using a first key and a second key to generate plaintext information; a step of making the information recorded by the recording means undecodable (for example, S191 to S194, or S250 to S254, or S494, S495, S485 to S489); Let the computer run the decryption step includes a step (e.g., S194, S233, or S475) of secretly storing the second key (e.g., the other symmetric key KA), The step of making the second key undecryptable is to make the second key undecryptable by updating the second key held in the holding step to another key (for example, a random number R) (for example, S190 to S194, or S250 to S254, or S494, S495, S485 to S489).
[0165] (Effect of Feature 1) According to Feature 1, it is possible to resolve as much as possible the dilemma of the trade-off between guaranteeing the authenticity of recorded information and guaranteeing the right to delete that information. (Feature 2) [Technical field]
[0166] Feature 2 relates to smart contracts used in blockchains, for example. [Background technology]
[0167] A smart contract is a computer protocol intended for smooth verification, condition confirmation, execution, implementation, and negotiation of a contract, and has been used in blockchains for some time. Smart contracts have long been known as a way to automate contracts, transactions, etc. (e.g., Patent No. 6403177). [Summary of Feature 2] [Problem that Feature 2 aims to solve]
[0168] In the field of smart contracts, there is a demand for advanced smart contracts that can execute legal acts such as various contracts or transactions, such as sales contracts and loan agreements, on behalf of the users themselves.
[0169] The purpose of Feature 2, which was devised in light of the above situation, is to provide an advanced smart contract that can perform legal acts on behalf of the user. [Means for solving the problem] The subject of Feature 2 can be expressed as, for example, the following items: (Item 1) A machine learning means (e.g., S80 to S82) that inputs information about legal acts performed by multiple natural persons or legal entities as data for machine learning and generates a general model; A personalization means (e.g., S86 to S88 or S94 to S98) for personalizing the general model into a model suitable for the user, the personalization being based on information about legal acts performed by the user; A computer system comprising: a smart contract generation means (e.g., S86 to S88, or S94 to S98) that generates a smart contract using the personalized model to perform legal acts on behalf of the user.
[0170] (Item 2) A personalization means (e.g., S80 to S82, S86 to S88, or S94 to S98) for personalizing a general model generated by inputting information on legal acts performed by multiple natural persons or legal entities as data for machine learning into a model suitable for a user, the personalization being based on information on the legal acts performed by the user; A computer system comprising: a smart contract generation means (e.g., S86 to S88, or S94 to S98) that generates a smart contract using the personalized model to perform legal acts on behalf of the user.
[0171] (Item 3) A personalization means (e.g., S80 to S82, S86 to S88, or S94 to S98) for personalizing a general model generated by inputting information on legal acts performed by multiple natural persons or legal entities as data for machine learning into a model suitable for a user, the personalization being based on information on the legal acts performed by the user; A computer system comprising: a service providing means (e.g., S99) that provides a service that performs legal acts on behalf of the user using the personalized model as a smart contract.
[0172] (Item 4) The computer system described in item 3 further comprises a reinforcement learning means (e.g., S105 to S108) that learns a strategy for maximizing the accumulation of rewards by giving rewards for legal acts performed in connection with the provision of services by the service providing means (e.g., S99) to the model that performed the service.
[0173] (Item 5) A computer system that performs reinforcement learning by simulating within a computer a predetermined theme (for example, a trading simulation in an investment market such as stock trading or futures trading, a company management simulation, or a consumer behavior simulation, assuming that a policy or law that the government is planning to adopt (for example, a reduced tax rate following a consumption tax increase, a revised immigration law, the UK's withdrawal from the EU (European Union), partial or full adoption of basic income, or amendment to Article 9 of the Japanese Constitution) is adopted), A selection means (e.g., S344, S345) for selecting a group of users belonging to a plurality of personas that match the theme of the simulation; a collection means (e.g., S346) for grouping the user group selected by the selection means into the plurality of personas and collecting information on legal acts performed by the user group for each group; A generation means (e.g., S347) that performs machine learning using the collected information on legal acts as training data to generate a group of trained smart contract models for each persona; and simulation means (e.g., S336 to S339) for executing a simulation in a computer of legal acts being performed between the generated trained smart contract models; The computer system, wherein the simulation means includes a reinforcement learning means (e.g., S336, S338) that learns a strategy for the trained smart contract model to maximize the accumulation of rewards for executed legal acts by giving the trained smart contract model rewards for the executed legal acts.
[0174] (Item 6) A computer system that performs a simulation in a computer to progress reinforcement learning, A computer system comprising a reinforcement learning means (e.g., S336, S338) that performs a simulation reinforcement learning process in which a simulation of legal acts being performed between a group of trained smart contract models generated by machine learning is executed within a computer, and a reward for the executed legal act is given to the trained smart contract model, so that the trained smart contract model learns a strategy for maximizing the accumulation of the reward.
[0175] (Item 7) Item 7. The computer system of item 6, further comprising a selection means (e.g., S340) for selecting a trained smart contract model to be actually used from the group of trained smart contract models based on the performance of the reinforcement learning results by the reinforcement learning means.
[0176] (Note) The "machine learning data" for generating the general model and the "machine learning data" for personalization need only include "information about legal acts," and may also include information other than "information about legal acts" (e.g., website access history, GPS location information, etc.). The "smart contract generation means" also encompasses cases where an artificial intelligence such as a personal assistant trained by machine learning using information about legal acts is made to act as a smart contract.
[0177] (Effect of Feature 2) Feature 2 makes it possible to provide advanced smart contracts that can perform various legal acts on behalf of the user.
[0178] (Feature 3) [Technical field] Feature 3 relates to a computer system that sets conditions such as policies and laws that the government is planning to adopt (e.g., reduced tax rates in conjunction with a consumption tax increase, revised Immigration Control Act, the UK's withdrawal from the EU (European Union), partial or full adoption of basic income, amendment of Article 9 of the Japanese Constitution, etc.), marketing-related conditions (e.g., setting prices and compensation for new products (including financial products and life insurance) and new services, promotional effects through various media, etc.), and investment market-related conditions (e.g., weather conditions in futures trading, monetary tightening policies in the stock market, etc.), and then performs simulations within the computer under those conditions to predict in advance what the simulation results will be. [Background technology]
[0179] As a computer system of this type, a new accounting method has been proposed that establishes a statement of changes in net assets to clarify future liabilities that citizens will have to bear and future available resources, thereby supporting policy-level decision-making, and clarifies asset fluctuations resulting from policy decisions for the fiscal year in question, while also enabling a simulation of future burdens on citizens (for example, JP 2006-155233). [Summary of Feature 3] [Problem that Feature 3 aims to solve]
[0180] In the field of such simulations, there is a demand for a computer system that can run simulations within a computer that faithfully mimic the activities of natural persons and corporations in the real world, and derive simulation results that are as close to the real world as possible.
[0181] The purpose of Feature 3, which was devised in light of the above situation, is to enable simulations that faithfully mimic the activities of natural persons and corporations in the real world. [Means for solving the problem]
[0182] The subject of Feature 3 can be expressed, for example, as the following items: (Item 1) A computer system that performs simulations within a computer under predetermined conditions (such as policies and laws that the government intends to adopt (e.g., reduced tax rates following a consumption tax increase, revised immigration laws, the UK's withdrawal from the EU (European Union), partial or full adoption of basic income, amendments to Article 9 of the Japanese Constitution, etc.), marketing-related conditions (e.g., setting prices and fees for new products (including financial products and life insurance) and new services, promotional effects by various media, etc.), investment market-related conditions (e.g., weather conditions in futures trading, monetary tightening policies in the stock market, etc.)), A selection means (e.g., S144, S145) for selecting a group of users belonging to a plurality of personas that match the conditions of the simulation; a collection means (e.g., S146) for grouping the user group selected by the selection means into the plurality of personas and collecting information on legal acts performed by the user group for each group; A generation means (e.g., S147) that performs machine learning using the collected information on legal acts as training data to generate a group of trained smart contract models for each persona; A simulation means (e.g., S136 to S139) that executes a simulation in a computer of performing legal acts between the generated trained smart contract models; a derivation means (e.g., S140) for deriving a result of the simulation by the simulation means, The computer system, wherein the simulation means includes a reinforcement learning means (e.g., S136, S138) that learns a strategy for the trained smart contract model to maximize the accumulation of rewards by giving the trained smart contract model rewards for executed legal acts.
[0183] (Item 2) A computer system for performing a simulation within a computer, a reinforcement learning means (e.g., S136, S138) that executes a simulation in a computer in which a legal act is performed between a group of trained smart contract models generated by machine learning (e.g., S144 to S146), and advances reinforcement learning in which the trained smart contract model learns a strategy to maximize the accumulation of rewards by providing the trained smart contract model with a reward for the performed legal act; A computer system comprising: a derivation means (e.g., S140) that derives the results of a simulation in which legal acts are performed between a group of trained smart contract models in which reinforcement learning by the reinforcement learning means has progressed. (Effect of Feature 3)
[0184] Feature 3 makes it possible to simulate as faithfully as possible the activities of natural persons and corporations in the real world. [Third embodiment]
[0185] Next, a third embodiment will be described. This third embodiment relates to a system that performs simulations in a simulation environment within a mirror world (cyberspace) that is made up of a digital twin of the real world, thereby deriving, for example, an optimal solution that predicts the future, deriving an optimal solution for incentive design in a DAO (Decentralized Autonomous Organization), or performing AI machine learning (for example, reinforcement learning).
[0186] Digital twins are digital representations of real-world entities and systems. A mirror world is a mirror image world composed of digital twins in which all information from the physical (real-world) world, such as countries, cities, societies, local governments, companies, and other organizations and people, is digitized. Specifically, a person's digital twin is composed of an assistant AI (hereinafter referred to as "personal AI") that has undergone machine learning (e.g., agent-based reinforcement learning) to learn knowledge from the person's life log, such as their behavior (both real and virtual), and assist them in choosing the best course of action. This reinforcement learning is multi-agent reinforcement learning, in which multiple personal AIs cooperate to learn reinforcement. A digital twin of a real-world organization, such as a company, is constructed using the personal AI of the people who make up that organization; a digital twin of that local government is constructed using the personal AI of the people who make up that local government; a digital twin of a city is constructed using the personal AI of the people who make up that city; and a digital twin of a country is constructed using the personal AI of the people who make up that country.
[0187] In this third embodiment, a mechanism is provided in which real-world organizations and people, such as nations, cities, societies, local governments, and companies, in the real world take the initiative to participate and cooperate in the construction of a mirror world as a simulation environment. Specifically, by conducting various simulations in the mirror world, the personal AI of the digital twin participating in the simulation is subjected to machine learning (e.g., reinforcement learning), and the more highly trained personal AI is returned (feedback) to the real world. Using the benefits of this as an incentive, real-world organizations and people, such as nations, cities, societies, local governments, and companies, are encouraged to take the initiative to participate and cooperate in the construction of the mirror world.
[0188] 28, mirror world data is stored in a data center 45 in which multiple mirror world servers (including storage servers) 46 are installed. The hardware configuration of the mirror world server 46 is similar to the hardware configuration of the user terminal 16 shown in FIG. 1, and therefore its illustration and description will not be repeated here. The entire mirror world 51, which is composed of digital twins (real country digital twins (e.g., Japan digital twin 53), city digital twins 54, society, local governments, organizations such as companies, people, and the Earth digital twin 52) in which all information about real countries (e.g., Japan 49), cities 50, societies, local governments, organizations such as companies, people, and the Earth 48 in the real world 47 has been digitized, is stored in the data center 45 as digital data.
[0189] The data center 45 uses this mirror world 51 as a simulation environment to perform simulations, e.g., to derive optimal solutions that predict the future through simulation optimization. Examples of simulations include trading simulations in investment markets such as stock and futures trading, company management simulations, or consumer behavior simulations, assuming that the aforementioned policies and laws that the government is considering adopting (e.g., reduced tax rates following a consumption tax increase, revised immigration laws, the UK's withdrawal from the EU (European Union), partial or full adoption of basic income, and amendments to Article 9 of the Japanese Constitution) are adopted. Simulations may also be used to simulate the promotion of new products (including financial products and life insurance) or new services through various media. The optimal solutions derived through simulation optimization are fed back (returned) to the real world, providing the benefits of the optimal solutions to the real world. Furthermore, trained personal AIs that have undergone machine learning (e.g., reinforcement learning) through simulations can be returned to the real world, enabling more advanced trained personal AIs to perform tasks.
[0190] The personal AI and the smart contract cooperate to form the aforementioned cooperative AI smart contract. This data center 45 is connected to the Internet 1 shown in Figures 1 and 24. In Figure 28, various blockchains 2, 3, and 4, SNS 19, key registration center 30, etc. are not shown.
[0191] FIG. 29 shows a specific example of a city digital twin 54 in the mirror world 51. Within a city 50 in the real world 47, there is ABC Corporation 56, a person named Taro 55, and Taro's family 56. The corresponding city digital twin 54 also includes ABC Corporation 59, Taro's digital twin (Taro's personal AI) 57, and Taro's family digital twin 58. City digital twin data consisting of this data is stored in the mirror world server 46. If there are changes to various objects in the real world 49, such as ABC Corporation 56, the person named Taro 55, and Taro's family 56 (for example, personnel transfers, employment, or retirement at the company, or marriage or childbirth for people), the corresponding digital twins are updated with the changes. This city digital twin data is stored in the data center 45 for each city to become the data for the digital twin 53 of the nation of Japan 49, and the city digital twin data for each country is stored in the data center 45 for each city to become the digital twin data for each country. All of this digital twin data combine to become the data for the digital twin 52 of Earth 48.
[0192] As a specific example, the mirror world server 46 stores, as Taro's digital twin (Taro's personal AI) 57, the name: Taro, AI identification number: 82km9, personal AI data, and Taro's personal data (e.g., life log, profile, preference data, electronic medical record data, vital data, etc.). As Taro's family digital twin 58, the names: Taro, Sakura, Shiro, family composition: husband, wife, eldest son, AI identification numbers: 82km9, 11zk9, gf43y. As ABC Co., Ltd. digital twin 59, the names: Taro, Hanako...Saburo, position: representative director, managing director, department head...regular employee, AI identification numbers: 82km9, ba935, 2es14,...9w1c2 are stored.
[0193] Flowcharts of the main routine programs of the user terminal 16 and the mirror world server 46 will be described with reference to Figures 30 to 33. Referring to Figure 30A, the CPU 10 of the user terminal 16 executes a member registration request process S555 for requesting registration for participation in the mirror world 51 as a simulation environment, a simulation preparation response process S556, and a simulation response process S557. The CPU 10 of the mirror world server 46 executes a member registration process 550, a simulation preparation process S551, and a simulation process S552.
[0194] The member registration process and the member registration request process are explained based on FIG. 30B. Both processes are for registering members who wish to participate as digital twins in a simulation using the mirror world 51 as a simulation environment. In the member registration request process, the CPU 10 of the user terminal 16 determines whether to apply for registration in S560. If it determines not to apply for registration, the member registration process ends and returns. If it determines to apply for registration, in S561, the CPU 10 transmits the specified information required for the registration application to the mirror world server 46. If there is a person who does not have a personal AI, the CPU 10 also transmits this fact and the person's blockchain address to the mirror world server 46. Specifically, the specified information required for the registration application includes, for a person's digital twin, the AI identification number and personal AI data of the person's personal AI; for a family's digital twin, the names, family composition, and respective AI identification numbers of the family members; and for a company's digital twin, the names, job titles, and respective AI identification numbers of the employees.
[0195] The CPU 10 of the mirror world server 46 receives this information in S565 and determines in S566 whether or not the user already owns a personal AI. If the information sent in S561 includes information indicating that the user does not own a personal AI, control proceeds to S567, where a process for generating and selling a personal AI is carried out. However, if the information does not include information indicating that the user does not own a personal AI, control proceeds to S568, where predetermined information including the AI identification number sent in S562 is registered in the mirror world 51.
[0196] The personal AI generation and sales process shown in S567 will be explained with reference to Figure 31. In S573, the CPU 10 of the mirror world server 46 collects from the blockchain the transaction data and posted data on SNS etc. recorded in the blockchain address received in S565 (the blockchain address of a user who does not have a personal AI). Next, in S574, machine learning is performed using the transaction data and posted data on SNS etc. as training data to generate a trained personal AI. Next, in S575, the trained personal AI is sold to the corresponding user.
[0197] The simulation preparation process shown in S551 and the simulation preparation response process shown in S555 are explained with reference to FIG. 32. In S577, the CPU 10 of the mirror world server 46 determines whether a simulation request has been received. If it determines that a simulation request has not been received, this simulation preparation process ends and returns. If it determines that a simulation request has been received, control proceeds to S578, where processing is performed to identify personal AIs and digital twins that match the requested simulation. For example, in the case of a consumption behavior simulation under a reduced tax rate following the aforementioned consumption tax increase, the personal AIs corresponding to general consumers are identified in proportions according to demographic statistics such as gender, age, region, and annual income, as well as manufacturer digital twins and retailer digital twins of consumer goods that are subject to the reduced tax rate. Next, processing is performed to request consent for the simulation from the identified personal AIs and digital twins. Specifically, the contents of the simulation are sent to the user terminals 16 of each user group corresponding to the identified personal AIs and digital twins, and a request is made as to whether or not they consent.
[0198] The CPU 10 of each user terminal 16 of the identified personal AI group and the user group corresponding to the digital twin receives the transmitted simulation content in S580, and determines in S581 whether or not the user agrees to participate as a member in the execution of the simulation. This determination may be made by the personal AI, or by the user himself / herself. If it is determined that the user does not agree, the simulation preparation response process ends and returns, but if it is determined that the user agrees, in S582, a response indicating that the user agrees is sent to the mirror world server 46.
[0199] The CPU 10 of the mirror world server 46 receives the request in S583 and determines whether the necessary amount of consent has been obtained to execute the requested simulation. If it is determined that consent has been obtained, in S584, the AIs and digital twins for which consent has been obtained are copied and registered in the mirror world 51 as simulation targets. The registered state is shown in FIG. 29, as described above.
[0200] On the other hand, if it is determined that consent has not been obtained from the necessary personal AIs and digital twins, control proceeds to S585, where a process is performed to set up personas (including personas corresponding to the digital twins of manufacturers and retailers) that match the missing personal AIs and digital twins, and in S586, a user group (including users working for manufacturers and retailers) belonging to each persona is selected, and in S587, the user groups belonging to each persona are grouped and transaction data (including transaction data for manufacturers and retailers) of the user groups is collected from the blockchain for each group, and in S588, machine learning is performed using the transaction data as learning data to generate and replenish trained personal AIs and digital twins for each persona, and then control proceeds to S584. S585 to S588 are the same processes as S344 to S347 in Figure 9(B), and detailed explanations will not be repeated here.
[0201] Next, specific control of the simulation processing shown in S552 and the simulation response processing shown in S557 will be described with reference to Figure 33. S593 to S595 are similar to the processing of S336, S338, and S339 in Figure 9 and S136, S138, and S139 in Figure 12, and detailed description will be omitted here. In S596, the simulation results are notified to the person who requested the simulation. Specifically, the simulation results are transmitted to the user terminal 16 of the person who requested the simulation. Next, in S597, each personal AI used in the simulation (including personal AIs engaged in digital twins of a company organization, etc.) is transmitted to the user terminal 16 of each owner.
[0202] The CPU 10 of the user terminal 16 receives the personal AI in S598 and determines whether to delete the received personal AI in S599. The received personal AI is a trained AI that has undergone reinforcement learning (machine learning) by participating in a simulation, and as a result, its performance has been improved and it is capable of performing advanced task processing. However, depending on the content of the simulation, the AI may have undergone reinforcement learning (machine learning) that the user does not want. In such cases, the CPU 10 determines YES in S599 and deletes the received personal AI in S601. On the other hand, if the simulation is as desired by the user and the received personal AI is determined to have undergone desirable reinforcement learning (machine learning), control proceeds to S600, where the received trained personal AI is overwritten and saved. As a result, the user has the advantage of being able to obtain a personal AI with improved performance that has undergone desirable reinforcement learning (machine learning). This advantage can be used as an incentive to encourage organizations and individuals in the real world, such as nations, cities, societies, local governments, and companies, to proactively participate and cooperate in building the mirror world. By overwriting and saving the personal AI using S600, the data is updated in the mirror world server 46 to the digital twin of the new personal AI after the overwriting and to the digital twin of the organization consisting of the new personal AI (see Figure 29). Note that instead of overwriting and saving, both the existing personal AI and the trained personal AI may be stored and used as needed.
[0203] Next, we will explain a system that derives optimal solutions for incentive design in a DAO by conducting simulations in a mirror world simulation environment, with reference to Figures 34 to 59. Figure 34(A) is a schematic diagram of a multi-service DAO construction system. For example, Bitcoin is a type of DAO, but nodes (miners) perform only one type of service: mining (competing for bookkeeping rights). Blocks are added by providing incentives to those who successfully mine Bitcoin, allowing the Bitcoin system to continue autonomously. In contrast, a DAO that provides multiple services is called a multi-service DAO. For example, in the case of a DAO for a company, there are multiple services such as material procurement, assembly, advertising, and sales. Determining the optimal incentive design, in what proportion and to what extent, rewards should be distributed to nodes that perform these multiple services, is a difficult problem. We will now explain a system that derives optimal solutions for incentive design in such a multi-service DAO.
[0204] 34(A), mirror world server 46, which stores multi-service DAO data, stores data on DAO agent 61, persona agent group 62 performing service 1, persona agent group 63 performing service 2, ..., persona agent group 64 performing service n. Furthermore, mirror world server 46 also stores the types of rewards r1, r2, ..., rn given to each persona agent group as a result of reinforcement learning.
[0205] The terminal 16 of the multi-service DAO builder downloads and installs the DAO agent 61, the necessary persona agents, and the types of rewards r1, r2, ..., rn from the mirror world server 46. The terminals 16 are the nodes 19 that make up the public chain 4. A digital twin 66 of the multi-service DAO 65 operated on the public chain 4 composed of these nodes 19 undergoes simulation reinforcement learning within the mirror world 51 to derive an optimal solution for incentive design in the multi-service DAO. This optimal solution for incentive design is applied to an actual multi-service DAO 65 in the real world 47, resulting in the creation of a multi-service DAO 65 with an optimal incentive design. This simulation reinforcement learning within the mirror world 51 is executed on the mirror world server 46. The control of this process is described below.
[0206] Referring to FIG. 34(B), the CPU 10 of the terminal 16 performs a simulation reinforcement learning preparation response process in S606, and performs a simulation reinforcement learning response process in S607.
[0207] The CPU 10 of the mirror world server 46 performs simulation reinforcement learning preparation processing in S611, performs simulation reinforcement learning processing in S612, and performs DAO agent reinforcement learning processing in S613.
[0208] Specific control of the simulation reinforcement learning preparation process shown in S611 and the simulation reinforcement learning preparation response process shown in S606 will be described with reference to Figure 35. In the simulation reinforcement learning preparation response process, the CPU 10 of the terminal 16 determines in S615 whether to request simulation reinforcement learning. If not, this simulation learning preparation response process ends and returns. If a request is made, control proceeds to S616, where multi-service DAO data is sent to the mirror world server 46 to make the request. This multi-service DAO data includes the type of service. For example, in the case of the innovation induction DAO described below in Figures 36 to 42, there are five types of services: idea generation, improvement generation, commercialization, infringement detection, and token purchase, and these services are sent.
[0209] The CPU 10 of the mirror world server 46 receives the request in S620 and sets persona groups that match each of the multi-services in S621. For example, in the case of the above-mentioned innovation-inducing DAO, the persona group for idea generation and improvement generation could be people who often come up with inventions, the persona group for commercialization could be people who are interested in commercialization, the persona group for infringement detection could be people who are knowledgeable about patent law and copyright law, and the persona group for token purchase could be people who are interested in investing.
[0210] Next, in S622, a user group belonging to each persona is selected. For example, in the case of the above-mentioned innovation-inducing DAO, possible users include a user group listed as an inventor of a patent application as a user group belonging to the idea generation and improvement generation persona group, a user group of company managers as a user group belonging to the commercialization persona group, a user group of patent attorneys and lawyers as a user group belonging to the infringement detection persona group, and a user group who has purchased virtual currency such as Bitcoin as a user group belonging to the token purchase persona group.
[0211] Next, in S623, users belonging to each persona are grouped and transaction data of the users for each group is collected from the blockchain. In S624, machine learning is performed using the transaction data as learning data to generate a group of trained persona agents for each persona. Both of these controls are similar to the processes in S346 and S347 in Figure 9(B) described above, and detailed explanations will not be repeated here. Next, in S625, the persona agents are deployed in the multi-service DAO 65 to generate a multi-service DAO digital twin 66, which is registered in the mirror world 51 as a simulation target. This state is shown in Figure 36.
[0212] Referring to Figure 36, a digital twin 66 of a multi-service DAO 65 consisting of a public chain is constructed in the mirror world 51. A group of persona agents is deployed to the multi-service DAO digital twin 66 in S625, with one persona agent deployed to each node. The identification numbers of these persona agents are classified by service (idea generation, improvement generation, commercialization, infringement detection, token purchase) and stored in the mirror world server 46. For example, kc29m,1w13a,...9nad8 is stored in the mirror world server 46 as the identification number of the persona agent for the idea generation service. This multi-service DAO 65 is the innovation-inducing DAO described above, and the multi-service DAO will be explained below using the innovation-inducing DAO as an example.
[0213] Furthermore, the mirror world server 46 also stores a DAO agent that provides rewards (incentives) for the actions of each persona agent. This DAO agent uses reinforcement learning (machine learning) to determine the distribution rate and amount of rewards to be provided, thereby deriving an optimal solution for incentive design. An outline of this reinforcement learning (machine learning) system is shown in Figure 37.
[0214] Referring to FIG. 37, when a group of persona agents submits an idea to Origin as actions a11, a12, ..., a1n, the environmental state S1 is input to the DAO agent 61, and rewards r11, r12, ..., r1n are given to the group of persona agents 67 that performed the idea generation service. This idea generation service is a broad concept that encompasses the creation of dreams, ideas, business plans, technical concepts, copyrighted works, etc. The environmental state S1 is also given to the group of persona agents 67 that performed the idea generation service. This environmental state S1 is, for example, the content of the Origin idea submission, the floating market price of token A1 given as a reward to each persona agent 68 that performed the idea generation service, etc.
[0215] When the persona agent group 68 performs actions a21, a22, ... a2, such as posting an improvement proposal, on the origin idea, the environmental state S2 is input to the DAO agent 61, and rewards r21, r22, ... r2n are given to the persona agent group 68 that performed the improvement proposal service. The environmental state S1 is also given to the persona agent group 68 that performed the improvement proposal service. This environmental state S2 is, for example, the content of the improvement proposal post, the number of "Likes" given to the improvement proposal post, etc. The entity that can give this "Like" is limited to, for example, only those (persona agent group 71) who purchased the token A1 given as a reward for posting the origin idea. The reason for limiting the entities that can give "Likes" to stakeholders (stakeholders) in this way is to prevent fraudulent activity. This is to prevent fraudulent activities such as a person who posted an improvement proposal (persona agent group 68) colluding with many other people (persona agents) to get many "Likes" if the number of people who can give "Likes" were to be unlimited. For the same reason, the people who can give "Likes" to persona agent group 69 who provided commercialization services and persona agent group 70 who provided infringement remediation services are also limited to only those who purchased token A1 (persona agent group 71).
[0216] When the persona agent group 69 that has performed commercialization services for the idea of the origin performs actions a31, a32, ..., a3n, the environmental state S3 is input to the DAO agent 61, and rewards r31, r32, ..., r3n are given to the persona agent group 69 that performed the commercialization services. Examples of actions a31, a32, ..., a3n of the persona agent group 69 include posting a business plan, posting information about the progress of commercialization, posting information about the status of actual commercialization, and posting information about the amount of revenue from the commercialized business. The environmental state S3 is also given to the persona agent group 69 that performed the commercialization services. This environmental state S3 is, for example, the number of "likes" given to the posting of the business plan, the posting of information about the progress of commercialization, the posting of the amount of revenue from the commercialized business, etc.
[0217] When the persona agent group 70 that performed infringement remediation services for the idea of the origin performs actions a41, a42, ..., a4n, the environmental state S4 is input to the DAO agent 61, and rewards r41, r42, ..., r4n are given to the persona agent group 70 that performed those infringement remediation services. Examples of the actions a31, a32, ..., a3n of the persona agent group 70 include posting a report of infringement discovery, posting a report of infringement remediation, and posting a license negotiation report. Furthermore, the actions a31, a32, ..., a3n of the persona agent group 70 may include posting a report of patent application and a report of patent rights, which are prerequisites for these services. The environmental state S4 is also given to the persona agent group 70 that performed the infringement remediation services. This environmental state S4 may be, for example, the number of "likes" given to the report of infringement discovery, posting a report of infringement remediation, posting a report of license negotiation, etc.
[0218] When the persona agents perform actions a51, a52, ..., a5n to purchase the token A1 granted in exchange for the origin's idea proposal service, the environmental state S5 is input to the DAO agent 61, and rewards r51, r52, ..., r5n are given to the persona agents 71 who performed the token purchase service. The environmental state S5 is also given to the persona agents 71 who performed the token purchase service. This environmental state S5 may be, for example, the number of tokens purchased (or the purchase amount). The persona agents 71 purchase the token A1 by consuming virtual currency (e.g., Ethereum's ETH, etc.). The purchased tokens can be converted (cash) into virtual currency according to the price at the floating exchange rate, and the virtual currency can be converted (cash) into legal tender such as yen or dollars according to the price at the floating exchange rate.
[0219] The rewards r1 to r5 to be granted to each persona agent group are determined by the DAO agent 61 based on the reward table (see Figure 39(A)). For persona agents who performed idea generation services, r1 = A1 + B1 · b + G1 · g; for persona agents who performed improvement services, r2 = A2 · e + B2 · b + G2 · g; for persona agents who performed commercialization services, r3 = A3 · e + B3 · b; for persona agents who performed infringement response services, r4 = A4 · e + B4 · b + G4 · g; and for persona agents who performed token purchase services, r5 = B5 · b + G5 · g.
[0220] Here, A2 to A4, B1 to B5, G1, G2, G4, and G5 are coefficients that the DAO agent 61 converges to the optimal value through reinforcement learning. A1 is tokens, g is license revenue, e is the number of "likes," and b is commercialization revenue.
[0221] Only the license income g or commercialization income b generated after each persona agent group 68 to 71 performs its service is considered as rewards r2 to r5. This is to prevent fraudulent acts such as providing improvement services or token purchase services to an origin idea that has already generated license income g or commercialization income b.
[0222] In addition, a group of persona agents who have performed improvement services, commercialization services, or infringement countermeasure services may also perform token purchase services. Furthermore, a group of persona agents 67 who have performed idea proposal services may also perform improvement services, commercialization services, or infringement countermeasure services.
[0223] The details of the DAO agent reinforcement learning process shown in S613 will be explained with reference to Figure 38. In this process, the DAO agent 61 performs reinforcement learning on its own to optimize rewards r1 to r5. In S630, the DAO agent 61 determines whether or not it has received each action a of the persona agent group. If not, the process proceeds to S632, but if it determines that it has been received, the control proceeds to S631, where the received actions a are stored.
[0224] In S632, it is determined whether a "Like" has been given, and if not, the process proceeds to S634, but if it is determined that a Like has been given, in S633, a Like e is stored for each persona agent. In S634, it is determined whether there has been commercialization profit, and if not, the process proceeds to S636, but if it is determined that there has been commercialization profit, in S635, commercialization profit b is stored. In S636, it is determined whether there has been licensing profit, and if not, the process proceeds to S638, but if it is determined that there has been licensing profit, in S637, the licensing profit g is stored.
[0225] In S638, it is determined whether or not it is time to calculate the reward. If it is not, the process proceeds to S640. If it is determined that it is, in S639, the reward table (Figure 39(A)) is referenced to calculate each of the rewards r1 to r5 and grant them to the corresponding persona agent. In S640, it is determined whether or not it is time to update the learning. If it is determined that it is time to update the learning, in S641, the total granted price TT of the tokens A1 granted as a reward and the current total price TB of the granted tokens in the floating market price are calculated. In S642, the reward R for the DAO agent is calculated from the value of TB / TT. For example, the value of TB / TT at the previous learning update is compared with the value of TB / TT at the current learning update. If the value of TB / TT at the current learning update is larger, a larger reward R is assigned; if it is smaller, a smaller reward R is assigned. As a result, the reward R that the DAO agent 61 can obtain increases if the total price TB of the token in the floating market price rises, and decreases if the total price TB of the token in the floating market price falls.
[0226] Next, in S643, the optimal policy π is obtained by TD learning based on the reward R. * A process is performed to obtain actions A1-A4, B1-B5, G1, G2, G4, and G5 according to the above formula, and in S644, A1-A4, B1-B5, G1, G2, G4, and G5 in the reward table are updated to the obtained actions A1-A4, B1-B5, G1, G2, G4, and G5. As a result, the DAO agent 61 learns optimal actions A1-A4, B1-B5, G1, G2, G4, and G5 for increasing the total price TB of the token in the floating market. Note that this learning goal is merely an example, and other learning goals may include increasing the number of original idea submissions, increasing the total number of submissions of original ideas and improvement ideas, increasing the number of commercializations, and increasing the total commercialization revenue.
[0227] Next, in S645, it is determined whether the reinforcement learning is complete, and if it is not yet complete, the process returns. If it is determined that the reinforcement learning is complete, in S646, the trained multi-service DAO is sent to the requester of the simulation reinforcement learning.
[0228] The client of the simulation reinforcement learning can operate a trained multi-service DAO (innovation-inducing DAO) 65 with optimized incentive design in the real world 47. As a result, in this multi-service DAO (innovation-inducing DAO) 65, the "persona agent groups 67-71" shown in Figure 37 become actual users, and optimally designed rewards (incentives) are distributed to the user groups performing each service by the trained DAO agent 61. During actual operation in this real world 47, the services provided by each post and the token purchase and sale transactions are recorded on the blockchain with timestamps. As a result, the blockchain acts as a notary public for the origin idea post and the improvement proposal post, making it easier to apply exceptions to lack of novelty (Article 30 of the Patent Act) and to implement measures against misappropriated applications (Articles 49, Paragraph 1, Line 7, 74, and 123, Paragraph 1, Item 2 of the Patent Act).
[0229] Furthermore, even when the multi-service DAO (innovation-inducing DAO) 65 is operated in the real world 47, the DAO agent 61 may continue machine learning (reinforcement learning) to achieve an even more optimal incentive design that matches the actual operating situation. Note that the trained persona agent groups 67-71 (trained persona agent groups according to Figures 40(A)(B) and 41(A)(B)) may also be included in the multi-service DAO (innovation-inducing DAO) 65 and sent to the requester of the simulation reinforcement learning, and each persona agent group 67-71 may function as an advisor to the user group performing each service. Furthermore, when the multi-service DAO (innovation-inducing DAO) 65 is operated in the real world 47, it may be a mixed-type multi-service DAO 65 in which each service is executed by both a user group and each persona agent group 67-71, or it may be a persona agent-operated multi-service DAO (innovation-inducing DAO) 65 in which each service is executed only by each persona agent group 67-71. Note that this innovation-inducing DAO is not limited to being generated through the simulation reinforcement learning in the mirror world described above, but may also be generated artificially by other methods, for example, based on an artificial design, and may not be limited to a DAO but may be an organization with a specific administrator or entity (for example, a regular corporation, etc.).
[0230] Next, the main routine of the reinforcement learning process performed by the persona agent will be described with reference to Figure 39(B). In S648, an idea proposal service execution process is performed, in S649, an improvement service execution process is performed, in S650, a commercialization service execution process is performed, in S651, an infringement response service execution process is performed, and in S652, a token purchase service execution process is performed.
[0231] Details of the idea generation service execution process shown in S648 will be explained with reference to Figure 40(A). In S655, it is determined whether or not to generate an idea, and if not, the process returns. If it is determined that an idea should be generated, in S656, the process of creating an idea is performed. This idea generation uses, for example, an AI called DABUS. For example, persona agent 67 and DABUS work together to generate the idea. In S657, the content of the idea generation post is generated, and in S658, the idea generation post action a1i is executed.
[0232] In S659, it is determined whether or not the reward r1i has been received from the DAO agent 61, and if not, the process returns. If it is determined that the reward r1i has been received, in S660, the optimal policy π * If the reward r1i received is satisfactory, the action a will continue to repeatedly come up with ideas, but if the reward r1i is not satisfactory, the action a will choose another action (for example, improvement service, commercialization service, infringement response service, token purchase service, or no service at all).
[0233] The details of the improvement service execution process shown in S649 will be explained with reference to Figure 40(B). In S664, it is determined whether or not to post an improvement plan, and if not, the process returns. If it is determined that an improvement plan should be posted, the process of creating an improvement plan is performed in S665. The improvement plan is created using, for example, an AI called DABUS. For example, the persona agent 68 and DABUS work together to create the improvement plan. In S666, the content of the improvement plan post is generated, and in S667, the improvement plan post action a2i is executed.
[0234] In S668, it is determined whether or not the reward r2i has been received from the DAO agent 61, and if not, the process returns. If it is determined that the reward r2i has been received, in S669, the optimal policy π *If the reward r2i received is satisfactory, the user will continue to submit improvement proposals, but if the reward r2i is not satisfactory, the user will choose another action (for example, idea generation service, commercialization service, infringement response service, token purchase service, or no service at all).
[0235] The details of the commercialization service execution process shown in S650 will be explained with reference to Figure 41 (A). In S674, it is determined whether or not to commercialize, and if not, the process returns. If it is determined to commercialize, a business plan is generated in S675, an act of posting the business plan a3i is executed in S676, the commercialization service is executed in S677, and an act of posting the progress status a3i is executed in S678. This act of posting the progress status a3i also includes posting the profits earned from the commercialization, as described above.
[0236] In S679, it is determined whether or not a reward r3i has been received from the DAO agent 61, and if not, the process returns. If it is determined that the reward r3i has been received, in S680, the optimal policy π * If the received reward r3i is satisfactory, the action a will continue to perform the commercialization service, but if the reward r3i is not satisfactory, the action a will select another action (for example, idea generation service, improvement service, infringement response service, token purchase service, or no service at all).
[0237] The details of the infringement remediation service execution process shown in S651 will be explained with reference to Figure 41(B). In S684, it is determined whether or not to execute the infringement remediation service. If not, the process returns. However, if it is determined that the infringement service should be executed, in S685 an investigation into infringement is conducted, and in S686 it is determined whether or not an infringement has been found. Note that, as mentioned above, a patent application and the act of obtaining a patent may be performed before the infringement investigation is conducted. If no infringement is found, the process returns. However, if it is determined that an infringement has been found, in S687 a warning letter is generated to the suspected infringer, in S688 an act of posting the warning letter a4i is executed, in S689 an infringement remediation action a4i, such as negotiation with the suspected infringer, is executed, and in S690 an act of posting the execution status a4i is executed.
[0238] Next, in S691, it is determined whether or not a reward r4i has been received from the DAO agent 61, and if not, the process returns. If it is determined that the reward r4i has been received, in S692, the optimal policy π * If the reward r4i received is satisfactory, the action a will continue to perform the commercialization service, but if the reward r4i is not satisfactory, the action a will select another action (for example, idea generation service, improvement service, commercialization service, token purchase service, or no service at all).
[0239] Next, the details of the token purchase service execution process shown in S652 will be explained with reference to Figure 42 (A). In S969, it is determined whether or not to purchase tokens, and if not, the process returns. If it is determined that a purchase is to be made, in S697, a token purchase action a5i is executed. Next, in S698, it is determined whether or not a reward r5i has been received from the DAO agent 61, and if not, the process returns. If it is determined that a reward r5i has been received, in S699, the optimal policy π is calculated by TD learning based on the received reward r5i. *If the remuneration r5i received is satisfactory, the action a will continue to be a commercialization service, but if the remuneration r5i is not satisfactory, the action a will choose another action (for example, an idea generation service, an improvement service, a commercialization service, an infringement response service, or no service at all).
[0240] Based on FIG. 42(B), the price fluctuations of tokens 72 at the fluctuating market price due to the purchase of tokens by persona agent group 71 will be explained. Persona agent 67, who provided idea generation services, was granted 50 tokens (market capitalization: 50,000 yen) 72 as reward A1, and persona agent 71a purchased a portion of those tokens 72 (10 tokens) by paying virtual currency equivalent to 10,000 yen. Next, persona agent 71b purchased the 10 tokens by paying virtual currency equivalent to 15,000 yen. As a result, the value of the 10 tokens rises to 15,000 yen. Persona agent 71c then purchased them by paying virtual currency equivalent to 20,000 yen. As a result, the value of the 10 tokens rises to 20,000 yen. Persona agent 71d then purchased them by paying virtual currency equivalent to 25,000 yen. As a result, the value of the 10 tokens rises to 25,000 yen. Persona agent 71e then purchased them by paying virtual currency equivalent to 30,000 yen. As a result, the value of 10 tokens will rise to 30,000 yen.
[0241] As a result, Persona Agent 67's 40 tokens (market capitalization: 40,000 yen) will rise to a market capitalization of 120,000 yen. These tokens rise in value in direct proportion to expected value, so the more popular the origin idea, the more popular (more likes) improvement proposals are posted, the more popular (more likes) commercializations are posted, and the more popular (more likes) infringement countermeasures are posted.
[0242] When the multi-service DAO 65 is actually operated in the real world 47, the "persona agent groups 67-71" in Figure 37 will become user groups in the real world, as mentioned above. In this case, not only the tokens 72 awarded to the group of users who provided idea-generating services, but also the tokens of the users who provided various services (hereinafter referred to as "My Tokens") may be traded. Other users who view the posted service content may purchase the poster's My Tokens in anticipation of the poster's performance, causing the price of the My Tokens to rise in the floating market. In this case, a portion of the poster's income may be distributed to My Token purchasers in proportion to the amount purchased. These My Tokens may be issued within the multi-service DAO 65, or they may be linked to My Tokens issued to users by a professional company that issues and distributes My Tokens. Users of the multi-service DAO 65 may then trade the My Tokens issued by the professional company. VALU, Inc. is currently one such professional company that issues and distributes My Tokens.
[0243] Next, a system in which multiple functional elements cooperate to build a single, well-coordinated DAO will be described with reference to Figures 43 to 59. Such a DAO will be referred to as an "element-integrated DAO" hereinafter.
[0244] Referring to FIG. 43, this element integration DAO allows existing corporate organizations, etc., in the real world 47 to be easily constructed using DAOs, with element DAOs generated and prepared in advance for each of the functional elements. In other words, modularized element DAOs are prepared for each functional element, and the desired element integration DAO can be easily constructed by selecting and combining the necessary element DAOs. The element DAO provider 73 is provided with a server 74 and an element DAO protocol DB 75. The element DAO protocol DB 75 stores, for example, element DAOs prepared for each functional element required for a company, element DAOs prepared for each functional element required for an NPO (nonprofit organization), and element DAOs prepared for each functional element required for a local government.
[0245] An element integration DAO builder receives an order from a client to build an element integration DAO and installs element DAOs corresponding to the required functional elements on PC terminal 76 via server 74. In the example of Figure 43, an A1 element DAO (including an A1 element agent), an A2 element DAO (including an A2 element agent), an A5 element DAO (including an A5 element agent), and an A9 element DAO (including an A9 element agent) are installed. Each of the A1 to A9 element agents is an AI that performs reinforcement learning (machine learning) to enable the corresponding element DAO to perform at its best.
[0246] A control agent is also installed on the PC terminal 76. This control agent controls the element agents of each element DAO to optimize the entire element integration DAO, and the control agent itself performs reinforcement learning (machine learning) to achieve overall optimization. Because each element agent is responsible for maximizing the performance of the element DAO it is responsible for, there is a risk that the element agents alone will fall into partial optimization and will not be able to achieve overall optimization of the element integration DAO. Therefore, a control agent is needed to control the element integration DAO so that it is optimized as a whole. This can be said to be the same as, for example, searching for a Pareto-optimal solution in an incomplete information game.
[0247] The element integration DAO installed on PC67 is installed on multiple terminals 16 to perform simulation reinforcement learning using the mirror world 51 as a simulation environment, and a blockchain digital twin 2T consisting of a private chain 2 with these terminals 16 as nodes 19 is generated within the mirror world 51.
[0248] The main routine of the simulation reinforcement learning of this element integration DAO will be explained based on Figure 44(A). The CPU 10 of the user terminal 16 of the requester who requests simulation reinforcement learning of the element integration DAO performs simulation reinforcement learning preparation response processing at S674 and simulation reinforcement learning response processing at S675. The CPU 10 of the mirror world server 46 performs simulation reinforcement learning preparation processing at S679 and simulation reinforcement learning processing at S680.
[0249] Details of the simulation reinforcement learning preparation response process shown in S674 and the simulation reinforcement learning preparation process shown in S679 will be described with reference to FIG. 44(B). In the simulation reinforcement learning preparation response process, the CPU 10 of the user terminal 16 determines whether to request simulation reinforcement learning at S679, and returns if not. If it is determined to request simulation reinforcement learning, the CPU 10 transmits the DAO data and personal AIs at S680 to request simulation reinforcement learning. The DAO data is the functional elements of the organization for which simulation reinforcement learning is desired. For example, in the case of a furniture assembly and sales company, these elements include a material procurement element, an assembly element, an advertising element, and a sales element. The personal AIs are the personal AIs of people actually engaged in the element integration DAO in the real world 47. If there are employees without personal AIs or if employees have not yet been decided, a personal AI matching the element integration DAO to be used for simulation reinforcement learning is generated and prepared, as described above with reference to S561, S565 to S568, S562, and S573 to S575.
[0250] In the simulation reinforcement learning preparation process, the CPU 10 of the mirror world server 46 determines in S683 whether or not a request for simulation reinforcement learning has been made, and returns if not. If it determines that a request for simulation reinforcement learning has been made, in S684, the CPU 10 copies the personal AIs and deploys them in the element integration DAO to generate an element integration DAO digital twin, which is registered in the mirror world 51 as a simulation target.
[0251] This state is shown in Figure 45. A digital twin 78 of an element integration DAO 77 in the real world 47 is registered in the mirror world 51. The element integration DAO digital twin 78 shown in Figure 45 is, for example, the element integration DAO digital twin 78 of a furniture assembly and sales company, and has the functional elements of material procurement, assembly, advertising, and sales, and the identification numbers of the personal AI groups of people engaged in each of these functional elements are stored in the mirror world server 46.
[0252] For this element-integrated DAO digital twin 78, the optimal incentive design is derived through simulation reinforcement learning. The outline of the reinforcement learning (machine learning) system is shown in Figure 46.
[0253] 46, in the element integrated DAO digital twin 78, corresponding to each functional element of material procurement, assembly, advertising, and sales, a material procurement element agent 80 and a group of personal AIs 84 in charge of material procurement, an assembly element agent 81 and a group of personal AIs 85 in charge of assembly, an advertising element agent 82 and a group of personal AIs 86 in charge of advertising, a material procurement element agent 80 and a group of personal AIs 84 in charge of material procurement, and a sales element agent 83 and a group of personal AIs in charge of sales are formed. A supervisory agent 79 supervises each of these element agents 80 to 83.
[0254] The group of personal AIs 84 in charge of material procurement perform actions a11, a12, ... a1n such as proposals in internal meetings, and finally execute the final collective action a1 on the group of digital twins 88 of material suppliers. The state S1 of the group of digital twins 88 of material suppliers for the action a1 is input to the material procurement element agent 80 and the group of personal AIs 84 in charge of material procurement. This state S1 is, for example, the number of materials requested and the response price for the action a1 of price negotiation. Note that each of the actions a11, a12, ... a1n of the group of personal AIs 84 in charge of material procurement is also input to the material procurement element agent 80, and the collective action a1 is also input to the control agent 79 and the material procurement element agent 80.
[0255] The material procurement element agent 80 calculates the performance p1 of the group of personal AIs 84 in charge of material procurement based on the action a1 and the state S1, and transmits the performance p1 to the control agent 79. The control agent 79 determines a reward r1 based on the performance p1, and transmits the reward r1 to the material procurement element agent 80. The material procurement element agent 80 determines a reward distribution rate based on each action a11, a12, ... a1n of the group of personal AIs 84 in charge of material procurement, and distributes the reward r1 to each personal AI 84 in charge of material procurement according to the reward distribution rate.
[0256] The group of personal AIs 85 responsible for assembly perform actions a21, a22, ... a2i such as proposals in internal meetings, and finally execute the combined action a2 on the group of digital twins 89 of the assembly equipment. The state S2 of the group of digital twins 89 of the assembly equipment for that action a2 is input to the assembly element agent 81 and the group of personal AIs 85 responsible for assembly. This state S1 is, for example, the power consumption of the group of digital twins 89 of the assembly equipment and the total working hours of the group of personal AIs 85 responsible for assembly engaged in the group of digital twins 89 of the assembly equipment. Note that each of the actions a21, a22, ... a2i of the group of personal AIs 84 responsible for assembly is also input to the assembly element agent 81, and the combined action a2 is also input to the supervision agent 79 and the assembly element agent 81.
[0257] The assembly element agent 81 calculates the performance p2 of the group of personal AIs 85 in charge of assembly based on the action a2 and the state S2, and sends this performance p2 to the control agent 79. The control agent 79 determines a reward r2 based on the performance p2, and sends this reward r2 to the assembly element agent 81. The assembly element agent 81 determines a reward distribution rate based on each of the actions a21, a22, ... a2i of the group of personal AIs 85 in charge of assembly, and distributes the reward r2 to the personal AIs 85 in charge of assembly according to this reward distribution rate.
[0258] The advertising staff's personal AI group 86 performs actions a51, a52, ... a5j such as proposals in internal meetings, and finally executes the final collected action a5 on the consumer's personal AI group 90. The state S5 of the consumer's personal AI group 90 with respect to the action a5 is input to the advertising element agent 82 and the advertising staff's personal AI group 86. This state S5 is, for example, whether or not the consumer purchased the product in response to the product recommendation action a5 to the consumer, and the purchase amount, etc. Note that each of the actions a51, a52, ... a5j of the advertising staff's personal AI group 86 is also input to the advertising element agent 82, and the collected action a5 is also input to the control agent 79 and the advertising element agent 82.
[0259] The advertising element agent 82 calculates the performance p5 of the advertising manager's personal AI group 86 based on the action a5 and the state S5, and transmits the performance p5 to the control agent 79. The control agent 79 determines a reward r5 based on the performance p5, and transmits the reward r5 to the advertising element agent 82. The advertising element agent 82 determines a reward distribution rate based on each of the actions a51, a52, ... a5j of the advertising manager's personal AI group 86, and distributes the reward r5 to the advertising manager's personal AI 86 in accordance with the reward distribution rate.
[0260] The sales personal AI group 87 performs actions a91, a92, ... a9m such as proposals in internal meetings, and finally executes the final aggregated action a9 on the store and consumer digital twin group 91. The state S9 of the store and consumer digital twin group 91 for the action a9 is input to the sales element agent 83 and the sales personal AI group 87. This state S9 is, for example, the total sales amount at the store. The state S5 of the material provider's digital twin 88 for the above action a5 is also input to the sales element agent 83. Note that each of the actions a91, a92, ... a9m of the sales personal AI group 87 is also input to the sales element agent 83, and the aggregated action a9 is also input to the supervision agent 79 and the sales element agent 83.
[0261] The sales element agent 83 calculates the performance p9 of the sales representative's personal AI group 87 based on the action a9 and states S5 and S9, and transmits the performance p9 to the control agent 79. The control agent 79 determines a reward r9 based on the performance p9, and transmits the reward r9 to the sales element agent 83. The sales element agent 83 determines a reward distribution rate based on each action a91, a92, ... a9m of the sales representative's personal AI group 87, and distributes the reward r9 to the sales representative's personal AI 87 according to the reward distribution rate.
[0262] The calculation method for each of the performances p1 to p9 and each reward distribution rate will be explained based on Figures 47(A), (B), (C) and 48(A). The material procurement element agent 80 stores the calculation algorithm for performance p1 and distribution rate as knowledge. The calculation algorithm for performance p1 and distribution rate will be explained based on Figure 47(A). The material procurement element agent 80 sets the current material purchase amount u (from the previous YES in S689 to the current YES) and the current inventory amount z as state S1, and calculates performance p1 using the calculation formula: performance p1 = {2(average purchase amount / u) + (z / average inventory amount)} / 3. The average purchase amount is the average purchase amount of materials from the start of the simulation reinforcement learning to the present. The average inventory amount is the average number of materials purchased in stock from the start of the simulation reinforcement learning to the present. As a result of this calculation formula, if the current material purchase amount u is cheaper, performance p1 will increase, and if the current inventory amount z is larger, performance p1 will decrease.
[0263] Furthermore, if p1≧1, the reward distribution rate is calculated in proportion to the degree of approval for the collective action a1. Conversely, if p1<1, the reward distribution rate is calculated in inverse proportion to the degree of approval for the collective action a1. This is not limited to being proportional or inversely proportional to the first power of the "degree of approval," but also includes being proportional or inversely proportional to the nth power of the "degree of approval," and the material procurement element agent 80 finds the optimal proportional or inversely proportional function by performing reinforcement learning (machine learning). Furthermore, the "degree of approval" is highest for the personal AI that proposed the collective action a1 itself, and the material procurement element agent 80 determines (calculates) the degree of approval for each personal AI based on each of the personal AI's actions a11, a12,... a1n.
[0264] The calculation algorithm for performance p2 and distribution rate stored as knowledge by assembly element agent 81 will be described with reference to Figure 47(B). Assemble element agent 81 sets the current power consumption e of the assembly equipment (from the previous YES in S710 to the current YES) and the current total work hours t of assembly workers as state S2, and calculates performance p2 using the formula: performance p2 = {(average power consumption / e) + (average total work hours / t)} / 2. The current total work hours t of assembly workers is the total work hours of the personal AI group 85 in charge of assembly engaged in the current assembly equipment digital twin group 89. The average power consumption is the average power consumption of the assembly equipment from the start of the simulation reinforcement learning to the present. The average total work hours is the average total work hours of assembly workers from the start of the simulation reinforcement learning to the present. As a result of this calculation formula, if the current power consumption e of the assembly equipment is reduced, performance p2 will increase, and if the current total work hours t of assembly workers is increased, performance p2 will decrease.
[0265] Furthermore, if p2 ≥ 1, the reward distribution rate is calculated in proportion to the degree of approval for the collective action a2. Conversely, if p2 < 1, the reward distribution rate is calculated in inverse proportion to the degree of approval for the collective action a2. This is not limited to being proportional or inversely proportional to the first power of the "degree of approval," but also includes being proportional or inversely proportional to the nth power of the "degree of approval," and the assembly element agent 81 finds the optimal proportional or inversely proportional function by performing reinforcement learning (machine learning). Furthermore, the "degree of approval" is highest for the personal AI that proposed the collective action a2 itself, and the assembly element agent 81 determines (calculates) the degree of approval for each personal AI based on each of the personal AI's actions a21, a22, ... a2i.
[0266] The calculation algorithm for the performance p5 and distribution rate stored as knowledge by the advertising element agent 82 will be explained based on Figure 47 (C). The advertising element agent 82 sets the total purchase amount k of the recommended consumer's personal AI this time (from the previous YES in S728 to the current YES) as state S5, and calculates performance p5 using the calculation formula performance p5 = k / average total purchase amount K of the recommended consumer's personal AI. The average total purchase amount K is the average of the total purchase amounts of the recommended consumer's personal AI group 90 from the start of simulation reinforcement learning to the present. As a result of this calculation formula, the higher the current total purchase amount k of the recommended consumer's personal AI, the higher the performance p5 will be.
[0267] Furthermore, if p5≧1, the reward distribution rate is calculated in proportion to the degree of approval for the collective action a5. Conversely, if p5<1, the reward distribution rate is calculated in inverse proportion to the degree of approval for the collective action a5. This is not limited to being proportional or inversely proportional to the first power of the "degree of approval," but also includes being proportional or inversely proportional to the nth power of the "degree of approval," and the advertising element agent 82 finds the optimal proportional or inversely proportional function by performing reinforcement learning (machine learning). Furthermore, the "degree of approval" is highest for the personal AI that proposed the collective action a5 itself, and the advertising element agent 82 determines (calculates) the degree of approval for each personal AI based on each of the personal AI's actions a51, a52,...a5j.
[0268] The calculation algorithm for the performance p5 and distribution rate stored as knowledge by the sales element agent 83 will be explained based on Figure 48(A). The sales element agent 83 sets the total sales amount h and average total sales amount H at the store this time (from the previous YES in S749 to the current YES) as state S9, and calculates performance p9 using this state S9 and the above state S5 using the calculation formula performance p9 = (hk) / (HK). The average total sales amount H is the average of the total sales amount at the store from the start of the simulation reinforcement learning to the present. Furthermore, k is the total purchase amount of the recommended consumer personal AI this time, and K is the average of the total purchase amount of the recommended consumer personal AI group 90 from the start of the simulation reinforcement learning to the present (see Figure 47(C) and its explanation). As a result of this calculation formula, the performance p9 increases as the value obtained by subtracting the total purchase amount k of the recommended consumer personal AI from the total sales amount h at the store this time increases. The total purchase amount k of the consumer personal AI that made the recommendation is the credit of the group of personal AIs 86 in charge of advertising, and the credit of the group of personal AIs 87 in charge of sales alone is the value obtained by subtracting the total purchase amount k of the consumer personal AI that made the recommendation from the total sales amount h at the store.
[0269] Furthermore, if p9≧1, the reward distribution rate is calculated in proportion to the degree of approval for the collective action a9. Conversely, if p9<1, the reward distribution rate is calculated in inverse proportion to the degree of approval for the collective action a9. This is not limited to being proportional or inversely proportional to the first power of the "degree of approval," but also includes being proportional or inversely proportional to the nth power of the "degree of approval," and the advertising element agent 83 finds the optimal proportional or inversely proportional function by performing reinforcement learning (machine learning). Furthermore, the "degree of approval" is highest for the personal AI that proposed the collective action a9 itself, and the sales element agent 83 determines (calculates) the degree of approval for each personal AI based on each of the personal AI's actions a91, a92,...a9m.
[0270] Next, the reward table 92 stored as knowledge by the control agent 79 will be described with reference to Fig. 48(B). This reward table 92 stores a calculation formula for the reward to be distributed by the control agent 79 to each element agent 80-83. The distributed reward is calculated by multiplying the coefficient by (profit for this period) by (performance sent from the target element agent) divided by (total performance sent from all element agents). Here, "this period" refers to the period from the previous YES in S675 to the current YES.
[0271] Specifically, the reward r1 to be distributed to the material procurement element agent 80 is r1=A1·Lt·p1 / (p1+p2+p5+p9). The reward r2 to be distributed to the assembly element agent 81 is r2=A2·Lt·p2 / (p1+p2+p5+p9). The reward r5 to be distributed to the advertising element agent 82 is r5=A5·Lt·p5 / (p1+p2+p5+p9). The reward r9 to be distributed to the sales element agent 83 is r9=A9·Lt·p9 / (p1+p2+p5+p9). Here, Lt is the profit for the current period, and A1, A2, A5, and A9 are coefficients representing the actions determined by the supervision agent 79.
[0272] Next, the specific contents of the simulation reinforcement learning process shown in S680 will be described with reference to Figure 49. Reinforcement learning process for the control agent is executed in S687, reinforcement learning process for the material procurement element agent is executed in S688, reinforcement learning process for the assembly element agent is executed in S689, reinforcement learning process for the advertising element agent is executed in S690, reinforcement learning process for the sales agent is executed in S691, reinforcement learning process for the material procurement personal AI is executed in S692, reinforcement learning process for the assembly personal AI is executed in S693, reinforcement learning process for the advertising personal AI is executed in S694, and reinforcement learning process for the sales personal AI is executed in S695.
[0273] The details of the control agent reinforcement learning process shown in S687 will be explained with reference to Fig. 50. In S699, the control agent 79 determines whether or not it has received each performance p sent from each element agent 80 to 83. If it has not received it, control proceeds to S671, but if it is determined that it has received it, it stores each received performance p in S670.
[0274] Next, in S671, it is determined whether or not each of the actions a1 to a9 has been received, and if not, control proceeds to S673. If it is determined that each of the actions a1 to a9 has been received, in S672, the received actions a1 to a9 are stored. Next, in S673, it is determined whether or not there has been input of the state S9 sent from the personal AI group 91 of the store and consumer, and if not, control proceeds to S675. If it is determined that there has been input, in S674, sales=ΣS9 is calculated.
[0275] Next, in S675, it is determined whether it is time to calculate the reward. If it is not, control proceeds to S677. However, if it is determined that it is time, in S676, the reward table 92 is referenced to calculate rewards r1, r2, r3, r5, and r9 and they are sent to the corresponding element agents 80 to 83.
[0276] Next, in S677, it is determined whether it is time to update the reinforcement learning (machine learning), and if not, it returns. If it is determined that it is time to update, in S687, the current period's profit Lt = sales - expenses is calculated. Next, in S679, each reward r1, r2, r5, r9 is calculated and distributed to the corresponding element agents 80 to 83.
[0277] Next, in S680, the reward R of the control agent 79 is calculated from the profit Lt. This reward R is proportional to the profit Lt. Next, in S681, the optimal policy π *Next, in S682, A1, A2, A5, and A9 in the reward table 92 are updated to the actions (coefficients) A1, A2, A5, and A9 calculated in S681. As a result, the control agent 79 learns the actions (coefficients) A1, A2, A5, and A9 that maximize the profit Lt.
[0278] The details of the material procurement element agent reinforcement learning process shown in S688 will be explained with reference to Figure 51. In S684, the material procurement element agent 80 performs information collection processing using a crawler. A crawler is a program that periodically acquires documents and images on the web and automatically creates a database. It is also called a "bot," "spider," or "robot."
[0279] The details of the information collection process by the crawler will be explained with reference to Fig. 52(A). In S702, the material procurement element agent 80 receives the information collected by the crawler while circulating the network. Next, in S703, the received information is stored in the material procurement DB 93.
[0280] 52(B) shows the information stored in the material procurement DB 93. As shown in the figure, the material procurement DB 93 stores various information necessary for material procurement operations, such as economic information, social information, weather information, inventory information, market information, and supplier information.
[0281] Returning to Figure 51, in S685, the material procurement element agent 80 determines whether or not actions a11, a12, ... a1n have been received from the personal AI group 84. If not, control proceeds to S687, but if it is determined that they have been received, in S686, the received actions a11, a12, ... a1n are stored. In S687, it is determined whether or not state S1 has been received from the material supplier digital twin group 88, and if not, control proceeds to S689. If it is determined that they have been received, the received S1 is stored in S688.
[0282] In S689, it is determined whether it is time to calculate performance p1, and if it is not time, control proceeds to S692. If it is determined that it is time to calculate, in S690, performance p1 = {2 (average purchase amount / u) + (z / average inventory quantity)} / 3 is calculated. The performance p1 is then sent to the control agent 79 (S691).
[0283] In S692, it is determined whether or not the reward r1 transmitted from the control agent 79 has been received, and if not, the process returns. If it is determined that the reward r1 has been received, the reward distribution rate is calculated based on the algorithm for the reward distribution rate shown in FIG. 47(A) (S693). In S694, the reward is multiplied by each distribution rate to calculate each reward r11, r12...r1n, and in S695, each reward r11, r12...r1n is given to each material procurement personal AI group 84. In S696, based on the received reward r1, the optimal policy π * In S697, the proportional function or the inverse proportional function is updated to the function obtained in S696. As a result, the material procurement element agent 80 learns the proportional function or the inverse proportional function that maximizes the performance p1.
[0284] Next, the details of the assembly element agent reinforcement learning process shown in S689 will be explained with reference to Figure 53. In S706, the assembly element agent 81 determines whether or not actions a21, a22, ... a2n have been received from the personal AI group 85. If not, control proceeds to S708, but if it is determined that they have been received, in S707, the received actions a21, a22, ... a2n are stored. In S708, it is determined whether or not a state S2 has been received from the assembly equipment digital twin group 89, and if not, control proceeds to S710. If it is determined that they have been received, in S709, the received state S2 is stored.
[0285] In S710, it is determined whether it is time to calculate performance p2, and if it is not time, control proceeds to S713. If it is determined that it is time to calculate, in S711, performance p2 = {(average power consumption / e) + (average total working hours / t)} / 2 is calculated. The performance p2 is sent to the control agent 79 (S712).
[0286] In S713, it is determined whether or not the reward r2 sent from the control agent 79 has been received, and if not, the process returns. If it is determined that the reward r2 has been received, the reward distribution rate is calculated based on the algorithm for the reward distribution rate shown in FIG. 47(B) (S714). In S715, the reward is multiplied by each distribution rate to calculate each reward r11, r12...r1n, and in S695, each reward r21, r22...r2i is given to the assembly personal AI group 86. In S717, based on the received reward r2, the optimal policy π * In S718, the proportional function or the inverse proportional function is updated to the function obtained in S717. As a result, the assembly element agent 81 learns the proportional function or the inverse proportional function that maximizes the performance p2.
[0287] The details of the advertising element agent reinforcement learning process shown in S690 will be explained with reference to Figure 54. In S723, the advertising element agent 83 performs information collection processing using a crawler. The details of this processing will be explained with reference to Figure 55(A). In S740, the advertising element agent 83 receives information collected by the crawler patrolling the Internet, and in S741 stores the received information in the advertising DB 94.
[0288] The collected data stored in the advertisement DB 94 is shown in Figure 55(B). The advertisement DB 94 stores various behavioral data of consumers such as Taro, Jiro, and Hanako. For example, in the case of Taro, it is determined that there is a high possibility that he will purchase furniture based on the information that he "ordered a detached house," and furniture advertisements are made to Taro. In the case of Jiro, it is determined that there is a high possibility that he will purchase furniture for his new home because he is getting married soon based on the information that he "purchased a couple's furniture," and furniture advertisements are made to Jiro.
[0289] Returning to Figure 54, in S742, the advertising element agent 83 determines whether or not it has received an action from each personal AI group 86. If it has not received an action, control proceeds to S726, but if it is determined that it has received an action, it stores each received action in S725. In S726, it determines whether or not it has received a status S5 sent from the consumer's personal AI group 90, and if it has not yet received an action, control proceeds to S728. If it is determined that it has received an action, it stores the received status S5 in S727.
[0290] In S728, it is determined whether it is time to calculate performance p5, and if it is not time to calculate performance p5, control proceeds to S731. If it is determined that it is time to calculate performance p5, in S729, performance p5 = k / average total purchase amount K of the personal AI of the consumer who made the recommendation is calculated. Next, in S730, performance P5 is sent to the control agent 79.
[0291] In S731, it is determined whether or not the reward r5 sent from the control agent 79 has been received, and if it has not yet been received, the process returns. If it is determined that the reward r5 has been received, in S732, the reward distribution rate is calculated based on the algorithm for the reward distribution rate shown in Figure 47 (C) (S732). In S733, the reward is multiplied by each distribution rate to calculate each reward r51, r52 ... r5j, and in S734, each reward r51, r52 ... r5j is granted to the advertising personal AI group 87. In S735, based on the received reward r5, the optimal policy π *In S736, the proportional function or the inverse proportional function is updated to the function obtained in S735. As a result, the advertising element agent 82 learns the proportional function or the inverse proportional function that maximizes the performance p5.
[0292] Next, the details of the sales element agent reinforcement learning process shown in S691 will be explained with reference to Figure 56. The sales element agent 84 performs information collection processing using a crawler in S744. The details of this processing will be explained with reference to Figure 57(A). In S760, the sales element agent 84 receives information collected by the crawler patrolling the Internet, and in S761, stores the received information in the sales DB 95. Furthermore, in S762, POS data from the store is stored in the sales DB 95.
[0293] The collected data stored in the sales DB 95 is shown in Figure 57(B). The sales DB 95 stores various data such as weather data and POS data. "By date" in weather information is a concept that includes by day of the week. Based on the weather information (weather and temperature data by date and time) and the POS data (sales product data by date and time), it is possible to rearrange displayed products, for example, taking into account the day of the week, time, and weather conditions.
[0294] Returning to Figure 56, in S745, the sales element agent 84 determines whether or not it has received an action from each personal AI group 87. If it has not received an action, control proceeds to S747, but if it is determined that it has received an action, it stores each received action in S746. In S747, it determines whether or not it has received a status S9 sent from the retailer and consumer personal AI group 91, and if it has not yet received an action, control proceeds to S749. If it is determined that it has received an action, it stores the received status S9 in S748.
[0295] In S749, it is determined whether it is time to calculate performance p9, and if it is not time to calculate performance p9, control proceeds to S752. If it is determined that it is time to calculate performance p9, in S750, performance p9 = (hk) / (HK) is calculated. Next, in S751, performance p9 is sent to the control agent 79.
[0296] In S752, it is determined whether or not the reward r9 sent from the control agent 79 has been received, and if it has not yet been received, the process returns. If it is determined that the reward r9 has been received, in S753, the reward distribution rate is calculated based on the algorithm for the reward distribution rate shown in FIG. 48(A) (S753). In S754, the reward is multiplied by each distribution rate to calculate each reward r91, r92...r9m, and in S755, each reward r91, r92...r9m is granted to the sales personal AI group 88. In S756, based on the received reward r9, the optimal policy π * In S757, the proportional function or the inverse proportional function is updated to the function obtained in S756. As a result, the sales element agent 83 learns the proportional function or the inverse proportional function that maximizes the performance p9.
[0297] Next, details of the material procurement personal AI reinforcement learning process shown in S692 will be explained with reference to Figure 58 (A). In S765, the material procurement personal AI group 84 determines whether to negotiate with the material supplier digital twin group 88, and if not, control proceeds to S770. If it is determined that negotiation will be performed, in S766, the stored data in the fundraising DB 93 is viewed, and action a1 is determined while holding an internal meeting with reference to the stored data (S767), and then negotiations are performed with the material supplier digital twin group 88 (S768). In S769, it is determined whether the negotiation has ended, and if not, the process returns to S766, and if not, the process circulates through a loop of S767 → S768 → S769 → S766. When it is determined in S769 that the negotiation has ended, control proceeds to S770.
[0298] In S770, it is determined whether or not rewards r11, r12, . . . r1n have been received from the material procurement element agent 80, and if not, the process returns. If it is determined that rewards have been received, in S771, the optimal policy π * The action a1i involves moving (changing jobs) to another company DAO digital twin (for example, DAO digital twin 59 of ABC Co., Ltd. in Figure 45) if the received reward r1i is not satisfactory. As a result of this reinforcement learning, each personal AI responsible for material procurement will learn actions that increase the aforementioned performance p1.
[0299] Next, details of the assembly personal AI reinforcement learning process shown in S693 will be explained based on Figure 58 (B). In S775, the assembly personal AI group 85 determines whether to hold an internal meeting, and if not, control proceeds to S779. If it is determined that a meeting should be held, in S776, each assembly personal AI holds an internal meeting and decides on action a2. Next, in S777, the assembly equipment digital twin group 89 is test-run in accordance with action a2, and the appropriateness of action a2 is verified. In S778, it is determined whether the meeting has ended, and if it has not yet ended, the process returns to S776, and the process circulates in a loop of S777 → S778 → S776. If action a2 is appropriate as a result of the test run in S777, S778 determines that the meeting has ended, and control proceeds to S779.
[0300] In S779, it is determined whether or not rewards r21, r22, . . . r2n have been received from the assembly element agent 81, and if not, the process returns. If it is determined that a reward has been received, in S780, the optimal policy π *The action a2i involves moving (changing jobs) to another company DAO digital twin (for example, DAO digital twin 59 of ABC Co., Ltd. in Figure 45) if the received reward r2i is not satisfactory. As a result of this reinforcement learning, each of the assembly personal AIs will learn actions that will increase the aforementioned performance p2.
[0301] Next, the details of the advertising personal AI reinforcement learning process shown in S694 will be explained with reference to Figure 59 (A). In S784, the advertising personal AI group 86 determines whether to hold an internal meeting, and if not, control proceeds to S789. If it is determined that a meeting should be held, in S785, each advertising personal AI holds an internal meeting and decides on action a5. Next, in S787, action a2 toward the consumer is executed. In S788, it is determined whether the meeting has ended, and if not, the process returns to S785 and goes through a loop of S786 → S787 → S788. When it is determined in S788 that the meeting has ended, control proceeds to S789.
[0302] In S789, it is determined whether or not rewards r51, r52, . . . r5j have been received from the advertising element agent 82, and if not, the process returns. If it is determined that rewards have been received, in S790, the optimal policy π * The action a5i involves moving (changing jobs) to another company DAO digital twin (for example, DAO digital twin 59 of ABC Co., Ltd. in Figure 45) if the received reward r5i is not satisfactory. As a result of this reinforcement learning, each advertising personal AI will learn actions that increase the aforementioned performance p5.
[0303] Next, details of the salesperson personal AI reinforcement learning process shown in S695 will be explained with reference to Figure 59 (B). In S791, the salesperson personal AI group 87 determines whether to hold an internal meeting, and if not, control proceeds to S795. If it is determined that a meeting should be held, in S792, each salesperson personal AI decides on an action a9 while holding an internal meeting. Next, in S793, it is determined whether the meeting has ended, and if not, the process returns to S792 and goes through a loop of S792 → S793 → S792. When it is determined in S793 that the meeting has ended, control proceeds to S794. In S794, the action decided in the meeting is carried out for the consumer and the store.
[0304] Next, in S795, it is determined whether rewards r91, r92...r9m have been received from the sales element agent 83, and if not, the process returns. If it is determined that rewards have been received, in S796, actions (a91, a92...a9m) according to the optimal policy π* are obtained through TD learning based on the received rewards. This action a9i includes moving (changing jobs) to another company DAO digital twin (for example, the DAO digital twin 59 of ABC Co., Ltd. in Figure 45) if the received reward r9i is not satisfactory. As a result of this reinforcement learning, each sales personal AI learns actions that increase the aforementioned performance p9.
[0305] After completing the simulation reinforcement learning, the element integrated DAO is operated as an actual organization in the real world 47. At that stage, the "personal AI groups 84-87" in Figure 46 will be managed by actual people (users) in the real world. At that time, the personal AI groups 84-87 that have completed the simulation reinforcement learning will act as advisors to the actual people (users), and will be able to provide the actual people (users) with the knowledge, experience, and know-how they have gained through the simulation reinforcement learning.
[0306] The construction of the element integration DAO described above shows the creation of an entire organization, such as a company, NPO, or local government, by combining element DAOs for each function. However, it is also possible to build only a part of the organization (for example, material procurement) using element DAOs, rather than the entire organization.
[0307] The above-mentioned programs that run on user terminals 16, etc. and various servers may be downloaded and installed from a predetermined website, etc., or may be recorded on a recording medium (non-transitory recording medium) such as a CD-ROM 99 and distributed, and those who purchase the CD-ROM 99, etc., can install the programs on the user terminals 16 and various servers (see Figure 60). [Variations]
[0308] (1) For example, in the digital twin data shown in Figure 29, names such as Taro, Jiro, Sakura, and Saburo may be pseudonyms (anonymous) from the perspective of protecting personal information, so that they can be identified as the same person but cannot be used to identify a specific individual. In this case, an AI identification number or blockchain address may be used as the pseudonym (anonymous). Similarly, digital twins of companies such as ABC Co., Ltd. may use pseudonyms (anonymous) for their company names (organization names), so that they can be identified as the same company (same organization) but cannot be used to identify a specific company (organization). Furthermore, multiple digital twins of humans may be created using multiple personal AIs for a single person. Furthermore, one digital twin for a single person may be composed of a collection of multiple personal AIs (for example, a collection of specialized personal AIs in various fields).
[0309] (2) In Figures 34 to 59, we have explained a system that derives an optimal solution for incentive design in a DAO by performing simulations using a multi-service DAO, which has multiple types of services.However, this is not limited to multi-service DAOs, and the system may also be one that derives an optimal solution for incentive design through simulations for DAOs that have only one type of service.
[0310] (3) In Figure 35, a trained persona agent group is generated for each persona, but the personal AIs of the user group belonging to each persona may be selected from the existing personal AIs registered in the mirror world 51 and used as the persona agent group. In this case, it is necessary to inquire with the user group belonging to each persona as to whether or not it is OK to use their personal AI in the simulation and obtain consent. The personal AIs of the user group who have given consent are copied and used in the simulation, and the trained personal AI group after the simulation is completed is sent to each corresponding user. If each user who receives it determines that the trained personal AI is useful (necessary), they overwrite and save the trained personal AI over their existing personal AI. Note that both the existing personal AI and the trained personal AI may be stored and used as needed.
[0311] (4) As a multi-agent reinforcement learning method, we have presented a master agent method in which a supervisory agent (master agent) responsible for overall optimization distributes rewards to each agent, and the supervisory agent itself also performs reinforcement learning to converge the reward distribution behavior to the optimal one. However, multi-agent reinforcement learning is not limited to this, and for example, D-learning, which can converge to the optimal solution under a Markov decision process, or Bucket Brigade or Profit Sharing as reinforcement learning algorithms in classifier systems can also be used.
[0312] (5) Simulations using digital twins are not limited to digital twins of people or organizations composed of people (e.g., corporations, NPOs, etc.). For example, for an object such as an AI-equipped machine or electrical appliance (e.g., an AI-equipped vacuum cleaner), a digital twin of the environment in which the object operates (e.g., the interior of a user's home in which an autonomously moving AI-equipped vacuum cleaner operates) can be generated in cyberspace, and the AI equipped in the object can be simulated in advance in the environmental digital twin to undergo reinforcement learning (machine learning), and the customized (personalized) object equipped with the trained AI can be provided to the relevant user.
[0313] Since it can resolve the contradictory dilemma between guaranteeing the authenticity of recorded information and guaranteeing the right to delete that information as much as possible, it can be used for information recording methods that are non-erasable, such as blockchain. [Explanation of symbols]
[0314] 1. Internet 2 Private Chain 3. Consortium Chain 4. Public Chain 12 HDD 16 User terminals 19 nodes 30 Key Registration Center 32 Key DB 46 Mirror World Server 51 Mirror World 52 Earth Digital Twin 53 Japan Digital Twin 54 Town Digital Twin 57 Taro Digital Twin 58 Taro Family Digital Twin 59 ABC Digital Twin Co., Ltd. 61 DAO Agent 72 tokens 78 DAO Digital Twin 79 Supervising Agent.
Claims
1. an encryption means for performing an encryption process for encrypting information to be recorded; a recording means for recording the information after the encryption process; a decryption means for decrypting the information recorded by the recording means using a first key and a second key to generate plaintext information; a decryption-disabling means for making the information recorded by the recording means into a decryption-disabled state in which the information cannot be decrypted, the decryption means includes second key secret storage means for storing the second key in secret, The decryption disablement means updates the second key held by the second key secret holding means to a different key, thereby disabling the decryption.
2. 2. The processing system according to claim 1, wherein said decryption means further includes first key distribution means for distributing said first key to a person who wishes to view the information.
3. 3. The processing system according to claim 1, further comprising a search unit that searches the information recorded by said recording unit without converting the information into plain text.
4. the information recorded by the recording means includes personal information; 4. The processing system according to claim 1, wherein said decryption-disabling means makes the personal information of the personal information owner in the decryption-disabled state in response to a request from the personal information owner.
5. performing an encryption process for encrypting information to be recorded; a decryption step of decrypting the information recorded by the recording means for recording the encrypted information using a first key and a second key to generate plaintext information; a step of rendering the information recorded by the recording means in an undecodable state so that the information cannot be decoded; Let the computer run the decrypting step includes a step of keeping the second key secret; The step of making the second key undecryptable is performed by updating the second key held in the holding step to another key.
Citation Information
Patent Citations
Instance weighted learning machine learning model
JP2016505974A
System, method and apparatus for demand-initiated intelligent negotiation agents in a distributed network
US20020046157A1
Learning Based on Simulations of Interactions of a Customer Contact Center
US20160189558A1
Continual reinforcement learning with a multi-task agent
US20190244099A1
Credibility management system and credibility management method
JP2018128723A
Cited By
Data certification method, data certification program, and data certification device
JP7892230B1