Device and method
The apparatus and method optimize coupon distribution in virtual spaces by using user behavioral data to adaptively distribute NFT coupons, enhancing guidance to relevant locations.
Patent Information
- Application Number
- PCT/JP2024/010820
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-25
AI Technical Summary
Existing virtual spaces lack an efficient method to distribute distribution information, such as coupon information, to users based on their behavioral history, leading to ineffective guidance towards desired stores or areas.
An apparatus and method that includes a virtual space providing unit and a character control unit to distribute distribution information to avatars in a virtual space based on user behavioral information in the real world, utilizing NFT coupons and reinforcement learning to optimize coupon distribution.
Enables efficient distribution of coupon information to users by adapting to their interests and preferences, guiding them to relevant stores or areas effectively.
Smart Images

Figure JP2024010820_25092025_PF_FP_ABST
Abstract
Description
Apparatus and method
[0001] The present disclosure relates to an apparatus and method for providing a virtual space.
[0002] Patent Document 1 describes an information processing device that includes an acquisition means for acquiring the movement history of one or more players in a virtual space, and a movement information generation means for generating movement information to be made to move an NPC (Non Player Character) from the movement history of one or more players.
[0003] JP 2023-148489 A
[0004] In the invention described in Patent Document 1, NPCs are operated based on their behavioral history in a virtual space, and are not operated in accordance with the user's actual actions. In recent virtual spaces, AI characters called NPCs have appeared. Users can enter the virtual space using avatars, and once multiple avatars have gathered in the virtual space, it is conceivable that AI characters operated by companies could be used to distribute distribution information such as coupon information. It is conceivable that this distribution information could be used to guide users to desired stores or areas, but it is desirable to distribute this information efficiently.
[0005] Therefore, an object of the present disclosure is to provide an apparatus and method that can efficiently distribute distribution information.
[0006] The device disclosed herein includes a virtual space providing unit that provides a virtual space including a user's avatar and a computer character, and a character control unit that controls the distribution of distribution information to the avatar for the computer character based on behavioral information of the user in the real world.
[0007] According to the present disclosure, distribution information can be distributed efficiently in a virtual space.
[0008] FIG. 1 is a diagram illustrating the system configuration of a virtual space provision system. FIG. 2 is a diagram illustrating user attribute information, NPC types, information indicating NFT coupon transaction history, and a proposal table. FIG. 3 is a block diagram illustrating the functional configuration of a virtual space provision device 100 according to the present disclosure. FIG. 4 is a flowchart illustrating the operation of the virtual space provision device 100. FIG. 5 is a flowchart illustrating negotiation processing. FIG. 6 is a block diagram illustrating the functional configuration of a virtual space provision device 100a having an optimized proposal model 113. FIG. 7 is a diagram illustrating the relationship between coupon information and rewards. FIG. 8 is a diagram applied to a specific example of reinforcement learning. FIG. 9 is a diagram illustrating past coupon receipt history. FIG. 10 is a schematic diagram illustrating the relationship between rewards and receipts. FIG. 11 is a diagram illustrating the granting of rewards when a dialogue is terminated midway. FIG. 12 is a schematic diagram illustrating Q-learning processing. FIG. 13 is a diagram illustrating an example of the hardware configuration of a virtual space provision device 100 according to an embodiment of the present disclosure.
[0009] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.
[0010] The virtual space provision system of the present disclosure includes a virtual space provision device 100, an NFT coupon issuing device 300, a coupon information DB 400, a blockchain PC 600, and a user attribute information DB 700. Figure 1 is a diagram showing the system configuration of the virtual space provision system. A user terminal 500 can access each device and server that make up this system.
[0011] The virtual space providing device 100 transmits data constituting the virtual space to the user terminal 500. The virtual space includes NPCs, which are corporate characters, and the user's avatar. The NPCs operate in accordance with control information for distributing coupon information to the user's avatar. The avatars operate in accordance with operation information from the user terminal 500.
[0012] The user terminal 500 receives data constituting a virtual space from the virtual space providing device 100, thereby providing the virtual space. The user can operate an avatar in the virtual space provided by the user terminal 500. The user can also receive coupon information from an NPC via the avatar in the virtual space. This NPC is a computer character that belongs to a company, and distributes coupon information to the user's avatar.
[0013] The NFT coupon issuing device 300 is a device that issues NFT coupons to users (avatars) of the user terminal 500 in a virtual space. The NFT coupon is coupon information based on non-fungible information written on the blockchain PC 600. This NFT coupon may include restrictions such as the coupon information being changed when it is transferred.
[0014] The coupon information DB 400 is a database that stores coupon information, which indicates coupon details such as discount rates.
[0015] The user terminal 500 is a smartphone, PC, or the like operated by a user. In the present disclosure, the user terminal 500 has a wallet 501 that stores a private key for handling NFTs owned by the user. In the present disclosure, the wallet 501 stores a private key for handling NFT coupons.
[0016] The blockchain PC 600 is composed of multiple computers (PCs) and is a group of PCs that uses blockchain technology to store issuance history information of NFT coupons granted to users of the user terminal 500.
[0017] The user attribute information DB 700 is a database that stores user attribute information. A user registers account information for cashless payment used in the real world or virtual space in advance in the user attribute information DB 700. Note that the account information is not limited to account information for cashless payment, as long as user attribute information is stored.
[0018] FIG. 2 is a diagram showing user attribute information, NPC type, information indicating NFT coupon transaction history, and a specific example of the proposal table 105.
[0019] The user attribute information DB 700 stores user attribute information. As shown in Fig. 2A, the user attribute information stores attribute information such as gender, age, and address in association with a user ID, as well as interest information such as hobbies, interests, and preferences.
[0020] As shown in FIG. 2(b), the virtual space providing device 100 (a storage unit, not shown) stores the type of NPC for each NPC ID as type information for the NPC. This information is used to identify the type of NPC belonging to a company in the virtual space. The type refers to the company name (or an identifier for identifying the company) and the type of business, but may also include other information.
[0021] The blockchain PC 600 also stores transaction history information for NFT coupons. As shown in FIG. 2(c), this transaction history information includes the type of coupon received for each user ID. Although not shown in the figure, it may also include receipt date and time information. Because this NFT coupon transaction history information is stored in the blockchain PC 600, it is difficult to tamper with.
[0022] The proposal table 105 contains information used when distributing coupon information. FIG. 2(d) shows a specific example of the proposal table 105. This proposal table stores, for each company, the coupon target, the initial value (discount rate), the maximum discount rate, and the priority level, in association with each other. For example, the coupon information with priority level 1 shown in FIG. 2(d) is a coupon for 10% off fries, and this coupon information is distributed when the NPC first attempts to distribute coupon information to an avatar.
[0023] The coupon information with priority 2 is a coupon for 10% off a hamburger, and this coupon information is treated as the next proposal. Note that the coupon information is generated within the range using the initial value and the maximum value, and the discount rate, etc. is set before the coupon information is distributed. This setting rule is arbitrary, and for example, it may be increased by 10% at the time of negotiation.
[0024] 3 is a block diagram showing the functional configuration of the virtual space providing device 100 of the present disclosure. As shown in the figure, the virtual space providing device 100 includes a virtual space providing unit 101, a character control unit 102, a transaction history information acquiring unit 103, a negotiation processing unit 104, and a proposal table 105.
[0025] The virtual space providing unit 101 is a part that provides a virtual space to the user terminal 500. The virtual space includes an avatar set for the user and an NPC (Non Player Character) that operates under computer control.
[0026] The character control unit 102 controls the actions of avatars and NPCs in the virtual space. The character control unit 102 controls the actions of avatars based on operation instructions from the user terminal 500, and controls the actions of NPCs according to predetermined rules.
[0027] The transaction history information acquisition unit 103 is a part that acquires transaction history information regarding coupon information of a specified user from the blockchain PC 600.
[0028] The negotiation processing unit 104 is a part that performs negotiation processing to gradually make it easier for an avatar to receive the content of coupon information to be distributed to the avatar. That is, the negotiation processing unit 104 refers to the proposal table 105 (see FIG. 2(d)) and performs processing to gradually change the content of the coupon information for the avatar. For example, if the avatar does not receive the initial coupon information, the negotiation processing unit 104 refers to the proposal table 105 and performs control to attempt to distribute coupon information in which the content of the coupon information has been changed to be more advantageous to the user in accordance with the priority order.
[0029] The proposal table 105 is a table that stores the proposal details and priority order of coupons. As shown in Fig. 2(d), the proposal table 105 stores the company, coupon, initial value, maximum discount rate, and priority order in association with each other. This information is predetermined and set by the coupon distributor.
[0030] Next, a description will be given of the operation of the virtual space providing device 100. FIG.
[0031] The virtual space providing unit 101 acquires the positions of the user's avatar and NPC in the virtual space (S101). The virtual space providing unit 101 then extracts avatar candidates within a predetermined range of a specific NPC (S102). The specific NPC is, for example, an NPC belonging to a certain store, and is a predetermined NPC. This NPC is, for example, a character that is about to distribute a coupon for a hamburger shop. Furthermore, the avatar candidates within the predetermined range are at least one avatar within range where the specific NPC can call out to them. The specific NPC is a predetermined NPC, and is an NPC that is a corporate character. The virtual space providing device 100 performs this operation for each corporate character.
[0032] Then, the transaction history information acquisition unit 103 acquires transaction history information from the blockchain PC 600 for each avatar candidate user using the private key of the wallet 501 of each user terminal 500 (S103).
[0033] In addition, the virtual space providing unit 101 may refer to the transaction history information of the NFT coupon of the blockchain PC 600 acquired by the transaction history information acquiring unit 103, and may not consider an avatar that has recently had a history of distributing coupons as an avatar candidate.
[0034] Furthermore, the virtual space providing unit 101 acquires interest information from the user attribute information DB 700 (S104).
[0035] The virtual space providing unit 101 then refers to the NPC type information and interest information to determine whether the user's interest information matches the type of the NPC (corporate avatar) (S105). For example, if the type of the specific NPC has a type indicating that it belongs to a hamburger shop (type is hamburger), and the user acquired as an avatar candidate likes hamburgers (hamburgers are listed as a favorite), it is determined that the user and the NPC match. Note that instead of or in addition to the interest information, transaction history information and type information may be compared. For example, if there is no receipt history for a type in the transaction history information, it may be determined that there is no match. Conversely, if there is a frequent receipt history for a type, it may be determined that there is a match.
[0036] Then, the negotiation processing unit 104 refers to the proposal table 105 and prepares an initial proposal for coupon information based on the transaction history information of past coupon information (S106). This coupon information is stored in the proposal table 105. The initial proposal is determined based on the initial value and maximum discount rate in the proposal table 105. For example, if this is the first time, coupon information with the initial discount rate set is prepared. If coupon information has been received multiple times, the negotiation processing unit 104 may prepare coupon information with the discount rate set at the time of receipt. If coupon information has not been received in the past, coupon information with a discount rate set within the maximum value range may be prepared.
[0037] The negotiation processing unit 104 transmits coupon information as an initial proposal to the user terminal 500 and presents the coupon to the avatar (S107).
[0038] The character control unit 102 then controls the NPC in the virtual space to distribute the coupon information to the avatar candidate. If the avatar does not accept the coupon (S108: NO), the negotiation processing unit 104 executes negotiation processing (S109). Specifically, the negotiation processing unit 104 refers to the proposal table 105 and prepares the next priority coupon information. Executing negotiation here includes gradually changing the coupon according to a predetermined rule. For example, this may involve increasing the discount rate printed on the coupon, or adding services or products that are eligible for the coupon (discount, etc.).
[0039] When the avatar receives the coupon (S110: YES), the avatar causes the NFT coupon issuing device 300 to issue an NFT coupon to the user (S111). If the avatar does not receive the coupon information (S110: NO), the avatar does not issue an NFT coupon, and therefore does not issue any instructions to the NFT coupon issuing device 300.
[0040] When the NFT coupon issuing device 300 issues an NFT coupon, it stores the transaction history of the coupon information in the blockchain PC 600 and issues the NFT coupon to the user terminal 500. Issuing an NFT coupon here means storing the image data of the NFT coupon and information on the coupon content (discount items, discount rate, etc.) in a server (not shown), storing the ownership history information in the blockchain PC 600, and storing information such as a private key for retrieving the NFT coupon or transaction history information in the wallet 501 of the user terminal 500.
[0041] In this way, the virtual space providing device 100 distributes coupon information to an avatar, and when the avatar receives the coupon information, the NFT coupon issuing device 300 can issue an NFT coupon. This NFT coupon is coupon information identified based on non-fungible information, and its transaction history information is recorded. Therefore, it is possible to distribute an appropriate NFT coupon according to the transaction history information.
[0042] By checking the transaction history information stored in the blockchain PC 600, the virtual space providing device 100 can determine the distribution status of coupon information to the avatar in the past, and can therefore control the distribution so that it is not redistributed to the user to whom it was most recently distributed.
[0043] Furthermore, when an NFT coupon is transferred to another person, the virtual space providing device 100 can reset the discount rate (individual discount rate) that was increased as a result of negotiations to a predetermined discount rate. Because NFT is a program, unlike existing coupons, restrictions can be imposed, such as resetting the discount rate to its initial value when the coupon is transferred. It is also possible to impose restrictions that make the coupon non-transferable. By configuring the coupons distributed to avatars as NFTs in this way, it is possible to restrict the transfer of the coupons.
[0044] On the other hand, if the avatar does not receive the coupon information of the initial proposal, the negotiation processing unit 104 of the virtual space providing device 100 can perform negotiation processing for the user.
[0045] 5 is a flowchart showing the negotiation process, which illustrates steps S108 to S110 in FIG.
[0046] The negotiation processing unit 104 of the virtual space providing device 100 transmits initially proposed coupon information to the user terminal 500 (S201). This initially proposed coupon information is based on the initial value of priority 1 in the proposal table 105, but the coupon content may be changed based on the maximum value depending on the receipt history. For example, for a user whose receipt history indicates that they would only accept coupons with a 20% discount, the negotiation processing unit 104 may generate and distribute coupon information for a 20% discount as an initial proposal.
[0047] The negotiation processing unit 104 then determines whether the user has accepted the coupon information in the initial proposal (S202). Whether or not the coupon information has been accepted is determined by the user terminal 500 sending information to that effect to the virtual space providing device 100.
[0048] If the coupon has been received, the process proceeds to step S111 in FIG. 4, and a request is made to the NFT coupon issuing device 300 to issue an NFT coupon.
[0049] If the user does not accept the coupon information, the negotiation processing unit 104 proposes the next proposal, which is coupon information that is more likely to be accepted (S204). That is, the negotiation processing unit 104 refers to the proposal table 105 and transmits the coupon information with the next proposal priority to the user terminal 500.
[0050] The negotiation processing unit 104 then repeats this process until the user receives the NPC or until there is no coupon information available as a next proposal (S202: NO, S203: YES). If there is no next proposal, negotiation may be conducted directly with the operator. That is, if there are no next proposals available in the proposal table 105, the negotiation processing unit 104 may switch the NPC to an operator of the company to which the NPC belongs so that the NPC can be operated, or may call the operator's avatar.
[0051] In this way, the coupon information can be changed in stages. By making these staged changes advantageous to the user, the user will be more likely to receive the coupon information.
[0052] The next proposed content could include the following coupon information:
[0053] Within the same coupon, the discount rate will be increased within the maximum value range ((It can be all the way up to the maximum, or somewhere in between). The same coupon means that the coupon target does not change. For example, if there is a discount on a hamburger, the hamburger does not change, but the discount rate changes.
[0054] Alternatively, the coupon may be switched to another coupon from the same company (cancelling the coupon that was presented and presenting a different coupon). Another coupon from the same company is a coupon for a different store in the same management group.
[0055] It is also possible to add another coupon (for example, to add fries to a hamburger coupon).
[0056] If there is nothing to present (the maximum discount has already been reached), the process may be terminated.
[0057] Furthermore, looking at the specific processing based on the proposal table 105 in Figure 2 (d), a proposal may be made based on the initial value of the coupon information indicated by priority 1, and the discount rate may be repeatedly increased until it reaches the maximum value indicated by priority 1 and then proposed (next proposal), or a proposal based on priority 2 may be made as the next proposal without increasing the discount rate.
[0058] In this way, the virtual space providing device 100 (negotiation processing unit 104) can use the proposal table 105 to change the priority of the coupon information and distribute coupon information that is more likely to be accepted by the user.
[0059] The above describes a rule-based coupon information generation process using the proposal table 105, but the presentation of coupon information can be performed in a more sophisticated manner.
[0060] That is, a proposal model optimized using reinforcement learning may be used to determine which action defined in the proposal table 105 should be taken.
[0061] 6 is a block diagram showing the functional configuration of a virtual space providing device 100a having an optimized proposed model 113. As shown in the figure, the virtual space providing device 100a includes a virtual space providing unit 101, a character control unit 102, a transaction history information acquisition unit 103, a negotiation history DB 111, a reinforcement learning unit 112, and a proposed model 113. This proposed model 113 holds information equivalent to the proposal table 105, and stores information equivalent to at least priority and proposed actions (coupon targets, initial values (discount rates), etc.).
[0062] When attempting to distribute coupon information to an avatar, the proposed model 113 can output appropriate behavior by acquiring and inputting the avatar's coupon reception status based on the avatar's coupon reception history. The character control unit 102 can generate coupon information based on the output behavior and have the NPC distribute the coupon information.
[0063] This proposed model 113 is updated (learned) by the reinforcement learning unit 112. Learning is performed based on the negotiation situation, and effective negotiations for the company are learned and coupon information is set based on, for example, the negotiation history information (whether or not a coupon was distributed recently, the past acceptance rate, etc.) stored in the negotiation history DB 111 and whether or not a coupon was presented based on the priority stored in the proposed model 113. The negotiation history information indicates the efficiency of coupon acceptance (acceptance rate and discount rate).
[0064] Next, the reinforcement learning unit 112 will be described. Basically, if the user receives the coupon information, it is considered a success, and if the user is successful, the reinforcement learning unit 112 receives a reward. The reinforcement learning unit 112 learns behavior (negotiation content) to maximize this reward. However, since companies want to distribute coupon information without reducing profits as much as possible, it is desirable to increase the reward when coupon information set to reduce profits is distributed. This allows the reinforcement learning unit 112 to learn how to distribute coupon information to users in a way that minimizes profit reduction as much as possible.
[0065] For example, the coupon distributor (company) wants to minimize the number of coupon types and the discount rate, so it sets a large reward for coupon information (with a small discount rate) that it wants users to receive.
[0066] On the other hand, users, contrary to the company, want to receive coupon information with a larger discount rate. Furthermore, users want to forcibly terminate the dialogue if the number of turns (the number of times the coupon information is re-presented) becomes unnecessarily long. Therefore, it is possible to set an additional reward if the dialogue ends in a short turn.
[0067] Therefore, the negotiation history DB 111 stores the avatar's past history of receiving coupon information (including cases where no coupon information was received) as a premise for reinforcement learning. The reinforcement learning unit 112 then inputs a state (S: State) to the proposed model 113, which is the target of reinforcement learning, and learns to output the optimal action (A: Action) in that state.
[0068] As a premise, the state indicates the negotiation status regarding coupon information when attempting to distribute it, such as whether or not to present the coupon information. In addition, the optimal actions for presenting the next proposal are considered to be as follows: (A) If there is nothing to present (the coupon has already been discounted to the maximum), terminate. (B) Increase the discount rate within the same coupon. (C) Switch to another coupon from the same company. (D) Add another coupon. Reinforcement learning is used to optimize which of these actions (A) to (D) should be taken.
[0069] FIG. 7 is a diagram of relationship information showing the relationship between coupon information and rewards. The reward relationship information is stored in the negotiation history DB 111, but may also be stored in another database (not shown). As shown in the figure, this reward relationship information associates coupon information, discount rates for coupon items (discount rates for hamburgers, cheese bars, fries, etc.), and rewards. Rewards are information used in reinforcement learning, and are learned so that the higher the reward, the higher the distribution priority. These rewards are set by the coupon distributor. Typically, a coupon distributor will want to distribute coupon information that they prefer (e.g., with a low discount rate), and will set a high reward for that preferred coupon information.
[0070] Figure 8 is a diagram applied to a specific example of reinforcement learning. As shown in the figure, the input indicates the negotiation status (internal state) of an NPC regarding the distribution of coupon information to a user, and indicates a state including information such as the distribution status of the coupon information and the discount rate for the coupon. The output (actions A to D in the figure) indicates the reward calculated when each of the NPC's possible actions is taken.
[0071] The states are as follows, and these negotiation states are input: (1) The presentation state of each coupon, which can be further broken down into priority order: (1.1) whether coupon information is presented or not, and (1.2) the discount rate.
[0072] For example, 1 is input for presentation and 0 for non-presentation. Furthermore, if the discount rate is 10%, 0.1 is input. (2) The internal state of the NPC may also include whether or not coupon information has been distributed to the user most recently. If it has been distributed, 1 is input, and if not, 0 is input. (3) Furthermore, the internal state of the NPC may also include the past coupon information reception rate of users. If the reception rate is 90%, 0.9 is input.
[0073] 9A shows the negotiation history of past coupon information. This negotiation history information of coupon information is stored in the negotiation history DB 111.
[0074] This coupon negotiation history information is information that associates a session ID, index, state, action, and reward. Numbers with the same session ID indicate a series of actions. The index is the state number. The state indicates by number whether a coupon was presented, whether the coupon was discounted, whether it was recently distributed, and the coupon acceptance rate of users in the past. The index is the number assigned to the content of the state, and if the content of the state is the same, the index will also be the same.
[0075] In FIG. 9(a), for example, (1.1.1) = 1 indicates that a coupon with priority 1 has been distributed. If this is 0, it indicates that no coupon has been distributed. (1.2.1) = 0.1 indicates that the discount rate for the coupon with priority 1 is 10%. For a 20% discount, it would be 0.2. Similarly, (1.1.2) = 0 indicates that a coupon with priority 2 has not been distributed. (1.2.2) = 0 indicates that the discount rate for the coupon with priority 2 is 0. (1.1.3) to (1.2.4) for priorities 3 and 4 are also information similar to the above. (2) = 0 indicates that a coupon has been distributed to the most recent user. (3) = 0.7 indicates that the coupon acceptance rate for past users is 70%.
[0076] FIG. 9(b) is a diagram visualizing the negotiation history. According to the negotiation history, when the session ID is 1 and the state is S0 in the index column, action B (increasing the coupon discount rate) is performed. Thereafter, when the state is S1, action B (further increasing the coupon discount rate) is performed, and when the state is S2, action A (terminating coupon distribution) is performed. The reward is the reward of the coupon information at the time of receipt. In state S2, a reward of 100 is given, and the reward is granted based on the diagram of the relationship between coupon distribution and rewards. In FIG. 9(b), action A is performed in state S2, indicating that the user received this coupon information because they had no next options.
[0077] Here, we will explain the problems encountered in learning and how to solve them. FIG. 10 is a schematic diagram showing the relationship between rewards and coupon receipt. Even in the same state (state S2 is a coupon for 10% off a hamburger), some users receive the coupon (reward 100) and others do not (reward 0). In the visualized negotiation history diagram of FIG. 10, symbol P indicates that in state S2, coupon information was presented to four users, and only one user accepted it. In other words, action A indicates that there is no next plan, and distribution of coupon information ends. At that point, the user receives the coupon information and is returned a reward of 100. On the other hand, there are also cases where the user does not accept the coupon, in which case a reward of 0 is returned. Taking such situations into consideration, an expected value is used as the reward during learning. Referring to the portion of symbol P in FIG. 9, the expected value is 100 / 4 = 25.
[0078] What this means is as follows: Companies want to distribute coupons at the lowest possible cost, but users may not accept them. On the other hand, users may want to receive coupons with the best possible conditions, but the company's motivation to distribute them (i.e., reward) may be zero. Therefore, by repeating this distribution process, learning can be expected to progress so that the conditions settle at just right (state S5 in Figure 10).
[0079] One problem with learning is the frequent proposals. It is undesirable for the user to receive repeated proposals over and over again. Therefore, we impose constraints. Specifically, additional rewards are awarded if the dialogue is short. For example, if the first proposal is +15, the second +10, the third +5, and the fourth (and subsequent) +0, we can expect learning to progress so that negotiations are shortened. Figure 11 is a diagram that illustrates this concept. In Figure 11, for example, if a user receives a coupon in state S0, 15 is added to the reward of the coupon at that time. Also, if a user receives a coupon in state S2, 5 is added to the reward of the coupon at that time. Although the diagram shows an example in which 5 is added even if the user does not receive a coupon, it could be set to 0. Furthermore, if an additional coupon is received, the reward is 0, but the expected value could be calculated with 5 added.
[0080] Next, the specific learning process in the reinforcement learning unit 112 will be described. As described above, states S0, S1, ..., actions = [A, B, C, D], and reward R are determined from the negotiation history. Note that this reward is the expected value described above, but is not limited to this. It may also be the reward actually given.
[0081] In the proposed model 113 of the present disclosure, Q(S, a) in Q-learning is configured as a reward function (Q function) that outputs a reward expected when action a is performed in state S. Fig. 12 is a diagram schematically illustrating the update process of the Q function.
[0082] As shown in Figure 12, given a state S and an action a (in this case [A, B, C, D]), the current state (at time t) is St, the selected action is a_t, and the reward Rt+1 is obtained, reaching the state St+1. The Q function is updated (learned) as follows: Q(St, a_t) = (1-α) Q(St, a_t) + α [ Rt+1 + γ Max( Q(St+1, a) ) ]
[0083] Here, α is the learning rate, and 0 indicates no learning at all. γ is the discount rate, which is a real number between 0 and 1. Also, Max(Q(St+1,)) indicates the maximum reward expected from the state of St+1.
[0084] While repeatedly selecting an action based on the proposed model 113, the character control unit 102 operates the NPC to perform the action selected in each of states S0, S1, .... The reinforcement learning unit 112 then updates the Q function for the results of that action. By repeating this process, an appropriate Q function can be obtained.
[0085] It is known that the Q function converges even when it is random, but it is also possible to select randomly with a certain probability ε, or to select an action that maximizes the current Q function (called ε-greedy).
[0086] The updated Q function is included in the proposed model 113. The Q function is a function that receives the status of negotiation of coupon information with an avatar and outputs a reward for each of actions A to D, and the proposed model 113 outputs the action with the highest reward.
[0087] The character control unit 102 generates coupon information based on the behavior output from the proposed model 113 and causes the NPC to distribute it.
[0088] In this way, by using the proposed model 113 based on reinforcement learning (for example, Q-learning), it is possible to determine efficient actions when distributing coupon information.
[0089] In the above reinforcement learning, appropriate actions are learned depending on the negotiation situation (internal state), but the initial value (discount rate) for these actions may also be learned.
[0090] Next, we will explain the effects of the virtual space providing device 100 of the present disclosure. The virtual space providing device 100 of the present disclosure includes a virtual space providing unit 101 that provides a virtual space including a user's avatar and a computer character, and a character control unit 102 that controls the distribution of distribution information to the avatar of a computer character (NPC) based on information about the user's behavior in the real world.
[0091] The character control unit 102 then acquires the user's interest information based on the behavior information and controls the NPC.
[0092] This configuration enables distribution of distribution information according to a user's behavior in the real world, thereby enabling efficient distribution of coupon information. For example, it is possible to grasp a user's behavior in the real world based on information such as the user's cashless payment, and accordingly grasp the user's interests. As a result, by distributing coupon information according to the user's behavior in the real world as distribution information, it is possible to efficiently guide the user to stores in the real world. For example, the coupon information is a coupon that the user can use when purchasing a product or service in the real world, and is information for providing discounts, etc. It is preferable that this information be of interest to the user, so that it is easy for the user to receive it.
[0093] The character control unit 102 may control computer characters such as NPCs based on the user's receipt history of distribution information. Some users may not receive distribution information such as coupon information. Therefore, it is efficient to perform distribution operations based on the receipt history.
[0094] Furthermore, when the negotiation processing unit 104 distributes coupon information to an avatar, if the avatar does not accept the coupon information, the negotiation processing unit 104 distributes modified coupon information (modified distribution information) of the coupon information. For example, the negotiation processing unit 104 performs negotiation processing such as increasing the discount rate in the coupon information or adding products eligible for discount.
[0095] The virtual space providing device 100 further includes a proposal table 105 (distribution information management table) that associates coupon information with the priority of its distribution. The negotiation processing unit 104 references the proposal table 105 and controls distribution according to the priority.
[0096] This allows an attempt to distribute corrected coupon information as a substitute even if the user does not receive the coupon information, thereby increasing the probability that the user will receive it.
[0097] The present disclosure further includes a proposal model 113 that calculates rewards for each of a plurality of suggested actions based on the negotiation status of the coupon information with the user, such as whether or not the coupon information is presented and the discount rate, and a plurality of suggested actions in response to the user's response to the coupon information (increasing the discount rate, switching to another product, etc.), and derives one suggested action based on the rewards. The character control unit 102 performs control based on the one suggested action derived by the proposal model 113.
[0098] This allows proposals to be made using machine-learned proposal models, enabling efficient distribution of coupon information.
[0099] The proposal model 113 is updated based on the negotiation status of the coupon information with the user, the plurality of proposed actions, and the reward when the coupon information is received. The proposal model 113 is trained by reinforcement learning.
[0100] The device and method of the present disclosure have the following configuration.
[0101] [1] A device comprising: a virtual space providing unit that provides a virtual space including a user's avatar and a computer character; and a character control unit that controls distribution of distribution information to the avatar to the computer character based on behavioral information of the user in the real world.
[0102] [2] The device according to [1], wherein the character control unit acquires interest information of the user based on the behavior information and controls the computer character.
[0103] [3] The device according to [1], wherein the character control unit controls the computer character based on the user's receipt history of the distribution information.
[0104] [4] The device according to any one of [1] to [3], wherein the distribution information is information based on interest information of the user.
[0105] [5] The device according to any one of [1] to [4], wherein the distribution information is coupon information that the user can use when receiving a product or service in the real world.
[0106] [6] The device according to any one of [1] to [5], further comprising a negotiation unit that controls the character control unit to perform a distribution operation of modified distribution information obtained by modifying the distribution information when the distribution operation of the distribution information is performed for the avatar and the avatar does not receive the distribution information.
[0107] [7] The device according to [6], further comprising a distribution information management table in which the distribution information and distribution priorities are associated with each other, and the negotiation unit refers to the distribution information management table to control distribution in accordance with the priorities.
[0108] [8] The device described in [1] further comprises a proposal model that calculates a reward for each of the multiple proposed actions based on a negotiation status regarding the information to be distributed to the user and multiple proposed actions in response to the user's response to the distributed information, and derives one proposed action based on the rewards, and the character control unit performs control based on the one proposed action derived by the proposal model.
[0109] [9] The device according to [8], wherein the proposal model is updated based on a negotiation status of the distribution information to the user, the plurality of proposed actions, and a reward if the distribution information is received.
[0110]
[10] A method comprising: a virtual space providing step of providing a virtual space including a user's avatar and a computer character; and a character control step of controlling the computer character to distribute information to the avatar based on behavioral information of the user in the real world.
[0111] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.
[0112] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.
[0113] For example, the virtual space providing device 100 (hereinafter including the virtual space providing device 100a) according to an embodiment of the present disclosure may function as a computer that performs processing of the virtual space providing method of the present disclosure. Fig. 13 is a diagram showing an example of the hardware configuration of the virtual space providing device 100 according to an embodiment of the present disclosure. The above-described virtual space providing device 100 may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.
[0114] In the following description, the term "device" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the virtual space providing device 100 may be configured to include one or more of the devices shown in the figure, or may be configured to exclude some of the devices.
[0115] Each function of the virtual space providing device 100 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.
[0116] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, registers, etc. For example, the character control unit 102, negotiation processing unit 104, etc. described above may be realized by the processor 1001.
[0117] The processor 1001 also loads programs (program code), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the character control unit 102 and the negotiation processing unit 104 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented on one or more chips. The programs may also be transmitted from a network via a telecommunications line.
[0118] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be referred to as a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing a virtual space providing method according to one embodiment of the present disclosure.
[0119] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.
[0120] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, or a communication module. The communication device 1004 may include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the virtual space providing unit 101 and the transaction history information acquiring unit 103 may be realized by the communication device 1004. The communication device 1004 may be implemented with a transmitter and a receiver that are physically or logically separated from each other.
[0121] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).
[0122] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.
[0123] The virtual space providing device 100 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.
[0124] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.
[0125] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.
[0126] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.
[0127] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).
[0128] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).
[0129] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.
[0130] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.
[0131] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.
[0132] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.
[0133] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.
[0134] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, may be expressed using relative values from a predetermined value, or may be expressed using other corresponding information. For example, a radio resource may be indicated by an index.
[0135] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.
[0136] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.
[0137] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.
[0138] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.
[0139] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.
[0140] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."
[0141] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.
[0142] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.
[0143] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.
[0144] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."
[0145] 100, 100a...virtual space providing device, 101...virtual space providing unit, 102...character control unit, 103...transaction history information acquisition unit, 104...negotiation processing unit, 105...proposal table, 111...negotiation history DB, 112...reinforcement learning unit, 113...proposed model, 113...proposed model, 400...coupon information DB, 500...user terminal, 501...wallet, 600...blockchain PC, 700...user attribute information DB.
Claims
1. A device comprising: a virtual space providing unit that provides a virtual space including a user's avatar and a computer character; and a character control unit that controls the distribution of information to the avatar to the computer character based on behavioral information or interest information of the user in the real world.
2. The device according to claim 1, wherein the character control unit acquires interest information of the user based on the behavior information and controls the computer character.
3. The device according to claim 1, wherein the character control unit controls the computer character based on the user's receipt history of the distributed information.
4. The device according to claim 1, wherein the distribution information is information based on interest information of the user.
5. The device according to claim 1, wherein the distribution information is coupon information that can be used by the user when purchasing a product or service in the real world.
6. The device according to claim 1, further comprising a negotiation unit that controls the character control unit to perform a distribution operation of modified distribution information obtained by modifying the distribution information when the distribution operation is performed for the avatar and the avatar does not receive the distribution information.
7. The device according to claim 6, further comprising a distribution information management table that associates the distribution information with a distribution priority, and the negotiation unit references the distribution information management table and controls distribution in accordance with the priority.
8. The device described in claim 1, further comprising a proposal model that calculates rewards for each of the multiple proposed actions based on a negotiation status regarding the information distributed to the user and multiple proposed actions in response to the user's response to the distributed information, and derives one proposed action based on the rewards, and the character control unit performs control based on the one proposed action derived by the proposal model.
9. The device of claim 8, wherein the proposal model is updated based on a negotiation status of the distribution information to the user, the plurality of proposal actions, and a reward if the distribution information is received.
10. A method comprising: a virtual space providing step of providing a virtual space including a user's avatar and a computer character; and a character control step of controlling the distribution of distribution information to the avatar to the computer character based on behavioral information of the user in the real world.
Citation Information
Patent Citations
Creating, maintaining, and growing a music-themed virtual world
JP2023527763A
Real world interaction with virtual world privileges
US20070073614A1