Information processing device, information processing method, and information processing program

The information processing apparatus addresses the challenge of limited data in offline reinforcement learning by generating supplementary data using domain knowledge and LLMs, facilitating effective policy learning for unknown states and actions in web advertising.

WO2025154457A1PCT designated stage expired Publication Date: 2025-07-24SONY GROUP CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/044508
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2024-12-17
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing offline reinforcement learning methods struggle with insufficient historical data, making it difficult to learn appropriate policies for unknown states and actions, particularly in applications like web advertising where data is limited.

Method used

An information processing apparatus that generates additional learning data using domain knowledge and Large Language Models (LLM) to supplement action history data, enabling effective offline reinforcement learning even with limited historical data.

Benefits of technology

Enables accurate policy learning for unknown states and actions, allowing reinforcement learning to be applied in scenarios where conventional methods fail, such as decision-making in marketing and web advertising automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024044508_24072025_PF_FP_ABST
    Figure JP2024044508_24072025_PF_FP_ABST
Patent Text Reader

Abstract

This information processing device comprises an acquisition unit that acquires first data that is behavior history data that is to be used as training data for offline reinforcement learning, a generation unit that generates second data that is data that is different from the first data and can be used as training data for the offline reinforcement learning, and a training unit that trains a model by offline reinforcement learning that uses the first data and the second data.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and information processing program

[0001] The present disclosure relates to an information processing device, an information processing method, and an information processing program.

[0002] Reinforcement learning is known as one of the learning algorithms in machine learning. For example, a method for training a character by a learning process using reinforcement learning has been proposed (for example, Patent Document 1).

[0003] Patent No. 7061238

[0004] However, the conventional technology has room for improvement in terms of offline reinforcement learning, which is a reinforcement learning method that uses only historical data, which is data collected in the past. For example, the conventional technology may have difficulty in properly performing offline reinforcement learning when the historical data used for offline reinforcement learning is insufficient. Therefore, it is desirable to make it possible to properly perform offline reinforcement learning even when the historical data is insufficient.

[0005] Therefore, the present disclosure proposes an information processing device, an information processing method, and an information processing program that enable offline reinforcement learning to be performed appropriately.

[0006] In order to solve the above problem, an information processing device of one embodiment according to the present disclosure includes an acquisition unit that acquires first data, which is behavioral history data used as learning data for offline reinforcement learning; a generation unit that generates second data, which is data other than the first data and which can be used as learning data for the offline reinforcement learning; and a learning unit that learns a model by the offline reinforcement learning using the first data and the second data.

[0007] 1 is a diagram illustrating an example of information processing according to an embodiment of the present disclosure. FIG. 2 is a diagram illustrating an example configuration of an information processing device according to an embodiment of the present disclosure. FIG. 3 is a flowchart illustrating a procedure of information processing according to a first algorithm. FIG. 4 is a flowchart illustrating a procedure of information processing according to the first algorithm. FIG. 5 is a diagram illustrating an overview of the first algorithm. FIG. 6 is a flowchart illustrating a procedure of information processing according to the second algorithm. FIG. 7 is a diagram illustrating an overview of the second algorithm. FIG. 8 is a diagram illustrating an overview of web advertising management. FIG. 9 is a diagram illustrating a flow of advertising management using a DSP. FIG. 10 is a diagram illustrating an example of data in the case of advertising management based on a monthly budget. FIG. 11 is a diagram illustrating an example of actions in the case of advertising management based on a monthly budget. FIG. 12 is a diagram illustrating an example of data in the case of advertising management based on a daily budget. FIG. 13 is a diagram illustrating an example of actions in the case of advertising management based on a daily budget. FIG. 14 is a diagram illustrating an example of a prompt used to generate second data. FIG. 15 is a diagram illustrating an example of a display related to domain knowledge. FIG. 16 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of an information processing device. FIG. 17 is a diagram illustrating an overview of online reinforcement learning and offline reinforcement learning.

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the information processing device, information processing method, and information processing program according to the present application are not limited to these embodiments. In addition, in the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The present disclosure will be described in the following order of items: 1. Embodiments 1-1. Technical Overview (Background, etc.) 1-2. Overview of Information Processing According to Embodiments of the Present Disclosure 1-3. Configuration of Information Processing Device According to Embodiments 1-4. Algorithm Examples 1-4-1. First Algorithm 1-4-1-1. Processing Example Based on First Algorithm 1-4-1-2. Overview of First Algorithm 1-4-2. Second Algorithm 1-4-2-1. Processing Example Based on Second Algorithm 1-4-2-2. Overview of Second Algorithm 1-5. Application Examples (Web Advertising Management) 1-5-1. Advertising Management Example by DSP 1-5-2. Example for Monthly Budget 1-5-3. Example for Daily Budget 1-5-4. Processing Example 1-5-5. Prompt Example 1-5-6. Example of Use of Domain Knowledge 1-5-7. Other Application Examples 2. Other Configuration Examples 3. Others 4. Effects of the Present Disclosure 5. Hardware Configuration

[0010] <1. Embodiment> <1-1. Technical Overview (Background, etc.)> Prior to describing information processing and the like according to embodiments of the present disclosure, a technical overview of reinforcement learning, etc., which is the premise of the information processing and the like, will be described. Note that detailed descriptions of conventional techniques such as ordinary machine learning will be omitted as appropriate.

[0011] Reinforcement learning is a field of machine learning (learning algorithm) in which an intelligent agent (also simply referred to as an "agent") in a certain environment observes the current state and determines what action to take in order to maximize the profit (also referred to as a "reward"). In this way, reinforcement learning performs a learning process based on three pieces of information: state, action (sometimes referred to as "action"), and reward (profit). For example, the state is information that indicates the situation in which the agent is placed. The action is information that indicates the behavior of the agent. The reward is information that indicates the evaluation (goodness) of the action in that state. Note that these points are the same as conventional reinforcement learning, so a detailed description will be omitted.

[0012] For example, as shown in FIG. 18, conventional reinforcement learning (also called "online reinforcement learning") that interacts with the environment online requires learning while interacting with the environment, which results in high costs. FIG. 18 is a diagram illustrating an overview of online reinforcement learning and offline reinforcement learning. In this way, online reinforcement learning allows an intelligent agent to learn a policy while trying out its own actions, but it is costly and in practice it is often not easy to try out actions. For example, in fields such as investment and medical treatment, the high costs and risks of trying out actions can make it difficult to apply online reinforcement learning.

[0013] 18, offline reinforcement learning learns a policy from existing behavior history data. The policy here is, for example, a function (model) used to determine which behavior should be taken in which state. For example, offline reinforcement learning can learn a policy from behavior history data.

[0014] As such, offline reinforcement learning has the advantage of broadening the range of applications because it does not require interaction with the environment. However, offline reinforcement learning has the problem of not being able to effectively learn states or actions that are not in the action history. As such, in offline reinforcement learning, in real situations, the number and patterns of action history may be limited, and in such cases it is difficult to select good actions in unknown states.

[0015] <1-2. Overview of Information Processing According to an Embodiment of the Present Disclosure> Therefore, the information processing device 100 executes information processing as outlined in FIG. 1. FIG. 1 is a diagram illustrating an example of information processing according to an embodiment of the present disclosure. The information processing illustrated in FIG. 1 is realized by the information processing device 100 (see FIG. 2). Details of the configuration and the like will be described later, but the information processing device 100 is an example of a computer (learning device) that executes processing related to offline reinforcement learning.

[0016] For example, Fig. 1 is a diagram showing an outline of the flow of processing related to offline reinforcement learning executed by an information processing device 100. Fig. 1 illustrates an example in which the application domain (field) of the processing related to offline reinforcement learning executed by the information processing device 100 is the operation of advertisement distribution on the Internet, such as the Web (also simply referred to as "advertising operation"). Note that Fig. 1 illustrates an outline of the flow of processing executed by the information processing device 100, and details of algorithms and the like related to specific processing will be described later.

[0017] The information processing device 100 generates second data, which is data other than first data, which is behavior history data, and is data that can be used as learning data for offline reinforcement learning (step S1). The information processing device 100 acquires the first data FD and generates second data SD other than the acquired first data FD. The information processing device 100 generates the second data SD using domain knowledge DD. For example, the domain knowledge DD is information indicating knowledge related to the field of advertising management.

[0018] The domain knowledge here refers to information indicating knowledge, insight, etc. For example, domain knowledge related to a field may include various information related to that field. In FIG. 1, the field is advertising management, and the domain knowledge DD includes various information related to advertising management.

[0019] The domain knowledge DD includes domain knowledge related to a state. For example, the domain knowledge DD includes information (knowledge) indicating the relationship between price and consumption amount related to advertising management. For example, the domain knowledge DD includes "if price weight increases, then digestive amount will increase" as knowledge related to a state in advertising management.

[0020] Furthermore, the domain knowledge DD includes domain knowledge regarding combinations of states and actions. For example, the domain knowledge DD includes domain knowledge indicating the actions to be taken in a certain state. For example, the domain knowledge DD includes information (knowledge) indicating the relationship between the target index value and the budget related to advertising operations. For example, the domain knowledge DD includes "if CPA (cost per acquisition) is lower than goal CPA, then increase daily goal budget" as insight regarding the combination of states and actions in advertising operations. Note that CPA (cost per acquisition) is an example of an advertising index, and is the value obtained by dividing the amount spent by the number of conversions (= amount spent / number of conversions).

[0021] The domain knowledge DD includes domain knowledge related to behavior. For example, the domain knowledge DD includes information (knowledge) that indicates how to change price weights related to advertising operations. For example, the domain knowledge DD includes "if you change price weight at the first time, then set price weight at 1.2" as insight related to behavior in advertising operations.

[0022] Furthermore, domain knowledge DD includes domain knowledge related to combinations of actions and rewards. For example, domain knowledge DD includes domain knowledge that indicates the relationship between a certain action and reward. For example, domain knowledge DD includes domain knowledge that indicates the relationship between frequency (number of times an ad is exposed) and click-through rate. For example, domain knowledge DD includes "if you increase frequency, then click-through rate decreases" as knowledge related to combinations of actions and rewards in advertising operations.

[0023] Note that the above is merely an example of domain knowledge, and domain knowledge for each field may include not only the above-mentioned knowledge about states, state-actions, actions, action-rewards, etc., but also various other information related to that field. For example, domain knowledge may include knowledge about state-action-reward combinations. For example, domain knowledge may include knowledge about state-action-previous state combinations.

[0024] 1 , the information processing device 100 generates the second data SD using the domain knowledge DD related to advertising management as described above. For example, the information processing device 100 generates data including states and actions corresponding to the states as the second data SD. For example, the information processing device 100 generates the second data SD by randomly generating states and generating actions corresponding to the generated states.

[0025] The information processing device 100 may generate the second data using various information other than domain knowledge. The information processing device 100 may generate the second data using LLMs (Large Language Models). In this case, the information processing device 100 may generate the second data by inputting a prompt to the LLM instructing it to generate the second data and causing the LLM to output second data corresponding to the input prompt. An example of generating second data using the LLM will be described later.

[0026] Then, the information processing device 100 learns a model through offline reinforcement learning using the first data and the second data (step S2). For example, the information processing device 100 learns a model based on a Student-Teacher structure. In FIG. 1, the information processing device 100 learns a student network SN and a teacher network TN through offline reinforcement learning using the first data and the second data.

[0027] In this way, the information processing device 100 generates data (second data) other than the behavior history data (first data) used as learning data for offline reinforcement learning. Such second data is data that can be used as learning data for offline reinforcement learning. The information processing device 100 then performs offline reinforcement learning using the first data and the second data. As a result, even if there is a small amount of behavior history data (first data), the information processing device 100 can appropriately execute the learning process by performing offline reinforcement learning using the second data. Therefore, the information processing device 100 can appropriately execute offline reinforcement learning.

[0028] For example, conventional offline reinforcement learning has a problem in that it cannot effectively learn unknown (state, action) pairs. Conventional offline reinforcement learning methods (also referred to as "existing methods") address this problem by imposing a penalty on actions that are not included in the behavior history data used for learning. However, the existing methods are unable to select good actions in unknown states. Meanwhile, the information processing device 100 generates data (second data) other than the behavior history data (first data) and uses the generated second data to solve the above-described problems with the existing methods and enable appropriate offline reinforcement learning.

[0029] Generally, the principle of machine learning is that learning is not possible in areas where there is no data, making it difficult to make accurate inferences. In particular, reinforcement learning, as a learning mechanism, tends to have an increasingly strong influence if some inference results are poor, with the influence of areas where there is no data being particularly large. As mentioned above, conventional technologies mainly use approaches such as imposing constraints to prevent actions that lead to unknown states (scenes), but the learned policy tends to be limited to the range of existing behavioral history, and learning is difficult when there is little behavioral history data.

[0030] For example, when the application area is a business-related field and the target is the automation of business behavior, the amount of behavioral history data is often quite limited. However, when adjusting the budget and distribution settings for advertising operations in marketing, it may be possible to acquire domain knowledge such as the insights and knowledge described above.

[0031] Therefore, the information processing device 100 uses domain knowledge in addition to behavioral history data to generate information (second data) in a state where there is no behavioral history data, and by utilizing the generated second data, the performance of offline reinforcement learning can be improved. In this way, the information processing device 100 can learn a highly accurate policy (model) by using domain knowledge even in a situation where there is little behavioral history data. For example, the information processing device 100 can learn a policy (model) that enables behavior that will increase future rewards even in a situation (situation) where there is no particular behavioral history data but domain knowledge.

[0032] This enables the information processing device 100 to apply reinforcement learning to areas where it has been difficult to apply it up until now, such as decision-making in marketing. As described above, conventional offline reinforcement learning cannot learn a good behavior policy in an unknown area of ​​existing behavior data, but the information processing device 100 can learn a model (policy) for selecting good behavior even in an unknown area by utilizing domain knowledge.

[0033] <1-3. Configuration of information processing device according to embodiment> Next, a configuration of an information processing device 100, which is an example of an information processing device that executes information processing according to an embodiment, will be described. Fig. 2 is a diagram showing an example of the configuration of an information processing device according to an embodiment of the present disclosure. For example, the information processing device 100 shown in Fig. 2 is an example of an information processing device.

[0034] 2, the information processing device 100 includes a communication unit 11, an input unit 12, an output unit 13, a storage unit 14, and a control unit 15. In the example of Fig. 2, the information processing device 100 includes an input unit 12 (e.g., a keyboard, a mouse, etc.) that accepts various operations from an administrator of the information processing device 100, and an output unit 13 (e.g., a liquid crystal display, etc.) that outputs (displays, etc.) various information.

[0035] The communication unit 11 is realized by, for example, a network interface card (NIC), a communication circuit, etc. The communication unit 11 is connected to a communication network N (a network such as the Internet) by wire or wirelessly, and transmits and receives information to and from other devices, etc. via the communication network N.

[0036] The input unit 12 receives various operations input from the user. The input unit 12 accepts input from the user. The input unit 12 accepts input of information used to generate a mesh model by the user. The input unit 12 may accept various operations from the user via a keyboard, mouse, or touch panel provided in the information processing device 100.

[0037] The output unit 13 outputs various types of information. The output unit 13 has a display device (display unit) such as a display, and displays various types of information. The output unit 13 displays the information generated by the generation unit 152. The output unit 13 displays the processing results using a model related to the learning process by the learning unit 153. The output unit 13 displays the processing results using the student network learned by the learning unit 153.

[0038] Furthermore, the output unit 13 is not limited to a display function and may have a function of outputting information in various ways. For example, the output unit 13 may have a function of outputting information as sound. For example, the output unit 13 may have an audio output unit such as a speaker that outputs sound.

[0039] The storage unit 14 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 14 has a first data storage unit 141 and a domain knowledge storage unit 142.

[0040] The first data storage unit 141 according to the embodiment stores first data, which is behavior history data used as learning data for offline reinforcement learning. For example, the first data storage unit 141 stores the first data including at least a state and an action corresponding to the state. For example, the first data storage unit 141 stores a combination of an action, a state, and a reward as the first data. For example, the first data storage unit 141 stores the first data including a state, an action corresponding to the state, a state before transitioning to the state when the action is performed, and a reward.

[0041] The first data storage unit 141 may store various types of information depending on the purpose, not limited to the above.

[0042] The domain knowledge storage unit 142 according to the embodiment stores information about domain knowledge, which is knowledge about each field. For example, the domain knowledge storage unit 142 stores information about domain knowledge about the field in which the model is used. The domain knowledge storage unit 142 stores information about the domain knowledge of each field in association with the field. For example, the domain knowledge storage unit 142 stores information about the domain knowledge of each field in association with information identifying the field (field ID).

[0043] The domain knowledge storage unit 142 is not limited to the above, and may store various types of information depending on the purpose.

[0044] The storage unit 14 also stores various other information besides the above. For example, the storage unit 14 stores information generated by the generation unit 152. The storage unit 14 stores second data generated by the generation unit 152. For example, the storage unit 14 stores a model used in processing. The storage unit 14 stores a model used to generate the second data. The storage unit 14 stores an LLM used to generate the second data.

[0045] Furthermore, for example, the storage unit 14 stores information used in processing by a first algorithm described below. For example, the storage unit 14 stores a function used in processing by the first algorithm. For example, the storage unit 14 stores a judgment condition used in processing by the first algorithm. For example, the storage unit 14 stores information used in processing by a second algorithm described below. For example, the storage unit 14 stores a function used in processing by the second algorithm. For example, the storage unit 14 stores a judgment condition used in processing by the second algorithm.

[0046] The storage unit 14 acquires information obtained by the information processing executed by the information processing device 100. For example, the storage unit 14 stores a model learned by the learning unit 153. The storage unit 14 stores a teacher network and a student network learned by the learning unit 153.

[0047] The control unit 15 is realized, for example, by a central processing unit (CPU), a micro processing unit (MPU), or the like executing a program (for example, an information processing program according to the present disclosure) stored inside the information processing device 100 using a random access memory (RAM) or the like as a working area. The control unit 15 is also a controller, and may be realized, for example, by an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA).

[0048] 2, the control unit 15 has an acquisition unit 151, a generation unit 152, a learning unit 153, and a transmission unit 154, and realizes or executes the functions and actions of the information processing described below. Note that the internal configuration of the control unit 15 is not limited to the configuration shown in FIG. 2, and may be any other configuration as long as it performs the information processing described below.

[0049] The acquisition unit 151 acquires various types of information. The acquisition unit 151 acquires various types of information from an external information processing device. The acquisition unit 151 acquires various types of information from the storage unit 14. The acquisition unit 151 acquires information accepted by the input unit 12.

[0050] The acquisition unit 151 acquires various types of information from the storage unit 14. The acquisition unit 151 acquires first data from the first data storage unit 141. The acquisition unit 151 acquires information related to domain knowledge from the domain knowledge storage unit 142. The acquisition unit 151 acquires a model to be used for processing from the storage unit 14. The acquisition unit 151 acquires an LLM from the storage unit 14.

[0051] The acquisition unit 151 acquires first data, which is behavior history data used as learning data for offline reinforcement learning. The acquisition unit 151 acquires domain knowledge, which is knowledge related to the field in which the model is used. The acquisition unit 151 acquires LLMs, which are used to generate second data. The acquisition unit 151 acquires various information generated by the generation unit 152.

[0052] The generation unit 152 performs a generation process to generate various types of information. The generation unit 152 generates various types of information based on the information acquired by the acquisition unit 151. The generation unit 152 generates various types of information based on the information stored in the storage unit 14.

[0053] The generation unit 152 generates second data, which is data other than the first data and is data that can be used as learning data for offline reinforcement learning. The generation unit 152 generates the second data using domain knowledge. With this configuration, the information processing device 100 can generate the second data using domain knowledge, thereby making it possible to appropriately perform offline reinforcement learning. The generation unit 152 generates the second data using an LLM. With this configuration, the information processing device 100 can generate the second data using an LLM, thereby making it possible to appropriately perform offline reinforcement learning.

[0054] The generation unit 152 generates second data including at least states and actions corresponding to the states. The generation unit 152 generates the second data by generating states and actions corresponding to the generated states. The generation unit 152 randomly generates states and generates actions corresponding to the generated states. With this configuration, the information processing device 100 can generate second data including states and actions corresponding to the states, thereby making it possible to appropriately perform offline reinforcement learning.

[0055] The generation unit 152 generates an action corresponding to a state based on the state and the domain in which the model is used. The generation unit 152 generates an action corresponding to the state based on the state and a generative model that generates an action corresponding to the state. With this configuration, the information processing device 100 can generate an appropriate action corresponding to the state, thereby making it possible to appropriately perform offline reinforcement learning.

[0056] The generation unit 152 may generate various types of information. The generation unit 152 may generate various types of information to be displayed. For example, the generation unit 152 may generate content such as content CT1. In this case, the generation unit 152 generates information (images) related to the screen by appropriately using various conventional techniques related to images. The generation unit 152 generates images by appropriately using various conventional techniques related to GUIs. For example, the generation unit 152 may generate images using CSS, JavaScript (registered trademark), HTML, or any language capable of describing information processing such as the above-described information display and operation reception.

[0057] The learning unit 153 learns various types of information. The learning unit 153 learns various types of information based on information from an external information processing device and information stored in the storage unit 14. The learning unit 153 learns various types of information based on the first data and the second data stored in the storage unit 14. The learning unit 153 stores a model generated by learning in the storage unit 14. The learning unit 153 stores a model updated by learning in the storage unit 14.

[0058] The learning unit 153 performs a learning process. The learning unit 153 performs various types of learning. The learning unit 153 learns various types of information based on the information acquired by the acquisition unit 151. The learning unit 153 learns (generates) a model. The learning unit 153 learns various types of information such as a model. The learning unit 153 generates a model through learning. The learning unit 153 learns the model using various machine learning techniques. For example, the learning unit 153 learns parameters of the model (network). The learning unit 153 learns the model using various machine learning techniques.

[0059] The learning unit 153 performs a learning process based on the learning data stored in the storage unit 14. The learning unit 153 performs a learning process using the first data stored in the storage unit 14. The learning unit 153 performs a learning process using the second data stored in the storage unit 14.

[0060] The learning unit 153 learns a model through offline reinforcement learning using the first data and the second data. With this configuration, the information processing device 100 can generate second data other than the first data, which is behavior history data, and perform offline reinforcement learning using the generated second data as well, thereby making it possible to appropriately perform offline reinforcement learning. The learning unit 153 learns a student network and a teacher network, which are models used in the inference process. With this configuration, the information processing device 100 can appropriately learn a student network and a teacher network, which are models used in the inference process, making it possible to appropriately perform offline reinforcement learning.

[0061] The learning unit 153 learns a model based on a first algorithm. The learning unit 153 learns a student network and a teacher network based on the first algorithm. For example, the learning unit 153 learns a student network SN and a teacher network TN through offline reinforcement learning processing based on the first algorithm. For example, the student network SN is a policy (model) used during inference processing. For example, the student network SN and the teacher network TN have the same network structure.

[0062] The learning unit 153 executes a first learning process to learn a teacher network using the second data, and executes a second learning process to learn a student network using the teacher network learned by the first learning process and the first data. With this configuration, the information processing device 100 can appropriately learn the student network through the first learning process and the second learning process, thereby enabling appropriate offline reinforcement learning. In the second learning process, if the action selected by the teacher network differs from the action selected by the student network, the learning unit 153 executes an update process for the teacher network. If the index value of the action selected by the student network is higher than the index value of the action selected by the teacher network, the learning unit 153 updates the teacher network. With this configuration, the information processing device 100 can appropriately update the teacher network, thereby enabling appropriate offline reinforcement learning.

[0063] The learning unit 153 learns a model based on the second algorithm. The learning unit 153 learns a student network and a teacher network based on the second algorithm. For example, the learning unit 153 learns the student network S N and the teacher network T N by offline reinforcement learning processing based on the second algorithm.

[0064] The learning unit 153 executes a first learning process to learn a teacher network and a student network using the second data, and learns a student network using the teacher network and student network learned by the first learning process and the first data. In the first learning process, the learning unit 153 replaces parameters of the student network with parameters of the teacher network learned using the second data.

[0065] The learning unit 153 executes a second learning process using the second data used in the first learning process. In the second learning process, the learning unit 153 updates the parameters of the student network using the parameters of the teacher network. In the second learning process, the learning unit 153 replaces the parameters of the student network with the average values ​​of the parameters of the teacher network and the parameters of the student network.

[0066] The transmitting unit 154 transmits various types of information. The transmitting unit 154 provides various types of information. The transmitting unit 154 provides various types of information to an external information processing device. The transmitting unit 154 transmits various types of information to an external information processing device. The transmitting unit 154 transmits information stored in the storage unit 14. The transmitting unit 154 transmits information stored in the first data storage unit 141. The transmitting unit 154 transmits information stored in the domain knowledge storage unit 142.

[0067] The transmitting unit 154 transmits the information generated by the generating unit 152. The transmitting unit 154 transmits the second model generated by the generating unit 152 to an external device. The transmitting unit 154 transmits the information learned by the learning unit 153. The transmitting unit 154 transmits the model learned by the learning unit 153 to an external device that performs inference processing using the model. The transmitting unit 154 transmits at least one of the student network and the teacher network learned by the learning unit 153 to the external device. With this configuration, the information processing device 100 can perform inference processing using a model that has been appropriately learned through offline reinforcement learning.

[0068] <1-4. Example Algorithm> Based on the above-mentioned premise, an algorithm for processing related to offline reinforcement learning executed by the information processing device 100 will now be described. The information processing device 100 executes processing related to offline reinforcement learning based on either the first algorithm or the second algorithm below. Note that explanations of points similar to those explained in FIG. 1 etc. will be omitted as appropriate.

[0069] <1-4-1. First Algorithm> First, a first algorithm, which is an example of an algorithm for processing related to offline reinforcement learning executed by the information processing device 100, will be described.

[0070] <1-4-1-1. Example of processing based on the first algorithm> The information processing procedure according to the first algorithm will be described below. The information processing procedure according to the first algorithm will be described using Fig. 3 and Fig. 4. Fig. 3 and Fig. 4 are flowcharts showing the information processing procedure according to the first algorithm. For example, Fig. 3 is a flowchart showing the procedure of a first learning process according to the first algorithm. Furthermore, Fig. 4 is a flowchart showing the procedure of a second learning process according to the first algorithm.

[0071] 3, the information processing device 100 executes a first learning process as shown in steps S101 to S105. For example, the information processing device 100 executes pre-learning of a teacher network as the first learning process.

[0072] First, the information processing device 100 generates an OOD (Out-of-Distribution) state (step S101). Here, OOD refers to, for example, information not included in the first data (behavioral history data). For example, the OOD state is an unknown state in the first data (behavioral history data). For example, the information processing device 100 randomly generates a state, and if the generated state is not included in the first data, the information processing device 100 also uses the generated state as the OOD state.

[0073] The information processing device 100 assigns an action to the state by using the domain knowledge (step S102). For example, the information processing device 100 selects information corresponding to the generated state from the domain knowledge, determines (generates) an action corresponding to the generated state by using the selected information (domain knowledge), and associates the determined action with the generated state.

[0074] The information processing device 100 accumulates the state-action pairs (step S103). For example, the information processing device 100 stores the state-action pairs, in which the generated states are associated with the determined actions, in the storage unit 14 as second data.

[0075] If the number of accumulated pairs is less than b (step S104: Yes), the information processing device 100 returns to step S101 and repeats the process. For example, if the number of accumulated state-action pairs, i.e., the number of second data, is less than b (for example, any value such as 10 or 100), the information processing device 100 repeats the process of generating second data.

[0076] If the number of accumulated pairs is not less than b (step S104: No), the information processing device 100 learns and stores a teacher network from the accumulated pairs (step S105) and ends the processing. For example, if the number of accumulated state-action pairs, i.e., the number of second data, is not less than b, i.e., is equal to or greater than b, the information processing device 100 learns a teacher network using the accumulated state-action pairs, i.e., the second data. Then, the information processing device 100 stores the learned teacher network in the storage unit 14 and ends the first learning process. Then, the information processing device 100 executes the second learning process.

[0077] 4, the information processing device 100 executes a second learning process as shown in steps S201 to S208. For example, after the first learning process, the information processing device 100 executes a learning process for learning a teacher network and a student network as the second learning process.

[0078] First, the information processing device 100 samples a certain number of (action, state, reward) tuples from existing data (step S201). For example, the information processing device 100 selects a certain number of samples from first data, which are combinations of action, state, and reward.

[0079] The information processing device 100 creates a buffer from which tuples to which the domain knowledge can be applied are extracted (step S202). For example, the information processing device 100 extracts data to which the domain knowledge can be applied from the first data selected as samples, and generates a first buffer containing the extracted data.

[0080] The information processing device 100 creates a buffer that extracts tuples for which the actions output by the teacher network and the student network do not match (step S203). For example, the information processing device 100 extracts data from the first buffer for which the actions output by the teacher network and the student network do not match, and creates a second buffer that includes the extracted data.

[0081] The information processing device 100 branches the processing for the buffer depending on whether or not the following conditions, including condition #1 and condition #2, are satisfied (step S204). For example, the information processing device 100 branches the processing for the buffer depending on whether or not both of the following conditions, condition #1 and condition #2, are satisfied.

[0082] Condition #1: The Q-value of the action selected by the student network is higher than the Q-value of the action selected by the teacher network. Condition #2: The Q-value of the student network is highly reliable.

[0083] If both Condition #1 and Condition #2 are satisfied (Step S204: Yes), the information processing device 100 updates the teacher network and sets the loss term by the teacher network to zero (Step S205). For example, if the Q value of the action selected by the student network is higher than the Q value of the action selected by the teacher network and the reliability of the Q value of the student network is high, the information processing device 100 updates the teacher network and sets the loss term by the teacher network to zero. Then, the information processing device 100 executes the process of Step S207.

[0084] If at least one of Conditions #1 and #2 is not satisfied (Step S204: No), the information processing device 100 sets a loss term by the teacher network (Step S206). For example, if at least one of the following conditions is not satisfied: the Q value of the action selected by the student network is higher than the Q value of the action selected by the teacher network, and the reliability of the Q value of the student network is high, the information processing device 100 sets a loss term by the teacher network. For example, if at least one of Conditions #1 and #2 is not satisfied, the information processing device 100 sets a loss term by the teacher network so that the action selected by the teacher network is more likely to be selected. For example, if at least one of Conditions #1 and #2 is not satisfied, the information processing device 100 sets a loss term by the teacher network so that the Q value of the action selected by the teacher network is increased and the Q values ​​of other actions are decreased. Then, the information processing device 100 executes the processing of Step S207.

[0085] The information processing device 100 updates the Q function (student network) using the (action, state, reward) tuple including the part to which domain knowledge cannot be applied, the loss term by the teacher network, and the loss term of Conservative Q (step S207). For example, the information processing device 100 executes a learning process using the first data including the data to which domain knowledge cannot be applied, and updates information related to the student network, such as the Q function for the student network.

[0086] If the number of updates of the Q function is less than N times (step S208: Yes), the information processing device 100 returns to step S201 and repeats the process. For example, if the number of updates of the Q function, i.e., the number of updates of the student network, is less than N times (for example, an arbitrary value such as 5, 50, etc.), the information processing device 100 repeats the process of updating the student network.

[0087] If the number of updates of the Q function is not less than N times (step S208: No), the information processing device 100 ends the process. For example, if the number of updates of the Q function, i.e., the number of updates of the student network, is not less than N times, i.e., is N times or more, the information processing device 100 ends the second learning process.

[0088] <1-4-1-2. Overview of the First Algorithm> Here, an overview of the first algorithm will be described using Fig. 5. Fig. 5 is a diagram showing an overview of the first algorithm. First algorithm processing FA in Fig. 5 shows an overview of the processing executed by the first algorithm.

[0089] For example, a first learning process PS11 in FIG. 5 shows an overview of the first learning process executed by the first algorithm. For example, the first learning process PS11 corresponds to the first learning process shown in the flowchart of FIG. 3. For example, the first learning process PS11 corresponds to learning of a teacher network TN (policy) using domain knowledge and second data. For example, as shown in FIG. 5, the information processing device 100 generates an action corresponding to a state from domain knowledge using the following formula (1), and uses second data including the generated action to learn the teacher network TN (π ω (corresponding to).

[0090]

[0091] For example, "a" on the left side of formula (1) corresponds to the action to be generated. Also, for example, "s" on the right side of formula (1) corresponds to the state used to generate the action, and "D" on the right side of formula (1) corresponds to a function that takes the state as input and outputs an action based on domain knowledge.

[0092] For example, the second learning process PS12 in Fig. 5 shows an overview of the second learning process executed by the first algorithm. For example, the second learning process PS12 corresponds to the second learning process shown in the flowchart of Fig. 4. For example, the second learning process PS12 corresponds to updating the student network SN (policy) through offline reinforcement learning by the teacher network TN.

[0093] 5, the information processing device 100 learns a model based on the first algorithm. In the first algorithm process FA, the information processing device 100 learns a student network and a teacher network based on the first algorithm.

[0094] In the first algorithm process FA, the information processing device 100 executes a first learning process to learn a teacher network using second data, and learns a student network using the teacher network learned by the first learning process and the first data by a second learning process.

[0095] In the first algorithm process FA, the information processing device 100 performs an update process on the teacher network if the action selected by the teacher network differs from the action selected by the student network in the second learning process. In the first algorithm process FA, the information processing device 100 updates the teacher network if the index value of the action selected by the student network is higher than the index value of the action selected by the teacher network. For example, the information processing device 100 updates the teacher network if the Q value of the action selected by the student network (Q function) is higher than the Q value of the action selected by the teacher network (Q function). For example, as shown in FIG. 5, the information processing device 100 determines the action to be selected by the teacher network TN using the following equation (2):

[0096]

[0097] For example, "a" on the left side of formula (2) t " corresponds to the action selected by the teacher network TN. Also, for example, "s" on the right side of the formula (2) corresponds to the state, and "π ω()" corresponds to (a function of) the teacher network TN. Furthermore, for example, the information processing device 100 determines the action to be selected by the student network SN using the following equation (3), as shown in FIG. 5.

[0098]

[0099] For example, "a" on the left side of formula (3) s " corresponds to the action selected by the student network SN. Also, for example, "Q θ " corresponds to the Q value, and "argmax" on the right side of equation (3) corresponds to a function that returns the action corresponding to the maximum Q value.

[0100] 5, in the first algorithm process FA, for example, knowledge is given as an if then rule from a state to an action for a subset / space of states. As described above, in the first algorithm process FA, in pre-learning of the teacher network, the information processing device 100 generates action-state pairs using action->state knowledge, and learns the teacher network using a technique such as Behavior Cloning.

[0101] For example, in the first algorithm process FA, the information processing device 100 executes a first process of adding a loss (loss term) that brings the behavior of the student network closer to that of the teacher network. Note that the information processing device 100 does not add a loss term if the prediction of the student network is valid and has a high degree of certainty when the domain knowledge is valid.

[0102] In the first algorithm process FA, the information processing device 100 executes a second process of learning from the student network when the prediction of the student network is valid and has a high degree of confidence in a state in which domain knowledge is valid. In the first algorithm process FA, the information processing device 100 executes normal offline reinforcement learning as a third process. For example, in the first algorithm process FA, the information processing device 100 executes mini-batch learning, which repeats the first process, the second process, and the third process.

[0103] By processing based on the first algorithm processing FA described above, information from the teacher network is used to learn the Q function (value function) of the student network in the intersection of the area covered by domain knowledge and the area covered by data, making it possible to select appropriate actions even in an OOD state.

[0104] <1-4-2. Second Algorithm> Next, a second algorithm will be described, which is an example of an algorithm for processing related to offline reinforcement learning executed by the information processing device 100. Note that the algorithm for processing related to offline reinforcement learning executed by the information processing device 100 is not limited to the first algorithm and the second algorithm, and may be any algorithm. For example, the processing related to offline reinforcement learning executed by the information processing device 100 may be processing based on a third algorithm other than the first algorithm and the second algorithm.

[0105] <1-4-2. Second Algorithm> Next, a description will be given of a second algorithm, which is an example of an algorithm for processing related to offline reinforcement learning executed by the information processing device 100. Note that descriptions of similar points to the above-described first algorithm and the like will be omitted as appropriate.

[0106] <1-4-2-1. Example of Processing Based on the Second Algorithm> The information processing procedure according to the second algorithm will be described below. The information processing procedure according to the second algorithm will be described using Fig. 6 and Fig. 7. Fig. 6 and Fig. 7 are flowcharts showing the information processing procedure according to the second algorithm. For example, Fig. 6 is a flowchart showing the procedure of a first learning process according to the second algorithm. Furthermore, Fig. 7 is a flowchart showing the procedure of a second learning process according to the second algorithm.

[0107] 6, the information processing device 100 executes a first learning process as shown in steps S301 to S307. For example, the information processing device 100 executes pre-learning of a teacher network as the first learning process.

[0108] First, the information processing device 100 generates an OOD (Out-of-Distribution) state (step S301). Note that steps S301 to S304 in Fig. 6 are the same as steps S101 to S104 in Fig. 3, and therefore detailed description thereof will be omitted.

[0109] The information processing device 100 assigns an action to a state using domain knowledge (step S302). The information processing device 100 accumulates state-action pairs (step S303). If the number of accumulated pairs is less than b (step S304: Yes), the information processing device 100 returns to step S301 and repeats the process.

[0110] If the number of accumulated pairs is not less than b (step S304: No), the information processing device 100 trains the teacher network from the accumulated pairs for one epoch (step S305). For example, if the number of accumulated state-action pairs, i.e., the number of second data, is not less than b, i.e., is b or more, the information processing device 100 trains the teacher network using each of the accumulated state-action pairs, i.e., the second data, once.

[0111] Then, the information processing device 100 copies the parameters of the teacher network to the parameters of the student network (step S306). For example, the information processing device 100 copies the parameters of the teacher network updated by the learning process in step S305 to the parameters of the student network. The student network and the teacher network have the same network structure, and the information processing device 100 replaces the values ​​of the parameters of the student network with the values ​​of the parameters of the teacher network, which have a common network structure.

[0112] Then, the information processing device 100 stores the data of the state-action pair consisting of b pieces of data (step S307), and ends the processing. For example, the information processing device 100 stores the data of the state-action pair consisting of b pieces of data as data D d The information processing device 100 then stores the value of the first learning process as "1" and ends the first learning process. Then, the information processing device 100 executes the second learning process.

[0113] 7, the information processing device 100 executes a second learning process as shown in steps S401 to S408. For example, after the first learning process, the information processing device 100 executes a learning process for learning a teacher network and a student network as the second learning process.

[0114] First, the information processing device 100 receives the data D d For example, the information processing apparatus 100 samples a batch (a subset of state-action pairs) from the second data D d Select a subset of state-action pairs from as samples.

[0115] The information processing device 100 updates the parameters of the teacher network using the sampled batch (step S402). For example, the information processing device 100 executes a learning process using the sampled batch to update the parameters of the teacher network.

[0116] The information processing device 100 calculates the average value of the parameters of the teacher network and the parameters of the student network, and replaces the average value as the parameters of the student network (step S403). For example, the information processing device 100 calculates the average value of the parameters of the teacher network and the parameters of the student network updated in step S402, and replaces the values ​​of the parameters of the student network with the calculated average value of the parameters.

[0117] The information processing device 100 samples a batch from the original learning data D (step S404). For example, the information processing device 100 selects data as a sample from the original learning data D, which is first data (behavior history data).

[0118] The information processing device 100 outputs actions from each of the teacher network and the student network and calculates a discrepancy measure (step S405). For example, the information processing device 100 outputs actions from each of the teacher network and the student network for each data in the original training data D and calculates a discrepancy measure based on the difference between the actions output by the teacher network and the student network. For example, the discrepancy measure is a value based on the cosine similarity between the actions of the teacher network and the actions of the student network. For example, the discrepancy measure is a value obtained by adding the cosine similarity between the actions of the teacher network and the actions of the student network to a predetermined constant (e.g., 1).

[0119] The information processing device 100 uses the discrepancy measure to learn a value function (Q function) and a policy (student network) for the student network (step S406). For example, the information processing device 100 uses the difference (discrepancy measure) between the actions of the teacher network and the student network calculated in step S405 to learn a value function (Q function) and a policy (student network). For example, the information processing device 100 uses the discrepancy measure calculated in step S405 to assign a vector weight to the value function (loss term). For example, the information processing device 100 assigns a vector weight based on the discrepancy measure calculated in step S405 to the value function, and learns a value function (Q function) and a policy (student network) for the student network.

[0120] If the teacher network and the student network have not converged (step S407: No), the information processing device 100 returns to step S401 and repeats the process. For example, if the degree to which the value of at least one of the parameters of the teacher network and the parameters of the student network changes due to learning is equal to or greater than a predetermined threshold, the information processing device 100 determines that convergence has not occurred and returns to step S401 and repeats the process.

[0121] If the teacher network and the student network have converged (step S407: Yes), the information processing device 100 returns the student network (step S408) and ends the processing. For example, if the degree of change in the value of at least one of the parameters of the teacher network and the parameters of the student network due to learning is less than a predetermined threshold, the information processing device 100 determines that convergence has occurred, stores the student network in the storage unit 14, and ends the second learning process. Note that if the number of iterations of the second learning process reaches a predetermined number, the information processing device 100 may store the student network in the storage unit 14 and end the second learning process.

[0122] <1-4-2-2. Overview of the Second Algorithm> Here, the second process will be described with reference to Fig. 8. Fig. 8 is a diagram showing an overview of the second algorithm. Second algorithm process SA in Fig. 8 shows an overview of the process executed by the second algorithm.

[0123] For example, a first learning process PS21 in FIG. 8 shows an overview of the first learning process executed by the second algorithm. For example, the first learning process PS21 corresponds to the first learning process shown in the flowchart of FIG. 6. For example, the first learning process PS21 corresponds to learning of a teacher network TN (policy). For example, as shown in FIG. 8, the information processing device 100 uses the following equation (4) to learn the teacher network TN (π δ (corresponding to).

[0124]

[0125] Equation (4) corresponds to the loss at the learning step t during the learning of the teacher network TN.t d " corresponds to the sampling size at the learning step t. The first term (term to be subtracted) of the subtraction formula on the right side of equation (4) is the data set D used in the first learning process PS21. d The second term (the term to be subtracted) of the subtraction formula on the right side of equation (4) corresponds to the action determined based on the state of the data (sample) of the i-th state-action pair and the teacher network TN. d corresponds to the action of the data (sample) of the i-th state-action pair.

[0126] For example, the second learning process PS22 in Fig. 8 shows an overview of the second learning process executed by the second algorithm. For example, the second learning process PS22 corresponds to the second learning process shown in the flowchart of Fig. 7. For example, the second learning process PS22 corresponds to updating the student network SN (policy) through offline reinforcement learning by the teacher network TN. For example, as shown in Fig. 8, the information processing device 100 determines the action to be selected by the student network SN using a value function (Q function) such as the following equation (5).

[0127]

[0128] For example, equation (5) is a state s sampled (selected) from the data group D used in the second learning process PS22 in the learning step t. t In addition, for example, the value function (Q function) when "π φ ()) corresponds to (a function of) the student network SN. φ (s t ) is the state s t This corresponds to the action selected by the student network SN in

[0129] For example, the information processing device 100 determines the action to be selected by the student network SN using the following formula (5) as shown in Fig. 8. Also, for example, the information processing device 100 executes the second learning process PS22 using a discrepancy measure such as the following formula (6) as shown in Fig. 8.

[0130]

[0131] The "·" (dot) in equation (6) represents a dot product, and "||·||" in equation (6) represents the Euclidean norm.

[0132] For example, equation (6) is a state s sampled (selected) from the data group D used in the second learning process PS22 in the learning step t. t This corresponds to a formula for calculating the discrepancy measure based on the cosine similarity between two actions when t " is the state s t In addition, for example, "a" in Equation (6) corresponds to a pair of actions. t The elements with a "^" (hat) above them are in state s t For example, the left side of equation (6) corresponds to the discrepancy measure in step S405 of FIG.

[0133] 8, the information processing device 100 learns a model based on the second algorithm. In the second algorithm processing SA, the information processing device 100 learns a student network and a teacher network based on the second algorithm.

[0134] In the second algorithm processing SA, the information processing device 100 executes a first learning process to learn a teacher network and a student network using the second data, and learns a student network using the teacher network and student network learned by the first learning process and the first data by the second learning process.

[0135] In the second algorithm process SA, the information processing device 100 replaces the parameters of the student network with the parameters of the teacher network learned using the second data in the first learning process. In the second algorithm process SA, the information processing device 100 executes the second learning process using the second data used in the first learning process.

[0136] In the second algorithm process SA, the information processing device 100 updates the parameters of the student network using the parameters of the teacher network in the second learning process. In the second algorithm process SA, the information processing device 100 replaces the parameters of the student network with the average value of the parameters of the teacher network and the parameters of the student network in the second learning process.

[0137] The algorithm of the offline reinforcement learning process executed by the information processing device 100 is not limited to the first and second algorithms described above, and may be any algorithm. For example, the offline reinforcement learning process executed by the information processing device 100 may be a process based on a third algorithm other than the first and second algorithms.

[0138] As described above, when using knowledge of desirable actions in a state, the information processing device 100 identifies a state (OOD State) that is not covered by existing behavioral history data (offline dataset). The information processing device 100 then generates actions for the OOD state using at least one of domain knowledge and LLM, and constructs a teacher network (policy) using the generated state-action pairs. The information processing device 100 then learns a student network (policy) using offline reinforcement learning techniques while using the output and parameters of the teacher network (policy). The teacher network may be a reinforcement learning policy network. The student network (policy) may include a Q function. The teacher network (policy) may also include a Q function.

[0139] As another method, for example, the information processing device 100 may add domain knowledge or loss terms (e.g., loss terms) based on LLM to the objective function when learning the student network (policy). The information processing device 100 may also use knowledge about states, actions, and rewards (e.g., domain knowledge, LLM) to create behavioral history data (states, actions, rewards, previous states) of the OOD and use it for offline reinforcement learning. Furthermore, if the knowledge is about state transitions (which state will the OOD transition to next when a certain action is taken in a certain state), then in model-based reinforcement learning, state transition data in the OOD may be created using knowledge and used for learning. Furthermore, if the knowledge is about the merits or demerits of only states or actions, then it may be expressed as loss terms in a loss function and used for learning state transitions or policies, respectively.

[0140] <1-5. Application Example (Web Advertising Operation)> Hereinafter, a specific application example (embodiment) of the process executed by the above-described information processing device 100 will be described, taking Web advertising operation as an example. Note that the process executed by the information processing device 100 is not limited to Web advertising operation and may be applied to any destination, but this point will be described later.

[0141] <1-5-1. Example of Advertising Management by DSP> First, with reference to FIG. 9 , advertising management by a DSP (Demand-Side Platform) will be described as an example of web advertising management to which processing executed by the information processing device 100 is applied. FIG. 9 is a diagram showing an overview of web advertising management. The DSP referred to here is an existing technology in advertising management, and detailed description will be omitted as appropriate. However, it is a web advertising distribution system (platform) in which an advertiser manages bidding, targeting, advertising space, etc. for multiple advertisements, aiming to optimize the effectiveness of the advertiser's advertisements. For example, a DSP is used to maximize conversions within a budget. For example, an advertiser (such as a person in charge of operations) uses a DSP to set a budget for a campaign, etc., advertising settings, etc., and manages web advertising.

[0142] When using DSP to run web ads, you can change the following settings, for example: ・Budget for each campaign ・Delivery type (remarketing, similar, etc.), its settings (similarity, etc.), bidding strategy, etc. ・Target (list of domains to be delivered, time period, etc.) ・Creative selection (advertising image, etc.)

[0143] Some of the difficulties in using DSPs for web advertising include: - Making decisions based on highly volatile delivery performance data and A / B test results - Responding appropriately to changing circumstances (creative staleness, other campaigns, competitive situations, etc.) - Adjusting settings based on the quirks of the DSP

[0144] In the case of web advertising operations by DSP, the information observed by the advertiser (such as the person in charge of operations) includes, for example, the following information: Amount spent on ad delivery (amount spent) Number of impressions Number of clicks Number of conversions Current settings

[0145] For example, advertising management by a DSP can be automated by utilizing offline reinforcement learning through processing executed by the information processing device 100. For example, advertising management by a DSP can achieve labor-saving and automation of operations, improve efficiency through advice, and support for new employees through offline reinforcement learning through processing executed by the information processing device 100.

[0146] Here, an example of the processing flow for advertising management using a DSP will be described using FIG. 10 . FIG. 10 is a diagram showing the flow of advertising management using a DSP. In FIG. 10 , the direction from left (Start on the left end) to right (End on the right end) corresponds to the flow of time. For example, the user operation manager in FIG. 10 corresponds to the advertiser (e.g., a person in charge of operations), and the DSP corresponds to the DSP used by the advertiser for advertising management. Also, for example, the system in FIG. 10 is a system that manages advertiser-side settings and the like in the advertising management of the advertiser using a DSP. For example, the system in FIG. 10 performs inference processing using a model (e.g., the policy model in FIG. 10 ) learned by the information processing device 100, and executes processing to automatically change settings and the like in the advertising management using the DSP based on the results of the inference processing.

[0147] For example, the process shown in FIG. 10 is executed at a predetermined timing, such as every other day. For example, the system in FIG. 10 requests the execution of data export. The DSP in FIG. 10 exports the data to the system. The system then formats the data and stores it in a DB (database). The system then retrieves the data from the DB and executes processing such as data merging. The system executes the policy based on the policy model. The system visualizes the policy content and displays the approval screen for the user operations manager in FIG. 10. Note that if the process is fully automated without the involvement of the user operations manager, the approval process for the user operations manager does not need to be performed.

[0148] If the user operations manager approves, the system requests the DSP to execute the setting change. The DSP changes the setting in response to the request from the system. The system then saves the execution results and displays the results. If the user operations manager does not approve, the system does not execute the setting change, and the advertising operation continues without changing the settings, and displays the results. The user operations manager then checks the results.

[0149] <1-5-2. Example of Monthly Budget> Here, an example of advertising management on a monthly budget will be described with reference to FIG. 11. FIG. 11 is a diagram showing an example of data when advertising is managed on a monthly budget. FIG. 11 shows an example of distribution data when advertising is managed using a DSP on a monthly budget. For example, the distribution data is aggregated daily and has the following format:

[0150] 11 includes date, campaign ID, monthly budget (yen), frequency, target conversion cost (yen), campaign feature amount, creative feature amount, similarity, consumption amount (yen), impression, click, conversion, CPA, etc. In this way, the distribution data includes information related to distribution settings, information related to feature amount, information related to distribution results, etc.

[0151] For example, among the distribution data, campaign ID, monthly budget (yen), frequency, and target conversion cost (yen) correspond to information about distribution settings. Also, among the distribution data, campaign features, creative features, similarity, and consumption amount (yen) correspond to information about features. Also, among the distribution data, impressions, clicks, conversions, and CPA correspond to information about distribution results.

[0152] For example, both the campaign feature and the creative feature are continuous vectors of a fixed dimension, and are generated from information about the campaign and information about the creative, respectively.

[0153] In advertising operations with a monthly budget as described above, the reinforcement learning problem setting might be, for example, as follows:

[0154] For example, the time unit is one day, and t = 1 to 30. A series of deliveries carried out over 30 days is called a trajectory. The state is information about the feature amount, information about the delivery settings, and information about the delivery results in FIG. 11. Although not shown in the example of FIG. 11, delivery results that have been aggregated from past delivery results, such as CPA calculated from the past five days, may also be used.

[0155] The reward is the number of conversions. In this case, the policy (model) is trained so as to maximize the number of conversions in 30 days. The actions are five-dimensional discrete actions as shown in FIG. 12. FIG. 12 is a diagram showing an example of actions in the case of advertising management based on a monthly budget.

[0156] As shown in Figure 12, the targets of the five-dimensional discrete actions include the monthly budget, the frequency of campaign A, the frequency of campaign B, the cost per conversion of campaign A, and the cost per conversion of campaign B. Furthermore, three types of actions, increase, decrease, and no change, are associated with each of the targets of the five-dimensional discrete actions as possible actions. In this case, one of the three types of actions can be taken for each of the five-dimensional discrete actions.

[0157] Furthermore, the constraints for each distribution setting when executing an action may include, for example, the following constraints. For example, a setting change is implemented to realize the content of the action so that the total budget amount remains the same and a constraint (condition) such as a frequency between 1 and 50 is observed. In this case, for example, if the monthly budget for campaign A is 200,000 yen and the monthly budget for campaign B is 100,000 yen, executing an action to increase the monthly budget for campaign B will change the monthly budget for campaign B to 120,000 yen and the monthly budget for campaign A to 180,000 yen, and although the total budget amount will not change, a setting change is implemented to realize the action.

[0158] <1-5-3. Example in the Case of Daily Budget> Next, an example of advertising management based on a daily budget will be described with reference to FIG. 13. FIG. 13 is a diagram showing an example of data in the case of advertising management based on a daily budget. FIG. 13 shows an example of distribution data in the case of advertising management using a DSP based on a daily budget. For example, the distribution data is aggregated daily and has the following format. Note that explanations of the same points in advertising management based on a daily budget as in advertising management based on a monthly budget will be omitted as appropriate.

[0159] The distribution data shown in Figure 13 includes date, campaign ID, daily budget (yen), frequency, target conversion cost (yen), campaign feature amount, creative feature amount, similarity, consumption amount (yen), impression, click, conversion, CPA, etc. As such, the distribution data includes information related to distribution settings, information related to feature amount, information related to distribution results, etc. The distribution data shown in Figure 13 is almost the same as the distribution data shown in Figure 11 except that the monthly budget (yen) has been changed to a daily budget (yen), so a detailed description will be omitted.

[0160] In advertising operations with a daily budget as described above, the reinforcement learning problem setting is, for example, as follows:

[0161] For example, the time unit is one day, and t = 1 to 30. A series of deliveries carried out over 30 days is called a trajectory. The state is information about the feature amount, information about the delivery settings, and information about the delivery results in FIG. 13. Although not shown in the example of FIG. 13, delivery results that have been aggregated from past delivery results, such as CPA calculated from the past five days, may also be used.

[0162] The reward is the number of conversions. In this case, the policy (model) is trained so as to maximize the number of conversions over 30 days. The actions are five-dimensional discrete actions as shown in FIG. 14. FIG. 14 is a diagram showing an example of actions in the case of advertising management based on a daily budget.

[0163] As shown in Figure 14, the targets of the five-dimensional discrete actions include the daily budget, the frequency of campaign A, the frequency of campaign B, the cost per conversion of campaign A, and the cost per conversion of campaign B. Furthermore, each of the targets of the five-dimensional discrete actions is associated with three types of possible actions: increase, decrease, and no change. In this case, one of the three types of actions can be taken for each of the five-dimensional discrete actions.

[0164] Furthermore, restrictions for each distribution setting when executing an action may include, for example, the following restrictions. For example, a setting change is implemented to realize the content of the action so that the total budget remains the same and a restriction (condition) is met, such as a frequency between 1 and 50. In this case, for example, if the total monthly budget is 450,000 yen, the daily budget would be 15,000 yen if evenly distributed. However, if an action is executed to increase the daily budget for campaign B, a setting change is implemented to realize the action within the total daily budget of 15,000 yen. For example, a setting change is implemented to change (daily budget for campaign A, daily budget for campaign B) = (10,000 yen, 5,000 yen) to (8,000 yen, 7,000 yen).

[0165] <1-5-4. Processing Example> Here, a processing example relating to offline reinforcement learning executed by the information processing device 100 will be described, taking the case of advertising management with the monthly budget described above as an example.

[0166] In the following example, assume there are 100 past delivery results, i.e., 100 trajectories. For example, assume there are multiple services or products targeted for advertising, with various budget sizes. For example, advertisements for new targets or delivery settings with budget sizes that have never been set before will be in an unknown state, and will therefore be in an OOD state.

[0167] In addition, the domain knowledge is assumed to be as follows: ・If there is a difference in CPA between two campaigns, reduce the monthly budget for the one with the higher CPA and increase the monthly budget for the one with the lower CPA ・Set the frequency low at first, then increase it ・If the target cost per conversion is high compared to the actual CPA, reduce the target cost per conversion

[0168] For example, the domain knowledge described above may be converted into a function that inputs a state and outputs an action, as will be described later. Note that, as shown in Fig. 12 etc., in the case of a discrete action, an algorithm for discrete value actions can be applied.

[0169] For example, the information processing device 100 generates second data (OOD data) as follows. It is assumed that the budget size and feature amount of the distribution target are determined in advance. First, the information processing device 100 performs the following process to generate an OOD state. For example, the information processing device 100 performs the following process to generate distribution settings and distribution results for one day.

[0170] For example, the information processing device 100 randomly generates monthly budgets for campaign A and campaign B. The information processing device 100 also randomly generates frequencies for campaign A and campaign B.

[0171] For example, the information processing device 100 divides the budget by 30 days and adds a random number to calculate the amount spent. Furthermore, the information processing device 100 randomly generates, from a predetermined range of values, a CPM (cost per mile), which is the cost per 1,000 impressions, a CTR (click-through rate), which is calculated by the number of clicks / the number of impressions, and a CVR (conversion rate), which is calculated by the number of conversions / the number of clicks, for the amount spent, and calculates the distribution results for each. Furthermore, the information processing device 100 randomly generates a target conversion cost based on the distribution results.

[0172] For the above information, the information processing device 100 randomly selects one piece of domain knowledge and generates an action using the domain knowledge, thereby enabling the information processing device 100 to generate a pair of an OOD state and an action as second data.

[0173] Note that the above is merely an example, and the information processing device 100 may generate the second data using various information. For example, the information processing device 100 may learn a state transition, a reward function, and a policy using existing data, execute the policy, and generate the second data by adding random numbers to the state and action for a certain step. This allows the information processing device 100 to appropriately learn a good policy (model) that will earn high rewards in the future for campaigns with budgets and distribution records that are not included in the distribution record data.

[0174] <1-5-5. Example of Prompt> Furthermore, the information processing device 100 may generate the second data using the LLM as described above. An example of this point will be described with reference to FIG. 15. FIG. 15 is a diagram showing an example of a prompt used to generate the second data. For example, FIG. 15 shows an example of a prompt when requesting the generation of an action corresponding to a state specified in the prompt.

[0175] Prompt PT1 in Fig. 15 includes information on the status of two campaigns A and B in display advertising distribution, such as distribution settings and distribution results, and patterns of possible actions, and includes content requesting the generation of information indicating actions to be taken in response to those statuses. In the example of Fig. 15, the campaign features and creative features are each keywords, and the LLM generates and outputs a response taking into account the meaning of the keywords.

[0176] The information processing device 100 generates state-action pairs as second data using the output of the LLM. For example, since the LLM answers are generated in text, the information processing device 100 performs a process of converting the LLM answers into actions and generates actions from states. The information processing device 100 then generates state-action pairs as second data by associating the specified states with the generated actions.

[0177] <1-5-6. Example of Use of Domain Knowledge> Note that the above-described process is merely an example, and domain knowledge may be used in various ways. For example, the information processing device 100 may efficiently use domain knowledge by the following process.

[0178] For example, the information processing device 100 may generate a function that generates an action from a state. In this case, for example, the information processing device 100 selects one or more elements of a distribution setting, a feature, and a distribution result. Then, the information processing device 100 sets conditions for them.

[0179] Then, the information processing device 100 selects an action when the condition is satisfied or a condition that the action satisfies. For example, the information processing device 100 may evaluate the domain knowledge by the following process. For example, the information processing device 100 calculates and displays the coverage rate of the domain knowledge. In this case, the information processing device 100 displays content CT1 as shown in FIG. 16. FIG. 16 is a diagram showing an example of a display related to domain knowledge.

[0180] 16 shows an example of a user interface (UI) for domain knowledge, coverage rate, accuracy, etc. The content CT1 may be displayed on a terminal device used by a user such as an advertiser. In this case, the information processing device 100 may transmit the content CT1 to the terminal device used by the user, and the terminal device that receives the content CT1 may display the content CT1.

[0181] 16 includes various information related to domain knowledge. For example, the content CT1 includes information such as a state, a state condition, and an action condition based on the domain knowledge. The content CT1 also includes information such as a coverage rate and accuracy based on the domain knowledge.

[0182] For example, content CT1 includes information indicating the percentage to which domain knowledge can be applied to all states of existing data. Content CT1 also includes information indicating the percentage to which domain knowledge can be applied to states to which existing domain knowledge included in existing data cannot be applied. Content CT1 also includes information indicating the percentage to which domain knowledge can be applied when the OOD state is sampled using the above-described method. Content CT1 also includes information indicating the percentage to which domain knowledge can be applied to states to which existing domain knowledge cannot be applied when the OOD state is sampled.

[0183] For example, the content CT1 includes information indicating the accuracy of the domain knowledge. For example, when there is a trajectory (expert data) generated by an expert among the existing data, the accuracy of the domain knowledge may be the percentage of matches between the action conditions of the domain knowledge and the actions of the data when the domain knowledge is applied to a state included in the trajectory.

[0184] The information processing device 100 generates content CT1 as shown in Fig. 16. In this case, the information processing device 100 may calculate various information included in the content CT1. The information processing device 100 calculates the coverage rate of domain knowledge. In Fig. 16, the information processing device 100 calculates the coverage rate of existing data to be 15%, the coverage rate of existing data (without knowledge) to be 8%, the coverage rate of OOD to be 19%, and the coverage rate of OOD (without knowledge) to be 11%.

[0185] The information processing device 100 also calculates the accuracy of the domain knowledge. In Fig. 16, the information processing device 100 calculates the accuracy to be 68% based on the match rate with the expert data. Note that the calculation of various pieces of information may be performed by a device (computing device) other than the information processing device 100, and the information processing device 100 may receive information calculated by the computing device from the computing device.

[0186] The information processing device 100 may also request domain knowledge that covers the uncovered OOD. In this case, the information processing device 100 may sample multiple OOD states using the method described above, display states that are not covered by existing domain knowledge, and prompt the user to input domain knowledge. The information processing device 100 may also display states in which the confidence level of the learned policy or Q function for the OOD is low.

[0187] <1-5-7. Other Application Examples> In the application examples described above, distribution setting operation in advertising operation has been described as an example, but the above-described processing can be applied to various processes, not limited to advertising operation. For example, various processes such as the learning process executed by the information processing device 100 can also be used in the following areas. For example, a model learned by the learning process of the information processing device 100 may be used for inference processing in the following areas.

[0188] For example, the process executed by the information processing device 100 may be applied to resource management. In this case, the action may be, for example, shifts of call center operators. For example, the process executed by the information processing device 100 may be applied to preventing churn (cancellation of service). In this case, the action may be, for example, sending coupons.

[0189] For example, the processing executed by the information processing device 100 may be applied to automatic pricing. In this case, the action may be price setting, etc. For example, the processing executed by the information processing device 100 may be applied to portfolio management of asset management. In this case, the action may be buying and selling of assets, etc.

[0190] For example, the process executed by the information processing device 100 may be applied to recommendation. In this case, the action may be content presentation, etc. For example, the process executed by the information processing device 100 may be applied to air conditioning setting control. In this case, the action may be setting change, etc.

[0191] For example, the processing executed by the information processing device 100 may be applied to inventory management. In this case, the action may be determining the amount of stock to be purchased, etc. For example, the processing executed by the information processing device 100 may be applied to preventing equipment failures and improving yields, etc. In this case, the action may be setting parameters for the equipment, etc.

[0192] For example, the processing performed by the information processing device 100 may be applied to medical care and healthcare. In this case, the behavior may be a health improvement measure such as exercise. For example, the processing performed by the information processing device 100 may be applied to preventing employees from quitting their jobs. In this case, the behavior may be encouraging employees.

[0193] 2. Other Configuration Examples The processes according to the above-described embodiments etc. may be implemented in various different forms (modifications) other than the above-described embodiments etc. For example, the system configuration is not limited to the above-described examples and may be in various forms.

[0194] For example, the processes according to the above-described embodiments may be implemented in various different forms (modified examples) other than the above-described embodiments. For example, an information processing system may have a learning device (e.g., information processing device 100) that executes a learning process, and a device (inference device) that executes an inference process using a model learned by the learning device. In this case, the information processing system may include the information processing device 100 and the inference device. Note that the above is just one example, and the information processing system may be realized in various configurations. For example, the information processing device 100 may be a device that executes the learning process and the inference process. In other words, in the information processing system, the learning device and the inference device may be integrated.

[0195] <3. Others> Furthermore, among the processes described in each of the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0196] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0197] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0198] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0199] 5. Hardware Configuration An information device such as the information processing device 100 according to each of the above-described embodiments is realized by a computer 1000 having a configuration such as that shown in FIG. 17 . FIG. 17 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of an information processing device such as the information processing device 100. The following description will be given using the information processing device 100 according to the embodiment as an example. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0200] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0201] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0202] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an information processing program according to the present disclosure, which is an example of program data 1450.

[0203] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0204] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0205] For example, when the computer 1000 functions as the information processing device 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an information processing program loaded onto the RAM 1200, thereby realizing the functions of the control unit 15 and the like. The information processing program according to the present disclosure and data in the storage unit 14 are stored in the HDD 1400. The CPU 1100 reads and executes program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0206] The present technology may also be configured as follows. (1) An information processing device including: an acquisition unit that acquires first data, which is behavior history data used as learning data for offline reinforcement learning; a generation unit that generates second data, which is data other than the first data and can be used as learning data for the offline reinforcement learning; and a learning unit that learns a model by the offline reinforcement learning using the first data and the second data. (2) The information processing device described in (1), in which the acquisition unit acquires domain knowledge, which is knowledge related to a field in which the model is used, and the generation unit generates the second data using the domain knowledge. (3) The information processing device described in (1) or (2), in which the acquisition unit acquires large language models (LLMs) used to generate the second data, and the generation unit generates the second data using the LLMs. (4) The information processing device described in any one of (1) to (3), in which the generation unit generates the second data including at least states and actions corresponding to the states. (5) The information processing device according to (4), wherein the generation unit generates the second data by generating the state and generating an action corresponding to the generated state. (6) The information processing device according to (5), wherein the generation unit randomly generates the state and generates an action corresponding to the generated state. (7) The information processing device according to (5) or (6), wherein the generation unit generates an action corresponding to the state based on the state and a domain in which the model is used. (8) The information processing device according to any one of (5) to (7), wherein the generation unit generates an action corresponding to the state based on the state and a generative model that generates an action corresponding to the state. (9) The information processing device according to any one of (1) to (8), wherein the learning unit learns a student network and a teacher network, which are the models used in the inference process.(10) The information processing device according to (9), wherein the learning unit performs a first learning process to learn the teacher network using the second data, and learns the student network through a second learning process using the teacher network learned through the first learning process and the first data. (11) The information processing device according to (10), wherein the learning unit performs an update process for the teacher network if an action selected by the teacher network differs from an action selected by the student network in the second learning process. (12) The information processing device according to (11), wherein the learning unit updates the teacher network if an index value of an action selected by the student network is higher than an index value of the action selected by the teacher network. (13) The information processing device according to (9), wherein the learning unit performs a first learning process to learn the teacher network and the student network using the second data, and learns the student network through a second learning process using the teacher network learned through the first learning process, the student network, and the first data. (14) The information processing device according to (13), wherein the learning unit replaces parameters of the student network with parameters of the teacher network learned using the second data in the first learning process. (15) The information processing device according to (13) or (14), wherein the learning unit executes the second learning process using the second data used in the first learning process. (16) The information processing device according to any one of (13) to (15), wherein the learning unit updates parameters of the student network using parameters of the teacher network in the second learning process. (17) The information processing device according to (16), wherein the learning unit replaces parameters of the student network with average values ​​of parameters of the teacher network and parameters of the student network in the second learning process. (18) The information processing device according to any one of (1) to (17), further comprising a transmission unit that transmits the model learned by the learning unit to an external device that performs inference processing using the model.(19) An information processing method including: acquiring first data, which is behavior history data used as learning data for offline reinforcement learning; generating second data, which is data other than the first data and usable as learning data for the offline reinforcement learning; and learning a model by the offline reinforcement learning using the first data and the second data. (20) An information processing program that causes a computer to execute: acquiring first data, which is behavior history data used as learning data for offline reinforcement learning; generating second data, which is data other than the first data and usable as learning data for the offline reinforcement learning; and learning a model by the offline reinforcement learning using the first data and the second data.

[0207] REFERENCE SIGNS LIST 100 Information processing device 11 Communication unit 12 Input unit 13 Output unit 14 Storage unit 141 First data storage unit 142 Domain knowledge storage unit 15 Control unit 151 Acquisition unit 152 Generation unit 153 Learning unit 154 Transmission unit

Claims

1. An information processing apparatus comprising: an acquisition unit that acquires first data which is action history data used as learning data for offline reinforcement learning; a generation unit that generates second data which is data other than the first data and is data that can be used as learning data for the offline reinforcement learning; and a learning unit that learns a model by the offline reinforcement learning using the first data and the second data.

2. The information processing apparatus according to claim 1, wherein the acquisition unit acquires domain knowledge which is knowledge related to the field in which the model is used, and the generation unit generates the second data using the domain knowledge.

3. The information processing apparatus according to claim 1, wherein the acquisition unit acquires an LLM (Large Language Models) used for generating the second data, and the generation unit generates the second data using the LLM.

4. The information processing apparatus according to claim 1, wherein the generation unit generates the second data including at least a state and an action corresponding to the state.

5. The information processing apparatus according to claim 4, wherein the generation unit generates the second data by generating the state and generating an action corresponding to the generated state.

6. The information processing apparatus according to claim 5, wherein the generation unit randomly generates the state and generates an action corresponding to the generated state.

7. The information processing apparatus according to claim 5, wherein the generation unit generates an action corresponding to the state based on the state and the domain in which the model is used.

8. The information processing apparatus according to claim 5, wherein the generation unit generates an action corresponding to the state based on the state and a generation model that generates an action corresponding to the state.

9. The information processing apparatus according to claim 1, wherein the learning unit learns a student network and a teacher network which are the models used for inference processing.

10. The information processing apparatus according to claim 9, wherein the learning unit executes a first learning process of learning the teacher network using the second data, and learns the student network by a second learning process using the teacher network learned by the first learning process and the first data.

11. The information processing apparatus according to claim 10, wherein in the second learning process, the learning unit executes an update process for the teacher network when the action selected by the teacher network is different from the action selected by the student network.

12. The information processing apparatus according to claim 11, wherein the learning unit updates the teacher network when the index value of the action selected by the student network is higher than the index value of the action selected by the teacher network.

13. The information processing apparatus according to claim 9, wherein the learning unit executes a first learning process of learning the teacher network and the student network using the second data, and learns the student network by a second learning process using the teacher network, the student network, and the first data learned by the first learning process.

14. The information processing apparatus according to claim 13, wherein in the first learning process, the learning unit replaces the parameters of the teacher network learned using the second data with the parameters of the student network.

15. The information processing apparatus according to claim 13, wherein the learning unit executes the second learning process using the second data used in the first learning process.

16. The information processing apparatus according to claim 13, wherein in the second learning process, the learning unit updates the parameters of the student network using the parameters of the teacher network.

17. The information processing apparatus according to claim 16, wherein in the second learning process, the learning unit replaces the parameters of the student network with the average value of the parameters of the teacher network and the parameters of the student network.

18. The information processing apparatus according to claim 1, further comprising a transmission unit that transmits the model learned by the learning unit to an external device that performs an inference process using the model.

19. An information processing method including: acquiring first data that is action history data used as learning data for offline reinforcement learning; generating second data that is data other than the first data and is data that can be used as learning data for offline reinforcement learning; and learning a model by the offline reinforcement learning using the first data and the second data. Obtaining first data that is action history data used as learning data for offline reinforcement learning; generating second data that is data other than the first data and is data that can be used as learning data for the offline reinforcement learning; and causing a computer to execute learning of a model by the offline reinforcement learning using the first data and the second data. An information processing program.

Citation Information

Patent Citations

  • Ramp control method based on offline reinforcement learning and macroscopic model

    CN114141029A

  • Arbitrary-angle inverted pendulum model training method based on reinforcement learning

    CN117313826A