Method and server for providing personalized recommendation based on reinforcement learning

By generating simulated users and using reinforcement learning models for training, the problem of lack of data in the reinforcement learning recommendation system is solved, personalized recommendation optimization under limited data conditions is achieved, and the efficiency and accuracy of the recommendation system are improved.

CN120677533APending Publication Date: 2025-09-19SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480014215.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-06
Filing Date
2024-01-15
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In reinforcement learning recommendation systems, the lack of data leads to inefficiency in model training and optimization of recommendation strategies. Existing technologies find it difficult to effectively utilize limited user data for personalized recommendations.

Method used

By generating simulated users, using reinforcement learning models to train and optimize recommendation strategies, generating personalized recommendation sessions, and generating virtual users based on user data, the simulated users continuously adjust and optimize the recommendation elements through simulated user status and reward feedback to provide personalized recommendations.

Benefits of technology

Under limited data conditions, the personalized recommendation effect of the recommendation system is improved, the accuracy of the recommendation and user experience are enhanced, the training time is reduced, and the adaptability and efficiency of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677533A_ABST
    Figure CN120677533A_ABST
Patent Text Reader

Abstract

The invention provides a method for providing personalized recommendation based on reinforcement learning. The method can comprise the steps of obtaining user data; generating a simulated user based on the user data, wherein the simulated user represents a virtual user corresponding to the actual user receiving the recommendation; determining an action based on the simulated user state, where the action determines a recommendation element to be included in a recommendation session to be provided to the user; updating the state of the simulation user; identifying an update state of the simulated user and a reward for the recommended element output from the simulated user; generating a personalized recommendation session by repeatedly determining the recommendation element based on an updated state of the simulated user and the reward; and outputting the personalized recommendation session.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method, a server, and an electronic device for providing personalized recommendations to users based on reinforcement learning using a user simulator. Background Art

[0002] Reinforcement learning is a field of machine learning in which an agent learns actions that maximize rewards while interacting with its environment. When reinforcement learning is applied to a recommendation system, the agent of the reinforcement learning model can learn a recommendation strategy that maximizes the rewards obtained in the process of taking actions while interacting with the user. At the same time, as in general machine learning techniques, the problem of lack of data is one of the important considerations in reinforcement learning. In some reinforcement learning algorithms, techniques such as transfer learning or pre-training are used to address the problem of lack of data. In view of the above situation, the present disclosure proposes a technology for accurately training and utilizing reinforcement learning while overcoming the problem of lack of data in reinforcement learning. Summary of the Invention

[0003] Solution to the problem

[0004] According to one aspect of the present disclosure, a method for providing a personalized recommendation session based on reinforcement learning, performed by a server, may be provided. The method may include obtaining user data. The method may include generating a simulated user based on the user data, the simulated user representing a virtual user corresponding to an actual user receiving recommendations. The method may include determining an action based on a state of the simulated user, wherein the action determines a recommendation element to be included in a recommendation session to be provided to the user. The method may include updating the state of the simulated user. The method may include identifying the updated state of the simulated user and a reward outputted from the simulated user for the recommendation element. The method may include generating a personalized recommendation session by repeatedly determining the recommendation element based on the updated state of the simulated user and the reward. The method may include outputting the personalized recommendation session.

[0005] According to one aspect of the present disclosure, a server for providing a personalized recommendation session based on reinforcement learning may be provided. The server may include: a communication interface; a memory storing at least one instruction; and at least one processor configured to execute the at least one instruction stored in the memory. The at least one processor may be configured to execute the at least one instruction to obtain user data. The at least one processor may be configured to execute the at least one instruction to generate a simulated user based on the user data, the simulated user representing a virtual user corresponding to an actual user receiving recommendations. The at least one processor may be configured to execute the at least one instruction to determine an action based on a state of the simulated user, wherein the action determines a recommendation element to be included in the recommendation session to be provided to the user. The at least one processor may be configured to execute the at least one instruction to update the state of the simulated user. The at least one processor may be configured to execute the at least one instruction to identify the updated state of the simulated user and a reward outputted from the simulated user for the recommendation element. The at least one processor may be configured to execute the at least one instruction to generate a personalized recommendation session by repeatedly determining the recommendation element based on the updated state of the simulated user and the reward. The at least one processor may be configured to execute the at least one instruction to output the personalized recommendation session.

[0006] According to one aspect of the present disclosure, a computer-readable recording medium may be provided, which stores a program for executing any one of the methods described above or to be described below to provide a personalized recommendation session based on reinforcement learning, the method being executed by a server. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 is a diagram schematically illustrating a server that provides a personalized recommendation session based on reinforcement learning according to an embodiment of the present disclosure.

[0008] Figure 2 is a flowchart for describing an operation performed by a server to provide a personalized recommendation session according to an embodiment of the present disclosure.

[0009] Figure 3 is a diagram illustrating an overall structure of a reinforcement learning model that provides exercise recommendations according to an embodiment of the present disclosure.

[0010] Figure 4 is a diagram for describing an operation of generating a simulated user performed by a server according to an embodiment of the present disclosure.

[0011] Figure 5 is a diagram for describing an operation of generating a personalized recommendation session performed by a server according to an embodiment of the present disclosure.

[0012] Figure 6is a diagram for describing an operation of retraining a simulated user performed by the server 2000 according to an embodiment of the present disclosure.

[0013] Figure 7A is a diagram for describing the operation of the user simulator according to an embodiment of the present disclosure.

[0014] Figure 7B is a diagram for describing the operation of the user simulator according to an embodiment of the present disclosure.

[0015] Figure 8 is a diagram for describing an operation of a training user simulator performed by a server according to an embodiment of the present disclosure.

[0016] Figure 9 is a diagram for describing an operation performed by a server according to an embodiment of the present disclosure to generate data for training a reinforcement learning model.

[0017] Figure 10A is a diagram for describing an operation performed by a server according to an embodiment of the present disclosure to obtain user data to provide a personalized recommendation providing service.

[0018] Figure 10B is a diagram for describing an operation performed by a server according to an embodiment of the present disclosure to obtain user data to provide a personalized recommendation providing service.

[0019] Figure 11A is a diagram for describing operations performed by a server according to an embodiment of the present disclosure for providing a personalized recommendation service.

[0020] Figure 11B is a diagram for describing an operation of obtaining feedback on a personalized recommendation providing service, performed by a server according to an embodiment of the present disclosure.

[0021] Figure 11C is a diagram for describing an operation of obtaining feedback on a personalized recommendation providing service, performed by a server according to an embodiment of the present disclosure.

[0022] Figure 11D is a diagram for describing an operation of obtaining feedback on a personalized recommendation providing service, performed by a server according to an embodiment of the present disclosure.

[0023] Figure 12 is a diagram for describing an operation of additionally providing information related to a provided recommendation session, performed by a server according to an embodiment of the present disclosure.

[0024] Figure 13 is a diagram for describing an example of a server providing a personalized recommendation session according to an embodiment of the present disclosure.

[0025] Figure 14 is a diagram for describing an example of a server providing a personalized recommendation session according to an embodiment of the present disclosure.

[0026] Figure 15 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.

[0027] Figure 16 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The terms used in this specification will be briefly described, and the present disclosure will be described in detail. In the present disclosure, the expression "at least one of a, b, or c" refers to "a," "b," "c," "a and b," "a and c," "b and c," "all of a, b, and c," or variations thereof.

[0029] Although the general terms currently in wide use are selected as the terms used in this disclosure while taking into account the functions of the present disclosure, they may vary according to the intentions of those skilled in the art, judicial precedents, the emergence of new technologies, etc. In addition, the terms arbitrarily selected by the applicant of this disclosure may also be used in specific cases. In such cases, their meanings will be described in detail in the detailed description of this disclosure. Therefore, the terms used in this disclosure must be defined based on the meaning of the terms and the content of the entire specification, rather than simply stating the terms themselves.

[0030] Unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" include plural referents. All terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art in which this specification is written. In addition, in this specification, although terms including ordinal numbers, such as "first," "second," etc., may be used herein to describe various components, the components should not be limited by these terms. These terms are only used to distinguish one component from another.

[0031] Throughout this specification, it should be understood that when a certain section "includes" a certain component, the section does not exclude another component, but may further include another component, unless the context clearly dictates otherwise. In addition, the terms "part," "section," "module," etc. used in this specification refer to a unit for processing at least one function or operation, which is implemented as hardware, software, or a combination of hardware and software.

[0032] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that a person skilled in the art can easily implement the present disclosure. However, the present disclosure is not limited to these embodiments and can be embodied in various other forms. In the accompanying drawings, parts not related to the description are omitted to clearly describe the present disclosure, and the same reference numerals represent the same elements throughout the specification.

[0033] Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings.

[0034] Figure 1 is a diagram schematically illustrating a server that provides a personalized recommendation session based on reinforcement learning according to an embodiment of the present disclosure.

[0035] refer to Figure 1 , the server 2000 according to the embodiment can provide personalized recommendations to users by using the reinforcement learning model 100.

[0036] In a general reinforcement learning system, an agent can determine actions and execute them while interacting with an environment. A reinforcement learning agent can be trained to select actions that maximize reward by identifying the current state and observing the environment.

[0037] In an embodiment, reinforcement learning model 100 may include a recommendation generator 102 and a user simulator 104. In the reinforcement learning model 100 of the present disclosure, recommendation generator 102 may act as an agent to determine recommended elements for a user. Furthermore, the actual user (or recommendation-providing application) receiving the recommended elements may correspond to the environment and interact with the agent.

[0038] In an embodiment, the reinforcement learning model 100 may generate a simulated user using a user simulator 104. A simulated user may refer to synthetic data representing a virtual user corresponding to an actual user. Furthermore, the user simulator 104 may be generative artificial intelligence. The user simulator 104 may have been pre-trained, and the user simulator 104 may have been trained to simulate a user to provide "simulated feedback" that is virtual feedback similar to feedback from an actual user. Embodiments of the present disclosure can provide the advantage of using only a small amount of actual user data to train the reinforcement learning model 100 by using simulated users.

[0039] In an embodiment, the recommendation generator 102 may perform an action to determine any recommended element from candidate recommended elements configured with recommendation elements m1 to mN, and the user simulator 104 may generate a simulated user that simulates an actual user, and the simulated user may interact with the recommendation generator 102. The recommendation generator 102 may observe, for example, the state of the simulated user and determine any one of the recommended elements m1 to mN based on the simulated user's state. The simulated user generated by the user simulator 104 may output a reward for the recommended element received from the recommendation generator 102, and the reinforcement learning model 100 may be trained to optimize the process of determining the action of selecting the recommended element to maximize the reward. In the present disclosure, the reinforcement model 100 may be a model pre-trained to provide recommendations to users, and the reinforcement model 100 may be referred to as a trained reinforcement learning model or a distributed reinforcement learning model. Furthermore, by omitting the modifier, the reinforcement model 100 may be simply referred to as a reinforcement learning model. Furthermore, the reinforcement learning model 100 may be retrained based on data collected while the user uses the reinforcement learning model 100.

[0040] In an embodiment, the server 2000 may provide at least one recommended session to the user by using the reinforcement learning model 100. For example, the server 2000 may generate a total of K recommended sessions, including session 1 (112), session 2 (114), ..., session K (116). In the present disclosure, the plurality of recommended sessions provided by the server 2000 may be referred to as a recommended session group 110.

[0041] In an embodiment, the server 2000 may provide the user with various categories of recommendation conversation groups 110. For example, the server 2000 may provide the user with recommendation conversation groups 110 corresponding to various categories, such as exercise recommendations, diet recommendations, and media content recommendations.

[0042] More specifically, for example, server 2000 may repeatedly generate exercise recommendation sessions to generate an exercise recommendation session group including multiple exercise recommendation sessions. In this case, the exercise recommendation session group may correspond to an entire exercise, which may be configured with multiple exercise recommendation sessions, and each of the multiple exercise recommendation sessions may be configured with an exercise recommendation element. Here, the exercise may indicate a specific movement of the exercise (e.g., squats, etc.). Server 2000 may provide the user with the exercise recommendation sessions, and the user may perform an exercise session by following the N exercises received from server 2000.

[0043] Hereinafter, unless otherwise specified, the example of a personalized recommendation session provided by the server 2000 according to the present disclosure to a user will be assumed to be an exercise recommendation scenario. However, this is merely an example for ease of description, and the personalized recommendations provided by the server 2000 according to the present disclosure can be applied to other recommendation categories besides exercise in the same / similar manner. That is, the present disclosure can be applied to various fields in which personalized recommendation sessions can be provided to users, and the present disclosure can provide personalized recommendation sessions to users by using users simulated based on reinforcement learning.

[0044] The operation of providing personalized recommendations to a user through reinforcement learning using a simulated user, performed by the server 2000 , will be described in more detail with reference to the accompanying drawings.

[0045] Figure 2 is a flowchart for describing an operation performed by a server according to an embodiment of the present disclosure to provide a personalized recommendation session.

[0046] Will refer to Figure 2 The overall operation of the server 2000 according to the present disclosure is described. Also, details of the operation of the server 2000 will be described with reference to the following drawings.

[0047] In operation S210, the server 2000 may obtain user data. The user data may be already stored in the server 2000 or may be received from the user's electronic device (mobile phone, etc.). The server 2000 may receive a recommendation service provision request from the user.

[0048] In an embodiment, user data may have been input by the user who will receive the recommendation session. User data may include a basic description for identifying the user's tendencies. For example, user data may include personal information about the user, such as gender and age. In addition, for example, user data may include user-input information related to the recommendation category. More specifically, in an exercise recommendation scenario, user data may include exercise style, exercise duration, focused muscles, exercise proficiency, etc. In addition, in a diet recommendation scenario for weight loss, user data may include information about food allergies, region, budget, weight, blood sugar level, dining locations, cooking accessibility, etc. In addition, in a long-term content recommendation scenario, user data may include genre preferences, language preferences, region, interests, hobbies, marital status, the presence of children, etc.

[0049] According to an embodiment, server 2000 may obtain user data corresponding to a recommendation category in order to provide recommendations of a specific category to the user. In this case, the user data may be input by the user. For example, server 2000 may obtain user data corresponding to a predetermined recommendation category. In addition, server 2000 may obtain user data and identify a recommendation category based on data elements included in the user data. Server 2000 may perform the following operations based on the user data to generate a recommendation session configured with multiple recommendation elements.

[0050] In operation S220 , the server 2000 may generate a simulated user representing a virtual user corresponding to an actual user who will receive a recommendation based on the user data.

[0051] In an embodiment, server 2000 can generate a simulated user by using a pre-trained generative artificial intelligence user simulator. In the reinforcement learning system according to the present disclosure, the actual user (or recommendation providing application) who will receive the recommendation can correspond to the environment. In addition, the generator that determines and provides the recommendation elements can correspond to the agent. As a detailed example, a simulated user can be obtained by simulating the actual user who will receive the exercise recommendation. The simulated user can be used to track the internal state of the actual user and manage the history of recommendations provided to the actual user. In addition, the simulated user can generate simulated feedback for training the reinforcement learning model. In this case, the user simulator may have been trained to generate simulated feedback similar to the feedback from the actual user based on the feedback data of the actual user.

[0052] In operation S230, the server 2000 may determine an action based on the state of the simulated user. Here, the action may be determining a recommended element from the recommended element candidates, and the action may be performed by the generator as an agent.

[0053] In an embodiment, to provide an exercise recommendation session, the server 2000 may determine an exercise as an exercise recommendation element related to exercise based on the simulated user's state. The exercise recommendation element included in the exercise recommendation session may refer to a basic exercise included in an exercise session. For example, the exercise recommendation element may include, but is not limited to, squats, burpees, etc.

[0054] The user's state can be divided into various segments (s=Concat(s1, s2, s3, ..., sn)). The segments of the user's state may include, for example, at least one of a user description d, a user preference p, a user's internal state i, or a recommendation history h.

[0055] In the example of a workout recommendation scenario, the user description may include personal information about the user and user-input information related to the recommendation category (eg, workout style, workout duration, focused muscles, workout proficiency, etc.).

[0056] Additionally, user preferences may include exercise preferences, exercise difficulty level preferences, variety preferences, feedback preferences, and the like.

[0057] In addition, the user's internal state may include elements that are affected when the user receives recommendations. For example, fitness improvement, real-time heart rate, muscle fatigue, etc. may be included in the user's internal state.

[0058] In addition, the recommendation history may include a history of recommended elements recommended to the user and feedback on the recommended elements. For example, exercise as a recommended element included in exercise and feedback on the exercise may be included in the recommendation history.

[0059] In operation S240, the server 2000 may update the state of the simulated user. This can be expressed as a natural update of the state, as the state is transformed according to the action according to the state transition function T(s, a, s'). The simulated user may be a virtual environment (virtual user) that simulates the actual environment (user). The state of the simulated user may be updated based on the state transition function, and the server 2000 may manage the interaction between the simulated user as the environment and the generator as the agent. The state transition function may be a function that defines how the agent, when interacting with the environment, transitions to the next state s' based on the current state s and the action a selected by the agent. In other words, the state transition function may include the probability that the current state transitions to the next state according to a specific action, and may be modeled during the training of the reinforcement learning model.

[0060] In an embodiment, according to the server 2000 selecting a sport as a recommendation element from among sport candidates, the user description d, user preference p, user internal state i, recommendation history h, etc. included in the state of the simulated user may be updated.

[0061] In operation S250 , the server 2000 may identify an updated state of the simulated user and a reward output from the simulated user.

[0062] The reward function R(s, a, s') can be a function that defines a numerical value provided when a state s transitions to the next state s'. In the present disclosure, the reward can be defined as the action corresponding to the agent observing the current state and determining a recommended element. That is, in the present disclosure, because the reward is independent of the state to which the environment transitions, the reward function can be defined as R(s, a) that is not affected by the next state s'. The agent can learn a strategy for determining actions based on the reward while interacting with the environment, and therefore, the agent will find better optimal actions as the request for recommendation session is repeated in the future.

[0063] Rewards can be designed to include negative rewards and positive rewards. For example, a negative reward may include a case where the user does not like the recommended element, and a positive reward may include a case where the user likes the recommended element.

[0064] In an embodiment, the server 2000 may provide a motion as a recommendation element included in a recommendation session in order to provide an exercise recommendation session, and identify a negative reward or a positive reward for the recommendation element.

[0065] For example, negative rewards may include cases where the user does not like the recommended exercise, where the user determines that the recommended exercise is too easy or too difficult, where the recommended exercise is excessively repeated, where the recommended exercise does not satisfy the user's exercise style, etc. However, negative rewards are not limited to the above examples.

[0066] For example, positive rewards may include a case where the user likes the recommended exercise, a case where the recommended exercise meets the user's preferred exercise difficulty level, a case where the recommended exercise improves exercise diversity in an exercise recommendation session, a case where the recommended exercise corresponds to a focused muscle in the user's data, etc. However, positive rewards are not limited to the above examples.

[0067] In operation S260 , the server 2000 may repeatedly determine recommendation elements based on the updated status and reward of the simulated user, thereby generating a personalized recommendation session.

[0068] In reinforcement learning, an agent can repeat a certain action while interacting with the environment until the reinforcement learning episode (sequence of actions) terminates.

[0069] In an embodiment, the server 2000 may repeatedly determine recommended elements based on the termination condition of the episode. For example, the server 2000 may determine recommended elements N times to generate a recommendation session including N recommended elements. More specifically, the server 2000 may generate an exercise recommendation session including N exercise recommendation elements.

[0070] In an embodiment, the server 2000 may repeatedly generate a recommendation session to generate a recommendation session group including multiple recommendation sessions. For example, the server 2000 may repeatedly generate an exercise recommendation session to generate an exercise recommendation session group including multiple exercise recommendation sessions. In this case, the exercise recommendation session group may correspond to an entire exercise, and the entire exercise may be configured as multiple exercise recommendation sessions, and each of the multiple exercise recommendation sessions may be configured with an exercise recommendation element.

[0071] In an embodiment, the server 2000 may provide a recommendation session to a user and obtain feedback from the user. The server 2000 may retrain a simulated user based on the feedback from the user. Whenever a recommendation session is generated so that the simulated user accurately simulates the behavior of an actual user, the server 2000 may retrain the simulated user, thereby providing the user with a gradually personalized recommendation session.

[0072] In operation S270, the server 2000 may output the personalized recommendation session. The server 2000 may transmit the personalized recommendation session to the user's electronic device (eg, mobile phone).

[0073] Figure 3 is a diagram illustrating an overall structure of a reinforcement learning model that provides exercise recommendations according to an embodiment of the present disclosure.

[0074] refer to Figure 3 , the actual user 300 who receives the exercise recommendation can interact with the reinforcement learning model 310 by using an application (e.g., a mobile application, a web application, etc.). The user can receive the exercise recommendation elements (actions) from the reinforcement learning model 310 and provide the reinforcement learning model 310 with feedback on the recommendation elements and exercise history (user history).

[0075] In an embodiment, a reinforcement learning model 310 that provides exercise recommendations may include a virtual environment (also referred to as an RL gym) 312 that is configured to represent the environment of the reinforcement learning model 310. In addition, the reinforcement learning model 310 may include an agent (also referred to as a generator) 316 that interacts with the virtual environment 312 or the user 300 as the actual environment.

[0076] In an embodiment, virtual environment 312 may include user simulator 314, which is synthetic data and functions to generate interactions. User simulator 314 may generate a simulated user as synthetic data, generate simulated feedback, and provide the simulated feedback to agent 316. In other words, user simulator 314 may generate synthetic data to train agent 316 using only limited real data.

[0077] In one embodiment, reinforcement learning model 310 may utilize a Markov decision process (MDP), a mathematical model that probabilistically models sequential decision problems. Agent 316 of reinforcement learning model 310 can select a specific action in each state and, when selecting an action in the current state, can determine the next state using a state transition function. The reward obtained in each state is determined by the action and can be defined to reflect the agent's ultimate goal (e.g., providing personalized exercise session recommendations). Virtual environment 312 can send the current state of the environment to agent 316, receive actions from agent 316, and then provide rewards and new states / observations to agent 316. For example, virtual environment 312 can send the current state of a simulated user to agent 316, receive exercise recommendation elements from agent 316, then update the simulated user's state and provide rewards and new states / observations to agent 316.

[0078] Hereinafter, states, actions, rewards, feedback, etc. will be described in detail to describe the operation of the reinforcement learning model 310 according to the present disclosure.

[0079] In one embodiment, virtual environment 312 can transfer a state to agent 316. That is, agent 316 can observe virtual environment 312 and identify a state. This state can include all the information needed to determine the reward at a given time and the transition to the subsequent state. The state can be divided into various segments (s = Concat(s1, s2, s3, ..., sn)).

[0080] In an exercise recommendation scenario, for example, some sections of the state may include but are not limited to at least one of user description d, user preference p, user internal state i, or recommendation history h.

[0081] The user description included in the status may include, but is not limited to, personal information related to the user (e.g., age, gender, etc.) and user-input information related to the exercises that are recommended categories (e.g., exercise goals, exercise style, exercise duration, focused muscles, exercise proficiency, fitness level, etc.).

[0082] The user preferences included in the state may include, but are not limited to, exercise preferences, exercise difficulty level preferences, variety preferences, heart rate preferences, feedback preferences (probability of providing a specific type of feedback), and the like.

[0083] The user's internal state included in the status may include elements that are affected when the user receives the recommendation. For example, the user's internal state may include, but is not limited to, fitness improvement, real-time heart rate, and muscle fatigue, etc.

[0084] The recommendation history included in the state may include a history of recommended elements recommended to the user and feedback on the recommended elements. For example, exercise as a recommended element included in exercise, feedback on the exercise, etc. may be included in the recommendation history.

[0085] In an embodiment, the segments included in the state may be classified as "observable" or "unobservable (or partially observable)". An observable state may be a state in which the agent can directly observe the exact state of the environment. For example, user descriptions (exercise goals, initial fitness level, exercise style, etc.) and recommendation history may be classified as observable. An unobservable or partially observable state may be a state in which the agent cannot directly observe the exact state of the environment, so the agent obtains information about the exact state of the environment through indirect observation or inference. For example, user preferences and user internal states may be classified as unobservable or partially observable. In this case, the MDP of the reinforcement learning model 310 may be a partially observable Markov decision process (POMDP).

[0086] In an embodiment, a state transition function may define how to transition from a current state to a next state when the agent 316 takes a specific action in a specific state. That is, the state transition function T(s, a, s') may be a function that defines how to transition from a current state s to a next state s' based on an action a selected by the agent 316 when the agent 316 interacts with the virtual environment 312 or the real user 300. In an embodiment, the segments included in a state may be classified as "static" or "dynamic."

[0087] For example, a user description and user preferences included in a state may be classified as static. In a state transition function, a static section may be defined as remaining the same even after a transition to the next state occurs.

[0088] For example, a user's internal state and the recommendation history included in that state can be classified as dynamic. In a state transition function, a dynamic segment can be defined as changing when a transition to the next state occurs. In this case, the dynamic segment can be classified as either "deterministic (p = 1)" or "non-deterministic (p ≠ 1)."

[0089] For example, recommendation history, as a dynamic segment, can be classified as deterministic because it is generated by adding a new recommendation element, determined by the current action, to the end of the recommendation session list. Furthermore, user feedback (such as likes / dislikes of recommendation elements) can be classified as non-deterministic because it is defined by probabilities. Similarly, state transitions within a user's internal state can also be classified as non-deterministic. Because the user's internal state is unobservable or partially observable, it can be defined and controlled by an internal model that can track and predict the user's internal state.

[0090] In an embodiment, the agent 316 may receive a reward as a result of the action selected. The reward function R(s, a, s') may be a function that defines a numerical value provided when a state s transitions to a next state s'. In the present disclosure, the reward may be defined as the action corresponding to the agent 316 observing the current state and determining a recommended element. That is, in the present disclosure, because the reward is independent of the state to which the environment transitions, the reward function may be defined as R(s, a) that is not affected by the next state s'. The agent 316 may learn a strategy for determining actions based on rewards while interacting with the environment (e.g., the virtual environment 312), and thus, the agent 316 will find better optimal actions as the request for a recommendation session is repeated in the future.

[0091] Rewards can be designed to include negative rewards and positive rewards. For example, a negative reward may include a case where the user does not like the recommended element, and a positive reward may include a case where the user likes the recommended element.

[0092] As detailed examples, negative rewards may include a case where the user does not like the recommended exercise, a case where the user determines that the recommended exercise is too easy or too difficult, a case where the recommended exercise is excessively repeated, a case where the recommended exercise does not satisfy the user's exercise style, etc. However, negative rewards are not limited to the above examples.

[0093] As detailed examples, positive rewards may include a case where the user likes the recommended exercise, a case where the recommended exercise meets the user's preferred exercise difficulty level, a case where the recommended exercise improves exercise diversity in an exercise recommendation session, a case where the recommended exercise corresponds to a focused muscle of the user's data, etc. However, positive rewards are not limited to the above examples.

[0094] In an embodiment, the virtual environment 312 may send an action mask to the agent 316 to increase the speed of the agent's 316 training process and impose restrictions on specific actions. For example, the action mask may mask a specific action with a 1 or a 0. In this case, an action indicated as 1 may be selectable, and an action indicated as 0 may be unselectable. The agent 316 may determine an action by considering only selectable actions. That is, the action mask may prevent the agent 316 from performing unnecessary actions, thereby reducing training time and facilitating effective training. For example, the action mask may restrict the agent's 316 actions. That is, the action mask may prevent the agent 316 from recommending exercises that the user dislikes or designates as unavailable.

[0095] Figure 4 is a diagram for describing an operation of generating a simulated user performed by a server according to an embodiment of the present disclosure.

[0096] In operation S410, server 2000 receives user data from a first user. The first user may be a user who wishes to receive a personalized recommendation session. For example, the first user may include a new user who has accessed server 2000 via an electronic device (e.g., a mobile phone) to use the personalized recommendation service. Upon the first user's access to server 2000, server 2000 may provide a user interface (e.g., a graphical user interface (GUI), a voice interface, etc.) for receiving user data. For example, server 2000 may provide a user interface that allows a first-time user to enter personal information such as age and gender. Alternatively, server 2000 may provide a user interface that allows the user to enter exercise-related information, including recommended categories such as exercise goals, exercise style, exercise duration, targeted muscles, exercise proficiency, and fitness level. For example, server 2000 may provide a questionnaire to obtain information related to exercise and generate user data based on the user's responses to the questionnaire.

[0097] In operation S420 , the server 2000 may cluster the second user based on the user data.

[0098] The second user may refer to a user other than the first user and may include, for example, an existing user using a personalized recommendation service. In this case, user data of the second user may exist. The server 2000 may obtain a cluster of second users having similar features to the first user based on the user data of the first user and the user data of the second user. For example, the server 2000 may select or extract key features from the user data of the first user and the second user and calculate similarity. In this case, various algorithms for measuring similarity (e.g., Euclidean distance measurement) may be used. The server 2000 may group users into clusters using a clustering algorithm (e.g., hierarchical clustering, K-means, etc.). The server 2000 may compare the user data of the first user with the cluster of the second user to identify the cluster to which the first user belongs.

[0099] In operation S430 , the server 2000 may generate a simulated user of the first user based on a simulated user corresponding to the cluster of the second user.

[0100] In an embodiment, because each user in the second user's cluster has previously been provided with personalized recommendation services, there may be simulated users corresponding to each second user. The server 2000 may set parameters to be included in the simulated user of the first user based on the parameters included in the simulated user corresponding to the second user's cluster. For example, the server 2000 may set parameters representing the user description p, user preference p, etc. of the simulated user of the first user.

[0101] In an embodiment, when server 2000 generates a simulated user for a first user, data for a second user may be insufficient. Server 2000 may identify the number of second users and, if the number of second users is a predetermined value or greater, cluster the second users. Furthermore, if the number of second users is less than the predetermined value, server 2000 may randomly set parameters to be included in the simulated user for the first user based on the user data of the first user.

[0102] Will refer to Figure 7A and Figure 7B The operation of the simulated user generated by the server 2000 through the user simulator is described in more detail.

[0103] Figure 5 is a diagram for describing an operation of generating a personalized recommendation session performed by a server according to an embodiment of the present disclosure.

[0104] refer to Figure 5 , operations S510 to S530 may correspond to Figure 2 In addition, in view of reinforcement learning, the generator may correspond to the agent, and the simulated user may correspond to the environment. Figure 5 The operations shown in FIG. 20 may be performed by the server 2000. For example, the server 2000 may perform data processing by using a user simulator or generator of a reinforcement learning model.

[0105] In operation S510, the server 2000 may output a recommendation element. Determining the recommendation element may correspond to the action of the agent. That is, the generator of the agent as the reinforcement learning model may output the recommendation element, and the output recommendation element may be stored in a list.

[0106] The server 2000 may select a recommended element from the candidate elements stored in the recommended element database 500. The generator may observe the simulated user and select a recommended element based on the simulated user's state. In this case, the recommended element may include various attributes. For example, in an exercise recommendation scenario, the exercise as a recommended element may include, but is not limited to, the attributes of the basic exercise, the difficulty level, the focused muscle, and the cardio / non-cardio condition.

[0107] A basic movement may be a basic movement for performing a particular type of exercise. The exercise may be performed in various ways using different variations of the basic movement or supplementary movements. For example, the basic movement may be a squat, in which case the squat may include a parallel squat, a full squat, a half squat, etc.

[0108] The difficulty level may have a predetermined range of values. A larger difficulty level value for an exercise may indicate a higher difficulty level for the exercise. For example, the difficulty level may be defined as a value from 1 to 4, and the difficulty level for a parallel squat may be defined as 1.4.

[0109] The focus muscles may refer to the muscles involved in performing the movement. For example, in the case of a squat, the focus muscles may include the quadriceps, hamstrings, etc.

[0110] Cardio / non-cardio can include whether the exercise is a cardiovascular exercise to improve cardiopulmonary function or a non-cardio exercise to strengthen muscles.

[0111] The server 2000 can output the recommended elements determined based on the state of the simulated user and the attributes of the recommended elements by using a reinforcement learning model. In other words, the action of selecting the recommended elements can be performed by the generator of the reinforcement learning model.

[0112] In operation S520, the simulated user may update the current state to the next state based on the recommendation element. That is, the simulated user may be transitioned to the next state based on the state transition function T(s, a, s'). For example, based on the simulated user receiving an exercise recommendation, the simulated user's dynamic state, such as the user's internal state i and recommendation history h, may be updated.

[0113] In operation S530, the simulated user may output an updated state and reward. The updated state and reward may be sent to the agent to allow the agent to determine the next action. For example, the agent may use the simulated user's updated state and reward to determine another exercise for the user to perform after performing a previously recommended exercise. The server 2000 may store the simulated user's previous and next states, actions, and rewards in the training database 502. The data stored in the training database 502 may be used to retrain the reinforcement learning model.

[0114] In an embodiment, the server 2000 may repeat operations S510 to S530 a predetermined number of times. For example, the server 2000 may repeat operations S510 to S530 N times, which is a predetermined number. In this case, N recommended elements may be stored in the recommended element list. When the server 2000 repeats operations S510 to S530 to output N recommended elements, it may be referred to as executing one episode.

[0115] In operation S540, server 2000 may output the completed recommended session to the user. For example, server 2000 may output an exercise session including N exercises to the user. In this case, the user can receive the exercise session from server 2000 and sequentially perform the exercises in the exercise session. In an embodiment, when server 2000 provides a recommended session, server 2000 may provide the user with information related to the recommended session. For example, when server 2000 provides an exercise recommendation session, server 2000 may provide the user with information such as images, videos, and audio, allowing the user to follow each exercise included in the exercise session. The information related to the recommended session provided by server 2000 may be output on the user's electronic device (e.g., a mobile phone).

[0116] In an embodiment, the server 2000 may repeat operation S540 of generating a recommendation session M times, where M is a predetermined number of times. The server 2000 may repeatedly generate the recommendation session to generate a recommendation session group configured with M recommendation sessions. For example, the server 2000 may repeatedly generate exercise recommendation sessions to generate an exercise recommendation session group configured with M recommendation sessions.

[0117] Figure 6 is a diagram for describing an operation of retraining a simulated user performed by the server 2000 according to an embodiment of the present disclosure.

[0118] In operation S610, the server 2000 may output at least one recommended conversation. Operation S610 may correspond to Figure 2 Operation S270 and Figure 5 Operation S540 is performed. For example, the server 2000 may generate a recommendation session configured with multiple recommendation elements and output the recommendation session, or may repeat the operation of generating a recommendation session and output multiple recommendation sessions. More specifically, the server 2000 may generate an exercise recommendation session group configured with exercise recommendation session 1, exercise recommendation session 2, ..., exercise recommendation session M. The operation of the server 2000 generating at least one recommendation session has been described above, and therefore, its repeated description will be omitted.

[0119] In operation S620, the server 2000 may receive feedback from the user regarding the recommended session. For example, when the server 2000 provides the recommended session, the server 2000 may receive feedback from the user. More specifically, when the user receives the recommended session and performs each of the exercises in accordance with the recommended session, the user may input immediate feedback. Alternatively, after the user receives the recommended session and performs all of the exercises in accordance with the recommended session, the user may input feedback regarding the recommended session.

[0120] In an embodiment, various types of feedback may be defined. Examples of feedback regarding a recommended exercise session may include, but are not limited to, completing the exercise, skipping the exercise, making the exercise more difficult or easier, a like / dislike response to the exercise, a score rating of the intensity of the exercise session, a score of satisfaction with the exercise session, etc.

[0121] In operation S630 , the server 2000 may retrain the simulated user based on feedback from the user.

[0122] In an embodiment, the reinforcement learning model may include a user simulator, which is a pre-trained generative artificial intelligence. The user simulator can generate a simulated user representing a virtual user corresponding to an actual user. In addition, the simulated user of the user simulator may have been trained to generate simulated feedback similar to feedback from an actual user. The server 2000 may perform retraining to improve the accuracy of the simulated feedback of the simulated user based on the feedback from the actual user obtained in operation S620.

[0123] In an embodiment, the server 2000 may repeatedly generate personalized recommendation sessions to generate a recommendation session group configured with multiple personalized recommendation sessions. In this case, the server 2000 may retrain the simulated user each time each of the multiple recommendation sessions is generated. That is, after the server 2000 provides a recommendation session, the server 2000 may provide the next recommendation session using the retrained user simulator.

[0124] In embodiments, a reinforcement learning model can be trained using simulated feedback output from a simulated user. As the reinforcement learning model provides a recommendation session to the simulated user and interacts with the simulated user, the reinforcement learning model can collect a recommendation history h, which includes actions and feedback regarding the actions. More specifically, the reinforcement learning model can provide an exercise recommendation session to the simulated user and collect a recommendation history h, which includes exercises included in the exercise recommendation session and feedback regarding the exercises. The recommendation history h can be one of the segments included in a state, and the simulated feedback can be used to train the agent to select the optimal action in each state while the reinforcement learning model interacts with the simulated user.

[0125] Will refer to Figure 7A and Figure 7B Describe simulated users and simulated feedback in more detail.

[0126] Figure 7A is a diagram for describing the operation of the user simulator according to an embodiment of the present disclosure.

[0127] In an embodiment, the server 2000 may train the user simulator so that the simulated user output generated by the user simulator is similar to the simulated feedback f' from the actual user. The server 2000 may adjust the hyperparameters of the user simulator to improve the operation of the user simulator. The hyperparameters may be variables used to control the operation of the user simulator, and by adjusting the hyperparameters, the simulated feedback f' of the user simulator may become similar to the feedback f from the actual user.

[0128] In an embodiment, states can be classified into observable states and unobservable (or partially observable) states. For example, in an exercise recommendation scenario, user description d (710) and recommendation history h (720) can be classified as observable states, and user preference p (730) and user internal state i (740) can be classified as unobservable states. Recommendation history h (720) can include action a (722) and feedback f (724). For example, recommendation history h = [(a1, f1), (a2, f2), ..., (an, fn)].

[0129] Observable states may be directly observable and known, and unobservable states may need to be inferred. The server 2000 may infer unobservable states by using a user simulator and generate simulation feedback f' based on observable states d and h and unobservable states p and i.

[0130] In an embodiment, the user simulator may include an encoder 700 for inferring unobservable states. The encoder 700 may be a neural network model that receives a user description d (710), an action a (722), and feedback f (724) as input and outputs a user preference p (730) and a user internal state i (740). Because the user internal state i (740) is determined based on a previous state, the encoder 700 may be implemented as a long short-term memory (LSTM) model for processing time series information. However, the implementation method of the encoder 700 is not limited thereto.

[0131] Will refer to Figure 7B The operations performed by the server 2000 to train the user simulator to output simulation feedback are further described.

[0132] Figure 7B is a diagram for describing the operation of the user simulator according to an embodiment of the present disclosure.

[0133] In an embodiment, encoding may be a process for obtaining a mapping from feedback f (724) to an unobservable state (i.e., (p, i) = Encode (f, o, a)). The server 2000 may use a user simulator to infer an unobservable state by encoding an observable state included in data of an actual user. That is, the user simulator may use the encoder 700 to encode the user description d (710), action a (722), and feedback f (724) as observable states, thereby inferring the user preference p (73) and the user internal state i (740) as unobservable states. The encoder 700 may include a plurality of parameters that can be adjusted through training.

[0134] In an embodiment, decoding may be a process for generating simulated feedback based on observable states and unobservable states (i.e., f'=Decode(z, d, a)). The server 2000 may use a user simulator to decode observable states and unobservable states, thereby generating simulated feedback representing feedback from a simulated user. That is, the user simulator may use a feedback generator 750 to decode the user description d (710), the action a (722), the user preference p (730), and the user internal state i (740), thereby generating simulated feedback. Here, the feedback generator 750 may correspond to a decoder. The feedback generator 750 may include a plurality of parameters that can be adjusted through training.

[0135] The server 2000 may generate simulated feedback f' and compare the simulated feedback f' with the feedback f from the user (724). The server 2000 may cause the simulated user to output feedback similar to the feedback from the actual user by minimizing the difference between the simulated feedback f' and the feedback f from the user (724). That is, the user simulator may be trained to find hyperparameters for minimizing the difference between the simulated feedback f' and the feedback f from the user (724).

[0136] In an embodiment, the server 2000 may obtain a data set D configured with data of actual users. r Dataset D r It can include recommendation histories (actions a (722) and feedbacks f (724)) generated by multiple different strategies (e.g., strategy A and strategy B) and user information (user description d (710)) that has received the recommendation. For example, dataset D r It may include a user description d and a recommendation history h=[(a1, f1), (a2, f2), ..., (an, fn)] corresponding to each of a plurality of users.

[0137] The server 2000 can train the data included in the dataset D by using the data corresponding to each user. rThat is, the server 2000 can use the user simulator of multiple users in the dataset D r The server 2000 performs encoding and decoding based on the data of each user in the training dataset to tune the hyperparameters of the user simulator. This may be done to reflect the unique state characteristics of each user, as users have different characteristics in terms of behavior patterns, preferences, etc. For example, the server 2000 may train the user simulator using data corresponding to user A among multiple users.

[0138] The server 2000 can train the user simulator by using a plurality of data sets configured with different users. For example, the server 2000 can train the user simulator so that the user simulator is trained by using a data set D configured with a first group of users. r1 and a dataset D configured with a second set of users r2 To reflect the characteristics of various user groups.

[0139] Figure 8 is a diagram for describing an operation of a training user simulator performed by a server according to an embodiment of the present disclosure.

[0140] In an embodiment, the user simulator of server 2000 can generate a simulated user. Furthermore, the user simulator can mimic an actual user and track and manage the user's overall behavior. For example, it can track whether the user likes or dislikes a particular exercise category, how exercise intensity affects the user's satisfaction, and can adjust the rate at which exercise intensity changes based on the user's fitness level. The user simulator can include hyperparameters 800 and perform the above operations using hyperparameters 800.

[0141] The server 2000 may tune the values ​​of the hyperparameters 800 so that the user simulator can mimic the behavior of the actual user, thereby optimizing the hyperparameters 800. For example, in order to make the simulated feedback 810 from the simulated user similar to the feedback 820 from the actual user, the server 2000 may adjust at least a portion of the hyperparameters 800 so that the preferences 830 of the simulated user are similar to the preferences of the actual user.

[0142] In an embodiment, the server 2000 may train a policy of the reinforcement learning model by using a user simulator in which an initial value of the hyperparameter 800 is set. Hereinafter, a process of training the reinforcement learning model and the user simulator will be described.

[0143] The server 2000 can provide recommended elements to actual users based on the strategy of the reinforcement learning model, and r In addition, the server 2000 can provide recommended elements to the simulated user based on the strategy of the reinforcement learning model, and in the simulated user dataset D s The actual user data set D is used to accumulate the simulated user data.r and simulated user dataset D s Each of them can include a user description and recommendation history (actions and feedback).

[0144] Server 2000 can simulate user data set D by using s To train the encoder 700 and decoder (feedback generator, not shown). The encoder 700 can be trained to infer unobservable states (including simulating the user's preferences 830), and the decoder can be trained to generate simulated feedback.

[0145] The server 2000 can generate simulated feedback 822 of an actual user by using the trained encoder 700 and the trained decoder. The server 2000 can optimize the hyperparameters 800 to minimize the difference between the simulated feedback 822 and the feedback 820 from the actual user. As a result of adjusting the hyperparameters 800, the encoder 700 can infer preferences close to the actual user's manually labeled preferences 842 when inferring the actual user's predicted preferences 840, and the decoder can generate feedback close to the feedback 820 from the actual user when generating the simulated feedback 822 of the actual user.

[0146] In an embodiment, the server 2000 may gradually train the user simulator. For example, as the number of users increases, the number of user data may also increase. Whenever the number of users increases by a predetermined standard, the server 2000 may repeatedly perform the above-mentioned operation of training the user simulator. For example, the server 2000 may use a data set D configured with multiple users. r1 To train the user simulator, and based on the number of pieces of user data that increases as time passes, the server 2000 can train the user simulator by using the dataset D including the increased users. r2 Alternatively, the server 2000 may retrain the user simulator based on various conditions. For example, the server 2000 may retrain the user simulator based on the passage of a predetermined time or when new data is obtained because the user has used the recommendation provision service. However, the conditions for the server 2000 to retrain the user simulator are not limited to the above examples.

[0147] In an embodiment, the server 2000 may use only a portion of the feedback included in the feedback from the actual user 820 as training data. The server 2000 may select a portion of the feedback included in the feedback from the actual user 820 based on a predetermined standard. For example, feedback may be classified into three types: positive feedback, negative feedback, and neutral feedback. More specifically, a case where the user enjoys an exercise may correspond to positive feedback, a case where the user changes the exercise to a more difficult exercise may correspond to negative feedback, and a case where the user completes the exercise may correspond to neutral feedback.

[0148] Meanwhile, neutral feedback among multiple feedback types may not accurately reflect the user's preference for the recommended element. For example, a user may provide neutral feedback on a recommended element even if he / she has previously provided positive / negative feedback on the recommended element, or even if he / she strongly likes or dislikes the recommended element. Therefore, the server 2000 can identify the type of feedback 820 from actual users and exclude specific types of feedback from the feedback types. For example, the server 2000 can perform a process of training a user simulator using the data remaining after excluding neutral feedback.

[0149] Figure 9 is a diagram for describing an operation performed by a server according to an embodiment of the present disclosure to generate data for training a reinforcement learning model.

[0150] refer to Figure 9 , Figure 9 The operations shown in FIG. 20 may be performed by the server 2000. For example, the server 2000 may perform data processing by using a user simulator or generator of a reinforcement learning model.

[0151] In operation S910, all simulated users may send their status to the generator (agent). In some embodiments, the user simulator may have generated multiple simulated users. In this case, the server 2000 may request all (or some) of the simulated users to operate. In response to the request from the server 2000 to operate all simulated users, each of the simulated users may send their current status to the generator to receive a recommendation session. For example, in an exercise recommendation scenario, the status may include at least one of the user description d, user preferences p, user internal state i, or recommendation history h.

[0152] In operation S920, the generator may output a recommended element to each of the simulated users and store the recommended element in a recommended element list. The generator may observe the status of each simulated user and select a recommended element based on the status of each simulated user. The generator may select a recommended element from the recommended element candidates stored in the recommended element database 900 and provide the selected recommended element to each simulated user. The same recommended element or different recommended elements may be provided to the respective simulated users.

[0153] In operation S930, each simulated user may update a current state based on the recommendation element. For example, according to each simulated user receiving an exercise recommendation, the dynamic state of each simulated user, such as the user internal state i and the recommendation history h, may be updated.

[0154] In operation S940, each simulated user may output an updated state and reward. The updated state and reward may be sent to the generator to enable the generator to determine the next action. For example, the generator may use the updated state and reward of each simulated user to determine another exercise that the user will perform after performing a previously recommended exercise. The server 2000 may store the previous state, next state, action, and reward of the simulated user in the training database 902. The data stored in the training database 902 may be used as data for training a reinforcement learning model.

[0155] Figure 10A is a diagram for describing an operation performed by a server according to an embodiment of the present disclosure to obtain user data to provide a personalized recommendation providing service.

[0156] In an embodiment, the server 2000 may provide a personalized recommendation service to the user. For example, the server 2000 may provide a personalized recommendation service through the application 1000, and the user may execute the application 1000 through his / her electronic device (e.g., a mobile phone) to access the server 2000 and receive the service.

[0157] In an embodiment, the server 2000 may obtain user data from the user through the application 1000. The server 2000 may provide various data and functions to enable the user to input user data through the application 1000. For example, the server 2000 may provide a graphical interface with a preset template to enable the user to easily input data, and may receive user input from the user.

[0158] In an embodiment, application 1000 may be an application that provides an exercise recommendation service. In this case, server 2000 may obtain user data related to exercise in order to provide exercise recommendations. For example, referring to first screen 1010 of application 1000 for user data input, server 2000 may receive user input for gender through application 1000 of the user's electronic device. Alternatively, for example, referring to second screen 1020 of application 1000 for user data input, server 2000 may receive user input for current height and weight through an application of the user's electronic device.

[0159] In an embodiment, the server 2000 may generate questions related to the recommended elements to obtain user data. For example, the server 2000 may generate questions related to exercise ability to obtain information about the user's exercise ability. More specifically, referring to the third screen 1030 of the application 1000 for user data input, the server 2000 may receive user input in response to a question output by an application on the user's electronic device (e.g., "Can you do more than 10 push-ups at a time?").

[0160] Meanwhile, the user data obtained by the server 2000 is not limited to the above examples. For example, based on the exercise recommendation scenario, the server 2000 may obtain user data including personal information related to the user (e.g., age, gender, etc.) and user input information related to the exercise as the recommended category (e.g., exercise goal, exercise style, exercise duration, focused muscles, exercise proficiency, fitness level, etc.).

[0161] Figure 10B is a diagram for describing an operation performed by a server according to an embodiment of the present disclosure to obtain user data to provide a personalized recommendation providing service.

[0162] In an embodiment, the server 2000 may obtain user data related to a recommendation category. For example, if the recommendation category is exercise, the server 2000 may generate a form for user data input related to exercise to provide exercise recommendations. For example, the server 2000 may generate a data input form such as "Create an exercise plan."

[0163] More specifically, referring to the fourth screen 1040 of the application 1000 for user data input, the server 2000 may generate a form for creating an exercise plan and receive user input from the user, such as selecting an exercise start date, an exercise style, etc. The exercise style may include, for example, high-intensity interval training (HIIT). Types of high-intensity interval training may include, but are not limited to, circuit training, Zuniga regimen, Tabata regimen, and Gibala regimen.

[0164] Alternatively, referring to the fifth screen 1050 of the application 1000 for user data input, the server 2000 may generate a form for formulating an exercise plan and select a user input from the user to select a focus muscle. The focus muscles may include, but are not limited to, abdominal muscles, back, biceps, chest, buttocks, hamstrings, front thighs, shoulders, and triceps.

[0165] Meanwhile, the recommendation categories of personalized recommendations provided by server 2000 are not limited to exercise. For example, recommendations for various categories such as dietary recommendations, media content recommendations, etc. can be provided. Server 2000 can provide the user with a form that allows the user to obtain user data corresponding to each recommendation category, and provide optimized recommendations for each category.

[0166] In an embodiment, when server 2000 receives user data, server 2000 may generate a simulated user based on the user data, representing a virtual user corresponding to the actual user receiving the recommendation. The simulated user may correspond to an environment in a reinforcement learning system, and at least a portion of the information included in the state may be set based on the user data obtained from the user. For example, a user description d may be set based on the user data, including personal information related to the user (e.g., age, gender, etc.) and user-input information related to the recommended exercise category (e.g., exercise goal, exercise style, exercise duration, focused muscles, exercise proficiency, fitness level, etc.). Alternatively, user preferences p may be set based on the user data, including exercise preferences, exercise difficulty level preferences, exercise variety preferences, heart rate preferences, feedback preferences (the probability of providing a specific type of feedback), and the like.

[0167] Figure 11A is a diagram for describing operations performed by a server according to an embodiment of the present disclosure for providing a personalized recommendation service.

[0168] In an embodiment, when the server 2000 receives a recommendation service provision request from a user, the server 2000 may generate a recommendation session by using a user simulator to generate a simulated user and repeatedly determining recommendation elements by simulating the interaction between the user and the recommendation generator. The operation of generating a recommendation session performed by the server 2000 has been described above, and therefore, a repeated description thereof will be omitted.

[0169] In an embodiment, the server 2000 can provide a personalized recommendation service through an application. In this case, the user can execute the application through his / her electronic device (eg, mobile phone, etc.) to access the server 2000 and receive a recommendation session.

[0170] A recommendation session can include multiple recommendation elements. For example, a recommendation session can be configured with N predetermined recommendation elements. More specifically, referring to first screen 1110 of the application providing the recommendation session, server 2000 can provide the user with an exercise recommendation session configured with 16 recommended exercises. First screen 1110 may include, but is not limited to, a list of recommended elements, recommended element items, and content related to the recommended elements (e.g., videos). More specifically, first screen 1110 may include a list of exercises, exercise items, execution times, and videos for following the exercises. Furthermore, first screen 1110 may include instructions for obtaining additional data from the user. More specifically, instructions such as "Connect your smartwatch to measure your heart rate" may be included in first screen 1110. In this case, server 2000 can receive sensor data from another electronic device of the user and, by additionally utilizing the sensor data, provide the user with a recommended exercise session. For example, server 2000 can receive heart rate data from the user's smartwatch and, using this heart rate data, provide the user with a recommended exercise session configured with recommended exercises.

[0171] Figure 11B is a diagram for describing an operation of obtaining feedback on a personalized recommendation providing service, performed by a server according to an embodiment of the present disclosure.

[0172] In an embodiment, the server 2000 may obtain feedback from the user who received the recommendation session. The server 2000 may modify at least a portion of the recommendation elements included in the current recommendation session based on the user feedback input by the user who received the recommendation session. Alternatively, the server 2000 may modify at least a portion of the recommendation elements included in the next recommendation session based on the user feedback input by the user who received the recommendation session.

[0173] In embodiments, user feedback can be used to modify a recommended element (or its attributes). For example, referring to the second screen 1120 of the application providing the recommendation session, the server 2000 may provide the user with at least one other recommended element, along with the currently recommended element, that can replace the currently recommended element. More specifically, the second screen 1120 may display "heel taps" as the currently recommended exercise, along with other recommended exercises that can replace "heel taps," such as "hollow body holds" or "crunches." Furthermore, when the server 2000 provides an alternative recommended element, it may provide the user with information related to the alternative recommended element. More specifically, the second screen 1120 may display the alternative recommended exercise "hollow body holds" with a difficulty level of "easy," and the alternative recommended exercise "crunches" with a difficulty level of "hard." Upon the user selecting an alternative recommended element, the server 2000 may modify the recommended element for the current or next recommendation session to the selected alternative recommended element.

[0174] Alternatively, the attributes of the recommendation element may be modified based on user feedback, which is not shown on the second screen 1120. For example, based on the user adjusting the difficulty level to "hard," the server 2000 may modify the exercise as in the above example, or may increase the difficulty level of the exercise by modifying the attributes of the exercise (e.g., number of executions, execution time, etc.) without modifying the exercise.

[0175] In an embodiment, modifying the recommendation session by reflecting user feedback in the server 2000 may include retraining and updating the reinforcement learning model.

[0176] Figure 11C is a diagram for describing an operation of obtaining feedback on a personalized recommendation providing service, performed by a server according to an embodiment of the present disclosure.

[0177] In an embodiment, the server 2000 may obtain feedback from users who have received a recommendation session. The server 2000 may modify at least a portion of the recommendation elements included in the next recommendation session based on user feedback input received after providing the recommendation session.

[0178] In embodiments, user feedback can be feedback used to modify recommended elements (or attributes of recommended elements). For example, referring to third screen 1130 of the application providing a recommendation session, server 2000 may provide the user with a feedback menu for the recommended elements included in the recommendation session after providing the recommendation session. The feedback menu for the recommended elements may be configured to provide feedback on each recommended element. More specifically, third screen 1130 may display a menu that provides feedback on exercises such as "reverse lunge," "assisted squat," "reverse plank," and "lateral plank" that were recommended in a recently provided exercise recommendation session. Based on the user feedback provided on the recommended elements, server 2000 may modify the recommended elements based on the feedback. For example, referring to third screen 1130, the user may provide feedback to lower the difficulty level of "assisted squat" and feedback to increase the difficulty level of "reverse plank." In this case, when server 2000 provides the next recommendation session, server 2000 may modify the recommended elements or adjust their attributes based on the user feedback.

[0179] In an embodiment, modifying the recommendation session by reflecting user feedback in the server 2000 may include retraining and updating the reinforcement learning model.

[0180] Figure 11D is a diagram for describing an operation of obtaining feedback on a personalized recommendation providing service, performed by a server according to an embodiment of the present disclosure.

[0181] Figure 11D An example of a portion of various methods in which the server 2000 provides a graphical interface to obtain feedback from a user is shown.

[0182] For example, referring to the fourth screen 1140 of the application providing the recommendation session, the server 2000 may provide a simple feedback menu for the recommendation session. The simple feedback may be, but is not limited to, a request for overall feedback on the recommendation session, rather than a request for feedback on each recommendation element included in the recommendation session. More specifically, after the server 2000 provides the exercise recommendation session, the server 2000 may generate questions requesting an evaluation of the exercise recommendation session. For example, the server 2000 may provide questions such as "Is the exercise intensity appropriate?" and "Are you satisfied with the recommended exercise?" and obtain user feedback in response to the questions. After the server 2000 obtains the user feedback, the server 2000 may modify at least a portion of the recommendation session based on the feedback.

[0183] By providing a simple feedback menu, the server 2000 may enable users to easily optimize a personalized recommendation session without the hassle of individually evaluating recommendation elements in the recommendation session.

[0184] As another example, referring to fifth screen 1150 of the application providing a recommendation session, server 2000 may provide a feedback menu that guides modification of the recommendation session. Guiding modification of the recommendation session may include, but is not limited to, guiding the user through modifiable recommendation elements and requesting feedback on whether to modify them. More specifically, server 2000 may guide the user through unlocking a more difficult exercise. For example, server 2000 may guide the user through the "Mountain Climber" exercise and unlock the next level of difficulty, "Twisted Mountain Climber," and obtain user feedback in response to whether the unlocking is required. Based on the user feedback obtained by server 2000, server 2000 may modify at least a portion of the recommendation session based on the feedback.

[0185] By providing the guide feedback menu, the server 2000 can enable the user to modify the recommended elements based on the guide when he or she wants to modify the recommended elements but does not know how to modify the recommended elements. Therefore, the server 2000 can enable the user to easily optimize the personalized recommendation session.

[0186] In an embodiment, modifying the recommendation session by reflecting user feedback in the server 2000 may include a process of retraining and updating the reinforcement learning model.

[0187] Figure 12 is a diagram for describing an operation of additionally providing information related to a provided recommendation session, performed by a server according to an embodiment of the present disclosure.

[0188] In an embodiment, when the server 2000 provides a personalized recommendation session to a user, the server 2000 may also provide the user with information related to the personalized recommendation session. The information related to the personalized recommendation session may include, but is not limited to, summary information of the provided recommendation session, information about the user receiving the recommendation session, etc.

[0189] The information related to the recommended session may depend on the category of the recommended session. Hereinafter, a case where the recommended category is exercise will be described as an example.

[0190] Referring to the first screen 1210 of the application providing the recommendation session, the first screen 1210 may include summary information of the exercise results. The summary information of the exercise results may include, but is not limited to, exercise time, exercise goal, heart rate results, and information about the sports included in the exercise recommendation session.

[0191] Referring to the second screen 1210 of the application providing the recommendation session, the second screen 1220 may include information indicating an exercise achievement level, which is information about the user who has received the exercise. The information indicating the exercise achievement level may include, but is not limited to, the exercise performed by the user, the achievement level of each exercise, the difficulty level of each exercise, etc.

[0192] Referring to the third screen 1230 of the application providing the recommendation session, the third screen 1230 may include a sports tree indicating the current fitness level, which is information about the user who has received the exercise. The sports tree may include, but is not limited to, sports unlocked according to the user's achievement level and information related to each sports.

[0193] According to an embodiment, the server 2000 can enhance the personalization effect by additionally providing information related to the personalized recommendation session. More specifically, Figure 12 As shown, the server 2000 can provide an exercise recommendation session and related information (eg, exercise results, achievement level, exercise tree suggesting the next exercise, etc.) based on the exercise result analysis, thereby enhancing the user's experience of receiving personalized recommendations.

[0194] Figure 13 is a diagram for describing an example of providing a personalized recommendation session performed by a server according to an embodiment of the present disclosure.

[0195] The server 2000 may provide personalized recommendation sessions for various categories. Figure 13 An example of personalized diet recommendation is shown. However, this is only an example to explain that the technical concept of the present disclosure is applicable to various recommendation categories and is not intended to limit the recommendation categories.

[0196] In an embodiment, providing a user with a personalized recommendation session for a particular recommendation category may require satisfying several constraints. As an example of constraints, for dietary recommendations, the user may need to be provided with a weekly meal plan, may need to follow the planned diet for several weeks, and may need to plan a variety of meals to reflect the user's preferences, allergies, and available food / ingredients. In addition, the various ingredients should not result in food waste, and the cost of maintaining the diet should remain within the budget. Server 2000 can provide a personalized dietary recommendation session to the user while satisfying various constraints by utilizing a reinforcement learning model that includes a user simulator.

[0197] In an embodiment, the server 2000 may provide a diet plan 1300 to a user by using a reinforcement learning model. The diet plan 1300 may include multiple diet recommendation sessions. For example, Figure 13 As shown, the diet plan 1300 may include diet Week 1 ( 1302 ), diet Week 2 ( 1304 ), ..., diet Week K ( 1306 ).

[0198] In an embodiment, the reinforcement learning model may include a diet recommendation generator 1310 that determines an action to generate a diet recommendation, and a user simulator 1320 that generates synthetic data to interact with the diet recommendation generator 1310 .

[0199] In the dietary recommendation scenario, User Simulator 1320 can generate a simulated user representing a virtual user receiving dietary recommendations. The simulated user can send a state to Diet Recommendation Generator 1310, which then determines the meals that constitute the recommendation. This state can contain all the information necessary to determine the reward at a given time and transition to subsequent states. The state can be divided into segments (s = Concat(s1, s2, s3, ..., sn)).

[0200] In the diet recommendation scenario, the state may include at least one of user description d, user preference p, user internal state i, or recommendation history h.

[0201] The user description included in the status may include, but is not limited to, personal information about the user (e.g., age, gender, etc.) and user-input information related to the diet being recommended (e.g., food allergies, regional food ingredient restrictions, the user's budget, the user's dining location, refrigeration requirements, access to cooking areas, initial weight, blood sugar, etc.).

[0202] User preferences included in the state may include, but are not limited to, food category preferences, food cooking time preferences, available ingredient preferences, recipe complexity, food / ingredient variety preferences, and feedback preferences (e.g., the probability that the user will provide feedback on skipping / substituting food / ingredients).

[0203] The user's internal state included in the state may include elements that are affected when the user receives a recommendation. For example, the user's internal state may include, but is not limited to, the user's weight, blood sugar, disliked ingredients, or the probability that the user will skip an ingredient when a complex recipe is recommended.

[0204] The recommendation history included in the state may include a history of recommendation elements recommended to the user and feedback on the recommendation elements. For example, food as a recommendation element included in a diet recommendation session and feedback on the food may be included in the recommendation history.

[0205] At the same time, the user's internal state and the recommendation history included in the state can be classified as dynamic. In the state transition function, a dynamic segment can be defined as changing according to the transition to the next state. In this case, the dynamic segment can be classified as "deterministic (p=1)" and "non-deterministic (p≠1)". For example, the recommendation history as a dynamic segment can be classified as deterministic because the recommendation history is generated by newly adding the recommendation element determined by the current action to the end of the recommendation session list. In addition, user feedback (such as like / dislike recommendation elements) can be classified as non-deterministic because user feedback is defined by probability. Similarly, the state transition of the user's internal state can also be classified as non-deterministic. Because the user's internal state is unobservable or partially observable, the user's internal state can be defined and controlled by an internal model that can track and predict the user's internal state.

[0206] Rewards can be designed to include both negative and positive rewards. For example, a negative reward could include a user disliking a recommended element, and a positive reward could include a user liking a recommended element. Rewards can be designed to reflect, for example, the variety of recommended foods for a meal, the cost of food ingredients, the nutritional diversity of the food, feedback on food likes / dislikes, and whether weight / blood sugar goals were achieved.

[0207] Meanwhile, in addition to the components of the reinforcement learning system using the user simulator as described above, specific operations including the interaction between the agent and the environment in the reinforcement learning system have been described with reference to the previous drawings, and therefore, repeated description thereof will be omitted.

[0208] Figure 14 is a diagram for describing an example of providing a personalized recommendation session performed by a server according to an embodiment of the present disclosure.

[0209] The server 2000 may provide personalized recommendation sessions for various categories. Figure 14 An example of personalized media content recommendation is shown. However, this is merely an example to explain that the technical concept of the present disclosure is applicable to various recommendation categories and is not intended to limit the recommendation categories.

[0210] In an embodiment, providing a personalized recommendation session for a specific recommendation category to a user may require satisfying several constraints. As examples of constraints, for media content recommendation, it may be necessary to recommend today's content to the user, the user may need to participate in the system for a long time, it may be necessary to reflect the user's preferences (type, playback time, viewing time zone), and it may be necessary to provide a variety of content. In addition, it may be necessary to provide a regular daily content viewing schedule, it may be necessary to integrate and recommend content provided from various content platforms (e.g., N content platform, Y content platform, etc.), and it may be necessary to recommend the next content based on previous content viewing history. Server 2000 can provide personalized media content recommendation sessions to users while satisfying various constraints by utilizing a reinforcement learning model that includes a user simulator.

[0211] In an embodiment, the server 2000 may provide media content recommendations 1400 to the user by using a reinforcement learning model. The media content recommendation 1400 may include multiple media content recommendation sessions. For example, Figure 4 As shown, media content Day 1 ( 1402 ), media content Day 2 ( 1404 ), . . . , media content Day K ( 1406 ) on the Kth day may be included in the media content recommendation 1400 .

[0212] In an embodiment, the reinforcement learning model may include a media content recommendation generator 1410 that determines an action to generate a media content recommendation, and a user simulator 1420 that generates synthetic data to interact with the media content recommendation generator 1410 .

[0213] In the media content recommendation scenario, the user simulator 1320 can generate a simulated user representing a virtual user receiving media content recommendations. The simulated user can transmit a state to the media content recommendation generator 1410, which can determine the content to be used as recommendation elements. This state can contain all the information required to determine the reward at a given time and transition to subsequent states. The state can be divided into various segments (s = Concat(s1, s2, s3, ..., sn)).

[0214] In the media recommendation scenario, the state may include at least one of user description d, user preference p, user internal state i, or recommendation history h.

[0215] The user description included in the status may include, but is not limited to, personal information about the user (e.g., age, gender, marital status, presence of children, interests, etc.) and user-input information related to media content as recommended categories (e.g., content genre, content language, actors, content region, previous residential history, etc.).

[0216] User preferences included in the state may include, but are not limited to, content genre preferences, actor preferences, region preferences, showtime preferences, content diversity preferences, and feedback preferences (eg, the probability that the user will provide feedback to skip / replace content).

[0217] The user's internal state included in the status may include elements that are affected when the user receives recommendations. For example, the user's internal state may include, but is not limited to, the user's tracked viewing time, service reconnection rate, content recommendation hit / miss ratio, content decision time, percentage of viewed content relative to total content, and content evaluation.

[0218] The recommendation history included in the state may include a history of recommendation elements recommended to the user and feedback on the recommendation elements. For example, content as a recommendation element included in a media content recommendation session and feedback on the content may be included in the recommendation history.

[0219] At the same time, the user's internal state and the recommendation history included in the state can be classified as dynamic. In the state transition function, a dynamic segment can be defined as changing according to the transition to the next state. In this case, the dynamic segment can be classified as "deterministic (p=1)" and "non-deterministic (p≠1)". For example, the recommendation history as a dynamic segment can be classified as deterministic because the recommendation history is generated by newly adding the recommendation element determined by the current action to the end of the recommendation session list. In addition, user feedback (such as like / dislike recommendation elements) can be classified as non-deterministic because user feedback is defined by probability. Similarly, the state transition of the user's internal state can also be classified as non-deterministic. Because the user's internal state is unobservable or partially observable, the user's internal state can be defined and controlled by an internal model that can track and predict the user's internal state.

[0220] Rewards can be designed to include both negative and positive rewards. For example, a negative reward could include a case where the user dislikes a recommended element, and a positive reward could include a case where the user likes a recommended element. Rewards can be designed to reflect, for example, the variety of content provided in a day, the variety of content provided across days, the number of pieces of content viewed, like / dislike feedback on the content, and whether the content was selected for recommendation.

[0221] Meanwhile, in addition to the components of the reinforcement learning system using the user simulator as described above, specific operations including the interaction between the agent and the environment in the reinforcement learning system have been described with reference to the previous drawings, and therefore, repeated description thereof will be omitted.

[0222] Figure 15 is a block diagram illustrating a configuration of a server according to an embodiment of the present disclosure.

[0223] In an embodiment, the server 2000 may include a communication interface 2100 , a memory 2200 , and a processor 2300 .

[0224] The communication interface 2100 may perform data communication with other electronic devices under the control of the processor 2300 .

[0225] The communication interface 2100 may include a communication circuit capable of performing data communication between the server 2000 and other devices through at least one of the data communication methods, wherein the data communication method includes, for example, wireless local area network (LAN), wireless fidelity (Wi-Fi), Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared communication (Infrared Data Association (IrDA)), Bluetooth low energy (BLE), near field communication (NFC), wireless broadband Internet (Wibro), world interoperability for microwave access (WiMAX), shared radio access protocol (SWAP), wireless Gigabit Alliance (WiGig) and radio frequency (RF) communication.

[0226] The communication interface 2100 can send / receive data to / from an external device for providing personalized recommendation services. For example, the communication interface 2100 can receive user data from a user's electronic device (e.g., a mobile phone, etc.), send a personalized recommendation session to the user's electronic device 2000, and receive user feedback from the user's electronic device.

[0227] Instructions, data structures, and program codes that can be read by the processor 2300 may be stored in the memory 2200. Operations performed by the processor 2300 may be implemented by executing instructions or codes of programs stored in the memory 2200.

[0228] The memory 2200 may include a flash memory type, a hard disk type, a multimedia card micro and a card type memory (e.g., a secure digital (SD) memory or an extreme digital (XD) memory), and may include a non-volatile memory and a volatile memory, the non-volatile memory including at least one of a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk or an optical disk, and a volatile memory such as a random access memory (RAM) or a static random access memory (SRAM).

[0229] The memory 2200 may store one or more instructions and / or programs to operate the server 2000 to provide a personalized recommendation session service. For example, the reinforcement learning model 2210 and the recommendation session management module 2220 may be stored in the memory 2200. The reinforcement learning model 2210 may include a user simulator 2212 and a generator 2214.

[0230] The processor 2300 may include the overall operation of the server 2000. For example, the processor 2300 may execute one or more instructions of a program stored in the memory 2200 to control the overall operation of the server 2000 to generate a personalized recommendation session. One or more processors 2300 may be provided.

[0231] Processor 2300 may be configured with, for example, at least one of a central processing unit, a microprocessor, a graphics processing unit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), an application processor, a neural processing unit, or an AI-specific processor designed with a hardware structure specifically for processing AI models, but is not limited thereto.

[0232] In an embodiment, the processor 2300 may use a reinforcement learning model 2210 to generate a personalized recommendation session. The reinforcement learning model 2210 may include a user simulator 2212 and a generator 2214. The user simulator 2212 may generate a simulated user representing a virtual user corresponding to an actual user receiving recommendations. The generator 2214 may generate recommendation elements included in the recommendation session provided to the user. In the reinforcement learning system according to the present disclosure, the actual user receiving the recommendation (or the recommendation-providing application) may correspond to the environment. Furthermore, the generator 2214, which determines and provides the recommendation elements, may correspond to an agent. In this case, the simulated user may be referred to as a virtual environment because the simulated user simulates the actual user. The simulated user can be used to track the actual user's internal state and manage the history of recommendations provided to the actual user. Furthermore, the simulated user can generate simulated feedback used to train the reinforcement learning model. In this case, the user simulator 2212 may have been trained to generate simulated feedback similar to feedback from the actual user based on feedback data from the actual user. The personalized recommendation session may be generated through interaction between the simulated user generated by the user simulator 2212 and the generator 2212.

[0233] The processor 2300 may provide a personalized recommended session (or a recommended session group) to the user by using the reinforcement learning model 2210. The processor 2300 may train the reinforcement learning model 2210 by using feedback from the user and / or synthetic data generated by the user simulator 2212.

[0234] In an embodiment, the processor 2300 may use the recommendation session management module 2220 to manage the recommendation session provided to the user. For example, the processor 2300 may use the recommendation session management module 2220 to manage the history of the recommendation session provided to the user, generate and manage information related to the recommendation session (e.g., the results of providing the recommendation session, summary information, information about the user receiving the recommendation session, etc.), and provide the information to the user.

[0235] Descriptions about operations of the reinforcement learning model 2210 and the recommendation session management module 2220 are included in those given with reference to previous drawings, and thus, repeated descriptions will be omitted.

[0236] Meanwhile, the modules and models stored in the memory 2200 are for the convenience of description and are not necessarily limited. Another module may be added to implement the above embodiment, and some of the above modules may be implemented as one module.

[0237] When the method according to an embodiment of the present disclosure includes multiple operations, the multiple operations can be performed by a single processor or multiple processors. For example, when the first operation, the second operation, and the third operation are performed by the method according to the embodiment, the first operation, the second operation, and the third operation can all be performed by the first processor, or the first operation and the second operation can be performed by the first processor (e.g., a general-purpose processor), while the third operation can be performed by the second processor (e.g., an AI-specific processor). As an example of a second processor, the AI-specific processor can perform operations for training / inferring an AI model. However, the embodiments of the present disclosure are not limited to this.

[0238] One or more processors according to the present disclosure may be implemented as a single-core processor or a multi-core processor.

[0239] In case that the method according to an embodiment of the present disclosure includes a plurality of operations, the plurality of operations may be performed by a single core or a plurality of cores included in one or more processors.

[0240] Figure 16 is a block diagram illustrating a configuration of an electronic device according to an embodiment of the present disclosure.

[0241] In an embodiment, the operations of the server 2000 described above may be performed by an electronic device 3000. The electronic device 3000 may be, for example, an electronic device of a user and may be implemented as various types of computing devices capable of operating a reinforcement learning model, such as a mobile phone, a tablet computer, etc. In other words, the operations of the present disclosure may be an on-device service that can be implemented by a single device. Thus, a user can receive personalized recommendation sessions using only their electronic device.

[0242] The electronic device 3000 may include a communication interface 3100, a memory 3200, and a processor 3300, and a reinforcement learning model 3210 and a recommendation session management module 3220 may be stored in the memory 3200. The reinforcement learning model 3210 may include a user simulator 322 and a generator 3214.

[0243] The operations of the communication interface 3100, the memory 3200 and the processor 3300 of the electronic device 3000 may be similar to those of Figure 15 The operations of the communication interface 2100, the memory 2200, and the processor 2300 of the server 2000 are described below, and therefore, repeated descriptions thereof will be omitted.

[0244] The present disclosure relates to a method for providing highly personalized recommendations to users based on reinforcement learning. Furthermore, the present disclosure relates to a method for training a high-performance reinforcement learning model using only a small amount of actual user data by using synthetic data generated by a user simulator. It should be noted that the technical objectives of the present disclosure are not limited to the above-mentioned technical objectives, and other technical objectives not mentioned will be apparent to those skilled in the art based on the description of this specification.

[0245] According to one aspect of the present disclosure, a method for providing a personalized recommendation session based on reinforcement learning, performed by a server, may be provided.

[0246] The method may include obtaining user data.

[0247] The method may include generating a simulated user based on the user data, the simulated user representing a virtual user corresponding to an actual user receiving the recommendation.

[0248] The method may include determining an action based on a simulated user state.

[0249] This action may determine recommendation elements to include in a recommendation session to be provided to the user.

[0250] The method may include updating the state of the simulated user.

[0251] The method may include identifying an updated state of the simulated user and a reward for the recommended element output from the simulated user.

[0252] The method may include generating a personalized recommendation session by repeatedly determining recommendation elements based on an updated state and reward of a simulated user.

[0253] The method may include outputting a personalized recommendation session.

[0254] The state of the simulated user may include at least one of a user description, a user preference, a user internal state, or a recommendation history.

[0255] The generation of the simulated user may include receiving user data from the first user.

[0256] The generating of the simulated user may include clustering the second user based on the user data.

[0257] The generating of the simulated user may include generating a simulated user of the first user based on the simulated users corresponding to the cluster of the second user.

[0258] The generating of the simulated users may include identifying a number of second users.

[0259] The generating of the simulated user may include randomly setting a parameter included in the simulated user of the first user based on the user data, based on the number of the second users being less than a predetermined value.

[0260] The method may include obtaining user feedback regarding the personalized recommendation session.

[0261] The method may include retraining the simulated user based on user feedback.

[0262] The method may include generating a recommendation session group configured with a plurality of recommendation sessions by repeatedly generating personalized recommendation sessions.

[0263] Whenever each of the plurality of personalized recommendation sessions is generated, retraining of the simulated user may be performed.

[0264] The generation of the simulated user may include generating the simulated user by using a user simulator that is pre-trained generative artificial intelligence.

[0265] The user simulator may have been trained to generate simulated feedback that is similar to feedback from actual users based on feedback data of actual users.

[0266] The reinforcement learning model that provides the personalized recommendation session may have been pre-trained using simulated feedback.

[0267] The personalized recommendation session may be an exercise recommendation session that includes exercise recommendation elements related to exercise.

[0268] According to an aspect of the present disclosure, a server for providing a personalized recommendation session based on reinforcement learning may be provided.

[0269] The server may include: a communication interface; a memory storing at least one instruction; and at least one processor configured to execute the at least one instruction stored in the memory.

[0270] The at least one processor may be configured to execute the at least one instruction to obtain user data.

[0271] At least one processor may be configured to execute at least one instruction to generate a simulated user based on the user data, the simulated user representing a virtual user corresponding to an actual user receiving the recommendation.

[0272] At least one processor may be configured to execute at least one instruction to determine an action based on a simulated user state.

[0273] This action may determine recommendation elements to include in a recommendation session to be provided to the user.

[0274] The at least one processor may be configured to execute at least one instruction to update a state of a simulated user.

[0275] The at least one processor may be configured to execute at least one instruction to identify an updated state of the simulated user and a reward for the recommended element output from the simulated user.

[0276] At least one processor may be configured to execute at least one instruction to generate a personalized recommendation session by repeatedly determining recommendation elements based on an updated state and a reward of a simulated user.

[0277] The at least one processor may be configured to execute at least one instruction to output a personalized recommendation session.

[0278] The state of the simulated user may include at least one of a user description, a user preference, a user internal state, or a recommendation history.

[0279] The at least one processor may be configured to further execute the at least one instruction to receive the user data from the first user.

[0280] The at least one processor may be configured to further execute the at least one instruction to cluster the second user based on the user data.

[0281] The at least one processor may be configured to further execute at least one instruction to generate a simulated user of the first user based on the simulated users corresponding to the cluster of the second user.

[0282] The at least one processor may be configured to further execute the at least one instruction to identify a number of the second users.

[0283] The at least one processor may be configured to further execute the at least one instruction to randomly set parameters included in the simulated user of the first user based on the user data and based on the number of the second users being less than a predetermined value.

[0284] The at least one processor may be configured to further execute at least one instruction to obtain user feedback regarding the personalized recommendation session.

[0285] The at least one processor may be configured to further execute at least one instruction to retrain the simulated user based on the user feedback.

[0286] The at least one processor may be configured to further execute the at least one instruction to generate a recommendation session group configured with a plurality of recommendation sessions by repeatedly generating the personalized recommendation session.

[0287] Whenever each of the plurality of personalized recommendation sessions is generated, retraining of the simulated user may be performed.

[0288] The at least one processor may be configured to further execute at least one instruction to generate a simulated user by using a user simulator that is a pre-trained generative artificial intelligence.

[0289] The user simulator may have been trained to generate simulated feedback that is similar to feedback from actual users based on feedback data from actual users.

[0290] The reinforcement learning model that provides the personalized recommendation session may have been pre-trained using simulated feedback.

[0291] The personalized recommendation session may be an exercise recommendation session that includes exercise recommendation elements related to exercise.

[0292] At the same time, the embodiments of the present disclosure may be implemented in the form of a recording medium including instructions executable by a computer, such as a program module executed by a computer. Computer-readable media can be any available media that can be accessed by a computer and can include volatile or non-volatile media and detachable or non-detachable media. In addition, computer-readable media may include computer storage media and communication media. Computer storage media may include volatile and non-volatile media and detachable and non-detachable media implemented by any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Communication media may include other data such as computer-readable instructions, data structures, or program modules that modulate the data signal.

[0293] Furthermore, computer-readable storage media may be provided in the form of non-transitory storage media. Here, the term "non-transitory storage medium" simply means that the storage medium is a tangible device and does not include signals (e.g., electromagnetic waves). However, the term does not distinguish between cases where data is semi-permanently stored in the storage medium and cases where data is temporarily stored in the storage medium. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0294] According to an embodiment, the methods according to various embodiments of the present disclosure may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., a compact disc read-only memory (CD-ROM)), distributed online via an app store (e.g., downloadable or uploadable), or distributed directly between two user devices (e.g., smartphones). When distributed online, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily generated or at least temporarily stored in a machine-readable storage medium, such as a memory on a manufacturer's server, an app store's server, or a relay server.

[0295] The foregoing description of the present disclosure is for illustrative purposes only, and it is apparent that a person skilled in the art may make various modifications to the present disclosure without changing the technical concept and essential features of the present disclosure. Therefore, it should be understood that the above embodiments are for illustrative purposes only in all aspects and are not intended to be limiting. For example, each component described as a single type may be implemented in a distributed type, and components described as distributed may be implemented in a combined form.

[0296] The scope of the present disclosure is shown by the appended claims rather than the detailed description, and should be interpreted as the meaning and scope of the claims and all modifications or modified forms derived from equivalent concepts thereof are included in the scope of the present disclosure.

Claims

1. A method for providing personalized recommendation sessions based on reinforcement learning, the method comprising: Obtain user data; generating a simulated user based on the user data, the simulated user representing a virtual user corresponding to an actual user receiving a recommendation; determining an action based on the simulated user state, wherein the action determines a recommendation element to be included in a recommendation session to be provided to the user; Updating the status of the simulated user; identifying an updated state of the simulated user and a reward for the recommended element output from the simulated user; generating a personalized recommendation session by repeatedly determining the recommendation elements based on the updated state of the simulated user and the reward; and The personalized recommendation session is output.

2. The method according to claim 1, wherein The state of the simulated user includes At least one of user description, user preferences, user internal state, or recommendation history.

3. The method according to claim 1, wherein The generation of the simulated user includes: receiving the user data from a first user; clustering the second user based on the user data; and A simulated user of the first user is generated based on a simulated user corresponding to the cluster of the second user.

4. The method according to claim 3, wherein: The generation of the simulated user includes: identifying the number of the second user; and Based on the number of the second users being less than a predetermined value, parameters included in the simulated user of the first user are randomly set based on the user data.

5. The method according to claim 1, Also includes obtaining user feedback regarding the personalized recommendation session; and The simulated user is retrained based on the user feedback.

6. The method according to claim 5, Also includes Generate a recommendation session group configured with multiple recommendation sessions by repeatedly generating the personalized recommendation session, in, Whenever each of the plurality of personalized recommendation sessions is generated, retraining of the simulated user is performed.

7. The method according to claim 1, wherein The generation of the simulated user includes The simulated user is generated by using a user simulator, which is a pre-trained generative artificial intelligence.

8. A server for providing personalized recommendation sessions based on reinforcement learning, the server comprising: Communication interface; a memory storing at least one instruction; and At least one processor configured to execute the at least one instruction stored in the memory to Get user data, generating a simulated user based on the user data, the simulated user representing a virtual user corresponding to the actual user receiving the recommendation, determining an action based on the simulated user state, wherein the action determines a recommendation element to be included in a recommendation session to be provided to the user, updating the state of the simulated user, identifying an updated state of the simulated user and a reward for the recommended element output from the simulated user, generating a personalized recommendation session by repeatedly determining the recommendation element based on the updated state of the simulated user and the reward, and The personalized recommendation session is output.

9. The server according to claim 8, wherein The state of the simulated user includes At least one of user description, user preferences, user internal state, or recommendation history.

10. The server according to claim 8, wherein The at least one processor is configured to further execute the at least one instruction to receiving the user data from the first user, clustering the second user based on the user data, and A simulated user of the first user is generated based on the simulated user corresponding to the cluster of the second user.

11. The server according to claim 10, wherein The at least one processor is configured to further execute the at least one instruction to identifying the number of the second user, Based on the number of the second users being less than a predetermined value, parameters included in the simulated users of the first user are randomly set based on the user data.

12. The server according to claim 8, wherein The at least one processor is configured to further execute the at least one instruction to obtaining user feedback on the personalized recommendation session, and The simulated user is retrained based on the user feedback.

13. The server according to claim 12, wherein The at least one processor is configured to further execute the at least one instruction to A recommendation session group configured with a plurality of recommendation sessions is generated by repeatedly generating the personalized recommendation session, and retraining of the simulated user is performed whenever each of the plurality of personalized recommendation sessions is generated.

14. The server according to claim 8, wherein The at least one processor is configured to further execute the at least one instruction to The simulated user is generated by using a user simulator, which is a pre-trained generative artificial intelligence. 15 . A computer-readable recording medium storing a program for executing the method according to claim 1 on a computer.