Method and electronic device for providing information using reinforcement learning
Patent Information
- Application Number
- PCT/KR2024/000122
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2024-01-03
- Publication Date
- 2025-05-22
AI Technical Summary
Online purchases face challenges in providing appropriate benefits to users due to lower purchase probability and repurchase probability compared to offline purchases, necessitating personalized promotions with discount rates that reflect user feedback.
An electronic device employs reinforcement learning to identify user purchase intent and infer product combinations and discount rates by applying data to AI models, considering user behavior, product relationships, and marketing data, and displays these combinations with adjusted discount rates based on user suitability and preferences.
This approach enhances the personalization of promotions, increasing the likelihood of product combinations being purchased by users and optimizing discount rates, thereby improving user engagement and sales in online services.
Smart Images

Figure KR2024000122_22052025_PF_FP_ABST
Abstract
Description
Method and electronic device for providing information using reinforcement learning
[0001] The disclosed embodiments relate to a method and an electronic device for providing product combinations and discount rates using reinforcement learning.
[0002] With technological advancements in online services and communication systems, personalized promotions are now offered in various forms and methods. Electronic devices learn to generate and deliver promotions tailored to each user.
[0003] Recently, with the advancement of artificial intelligence technology, discount rates can be calculated and presented to users for each product. For online purchases, compared to offline purchases, online users are less likely to purchase the products they browse and repurchase, making it necessary to provide more tailored benefits to users. Furthermore, considering the low repurchase rate among online shoppers, it's necessary to reflect user feedback on promotions to recommend products and offer promotions with discount rates.
[0004] According to one embodiment of the present disclosure, a method for an electronic device to provide a service to a user may be provided. The method may include obtaining data related to at least one of a user, multiple products, or marketing. The method may include identifying the user's purchasing intent based on the data. The method may include applying the user's purchasing intent and the data to an AI model to identify at least one product combination including at least two products from the multiple products and a discount rate for the at least one product combination. The method may include displaying the at least one product combination and the discount rate.
[0005] According to one embodiment of the present disclosure, an electronic device for providing a service to a user may be provided. The electronic device may include a communication unit, a memory storing one or more instructions, and at least one processor executing one or more instructions stored in the memory. The at least one processor may, by executing the one or more instructions, obtain data related to at least one of a user, a plurality of products, or marketing. The at least one processor may identify a user's purchasing intent based on the data. The at least one processor may identify at least one product combination including at least two products among a plurality of products and a discount rate for the at least one product combination by applying the user's purchasing intent and the data to an AI model. The at least one processor may display the at least one product combination and the discount rate.
[0006] According to one embodiment of the present disclosure, a computer-readable recording medium having recorded thereon a program for causing an electronic device to perform any one of the aforementioned and hereinafter described methods for providing a service to a user may be provided.
[0007] FIG. 1 is a diagram illustrating a method for an electronic device to provide a promotion according to one embodiment of the present disclosure.
[0008] FIG. 2 is a drawing for explaining an example of a promotion provided by an electronic device according to one embodiment of the present disclosure.
[0009] FIG. 3 is a flowchart illustrating a method for an electronic device to provide a promotion according to one embodiment of the present disclosure.
[0010] FIG. 4 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to provide a promotion using reinforcement learning.
[0011] FIG. 5 is a diagram illustrating an example of data obtained through interaction between a user and an online service according to one embodiment of the present disclosure.
[0012] FIG. 6 is a diagram illustrating a method for inferring a product combination according to one embodiment of the present disclosure.
[0013] FIG. 7 is a diagram for explaining an inference process among reinforcement learning algorithms according to one embodiment of the present disclosure.
[0014] FIG. 8 is a diagram illustrating an example of providing a promotion according to a user's priority according to one embodiment of the present disclosure.
[0015] FIG. 9 is a diagram for explaining time information included in data according to one embodiment of the present disclosure.
[0016] FIG. 10 is a diagram for explaining a method of considering time information in an algorithm according to one embodiment of the present disclosure.
[0017] FIG. 11 is a diagram illustrating an algorithm for providing a promotion by an electronic device according to one embodiment of the present disclosure.
[0018] FIG. 12 is a schematic block diagram of an electronic device according to one embodiment of the present disclosure.
[0019] Hereinafter, embodiments of the present disclosure are described in detail with reference to the attached drawings.
[0020] The present disclosure may be subject to various modifications and embodiments. Specific embodiments are illustrated in the drawings and described in detail herein. However, this is not intended to limit the embodiments of the present disclosure, and it should be understood that the present disclosure encompasses all modifications, equivalents, and alternatives falling within the spirit and technical scope of the various embodiments.
[0021] In describing the embodiments, detailed descriptions of related known technologies are omitted if they are deemed to unnecessarily obscure the gist of the present disclosure. Furthermore, numbers (e.g., "first," "second," etc.) used throughout the description of the specification are merely identifiers used to distinguish one component from another.
[0022] The terms used in the embodiments of this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on the meanings of the terms and the overall content of the present disclosure.
[0023] The scope of the present disclosure may be indicated by the claims that follow rather than the detailed description above. Various features mentioned in one claim category of the present disclosure (e.g., in a method claim) may also be claimed in another claim category (e.g., in a system claim). Furthermore, an embodiment of the present disclosure may include not only combinations of features specified in the appended claims, but also various combinations of individual features within the claims. The scope of the present disclosure should be interpreted to include all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts.
[0024] In addition, components expressed as 'unit', 'module', etc. in the present disclosure may be two or more components combined into one component, or one component may be divided into two or more components with more detailed functions. These functions may be implemented by hardware or software, or a combination of hardware and software. In addition, each component described below may additionally perform some or all of the functions performed by other components in addition to its own main function, and of course, some of the main functions performed by each component may be dedicated and performed by other components.
[0025] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art described herein.
[0026] Throughout this disclosure, unless specifically stated otherwise, "or" is inclusive and not exclusive. Thus, unless explicitly stated otherwise or the context dictates otherwise, "A or B" can mean "A, B, or both." As used herein, the phrases "at least one of" or "one or more of" can mean that different combinations of one or more of the listed items can be used, or that only any one of the listed items is required. For example, "at least one of A, B, and C" can include any of the following combinations: A, B, C, A and B, A and C, B and C, or A and B and C.
[0027] It will be appreciated that each block of the flowchart drawings and combinations of the flowchart drawings can be performed by computer program instructions. These computer program instructions can be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, such that the instructions, when executed by the processor of the computer or other programmable data processing equipment, create a means for performing the functions described in the flowchart block(s). These computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to perform the functions in a specific manner, such that the instructions stored in the computer-available or computer-readable memory can produce an article of manufacture that includes instruction means for performing the functions described in the flowchart block(s). Since the computer program instructions may be installed on a computer or other programmable data processing device, a series of operational steps may be performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for performing the functions described in the flowchart block(s) may also provide steps for performing the functions described in the flowchart block(s).
[0028] Additionally, each block may represent a module, segment, or portion of code that contains one or more executable instructions for performing a specific logical function(s). It should also be noted that in some alternative implementation examples, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on their respective functions.
[0029] The artificial intelligence-related functions according to the present disclosure are operated via a processor and memory. The processor may be comprised of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a Digital Signal Processor (DSP), a graphics-only processor such as a GPU or a Vision Processing Unit (VPU), or an artificial intelligence-only processor such as an NPU. One or more processors control the processing of input data according to predefined operating rules or artificial intelligence models stored in memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0030] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is trained using a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed on the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0031] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0032] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In addition, for the purpose of clearly explaining the present disclosure in the drawings, parts irrelevant to the description are omitted, and similar parts are designated with similar reference numerals throughout the specification.
[0033] The terms used in this disclosure will be briefly explained, and an embodiment of the present invention will be specifically described.
[0034] The terms described below are defined based on their functions within the present invention and may vary depending on the intent or custom of the user or operator. Therefore, their definitions should be based on the contents of this specification.
[0035] In this disclosure, "user's intent to purchase" may refer to the user's intent to purchase unspecified products in general, rather than a specific product. In other words, it may refer to the user's intent to purchase.
[0036] In the present disclosure, 'reward value' may mean a reward reflected as feedback for performing reinforcement learning.
[0037] In this disclosure, 'product group' refers to a group of products grouped by type, and may mean a group of products with a similarity level above a certain level.
[0038] In this disclosure, a 'flagship' product may mean a representative product or main product among multiple products.
[0039] In this disclosure, 'logit value' may mean decision confidence.
[0040] In the present disclosure, 'action space' may mean a set of all possible actions in a given environment.
[0041] In this disclosure, a 'hidden layer' exists between the input layer and the output layer of an AI model, and may mean a layer that forms a single neural network by continuously connecting models.
[0042] FIG. 1 is a diagram illustrating a method for an electronic device to provide a promotion according to one embodiment of the present disclosure.
[0043] Referring to FIG. 1, a method for providing a promotion to a user may include an electronic device (10). In one embodiment of the present disclosure, the electronic device (10) may determine and provide a promotion (160) through an AI model (140, 150) utilizing reinforcement learning based on data obtained from an online service (110) or an intranet (120) with which the user (20) interacted.
[0044] An electronic device (10) according to one embodiment of the present disclosure may be a device including a display. The electronic device (10) may be a device that outputs acquired promotions through a display. For example, the electronic device (10) in the present disclosure may include, but is not limited to, a smart TV, a smart phone, a tablet PC, a laptop computer, an e-book terminal, a digital broadcasting terminal, a Personal Digital Assistant (PDA), a Portable Multimedia Player (PMP), and the like. The electronic device (110) may be implemented as various types and forms of electronic devices including a display.
[0045] An electronic device (10) can obtain user-related data from an online service (110) that interacts with the user (20). Furthermore, the electronic device (10) can obtain data related to multiple users. The user-related data may include, but is not limited to, user behavioral data, location information related to the user's electronic device or location information using a tracking service, channel information from which the user was accessed (e.g., a website or social media page), device information accessed, campaign information that brought the user in, and the like. Furthermore, the data may include all data generated when the user (20) interacts with the online service (110). Examples of specific data will be described below with reference to FIG. 5.
[0046] The electronic device (10) can obtain product or marketing-related data from the intranet (120). Product-related data may include, but is not limited to, information about the product disclosed to users, the product group to which the product belongs, the product name, the product ID, and the product price. Marketing-related data may include, but is not limited to, the type of campaign, marketing or sales statistics resulting from the campaign, the marketing execution period, the discount rate based on the marketing, the campaign ID, information related to social media advertising, and information related to email marketing.
[0047] Additionally, the data acquired by the electronic device (10) may include the time the user (20) maintained the page within the session, the time interval between pages, etc. The electronic device (10) may store the acquired data in a database (130).
[0048] In one embodiment of the present disclosure, the electronic device (10) may perform preprocessing on acquired data or data stored in a database (130) to fit an AI model. Preprocessing of data may convert data into a form understandable by the AI model or improve the quality of the data to enhance the learning effect by utilizing high-level learning data. The electronic device (10) may convert the acquired data into a one-dimensional vector or integer value to generate and utilize a single one-dimensional vector. The electronic device (10) may use the preprocessed data as input for an AI model for identifying purchase intent or an AI model for inferring product combinations and discount rates. Furthermore, in one embodiment of the present disclosure, the electronic device (10) may store preprocessed data in the database (130) or update the stored data.
[0049] In one embodiment of the present disclosure, the electronic device (10) can estimate the purchase intent of a user (20) through a purchase intent estimation model (140) that has as input values data acquired or preprocessed or data stored in a database (130). The term "user purchase intent" in the present disclosure may refer to the user's purchase intent for unspecified products in general, rather than the user's purchase intent for a specific product. The specific product may refer to a product already purchased by the user. The purchase intent estimation model (140) may be implemented as an algorithm based on the acquired data. The purchase intent estimation model (140) may be a deep neural network-based classification inference model. The purchase intent estimation model (140) may identify that the user has a purchase intent if the probability that the user will purchase the product is greater than or equal to a preset probability. If the purchase intent estimation model (140) identifies that the user has a purchase intent, it may output '1' (or '0') as an output value, and if it identifies that the user has no purchase intent, it may output '0' (or '1') as an output value. The electronic device (10) can infer the purchase intention of the user (20) until a certain number of input and output values are obtained through the purchase intention estimation model (140).
[0050] In one embodiment of the present disclosure, if the electronic device (10) determines that the user has a purchasing intent, it can provide a promotion (160) to the user (20) using a product combination and discount rate inference model (150) having data about the user as input values.
[0051] The purchase intent estimation model (140) and the product combination and discount rate inference model (150) may be different AI models. Furthermore, the purchase intent estimation model (140) may be implemented using a simple algorithm, while the discount rate inference model (150) may be implemented using an AI model. Since they operate independently, the product combination and discount rate inference model (150) can be trained only on data indicating purchase intent, thereby increasing sample efficiency and reducing training time.
[0052] In one embodiment of the present disclosure, the product combination and discount rate inference model (150) can calculate user-specific preferred product groups and models. By inferring product combinations and discount rates from the same model, the weights of the deep neural network can interact. A product combination can include two or more products and products belonging to different product groups. The product combination and discount rate inference model (150) can perform reinforcement learning by reflecting user feedback based on product combinations or discount rates of product combinations through a reward function.
[0053] In one embodiment of the present disclosure, the product combination and discount rate inference model (150) can identify product combinations based on relationships between different product groups. These relationships can be categorized into preferences between product groups based on the customer experience journey, preferences between product groups based on inflow channels or networks, and preferences between product groups based on inflow digital marketing campaigns.
[0054] The product combination and discount rate inference model (150) can infer a discount rate within a valid integer range for a product combination if the user suitability for the obtained product combination exceeds a preset value. Discount rate inference can be performed by considering the margin for each product combination, user preferences, and constraints regarding campaigns or promotions. A specific product combination and discount rate inference method using reinforcement learning is described below in FIGS. 6 and 7.
[0055] In one embodiment of the present disclosure, an electronic device (10) may provide a user with a promotion (160) that includes a product combination and a discount rate. The electronic device (10) may provide information about the same product combination and discount rate to multiple users, and may change the display order by calculating a priority for each product combination included in the promotion. The electronic device (10) may determine the priority by multiplying the user's suitability for the product combination by the logit value (decision-making confidence) of the discount rate to be applied to each product combination. The logit value may correspond to the logarithm of the probability P of a specific event occurring. A specific example will be described below with reference to FIG. 8.
[0056] FIG. 2 is a drawing for explaining an example of a promotion provided by an electronic device according to one embodiment of the present disclosure.
[0057] Referring to FIG. 2, the electronic device (10) can infer product combinations (220, 230, 240) and discount rates for multiple products (210a, 210b, 210c, 210d) and provide them to the user.
[0058] In one embodiment of the present disclosure, the electronic device (10) can obtain data related to at least one of a user, a plurality of products (210a, 210b, 210c, 210d), or marketing.
[0059] In one embodiment of the present disclosure, user-related data may include, but is not limited to, user behavioral data, location information using the connected IP, channel information from which the user was accessed, device information from which the user was accessed, and campaign information that brought the user in. Furthermore, user-related data may include data on user interactions using online services. Data related to multiple products (120a, 120b, 120c, 120d) may include, but is not limited to, information about products disclosed to the user, product groups to which the products belong, product names, product IDs, and product prices. Furthermore, marketing-related data may include, but is not limited to, information about campaign types, sales statistics, marketing execution periods, discount rates based on marketing, campaign IDs, information related to social media advertising, and information related to email marketing.
[0060] The data acquired by the electronic device (10) may include information about the time the user stayed on a page within a session and the time interval between pages.
[0061] In one embodiment of the present disclosure, the electronic device (10) can determine a user's purchasing intent based on acquired data. The term "user purchasing intent" in the present disclosure may refer to a user's purchasing intent for unspecified products in general, rather than a user's purchasing intent for a specific product. The electronic device (10) can identify a user's purchasing intent using an algorithm. Furthermore, the electronic device (10) can identify a user's purchasing intent using an AI model. The AI model can output an indication of the user's likelihood of purchasing one or more products.
[0062] The electronic device (10) can extract and convert data from the acquired data for estimating the user's purchasing intent. The electronic device (10) can use the converted data as input to an AI model to infer the user's purchasing intent. For example, the electronic device (10) can infer that the user has a purchasing intent when the user places a product in the shopping cart, adds the product to the list of products of interest, or views a specific product a certain number of times or more.
[0063] The electronic device (10) can output '1' (or '0') as an output value when the user has a purchase intention, and '0' (or '1') as an output value when the user has no purchase intention. The electronic device (10) can infer the user's purchase intention until a certain number of input and output values are obtained.
[0064] In one embodiment of the present disclosure, if a user has a purchase intent, the electronic device (10) can apply the acquired data to an AI model to obtain product combinations (220, 230, 240) containing two or more products and discount rates for the product combinations. The AI model used to identify product combinations and discount rates may be different from the AI model used to estimate the user's purchase intent to increase sampling efficiency. In other words, by inferring product combinations and discount rates based only on data from users with a purchase intent and then providing promotions, learning efficiency can be improved.
[0065] In one embodiment of the present disclosure, an electronic device (10) can obtain a product combination (220, 230, 240) including two or more products by using a reinforcement learning model for a plurality of products (210a, 210b, 210c, 210d). The products included in the product combination may belong to different product groups.
[0066] For example, the electronic device (10) can infer product combination 1 (220) including product 1 (210a) and product 2 (210b) among multiple products (210a, 210b, 210c, 210d). In addition, the electronic device (10) can infer product combination 2 (230) including product 1 (210a) and product 3 (210c) and product combination 3 (240) including product 1 (210a), product 2 (210b), and product 3 (210c). Product 1 (210a), product 2 (210b), and product 3 (210c) belong to different product groups, and the electronic device (10) can identify the product combination based on the acquired data.
[0067] In one embodiment of the present disclosure, the electronic device (10) can maintain the time complexity of the algorithm for inferring a product combination at O(n). That is, when identifying a product combination containing n products, the electronic device (10) can maintain the complexity at O(n) by adding products included in the product combination one by one, starting from product 1.
[0068] In one embodiment of the present disclosure, the product combination and discount rate inference model (150) can identify product combinations based on relationships between different product groups. These relationships can be categorized into preferences between product groups based on the customer experience journey, preferences between product groups based on inflow channels or networks, and preferences between product groups based on inflow digital marketing campaigns.
[0069] In one embodiment of the present disclosure, the electronic device (10) can identify a discount rate for each product combination (220, 230, 240) if the user suitability for the identified product combination (220, 230, 240) is greater than or equal to a preset value. In one embodiment, the user suitability for the identified product combination may indicate the likelihood that the user will purchase the product combination. In one embodiment of the present disclosure, the electronic device (10) can identify a discount rate for the product combination by using the product combination as an input value of a hidden layer of an AI model. The hidden layer may refer to a layer located between the input layer and the output layer of the AI model. The electronic device (10) can identify a discount rate within a valid integer range by considering the margin for each product combination, the user's preference, and restrictions on campaigns or promotions using a reinforcement learning model. For example, when considering the margin for product 1 (210a), if a discount rate exceeding 20% is invalid, the discount rate for the product combination (220, 230, 240) including product 1 (210a) cannot exceed 20%. Or, when considering the flagship for product 3 (210c), if a discount rate exceeding 15% is invalid, the discount rate for the product combination (230, 240) including product 3 (210c) cannot exceed 15%.
[0070] The electronic device (10) provides the same information to multiple users regarding the acquired product combinations and corresponding discount rates, but can change the display order of the product combinations and discount rates based on the user's preferences. Furthermore, if the user's preference for a specific product is below a preset value, the electronic device (10) may not provide promotional information including that product. A specific example of providing promotional information is described below in FIG. 8.
[0071] FIG. 3 is a flowchart illustrating a method for an electronic device to provide a promotion according to one embodiment of the present disclosure.
[0072] In step S310, the electronic device (10) may acquire data. In one embodiment of the present disclosure, the data may refer to data related to at least one of a user, a product, or marketing. The data may refer to data acquired by the electronic device (10) from an online service or intranet with which the user interacted.
[0073] User-related data may include, but is not limited to, user behavior data, location information using the connected IP (e.g., location information obtained through the location of the user's electronic device or a tracking service), information on the channel through which the user was accessed (e.g., website, social media page), information on the device accessed, and information on the campaign that brought the user in. Furthermore, user-related data may include data on the user's interactions with online services. Product-related data may include, but is not limited to, information on products disclosed to the user, the product group to which the product belongs, the product name, the product ID, and the product price. Furthermore, marketing-related data may include, but is not limited to, the type of campaign, marketing or sales statistics resulting from the campaign, the duration of the marketing, the discount rate according to the marketing, the campaign ID, information related to social media advertising, and information related to email marketing. Furthermore, data acquired by the electronic device (10) may include the time the user spent on a page within a session, the time interval between pages, etc.
[0074] In one embodiment of the present disclosure, the electronic device (10) can perform preprocessing on the acquired data. The electronic device (10) can convert the acquired data into a one-dimensional vector or integer value to generate a single one-dimensional vector. Preprocessing on the data can convert the data into a form that an AI model can understand or improve its quality to enable effective learning using high-level learning data. Specifically, for data such as characters or words, a one-hot encoding vector can be used to convert the size of a word set into the dimension of a vector and utilize it. For image data, components can be classified by pixel, the size of the image can be normalized, a convolution feature map can be generated, and ultimately utilized in the form of a one-dimensional vector. The electronic device (10) can generate a single one-dimensional vector through preprocessing on the acquired data, perform a normalization operation on each element, and utilize it as an input for an AI model.
[0075] In one embodiment of the present disclosure, the electronic device (10) can perform preprocessing by classifying data element types. For example, the electronic device (10) can classify behavioral data, such as clicks, scrolls, adding to cart, and payments, which occur while a user uses an online service, by type. Furthermore, the electronic device (10) can classify data by the URL of the web page or by the product viewed.
[0076] In step S320, the electronic device (10) can identify the user's purchasing intent based on the data. The electronic device (10) can identify the user's purchasing intent through an algorithm using the data acquired in step S310. In one embodiment of the present disclosure, the electronic device (10) can identify the user's purchasing intent using an AI model. The term "user purchasing intent" in the present disclosure may refer to the user's purchasing intent for unspecified products in general, rather than the user's purchasing intent for a specific product.
[0077] In one embodiment of the present disclosure, an electronic device (10) can extract and convert data for identifying a user's purchasing intent. The electronic device (10) can set criteria for identifying the user's purchasing intent, classify data according to the set criteria, and use the data as input for an algorithm. For example, the criteria for identifying the user's purchasing intent by the electronic device (10) may include, but are not limited to, the user's online service usage or frequency, user behavior patterns based on the user's interests, whether or not the user has viewed promotions or benefits, or user behavior patterns associated with the process of purchasing a product. The user behavior patterns may correspond to, for example, the frequency of purchasing one or more products, the amount spent on products belonging to a specific category over a certain period of time, the number of times the product is viewed over a certain period of time, and the like. Specific examples will be described below with reference to FIG. 5.
[0078] In one embodiment of the present disclosure, an electronic device may utilize a deep neural network-based classification inference model to identify a user's purchasing intent. Input to the AI model may include multiple online users, product or marketing-related data, or preprocessed data.
[0079] If the electronic device (10) identifies that the user has a purchase intention, it can output '1' (or '0') as an output value, and if it identifies that the user has no purchase intention, it can output '0' (or '1') as an output value. The electronic device (10) can infer the user's purchase intention until a certain number of input and output values are obtained.
[0080] In step S330, the electronic device (10) can identify at least one product combination and a discount rate for the product combination using an AI model. In one embodiment of the present disclosure, if the electronic device (10) determines that the user has a purchase intent, the electronic device (10) can identify the product combination and the discount rate for the product combination using the AI model. The electronic device (10) can identify the discount rate for the product combination using the product combination as an input value of the hidden layer of the AI model. If the electronic device (10) determines that the user has a purchase intent and the user has searched for products belonging to two or more product groups, the electronic device (10) can identify the product combination using the AI model.
[0081] The AI model for identifying product combinations and discount rates can operate independently from the algorithm for identifying users' purchasing intent. By operating the AI model for identifying product combinations and discount rates only for users with purchasing intent, the sampling efficiency for model learning can be increased and training time can be shortened. The electronic device (10) can infer product combinations using a recurrent VAE and a reinforcement learning model. A detailed description of the method is provided below in FIG. 6.
[0082] In one embodiment of the present disclosure, the electronic device (10) can use an AI model to calculate a user-specific preferred product group, product, or preferred model. A product combination may include two or more products, and may include products belonging to different product groups. The electronic device (10) can acquire a product combination by reflecting user feedback based on the discount rate of the product combination.
[0083] In one embodiment of the present disclosure, the electronic device (10) can classify the relationships between different product groups into preferences between product groups based on the customer experience journey, preferences between product groups based on inflow channels or networks, and preferences between product groups based on inflow digital marketing campaigns. Based on the relationships between the classified product groups, the electronic device (10) can infer a combination of products appropriate for the user's characteristic distribution. In other words, based on the classified preference information, the electronic device (10) can detect changes in the user's preferences in real time and learn based on the preferences.
[0084] In one embodiment of the present disclosure, the electronic device (10) can sequentially identify products included in a product combination, one by one. That is, when one product is included, a combination including two products can be formed by adding one more product. Thereafter, a combination including three products can be formed by adding one more product to the product combination including two products, thereby sequentially obtaining product combinations. By sequentially identifying products included in a product combination, the electronic device (10) can maintain the complexity of the algorithm at O(n). The electronic device (10) can repeatedly perform the product identification operation until no products satisfying a preset criterion exist or until a preset number of products is satisfied.
[0085] The electronic device (10) can exclude products in the same category as those included in the product combination from the candidate list. Furthermore, the electronic device (10) can calculate a score (e.g., logit score) for each product combination and sort the scores in order to infer a discount rate for each product combination. The specific method by which the electronic device (10) infers product combinations is described below in FIG. 6.
[0086] In one embodiment of the present disclosure, product combination inference and discount rate inference are performed within the same AI model, allowing the weights of the deep neural network to interact with each other. For each product combination obtained, the electronic device (10) can obtain the discount rate for the product combination in descending order of the scores obtained for each product combination. The identification of the product combination and the discount rate for the product combination can be performed through the same AI model. By performing operations within the same AI model, feedback regarding both the product combination and the discount rate can be reflected to perform subsequent identification of the product combination and discount rate.
[0087] In one embodiment of the present disclosure, the electronic device (10) can identify a discount rate only for product combinations among the identified product combinations whose user suitability is greater than or equal to a preset value.
[0088] In one embodiment of the present disclosure, the electronic device (10) can calculate a discount rate for each combination using a reinforcement learning model that infers a discount rate. When calculating the discount rate, the AI model can learn a valid integer range of discount rates using backpropagation based on the margin for each product combination, user preferences, and constraints regarding campaigns or promotions. Specifically, the electronic device (10) can infer a discount rate for a product combination calculated by considering the cost of each of the multiple products included in the product combination, or by considering whether the product is a flagship product.
[0089] The electronic device (10) can obtain a discount rate within a valid integer range using an action space in reinforcement learning. A specific method for identifying the discount rate is described below in FIG. 7.
[0090] In step S340, the electronic device (10) may provide the acquired product combination and discount rate. The electronic device (10) may provide the user with the promotional product combination in order of valid discount rates.
[0091] In one embodiment of the present disclosure, an electronic device (10) provides information on the same product combination and discount rate to multiple users, but can change the display order of the product combination and discount rate based on the user's preferences. For example, for a user with a high interest in product A, promotional information for a product combination that includes product A may be provided preferentially over other product combinations. A specific example of an electronic device providing a promotion is described below in FIG. 8.
[0092] The electronic device (10) can provide information on the product combination and discount rate during the promotion period, and when the promotion is renewed thereafter, can provide information on the updated product combination and discount rate.
[0093] The electronic device (10) can perform reinforcement learning model training by reflecting the user's feedback information using a reward function in response to promotional information provided to the user. The training process of the reinforcement learning model using the reward function is described below in FIG. 4.
[0094] Figure 4 is a flowchart illustrating a method for an electronic device, according to one embodiment of the present disclosure, to provide a promotion using reinforcement learning. For the sake of brevity, any details overlapping with those in Figure 3 are omitted.
[0095] In step S405, the electronic device (10) can perform learning on a product combination inference model and a reinforcement learning model. The learning model can perform learning through an AI model, and will be discussed in detail below.
[0096] In step S410, the electronic device (10) may acquire data through an online service or intranet with which the user interacted. Specifically, the electronic device (10) may collect real-time user session information or marketing campaign or promotional information associated with the user session. Real-time user session information may include, but is not limited to, the websites visited by the user, the products viewed, the time spent on each visited web page, location information based on the access IP, information about the access device, or information about the channel through which the user entered the information.
[0097] The electronic device (10) can use the acquired data as an input value of the algorithm.
[0098] In step S415, the electronic device (10) may process the collected unstructured information and convert it into numerical information. The electronic device may perform preprocessing on the acquired data. Preprocessing on the data may convert the data into a form that the electronic device can understand or improve its quality to enable effective learning using high-level learning data. In one embodiment of the present disclosure, the electronic device (10) may identify a user's purchasing intent using an AI model. The electronic device may preprocess the data into a form that the AI model can understand to enable effective learning. In other words, the electronic device may convert or normalize the data into a form suitable as an input value for the AI model and use it. In one embodiment of the present disclosure, the electronic device (10) may perform preprocessing by classifying the data by type.
[0099] In step S420, the electronic device (10) can identify the user's intent to purchase based on preset criteria. For example, if the user's purchase probability (x%) calculated by the electronic device (10) is α% or higher or if there are products in the shopping cart, the user can be identified as having intent to purchase.
[0100] In one embodiment of the present disclosure, the electronic device (10) sets criteria for identifying a user's purchasing intent, classifies data based on the set criteria, and uses the data as input values for an algorithm. A specific example is described below with reference to FIG. 5.
[0101] In one embodiment of the present disclosure, the electronic device (10) can identify a user's purchasing intent using an AI model. The AI model may be a deep neural network-based classification inference model that uses acquired data as input values. If the electronic device (10) identifies that the user has a purchasing intent, it can output '1' (or '0') as an output value, and if the electronic device (10) identifies that the user has no purchasing intent, it can output '0' (or '1') as an output value. The electronic device (10) can infer the user's purchasing intent until a certain number of input and output values are acquired.
[0102] In one embodiment of the present disclosure, if the electronic device (10) determines that the user has a purchase intent, and if the user has viewed two or more different product groups, the electronic device (10) can identify product combinations and discount rates through an AI model. If the electronic device (10) determines that the user has no purchase intent or if the user has only viewed products belonging to the same product group, step S410 of collecting user, marketing, or campaign information can be performed without performing training on the AI model to improve training efficiency.
[0103] In step S425, the electronic device (10) can infer a promotional product combination. If the electronic device (10) identifies that the user has a purchasing intent and the user has searched for products belonging to two or more different product groups, the electronic device (10) can infer a product combination comprised of two or more products belonging to different product groups using an AI model.
[0104] The AI model for identifying product combinations can operate independently from the algorithm or AI model for identifying users' purchasing intent. This is to increase sample efficiency and shorten training time by training the AI model for identifying product combinations and discount rates only for users with purchasing intent. The electronic device (10) can infer product combinations using an AI model that applies a recurrent VAE and reinforcement learning model. A detailed description of the method is provided below in FIG. 6.
[0105] In one embodiment of the present disclosure, the electronic device (10) can maintain the number of nodes for inferring product combinations as O(n). That is, the time complexity of the algorithm can be maintained at n by gradually increasing the number of products included in the product combination. For example, if one product is included in the product combination, an operation can be performed to identify one additional product so that two products are included, and then another product can be identified so that three products are included. The electronic device (10) can repeatedly perform the operation of identifying products until no products that satisfy the set criteria exist or until a preset number of products is satisfied.
[0106] In one embodiment of the present disclosure, the electronic device (10) can classify the relationships between different product groups into preferences between product groups based on the customer experience journey, preferences between product groups based on inflow channels or networks, and preferences between product groups based on inflow digital marketing campaigns. Based on the relationships between the classified product groups, the electronic device (10) can infer a combination of products appropriate for the user's characteristic distribution. In other words, based on the classified preference information, the electronic device (10) can detect changes in the user's preferences in real time and learn based on the preferences.
[0107] In one embodiment of the present disclosure, an electronic device (10) can obtain feedback on a user's product combination and discount rate using a reward function using a learning model utilizing reinforcement learning. The electronic device (10) can identify a product combination based on the obtained feedback or reward value.
[0108] In step S430, the electronic device (10) can identify whether the user's suitability (y) for the acquired product combination is greater than or equal to a preset value (β). If the user's suitability for the presented product combination is less than β, the electronic device (10) can perform step S410, which collects information about the user and marketing campaigns, promotions, etc. associated with the user without inferring a discount rate for the product.
[0109] In one embodiment of the present disclosure, the electronic device (10) can calculate the user's suitability for the acquired product combination based on the acquired data. For example, the user's level of interest in the products included in the acquired product combination can be calculated based on factors such as the number of times the user has viewed the product, whether the product has been added to the shopping cart, and whether the user has a payment history for the product. Based on the user's level of interest in the products included in the product combination, the electronic device (10) can obtain the user's suitability (y) for the product combination and compare it with a preset value.
[0110] In step S435, the electronic device (10) can infer a discount rate within an integer range for the acquired product combination. If the user suitability for the product combination acquired in step S425 is greater than or equal to a preset value, the electronic device (10) can infer a discount rate for the corresponding product combination using an AI model. The electronic device (10) can learn the product combination and the discount rate for the product combination from the same model, allowing the weights of the deep neural network to interact.
[0111] In one embodiment of the present disclosure, the electronic device (10) can identify a valid range of discount rates through reinforcement learning, taking into account margins for product combinations, user preferences, and constraints on campaigns or promotions. For example, the electronic device (10) can infer a discount rate by considering margins based on the cost of multiple products included in the product combination and whether each product is a flagship product.
[0112] In one embodiment of the present disclosure, an electronic device (10) can obtain feedback on a user's product combination and discount rate using a reward function using a learning model utilizing reinforcement learning. The electronic device (10) can identify the discount rate for a product combination by reflecting the obtained feedback or reward value.
[0113] The electronic device (10) can obtain a discount rate within a valid integer range by utilizing an action space in reinforcement learning. A detailed method thereof will be described below in FIG. 7.
[0114] The logit value for the discount rate inferred by the electronic device (10) for a product combination is defined as "t." The logit value for the discount rate is a numerical value indicating the degree to which a high reward is expected, and can indicate decision-making confidence. Hereinafter, using "t," the priority for each promotion can be calculated, a reward function can be set, and the obtained reward value can be reflected in the model to perform reinforcement learning.
[0115] In step S440, the electronic device (10) may provide promotions to the user according to priority. The promotions may include the product combinations obtained in steps S425 and S435 and the discount rates corresponding to the product combinations. The priority may be calculated as the product of the user's suitability for the product combination and the decision-making confidence derived from the discount rate inference. Based on the calculated priority, the electronic device (10) may perform session-specific personalization to provide promotions to the user.
[0116] In one embodiment of the present disclosure, the priority for a promotion can be calculated according to the following mathematical formula (1).
[0117] … mathematical formula (1)
[0118] In mathematical formula (1), 'x' represents the user's purchase probability, 'y' represents the user's suitability for the product combination, and 't' represents the logit value of the discount rate inferred for the product combination. In mathematical formula (1), the higher the user's purchase probability, the larger the value of (1+x) is over 1, and the higher the user's suitability for the product combination, the larger the value of (1+y) is over 1, and the higher the priority. In addition, it can be confirmed that the larger the value of 't', which numerically represents the degree to which a high reward for the discount rate is expected, the higher the priority as it has a value closer to 1.
[0119] In one embodiment of the present disclosure, the electronic device (10) may provide promotions to users in descending order of priority according to mathematical formula (1). That is, the electronic device (10) may provide information on the same product combination and discount rate to multiple users, but may change the display order of the product combination and discount rate according to the priority of each user. Furthermore, depending on the user's priority, some product combinations included in the promotion may not be displayed. For example, to a user with a high interest in product A, promotional information for a product combination including product A may be provided preferentially over other product combinations. Furthermore, if the user's suitability for product combination B is below a preset value, the electronic device may not provide promotional information for product combination B to the user.
[0120] In step S445, the electronic device (10) can identify whether the user has viewed a product combination through the presented promotion. Depending on whether the user has viewed a product through the promotion, the electronic device (10) can reflect the user's feedback on the promotion through a reward function.
[0121] If the user does not view the product through the presented promotion, the electronic device (10) can reflect the reward value in model learning using mathematical expression (2) in step S450.
[0122] … Mathematical formula (2)
[0123] The variables in Equation (2) are the same as in Equation (1). Since the user does not view the product through the promotion, the reward value may be less than 1. Specifically, The value approaches 1 as the value of t increases, but is less than 1. In addition, since the user's purchase probability for x is less than 1, the reward value obtained by mathematical expression (2) is less than 1. That is, if the user did not search for the product combination through the promotion, it corresponds to negative feedback about the promotion, and therefore, a reward value less than 1 is reflected as the reward of reinforcement learning.
[0124] If the user searches for a product combination through the presented promotion, the electronic device (10) can determine whether the user ultimately purchases the product combination through the presented promotion in step S455. Depending on whether the user ultimately purchases the product combination through the presented promotion, the reward value reflected in model learning may be calculated differently. If the user does not ultimately purchase the product combination through the promotion, the reward value may be calculated using mathematical formula (3) in step S460. If the user ultimately purchases the product combination through the promotion, the reward value may be calculated using mathematical formula (4) in step S465 and reflected as the reward of the reinforcement learning model.
[0125] … Mathematical formula (3)
[0126] … Mathematical formula (4)
[0127] The variables in Equations (3) and (4) are identical to those in Equation (1). As users search for product combinations through promotions, the base of the exponential function in Equations (3) and (4) is the same (1+x), but the exponents differ (y and (1+y). Assuming that the variables are identical due to the different exponent values, we can see that the reward value obtained through Equation (4) is greater than that obtained through Equation (3). In other words, when a user ultimately purchases the corresponding product combination through a promotion, a larger reward value is reflected in reinforcement learning.
[0128] The reward values obtained through steps S450, S460, and S465 are reflected in model learning, and in step S470, the electronic device (10) can identify whether experience has been accumulated as much as the batch size. If sufficient experience has been accumulated, the electronic device (10) can perform model learning according to step S405, and if experience is insufficient compared to the batch size, the electronic device (10) can repeatedly perform the steps from step S410, which collects promotion and campaign information related to users and user sessions, until experience has been accumulated as much as the batch size.
[0129] FIG. 5 is a diagram illustrating an example of data obtained through interaction between a user and an online service according to one embodiment.
[0130] Referring to FIG. 5, in order to identify the user's purchasing intent, the electronic device (10) classifies the element types of the acquired data, performs preprocessing, and can use them as input values.
[0131] In one embodiment of the present disclosure, the electronic device (10) can perform preprocessing by classifying the element types of acquired data. The electronic device (10) can classify behavioral data, such as clicks, scrolls, adding to cart, and payments, by type, as events occurring while the user uses an online service. Preprocessing can be performed on the classified data and a benchmark score can be set for each data to identify the user's purchasing intent.
[0132] If a user is directed to a pre-order event by clicking (510) the pre-order event button, the electronic device (10) can obtain the user's inflow path and the URL of the web page, etc. In addition, whether the user clicks (510) the pre-order button to enter the purchase page (520) and clicks (530) the purchase button within the page can be used as data for identifying the user's purchase intent. In one embodiment, the user's inflow path may include a series of interactions with the web page (e.g., clicking links, selecting options, etc.).
[0133] A user can select a model option they wish to purchase within the purchase page (540). The model option selected by the user represents product-related data, including but not limited to the product group to which the product belongs, the product ID, and the product price. Furthermore, when a user clicks on various products to select a model option, or the time the user spends on the page to select a product (550), this data can be used to identify the user's purchasing intent, determine the user's suitability for a product combination, or identify a valid discount rate.
[0134] The user can select a color (560) and a payment method (570) for the selected model. For example, if the user selects a card or Samsung Pay, the electronic device (10) may additionally provide information on benefits based on the card type or Samsung Pay when offering a promotion.
[0135] When a user repeats the same purchase pattern (580), data indicating that an error has occurred within the page can be obtained, and when the user makes a final purchase click (590), the final purchase history can be used as data for identifying the user's purchase intent or for determining the user's suitability for a product combination.
[0136] FIG. 6 is a diagram illustrating a method for inferring a product combination according to one embodiment of the present disclosure.
[0137] Referring to Figure 6, a method for inferring a product combination consisting of K+1 products by adding 1 product to a product combination consisting of K products is illustrated.
[0138] In one embodiment of the present disclosure, the electronic device (10) can maintain the number of nodes for inferring product combinations as O(n). That is, the time complexity of the algorithm can be maintained as n by gradually increasing the number of products included in the product combination. Specifically, when configuring a product combination including k products among n products, In contrast to the time complexity of n, the time complexity can be maintained as n by inferring the product combination by adding one product at a time for the product combination consisting of K products.
[0139] In one embodiment of the present disclosure, the electronic device (10) can classify the relationships between different product categories into preferences between product categories based on the customer experience journey, preferences between product categories based on inflow channels or networks, and preferences between product categories based on inflow digital marketing campaigns. Based on the relationships between the classified product categories, the electronic device (10) can infer product combinations.
[0140] Hereinafter, f(K, i, j) denotes the product combination with the highest user relevance when the product category index of the last added product among the combinations consisting of K products is i and the product index is j. Vector for combinations consisting of K products (610a) may include vectors relating to product groups (620a) and products (630a). For example, (610a) has as elements the value 0100 for the product group (620a) and the value 001000 for the product (630a).
[0141] The electronic device (10) uses a reinforcement learning model to infer product combinations. For (610a), one product was added based on the acquired data. (610b) can be inferred. Products added to a product combination can be identified by excluding products with an index lower than the product index of the last product, and belonging to a different product category from the products included in the existing product combination.
[0142] Specifically, the identification of a product combination can be performed using an action space. Hereinafter, in the present disclosure, the action space may refer to a set of all possible actions in a given environment. In one embodiment of the present disclosure, the (K+1)th product may include products belonging to a different product group from the (K)th product. That is, the electronic device (10) may exclude products belonging to the same product group as the (K)th product from the action space. In one embodiment of the present disclosure, the (K+1)th product may have a value greater than the product index of the (K)th product. That is, the electronic device (10) may exclude products having a product index equal to or less than the product index of the (K)th product from the action space.
[0143] for example, For (610a), the electronic device (10) reflects f(K,i,j) to be the category of the Kth product. For items of the same category as (620a), exclude them from the action space. (620b) can be obtained. In addition, the electronic device (10) reflects f(K, i, j) to obtain the product index of the Kth product. (630a) Products with indices less than or equal to are excluded from the action space. (630b) can be obtained. In the action space, the electronic device (10) determines the K+1th product through f(K, i, j), (610b) can be obtained.
[0144] In one embodiment of the present disclosure, the electronic device (10) can recursively perform product combination inference until there are no valid products in the action space or until a preset number of products included in the product combination is satisfied.
[0145] Below, the specific operation method of the algorithm is explained in Fig. 7.
[0146] Figure 7 is a diagram illustrating the inference process of a reinforcement learning algorithm according to one embodiment of the present disclosure. Specifically, it is a diagram of the inference process of an algorithm for increasing the efficiency of neural network learning in a discrete action space.
[0147] Referring to FIG. 7, the electronic device (10) can select a valid action space by using the action space (710) and bit manipulation (720) in reinforcement learning.
[0148] Bit manipulation refers to the algorithmic manipulation of bits or data fragments shorter than words. Bit manipulation can eliminate or reduce the need to repeat data structures and enable high-speed data manipulation.
[0149] In one embodiment of the present disclosure, an electronic device (10) may utilize an action space when identifying a product combination or a discount rate for a product combination. Specifically, when the electronic device (10) identifies a product to be added to a product combination consisting of K products, the (K+1)th product may belong to a different product group from the products included in the product combination and have a product index greater than that of the (K)th product. To identify the (K+1)th product satisfying the above conditions, the electronic device (10) may utilize an action space to identify valid products.
[0150] Additionally, when the electronic device (10) identifies a valid range of discount rates for a product combination, it can exclude invalid discount rates from the action space through bit manipulation. For example, the electronic device (10) can calculate a range of possible discount rates for a total of 101 valid integer ranges of discount rates ranging from 0% to 100% depending on the product combination, and exclude from the learning range any discount rates that are not valid.
[0151] Existing methods that assign negative reward values when an impossible action space is selected, thereby reducing the probability of selecting that node, result in learning inefficiencies, as the probability of selecting an impossible node itself increases. Therefore, in one embodiment of the present disclosure, the logit value (decision confidence) for invalid nodes is set to negative infinity, thereby blocking the propagation of values within the neural network.
[0152] Among reinforcement learning algorithms, value-based reinforcement learning algorithms make decisions based on values. Specifically, they learn a Q function to obtain an optimal Q function, and then make decisions based on this optimal Q function. In value-based reinforcement learning algorithms, values set to negative infinity have a zero probability of being selected, so they can be excluded from model inference and learning.
[0153] Furthermore, among reinforcement learning algorithms, policy-based reinforcement learning algorithms, unlike value-based reinforcement learning, do not utilize value functions in decision-making. Instead, they directly learn policies and make decisions based on the learned policies. In policy-based reinforcement learning, an objective function is set, and the optimal policy has a higher objective function value. Accordingly, reinforcement learning can be performed so that the objective function increases over time. In the case of policy-based algorithms, the action-probability distribution function The action is determined accordingly. Therefore, it is a necessary and sufficient condition to output a probability of 0 for actions for which the probability distribution is invalid. To convert the logit value into a probability value, the softmax function of the above mathematical expression (5) can be used.
[0154] … Mathematical formula (5)
[0155] In mathematical formula (5) If the value of has negative infinity (-∞), Since it converges to 0, invalid actions can be excluded from the model's inference and learning.
[0156] FIG. 8 is a diagram illustrating an example of providing a promotion according to a user's priority according to one embodiment of the present disclosure.
[0157] Referring to FIG. 8, it can be confirmed that the order in which multiple product combinations (820, 830, 840) are displayed for multiple users (810a, 810b, 810c) is different.
[0158] The electronic device (10) can identify product combinations (820, 830, 840) and discount rates for each product combination using an AI model utilizing reinforcement learning for multiple products. The electronic device (10) can infer product combinations including multiple products based on data acquired from an online service or an intranet, and if the user suitability for the product combination is greater than a preset value, the electronic device can infer a valid discount rate for the product combination.
[0159] The electronic device (10) can calculate a priority value for each product combination based on the user's product purchase probability, the user's suitability for the product combination, and the logit value (decision confidence) for the discount rate, for promotional information including product combinations and discount rates. The electronic device (10) can provide the same product combination and discount rate information to multiple users, but can change the display order of the product combinations based on the acquired priority values. Furthermore, in one embodiment of the present disclosure, the electronic device (10) can not provide a product combination if the user's priority value for the product combination is below a preset value.
[0160] For example, for user 1 (810a), all of the inferred product combinations 1 (820), 2 (830), and 3 (840) may be provided. In addition, if product combinations 1 (820), 2 (830), and 3 (840) have high priority values in that order, the electronic device (10) may display them in ascending order of priority. For user 2 (820a), since the priority value for product combination 1 (820) among the inferred product combinations is lower than a preset value, information on the corresponding product combination may not be provided. In the case of user 3 (810c), if the interest in product 3 is high among product combinations 1, 2, and 3, and accordingly, the priority values for product combinations are high in the order of product combination 2 (830), product combination 3 (840), and product combination 1 (820), the electronic device (10) can display the priorities in ascending order.
[0161] FIG. 9 is a diagram for explaining time information included in data according to one embodiment of the present disclosure.
[0162] Referring to FIG. 9, when a user searches for multiple products (910a, 910b, 910c, 910d) over time, the electronic device (10) can obtain information about the time or timing.
[0163] In one embodiment of the present disclosure, an electronic device (10) may obtain data from an online service or an intranet. The data may include data related to a user, a product, or marketing, and may include time information obtained as a result of a user's interaction with the online service.
[0164] In one embodiment of the present disclosure, the electronic device (10) may set different weights for each piece of data based on acquired time information. Specifically, the data may include information (920a, 920b, 920c, 920d) regarding the timing at which a user viewed a specific product. The further the timing at which a specific product was viewed is from the predicted timing (940), the lower the weight may be set for the event of viewing the product. The predicted timing (940) may include the timing at which the electronic device (10) identifies the user's purchasing intent, the timing at which the product combination is identified, the timing at which the user's suitability for the product combination is identified, or the timing at which the discount rate is identified.
[0165] For example, the time furthest from the predicted timing (940) The weight can be set low for information about product 5 (910a) searched. The weight can be set high for information about product 3 (910d) searched closest to the predicted timing (940).
[0166] In one embodiment of the present disclosure, the electronic device (10) can identify a user's preference for a specific product based on acquired time information. Specifically, the data may include information (930a, 930b, 930c) indicating the time the user spent on a web page while browsing a specific product. The longer the user spent on the web page, the higher the user's preference for the product can be identified.
[0167] For example, for multiple products (910a, 910b, 910c, 910d) viewed before the predicted timing (940), the time spent on the page while viewing each product is the time spent while viewing product 3 (910b). Based on the information about (930b), the user's preference for product 3 (910b) can be identified as high.
[0168] In one embodiment of the present disclosure, the electronic device (10) can utilize the acquired time information as data for identifying the user's purchasing intent, data for identifying a product combination, data for determining the user's suitability for the product combination, data for identifying a discount rate for the product combination, or data for providing a promotion. For example, when the electronic device identifies a product combination, the electronic device can use the time closest to the timing (940) for predicting the products included in the product combination. A high weight can be set for product 3 (910d) viewed at (920d). In addition, the electronic device can set the product that has been viewed for the longest time before the predicted timing (950). (930b) It can be identified that the user has a high preference for the viewed product 3 (910b). Accordingly, the electronic device can identify product 3 as a product (950) to be included in the product combination.
[0169] FIG. 10 is a diagram for explaining a method of considering time information in an algorithm according to one embodiment of the present disclosure.
[0170] Referring to Figure 10, two methods of applying additional data on time information in a method of training AI are proposed.
[0171] In one embodiment of the present disclosure, the temporal information may include temporal information obtained as a result of a user's interaction with an online service. Specifically, the temporal information may include, but is not limited to, predicted timing, the timing of an event occurrence, the time spent within a specific session, and the time interval between pages. Below, a method of utilizing temporal information for AI model training is described.
[0172] The first method (1010) involves adding time information to the product itself. Product information may include, but is not limited to, information about the product disclosed to the user, the product group to which the product belongs, the product name, the product ID, and the product price. In one embodiment of the present disclosure, the method for adding time information to the product itself may include, but is not limited to, the timing of the product being viewed or the time spent on the page where the product is viewed.
[0173] For example, the electronic device (10) may add information about how far back in time the product was viewed from the predicted timing to the product information itself. Furthermore, the electronic device (10) may add information about the time the user spent viewing the product to the product information itself.
[0174] In one embodiment of the present disclosure, the electronic device (10) can normalize product information and time information to create a single set. The electronic device (10) can divide the total time into N groups. Based on the divided time, the electronic device (10) can reflect time information into the product information.
[0175] The second method (1020) allows electronic devices to consider time information when exchanging information in a Graph Neural Network (GNN). In the present disclosure, GNN is a graph neural network that can be used for training an AI model. The electronic device (10) can divide the entire time into N groups. When training the GNN, the electronic device (10) can reflect time information in product information. For example, when the electronic device identifies a product combination, information can be exchanged between nodes representing each product by reflecting the timing of product viewing or the time spent on the page to view the product.
[0176] In one embodiment of the present disclosure, the electronic device (10) may perform learning using the two methods described above by reflecting time information, but is not limited thereto. Based on the learning results, the electronic device (10) may identify a product combination or a discount rate.
[0177] FIG. 11 is a diagram illustrating an algorithm for providing a promotion by an electronic device according to one embodiment of the present disclosure.
[0178] Referring to Figure 11, the entire process of identifying a user's purchasing intent, product combinations, and discount rates is illustrated. For the sake of brevity, any details that overlap with the above description will be omitted.
[0179] In one embodiment of the present disclosure, the electronic device (10) may obtain data (1105) from an online service or intranet with which the user interacts. The data may include, but is not limited to, information about the user, products, marketing, or time.
[0180] The electronic device (10) can perform preprocessing (1110) on the acquired data to make it suitable for an AI model. Preprocessing data may involve processing unstructured information and converting it into numerical information. The electronic device (10) can convert data into a form understandable to the AI model or improve the quality of the data to enable effective learning through high-quality learning data.
[0181] The electronic device (10) can use the acquired data or the preprocessed data as input values of the AI model. The electronic device (10) can use the user purchase intention inference model ( )(1115) can identify the user's purchase intent. If the user purchase intent inference model (1115) identifies that the user has a purchase intent, it outputs '1' (or '0') as the output value. If the user has no intention of purchasing, the output value is '0' (or '1'). It can be output as a user purchase intention inference model. (1115) can infer the user's purchasing intention until a certain number of input and output values are obtained.
[0182] Electronic device (10) adds time information to the user's purchase intention and information about the product Status information via (1125) (1120) can be updated. For data that adds time information to product information, refer to the description of FIG. 12 above. In one embodiment of the present disclosure, the electronic device (10) is a product combination consisting of K products. ) status information 1120) can be updated recursively until the product combination is identified. Status information (1120) is a shared layer shared by models that infer product combinations and discount rates. (1130) can be transmitted. In one embodiment of the present disclosure, the shared layer (1130) may mean a hidden layer. The electronic device (10) is a shared layer (1130) can be used to implement a neural network by sequentially connecting the input and output layers of an AI model. Since inferences about product combinations and discount rates are performed through the same AI model, a shared layer (1130) can share the result value of each operation. Shared layer (1130) When inferring product combinations and discount rates, inference can be performed by interacting with the weights of a deep neural network. Shared layer (1130) A model for inferring product combinations using information (1135) can be used as an input value.
[0183] In one embodiment of the present disclosure, a model for inferring product combinations (1135) identifies the K+1th product to be added to the product mix and updates the product mix ( )(1140) can be identified. Updated product combination( )(1140) is status information (1120) is transmitted, and the product combination can be inferred recursively. In addition, the updated product combination ( )(1140), if the user suitability is higher than the preset value, the electronic device (10) is used to combine the corresponding products ( )(1140) can be inferred as a discount rate.
[0184] In one embodiment of the present disclosure, a model for inferring a discount rate (1145) can infer a discount rate within a valid integer range for a product combination. A model for inferring discount rates. (1145) can identify a valid range of discount rates by considering information about the product (e.g., price, margin, whether it is a flagship product, etc.), user preferences, and restrictions on campaigns or promotions. A model for inferring discount rates. (1145) can identify the logit value 't' (1150), which is the confidence level of decision-making derived from the discount rate and discount rate inference.
[0185] In one embodiment of the present disclosure, an electronic device (10) may provide a promotion (1155) including a product combination and a discount rate. The electronic device (1155) may provide information on the same product combination and discount rate to multiple users, but may change the display order of the product combination and discount rate according to the user's preference.
[0186] In one embodiment of the present disclosure, the electronic device (10) may calculate a reward value (1160) based on the user's feedback regarding the provided promotion (1155). The electronic device (10) may calculate different reward values depending on whether the user has searched for the corresponding product combination through the promotion (1155) or purchased the corresponding product combination.
[0187] In one embodiment of the present disclosure, the electronic device (10) has status information (1120), product combination ( )(1140) and the experience information (1165) can be obtained based on the discount rate for the product combination and the logit value 't' (1150) for the discount rate and the reward value (1160) according to the user's feedback. The electronic device (10) can obtain the experience information (1165) based on the user's purchase intention inference model. (1115), shared layer (1130), a model for inferring product combinations (1135), a model for inferring discount rates (1145) can be applied to perform reinforcement learning.
[0188] FIG. 12 is a schematic block diagram of an electronic device according to one embodiment of the present disclosure.
[0189] Referring to FIG. 12, an electronic device (10) according to an embodiment of the present disclosure may include a communication unit (1210) for performing communication with an external device (not shown), at least one processor (1220) for performing at least one instruction, and a memory (1230) for storing at least one instruction. However, not all of the illustrated components are essential components. The electronic device (10) may be implemented with more components than the illustrated components, or may be implemented with fewer components.
[0190] The communication unit (1210) can communicate with external devices (not shown) via a wired or wireless network. Here, the external devices (not shown) are devices capable of transmitting and receiving content using channels, and may include broadcasting station servers, content storage devices, display devices, etc.
[0191] A communication unit (1210) according to one embodiment of the present disclosure includes at least one communication module, such as a short-range communication module, a wired communication module, a mobile communication module, a broadcast reception module, etc. Here, at least one communication module means a communication module capable of transmitting and receiving data through a network that follows a communication standard, such as a tuner that performs broadcast reception, Bluetooth, WLAN (Wireless LAN) (Wi-Fi), Wibro (Wireless broadband), Wimax (World Interoperability for Microwave Access), CDMA, WCDMA, etc. The communication unit (1210) according to one embodiment of the present disclosure can receive data related to a user, a product, or marketing from an online service or intranet with which a user has interacted.
[0192] The processor (1220) typically controls the overall operation of the electronic device (10). For example, the processor (1220) may use data received through the communication unit (1210) to identify a user's purchasing intent, product combination, the user's suitability for the product combination, a discount rate for the product combination, or priorities for the product combination and discount rate by executing instructions stored in the memory (1230). In one embodiment of the present disclosure, the processor (1220) may perform preprocessing on the acquired data to fit an AI model. The processor (1220) may input data into the AI model and obtain the user's purchasing intent, product combination, and discount rate as output values. The processor (1220) may perform reinforcement learning by reflecting the user's feedback information. In one embodiment of the present disclosure, the processor (1220) may be implemented as a plurality of processors.
[0193] The memory (1230) may store program instructions or codes executed by the processor (1220) and may store input / output data (e.g., data regarding users, products, or marketing, user feedback, etc.). In one embodiment of the present disclosure, the memory (1230) may be implemented as a plurality of memories.
[0194] The memory (1230) may include at least one type of storage medium among a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a magnetic memory, a magnetic disk, and an optical disk.
[0195] In one embodiment of the present disclosure, memory (1220) may store data acquired via the communication unit (1210) or data preprocessed for application to an AI model. The data may include data regarding users, products, and marketing, and may include, but is not limited to, time information acquired when a user interacts with an online service.
[0196] In one embodiment of the present disclosure, the memory (1230) may include a plurality of modules. The purchase intent identification module (1240) may identify the user's purchase intent using acquired data or preprocessed data as input values. The purchase intent identification module (1240) may output '1' (or '0') if it determines that the user has a purchase intent, and '0' (or '1') if it determines that the user does not have a purchase intent.
[0197] The product combination and discount rate identification module (1250) can identify a product combination including multiple products if it is determined that the user has a purchasing intent. The product combination identification module (1250) can identify a product combination including two or more products belonging to different product groups using acquired data, preprocessed data, or user feedback information. In addition, the product combination and discount rate identification module (1250) can identify a discount rate within a valid integer range by considering information about the corresponding products (e.g., product price, margin, flagship product), etc., if the user's suitability for the identified product combination is above a preset value.
[0198] The feedback-based reward value acquisition module (1260) can acquire a reward value based on the feedback obtained by providing the user with a promotion that includes the identified product combination and discount rate. The reward value may have different values depending on whether the user has viewed the product combination through the promotion or made a payment. The product combination and discount rate can be adjusted through reinforcement learning using the feedback-based reward value.
[0199] In one embodiment of the present disclosure, a method for providing a service to a user by an electronic device (10) may include acquiring data related to at least one of a user, multiple products, or marketing, and identifying the user's purchasing intent based on the acquired data. According to one embodiment, the method may identify at least one product combination including two or more products among multiple products and a discount rate for the at least one product combination by applying the identified user's purchasing intent and data to an AI model. According to one embodiment, the method may provide at least one product combination and a discount rate.
[0200] In one embodiment of the present disclosure, the method can identify at least one product combination and discount rate using a hidden layer of an AI model.
[0201] In one embodiment of the present disclosure, the data may include at least one of the user's behavioral data, data related to a web page, data about a viewed product, data about the marketing, or data about time.
[0202] In one embodiment of the present disclosure, the AI model may be a reinforcement learning artificial intelligence model trained to infer at least one product combination and discount rate for a plurality of products.
[0203] In one embodiment of the present disclosure, the method may receive a reward value based on a user's feedback on at least one product combination and a discount rate based on a reward function, and adjust the discount rate based on the received reward value.
[0204] In one embodiment of the present disclosure, the method may identify a product combination based on data including at least one of a user's preference for a product group, a preference for a product group by a path through which the user entered, a preference for a product group by digital marketing provided to the user, or a preference according to the discount rate.
[0205] In one embodiment of the present disclosure, the method can identify a discount rate when a user suitability for at least one product combination is greater than or equal to a preset value.
[0206] In one embodiment of the present disclosure, the method can identify a user's priorities for at least one product combination and discount rate. According to one embodiment, the method can determine the order of placement of at least one product combination and discount rate based on the identified priorities, and provide at least one product combination and discount rate based on the order of placement.
[0207] In one embodiment of the present disclosure, the method can identify at least one product combination by applying a reward value and data based on user feedback to an AI model.
[0208] In one embodiment of the present disclosure, the data may include data relating to at least one of an offline promotion, a marketing campaign, or a channel-specific advertisement.
[0209] In one embodiment of the present disclosure, the discount rate may be identified based on the margin or flagship of multiple products.
[0210] In one embodiment of the present disclosure, the data may include times when a user interacts with multiple products.
[0211] In one embodiment of the present disclosure, the data may include a point in time for identifying a product combination or a point in time for identifying a discount rate, and a weight may be set for each piece of data based on the point in time.
[0212] In one embodiment of the present disclosure, at least one product combination may include two or more products belonging to different product groups among a plurality of products.
[0213] In one embodiment of the present disclosure, an electronic device providing a service to a user may include a communication unit, a memory storing one or more instructions, and at least one processor executing one or more instructions. The at least one processor may obtain data related to at least one of a user, a plurality of products, or marketing, and identify the user's purchasing intent based on the data. According to one embodiment, the electronic device may identify at least one product combination including two or more products among a plurality of products and a discount rate for the at least one product combination by applying the identified user's purchasing intent and data to an AI model. According to one embodiment, the electronic device may provide at least one product combination and a discount rate.
[0214] In one embodiment of the present disclosure, a computer-readable recording medium records a program for executing on a computer to perform the above-described methods.
[0215] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0216] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
Claims
1. In the method of providing a service to a user by an electronic device, Step (S310) of obtaining data related to at least one of the above users, multiple products, or marketing; Step (S320) of identifying the user's purchasing intent based on the above data; Step (S330) of identifying at least one product combination including at least two products among the plurality of products and a discount rate for the at least one product combination by applying the purchase intention of the identified user and the data to the AI model; and A method comprising: a step (S340) of displaying at least one product combination and the discount rate; 2. In paragraph 1, A method wherein the data includes at least one of the user's behavioral data, data related to a web page, data about a viewed product, data about the marketing, or data about time corresponding to the behavioral data.
3. In either paragraph 1 or paragraph 2, A method wherein the above AI model is a reinforcement learning artificial intelligence model trained to infer at least one product combination and discount rate for multiple products.
4. In any one of paragraphs 1 to 3, The step of identifying at least one product combination and the discount rate is: A step of receiving a reward value according to the user's feedback on the at least one product combination and the discount rate based on the reward function; and A method comprising: a step of adjusting the discount rate based on the compensation value; 5. In any one of paragraphs 1 to 4, The step of identifying at least one product combination and the discount rate is: A method for identifying based on data including at least one of the user's preference for a product group, the user's preference for a product group by route through which the user entered, the user's preference for a product group by digital marketing provided to the user, or the user's preference according to the discount rate.
6. In any one of paragraphs 1 to 5, A method for identifying the discount rate when the user suitability for at least one of the above product combinations is greater than or equal to a preset value.
7. In any one of paragraphs 1 to 6, The step of displaying at least one product combination and the discount rate is: A step of identifying the user's priority for at least one product combination and the discount rate; A step of determining the arrangement order of the at least one product combination and the discount rate based on the above priorities; and A method further comprising: a step of displaying the at least one product combination and the discount rate based on the arrangement order; 8. In any one of paragraphs 1 to 7, The step of identifying at least one product combination and the discount rate is: A method for identifying at least one product combination by applying a reward value and the data according to the feedback of the user to the AI model.
9. In any one of paragraphs 1 to 8, A method wherein the above data includes a time at which the user interacts with the plurality of products.
10. In any one of paragraphs 1 to 9, The above data includes a first point in time identifying the product combination or a second point in time identifying the discount rate, A method in which a weight is set for each of the data based on the first point in time or the second point in time.
11. For electronic devices that provide services to users, Communication Department (1210); A memory (1230) storing one or more instructions; and At least one processor (1220) for executing one or more of the above instructions, The at least one processor (1220) executes the one or more instructions, Obtain data related to at least one of the above users, multiple products or marketing, Based on the above data, identify the user's purchasing intent, By applying the purchase intention of the identified user and the data to the AI model, at least one product combination including at least two products among the plurality of products and a discount rate for the at least one product combination are identified, An electronic device that displays at least one product combination and the discount rate.
12. In paragraph 11, At least one processor (1220) above, Based on the reward function, a reward value is received according to the user's feedback on the at least one product combination and the discount rate, An electronic device that adjusts the discount rate based on the above compensation value.
13. In either of paragraphs 11 or 12, At least one processor (1220) above, Identifying the user's priorities for at least one of the above product combinations and the above discount rate, Based on the above priorities, the arrangement order of the at least one product combination and the discount rate is determined, An electronic device that displays at least one product combination and the discount rate based on the above arrangement order.
14. In any one of paragraphs 11 to 13, An electronic device wherein the above data includes the time at which the user interacts with the plurality of products.
15. A computer-readable recording medium having recorded thereon a program for executing the method of paragraph 1 on a computer.
Citation Information
Patent Citations
Information processing device, information processing method and information processing program
JP2020064424A
Method for selling bundle discount commodities in the electronic commerce and computer readable record medium on which a program therefor is recorded
KR1020100001212A
Electronically commodity selling system and service method of it
KR1020150108803A
Server providing product sales service and operation method thereof
KR102314730B1
Item information offering method of electronic apparatus
KR102441999B1