Method for automatically generating training email subjects using reinforcement learning
Patent Information
- Application Number
- KR1020250061333
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2045-05-12
Smart Images

Figure 112025052783916-PAT00009_ABST
Abstract
Description
Technology Field
[0001] This specification relates to the field of artificial intelligence (AI) and security technology, and concerns a method for selecting optimal training email topics by applying Deep Reinforcement Learning, taking into account user preferences, timeliness, and social interest. Background Technology
[0003] Existing training email systems rely on static topic selection methods. It is common practice for experts to directly analyze the latest issue emails to select topics, or to use recurring issues such as year-end tax settlements or summer vacations. Additionally, while new topics are sometimes selected based on email subjects that users frequently clicked on in the past, this approach has limitations in responding to changes in user interests or the latest threats.
[0004] Phishing email attacks are evolving rapidly, and it is highly likely that entirely new types of phishing emails will emerge. Existing systems are not updated until experts encounter and analyze the latest phishing email information, which can delay responses to new threats. This structure increases the risk of users being exposed to the latest phishing attacks.
[0005] Existing training email systems primarily performed static recommendations based on historical click-through rates. This makes it difficult to reflect real-time changing user interests or the latest security issues. Consequently, it is difficult to select topics that users are genuinely interested in, and user engagement may decrease as standardized training topics are repeated. The problem to be solved
[0007] The purpose of this specification is to provide an adaptive reinforcement learning-based method for selecting customized email topics by learning the individual preferences and behavioral patterns of users.
[0008] Furthermore, the purpose of this specification is to implement an algorithm for topic selection that dynamically evaluates and reflects the timeliness, importance, and social interest of emails.
[0009] The technical problems that this specification aims to solve are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this specification belongs from the detailed description of the specification below. means of solving the problem
[0011] One aspect of the present specification may include a method for a server to perform reinforcement learning to generate a subject for a training email, comprising the steps of: obtaining a user state through a database, wherein the user state includes characteristic information of the user to whom the training email is delivered, and inputting the state into a reinforcement learning model (Q-Network); and determining a policy of the reinforcement learning model and recommending a subject for the training email.
[0012] Additionally, it may include a step of measuring a reward based on the subject of the training email; wherein the reward is determined in relation to the user's feedback, and the reinforcement learning model is updated based on the reward.
[0013] In addition, the policy of the above reinforcement learning model can be based on an ε-greedy policy.
[0014] In addition, the policy of the above reinforcement learning model can recommend a random topic with a probability of ε and recommend the topic with the highest Q value with a probability of (1 - ε).
[0015] Additionally, the step of updating the reinforcement learning model based on the above reward may include: storing the user's response as an experience in a replay memory based on the above reward; mini-batch sampling the experience using the replay memory; and updating the reinforcement learning model based on the mini-batch sampling.
[0016] Additionally, the method may further include the step of generating content for the training email using the subject of the recommended training email; and the step of generating the training email based on the content.
[0017] In addition, the step of measuring a reward based on the subject of the training email may further include the step of determining the weight of the reward based on user information or environment information.
[0018] Another aspect of the present specification comprises, in a server performing reinforcement learning to generate a subject for a training email, a communication module; a memory; and a processor for functionally controlling the communication module and the memory; wherein the processor obtains a user state through a database, the user state includes characteristic information of the user to whom the training email is delivered, inputs the state to a reinforcement learning model (Q-Network), determines a policy of the reinforcement learning model, and can recommend a subject for the training email. Effects of the invention
[0020] According to the embodiments of the present specification, an adaptive reinforcement learning-based method for selecting customized email topics by learning the individual preferences and behavioral patterns of a user can be provided.
[0021] In addition, the purpose of this specification is to implement an algorithm for topic selection that dynamically evaluates and reflects the timeliness, importance, and social interest of emails.
[0022] The effects obtainable in this specification are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which this specification belongs from the description below. Brief explanation of the drawing
[0024] FIG. 1 is a block diagram for illustrating an electronic device related to the present specification. FIG. 2 is a block diagram of an AI device according to one embodiment of the present specification. FIG. 3 illustrates a training email automatic generation system to which the present specification may be applied. FIG. 4 is an example of a reinforcement learning system to which the present specification may be applied. FIG. 5 illustrates a reinforcement learning method of a reinforcement learning module that can be applied to the present specification. FIG. 6 illustrates a reinforcement learning algorithm to which the present specification may be applied. FIG. 7 illustrates a compensation function to which the present specification can be applied. FIG. 8 illustrates a method for automatically generating training emails that can be applied to the present specification. Specific details for implementing the invention
[0025] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Identical or similar components regardless of drawing symbols are assigned the same reference number, and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably solely for the ease of drafting the specification and do not have distinct meanings or roles in themselves. Furthermore, in describing the embodiments disclosed in this specification, if it is determined that a detailed description of related prior art could obscure the essence of the embodiments disclosed in this specification, such detailed description will be omitted. Additionally, the attached drawings are intended only to facilitate understanding of the embodiments disclosed in this specification; the technical concept disclosed in this specification is not limited by the attached drawings, and it should be understood that they include all modifications, equivalents, and substitutions that fall within the concept and technical scope of this specification.
[0026] Terms including ordinal numbers, such as first, second, etc., may be used to describe various components, but said components are not limited by said terms. These terms are used solely for the purpose of distinguishing one component from another.
[0027] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0028] Singular expressions include plural expressions unless the context clearly indicates otherwise.
[0029] In this application, terms such as “comprising” or “having” are intended to specify the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0030] FIG. 1 is a block diagram for illustrating an electronic device related to the present specification.
[0031] The above electronic device (100) may include a wireless communication unit (110), an input unit (120), a sensing unit (140), an output unit (150), an interface unit (160), a memory (170), a control unit (180), and a power supply unit (190), etc. Since the components illustrated in FIG. 1 are not essential for implementing the electronic device, the electronic device described herein may have more or fewer components than those listed above.
[0032] More specifically, among the above components, the wireless communication unit (110) may include one or more modules that enable wireless communication between the electronic device (100) and a wireless communication system, between the electronic device (100) and another electronic device (100), or between the electronic device (100) and an external server. Additionally, the wireless communication unit (110) may include one or more modules that connect the electronic device (100) to one or more networks.
[0033] This wireless communication unit (110) may include at least one of a broadcast receiving module (111), a mobile communication module (112), a wireless internet module (113), a short-range communication module (114), and a location information module (115).
[0034] The input unit (120) may include a camera (121) or video input unit for inputting a video signal, a microphone (122) or audio input unit for inputting an audio signal, and a user input unit (123, e.g., a touch key, a mechanical key, etc.) for receiving information from a user. Voice data or image data collected from the input unit (120) may be analyzed and processed into a control command by the user.
[0035] The sensing unit (140) may include one or more sensors for sensing at least one of information within the electronic device, information about the surrounding environment surrounding the electronic device, and user information. For example, the sensing unit (140) may include at least one of a proximity sensor (141), an illumination sensor (142), a touch sensor, an acceleration sensor, a magnetic sensor, a gravity sensor (G-sensor), a gyroscope sensor, a motion sensor, an RGB sensor, an infrared sensor (IR sensor: infrared sensor), a fingerprint sensor (finger scan sensor), an ultrasonic sensor, an optical sensor (e.g., see camera (121)), a microphone (see 122), a battery gauge, an environmental sensor (e.g., a barometer, a hygrometer, a thermometer, a radiation detection sensor, a heat detection sensor, a gas detection sensor, etc.), and a chemical sensor (e.g., an electronic nose, a healthcare sensor, a biometric sensor, etc.). Meanwhile, the electronic device disclosed in this specification can utilize information sensed by at least two of these sensors in combination.
[0036] The output unit (150) is intended to generate output related to sight, hearing, or touch, and may include at least one of a display unit (151), an acoustic output unit (152), a haptic module (153), and an optical output unit (154). The display unit (151) may form a layered structure with a touch sensor or be formed integrally to implement a touch screen. Such a touch screen functions as a user input unit (123) that provides an input interface between the electronic device (100) and the user, and at the same time can provide an output interface between the electronic device (100) and the user.
[0037] The interface section (160) serves as a passage for various types of external devices connected to the electronic device (100). This interface section (160) may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. In response to an external device being connected to the interface section (160), the electronic device (100) can perform appropriate control related to the connected external device.
[0038] Additionally, the memory (170) stores data that supports various functions of the electronic device (100). The memory (170) can store a number of application programs (or applications) running on the electronic device (100), data for the operation of the electronic device (100), and commands. At least some of these application programs may be downloaded from an external server via wireless communication. Also, at least some of these application programs may exist on the electronic device (100) from the time of shipment for the basic functions of the electronic device (100) (e.g., phone incoming and outgoing functions, message receiving and outgoing functions). Meanwhile, the application programs may be stored in the memory (170), installed on the electronic device (100), and driven by the control unit (180) to perform the operation (or function) of the electronic device.
[0039] In addition to operations related to the application program, the control unit (180) typically controls the overall operation of the electronic device (100). The control unit (180) can provide or process appropriate information or functions to the user by processing signals, data, information, etc. that are input or output through the components described above, or by running an application program stored in memory (170).
[0040] Additionally, the control unit (180) can control at least some of the components examined together with FIG. 1 in order to run an application program stored in memory (170). Furthermore, the control unit (180) can operate at least two or more of the components included in the electronic device (100) in combination with each other to run the application program.
[0041] The power supply unit (190) receives external power and internal power under the control of the control unit (180) and supplies power to each component included in the electronic device (100). This power supply unit (190) includes a battery, and the battery may be a built-in battery or a replaceable battery.
[0042] At least some of the above components may operate in cooperation with each other to implement the operation, control, or control method of an electronic device according to various embodiments described below. Additionally, the operation, control, or control method of the electronic device may be implemented on the electronic device by running at least one application program stored in the memory (170).
[0043] In this specification, the electronic device (100) may be collectively referred to as a server, and the server may include a cloud server. Additionally, the terminal may include all or part of the configuration of the electronic device (100), and may include a tablet PC.
[0044] FIG. 2 is a block diagram of an AI device according to one embodiment of the present specification.
[0045] The AI device (20) may include an electronic device including an AI module capable of performing AI processing, or a terminal including the AI module. Additionally, the AI device (20) may be configured to be included as at least a part of the configuration of the electronic device (100) shown in FIG. 1 to perform at least a part of the AI processing together.
[0046] The above AI device (20) may include an AI processor (21), memory (25) and / or a communication unit (27).
[0047] The above AI device (20) is a computing device capable of learning a neural network and can be implemented as various electronic devices such as a terminal, desktop PC, laptop PC, tablet PC, etc.
[0048] The AI processor (21) can train a neural network using a program stored in memory (25). In particular, the AI processor (21) can automatically select a subject for a training email based on reinforcement learning and create an artificial intelligence model that can optimize email content by learning user responses.
[0049] For example, the AI processor (21) can automatically select the subject of a training email by analyzing user characteristics, trends by country and generation, and the latest security threat data, and can continuously optimize the most effective subject by learning user responses (views, clicks, etc.) in real time through a reinforcement learning algorithm. In addition, it can automatically generate an email subject and body content that match the subject by utilizing LLM, and provide a message optimized for the user by using a designated prompt.
[0050] User response data (e.g., viewing rate, click rate, response time, etc.) is collected and analyzed by an AI processor (21), and can be reflected in a reinforcement learning algorithm to improve performance when generating the next training email. In particular, the optimal topic and content are automatically learned by calculating a reward based on user response, and the generated email is automatically converted into an EML (Electronic Mail) file and provided in a format optimized for email clients (e.g., Outlook, Gmail, etc.).
[0051] In this process, the effectiveness of the training email is maintained by reflecting the security rules of each email client, and the AI processor (21) can provide an intelligent security training system that can increase the efficiency of the training email system and respond quickly to the latest phishing techniques through this automated process.
[0052] Meanwhile, the AI processor (21) that performs the functions described above may be a general-purpose processor (e.g., CPU), but may be an AI-dedicated processor for artificial intelligence learning (e.g., GPU, graphics processing unit).
[0053] The memory (25) can store various programs and data required for the operation of the AI device (20). The memory (25) can be implemented as non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), or solid-state drive (SDD). The memory (25) is accessed by the AI processor (21), and the AI processor (21) can perform reading / writing / modification / deletion / updating of data. Additionally, the memory (25) can store a neural network model (e.g., a deep learning model) generated through a learning algorithm for data classification / recognition according to one embodiment of the present specification.
[0054] Meanwhile, the AI processor (21) may include a data learning unit that learns a neural network for data classification / recognition. For example, the data learning unit may learn a deep learning model by acquiring training data to be used for learning and applying the acquired training data to a deep learning model.
[0055] The communication unit (27) can transmit the AI processing results by the AI processor (21) to an external electronic device.
[0056] Here, external electronic devices may include other terminals and / or servers.
[0057] Meanwhile, although the AI device (20) illustrated in FIG. 2 is described by functionally separating it into an AI processor (21), memory (25), and communication unit (27), the aforementioned components may be integrated into a single module and referred to as an AI module or an artificial intelligence (AI) model.
[0058] FIG. 3 illustrates a training email automatic generation system to which the present specification may be applied.
[0059] Referring to FIG. 3, the training email automatic generation system (300) in this specification may be implemented as a physical server or a cloud server and may include a data collector, a DB (310), a subject selection module (320), an LLM Agent (330), and an EMI generator (340).
[0060] Data Collector It can collect data in real time from various sources on the internet (e.g., websites, social media, news, blogs, etc.). For example, this module can automatically collect the latest security issues, phishing techniques, and topics of user interest needed to create training emails, and can continuously update the DB (310) with information needed for topic selection and email creation. The collected data can be refined and stored in the DB (310).
[0061] In addition, the data collector can be configured to collect customized data based on users' specific interests or the profiles of training subjects. For example, it can prioritize the collection of information related to the latest financial fraud techniques for users in the financial sector, or prioritize the latest security threats occurring in a specific country for users in that country.
[0062] DB(310)The DB (310) is a database module that stores and manages data collected from a data collector. This DB (310) stores various information for generating training emails in a structured form and can classify and manage the collected data by user characteristics, latest security threats, trends, etc. This may include phishing-related data, user response data (e.g., open rate, click rate), user feedback data, etc. Additionally, the DB (310) can be linked with the topic selection module (320) to provide data necessary for the reinforcement learning-based topic selection process. For example, email topics that have recorded a high click rate in a specific user group are used as training data so that they can be selected preferentially in reinforcement learning, which can contribute to continuously improving the performance of the training system (300).
[0063] Topic selection module (320) The module can perform the role of automatically selecting the most effective training email topics for users based on data provided from the DB (310). This module learns user response data in real time by applying a reinforcement learning algorithm and can dynamically determine customized topics for each user group. The topic selection module (320) can select optimal topics by reflecting user characteristics (e.g., country, occupation, age, etc.), the latest trends, and security threats. Additionally, the topic selection module (320) evaluates the performance of topics based on user responses (e.g., open rate, click rate, etc.), and topics that record high performance receive high rewards in reinforcement learning, increasing the likelihood of being selected preferentially. This module also transmits the selected topics to the LLM Agent (330) to ensure smooth email creation.
[0064] LLM Agent(330)It can generate training email content that matches the topic received from the topic selection module by utilizing large-scale language models such as Llama, GPT, and Deep Seek. This module can automatically generate an email subject and body optimized for the user through a specified prompt, and can provide customized messages that reflect the user's characteristics and current issues. The LLM Agent (330) can generate not only the text of the training email but also images, links, interactive elements, etc., and can generate content optimized so that it can be used smoothly in various email clients (e.g., Outlook, Gmail, etc.). The generated email content is delivered to the EML generator (340).
[0065] EML generator (340) The module is a module that converts email content generated by the LLM Agent (330) into an EML file format to generate a final training email. An EML file is a standard format that can be used in email clients and can include an email subject, body, images, links, etc. Additionally, the EML generator (340) can also include links or images that can track user reactions to enable the measurement of training performance. The generated EML file is delivered to the user, and the performance of the system (300) can be continuously improved based on customer feedback.
[0066] FIG. 4 is an example of a reinforcement learning system to which the present specification may be applied.
[0067] Reinforcement learning is a learning method that enables an agent to learn optimal actions by interacting with its environment. The agent selects an action in a specific situation (state) and receives a reward as a result. To maximize this reward, the agent learns through experience and can acquire the optimal course of action over time.
[0068] Referring to FIG. 4, the topic selection module (320) may include a preprocessing module for reinforcement learning and a reinforcement learning module.
[0069] Preprocessing module The preprocessing module can perform the role of converting data stored in the DB (310) into a form that can be learned by a reinforcement learning model. The preprocessing module can extract features such as user characteristics (job, years of experience, etc.), click time information, and clicked topics from user click data, and can identify keywords related to the latest issues from social trend data and convert them into topic vectors.
[0070] In addition, the preprocessing module performs data quality management, such as data cleaning, duplicate removal, and deletion of unnecessary information. For example, incorrect click records or duplicate data in user click data are filtered out to prevent interference with learning. The preprocessed data is converted into states, actions, and rewards that can be used by the reinforcement learning module.
[0071] Reinforcement learning module It can automatically generate optimal training email topics based on user behavior patterns and the latest trend data. This module uses a reinforcement learning algorithm (DQN, Deep Q-Network) to recommend training topics and calculates the Q-value (predicted click-through rate) for each topic to present the optimal topic to the user.
[0072] For example, a reinforcement learning module can select a topic (Action) based on an initial state (State), evaluate the reward based on whether the user clicks (Reward), and predict the optimal topic through a Q-network. Subsequently, the reinforcement learning module can update the experience memory (Replay Memory) using user click data and optimize the performance of the AI model through learning.
[0073] The reinforcement learning module learns in real-time and can continuously improve the performance of recommendation topics based on user feedback. This allows for the immediate reflection of the latest phishing attacks or user interests.
[0074] FIG. 5 illustrates a reinforcement learning method of a reinforcement learning module to which the present specification can be applied.
[0075] A reinforcement learning environment can refer to a space where agents can perform learning through interaction.
[0076] Referring to Fig. 5, the reinforcement learning module can configure a reinforcement learning environment to perform optimal topic recommendations that reflect user behavior patterns and the latest trends.
[0077] For example, a reinforcement learning environment can be configured as follows:
[0078] 1. State
[0079] User characteristics: Reflects individual user characteristics such as job function, years of experience, region, and gender.
[0080] Issue Characteristics: Reflects the results of topic clustering regarding currently popular issues, the latest phishing topics, and social trends.
[0081] Temporal context: Year, season, period of specific events (e.g., year-end tax settlement, summer vacation, etc.)
[0082] Social Trends: Latest news, community issues, and trend data collected from social media
[0083] Diversity Indicators: Duplication rate of recently recommended topics, prevention of repeated topic recommendations
[0084] 2. Action
[0085] Topic Selection: Select a topic for training emails, and topic recommendations reflecting user characteristics and the latest issues are available.
[0086] Diversity Adjustment: Minimizes repetition of the same topic by controlling the range and variation of recommended topics.
[0087] 3. Reward
[0088] Engagement: Grants a positive reward (+1) when the user clicks or opens an email.
[0089] Interaction: Provides additional rewards when the user clicks a link in the email.
[0090] Diversity: Rewards increase as repetition of the same topic decreases, and weights increase as diversity increases.
[0091] Timeliness: Higher rewards for topics related to recent issues (e.g., year-end tax settlement topics during the year-end tax season)
[0092] 4. Environment
[0093] User behavior: Whether the user clicks on an email, whether the email content is viewed, etc.
[0094] News Trend API: Collects data on the latest news, social media trends, and phishing attacks
[0095] Phishing Campaign Statistics: Latest Phishing Attack Cases and User Response Statistics
[0096] 5. Reinforcement Learning Model
[0097] Q-Network (Deep Q-Network, DQN): Calculates the predicted probability (Q-value) of each topic (behavior) being clicked based on the user state.
[0098] Policy: Based on ε-greedy, topic selection is performed according to the Q value, and the ε value balances exploration (random topic recommendation) and exploitation (optimal topic recommendation).
[0099] The reinforcement learning module can automatically generate topics through the configured reinforcement learning environment.
[0100] As mentioned above, the reinforcement learning environment can be designed to recommend topics optimized for the user by reflecting user behavior patterns and the latest trends in real time. For example, the reinforcement learning module can predict the click probability for each topic in a Q-Network based on the user state and recommend topics to the user according to a policy.
[0101] The reinforcement learning module can continuously improve topic recommendation performance by comprehensively analyzing various data, such as user characteristics, temporal context, social trends, and the latest phishing topics. By observing the initial state, it identifies user characteristics (job function, years of experience, region, etc.) and current issues, and inputs this into a Q-Network to calculate Q-values for each topic (behavior). For example, the Q-values of related topics may increase during the year-end tax settlement season, which enables the automatic recommendation of topics that users are more likely to click on frequently.
[0102] In addition, through the ε-greedy policy, the reinforcement learning module can maintain a balance between exploration (random topic recommendation) and exploitation (optimal topic recommendation). The ε value is set high during the initial stages of the reinforcement learning module's training to explore various topics, and as training progresses, it gradually decreases to focus recommendations on optimal topics. This enables the reinforcement learning module to learn increasingly efficiently through user response data.
[0103] User actions (e.g., clicking, viewing, link clicking) are stored as feedback for the reinforcement learning module to learn from, and data stored in the Replay Memory can be utilized for Q-Network training through mini-batch sampling. This enables the reinforcement learning module to continuously improve training performance by iteratively learning user responses. The Q-Network can predict optimal topics by learning user behavior patterns and maintain learning stability through the Target Network.
[0104] FIG. 6 illustrates a reinforcement learning algorithm to which the present specification may be applied.
[0105] Referring to Fig. 6, the reinforcement learning algorithm of the reinforcement learning module for automatically generating user-customized training email topics in Fig. 5 described above is illustrated in more detail.
[0106] The reinforcement learning module observes the user's initial state (S6010).
[0107] The reinforcement learning module observes the user's state through the database. The system defines the current state based on user characteristics, time information, social trends, issue characteristics, and diversity indicators. User characteristics may include job function, years of experience, region, and gender. Time information can reflect how user interests may vary over time based on the year, season, or specific event periods (e.g., year-end tax settlement, summer vacation). Additionally, social trends can reflect user interests in real time by including the latest news, community issues, and trend data collected from social media. Issue characteristics refer to the clustering results of recent phishing topics or social issues, enabling users to learn about the latest threats. Diversity indicators can be used to prevent repetitive topic recommendations by evaluating the duplication rate of recently recommended topics. Through this process, the reinforcement learning module can identify the user and the current environment, and the observed state information can be passed to the reinforcement learning model (Q-Network).
[0108] The reinforcement learning module inputs the state into the reinforcement learning model (Q-Network) and calculates the Q value (S6020).
[0109] The reinforcement learning module inputs the observed state vector into the Q-Network. The Q-Network is a deep neural network capable of calculating a Q-value for each state based on topic selection (action). For example, the Q-value quantifies the probability that each topic will be clicked by the user, enabling the system to recommend the optimal topic.
[0110] If the state is "Year-end tax settlement season, Job: Office worker", the Q-Network can calculate the following Q value.
[0111] Q(s, a1) = 0.77 (Phishing topic related to year-end tax settlement)
[0112] Q(s, a2) = 0.1 (Hacking email alert topic)
[0113] In this context, the Q-value reflects the user's interest or click probability, and a higher Q-value may indicate that the topic is more attractive to the user. Through training, the Q-Network can predict these Q-values with increasing accuracy and continuously improve its performance using user click data. For example, training of the Q-Network can be performed based on a loss function, and as the loss value decreases, the prediction of the Q-value becomes more accurate.
[0114] The reinforcement learning module determines the policy of the reinforcement learning model and recommends a topic (S6030).
[0115] The reinforcement learning module recommends training email topics for the user to view based on a policy. For example, the reinforcement learning module uses an ε-greedy policy, which can be designed to maintain a balance between exploration and exploitation. In an ε-greedy policy, the ε value can adjust the probability between random topic recommendations (exploration) and the topic recommendation with the highest Q value (exploitation).
[0116] Random topic recommendation (search) with probability ε
[0117] Recommendation of the topic with the highest Q value with a probability of (1 - ε) (Application)
[0118] For example, when ε = 0.5, the reinforcement learning module can recommend a random topic with a 50% probability and the topic with the highest Q-value with the remaining 50% probability. This policy allows the reinforcement learning module to explore various topics during the initial stages of training and enables recommendations to increasingly center on the optimal topic as training progresses. Through exploration, the reinforcement learning module can test new topics and, through utilization, recommend topics with good user response more frequently. This maximizes the performance of personalized topic recommendations.
[0119] The reinforcement learning module measures the reward based on the topic (S6040).
[0120] Recommended topics are provided to users as training emails via the LLM Agent, and the reinforcement learning module can measure rewards based on user responses. For example, a positive reward (+1) is granted if the user clicks the email or opens a link, while a reward of 0 is applied if they do not click. This reward can be used as an indicator to evaluate whether the system accurately predicted the user's interest. The reward can comprehensively reflect various factors, not just user clicks.
[0121] for example,
[0122] Engagement: +1 when the user clicks or opens the email
[0123] Diversity: Rewards increase as repetition of the same topic decreases.
[0124] Timeliness: Higher rewards for newer topics (e.g., year-end tax settlement topics during the year-end tax season)
[0125] This reward system enables the topic selection module to recommend topics that reflect user interests more frequently and can continuously optimize the performance of the reinforcement learning model.
[0126] FIG. 7 illustrates a compensation function to which the present specification can be applied.
[0127] Referring to Fig. 7, the reward function of a reinforcement learning-based training email topic automatic generation system can be structured to grant rewards by comprehensively evaluating user behavior and environmental factors. For example, the reward value can be dynamically determined by reflecting the user's behavior patterns, recent issues, regular issues, and personal characteristics of the user.
[0128] The reinforcement learning module determines reward weights based on user information and / or environment information (S7010).
[0129] The reinforcement learning module can determine initial reward weights by verifying user and environmental information through the database. For example, environmental information includes personal user characteristics such as date, user location, and user job function. This information provides important foundational data for topic recommendations and can be utilized as basic data to provide customized topics to the user.
[0130] Through date information, reinforcement learning modules can identify the current season, year, and specific event periods (e.g., year-end tax settlement, summer vacation), which can act as key variables that influence user interests. User information reflects personal characteristics such as job function, years of experience, region, and gender, which can be utilized to understand user behavior patterns and recommend personalized topics.
[0131] For example, if the user is in the tax season, the reinforcement learning module can set the reward weights for tax-related topics higher. On the other hand, if it is the summer vacation season, travel or security-related topics may receive higher rewards.
[0132] The reinforcement learning module modifies the reward based on periodic issues and / or recent issues (S7020).
[0133] The reinforcement learning module can increase or decrease rewards based on periodic and recent issues via the database. For example, periodic issues refer to those that occur periodically by country or region, and may include year-end tax settlements, vacations, paydays, and security checks. These periodic issues are important factors for easily predicting user interest, and the reinforcement learning module can automatically detect them and increase the reward values for related topics.
[0134] Latest issues include real-time news, social media trends, and recent phishing cases, and the reinforcement learning module can adjust reward values by evaluating whether these issues are current trends. For example, if recent phishing attacks are occurring in a specific manner, the reward value for that topic may increase.
[0135] Through this, the reinforcement learning module can improve the accuracy of topic recommendations by reflecting the likelihood that users will be interested in the latest trends.
[0136] The reinforcement learning module modifies the reward based on the user's personal issues (S7030).
[0137] The reinforcement learning module can weight or subtract reward values based on the user's personal issues via the database. These personal issues include the user's job function, years of experience, welfare-related information, and history of security breaches, which can be used to recommend topics tailored to the user's characteristics. For example, if a user works in the financial industry, reward values for topics related to financial fraud or security may be set higher.
[0138] In addition, if a user has a history of security breaches, rewards for security awareness training topics may increase, and if the user's job is IT-related, reward values for topics related to the latest technology trends or cybersecurity may be higher. These personalized rewards enable the system to recommend customized topics more accurately.
[0139] The reinforcement learning module modifies the reward based on the user's click pattern and / or community trends (S7040).
[0140] The reinforcement learning module can ultimately adjust reward values based on user click patterns and community trends through the database. For example, the module can analyze topics frequently clicked by users in the past and increase the reward values for topics with high click-through rates. This can improve the accuracy of personalized topic recommendations by reflecting user interests.
[0141] Furthermore, by reflecting community trends, the reward values for topics mentioned in the latest social media issues or news can be additionally weighted. For example, if a security issue has recently become a hot topic on social media, the reward value for that topic may increase. This allows for providing users with topics that reflect the latest trends and enables them to quickly acquire the latest information.
[0142] Again, referring to FIG. 6, the reinforcement learning module stores the experience based on the reward and samples a mini-batch (S6050).
[0143] The reinforcement learning module can store user responses (Transition: s, a, s', r) in the experience memory (Replay Memory). The experience memory is a storage that allows the system to store and learn from past user response data. For example, topics clicked and unclicked by the user, and the click-through rate for each topic can be stored.
[0144] The main advantage of experience memory is that it allows for random sampling regardless of chronological order. N random experiences can be sampled at regular intervals through mini-batch sampling and utilized for Q-Network training. For example, 66 experiences out of 1,000 past experiences can be randomly sampled for training. This mini-batch training ensures the diversity of the training data and prevents the system from becoming dependent on specific users or topics. Furthermore, experience memory-based training can enhance the stability of Q-Network training.
[0145] The reinforcement learning module updates the reinforcement learning model (S6060).
[0146] The reinforcement learning module can update the Q-Network based on learned experience data through mini-batch sampling. For example, it can estimate the expected future reward by calculating the TD (Target) value for each sample and adjust the Q-Network weights by minimizing the loss function.
[0147] Here, the TD(Target) calculation can be defined as follows:
[0148]
[0149] Here, γ (discount rate) determines the importance of future rewards, and at regular intervals, the target network ( _target) to Q-Network( Learning stability can be guaranteed by synchronizing with ). The ε value can also be gradually reduced, decreasing the probability of random recommendations and increasing the probability of optimal topic recommendations. Through iterative learning, the reinforcement learning module enables the system to continuously improve user-customized topic recommendation performance.
[0150] FIG. 8 illustrates a method for automatically generating training emails that can be applied to the present specification.
[0151] Referring to FIG. 8, the system (300) can automatically generate a training email by applying the aforementioned reinforcement learning and deliver it to the user through a terminal.
[0152] The system receives a command from the terminal for a request to create a training email (S8010).
[0153] Customers transmit a request to the system to generate training emails via their terminals, and the system receives customer data along with the request. For example, customer data may include user characteristics (occupation, age, country) and past training performance (open rate, click-through rate). This data provides basic information for generating training emails and lays the foundation for providing customized training emails to user groups. The more abundant and specific the customer data, the higher the accuracy of topic selection. For instance, the latest topics for financial fraud prevention can be selected based on data from a financial industry user group, whereas general security training topics may be selected if data is insufficient. If necessary, the system can obtain additional data through the S8020.
[0154] The system collects data to generate training emails through a data collector and stores it in a DB (S8020).
[0155] If necessary, the data collector can gather data such as the latest security issues, phishing techniques, and topics of user interest required for generating training emails from various sources (e.g., websites, social media, news, blogs, etc.) via the Internet. This data can be updated regularly. The collected data can be refined so that it can be used for topic selection. Subsequently, the collected data is stored in a database.
[0156] For example, when generating training emails for financial industry professionals, the data collector can automatically gather the latest financial fraud cases, cloud security issues, phishing email templates, and more. This data lays the foundation for effectively selecting topics for the training emails to be provided to users.
[0157] The system selects the subject of the training email through the subject selection module (S8030).
[0158] The topic selection module selects the optimal training email topic by applying a reinforcement learning algorithm based on the collected data.
[0159] The system post-processes the selected topic (S8040).
[0160] The system may further include a topic post-processing module and an image update module. The topic post-processing module can determine whether additional HTML formatting is required for the selected topic or if images need to be included. Additionally, if HTML analysis is necessary, it can automatically generate an HTML structure or modify existing HTML templates in a manner similar to an LLM Agent. If images are required, the image update module can generate images suitable for the topic or retrieve them from the database to add them.
[0161] For example, for the topic of financial fraud prevention, you can include financial company logos in HTML format and add images explaining recent financial fraud cases. These HTML and image settings are configured to maximize user response.
[0162] The system generates the content of a training email through an LLM Agent based on the subject (S8050).
[0163] The LLM Agent can generate training email content based on selected topics and HTML and image information delivered from the topic post-processing module. Using large-scale language models such as GPT, Llama, and Deep Seek, the LLM Agent automatically generates email subject lines and bodies, and can create customized content that reflects user characteristics and the latest security threats.
[0164] Thus, the structure of separating topic generation from content generation and applying reinforcement learning solely to topic selection was designed due to the inherent limitations of reinforcement learning. Reinforcement learning is an algorithm that learns dynamically within a State, Action, and Reward structure; actions selected within a specific state are evaluated for performance through rewards, and this evaluation is optimized through learning. However, while content generation relies heavily on text context, grammar, and the creativity of the message to be conveyed to the user, reinforcement learning struggles to effectively learn this complex language generation process. In particular, text generation allows for an infinite number of expressions, and user responses (views, clicks, etc.) cannot clearly evaluate performance at the sentence level, making them unsuitable for reinforcement learning training. On the other hand, topic selection can be directly evaluated using quantitative performance metrics (click-through rates, view rates), and topic performance can be optimized through reinforcement learning tailored to specific user groups. Therefore, this system adopts a structure where reinforcement learning is applied only to topic selection, while content generation is performed via an LLM Agent using fixed prompts, thereby maximizing the efficiency of reinforcement learning and maintaining the stability of content generation.
[0165] Table 1 below is an example of a prompt for an LLM Agent to generate the content of a training email.
[0166] Referring to Table 1, the prompt can be configured to automatically generate email subjects and contents in JSON format.
[0167] Please select an email subject that people related to the following reference data would click on and an email content corresponding to that subject. The subject and the email content corresponding to the subject should be written in Korean, and the data format should be in the following json format. { "subject" : "", "contents" : ""} Reference data:
[0168] Table 2 below shows examples of data returned when the LLM Agent generates the content of a training email.
[0169] Referring to Table 2, the returned data includes an email subject and contents in JSON format and can be dynamically generated based on user characteristics or topics.
[0170] { "subject": "Guidance on Workforce Restructuring and Future Direction Following AI Adoption", "contents": "Hello, employees. First, I am heavy-hearted to share the significant changes currently taking place at our company. Kakao Corporation has recently implemented an organizational restructuring aimed at enhancing AI-based services and maximizing efficiency. In this process, new hires have been temporarily suspended, and we have initiated voluntary resignation procedures for approximately 10% of our total workforce. This decision was by no means made lightly; it is the result of intense internal discussions. We ask for your understanding that this is an unavoidable measure to respond to the rapidly changing market and technological environment, and to secure the company's sustainability and competitiveness. We plan to provide separate announcements regarding the details and support measures related to these actions, and all procedures will be conducted as transparently and fairly as possible. Furthermore, we plan to provide practical assistance to those affected by establishing various measures, such as career transition support, psychological counseling, and re-employment programs. The transition to AI and efficiency is a trend of a new era that we must prepare for together. As such, cooperation and trust with the remaining members have become even more important. We [will address] this change We believe that this will serve as an opportunity for us to be reborn as a stronger and more flexible organization. We anticipate that you may have many concerns and questions, so please feel free to contact the HR team or reply to this email at any time, and we will respond faithfully. Thank you. - The Management Team -
[0171] The system generates and distributes training emails based on the content through an EML generator (S8060). Content generated by the LLM Agent can be converted into an EML file format by the EML generator.
[0172] FIG. 9 is an example of a training email to which the present specification may be applied.
[0173] Referring again to FIG. 8, the EML file is a standard format usable by email clients (Outlook, Gmail, etc.) and can include an email subject, body, images, links, etc. The HTML format is automatically optimized for each client. Subsequently, the generated EML file can be sent to the training target customers through an email distribution service.
[0174] The system receives feedback corresponding to a training email from the terminal and performs reinforcement learning (S8070).
[0175] The system collects responses to training emails from users (e.g., views, clicks, etc.) and analyzes training result data received from customers.
[0176] Here, the collected feedback may be as follows.
[0177] 1. Customer Information: Company name, job category, average years of experience, etc.
[0178] 2. Clicked customer data: Annual leave [optional], Email viewing time [optional]
[0179] 3. Click-through Rate: Click-through rate of training conducted by the client [optional], percentage of click-through customers by years of experience (10 years n%, 5 years m%...) [optional]
[0180] After being refined, this feedback data can be passed to the topic selection module and utilized for learning.
[0181] For example, the topic selection module can evaluate the performance of each topic based on feedback and update the Q-value. Topics with high Q-values are preferentially selected in subsequent training, while low-performance topics are gradually optimized so that they are not selected.
[0182] Furthermore, this specification can significantly reduce maintenance costs for the training system by minimizing user intervention. Through automated topic selection, email generation, user response analysis, and performance learning processes, the system can autonomously generate and improve optimal training emails without the need for an administrator, which represents a clear structural and functional difference compared to existing manual-based training systems. In particular, by automatically providing user-customized training through reinforcement learning, it is possible to realize real-time capabilities, automation, and individual optimization that were not achievable with conventional technology.
[0183] The foregoing specification may be implemented as computer-readable code on a medium on which a program is recorded. A computer-readable medium includes all types of recording devices in which data that can be read by a computer system is stored. Examples of computer-readable media include Hard Disk Drives (HDDs), Solid State Disks (SSDs), Silicon Disk Drives (SDDs), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, optical data storage devices, etc., and also include implementations in the form of carrier waves (e.g., transmission over the Internet). Accordingly, the above detailed description should not be interpreted restrictively in all respects and should be considered exemplary. The scope of this specification should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of this specification are included within the scope of this specification.
[0184] Furthermore, although the above description has focused on the services and embodiments, this is merely illustrative and does not limit the scope of this specification. Those skilled in the art will understand that various modifications and applications not exemplified above are possible without departing from the essential characteristics of the services and embodiments. For example, each component specifically shown in the embodiments may be modified and implemented. Differences related to such modifications and applications should be interpreted as being included within the scope of this specification as defined in the appended claims.
Claims
Claim 1 A method for performing reinforcement learning to generate a topic for a personalized training email for responding to security threats, comprising: a step of obtaining a user state through a database, wherein the user state includes characteristic information of the user to whom the training email is delivered, wherein the characteristic information includes the user's job, years of experience, region, or gender; a step of inputting the state into a reinforcement learning model (Q-Network); and a step of determining a policy of the reinforcement learning model and recommending a topic for the training email; a step of measuring a reward based on the topic for the training email, wherein the reward is determined by data related to the user's response, and a step of updating the reinforcement learning model based on the reward; a step of generating content for the training email using the recommended topic for the training email based on the updated reinforcement learning model; and a step of generating the training email based on the content. Claim 2 delete Claim 3 In claim 1, the method of execution is based on the ε-greedy policy of the policy of the reinforcement learning model. Claim 4 In paragraph 3, the policy of the reinforcement learning model is to recommend a random topic with a probability of ε and to recommend the topic with the highest Q value with a probability of (1 - ε). Claim 5 A method for performing, wherein, in claim 1, the step of updating the reinforcement learning model based on the reward comprises: the step of storing the user's response as an experience in a replay memory based on the reward; the step of mini-batch sampling the experience using the replay memory; and the step of updating the reinforcement learning model based on the mini-batch sampling. Claim 6 delete Claim 7 A method of execution according to claim 1, wherein the step of measuring a reward based on the subject of the training email further comprises the step of determining the weight of the reward based on user information or environment information. Claim 8 A server performing reinforcement learning to generate a topic for a personalized training email for responding to security threats comprises: a communication module; a memory; and a processor for functionally controlling the communication module and the memory; wherein the processor obtains a user state through a database, the user state includes characteristic information of the user to whom the training email is delivered, the characteristic information includes the user's job, years of service, region, or gender; inputs the state into a reinforcement learning model (Q-Network), determines a policy of the reinforcement learning model, recommends a topic for the training email, measures a reward based on the topic for the training email, wherein the reward is determined by data related to the user's response, updates the reinforcement learning model based on the reward, generates content for the training email using the recommended topic for the training email based on the updated reinforcement learning model, and generates the training email based on the content.
Citation Information
Patent Citations
Training neural networks using prioritized experience memory
KR102191444B1
Method and Platform System for Simulation Training of Malicious Mail
KR102327773B1
Method and system for automatically generating trend-compliant webzines based on artificial intelligence
KR102759025B1