Mail reply content generation method based on knowledge base and computer program product
By using a knowledge-based email reply method, which generates email replies using similarity thresholds and generalization ranges, the problems of low email reply efficiency and inaccurate information disclosure in existing technologies are solved, and efficient and personalized email replies and information control are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI SHENGYUN INTELLIGENT TECH CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-08
AI Technical Summary
Existing email reply methods are inefficient, lack personalization, and cannot effectively control the extent of information disclosure. The replies generated by existing intelligent reply systems are too generic and cannot adapt to diverse email content.
By acquiring data on emails to be replied to, determining similarity thresholds and generalization ranges, and using AI to generate reply emails that conform to the user's knowledge base, the degree of information disclosure is controlled, and the relevance and personalization of the reply content are enhanced.
It enables the automatic generation of efficient and personalized email replies, while taking into account information privacy controls, avoiding empty content, and improving the relevance of email replies and the accuracy of information disclosure.
Smart Images

Figure CN121998604A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a knowledge base-based method for generating email reply content and a computer program product. Background Technology
[0002] This section is intended to provide background or context for the embodiments disclosed herein. The description herein is not intended to imply that it is prior art simply because it is included in this section.
[0003] With the rapid growth of email communication, efficient email processing has become crucial for improving work efficiency. Currently, the main methods for replying to emails include: manual word-by-word replies, which are inefficient and highly dependent on human intervention; fixed template replies, which lack personalization and targeting; and rule-based automated reply systems, which are inflexible and unable to adapt to diverse email content.
[0004] In recent years, artificial intelligence (AI) technology has been applied to automated email replies. For example, some intelligent reply systems generate reply suggestions using natural language processing, and platforms such as Alibaba Cloud also provide similar automated services. However, these existing technologies have significant drawbacks: First, the generated replies are often generic and difficult to tailor to specific questions; second, the systems cannot intelligently control the degree of information disclosure, which may lead to over-disclosure of confidential information or under-disclosure in business communications, affecting communication efficiency. Summary of the Invention
[0005] Therefore, it is necessary to provide a knowledge-based email reply content generation method and computer program product that can automatically and effectively reply to emails while maintaining the accuracy of information control, in order to address the aforementioned technical problems.
[0006] Firstly, this disclosure provides a method for generating email reply content based on a knowledge base. The method includes:
[0007] Retrieve data of emails awaiting reply;
[0008] Based on the email data to be replied to, a first threshold and a generalization range are determined; the first threshold is the similarity threshold between the data in the user's knowledge base and the email data to be replied to; the generalization is the ratio of the number of words in the generalized text to the number of words in the original text.
[0009] Calculate the similarity between the data in the user knowledge base and the data of the email to be replied to, and obtain the matching similarity.
[0010] Filter data from the user knowledge base whose similarity exceeds the first threshold to obtain target data;
[0011] The AI is invoked to generate a reply email that conforms to the stated generalization range based on the target data and the email data to be replied to.
[0012] Optionally, the first threshold and the generalization range are obtained by the following method:
[0013] Obtain background data; the background data includes: relationship data between the party to be responded to and the user, identity data of the party to be responded to, and identity data of the user;
[0014] The AI is invoked to determine the first threshold and the generalization range based on the email data to be replied to and the background data.
[0015] Optionally, the user knowledge base includes an enterprise knowledge base and / or a personal knowledge base; the method further includes:
[0016] The first threshold and generalization range determined by AI are sent to the user, and subsequent steps are carried out after the user modifies or confirms them.
[0017] Optionally, the first threshold and the generalization range are obtained by the following method:
[0018] The user determines the first threshold and the generalization range based on the email data to be replied to.
[0019] Optionally, the method further includes:
[0020] At least two privacy levels are preset, and each privacy level corresponds to a set first threshold and a range of generalization.
[0021] Users determine the first threshold and the generalization range by selecting a privacy level.
[0022] Optionally, the privacy level includes a highest privacy level and a lowest privacy level;
[0023] At the highest privacy level, the generalization range is 0, and the reply email does not contain any content related to the target data;
[0024] At the lowest privacy level, the generalization range is 100%, and the reply email contains all the content of the target data.
[0025] Optionally, the method further includes:
[0026] Retrieve a user's sent email history;
[0027] The AI is required to adjust the language style of the reply email to match the style of the historical emails sent.
[0028] Optionally, the method further includes:
[0029] The similarity algorithm is used to calculate the similarity between the data in the user knowledge base and the email data to be replied to; the range of the generalization degree is 0~100%; the range of the first threshold is 0~100%.
[0030] Optionally, the method further includes:
[0031] A first control bar and a second control bar are set in the user interface; the first control bar is used for the user to set the first threshold; the second control bar is used for the user to set the generalization range.
[0032] Secondly, this disclosure also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0033] Retrieve data of emails awaiting reply;
[0034] Based on the email data to be replied to, a first threshold and a generalization range are determined; the first threshold is the similarity threshold between the data in the user's knowledge base and the email data to be replied to; the generalization is the ratio of the number of words in the generalized text to the number of words in the original text.
[0035] Calculate the similarity between the data in the user knowledge base and the data of the email to be replied to, and obtain the matching similarity.
[0036] Filter data from the user knowledge base whose similarity exceeds the first threshold to obtain target data;
[0037] The AI is invoked to generate a reply email that conforms to the stated generalization range based on the target data and the email data to be replied to.
[0038] The aforementioned knowledge base-based email reply content generation method and computer program product, by calling AI to generate reply emails based on user knowledge bases, can enhance the relevance of reply email content and avoid empty content. When calling user knowledge bases, the content similarity threshold, i.e., the first threshold, is used to control the data calling scope of user knowledge bases while ensuring that the called data is relevant to the email to be replied to. The generality parameter is used to limit the specificity of the reply email content, which can achieve the effect of automatically generating effective reply emails while taking into account the control of information privacy. Attached Figure Description
[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0040] Figure 1This is a flowchart illustrating a knowledge base-based email reply content generation method in one embodiment. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this disclosure clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this disclosure.
[0042] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0043] The knowledge base-based email reply content generation method provided in this disclosure can be applied to scenarios where AI is used for automatic email replies. For example, one scenario includes a terminal, a server, and a data storage system. The terminal can communicate with the server via a network. The data storage system can store data that the server needs to process. The data storage system can be integrated into the server or placed on a cloud server or other network server. The terminal, as a data acquisition end, can acquire the email data to be replied to. The server can determine a first threshold and a generalization range based on the email data to be replied to, with or without user participation. The first threshold is a similarity threshold between data in the user's knowledge base and the email data to be replied to. The generalization rate is the ratio of the number of characters in the generalized text to the number of characters in the original text. The server can calculate the similarity between the data in the user's knowledge base and the email data to be replied to, obtaining a matching similarity. Then, the server filters data in the user's knowledge base whose matching similarity exceeds the first threshold to obtain target data. Finally, the server can call AI to generate a reply email that conforms to the generalization range based on the target data and the email data to be replied to. The terminals can be, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. Servers can be implemented using independent servers or server clusters composed of multiple servers.
[0044] In one embodiment, such as Figure 1As shown, a knowledge base-based method for generating email reply content is provided. Taking the application of this method in the aforementioned scenario as an example, the method includes the following steps:
[0045] Step 102: Obtain the data of emails to be replied to.
[0046] Among them, pending email data can refer to data containing the specific content information of the emails to be replied to.
[0047] Specifically, the data for the email to be replied to can include data about the sender and the body of the email. Data about the sender can refer to data containing information about the sender of the email to be replied to. The body of the email can refer to data containing information about the body of the email to be replied to. The data for the email to be replied to can include the sender's information and the content of the email body. The server can obtain the data for the email to be replied to in various ways, such as retrieving the content from the user's mailbox via a data terminal. Alternatively, the user can actively send the data to the server. Or, the user can grant the server permission to access the content of emails in their mailbox, and the server can automatically retrieve the content when it detects a new email in the mailbox.
[0048] Step 104: Based on the email data to be replied to, determine a first threshold and a generalization range; the first threshold is the similarity threshold between the data in the user's knowledge base and the email data to be replied to; the generalization is the ratio of the number of characters in the generalized text to the number of characters in the original text.
[0049] Here, "generalization range" can refer to the range of values that a generalization can take. "User knowledge base" can refer to the user's database that provides content support for automatically generating reply emails.
[0050] Specifically, the generalization range is generally a numerical range, but it can also be a numerical point. The user knowledge base can include at least one of an enterprise knowledge base and a personal knowledge base. The first threshold is used to control the scope of data called from the user knowledge base. When the first threshold is set high, the server calls less data, only calling data in the user knowledge base that has a high similarity (or relevance) to the data in the email to be replied to; when the first threshold is set low, data in the user knowledge base that has a low similarity to the data in the email to be replied to will also be called, and the server calls more data. The generalization is used to control the level of detail of the newly generated text content compared to the content in the user knowledge base when calling data from the user knowledge base to generate new text. Generally, the smaller the generalization value, the more general the newly generated text content; the larger the generalization value, the more specific the newly generated text content. The first threshold and generalization range used in the process of generating the reply email can be determined by the user or AI based on the specific content of the email to be replied to. Step 104 is not necessarily completed before step 106; for example, it can be completed after step 106, or simultaneously with step 106.
[0051] Step 106: Calculate the similarity between the data in the user knowledge base and the data of the email to be replied to, and obtain the matching similarity.
[0052] Specifically, there are many algorithms in the prior art for calculating text similarity, such as cosine similarity algorithm, Euclidean distance algorithm, semantic similarity algorithm based on neural network model, etc. Similarity algorithms can be used to calculate the similarity between various data in the user's knowledge base and the data in the email to be replied to; the calculation result is called the matching similarity. When using the solution in this disclosure, those skilled in the art can choose an appropriate similarity calculation algorithm according to the actual situation; this disclosure does not impose any restrictions.
[0053] Step 108: Filter the data in the user knowledge base whose similarity exceeds the first threshold to obtain the target data.
[0054] Specifically, based on a first threshold, the user knowledge base data to be invoked is determined through filtering using matching similarity. Data in the user knowledge base with a matching similarity exceeding the first threshold, after filtering, is referred to as target data.
[0055] Step 110: Call the AI to generate a reply email that conforms to the generalization range based on the target data and the email data to be replied to.
[0056] Specifically, the target data and the email data to be replied to can be sent together to the AI, requesting the AI to generate a reply email. The content related to the target data in the reply email is a summary of the target data, and the degree of summary of this content conforms to the stated range. It should be noted that the reply email generated by the AI may contain content unrelated to the target data, such as "Hello," "Thank you," "Sincerely," "Welcome," and a review of the content to be replied to. This disclosure does not impose restrictions on this part of the content, and it is unrelated to the degree of summary in this disclosure.
[0057] In the aforementioned knowledge base-based email reply content generation method, calling AI to generate reply emails based on the user's knowledge base can enhance the relevance of the reply email content and avoid empty content. When calling the user's knowledge base, the content similarity threshold, i.e., the first threshold, is used to control the data calling scope of the user's knowledge base while ensuring that the called data is relevant to the email to be replied to. The generality parameter is used to limit the specificity of the reply email content, which can achieve the effect of automatically generating effective reply emails while taking into account the control of information privacy.
[0058] In one embodiment, the first threshold and the generalization range are obtained by the following method:
[0059] Obtain background data; the background data includes: relationship data between the party to be responded to and the user, identity data of the party to be responded to, and identity data of the user;
[0060] The AI is invoked to determine the first threshold and the generalization range based on the email data to be replied to and the background data.
[0061] Specifically, the initial threshold and generalization range for replying to emails can be automatically generated by the system. For example, in cases with background data, when the identity data of the party to be replied to, the user, and the relationship between them can be obtained, AI can be used to analyze the identities and relationships of both parties, automatically generating the initial threshold and generalization range. The generated results can be sent to the user for confirmation or modification, and subsequent processing steps can only proceed after the user's approval.
[0062] In this embodiment, by using AI to automatically determine the first threshold and generalization range based on background data, the privacy control effect of automatically controlling the content of reply emails can be achieved.
[0063] In one embodiment, the first threshold and the generalization range are obtained by the following method:
[0064] The user determines the first threshold and the generalization range based on the email data to be replied to.
[0065] Specifically, the system can interact with the user, allowing the user to understand the specific content of the email data to be replied to, and then set a specific first threshold and generalization range.
[0066] In this embodiment, by allowing users to set a first threshold and a generalization range, users can more accurately control the privacy of their email replies.
[0067] In one embodiment, the method further includes:
[0068] At least two privacy levels are preset, and each privacy level corresponds to a set first threshold and a range of generalization.
[0069] Users determine the first threshold and the generalization range by selecting a privacy level.
[0070] Specifically, if a user's needs for controlling email privacy are relatively fixed, and the first threshold and generalization range only need to be selected from a few fixed parameters, then privacy levels can be preset in the system. Each privacy level has its own fixed first threshold and generalization range. This way, after understanding the content of the email to be replied to, the user only needs to select a privacy level to determine the first threshold and generalization range, making the selection process more convenient and intuitive. There should be at least two privacy levels, for example, 2 to 10. Generally, the higher the privacy level, the larger the first threshold, and the smaller the generalization value within the generalization range. This means less data can be accessed from the user database, and the email content generated based on the target data is more general. Users can set the number of privacy levels and the corresponding first threshold and generalization range for each privacy level according to their actual needs.
[0071] In one embodiment, the privacy level includes a maximum privacy level and a minimum privacy level;
[0072] At the highest privacy level, the generalization range is 0, and the reply email does not contain any content related to the target data;
[0073] At the lowest privacy level, the generalization range is 100%, and the reply email contains all the content of the target data.
[0074] The highest privacy level can refer to the level with the strictest privacy controls among all privacy levels. The lowest privacy level can refer to the level with the most lenient privacy controls among all privacy levels.
[0075] Specifically, the generalization range can be a numerical value. The highest privacy level can be set to a generalization of 0, meaning the AI-generated response email will not contain any content related to the target data. Conversely, the lowest privacy level can be set to a generalization of 100%, meaning the AI-generated response email will contain all content related to the target data.
[0076] In this embodiment, by defining specific values for the generality of the highest and lowest privacy levels, and the impact of these specific values on the content of reply emails, privacy level settings can be standardized, providing a benchmark for setting privacy levels.
[0077] In one embodiment, the method further includes:
[0078] Retrieve a user's sent email history;
[0079] The AI is required to adjust the language style of the reply email to match the style of the historical emails sent.
[0080] Among them, historical emails can refer to emails sent by a user in the past that contain information about the user's email sending style.
[0081] Specifically, a user's past emails can be sent to AI, allowing the AI to learn the user's email writing style and then adjust the wording of reply emails to match that style. This can increase user satisfaction with the content of the reply emails.
[0082] In one embodiment, the method further includes:
[0083] The similarity algorithm is used to calculate the similarity between the data in the user knowledge base and the email data to be replied to; the range of the generalization degree is 0~100%; the range of the first threshold is 0~100%.
[0084] Specifically, the generalization range can be any range from 0% to 100% or a single value, such as 0, 0.1% to 1%, 2% to 3%, 1% to 10%, 15% to 30%, or 100%. The first threshold can be any value from 0% to 100%, such as 0, 10%, 50%, 60%, 80%, or 90%. When the first threshold is 0, similarity calculation and target data filtering are not performed; these two steps are completed by default, and all data in the user database is retrieved when generating the reply email. The first threshold is generally not 100%; when the first threshold is set to 100%, the system will not retrieve any data from the user database.
[0085] In this embodiment, by setting the range of generalization and the range of the first threshold, the settings of these two parameters can be standardized.
[0086] In one embodiment, the method further includes:
[0087] A first control bar and a second control bar are set in the user interface; the first control bar is used for the user to set the first threshold; the second control bar is used for the user to set the generalization range.
[0088] In this embodiment, by setting a first control bar and a second control bar in the system-user interaction interface, the user can set the first threshold and generalization range through mouse or touch, which facilitates parameter setting.
[0089] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0090] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps:
[0091] Retrieve data of emails awaiting reply;
[0092] Based on the email data to be replied to, a first threshold and a generalization range are determined; the first threshold is the similarity threshold between the data in the user's knowledge base and the email data to be replied to; the generalization is the ratio of the number of words in the generalized text to the number of words in the original text.
[0093] Calculate the similarity between the data in the user knowledge base and the data of the email to be replied to, and obtain the matching similarity.
[0094] Filter data from the user knowledge base whose similarity exceeds the first threshold to obtain target data;
[0095] The AI is invoked to generate a reply email that conforms to the stated generalization range based on the target data and the email data to be replied to.
[0096] In one embodiment, the computer program, when executed by a processor, implements the steps in any of the above method embodiments.
[0097] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0098] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this disclosure can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (RRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this disclosure may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this disclosure may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0099] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0100] The embodiments described above are merely illustrative of several implementations of this disclosure, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent disclosure. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this disclosure, and these all fall within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the appended claims.
Claims
1. A method for generating email reply content based on a knowledge base, characterized in that, The method includes: Retrieve data of emails awaiting reply; Based on the email data to be replied to, a first threshold and a generalization range are determined; the first threshold is the similarity threshold between the data in the user's knowledge base and the email data to be replied to; the generalization is the ratio of the number of words in the generalized text to the number of words in the original text. Calculate the similarity between the data in the user knowledge base and the data of the email to be replied to, and obtain the matching similarity. Filter data from the user knowledge base whose similarity exceeds the first threshold to obtain target data; The AI is invoked to generate a reply email that conforms to the stated generalization range based on the target data and the email data to be replied to.
2. The method according to claim 1, characterized in that, The first threshold and the generalization range are obtained by the following method: Obtain background data; the background data includes: relationship data between the party to be responded to and the user, identity data of the party to be responded to, and identity data of the user; The AI is invoked to determine the first threshold and the generalization range based on the email data to be replied to and the background data.
3. The method according to claim 2, characterized in that, The user knowledge base includes an enterprise knowledge base and / or a personal knowledge base; the method further includes: The first threshold and generalization range determined by AI are sent to the user, and subsequent steps are carried out after the user modifies or confirms them.
4. The method according to claim 1, characterized in that, The first threshold and the generalization range are obtained by the following method: The user determines the first threshold and the generalization range based on the email data to be replied to.
5. The method according to claim 4, characterized in that, The method further includes: At least two privacy levels are preset, and each privacy level corresponds to a set first threshold and a range of generalization. Users determine the first threshold and the generalization range by selecting a privacy level.
6. The method according to claim 5, characterized in that, The privacy levels include the highest privacy level and the lowest privacy level; At the highest privacy level, the generalization range is 0, and the reply email does not contain any content related to the target data; At the lowest privacy level, the generalization range is 100%, and the reply email contains all the content of the target data.
7. The method according to claim 1, characterized in that, The method further includes: Retrieve a user's sent email history; The AI is required to adjust the language style of the reply email to match the style of the historical emails sent.
8. The method according to claim 1, characterized in that, The method further includes: The similarity algorithm is used to calculate the similarity between the data in the user knowledge base and the email data to be replied to; the range of the generalization degree is 0~100%; the range of the first threshold is 0~100%.
9. The method according to claim 4, characterized in that, The method further includes: A first control bar and a second control bar are set in the user interface; the first control bar is used for the user to set the first threshold; the second control bar is used for the user to set the generalization range.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.