AI reminding system based on voice conversion and text entry

Through an AI reminder system based on voice conversion and text entry, the outgoing call notifications of account managers are automatically processed, solving the problem of low efficiency of notification matters, realizing time saving and improving service quality.

CN120201127APending Publication Date: 2025-06-24福建省烟草公司漳州市公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510378374.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The account manager notifies many things and is less efficient in his daily work, which leads to the mechanized work squeezing personalized service time and low notification efficiency.

Method used

An AI reminder system based on voice conversion and text entry is provided, including a voice generation engine, an out-of-call execution unit, a communication result analysis and processing unit, a data storage unit and an encryption unit, through which automated out-of-call reminder and result analysis are realized.

Benefits of technology

Save time for account managers to notify matters, improve notification efficiency, improve service quality, and improve customer satisfaction through personalized outbound voice and flexible communication strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201127A_ABST
    Figure CN120201127A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic voice communication, in particular to an AI reminding system based on voice conversion and text entry, which comprises a front-end client and a rear-end function module, the rear-end function module comprises a voice generation engine, an outbound execution unit, a communication result analysis and processing unit, a data storage unit and an encryption unit, an outbound execution strategy is set, and the communication result analysis and processing unit is encrypted. The outbound call execution unit is enabled to automatically perform voice call reminding on batch clients based on the execution strategy, so that the time of notifying items by a client manager is saved, the notification efficiency is improved, the time of providing a large amount of available personalized services is saved, and the improvement of the service quality is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of automatic voice communication technology, and specifically relates to an AI reminder system based on voice conversion and text entry. Background Art

[0002] Currently, the average number of customers per customer manager is relatively large, and there are many notification matters and low efficiency in daily work, resulting in mechanical work squeezing the time for personalized services. After investigation, the average monthly call duration for notification matters is about 6.21 hours, and the average duration of visiting customers for notification matters is 2257.6 minutes. The time consumed by customer managers on notification matters accounts for about 14.18% of the total customer service duration, seriously squeezing the time for customer managers to handle other matters, and the notification efficiency is low. Summary of the Invention

[0003] This application provides an AI reminder system based on voice conversion and text entry to solve or partially solve the problems raised in the above background art.

[0004] This application provides an AI reminder system based on voice conversion and text entry, including: a front-end client and a back-end function module. The back-end function module includes a voice generation engine, an outbound execution unit, a communication result analysis and processing unit, a data storage unit, and an encryption unit. The functions of each unit are as follows:

[0005] The voice generation engine receives a text set and an artificial recording set, generates a first voice set based on text transcription, and matches and combines the transcribed voices in the first voice set with the recordings in the artificial recording set to generate outbound voices;

[0006] The outbound execution unit selects an outbound voice based on the notification type and customer information and links to the outbound API interface of the communication service provider to implement outbound execution. It receives the outbound result information returned by the outbound API interface, sets the outbound execution strategy, and enables the outbound execution unit to automatically perform voice call reminders on a batch of customers based on the execution strategy;

[0007] The communication result analysis and processing unit receives the outbound result information and performs statistics to generate a call record list, and analyzes the historical call records of each customer to construct customer tags, where the customer tags include communication strategies;

[0008] The data storage unit is provided with a database for data storage;

[0009] The encryption unit performs encryption processing on sensitive data.

[0010] Preferably, the specific method for the voice generation engine to generate outbound voices based on text and recordings is as follows:

[0011] S1: Invoke the TTS engine to transcribe the text collection to generate the first voice collection, and set voice tags for the collection elements based on the notification type;

[0012] S2: Set the voice template, which includes a first interaction segment, a business segment, and a second interaction segment set in sequence. Among them, the first interaction segment and the second interaction segment use artificial recordings, the business segment uses the voice converted from text, the first interaction segment is associated with customer information, the business segment is associated with the notification type, and the second interaction segment is the closing statement associated with the notification type;

[0013] S3: Based on the customer information and the notification type, select recordings from the artificial recording collection to match the first interaction segment and the second interaction segment, select the transcribed voice from the first voice collection as the business segment based on the notification type, fill the voice template, and generate the outbound voice corresponding to [customer + notification type].

[0014] Preferably, in step S1, the TTS engine selects an open-source TTS engine interface, and the format of the artificial recording is selected as the MP3 format of 320 kbps.

[0015] Preferably, in step S1, the notification types at least include: unordered notification, important event notice, and temporary matter notice.

[0016] Preferably, the communication result analysis and processing unit receives the outbound call result information and conducts statistics to generate a call record list. The outbound call result information at least includes whether the call is connected and the call duration information. The element list of the call record list at least includes customer information, whether the call is connected, and the call duration.

[0017] Preferably, the specific method for the communication result analysis and processing unit to analyze the historical call records of each customer and construct customer tags is as follows:

[0018] Statistical connection rate and average call duration of the outbound voice of each customer for each type of notification. If the connection rate is greater than the threshold n1 and the average call duration is greater than the threshold n2, the customer tag is marked as normal, and no additional communication strategy needs to be set;

[0019] If the connection rate is less than or equal to the threshold n1 and the average connection duration is greater than the threshold n2, it is marked as a type 1 key customer, and it is recommended to adjust the notification period;

[0020] If the connection rate is greater than the threshold n1 and the average connection duration is less than or equal to the threshold n2, it is marked as a type 2 key customer, and it is recommended to adjust the voice template;

[0021] If the connection rate is less than or equal to the threshold n1 and the average connection duration is less than or equal to the threshold n2, it is marked as a type 3 key customer, and it is recommended to conduct an offline visit.

[0022] Preferably, the threshold n2 is set based on the notification type, and n2 = Ts i + Ratei * Tmi, where Ts i is the duration of the first interaction segment of the outbound voice of the i-type notification of the corresponding customer, Tmi is the duration of the service segment of the i-type notification of the corresponding customer, Ratei is a preset proportional threshold, which is set based on the content of the service segment, and i is the notification type.

[0023] Preferably, the specific method for the encryption unit to encrypt sensitive data is as follows:

[0024] Data collection stage: When collecting sensitive data, encrypt it immediately and then store it in the database.

[0025] Data usage stage: When it is necessary to use the encrypted data for voice calls, decrypt it first and then use the decrypted data.

[0026] Preferably, the encryption unit uses the AES encryption algorithm and padding method, and uses the KMS key management system to store and manage keys, including regularly replacing keys.

[0027] Preferably, the front-end client uses the Vue framework, the back-end functional module uses the python language for system code development, and the data storage unit uses the relational database mysql for user data processing and recording.

[0028] Compared with the prior art, the beneficial effects of the present application are as follows:

[0029] (1) Based on the outbound execution unit, the present application makes batch outbound notifications to customers, saves the time of the customer manager for notification matters, improves the notification efficiency, saves a large amount of time that can be used to provide personalized services, and is beneficial to improving the service quality.

[0030] (2) Through the voice generation engine, the present application generates personalized outbound voices based on the combination of text transcription and manual recording, which is beneficial to improving customer satisfaction and ensuring the notification effect.

[0031] (3) Through the communication result analysis and processing unit, the present application analyzes and statistics the outbound results, constructs customer tags, screens key customers based on the outbound results and adjusts the communication strategy, ensuring the flexibility and effectiveness of the notification reminder. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The present application will be further described below with reference to the drawings and embodiments.

[0033] Figure 1 is a schematic diagram of the overall system composition of the present application,

[0034] Figure 2 is a software framework diagram of an embodiment of the present application. Detailed implementation manners

[0035] As used in the specification and claims, certain terms are used to refer to specific components. Those skilled in the art should understand that hardware manufacturers may use different terms to refer to the same component. The specification and claims do not distinguish components by the difference in names, but by the difference in functions of the components. As used throughout the specification and claims, "comprising" is an open-ended term and should be interpreted as "including but not limited to". "Substantially" means within an acceptable error range. Those skilled in the art can solve the technical problems within a certain error range and basically achieve the technical effects.

[0036] In the description of the present application, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "front", "rear", "left", "right", "horizontal", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application.

[0037] In the present application, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected to", "fixed", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.

[0038] Embodiment 1

[0039] As Figures 1 to 2 shown, the present application provides an AI reminder system based on voice conversion and text input, including a front-end client and a back-end function module. The back-end function module includes a voice generation engine, an outbound execution unit, a communication result analysis and processing unit, a data storage unit, and an encryption unit. The functions of each unit are as follows:

[0040] The voice generation engine receives a text set and an artificial recording set, generates a first voice set based on text transcription, and matches and combines the transcribed voices in the first voice set with the recordings in the artificial recording set to generate outbound voices;

[0041] The outbound execution unit selects an outbound voice based on the notification type and customer information and links to the outbound API interface of the communication service provider to implement outbound execution, and receives the outbound result information returned by the outbound API interface;

[0042] The communication result analysis and processing unit receives the outbound call result information, conducts statistics, generates a call record list, and analyzes the historical call records of each customer to construct customer tags, where the customer tags include communication strategies;

[0043] The data storage unit is provided with a database for data storage;

[0044] The encryption unit performs encryption processing on sensitive data.

[0045] Users upload texts and recordings through the front-end client, control the outbound call execution unit to make voice call reminders to customers, and can set outbound call execution strategies, enabling the outbound call execution unit to automatically make voice call reminders to a batch of customers based on the execution strategies. The outbound call results, namely information such as the call record list and customer tags, can be viewed through the front-end client, saving the reminder time of the customer manager, improving efficiency, and enhancing the reminder effect.

[0046] Specifically, the specific method for the voice generation engine to generate outbound call voices based on texts and recordings is as follows:

[0047] S1: Call the TTS engine to transcribe the text set to generate a first voice set, and set voice tags for the set elements based on the notification type;

[0048] S2: Set a voice template, which includes a first interaction segment, a business segment, and a second interaction segment set in sequence. Among them, the first interaction segment and the second interaction segment use artificial recordings, the business segment uses the voice converted from text, the first interaction segment is associated with customer information, the business segment is associated with the notification type, and the second interaction segment is an ending sentence associated with the notification type;

[0049] S3: Based on the customer information and the notification type, select recordings from the artificial recording set to match the first interaction segment and the second interaction segment, select the transcribed voice from the first voice set as the business segment based on the notification type, fill the voice template, and generate the outbound call voice corresponding to [customer + notification type].

[0050] In step S1, the TTS engine selects an open-source TTS engine interface. TTS engine interfaces for text-to-speech with excellent performance are provided by manufacturers such as Baidu Voice, iFlytek, and Alibaba Cloud Voice. This application selects the TTS engine interface of China Unicom Voice Service; the format of the artificial recording is selected as the MP3 format of 320 kbps.

[0051] In step S1, the voice tag of the first voice can be set according to the file name of the text. The notification type at least includes: unordered notice, important event notice, temporary matter notice. The important matter notice is an event that needs to be carried out every month, and the temporary matter is a non-periodic temporary matter, such as the adjustment of the order date and other matters.

[0052] In step S2, the configuration method of the artificial recording and the transcribed speech in the voice template is obtained based on the actual customer experience research. Taking a simple unordered notice as an example, "Dear xx customer, hello, yy Tobacco reminds you that xxxdate is your ordering day. The system detects that you have not placed an order yet. Please do not miss the ordering time and affect the normal sales. yy Tobacco wishes you a prosperous business." The front part "Dear xx customer, hello" and the tail part "Please do not miss the ordering time and affect the normal sales. yy Tobacco wishes you a prosperous business" are the first interaction segment and the second interaction segment, providing full emotional value through artificial recording, and the middle paragraph involving business information uses transcribed speech to accurately notify the specific business information to the customer.

[0053] Specifically, the outbound execution unit automatically makes voice call reminders to a batch of customers through the outbound execution strategy. The outbound execution strategy is preset, mainly setting the customer group and the time threshold. In this application, China Unicom is selected as the communication service provider.

[0054] Specifically, the communication result analysis and processing unit receives the outbound result information and conducts statistics to generate a call record list. The outbound result information at least includes information such as whether the call is connected and the call duration. The element columns of the call record list at least include customer information, whether the call is connected, the call duration, etc. The specific method for analyzing the historical call records of each customer and constructing customer tags is as follows:

[0055] Statistical analysis is carried out on the connection rate and average call duration of the outbound voice of each customer for each type of notice. If the connection rate is greater than the threshold n1 and the average call duration is greater than the threshold n2, the customer tag is marked as normal and no additional communication strategy needs to be set;

[0056] If the connection rate is less than or equal to the threshold n1 and the average connection duration is greater than the threshold n2, it is marked as a type 1 key customer, and it is recommended to adjust the notification period;

[0057] If the connection rate is greater than the threshold n1 and the average connection duration is less than or equal to the threshold n2, it is marked as a type 2 key customer, and it is recommended to adjust the voice template;

[0058] If the connection rate is less than or equal to the threshold n1 and the average connection duration is less than or equal to the threshold n2, it is marked as a type 3 key customer, and it is recommended to conduct an on-site visit.

[0059] The threshold n2 is set based on the notification type, n2 = Tsi + Ratei * Tmi, where Tsi is the duration of the first interaction segment of the outbound voice of the i-type notice of the corresponding customer, Tmi is the duration of the business segment of the i-type notice of the corresponding customer, Ratei is the preset proportional threshold, which is set based on the content of the business segment, and i is the notification type, that is, it is considered that the notification is completed if the customer receives the core vocabulary of the business segment.

[0060] Specifically, the specific method for the encryption unit to encrypt sensitive data is as follows:

[0061] Data collection stage: When collecting sensitive data, encrypt it immediately and then store it in the database.

[0062] Data usage stage: When using encrypted data for voice calls, decrypt it first and then use the decrypted data.

[0063] Furthermore, use the AES encryption algorithm and padding method, and use the KMS key management system to store and manage keys, including regularly replacing keys.

[0064] The front-end client of this application uses the Vue framework, the back-end functional module uses the Python language for system code development, and the data storage unit uses a relational database (MySQL) to process and record user data.

[0065] After statistics, the notification duration for the account manager to use this system to batch notify customers is about 2.34 hours, less than 3 hours. Compared with 6.21 hours before using the system, the efficiency has increased by 165.38%.

[0066] The embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the present application is not limited to the above embodiments. Within the knowledge scope of those of ordinary skill in the art, various changes can be made without departing from the purpose of the present application.

Claims

1. AI reminder system based on voice conversion and text input, characterized by: include: The front-end client and back-end functional modules include a voice generation engine, an outbound call execution unit, a communication result analysis and processing unit, a data storage unit, and an encryption unit. The functions of each unit are as follows: A speech generation engine receives a text set and a manual recording set, generates a first speech set based on text transcription, and matches and combines the transcribed speech in the first speech set with the recording in the manual recording set to generate an outbound call speech; The outbound call execution unit selects outbound call voice based on notification type and customer information and links the outbound call API interface of the communication service provider to implement outbound call execution, receives the outbound call result information returned by the outbound call API interface, sets the outbound call execution strategy, and enables the outbound call execution unit to automatically make voice call reminders to batch customers based on the execution strategy; The communication result analysis and processing unit receives and counts the outbound call result information, generates a call record list, and analyzes each customer's historical call record to construct a customer tag, which includes a communication strategy. A data storage unit, provided with a database for data storage; Encryption unit, encrypts sensitive data.

2. The AI ​​reminder system based on voice conversion and text input according to claim 1, characterized in that: The specific method of the speech generation engine generating outbound speech based on text and recording is as follows: S1: calling the TTS engine to transcribe the text set into a first voice set, and setting voice tags for the set elements based on the notification type; S2: Setting a voice template, the voice template includes a first interaction segment, a business segment, and a second interaction segment, which are set in sequence. The first interaction segment and the second interaction segment use manual recordings, the business segment uses text-converted speech, the first interaction segment is associated with customer information, the business segment is associated with a notification type, and the second interaction segment is a closing statement associated with the notification type; S3: Based on the customer information and notification type, a recording is selected from the manual recording set to match the first interaction segment and the second interaction segment. Based on the notification type, a transcribed voice is selected from the first voice set as the business segment, and a voice template is filled to generate an outbound voice corresponding to [customer+notification type].

3. The AI ​​reminder system based on voice conversion and text input according to claim 2, characterized in that: In step S1, the TTS engine selects an open source TTS engine interface, and the manual recording format selects a 320 kbps MP3 format.

4. The AI ​​reminder system based on voice conversion and text input according to claim 2, characterized in that: In step S1, the notification types include at least: unordered notification, important event notification, and temporary event notification.

5. The AI ​​reminder system based on voice conversion and text input according to claim 1, characterized in that: The communication result analysis processing unit receives and counts the outbound call result information to generate a call record list. The outbound call result information at least includes whether the call is connected and the call duration. The element columns of the call record list at least include customer information, whether the call is connected and the call duration.

6. The AI ​​reminder system based on voice conversion and text input according to claim 5, characterized in that: The communication result analysis processing unit analyzes each customer's historical call records and constructs a customer tag in the following specific method: Statistics are collected on the connection rate and average answering time of each customer's outbound voice calls for each type of notification. If the connection rate is greater than the threshold n1 and the average answering time is greater than the threshold n2, the customer label is marked as normal, and no additional communication policy needs to be set. If the connection rate is less than or equal to threshold n1 and the average connection time is greater than threshold n2, it is marked as a type 1 key customer and it is recommended to adjust the notification period; If the connection rate is greater than the threshold n1 and the average connection time is less than or equal to the threshold n2, the customer is marked as a type 2 key customer and it is recommended to adjust the voice template; If the connection rate is less than or equal to threshold n1 and the average connection time is less than or equal to threshold n2, the customer is marked as type 3 key customer and an offline visit is recommended.

7. The AI ​​reminder system based on voice conversion and text input according to claim 6, characterized in that: The threshold n2 is set based on the notification type, n2=Tsi+Ratei*Tmi, where Tsi is the first interaction segment duration of the outbound voice of the corresponding customer's i-type notification, Tmi is the business segment duration of the corresponding customer's i-type notification, Ratei is a preset ratio threshold, which is set based on the content of the business segment, and i is the notification type.

8. The AI ​​reminder system based on voice conversion and text input according to claim 1, characterized in that: The specific method for the encryption unit to encrypt sensitive data is as follows: Data collection stage: When collecting sensitive data, it is immediately encrypted and then stored in the database; Data usage stage: When encrypted data is needed for voice dialing, decryption is performed first, and then the decrypted data is used.

9. The AI ​​reminder system based on voice conversion and text input according to claim 8, characterized in that: The encryption unit uses the AES encryption algorithm and padding method, and uses the KMS key management system to store and manage keys, including regular key replacement.

10. The AI ​​reminder system based on voice conversion and text input according to claim 1, characterized in that: The front-end client adopts the Vue framework, the back-end functional module adopts the python language for system code development, and the data storage unit uses the relational database mysql for user data processing and recording.