Outbound voice generation method and device, storage medium and electronic equipment

By obtaining the personalized parameter data of the target account, generating a target account portrait, determining characteristic parameters, and seamlessly splicing with fixed speech audio streams, the problem of insufficient personalized services in the intelligent outbound call system is solved, personalized voice interaction is achieved, and user experience is improved.

CN120475104APending Publication Date: 2025-08-12CHINA TELECOM BESTPAY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510693757.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing intelligent outbound call system lacks the ability to accurately identify and respond to personalized account needs, which leads to the standardized broadcast content during the call that cannot meet the personalized needs of different accounts.

Method used

By obtaining the personalized parameter data of the target account, generating a target account portrait, determining the target feature parameters, generating a pre-synthesis audio stream that matches the target feature parameters, and splicing it with the pre-recorded fixed speech audio stream to form a personalized target audio stream for broadcasting.

Benefits of technology

It realizes highly personalized voice interaction, improves the user experience and satisfaction of outgoing call calls, and solves the problem of insufficient personalized services in smart outgoing call.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475104A_ABST
    Figure CN120475104A_ABST
Patent Text Reader

Abstract

The invention discloses an outbound voice generation method and device, a storage medium and electronic equipment. The method relates to the field of artificial intelligence, and comprises the following steps: in a process of performing voice interaction with a target account, obtaining personalized parameter data of the target account; generating a target account portrait based on the personalized parameter data; determining a target feature parameter corresponding to the target account based on the target account portrait; generating a pre-synthesized audio stream matched with the target feature parameter; and splicing the pre-synthesized audio stream and a pre-recorded fixed verbal skill audio stream to obtain a target audio stream, and broadcasting the target audio stream to the target account. According to the method and the device, the technical problem of insufficient personalized services in intelligent outbound calling in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to an outbound voice generation method, device, storage medium and electronic device. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, intelligent communications are playing an increasingly important role in business communications. As a key application of intelligent communications technology, intelligent outbound call systems provide businesses with efficient customer communication and service solutions through automated calls and voice interaction. However, these outbound call systems often lack the ability to accurately identify and respond to individual customer needs, resulting in standardized announcements during calls that fail to meet the personalized needs of different accounts.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device, storage medium and electronic device for generating an outbound voice, so as to at least solve the technical problem of insufficient personalized service in intelligent outbound calling in related technologies.

[0005] According to one aspect of an embodiment of the present invention, a method for generating outbound voice is provided, comprising: obtaining personalized parameter data of the target account during voice interaction with a target account; generating a target account portrait based on the personalized parameter data; determining target feature parameters corresponding to the target account based on the target account portrait; generating a pre-synthesized audio stream that matches the target feature parameters; splicing the pre-synthesized audio stream and a pre-recorded fixed-speech audio stream to obtain a target audio stream, and broadcasting the target audio stream to the target account.

[0006] According to another aspect of an embodiment of the present invention, an outbound voice generation device is also provided, including: a data acquisition module, used to obtain personalized parameter data of the target account during voice interaction with the target account; an account portrait generation module, used to generate a target account portrait based on the personalized parameter data; a parameter determination module, used to determine the target feature parameters corresponding to the target account based on the target account portrait; an audio stream synthesis module, used to generate a pre-synthesized audio stream matching the target feature parameters; an audio stream splicing module, used to splice the pre-synthesized audio stream and a pre-recorded fixed speech audio stream to obtain a target audio stream, and broadcast the target audio stream to the target account.

[0007] According to another aspect of an embodiment of the present invention, a non-volatile storage medium is provided, wherein the non-volatile storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executed by any one of the outbound voice generation methods.

[0008] According to another aspect of an embodiment of the present invention, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement any one of the outbound voice generation methods.

[0009] In an embodiment of the present invention, personalized parameter data of the target account is obtained during voice interaction with the target account; a target account portrait is generated based on the personalized parameter data; target characteristic parameters corresponding to the target account are determined based on the target account portrait; a pre-synthesized audio stream matching the target characteristic parameters is generated; the pre-synthesized audio stream and a pre-recorded fixed speech audio stream are spliced to obtain a target audio stream, and the target audio stream is broadcast to the target account, thereby achieving the purpose of dynamically obtaining and analyzing personalized parameter data of the target account, generating and optimizing the target account portrait, and then accurately determining account characteristics, pre-synthesizing personalized pre-synthesized audio stream, and seamlessly splicing it with the fixed speech audio stream to form the target audio stream, thereby achieving highly personalized voice interaction and improving the user experience and satisfaction of outbound calls, thereby solving the technical problem of insufficient personalized services in related technologies in intelligent outbound calls. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0011] Figure 1 is a flow chart of a method for generating an outbound voice according to an embodiment of the present invention;

[0012] Figure 2 1 is a schematic diagram of an optional fixed speech audio stream according to an embodiment of the present invention;

[0013] Figure 3 is a schematic diagram of an optional pre-synthesized audio stream according to an embodiment of the present invention;

[0014] Figure 4 is a schematic diagram of an optional target audio stream according to an embodiment of the present invention;

[0015] Figure 5 is a flow chart of an optional outbound voice generation method according to an embodiment of the present invention;

[0016] Figure 6 FIG. 4 is a schematic diagram of an outbound voice generation device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0018] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0019] According to an embodiment of the present invention, a method embodiment for generating an outbound voice is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0020] Figure 1 FIG. 1 is a flow chart of a method for generating an outbound voice according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps:

[0021] Step S102, during the voice interaction with the target account, obtaining personalized parameter data of the target account;

[0022] Optionally, when the intelligent outbound calling system initiates or receives a call with a target account, it collects personalized parameter data related to that account in real time. This personalized parameter data may include, but is not limited to, the account's geographic location, historical transaction records, preferences, and historical interaction data (such as previous phone calls and online interactions). The intelligent outbound calling system can obtain rich information about the target account, providing data support for subsequent generation of personalized broadcast content and dynamic adjustment of call strategies, thereby achieving more efficient and accurate customer service.

[0023] Optionally, after the personalized parameter data is collected, the collected personalized parameter data can be cleaned to remove invalid or redundant information for subsequent analysis and utilization.

[0024] Step S104: Generate a target account profile based on the personalized parameter data;

[0025] Optionally, key features can be extracted from the personalized parameter data. Based on this feature extraction, the target account's relevant information is compared and analyzed with historical account data to identify similarities and differences between the target account and other accounts, thereby constructing a profile that includes information such as account preferences, behavioral habits, and preferences. This target account profile is not just a static description; it can also include dynamically updated elements to reflect real-time changes in account needs and behaviors.

[0026] In an optional embodiment, when the personalized parameter data includes account information of the target account, a target account portrait is generated based on the personalized parameter data, including: obtaining multiple groups of account history portrait information included in the user behavior library; classifying and processing the multiple groups of account history portrait information according to account characteristics to obtain multiple categories of account history portrait information; determining a first account portrait group that matches the personalized parameter data from the multiple categories of account history portrait information; cross-matching the first account portrait with the multiple categories of account history portrait information to obtain a second account portrait group; and generating a target account portrait based on the second account portrait group.

[0027] Optionally, a large amount of historical interaction data from the target account and other accounts is collected, including but not limited to transaction records, service requests, and feedback reviews. This data forms an account behavior library. Multiple sets of historical account profile information are extracted from the behavior library. Each set of profile information reflects the behavioral patterns and preferences of a specific account in past interactions. Data mining and machine learning algorithms can be used to classify these extracted sets of account profile information. Classification can be based on, but is not limited to, characteristics such as the account's spending habits, interests, preferences, and service needs. This classification categorizes accounts into several categories with similar characteristics, such as "Category A Account" and "Category B Account," which facilitates subsequent personalized analysis and the development of advertising strategies. By comparing the target account's personalized parameter data with the classified account historical profile information, the category with the characteristics most similar to the target account is identified, thereby determining the first set of account profiles. This set of profiles contains historical behavior and preference information for accounts similar to the target account, providing rich context for generating personalized advertising content. After determining the first account profile group, we further analyze the correlations and differences between the first account profile group and other categories of account profile information. Through cross-matching technology, we can integrate the characteristics of different account categories to form a more comprehensive and accurate second account profile group. This captures a wider range of behavioral patterns and potential needs, making the target account profile more rich and three-dimensional. Based on the information of the second account profile group, we further generate a highly personalized target account profile. The target account profile not only contains the target account's basic information and historical behavior records, but also incorporates the preferences and needs of similar accounts. It can more accurately predict the target account's expectations and reactions, thereby guiding subsequent parameter adjustments and voice synthesis processes, ensuring the personalization and effectiveness of intelligent outbound calls. This method can achieve a deeper understanding of the target account and provide a more personalized and demand-oriented service experience for the target account, thereby improving outbound call effectiveness and customer satisfaction.

[0028] In an optional embodiment, when the personalized parameter data also includes historical interaction data of the target account, a target account portrait is generated based on the second account portrait group, including: obtaining the target account portrait based on the second account portrait group and the historical interaction data.

[0029] Optionally, after generating the second account profile group, the profile information included in the second account profile group is combined with the target account's specific historical interaction data. This historical interaction data may include, but is not limited to, the account's past communication records, service requests, operational behaviors, feedback, and evaluations. Using data fusion techniques, the common features in the second account profile group are cross-compared and analyzed with the personalized information in the target account's historical interaction data. The target account data can be obtained by analyzing the second account profile group and historical interaction data through, but is not limited to, cluster analysis and association rule mining. The goal is to identify the target account's specific behavioral patterns, preferences, and potential needs from this historical data.

[0030] Optionally, during the analysis based on the second account profile group and historical interaction data, personalized parameters displayed by the target account in historical interactions can be identified, such as preferences for specific products or services, communication methods, etc. These parameters will be extracted and adjusted and optimized based on their importance and frequency in the historical interaction data to ensure that they are fully considered when generating the target account profile.

[0031] Ultimately, based on the fusion and analysis of the second account profile group and historical interaction data, a more comprehensive and personalized target account profile is constructed. This profile not only incorporates common features inherited from the second account profile group but also incorporates the target account's historical interaction data, better reflecting the target account's unique needs and behavioral preferences. This generated target account profile will guide subsequent parameter adjustments and personalized speech synthesis to provide more accurate and personalized outbound call services.

[0032] It's important to note that as target accounts continue to interact with the intelligent outbound call system, their historical data will be constantly updated, and accordingly, the target account profile should be dynamically adjusted. The latest historical interaction data can be incorporated into the target account profile on a regular or real-time basis to ensure the timeliness and accuracy of the profile information, thereby continuously optimizing the personalization level of outbound call service.

[0033] In the above method, based on the comprehensive information of the target account (including historical behavior and real-time interaction data), a more detailed and realistic target account portrait is generated, which can provide a solid data foundation for subsequent personalized parameter adjustment and voice synthesis, and significantly improve the personalization of outbound call services and customer satisfaction.

[0034] Step S106: determining target characteristic parameters corresponding to the target account based on the target account profile;

[0035] Optionally, after generating a target account profile, the detailed information within the profile can be further analyzed to identify target feature parameters that directly influence subsequent audio generation. These target feature parameters may include, but are not limited to, the account's geographic location, transaction preferences, historical consultation hotspots, and emotional state, directly reflecting the account's personalized needs and communication preferences.

[0036] In an optional embodiment, based on the target account portrait, the target characteristic parameters corresponding to the target account are determined, including: using multiple characteristic parameters selected from multiple categories of account historical portrait information as multiple initial clustering centers; clustering the multiple categories of account historical portrait information based on the multiple initial clustering centers to obtain clustering results; and classifying the target account portrait based on the clustering results to obtain the target characteristic parameters.

[0037] Optionally, multiple representative feature parameters are selected from multiple categories of historical account profile information as initial cluster centers. These feature parameters can cover multiple account dimensions, such as location, spending habits, and preferences. They serve as the starting point for cluster analysis, identifying similarities and differences between accounts. Machine learning clustering algorithms (such as k-means and hierarchical clustering) are used to cluster the historical account profile information. Based on the initial cluster centers, the clustering process calculates the distance between the account profile information and each center point, assigning the account information to the category containing the closest center point, resulting in a clustering result. This result further subdivides the account into multiple small groups with similar characteristics, helping to extract more specific target feature parameters. After obtaining the clustering results, the target account profile is classified to identify the cluster category or categories to which it belongs. Based on the target account profile's classification in the clustering results, the best matching feature parameter set, namely the target feature parameters, can be extracted. These parameters, derived from the target account profile and combined with big data analysis results, reflect the specific needs and preferences of the target account and serve as a key basis for the subsequent development of personalized service strategies. This approach leverages complex clustering algorithms and data analysis techniques to accurately identify target account characteristics from a wealth of historical account profile information, providing powerful support for personalized outbound call services. This approach not only reduces manual configuration workload but also significantly improves the quality and efficiency of outbound call services, providing a superior user experience.

[0038] In an optional embodiment, the target account portrait is classified based on the clustering results to obtain target feature parameters, including: classifying the target account portrait based on the clustering results to obtain first feature parameters; obtaining recommended broadcast parameters during the voice interaction process; cross-comparing the first feature parameters and the recommended broadcast parameters to obtain target feature parameters.

[0039] Optionally, by classifying the target account profile based on clustering results, common characteristics of groups of accounts highly similar to the target account can be identified. These characteristics constitute the first feature parameters, helping to extract core parameters directly related to the target account from big data, laying a more specific and focused foundation for subsequent personalized services. Obtaining recommended announcement parameters during voice interaction allows the target account's immediate feedback and dynamic behavior to be considered during real-time calls, providing real-time and dynamic support for intelligent outbound calling services. Recommended announcement parameters can be dynamically generated based on factors such as the target account's emotional changes and the depth of the conversation during the call, improving real-time adaptability to outbound call scenarios. Cross-comparing the first feature parameters with the recommended announcement parameters aims to determine the final target feature parameters by integrating static data analysis results with dynamic interaction feedback. This process identifies differences and intersections between the two, ensuring that outbound calling services incorporate both the target account's historical preferences and the immediate needs of the call, thereby providing more precise and personalized services. Refined target feature parameter generation also avoids providing irrelevant or uninteresting information to the target account, reducing the possibility of ineffective communication and improving outbound call efficiency and targeted service.

[0040] In the above methods, through the combination of static data analysis and dynamic interactive feedback, deep personalization of outbound call services can be achieved, which can not only enhance customer experience but also improve service efficiency.

[0041] Step S108, generating a pre-synthesized audio stream matching the target feature parameters;

[0042] Optionally, based on the characteristic parameters in the target account profile, content points that the account may be interested in are analyzed. For example, if the account profile shows that it has special needs for financial services, this characteristic parameter will be identified, and preset financial-related speech content will be selected for matching. This process ensures that the generated pre-synthesized audio stream is closely related to the interests and needs of the account, improving the targeting and efficiency of outbound calls. After determining the content points that match the account characteristic parameters, these contents are converted into text format. For example, a preset text template can be called from the database, or personalized text content can be dynamically generated based on the characteristic parameters. For example, for accounts that prefer financial security knowledge, the system will generate a text containing the latest financial security tips.

[0043] Optionally, you can use Text-to-Speech (TTS) technology to convert the text generated in the previous step into a pre-synthesized speech stream. TTS technology converts text information into a natural and fluent speech broadcast. It can also adjust parameters such as speech rate and intonation to better align the speech with the target account's preferences, enhancing communication effectiveness. For example, for accounts that prefer a fast-paced communication, you can increase the speech rate appropriately.

[0044] Through the above approach, we can generate pre-synthesized audio streams that closely match the target account's characteristics, providing a more personalized and attentive service experience during actual calls with the account. This not only helps improve account satisfaction, but also increases the conversion rate and efficiency of outbound calls.

[0045] Step S110: splice the pre-synthesized audio stream and the pre-recorded fixed speech audio stream to obtain a target audio stream, and broadcast the target audio stream to the target account.

[0046] Optional, pre-recorded fixed-script audio streams form the foundation for standardized communication in intelligent outbound call systems. These audio streams contain customer service personnel's pre-prepared scripts for different scenarios and purposes, such as greetings, product introductions, and service descriptions. Pre-synthesized audio streams generated based on target feature parameters can reflect the user's personalized needs. The target audio stream generated in this way not only meets the user's business needs but also their personalized experience. This highly customized target audio stream, which conforms to professional customer service standards, can significantly improve the quality and efficiency of outbound call service, enhancing user interactivity and customer experience.

[0047] In an optional embodiment, a pre-synthesized audio stream and a pre-recorded fixed-speech audio stream are spliced to obtain a target audio stream, including: reading slot information in the fixed-speech audio stream; and splicing the pre-synthesized audio stream to the fixed-speech audio stream according to the slot information to obtain the target audio stream.

[0048] Optional, pre-recorded fixed-script audio streams form the foundation for standardized communication within intelligent outbound call systems. These streams contain pre-prepared scripts for customer service personnel for different scenarios and purposes, such as greetings, product introductions, and service descriptions. Before splicing the audio, it is necessary to identify the slots within the fixed-script audio stream where personalized information should be inserted—that is, the locations within the speech segments where the parameters should be embedded. Pre-synthesized audio streams generated based on the target feature parameters are precisely inserted into the corresponding slots within the fixed-script audio stream. Figure 2 is a schematic diagram of an optional fixed speech audio stream according to an embodiment of the present invention. Figure 3 is a schematic diagram of an optional pre-synthesized audio stream according to an embodiment of the present invention. Figure 4 This is a schematic diagram of an optional target audio stream according to an embodiment of the present invention. Figure 2 The slot information in the fixed speech audio stream will be Figure 3 Insert the pre-synthesized audio stream in the corresponding position and you can get the following Figure 4The target audio stream is shown. For example, if the fixed script is "Dear customer, we have noticed that you are very interested in {product name} recently. We will introduce it to you in detail...", the system will insert the personalized pre-synthesized audio stream corresponding to the "product name" into the curly bracket position to achieve seamless embedding of customized content. Utilize audio splicing technology, such as using tools such as the Free Multimedia Framework (ffmpeg), to seamlessly splice the pre-synthesized audio stream with the fixed script audio stream to generate a complete target audio stream. This process not only ensures the continuity and naturalness of the audio content, but also handles issues such as volume balance, tone matching, and silent segment transitions between different audio streams to ensure the quality of the final audio stream.

[0049] Optionally, further optimization can be performed on the spliced target audio stream, such as sound quality enhancement and background noise elimination, to ensure clear audio during transmission and provide a good call experience. Furthermore, audio stream transmission delay can be monitored to ensure real-time performance and avoid unnecessary waiting during calls due to technical issues.

[0050] Optionally, after generating the target audio stream, the generated target audio stream is broadcast to the target account in real time through the intelligent outbound calling system. During the broadcast process, the account's response (such as feedback information obtained by voice recognition technology) is continuously monitored, and subsequent voice generation strategies are adjusted based on the feedback, optimizing the service process to achieve the best communication effect and customer satisfaction.

[0051] Through the above method, standardized fixed scripts can be combined with personalized pre-synthesized audio streams to generate target audio streams that meet professional customer service standards and are highly customized. This can significantly improve the quality and efficiency of outbound call services and enhance the interactivity with accounts and customer experience.

[0052] In an optional embodiment, a pre-synthesized audio stream and a pre-recorded fixed-speech audio stream are spliced to obtain a target audio stream, including: obtaining emotional state information of the target account during voice interaction; adjusting the pre-synthesized audio stream based on the emotional state information to obtain an adjusted audio stream; and splicing the adjusted audio stream with the fixed-speech audio stream to obtain a target audio stream.

[0053] Optionally, during voice interaction between the intelligent outbound calling system and the target account, emotion recognition technology can be used to monitor the target account's emotional state in real time, including but not limited to happiness, anxiety, and confusion. This emotional state information can be acquired through, but is not limited to, voice analysis, intonation recognition, and keyword detection, providing important information for subsequent audio stream adjustments. Based on the target account's emotional state information, the pre-synthesized audio stream can be adjusted in real time. For example, if the target account is detected to be anxious, the speed, pitch, and tone of the pre-synthesized audio stream can be adjusted to be slower and gentler to soothe the target account. This dynamic adjustment ensures that the audio content adapts to the target account's emotional state, improving communication effectiveness. After the pre-synthesized audio stream is adjusted, it is seamlessly spliced with the pre-recorded fixed-script audio stream to form the final target audio stream. The fixed-script audio stream contains the unchanging components of the outbound call, such as the greeting and closing remarks, while the adjusted audio stream contains the components that require personalization, such as specific event information and account-specific parameters. The splicing process ensures the coherence and naturalness of call content while also taking into account the emotional state of the target account, thereby enhancing the personalization of outbound service and customer satisfaction. By dynamically adjusting the audio stream, it can better adapt to the communication needs and emotional state of different target accounts, making outbound service more personalized and attentive. This adjustment not only alleviates negative emotions in target accounts but also enhances positive emotions, thereby increasing customer satisfaction with the service and promoting better communication results. Furthermore, real-time emotional state monitoring and dynamic audio stream adjustment can reduce invalid calls and call interruptions caused by miscommunication, thereby improving the overall efficiency and success rate of outbound calling campaigns.

[0054] Through the above steps S102 to S108, it is possible to dynamically obtain and analyze the personalized parameter data of the target account, generate and optimize the target account portrait, and then accurately determine the account characteristics, pre-synthesize the personalized pre-synthesized audio stream, and seamlessly splice it with the fixed speech audio stream to form the target audio stream, thereby achieving highly personalized voice interaction and improving the user experience and satisfaction of outbound calls, thereby solving the technical problem of insufficient personalized services in related technologies in intelligent outbound calls.

[0055] Based on the above embodiments and optional embodiments, the present invention proposes an optional implementation mode: Figure 5 FIG. 1 is a flow chart of an optional outbound voice generation method according to an embodiment of the present invention. Figure 5 As shown, the method includes:

[0056] S1, Data Collection and Analysis: First, collect the personalized parameter data of the target account, including geographic location, preference information, historical interaction records, etc. Through the data analysis algorithm model, conduct in-depth mining of the account's personalized parameter data to identify account characteristics and needs. Specifically, it includes:

[0057] S11, extracting the account history profile information from the account behavior database and cleaning the account profile information data according to information such as account behavior and intention type;

[0058] S12, classify and grade the cleaned account portrait information, and divide the accounts into A, B, C... categories according to their characteristics, and obtain the historical portrait information of multiple categories of accounts as the graded classification information Data.

[0059] S2, Account Profile Construction: Based on the account information collected from the account personalized parameter data, a target account profile is constructed, including information such as the account's preferences, behavioral habits, and preferences. This helps the system better understand the user and provides a basis for personalized parameter broadcasting. Specifically, it includes:

[0060] S21, re-extract features from the hierarchical classification information Data in step S1, and construct a corresponding first account profile group Dg, analyzing the similarity between the current account and other accounts;

[0061] S22, cross-match Dg with historical account portraits, dynamically adjust the content in Dg, and form a new account portrait as the target account portrait.

[0062] S23, combines historical voice and other information to build a more comprehensive account portrait information, updates the adjusted Dg (i.e., target account portrait) to the account behavior library and saves it persistently, and recommends personalized parameter broadcast content that meets the corresponding account preferences.

[0063] S3, Parameter Adjustment Algorithm (ADAL): Utilizes machine learning algorithms and data mining techniques to analyze and model individualized account parameter data, enabling intelligent identification and prediction of account needs. The parameter adjustment algorithm dynamically adjusts broadcast content based on account characteristics. Specifically, it includes:

[0064] S31: Pass the account profile information classified and graded in step S1 and Dg (i.e., the target account profile) adjusted in step S2 to the ADAL parameter adjustment algorithm;

[0065] S32, the ADAL algorithm takes the historical account portraits in the classified and graded account portrait information as the data set Ds, and randomly selects n feature parameters from Ds as n center points.

[0066] S33, respectively calculate the distance Dis from each data point in Ds to the center point, and perform secondary classification on the target account profile based on Dis to obtain Da, and perform data normalization to obtain the parameters Pt required for splicing;

[0067] S34: Cross-check the parameters Pt to be spliced and the recommended broadcast parameters to form the final parameters P to be spliced.

[0068] S4, parameter broadcast strategy optimization: adjust the parameter broadcast strategy in real time based on account feedback and behavior, and continuously optimize the broadcast content. By monitoring the response and interaction of the account, the call effect and service quality are improved. Specifically including:

[0069] S41, adjusting parameters in real time to pre-synthesize the audio stream based on the account's mood, status, and other information during the call;

[0070] S42, temporarily storing the determined parameters to obtain a parameter set Ps, and pre-synthesizing the parameters in Ps to form a parameter-speech stream mapping relationship, and obtaining a pre-synthesized audio stream.

[0071] S5, Application of Speech Synthesis Technology: During actual calls, speech synthesis and audio splicing technologies are used to convert text content into speech and then splice it with personalized parameter audio to achieve voice broadcast of personalized parameters. This improves the naturalness and fluency of calls. Specifically, this includes:

[0072] The pre-synthesized audio stream in step S4 is spliced with the pre-recorded node fixed speech audio stream to form a complete speech audio stream, and the spliced voice stream is broadcast to the target account as a complete speech.

[0073] During a real-time call, read the fixed-speech audio file to obtain the fixed-speech audio stream L1, read the variable slot information, send the slot variable information to the TTS service to obtain the pre-synthesized audio stream L2 corresponding to the slot, and use ffmpeg to splice the fixed-speech audio stream L1 and the pre-synthesized audio stream L2 to form a complete node voice stream, as shown in the following example:

[0074] Complete node script: Dear {honorific title}, hello, we are xxxx, and we are calling to inform you of {event content}…

[0075] The above scripts are configured by the operators. The system generates the content corresponding to the slot according to steps S1 and S2, and uses TTS synthesis technology for pre-synthesis to form a parameter voice stream mapping {'honorific title': voice stream, "activity content": voice stream, other personalized parameters}. Except for the variables, all other parts can be saved separately in the form of recording files. When broadcasting to this node, the system no longer needs to perform TTS synthesis separately. It only needs to read the audio stream at the corresponding position and the audio stream of the parameter mapping, and use ffmpeg to splice two or more audio streams to reduce the time difference between the node fixed script audio and the intermediate variable synthesized audio broadcast caused by TTS real-time synthesis, which affects the account experience.

[0076] It should be noted that in this embodiment, by introducing a personalized parameter insertion announcement feature, the intelligent outbound call system can dynamically adjust announcement content during calls based on the user's specific attributes and needs, thereby providing personalized services for different users. Specifically, the system automatically adjusts call content based on parameters such as the user's location, preferences, and historical interaction records, making the call process more tailored to user needs and improving user experience and satisfaction.

[0077] Through data analysis and machine learning algorithms, we can quickly identify user characteristics and needs and generate personalized call content for each user. At the same time, we can also monitor user feedback and behavior in real time, continuously optimize parameter broadcast strategies, and improve call effectiveness and service quality.

[0078] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. For example, an interface is set up between this system and the relevant user or organization. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving the consent information fed back by the aforementioned user or organization.

[0079] This embodiment also provides an outbound voice generation device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated here. As used below, the terms "module" and "device" may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0080] According to an embodiment of the present invention, there is also provided an embodiment of a device for implementing the above-mentioned outbound voice generation method. Figure 6 : is a structural diagram of an outbound voice generating device according to an embodiment of the present invention. Figure 6As shown, the outbound voice generation device includes: a data acquisition module 600, an account portrait generation module 602, a parameter determination module 604, an audio stream synthesis module 606, and an audio stream splicing module 608, wherein:

[0081] The data acquisition module 600 is used to acquire personalized parameter data of the target account during voice interaction with the target account;

[0082] An account profile generation module 602 , connected to the data acquisition module 600 , is configured to generate a target account profile based on the personalized parameter data;

[0083] The parameter determination module 604 is connected to the account profile generation module 602 and is used to determine the target characteristic parameters corresponding to the target account based on the target account profile;

[0084] The audio stream synthesis module 606 is connected to the parameter determination module 604 and is used to generate a pre-synthesized audio stream that matches the target characteristic parameters;

[0085] The audio stream splicing module 608 is connected to the audio stream synthesis module 606, and is used to splice the pre-synthesized audio stream and the pre-recorded fixed speech audio stream to obtain the target audio stream, and broadcast the target audio stream to the target account.

[0086] In an embodiment of the present invention, a data acquisition module 600 is provided for acquiring personalized parameter data of a target account during voice interaction with the target account; an account portrait generation module 602 is connected to the data acquisition module 600 and is used to generate a target account portrait based on the personalized parameter data; a parameter determination module 604 is connected to the account portrait generation module 602 and is used to determine the target characteristic parameters corresponding to the target account based on the target account portrait; an audio stream synthesis module 606 is connected to the parameter determination module 604 and is used to generate a pre-synthesized audio stream that matches the target characteristic parameters; and an audio stream splicing module 608 is provided. , connected to the audio stream synthesis module 606, is used to splice the pre-synthesized audio stream and the pre-recorded fixed-speech audio stream to obtain the target audio stream, and broadcast the target audio stream to the target account, thereby achieving the purpose of dynamically obtaining and analyzing the personalized parameter data of the target account, generating and optimizing the target account portrait, and then accurately determining the account characteristics, pre-synthesizing the personalized pre-synthesized audio stream, and seamlessly splicing it with the fixed-speech audio stream to form the target audio stream, thereby achieving highly personalized voice interaction and improving the user experience and satisfaction of outbound calls, thereby solving the technical problem of insufficient personalized services in related technologies in intelligent outbound calls.

[0087] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0088] It should be noted that the data acquisition module 600, account profile generation module 602, parameter determination module 604, audio stream synthesis module 606, and audio stream splicing module 608 correspond to steps S102 to S110 in the embodiment. The examples and application scenarios implemented by these modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment. It should be noted that the above modules, as part of the device, can be run on a computer terminal.

[0089] It should be noted that the optional or preferred implementation of this embodiment can be found in the relevant description in the embodiment, which will not be repeated here.

[0090] The above-mentioned outbound voice generation device can also include a processor and a memory. The above-mentioned data acquisition module 600, account portrait generation module 602, parameter determination module 604, audio stream synthesis module 606, audio stream splicing module 608, etc. are all stored in the memory as program modules, and the processor executes the above-mentioned program modules stored in the memory to realize the corresponding functions.

[0091] The processor includes a core, which retrieves corresponding program modules from memory. There can be one or more cores. Memory may include non-permanent memory in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.

[0092] According to an embodiment of the present application, an embodiment of a non-volatile storage medium is also provided. Optionally, in this embodiment, the non-volatile storage medium includes a stored program, wherein when the program is executed, the device containing the non-volatile storage medium is controlled to execute any of the above-mentioned outbound voice generation methods.

[0093] Optionally, in this embodiment, the non-volatile storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group, and the non-volatile storage medium includes a stored program.

[0094] Optionally, when the program is running, the device where the non-volatile storage medium is located is controlled to perform the following functions: in the process of voice interaction with the target account, personalized parameter data of the target account is obtained; based on the personalized parameter data, a portrait of the target account is generated; based on the target account portrait, the target characteristic parameters corresponding to the target account are determined; a pre-synthesized audio stream that matches the target characteristic parameters is generated; the pre-synthesized audio stream and the pre-recorded fixed-speech audio stream are spliced to obtain the target audio stream, and the target audio stream is broadcast to the target account.

[0095] According to an embodiment of the present application, an embodiment of a processor is further provided. Optionally, in this embodiment, the processor is configured to run a program, wherein the program, when running, executes any of the above-mentioned methods for generating an outbound voice.

[0096] According to an embodiment of the present application, an embodiment of a computer program product is also provided. When executed on a data processing device, it is suitable for executing a program that initializes any one of the steps of the above-mentioned outbound voice generation method.

[0097] Optionally, the above-mentioned computer program product, when executed on a data processing device, is suitable for executing a program initialized with the following method steps: obtaining personalized parameter data of the target account during voice interaction with the target account; generating a target account portrait based on the personalized parameter data; determining the target feature parameters corresponding to the target account based on the target account portrait; generating a pre-synthesized audio stream that matches the target feature parameters; splicing the pre-synthesized audio stream and the pre-recorded fixed-speech audio stream to obtain the target audio stream, and broadcasting the target audio stream to the target account.

[0098] An embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and runnable on the processor. When the processor executes the program, the following steps are implemented: in the process of voice interaction with a target account, personalized parameter data of the target account is obtained; based on the personalized parameter data, a target account portrait is generated; based on the target account portrait, target feature parameters corresponding to the target account are determined; a pre-synthesized audio stream matching the target feature parameters is generated; the pre-synthesized audio stream and a pre-recorded fixed-speech audio stream are spliced to obtain a target audio stream, and the target audio stream is broadcast to the target account.

[0099] The above sequence of the embodiments of the present invention is for description only and does not represent the superiority or inferiority of the embodiments.

[0100] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0101] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the above modules can be a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, modules or indirect coupling or communication connection of modules, which can be electrical or other forms.

[0102] The modules described above as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.

[0103] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules.

[0104] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable non-volatile storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned non-volatile storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, and other media that can store program codes.

[0105] The above are only preferred embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for generating an outbound voice, characterized in that: include: During voice interaction with a target account, obtaining personalized parameter data of the target account; Generate a target account profile based on the personalized parameter data; Determining target characteristic parameters corresponding to the target account based on the target account portrait; generating a presynthesized audio stream matching the target characteristic parameters; The pre-synthesized audio stream and the pre-recorded fixed speech audio stream are spliced to obtain a target audio stream, and the target audio stream is broadcast to the target account.

2. The method according to claim 1, characterized in that In a case where the personalized parameter data includes account information of the target account, generating a target account profile based on the personalized parameter data includes: Obtain multiple groups of account history profile information included in the user behavior database; Classifying the multiple groups of account history profile information according to account characteristics to obtain multiple categories of account history profile information; Determining a first account profile group that matches the personalized parameter data from the multiple categories of account history profile information; Cross-matching the first account profile with the historical account profile information of the multiple categories to obtain a second account profile group; Based on the second account portrait group, the target account portrait is generated.

3. The method according to claim 2, characterized in that In a case where the personalized parameter data further includes historical interaction data of the target account, generating the target account profile based on the second account profile group includes: Based on the second account portrait group and the historical interaction data, the target account portrait is obtained.

4. The method according to claim 2, characterized in that The determining, based on the target account profile, target characteristic parameters corresponding to the target account includes: Using multiple feature parameters selected from the multiple categories of account history profile information as multiple initial clustering centers; Based on the multiple initial cluster centers, clustering the multiple categories of account history profile information to obtain clustering results; The target account portrait is classified based on the clustering result to obtain the target feature parameters.

5. The method according to claim 4, characterized in that The classifying the target account profile based on the clustering result to obtain the target feature parameter includes: Classifying the target account profile based on the clustering result to obtain a first feature parameter; Obtaining the recommended broadcast parameters during the voice interaction process; The first characteristic parameter and the recommended broadcast parameter are cross-compared to obtain the target characteristic parameter.

6. The method according to claim 1, characterized in that The step of splicing the pre-synthesized audio stream and the pre-recorded fixed speech audio stream to obtain a target audio stream includes: Reading slot information in the fixed speech audio stream; The pre-synthesized audio stream is spliced to the fixed speech audio stream according to the slot information to obtain the target audio stream.

7. The method according to any one of claims 1 to 6, characterized in that The step of splicing the pre-synthesized audio stream and the pre-recorded fixed speech audio stream to obtain a target audio stream includes: Acquiring emotional state information of the target account during the voice interaction process; Adjusting the pre-synthesized audio stream based on the emotional state information to obtain an adjusted audio stream; The adjusted audio stream and the fixed speech audio stream are spliced to obtain the target audio stream.

8. An outbound voice generation device, characterized in that: include: A data acquisition module, configured to acquire personalized parameter data of a target account during voice interaction with the target account; An account profile generation module, configured to generate a target account profile based on the personalized parameter data; A parameter determination module, configured to determine target characteristic parameters corresponding to the target account based on the target account profile; An audio stream synthesis module, configured to generate a pre-synthesized audio stream matching the target characteristic parameters; The audio stream splicing module is used to splice the pre-synthesized audio stream and the pre-recorded fixed speech audio stream to obtain a target audio stream, and broadcast the target audio stream to the target account.

9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executed by the outbound voice generation method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The invention comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the outbound voice generation method according to any one of claims 1 to 7.