Telephone collection method, device and voice robot with adaptive dialect voice switching
By identifying the old dialect used by the user during the telephone call and switching to the corresponding voice synthesis model, the problem that the voice robot cannot recognize the user's dialect is solved, achieving a more efficient power bill call effect.
Patent Information
- Application Number
- CN202210947087.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2042-08-08
AI Technical Summary
Existing voice robots cannot accurately identify the dialect used by users when urging electricity bills, resulting in inefficient payment and making it difficult for users to understand the content of the payment.
The detection statement is sent through a pre-constructed new dialect speech synthesis model, identify the old dialect of the user terminal, and use the old dialect speech synthesis model to send transition and call-up statements to ensure that the user understands the call-up content.
It improves the success rate and efficiency of telephone call requests, ensures that users are aware of the call request information, and reduces the possibility of failure of call requests.
Smart Images

Figure CN115376489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech recognition technology, and in particular to a telephone payment collection method, device and voice robot capable of adaptively switching dialect voices. Background Art
[0002] In most power supply stations, some users fail to pay or owe electricity bills every month. If the power is forcibly cut off for these users, it may cause unpredictable consequences for them and even affect the relationship between electricity users and the power grid company. Therefore, manual telephone calls are mainly used to urge payment.
[0003] In recent years, to improve the efficiency of payment collection calls, voice robots using intelligent voice technologies such as TTS (Text To Speech) have been introduced to replace manual phone collection. However, a large portion of the population of large cities has migrated from other areas, resulting in a mix of dialects within a city. While existing TTS technology can personalize dialects, voice robots do not know the user's familiar dialect during the initial collection call and can only communicate in Mandarin. This cannot guarantee that the user can fully understand Mandarin and understand the collection content, which can easily lead to collection failures and hinder further improvement in payment collection efficiency. Summary of the Invention
[0004] In order to overcome the defects of the existing technology, the present invention provides a telephone collection method, device and voice robot that adaptively switches dialect voices. The method, device and voice robot can first accurately identify the dialect used by the user during the telephone collection process, and then use the dialect used by the user to inform the user of the collection content, effectively ensuring that the user knows the collection content, improving the success rate of telephone collection, and further improving the efficiency of telephone collection.
[0005] In order to solve the above technical problems, in a first aspect, an embodiment of the present invention provides a telephone collection method with adaptive dialect switching, comprising:
[0006] When establishing a voice communication channel with a user terminal, synthesizing a probe sentence using a pre-built speech synthesis model of a new dialect, and sending the probe sentence to the user terminal;
[0007] receiving a reply sentence sent by the user terminal, and identifying the old dialect corresponding to the reply sentence as a target dialect using a pre-built classification model of the old dialect;
[0008] synthesizing a first transition sentence using a speech synthesis model of the new dialect, synthesizing a second transition sentence using a speech synthesis model of the old dialect corresponding to the target dialect, obtaining a transition sentence based on the first transition sentence and the second transition sentence, and sending the transition sentence to the user terminal;
[0009] A payment reminder sentence is synthesized by using the speech synthesis model of the old dialect, and the payment reminder sentence is sent to the user terminal.
[0010] Furthermore, when establishing a voice communication channel with a user terminal, synthesizing a probe sentence using a pre-built speech synthesis model of a new dialect, and before sending the probe sentence to the user terminal, the method further includes:
[0011] Collecting corpora for the city where the user terminal is located, inputting each corpus into the classification model of the old dialect, and obtaining an old dialect probability set for each corpus; wherein the old dialect probability set for the corpus includes the probability that the corpus belongs to various old dialects;
[0012] When at least two probability values in the old dialect probability set of any of the corpora are within a preset grayscale range, the corresponding corpora are used as new dialect corpora, and training corpora are selected from all the new dialect corpora. Based on all the training corpora, a speech synthesis model of the new dialect is trained.
[0013] Furthermore, when establishing a voice communication channel with a user terminal, synthesizing a probe sentence using a pre-built speech synthesis model of a new dialect and sending the probe sentence to the user terminal includes:
[0014] When establishing a voice communication channel with the user terminal, determining whether the user terminal meets the Mandarin conversation conditions based on the user information of the user terminal;
[0015] When the user terminal does not meet the Mandarin dialogue condition, the probe sentence is synthesized by using the speech synthesis model of the new dialect, and the probe sentence is sent to the user terminal.
[0016] Furthermore, the detection sentence is synthesized by using a pre-built speech synthesis model of the new dialect, specifically:
[0017] Writing the user information of the user terminal into a pre-stored detection template to generate a detection text;
[0018] The detection text is converted into the new dialect speech by using the speech synthesis model of the new dialect to obtain the detection sentence.
[0019] Furthermore, after sending the detection statement to the user terminal, the method further includes:
[0020] When the reply sentence is not received within a preset period of time, a new probe sentence is synthesized using the speech synthesis model of the new dialect, and the new probe sentence is sent to the user terminal.
[0021] Furthermore, the transition statement obtained according to the first transition statement and the second transition statement is specifically:
[0022] The first transition statement and the second transition statement are linearly fused to obtain the transition statement.
[0023] Furthermore, the linear fusion of the first transition statement and the second transition statement to obtain the transition statement is specifically:
[0024] performing a weighted summation of the first transition statement and the second transition statement according to the weights of the first transition statement and the second transition statement to obtain a fused transition statement;
[0025] The fused transition sentence is smoothed to obtain the transition sentence.
[0026] Furthermore, the payment collection statement is synthesized by the speech synthesis model of the old dialect, specifically:
[0027] Writing the payment reminder information of the user terminal into a pre-stored payment reminder template to generate a payment reminder text;
[0028] The payment demand text is converted into target dialect speech through the speech synthesis model of the old dialect to obtain the payment demand sentence.
[0029] In a second aspect, an embodiment of the present invention provides a telephone collection device capable of adaptively switching between dialect voices, comprising:
[0030] a probe sentence sending module, configured to synthesize a probe sentence using a pre-built speech synthesis model of a new dialect when establishing a voice communication channel with a user terminal, and send the probe sentence to the user terminal;
[0031] a target dialect identification module, configured to receive a reply sentence sent by the user terminal and identify the old dialect corresponding to the reply sentence as a target dialect using a pre-built classification model of the old dialect;
[0032] a transition sentence sending module, configured to synthesize a first transition sentence using a speech synthesis model of the new dialect, synthesize a second transition sentence using a speech synthesis model of the old dialect corresponding to the target dialect, obtain a transition sentence based on the first transition sentence and the second transition sentence, and send the transition sentence to the user terminal;
[0033] The payment reminder statement sending module is used to synthesize the payment reminder statement through the speech synthesis model of the old dialect and send the payment reminder statement to the user terminal.
[0034] In a third aspect, an embodiment of the present invention provides a voice robot, wherein the voice robot is internally provided with the above-mentioned telephone collection device capable of adaptively switching dialect voices.
[0035] The embodiments of the present invention have the following beneficial effects:
[0036] When establishing a voice communication channel with a user terminal, a detection sentence is synthesized by a pre-built speech synthesis model of the new dialect, and the detection sentence is sent to the user terminal; a reply sentence sent by the user terminal is received, and the old dialect corresponding to the reply sentence is identified as the target dialect by a pre-built classification model of the old dialect; a first transition sentence is synthesized by the speech synthesis model of the new dialect, a second transition sentence is synthesized by the speech synthesis model of the old dialect corresponding to the target dialect, and a transition sentence is obtained based on the first transition sentence and the second transition sentence, and the transition sentence is sent to the user terminal; a collection demand sentence is synthesized by the speech synthesis model of the old dialect, and the collection demand sentence is sent to the user terminal, thereby completing the telephone collection demand. Compared with the existing technology, the embodiments of the present invention first use the new dialect detection to obtain the reply statement sent by the user terminal to identify the old dialect used by the user when establishing a voice communication channel with the user terminal, and then transition and switch to use the old dialect used by the user to send a collection statement to the user terminal to inform the user of the collection content. In the process of telephone collection, it is possible to first accurately identify the dialect used by the user, and then use the dialect used by the user to inform the user of the collection content, effectively ensuring that the user knows the collection content, improving the success rate of telephone collection, and further improving the efficiency of telephone collection. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 1. A flowchart of a telephone payment collection method with adaptively switching dialect voices in a first embodiment of the present invention;
[0038] Figure 2 Schematic diagram of the structure of the classification model of the old dialect used as an example in the first embodiment of the present invention;
[0039] Figure 3 This is a schematic structural diagram of a telephone payment collection device capable of adaptively switching between dialect voices in a second embodiment of the present invention;
[0040] Figure 4 Schematic diagram of the structure of a voice robot in the third embodiment of the present invention. DETAILED DESCRIPTION
[0041] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0042] It should be noted that the step numbers in this article are only for the convenience of explaining the specific embodiment and do not limit the order in which the steps are executed. The method provided in this embodiment can be executed by relevant terminal devices, and the following description uses a voice robot as the execution subject as an example.
[0043] like Figure 1 As shown, the first embodiment provides a telephone collection method with adaptively switching dialect voices, including steps S1 to S4:
[0044] S1. When establishing a voice communication channel with a user terminal, synthesize a probe sentence using a pre-built speech synthesis model of a new dialect and send the probe sentence to the user terminal;
[0045] S2. Receive a reply sentence sent by the user terminal, and identify the old dialect corresponding to the reply sentence as the target dialect using a pre-built old dialect classification model;
[0046] S3. Synthesizing a first transition sentence using a speech synthesis model for the new dialect, synthesizing a second transition sentence using a speech synthesis model for the old dialect corresponding to the target dialect, obtaining a transition sentence based on the first transition sentence and the second transition sentence, and sending the transition sentence to the user terminal.
[0047] S4. Synthesize a payment reminder statement using the old dialect speech synthesis model and send the payment reminder statement to the user terminal.
[0048] It should be noted that user terminals include mobile phones, landlines and other communication devices capable of voice interaction held by users who have not paid or are in arrears with their electricity bills.
[0049] When people from different regions work, live, and study in the same city, it's a process of cultural integration. The dialects used by people from different regions also converge in this process, forming a new dialect that lies somewhere between the old dialects, making it somewhat understandable to both sides. For example, many elderly people in Guangzhou are accustomed to speaking Cantonese. When they go to the market to buy groceries, they might encounter people from Hunan or Guizhou. To keep things simple, they more or less imitate their dialects, creating a language commonly known as "bo dong gua."
[0050] Based on this situation, this embodiment pre-constructs speech synthesis models for multiple old dialects, such as the old dialect TTS model corresponding to Cantonese and the old dialect TTS model corresponding to Minnan, constructs a classification model for old dialects, and constructs speech synthesis models for new dialects, such as the new dialect TTS model, to conduct three-stage telephone collection.
[0051] Among them, the old dialect TTS model can directly select the old dialect TTS model trained by a third party. Figure 2 As shown, the classification model of the old dialect has a backbone network and multiple branch networks. The backbone network includes a TDNN (time-delay neural network) and two Res_block (residual blocks) for extracting common features. Each branch network includes a Conv (convolutional layer) and an FC (fully connected layer) for outputting the probability that the audio belongs to Cantonese, Minnan or another dialect. Specifically, some conventional features are extracted from the audio signal, such as Mel cepstral coefficient (MFCC), zero-crossing rate and short-time energy (STE). TDNN converts these features from two-dimensional to three-dimensional to facilitate subsequent Res_block and Conv processing, and finally FC outputs the probability.
[0052] As an example, in step S1, a user terminal's phone number is dialed. When a voice communication channel is established with the user terminal, a probe sentence is synthesized using a pre-built speech synthesis model in the new dialect and sent to the user terminal. The probe sentence primarily serves to greet the user and does not involve any payment reminders.
[0053] It is understandable that synthesizing the detection sentence through the speech synthesis model of the new dialect and sending the detection sentence to the user terminal can take into account that most users can understand the new dialect to a certain extent, ensuring that the user knows the detection content and responds.
[0054] In step S2, a reply sentence sent by the user terminal is received, and the reply sentence is input into a pre-built classification model of the old dialect to obtain the probability that the reply sentence belongs to various old dialects, and the old dialect with the highest probability is selected as the target dialect.
[0055] In step S3, a first transitional sentence is synthesized using the speech synthesis model for the new dialect, and a second transitional sentence is synthesized using the speech synthesis model for the old dialect corresponding to the target dialect. The first and second transitional sentences are then merged to generate a transitional sentence, which is then sent to the user terminal. The transitional sentence primarily serves to transition from the new dialect to the old dialect and does not involve payment reminders.
[0056] In step S4, a payment reminder statement is synthesized using the old dialect speech synthesis model and sent to the user terminal to inform the user of the payment reminder content, completing the telephone payment reminder. The main function of the payment reminder statement is to inform the user of the payment reminder content and remind the user to pay as soon as possible.
[0057] This embodiment first uses the new dialect detection to obtain the reply statement sent by the user terminal to identify the old dialect used by the user when establishing a voice communication channel with the user terminal, and then transitions to use the old dialect used by the user to send a collection statement to the user terminal to inform the user of the collection content. In this way, the dialect used by the user can be accurately identified during the telephone collection process, and the user can be informed of the collection content in the dialect used by the user, which effectively ensures that the user knows the collection content, improves the success rate of telephone collection, and further improves the efficiency of telephone collection.
[0058] In a preferred embodiment, when establishing a voice communication channel with a user terminal, a detection sentence is synthesized by a pre-constructed speech synthesis model of a new dialect, and before sending the detection sentence to the user terminal, it also includes: collecting corpora for the city where the user terminal is located, inputting each corpus into the classification model of the old dialect respectively, and obtaining the old dialect probability set of each corpus; wherein the old dialect probability set of the corpus includes the probability that the corpus belongs to various old dialects; when there are at least two probability values in the old dialect probability set of any corpus within a preset grayscale range, the corresponding corpus is used as the new dialect corpus, and training corpora are selected from all the new dialect corpora, and the speech synthesis model of the new dialect is trained based on all the training corpora.
[0059] As an example, in the process of building the classification model of the old dialect, pure old dialects are selected from some public old dialect corpora to train the classification model of the old dialect, and the loss function is Softmax.
[0060] In the process of constructing the speech synthesis model of the new dialect, considering that the new dialect is difficult to define without a definite standard, a fuzzy classification method is used to collect new dialect data.
[0061] Corpus data is collected from various channels in the city where the user terminal is located, such as the customer service center of the power grid, the customer service telephone number in the business hall, the collection telephone number, etc., and each corpus is input into the trained old dialect classification model to obtain the old dialect probability set of each corpus. The old dialect probability set of the corpus includes the probability of the corpus belonging to various old dialects.
[0062] For each corpus, it is determined whether the value of each probability in the old dialect probability set of the corpus is within the preset grayscale range, such as 0.35-0.65. It can be understood that if the value of the probability that the corpus belongs to Cantonese is within the preset grayscale range, it means that the confidence that the corpus belongs to Cantonese is average, and it is difficult to define whether the corpus belongs to Cantonese. If all the probability values in the old dialect probability set of the corpus are not within the preset grayscale range, then the corpus is considered to belong to the old dialect, and the old dialect with the highest probability is selected as the dialect corresponding to the corpus; if at least two probability values in the old dialect probability set of the corpus are within the preset grayscale range, then the corpus is considered to belong to the new dialect, and its evolution process is from the old dialect with high probability to the old dialect with low probability.
[0063] Based on the census results of the city where the user terminal is located, the registered residence of the permanent residents in the city is analyzed to determine the regional distribution of the permanent population. By associating the regions with dialects, the distribution probability of various old dialects in the city can be obtained. The distribution probability of an old dialect in a city can be calculated by the percentage of the registered population speaking the old dialect among the permanent population of the city.
[0064] For different old dialects, k-means can be used for clustering to obtain old dialects with a wider range and closer pronunciation. For example, the Cantonese in northern Guangdong has some Hakka accents, but is basically the same as the Cantonese in Guangzhou and Foshan. They can be clustered and considered as one old dialect.
[0065] After the aggregation of old dialects, only two commonly used old dialects remain. For example, in Guangzhou, Cantonese and Mandarin are spoken, and the evolutionary direction is from Cantonese to Mandarin. In this case, the two old dialects with the highest distribution probability are selected. The evolutionary process is from the old dialect with high distribution probability to the old dialect with low distribution probability. Based on this evolutionary process, appropriate training data is selected from all new dialect corpora to train a speech synthesis model for the new dialect, such as a new dialect TTS model.
[0066] This embodiment helps to further improve the efficiency of telephone collection by reasonably selecting new dialect materials to train the speech synthesis model of the new dialect.
[0067] In a preferred embodiment, when establishing a voice communication channel with a user terminal, a probe sentence is synthesized by a pre-constructed speech synthesis model of a new dialect, and the probe sentence is sent to the user terminal, including: when establishing a voice communication channel with the user terminal, judging whether the user terminal meets the Mandarin conversation conditions based on the user information of the user terminal; when the user terminal does not meet the Mandarin conversation conditions, synthesizing the probe sentence by a speech synthesis model of the new dialect, and sending the probe sentence to the user terminal.
[0068] As an example, considering that after the implementation of modern education, the younger generation can directly use Mandarin for conversation, the Mandarin conversation condition can be pre-defined as the user's age not exceeding 45 years old.
[0069] When establishing a voice communication channel with a user terminal, the user information of the user terminal is obtained and the user age is extracted from the user information. If the user age is over 45 years old, it is determined that the user terminal does not meet the conditions for Mandarin conversation. A detection sentence is synthesized using the speech synthesis model of the new dialect and sent to the user terminal. If the user age is not more than 45 years old, it is determined that the user terminal meets the conditions for Mandarin conversation and Mandarin can be directly used for voice interaction.
[0070] This embodiment adaptively switches to a dialect voice to make payment reminders over the phone when the user terminal does not meet the Mandarin conversation conditions. That is, the new dialect is first detected to obtain the reply statement sent by the user terminal to identify the old dialect used by the user, and then the old dialect used by the user is transitionally switched to send a payment reminder statement to the user terminal to inform the user of the payment content. When the user terminal meets the Mandarin conversation conditions, Mandarin voice is directly used to make payment reminders over the phone, which is conducive to further improving the efficiency of payment reminders over the phone.
[0071] In a preferred embodiment, the synthesis of the detection sentence by the pre-constructed speech synthesis model of the new dialect is specifically as follows: the user information of the user terminal is written into the pre-stored detection template to generate the detection text; and the detection text is converted into the new dialect speech by the speech synthesis model of the new dialect to obtain the detection sentence.
[0072] As an example, the main function of the detection sentence is to greet the user and does not involve collection of payment. The user information of the user terminal is obtained, the user name, user gender, etc. are extracted from the user information, a detection template is selected from the detection template database, the extracted user information such as user name and user gender are written into the detection template, a detection text is generated, and the detection text is converted into the new dialect speech through the new dialect speech synthesis model to obtain the detection sentence. For example, if the selected detection template is "user's last name + Mr. / Ms. Hello, please introduce yourself", the final detection sentence is "Dear Mr. Liu, hello, I am Xiaomei, a customer service representative of Guangzhou Power Grid. I am sorry to bother you."
[0073] By pre-designing and storing the detection template, this embodiment can directly write the user information of the user terminal into the pre-stored detection template, quickly generate the detection text, and is conducive to further improving the efficiency of telephone collection.
[0074] In a preferred embodiment, after sending the probe sentence to the user terminal, the method further includes: when no reply sentence is received within a preset time period, synthesizing a new probe sentence through a speech synthesis model of the new dialect, and sending the new probe sentence to the user terminal.
[0075] As an example, after sending a probe statement to the user terminal, it is continuously monitored whether a reply statement sent by the user terminal is received. If no reply statement is received within a preset time period, a new probe statement is synthesized through the speech synthesis model of the new dialect. Specifically, another probe template is selected from the probe template database, and the extracted user information such as user name and user gender is written into the other probe template to generate a new probe text. The new probe text is converted into the new dialect speech through the speech synthesis model of the new dialect to obtain a new probe statement, and the new probe statement is sent to the user terminal. The above operations are repeated until a reply statement is received or communication with the user terminal is disconnected.
[0076] This embodiment synthesizes a new detection sentence through a speech synthesis model of a new dialect when no reply sentence is received from the user terminal within a preset time period, and sends the new detection sentence to the user terminal. This ensures that the reply sentence sent by the user terminal is obtained to accurately identify the dialect used by the user, which is conducive to further improving the efficiency of telephone collection.
[0077] In a preferred embodiment, obtaining the transition statement based on the first transition statement and the second transition statement is specifically: linearly fusing the first transition statement and the second transition statement to obtain the transition statement.
[0078] In a preferred embodiment, the linear fusion of the first transition statement and the second transition statement to obtain the transition statement is specifically as follows: according to the weights of the first transition statement and the second transition statement, the first transition statement and the second transition statement are weighted summed to obtain the fused transition statement; and the fused transition statement is smoothed to obtain the transition statement.
[0079] As an example, the primary function of a transition statement is to transition from the new dialect to the old dialect and does not involve any demand for payment. The speech signal of the first transition statement has a weight that decreases over time, while the speech signal of the second transition statement has a weight that increases over time. Based on the current weights of the first and second transition statements, a weighted sum of the first and second transition statements is performed to produce a fused transition statement. This fused transition statement is then smoothed to produce a transition statement.
[0080] This embodiment obtains a transition sentence by linearly fusing the first transition sentence and the second transition sentence, thereby ensuring smooth switching from the new dialect to the old dialect used by the user, that is, the target dialect.
[0081] In a preferred embodiment, the synthesis of the collection statement by the speech synthesis model of the old dialect is specifically as follows: writing the collection information of the user terminal into a pre-stored collection template to generate a collection text; converting the collection text into the target dialect speech by the speech synthesis model of the old dialect to obtain the collection statement.
[0082] For example, the primary purpose of a payment reminder is to inform users of the payment details and remind them to pay as soon as possible. The payment reminder information from the user terminal is obtained, a payment reminder template is selected from a payment reminder template database, and the payment reminder information from the user terminal is written into the template to generate a payment reminder text. The payment reminder text is then converted into the target dialect using a speech synthesis model in the old dialect to produce the payment reminder sentence.
[0083] For example, there are three collection templates in the collection template database, as follows:
[0084] Payment reminder template 1 is "Dear xx, you owe xx yuan for electricity. To prevent power outages, please pay your electricity bill as soon as possible at the local power supply business office or pay online."
[0085] Payment reminder template 2 is "Dear xx, you owe xx yuan for electricity. Account number xx, account name xx. To prevent power outages, please go to the local State Grid power supply business office as soon as possible to pay the electricity bill. Thank you for your cooperation."
[0086] Collection template 3 is "Dear xx, your total electricity consumption in xx month of xx year is xx, the total electricity bill is xx, the current balance is 0.00 yuan, and the amount of outstanding bills is xx. In order to prevent power outages from causing inconvenience to your life, please pay the electricity bill as soon as possible. Please call xx for details."
[0087] This embodiment pre-designs and stores the collection reminder template, and can directly write the collection reminder information of the user terminal into the pre-stored collection reminder template to quickly generate the collection reminder text, which is conducive to further improving the efficiency of telephone collection reminders.
[0088] Based on the same inventive concept as the first embodiment, the second embodiment provides Figure 3 The device for collecting payments by telephone, which is capable of adaptively switching between dialect voices, comprises: a detection sentence sending module 21, which is used to synthesize a detection sentence through a pre-built speech synthesis model of a new dialect when establishing a voice communication channel with a user terminal, and send the detection sentence to the user terminal; a target dialect recognition module 22, which is used to receive a reply sentence sent by the user terminal, and identify the old dialect corresponding to the reply sentence as the target dialect through a pre-built classification model of the old dialect; a transition sentence sending module 23, which is used to synthesize a first transition sentence through a speech synthesis model of the new dialect, and synthesize a second transition sentence through a speech synthesis model of the old dialect corresponding to the target dialect, and obtain a transition sentence based on the first transition sentence and the second transition sentence, and send the transition sentence to the user terminal; a collection statement sending module 24, which is used to synthesize a collection statement through a speech synthesis model of the old dialect, and send the collection statement to the user terminal.
[0089] In a preferred embodiment, the telephone collection device that adaptively switches dialect voices also includes: a speech synthesis model construction module for the new dialect, which is used to synthesize a detection sentence through a pre-constructed speech synthesis model of the new dialect when establishing a voice communication channel with the user terminal, and before sending the detection sentence to the user terminal, collect corpus for the city where the user terminal is located, input each corpus into the classification model of the old dialect respectively, and obtain the old dialect probability set of each corpus; wherein the old dialect probability set of the corpus includes the probability that the corpus belongs to various old dialects; when there are at least two probability values in the old dialect probability set of any corpus within a preset grayscale range, the corresponding corpus is used as the new dialect corpus, and training corpus is selected from all the new dialect corpus, and the speech synthesis model of the new dialect is trained based on all the training corpus.
[0090] In a preferred embodiment, when establishing a voice communication channel with a user terminal, a probe sentence is synthesized by a pre-constructed speech synthesis model of a new dialect, and the probe sentence is sent to the user terminal, including: when establishing a voice communication channel with the user terminal, judging whether the user terminal meets the Mandarin conversation conditions based on the user information of the user terminal; when the user terminal does not meet the Mandarin conversation conditions, synthesizing the probe sentence by a speech synthesis model of the new dialect, and sending the probe sentence to the user terminal.
[0091] In a preferred embodiment, the synthesis of the detection sentence by the pre-constructed speech synthesis model of the new dialect is specifically as follows: the user information of the user terminal is written into the pre-stored detection template to generate the detection text; and the detection text is converted into the new dialect speech by the speech synthesis model of the new dialect to obtain the detection sentence.
[0092] In a preferred embodiment, the detection sentence sending module 21 is also used to synthesize a new detection sentence through a speech synthesis model of the new dialect and send the new detection sentence to the user terminal when no reply sentence is received within a preset time period after the detection sentence is sent to the user terminal.
[0093] In a preferred embodiment, obtaining the transition statement based on the first transition statement and the second transition statement is specifically: linearly fusing the first transition statement and the second transition statement to obtain the transition statement.
[0094] In a preferred embodiment, the linear fusion of the first transition statement and the second transition statement to obtain the transition statement is specifically as follows: according to the weights of the first transition statement and the second transition statement, the first transition statement and the second transition statement are weighted summed to obtain the fused transition statement; and the fused transition statement is smoothed to obtain the transition statement.
[0095] In a preferred embodiment, the synthesis of the collection statement by the speech synthesis model of the old dialect is specifically as follows: writing the collection information of the user terminal into a pre-stored collection template to generate a collection text; converting the collection text into the target dialect speech by the speech synthesis model of the old dialect to obtain the collection statement.
[0096] Based on the same inventive concept as the second embodiment, the third embodiment provides Figure 4 A voice robot is shown, in which a telephone collection device capable of adaptively switching between dialect voices as described in the second embodiment is provided, and can achieve the same beneficial effects as the second embodiment.
[0097] In summary, the implementation of the embodiments of the present invention has the following beneficial effects:
[0098] The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone. The invention provides a method for collecting payment by telephone.
[0099] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
[0100] Those skilled in the art will appreciate that all or part of the processes in the above embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
Claims
1. A telephone collection method with adaptive dialect voice switching, characterized in that: include: When establishing a voice communication channel with a user terminal, synthesizing a probe sentence using a pre-built speech synthesis model of a new dialect, and sending the probe sentence to the user terminal; receiving a reply sentence sent by the user terminal, and identifying the old dialect corresponding to the reply sentence as a target dialect using a pre-built classification model of the old dialect; synthesizing a first transition sentence using a speech synthesis model of the new dialect, synthesizing a second transition sentence using a speech synthesis model of the old dialect corresponding to the target dialect, obtaining a transition sentence based on the first transition sentence and the second transition sentence, and sending the transition sentence to the user terminal; A payment reminder sentence is synthesized by using the speech synthesis model of the old dialect, and the payment reminder sentence is sent to the user terminal.
2. The telephone collection method with adaptive dialect switching as claimed in claim 1, characterized in that: When establishing a voice communication channel with a user terminal, synthesizing a probe sentence using a pre-built speech synthesis model of a new dialect, and before sending the probe sentence to the user terminal, the method further includes: Collecting corpora for the city where the user terminal is located, inputting each corpus into the classification model of the old dialect, and obtaining an old dialect probability set for each corpus; wherein the old dialect probability set for the corpus includes the probability that the corpus belongs to various old dialects; When at least two probability values in the old dialect probability set of any of the corpora are within a preset grayscale range, the corresponding corpora are used as new dialect corpora, and training corpora are selected from all the new dialect corpora. Based on all the training corpora, a speech synthesis model of the new dialect is trained.
3. The telephone collection method with adaptive dialect switching as claimed in claim 1, characterized in that: When establishing a voice communication channel with a user terminal, synthesizing a probe sentence using a pre-built speech synthesis model of a new dialect, and sending the probe sentence to the user terminal, includes: When establishing a voice communication channel with the user terminal, determining whether the user terminal meets the Mandarin conversation conditions based on the user information of the user terminal; When the user terminal does not meet the Mandarin dialogue condition, the probe sentence is synthesized by using the speech synthesis model of the new dialect, and the probe sentence is sent to the user terminal.
4. The telephone collection method with adaptive dialect switching as claimed in claim 1, characterized in that: The detection sentence is synthesized by using the pre-built speech synthesis model of the new dialect, specifically: Writing the user information of the user terminal into a pre-stored detection template to generate a detection text; The detection text is converted into the new dialect speech by using the speech synthesis model of the new dialect to obtain the detection sentence.
5. The telephone collection method with adaptive dialect switching as claimed in claim 1, characterized in that: After sending the detection statement to the user terminal, the method further includes: When the reply sentence is not received within a preset period of time, a new probe sentence is synthesized using the speech synthesis model of the new dialect, and the new probe sentence is sent to the user terminal.
6. The telephone collection method with adaptive dialect switching as claimed in claim 1, characterized in that: The transition statement obtained according to the first transition statement and the second transition statement is specifically: The first transition statement and the second transition statement are linearly fused to obtain the transition statement.
7. The telephone collection method with adaptive dialect switching as claimed in claim 6, characterized in that: The linear fusion of the first transition statement and the second transition statement to obtain the transition statement is specifically: performing a weighted summation of the first transition statement and the second transition statement according to the weights of the first transition statement and the second transition statement to obtain a fused transition statement; The fused transition sentence is smoothed to obtain the transition sentence.
8. The telephone collection method with adaptive dialect switching as claimed in claim 1, characterized in that: The payment collection statement synthesized by the speech synthesis model of the old dialect is specifically: Writing the payment reminder information of the user terminal into a pre-stored payment reminder template to generate a payment reminder text; The payment demand text is converted into target dialect speech through the speech synthesis model of the old dialect to obtain the payment demand sentence.
9. A telephone collection device capable of adaptively switching between dialect voices, characterized in that: include: a probe sentence sending module, configured to synthesize a probe sentence using a pre-built speech synthesis model of a new dialect when establishing a voice communication channel with a user terminal, and send the probe sentence to the user terminal; a target dialect identification module, configured to receive a reply sentence sent by the user terminal and identify the old dialect corresponding to the reply sentence as a target dialect using a pre-built classification model of the old dialect; a transition sentence sending module, configured to synthesize a first transition sentence using a speech synthesis model of the new dialect, synthesize a second transition sentence using a speech synthesis model of the old dialect corresponding to the target dialect, obtain a transition sentence based on the first transition sentence and the second transition sentence, and send the transition sentence to the user terminal; The payment reminder statement sending module is used to synthesize the payment reminder statement through the speech synthesis model of the old dialect and send the payment reminder statement to the user terminal.
10. A voice robot, characterized in that: The voice robot is internally provided with a telephone collection device capable of adaptively switching dialect voices as described in claim 9.
Citation Information
Patent Citations
Service guiding method and apparatus thereof
CN107393530A
Speech synthesis model training method and speech synthesis method
CN113450756A