Methods and systems for broadcasting sales messages for financial products, and computer equipment.
By adding attribute labels to polyphonic characters and numbers, and combining hidden Markov models and natural language processing models, the problem of inaccurate pronunciation of polyphonic characters and numbers in online sales of financial products has been solved, achieving higher pronunciation accuracy and improved customer experience.
Patent Information
- Application Number
- CN202211242388.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-10-11
AI Technical Summary
In the online sales of financial products, the lack of unified rules for the pronunciation of polyphonic characters and numbers affects the compliance of the sales process and the customer experience.
By adding attribute labels to polyphonic characters and numbers, and combining hidden Markov models and natural language processing models, TTS technology is used for voice broadcasting to ensure pronunciation accuracy.
This improved the standardization and normalization of the sales process for financial products, enhanced customer experience, and reduced manual maintenance costs.
Smart Images

Figure CN115547295B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular to a method and system for broadcasting voice messages for the sale of financial products, as well as computer equipment. Background Technology
[0002] With the rapid development of mobile internet technology, the banking and other financial industries are gradually moving towards online and intelligent operations. For higher-risk wealth management products, remote face-to-face verification, including risk disclosure, is still required during online sales, as mandated by regulations. To standardize online sales processes, banks often develop standardized sales and risk disclosure scripts, which are broadcast during the sales process and require customer confirmation. However, these scripts may contain polyphonic characters or numbers. Currently, there are no unified rules for the pronunciation of polyphonic characters; the judgment of pronunciation is usually based on daily experience or conventional usage. Therefore, the correct pronunciation of polyphonic characters or numbers, and its accuracy, significantly impacts the compliance of the sales process and the customer experience. Summary of the Invention
[0003] In view of this, it is necessary to provide a method, system, and computer equipment for broadcasting the sales voice of financial products, which can effectively improve the accuracy of pronunciation.
[0004] In a first aspect, embodiments of this application provide a method for broadcasting sales voice messages for financial products, the method comprising:
[0005] Read the sales script text;
[0006] Determine whether there are preset characters in the sales script text, wherein the preset characters include polyphonic characters and numbers;
[0007] When the preset text exists in the sales script, add attribute tags to the preset text;
[0008] Determine whether the preset text has a definite pronunciation based on the attribute tags;
[0009] When the preset character has no definite pronunciation, the pronunciation of the preset character is recognized according to a first recognition model, or a first recognition model and a second recognition model, wherein the first recognition model and the second recognition model are different; and
[0010] The sales script text is converted into sales speech based on the pronunciation of the preset text, and the sales speech is then broadcast.
[0011] Secondly, embodiments of this application provide a computer device, the computer device comprising:
[0012] Memory, used to store program instructions; and
[0013] A processor is used to execute the program instructions to implement the method for broadcasting the sales voice of financial products as described above.
[0014] Thirdly, embodiments of this application provide a voice broadcasting system for the sale of financial products, the voice broadcasting system for the sale of financial products comprising:
[0015] The reading module is used to read sales script text;
[0016] The first judgment module is used to determine whether there are preset characters in the sales script text, wherein the preset characters include polyphonic characters and numbers;
[0017] The tag module is used to add attribute tags to the preset text when the preset text exists in the sales script text;
[0018] The second judgment module is used to determine whether the preset text has a definite pronunciation based on the attribute label;
[0019] A recognition module is configured to recognize the pronunciation of the preset text according to a first recognition model, or a first recognition model and a second recognition model, when the preset text does not have a definite pronunciation; wherein the first recognition model and the second recognition model are different; and
[0020] The execution module is used to convert the sales script text into sales speech based on the pronunciation of the preset text, and then broadcast the sales speech.
[0021] The aforementioned method, system, and computer equipment for broadcasting sales voice messages for financial products involve adding attribute tags to preset text, i.e., polyphonic characters or numbers, and determining the pronunciation of the preset text based on these attribute tags. If the pronunciation of the preset text cannot be determined, it is determined based on a first recognition model, or a first recognition model and a second recognition model. The sales script text containing the preset text with the determined pronunciation is then converted into sales voice for broadcast. This method for broadcasting sales voice messages for financial products builds upon existing decision tree models or hidden Markov models, combining natural language processing models, text-to-speech technology, and the characteristics of the financial industry. It employs NLP technology to create a rule-based algorithm model to distinguish and label scenarios where polyphonic characters and numbers appear; and uses TTS technology for voice broadcasting, playing corresponding pronunciations in a tagged manner. This reduces model training costs while improving the accuracy of text pronunciation, thereby increasing the accuracy of the information conveyed to customers and standardizing the financial product sales process, ultimately improving customer experience. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0023] Figure 1 A flowchart illustrating the method for broadcasting sales voice messages for financial products provided in this application embodiment.
[0024] Figure 2 The first sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment.
[0025] Figure 3 The second sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment.
[0026] Figure 4 The third sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment.
[0027] Figure 5 The fourth sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment.
[0028] Figure 6 The fifth sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment.
[0029] Figure 7 This is a schematic diagram illustrating a scenario for the voice broadcasting method for selling financial products provided in this application embodiment.
[0030] Figure 8 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application.
[0031] Figure 9 This is a schematic diagram of the internal structure of the voice broadcasting system for selling financial products provided in this application embodiment.
[0032] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0033] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0034] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar planned objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data are interchangeable where appropriate; in other words, the described embodiments are implemented according to a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, may also include other content; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] It should be noted that the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of that feature. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0036] Please refer to the following: Figure 1 and Figure 7 , Figure 1 A flowchart illustrating the method for broadcasting sales voice messages for financial products provided in this application embodiment. Figure 7 This is a schematic diagram illustrating a scenario for the method of broadcasting sales voice messages for financial products provided in this application embodiment. The method for broadcasting sales voice messages for financial products is applied in the fintech field and is used to broadcast pre-set sales scripts when selling financial products online. Figure 7Taking the illustrated application scenario as an example, a communication connection is established between the broadcasting platform 31 and the client 32. In this embodiment, the broadcasting platform 31 is used to execute a method for broadcasting the sales voice of financial products. The relevant functions of the broadcasting platform 31 can be implemented by a single device, multiple devices working together, or one or more functional modules within a single device; no specific limitations are imposed here. It is understood that the aforementioned functions can be network elements in hardware devices, software functions running on dedicated hardware, a combination of hardware and software, or virtualization functions instantiated on a platform (e.g., a cloud platform).
[0037] The specific steps for broadcasting sales announcements for financial products include the following.
[0038] Step S102: Read the sales script text.
[0039] The broadcasting platform 31 reads the sales script text. This sales script text corresponds to the financial product and is pre-set based on the product. The sales script text includes several sentences about the corresponding financial product.
[0040] Step S104: Determine whether there are preset texts in the sales script.
[0041] The broadcasting platform 31 determines whether there are pre-set characters in the sales script. These pre-set characters include polyphonic characters and numbers. It's understandable that polyphonic characters have different pronunciations depending on the context; in the financial field, the pronunciation of numbers representing monetary amounts differs from the pronunciation of ordinary numbers.
[0042] If the sales script contains preset text, proceed to step S106; if the sales script does not contain preset text, proceed to step S114.
[0043] Step S106: Add attribute tags to the preset text.
[0044] The broadcasting platform 31 adds attribute tags to the preset text. Specifically, the broadcasting platform 31 adds attribute tags to the preset text based on the context in which the preset text is located. Each preset text has a corresponding attribute tag.
[0045] The specific process of adding attribute tags to preset text will be described in detail below.
[0046] Step S108: Determine whether the preset text has a definite pronunciation based on the attribute tags.
[0047] The broadcasting platform 31 determines whether the preset text has a definite pronunciation based on its attribute tags. In this embodiment, attribute tags corresponding to numbers include, but are not limited to, phone numbers, finance, and assets; attribute tags corresponding to polyphonic characters include, but are not limited to, personal names and place names. For example, if the attribute tag of the number "80088208" is a phone number, then the corresponding pronunciation is "ba ling ling ba ba er ling ba"; if the attribute tag of the number "80088208" is finance, then the corresponding pronunciation is "ba qian ling ba wan ba qian er bai ling ba". That is to say, in scenarios representing monetary amounts, numbers with finance or asset attribute tags are pronounced as the entire numerical value, rather than as individual digits; in phone number scenarios, numbers with phone number attribute tags are pronounced as the phone number.
[0048] If the preset text does not have a definite pronunciation, proceed to step S110; if all preset texts have a definite pronunciation, proceed to step S112.
[0049] Step S110: Recognize the pronunciation of the preset text according to the first recognition model, or the first recognition model and the second recognition model.
[0050] The broadcasting platform 31 recognizes the pronunciation of preset text based on a first recognition model, or the broadcasting platform 31 recognizes the pronunciation of preset text based on a first recognition model and a second recognition model. The first recognition model and the second recognition model are different. It is understood that both the first and second recognition models are pre-trained recognition models.
[0051] In this embodiment, the first recognition model is a Hidden Markov Model (HMM) model, and the second recognition model is a Natural Language Processing (NLP) model.
[0052] In some feasible embodiments, the first identification model can be a decision tree model, or a combination of a hidden Markov model and a decision tree model, without limitation.
[0053] The specific process of recognizing the pronunciation of a preset text based on the first recognition model, or the first and second recognition models, will be described in detail below.
[0054] Step S112: Convert the sales script text into sales voice according to the preset pronunciation of the text, and broadcast the sales voice.
[0055] The broadcasting platform 31 uses text-to-speech (TTS) technology to convert all the text in the sales script into sales speech based on the pronunciation of the preset text, and then broadcasts it.
[0056] Step S114: Convert the sales script text into sales voice and broadcast the sales voice.
[0057] The broadcasting platform 31 uses text-to-speech (TTS) technology to directly convert all the text in the sales script into sales speech and broadcast it.
[0058] In the above embodiments, attribute tags are added to preset text, i.e., polyphonic characters or numbers, and the pronunciation of the preset text is determined by the attribute tags. If the pronunciation of the preset text cannot be determined, the pronunciation of the preset text is determined according to the first recognition model, or the first recognition model and the second recognition model. Then, the sales script text containing the preset text with the determined pronunciation is converted into sales voice for broadcast. The method of broadcasting sales voice for financial products is based on existing decision tree models or hidden Markov models, combined with natural language processing models, text-to-speech technology and the characteristics of the financial industry. It uses NLP technology to create a rule algorithm model to distinguish and label the scenarios in which polyphonic characters and numbers appear; it uses TTS technology for voice broadcasting, playing the corresponding pronunciation in a tagged manner, reducing the model training cost while improving the accuracy of text pronunciation, improving the accuracy of text pronunciation during voice broadcasting, and thus improving the accuracy of information conveyed to customers, making the sales process of financial products standardized and regulated, and improving customer experience. In addition, it effectively reduces manual maintenance costs and the possibility of human error, and the same script does not need to maintain multiple sets of content.
[0059] Please refer to the following: Figure 2 This is the first sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment. Step S106 specifically includes the following steps.
[0060] Step S202: Obtain scene information of preset text.
[0061] The broadcasting platform 31 acquires scene information of preset text. In this embodiment, scene information corresponding to numbers includes, but is not limited to, monetary scenes, telephone number scenes, etc.; scene information corresponding to polyphonic characters includes, but is not limited to, personal name scenes, place name scenes, etc.
[0062] Step S204: Add attribute tags to the preset text based on the preset weight dictionary and scene information.
[0063] The broadcasting platform 31 adds attribute tags to preset text based on a preset weight dictionary and scene information. In this embodiment, the preset weight dictionary includes an attribute tag library and a polyphonic character dictionary. The attribute tag library includes several attribute tags. The broadcasting platform 31 determines the attribute tags of the preset text from the attribute tag library based on the polyphonic character dictionary and the scene information of the preset text.
[0064] In the above embodiments, a preset weight dictionary and configuration rules are created. The attribute tag library and polyphonic character dictionary in the preset weight dictionary are used to contextualize and attribute-tagged polyphonic characters, numbers, etc., that is, attribute tags are added, so that the pronunciation of the preset text can be determined according to the context information of the preset text.
[0065] Please refer to the following: Figure 3 This is the second sub-flowchart of the method for broadcasting the sales voice of financial products provided in this application embodiment. In step S110, recognizing the pronunciation of the preset text according to the first recognition model specifically includes the following steps.
[0066] Step S302: Determine whether the preset text has context based on the first recognition model.
[0067] The broadcasting platform 31 determines whether the preset text has context based on the first recognition model. It is understandable that the first recognition model is capable of recognizing the context of the preset text.
[0068] When the preset text has context, proceed to step S304.
[0069] Step S304: Determine the pronunciation of the preset text based on the context of the preset text.
[0070] The broadcasting platform 31 determines the pronunciation of the preset text based on its context. It is understandable that if the preset text has context, its pronunciation can be determined.
[0071] In this embodiment, the pronunciation of all preset characters is marked using labels.
[0072] Step S306: When all preset texts have a definite pronunciation, add a mark to be broadcast to the sales script text.
[0073] When all preset texts have context, all preset texts have a definite pronunciation. When all preset texts have a definite pronunciation, the broadcasting platform 31 adds a "to be broadcast" mark to the sales script text. That is, the sales script text with the "to be broadcast" mark can be converted into sales speech and broadcast by the broadcasting platform 31 using TTS technology.
[0074] In the above embodiments, the first recognition model can recognize polyphonic characters, thereby enabling the phonetic annotation of preset text with sufficient context to obtain a definite pronunciation.
[0075] Please refer to the following: Figure 4 This is the third sub-flowchart of the method for broadcasting the sales voice of financial products provided in this application embodiment. In step S110, recognizing the pronunciation of the preset text according to the first recognition model and the second recognition model specifically includes the following steps.
[0076] Step S302: Determine whether the preset text has context based on the first recognition model.
[0077] The broadcasting platform 31 determines whether the preset text has context based on the first recognition model. It is understandable that the first recognition model is capable of recognizing the context of the preset text.
[0078] When the preset text has context, proceed to step S304; when the preset text does not have context, proceed to step S402.
[0079] Step S304: Determine the pronunciation of the preset text based on the context of the preset text.
[0080] When the preset text has context, the broadcasting platform 31 determines the pronunciation of the preset text based on the context. It can be understood that if the preset text has context, the pronunciation of the preset text can be determined.
[0081] Step S402: Based on the second recognition model and the context of the preset text, determine whether the preset text has a definite pronunciation.
[0082] When the preset text lacks context, the broadcasting platform 31 determines whether the preset text has a definite pronunciation based on the second recognition model and the context of the preset text. It is understandable that the second recognition model can determine the pronunciation of the preset text by incorporating context information.
[0083] In this embodiment, the pronunciation of all preset characters is marked using labels.
[0084] When all preset characters have a definite pronunciation, proceed to step S306.
[0085] Step S306: Add a "to be broadcast" marker to the sales script text.
[0086] When all preset texts have context, all preset texts have a definite pronunciation. When all preset texts have a definite pronunciation, the broadcasting platform 31 adds a "to be broadcast" mark to the sales script text. That is, the sales script text with the "to be broadcast" mark can be converted into sales speech and broadcast by the broadcasting platform 31 using TTS technology.
[0087] In the above embodiments, the first recognition model can recognize polyphonic characters, thereby enabling the phonetic annotation of preset texts with sufficient context to obtain a definite pronunciation. For preset texts whose pronunciation cannot be determined by the first recognition model, the second recognition model can combine the usage scenario of the entire text paragraph with other elements of the financial product, such as the article title where the preset text is located, the interface called by the system, and the channel for customer use, to infer the entire environmental scenario where the polyphonic character is located. For preset texts that cannot be recognized when they appear alone in the text paragraph, the model can infer their semantics and determine their pronunciation.
[0088] Leveraging the advantages of existing text-to-speech technology and Hidden Markov Models (HMMs), and addressing the issue of low accuracy of HMMs in handling polyphonic characters without context, this paper trains and learns a natural language processing model to generate contextualized labels for pronunciation determination. Furthermore, these contextualized labels can also be applied to other scenarios requiring voice broadcasting.
[0089] Please refer to the following: Figure 5 This is the fourth sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment. After executing step S402, the method for broadcasting sales voice messages for financial products further includes the following steps.
[0090] Step S502: When the preset text does not have a definite pronunciation, a prompt message is sent.
[0091] When neither the first recognition model nor the second recognition model can determine the pronunciation of the preset text, the broadcasting platform 31 sends a prompt message. In this embodiment, the prompt message is used to prompt the administrator (not shown) of the broadcasting platform 31 to intervene.
[0092] Step S504: Determine whether annotation information has been received.
[0093] The broadcasting platform 31 determines whether it has received the annotation information. Understandably, when the administrator of the broadcasting platform 31 sees the prompt information, they can annotate the preset text based on the prompt, thus generating annotation information. This annotation information includes the pronunciation of the preset text.
[0094] When the annotation information is received, step S506 is executed.
[0095] Step S506: Confirm that all preset characters have a definite pronunciation.
[0096] When the annotation information is received, the broadcasting platform 31 can determine the pronunciation of the corresponding preset text based on the annotation information, so that all preset texts have a definite pronunciation.
[0097] In the above embodiments, an annotation process is established. If there are any special unrecognized preset characters, the pronunciation can be corrected through the annotation process to obtain a definite pronunciation.
[0098] Please refer to the following: Figure 6 This is the fifth sub-flowchart of the method for broadcasting sales voice messages for financial products provided in this application embodiment. After executing step S502, the method for broadcasting sales voice messages for financial products further includes the following steps.
[0099] Step S602: When the annotation information is received, the first recognition model and the second recognition model are updated according to the annotation information.
[0100] When receiving annotation information, the broadcasting platform 31 can also update the first recognition model and the second recognition model based on the annotation information. That is to say, the broadcasting platform 31 can use the annotation information to retrain the first recognition model and the second recognition model.
[0101] Step S604: Update the preset weight dictionary based on the first recognition model, the second recognition model, and the annotation information.
[0102] The broadcasting platform 31 updates the preset weight dictionary based on the first and second recognition models after retraining and the annotation information.
[0103] In the above embodiments, a correction process is established, and the newly corrected pronunciations are integrated into the first and second recognition models for continued learning and training, enabling the models to learn independently and continuously improve. Unrecognized characters or those requiring special pronunciations are manually corrected and annotated for pronunciation, allowing the model to learn by association and gradually improve the accuracy of pronunciations of other related characters. The addition of a special annotation process enables the process to continuously learn independently, improving scalability and allowing the model to learn and train sustainably. This effectively improves the accuracy of pronunciations of polyphonic characters and numbers while reducing manual maintenance costs.
[0104] Please refer to the following: Figure 8 This is a schematic diagram of the internal structure of the computer device provided in this application embodiment. The computer device 10 includes a memory 11 and a processor 12. The memory 11 is used to store program instructions, and the processor 12 is used to execute the program instructions to implement the above-described method for broadcasting the sales voice of financial products.
[0105] In some embodiments, the processor 12 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program instructions stored in the memory 11.
[0106] The memory 11 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of a computer device, such as a hard disk. In other embodiments, the memory 11 may be an external storage device of a computer device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., provided on the computer device. Furthermore, the memory 11 may include both internal and external storage units of the computer device. The memory 11 can be used not only to store application software and various types of data installed on the computer device, such as code for implementing voice broadcasting methods for financial product sales, but also to temporarily store data that has been output or will be output.
[0107] Please refer to the following: Figure 9 This is a schematic diagram of the internal structure of the voice broadcasting system for selling financial products provided in this application embodiment. The voice broadcasting system 20 for selling financial products includes a reading module 21, a first judgment module 22, a tag module 23, a second judgment module 24, a recognition module 25, and an execution module 26.
[0108] Reading module 21 is used to read sales script text.
[0109] The reading module 21 reads the sales script text. This sales script text corresponds to the financial product and is pre-set according to the financial product. The sales script text includes several sentences about the corresponding financial product.
[0110] The first judgment module 22 is used to determine whether there is preset text in the sales script text.
[0111] The first judgment module 22 determines whether there are preset characters in the sales script text. These preset characters include polyphonic characters and numbers. It is understandable that polyphonic characters have different pronunciations depending on the context; in the financial field, the pronunciation of numbers representing monetary amounts differs from the pronunciation of ordinary numbers.
[0112] Tag module 23 is used to add attribute tags to preset text when preset text exists in the sales script text.
[0113] The tag module 23 adds attribute tags to the preset text. Specifically, the tag module 23 adds attribute tags to the preset text based on the context in which the preset text is located. Each preset text has a corresponding attribute tag.
[0114] The second judgment module 24 is used to determine whether the preset text has a definite pronunciation based on the attribute label.
[0115] The second judgment module 24 determines whether the preset text has a definite pronunciation based on its attribute tags. In this embodiment, attribute tags corresponding to numbers include, but are not limited to, phone numbers, finance, assets, etc.; attribute tags corresponding to polyphonic characters include, but are not limited to, personal names, place names, etc. For example, if the attribute tag of the number "80088208" is a phone number, then the corresponding pronunciation is "ba ling ling ba ba er ling ba"; if the attribute tag of the number "80088208" is finance, then the corresponding pronunciation is "ba qian ling ba wan ba qian er bai ling ba". That is to say, in the scenario representing a monetary amount, numbers with the finance or asset attribute tag are pronounced as the entire numerical value, rather than as a single digit; in the phone number scenario, numbers with the phone number attribute tag are pronounced as the phone number.
[0116] The recognition module 25 is used to recognize the pronunciation of the preset text according to the first recognition model, or the first recognition model and the second recognition model, when the preset text does not have a definite pronunciation.
[0117] The recognition module 25 recognizes the pronunciation of the preset text based on the first recognition model, or the recognition module 25 recognizes the pronunciation of the preset text based on the first recognition model and the second recognition model. The first recognition model and the second recognition model are different. It is understood that both the first recognition model and the second recognition model are pre-trained recognition models.
[0118] In this embodiment, the first recognition model is a Hidden Markov Model (HMM) model, and the second recognition model is a Natural Language Processing (NLP) model.
[0119] In some feasible embodiments, the first identification model may also be a combination of a hidden Markov model and a decision tree model, which is not limited here.
[0120] The execution module 26 is used to convert the sales script text into sales voice according to the pronunciation of the preset text, and to broadcast the sales voice.
[0121] The execution module 26 uses text-to-speech (TTS) technology to convert all the text in the sales script into sales speech based on the pronunciation of the preset text and then broadcast it.
[0122] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
[0123] The above-listed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A method for broadcasting sales voice messages for financial products, characterized in that, The method for broadcasting the sales voice message for the financial product includes: Read the sales script text; Determine whether there are preset characters in the sales script text, wherein the preset characters include polyphonic characters and numbers; When the preset text exists in the sales script, add attribute tags to the preset text; Determine whether the preset text has a definite pronunciation based on the attribute tags; When the preset character has no definite pronunciation, the pronunciation of the preset character is identified according to a first recognition model and a second recognition model, wherein the first recognition model and the second recognition model are different; and The sales script text is converted into sales speech based on the pronunciation of the preset text, and the sales speech is broadcast. Specifically, identifying the pronunciation of the preset text based on the first recognition model and the second recognition model includes: determining whether the preset text has context based on the first recognition model; when the preset text has context, determining the pronunciation of the preset text based on the context; when the preset text does not have context, determining whether the preset text has a definite pronunciation based on the second recognition model combined with the context of the preset text; and when all the preset texts have a definite pronunciation, adding a "to be broadcast" marker to the sales script text. Specifically, adding attribute tags to the preset text includes: obtaining scene information of the preset text, wherein scene information corresponding to numbers includes monetary scenes and telephone number scenes; scene information corresponding to polyphonic characters includes personal name scenes and place name scenes; and adding attribute tags to the preset text according to a preset weight dictionary and the scene information.
2. The method for broadcasting sales voice messages for financial products as described in claim 1, characterized in that, After determining whether the preset text has a definite pronunciation based on the second recognition model and the context of the preset text, the method for broadcasting the sales voice of the financial product further includes the following steps: When the preset text does not have a definite pronunciation, a prompt message is sent; Determine whether annotation information has been received, wherein the annotation information includes the pronunciation of the preset text; and When the annotation information is received, it is confirmed that all the preset characters have a definite pronunciation.
3. The method for broadcasting sales voice messages for financial products as described in claim 2, characterized in that, After determining whether the annotation information has been received, the method for broadcasting the sales voice message for the financial product further includes: When the annotation information is received, the first recognition model and the second recognition model are updated according to the annotation information; and The preset weight dictionary is updated based on the first recognition model, the second recognition model, and the annotation information.
4. The method for broadcasting sales voice messages for financial products as described in claim 1, characterized in that, After determining whether preset text exists in the sales script, the method for broadcasting the sales voice of financial products further includes: When the preset text is not present in the sales script, the sales script is converted into a sales voice and the sales voice is broadcast.
5. The method for broadcasting sales voice messages for financial products as described in claim 1, characterized in that, After determining whether the preset text has a definite pronunciation based on the attribute tags, the method for broadcasting the sales voice of financial products further includes: When all the preset texts have a definite pronunciation, the sales script text is converted into sales voice and the sales voice is broadcast.
6. A computer device, characterized in that, The computer device includes: Memory, used to store program instructions; and A processor is configured to execute the program instructions to implement the method for broadcasting the sales voice of financial products as described in any one of claims 1 to 5.
7. A voice broadcasting system for the sale of financial products, characterized in that, The voice broadcasting system for the sales of the financial products includes: The reading module is used to read sales script text; The first judgment module is used to determine whether there are preset characters in the sales script text, wherein the preset characters include polyphonic characters and numbers; The tag module is used to add attribute tags to the preset text when the preset text exists in the sales script text. Specifically, adding attribute tags to the preset text includes: obtaining the scenario information of the preset text, wherein the scenario information corresponding to numbers includes monetary scenarios and telephone number scenarios; the scenario information corresponding to polyphonic characters includes personal name scenarios and place name scenarios; and adding attribute tags to the preset text according to a preset weight dictionary and the scenario information. The second judgment module is used to determine whether the preset text has a definite pronunciation based on the attribute label; The recognition module is used to recognize the pronunciation of the preset text according to a first recognition model and a second recognition model when the preset text does not have a definite pronunciation. The first recognition model and the second recognition model are different. Specifically, recognizing the pronunciation of the preset text according to the first and second recognition models includes: determining whether the preset text has context based on the first recognition model; determining the pronunciation of the preset text based on its context when the preset text has context; determining whether the preset text has a definite pronunciation based on the second recognition model combined with the context of the preset text when the preset text does not have context; and adding a "to be broadcast" marker to the sales script text when all the preset text has a definite pronunciation. The execution module is used to convert the sales script text into sales speech based on the pronunciation of the preset text, and then broadcast the sales speech.
Citation Information
Patent Citations
Polyphone disambiguation method and device, electronic equipment and readable storage medium
CN114550691A