Intelligent robot voice information generation method and device, electronic equipment and medium
By generating voice information based on case attributes and communication templates using intelligent robots, the problem of rigid interaction in existing technologies has been solved, communication quality has been improved, and the hang-up rate has been reduced.
Patent Information
- Application Number
- CN202510265580.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing question-and-answer response mode of intelligent robots is based on fixed template configuration, which results in rigid interaction with the communication object and makes it impossible to adjust the wording or question-and-answer strategy according to the emotional state.
By acquiring a list of cases to be processed by the intelligent robot, and based on the correspondence between different case stages and communication templates, the target communication template is determined, and voice information, including timbre and volume, is generated according to case attributes. The script examples are adjusted in real time to improve communication quality.
It improved the communication quality of outbound calls made by intelligent robots, reduced the hang-up rate of the callers, and achieved more natural and effective interaction.
Smart Images

Figure CN119814926B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of robot telephone communication, in particular to a voice information generation method and device of an intelligent robot, an electronic device and a medium. BACKGROUND
[0002] More and more outbound intelligent robots have been put into practice in real business scenarios and have good market feedback, replacing a large number of manpower. However, in the prior art, the question and answer response mode of the intelligent robot is still based on fixed template configuration. Although the convenient front-end visual interface allows configuration of multiple templates, only one can be enabled in a conversation. This results in the interaction between the intelligent robot and the communication object being stereotyped and rigid, and the intelligent robot being unable to adjust its rhetoric or question and answer strategy according to the emotional state of the communication object. SUMMARY
[0003] The purpose of the embodiments of the present application is to provide a voice information generation method and device of an intelligent robot, an electronic device and a medium, to solve the above problems existing in the prior art and improve the quality of outbound calls of the intelligent robot.
[0004] In a first aspect, a voice information generation method of an intelligent robot is provided, which can include:
[0005] Obtaining a case list to be processed by the intelligent robot, the case list including a plurality of target cases and target stages and case attributes of the corresponding target cases;
[0006] For any target case, determining a target communication template corresponding to the target stage of the target case based on the correspondence relationship between different case stages and different communication templates configured; the target communication template includes variable information and fixed information;
[0007] Based on the case attributes, determining actual information corresponding to the variable information, and determining a rhetoric instance according to the actual information and the fixed information;
[0008] Processing the rhetoric instance according to the case attributes to generate voice information of the intelligent robot, so that the intelligent robot communicates with the communication object in the case attributes according to the voice information and the rhetoric instance; wherein the voice information includes tone and volume.
[0009] In a possible implementation, the configuration process of the correspondence relationship between different case stages and different communication templates includes:
[0010] Obtaining a communication completion rate of a plurality of different historical communication templates corresponding to each historical case stage in a configured historical period;
[0011] For any historical case stage, the historical communication template corresponding to the maximum communication completion rate in the plurality of communication completion rates of the historical case stage is determined as the target historical communication template of the historical case stage;
[0012] Based on the correspondence between each historical case stage and the corresponding target historical communication template, the correspondence between different case stages and different communication templates is determined.
[0013] In one possible implementation, after determining the voice information of the intelligent robot, the method further includes:
[0014] The dialogue instance is divided to determine a plurality of sub-contents and a communication order of the plurality of sub-contents;
[0015] When the intelligent robot communicates with the communication object in the case attribute according to the current sub-content, first feedback information of the communication object is obtained;
[0016] According to the first feedback information, a target sub-content in the plurality of sub-contents is determined.
[0017] In one possible implementation, according to the first feedback information, the target sub-content in the plurality of sub-contents is determined, including:
[0018] The first feedback information is processed by a configured preset ASR model to determine a feedback type corresponding to the first feedback information:
[0019] If the feedback type is interesting and the current sub-content has not been communicated, the current sub-content is determined as the target sub-content;
[0020] If the feedback type is interesting and the current sub-content has been communicated, according to the communication order of the plurality of sub-contents, a next sub-content of the current sub-content is determined as the target sub-content.
[0021] In one possible implementation, determining the feedback type corresponding to the first feedback information further includes:
[0022] If the feedback type is not interesting, a target level of not interesting is determined;
[0023] According to a correspondence between different levels and different appeasement information, target appeasement information corresponding to the target level is determined, so that the intelligent robot communicates with the communication object according to the target appeasement information.
[0024] In one possible implementation, after determining the target appeasement information corresponding to the target level, the method further includes:
[0025] Second feedback information of the intelligent robot after communicating with the communication object according to the target appeasement information is obtained;
[0026] determining a communication willingness of the communication object according to the second feedback information;
[0027] If the communication willingness is to continue the communication, determining a sub-content irrelevant to the current sub-content as a target sub-content in the plurality of sub-contents.
[0028] In a possible implementation, the determining of the communication willingness of the communication object further includes:
[0029] If the communication willingness is to refuse the communication, determining an end communication template corresponding to the refusal of the communication, so that the intelligent robot ends the current communication according to the end communication template.
[0030] In a second aspect, a voice information generation apparatus of an intelligent robot is provided, which can include:
[0031] An acquisition unit is configured to acquire a case list to be processed by the intelligent robot, the case list including a plurality of target cases, target stages at which the target cases are located, and case attributes of the target cases;
[0032] A determination unit is configured to determine, for any target case, a target communication template corresponding to a target stage at which the target case is located based on a correspondence relationship between different case stages and different communication templates configured;
[0033] and determine actual information corresponding to the variable information based on the case attributes, and determine a dialogue instance according to the actual information and the fixed information;
[0034] A generation unit is configured to process the dialogue instance according to the case attributes, and generate voice information of the intelligent robot, so that the intelligent robot communicates with a communication object in the case attributes according to the voice information and the dialogue instance; wherein the voice information includes tone and volume.
[0035] In a third aspect, an electronic device is provided, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0036] The memory is configured to store a computer program;
[0037] The processor is configured to execute the program stored on the memory, and implement the method steps of any of the first aspect.
[0038] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the method steps of any of the first aspect.
[0039] The application provides a voice information generation method of an intelligent robot, which comprises the following steps: obtaining a case list to be processed by the intelligent robot, wherein the case list comprises a plurality of target cases, target stages in which the target cases are located, and case attributes; for any target case, determining a target communication template corresponding to the target stage in which the target case is located based on a correspondence relationship between different case stages and different communication templates; the target communication template comprises variable information and fixed information; determining actual information corresponding to the variable information based on the case attributes, and determining a dialogue instance according to the actual information and the fixed information; processing the dialogue instance according to the case attributes, and generating voice information of the intelligent robot, so that the intelligent robot communicates with a communication object in the case attributes according to the voice information and the dialogue instance; wherein the voice information comprises tone and volume. The application can improve the communication quality of the intelligent robot and reduce the hang-up rate of the communication object. BRIEF DESCRIPTION OF DRAWINGS
[0040] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments of the application. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be regarded as a limitation to the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.
[0041] Figure 1 The system architecture diagram of the voice information generation method of the intelligent robot provided by the embodiments of the application is shown in the figure.
[0042] Figure 2 The flowchart of the voice information generation method of the intelligent robot provided by the embodiments of the application is shown in the figure.
[0043] Figure 3 The structural diagram of the voice information generation device of the intelligent robot provided by the embodiments of the application is shown in the figure.
[0044] Figure 4 The structural diagram of the electronic device provided by the embodiments of the application is shown in the figure. DETAILED DESCRIPTION
[0045] The technical solutions of the embodiments of the application will be described clearly and completely in the following with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only some of the embodiments of the application, and not all the embodiments. Based on the embodiments of the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0046] The voice information generation method of the intelligent robot provided in the embodiments of the present application can be applied in Figure 1 The system architecture is shown in FIG. 1. Figure 1 The system can include a terminal, a server and an intelligent robot. The server can be a physical server, a server cluster composed of multiple physical servers or a distributed system, and can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a mobile phone, a smart phone, a notebook computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer (PAD), etc. user equipment (UE), a handheld device, a vehicle-mounted device, a wearable device, a computing device or other processing device connected to a wireless modem, a mobile station (MS), a mobile terminal (Mobile Terminal), etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0047] The terminal is configured to obtain a case list to be processed, and send the case list to the server.
[0048] The server is configured to obtain the case list to be processed sent by the terminal, execute the voice information generation method of the intelligent robot provided in the present application, and send the generated voice information to the controller of the intelligent robot.
[0049] The controller of the intelligent robot is configured to control the intelligent robot to communicate with a communication object according to the voice information.
[0050] More and more outbound intelligent robots have been put into practice in real business scenarios and have good market feedback, replacing a large number of manpower. However, in the prior art, the question and answer response mode of the intelligent robot is still based on fixed template configuration. Although the convenient front-end visual interface allows configuration of multiple templates, only one can be enabled in a conversation. This results in the interaction between the intelligent robot and the communication object being stereotyped and rigid, and the intelligent robot being unable to adjust its rhetoric or question and answer strategy according to the emotional state of the communication object.
[0051] Therefore, the present application provides a voice information generation method of an intelligent robot to solve the above problems in the prior art and improve the quality of outbound calls of the intelligent robot.
[0052] The preferred embodiments of the present application are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are merely intended to illustrate and explain the present application, and are not intended to limit the present application, and the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0053] Figure 2 A flowchart of a voice information generation method of an intelligent robot is provided for the embodiments of the present application. As shown in Figure 2 , the method can include:
[0054] Step S210, obtaining a case list to be processed by the intelligent robot.
[0055] The case list includes a plurality of target cases and target stages and case attributes of the corresponding target cases.
[0056] The case attributes include case categories, telephone numbers of communication objects, places of origin of the telephone numbers, genders and ages of the communication objects, etc.
[0057] Specifically, the case list includes a plurality of target cases to be processed of the same category, and the target stages of the target cases to be processed can be the same or different.
[0058] Step S220, for any target case, determining a target communication template corresponding to the target stage of the target case based on the correspondence between different case stages and different communication templates.
[0059] Specifically, before performing step S220, the method further includes:
[0060] configuring the correspondence between different case stages and different communication templates, which can be specifically:
[0061] obtaining the communication completion rates of a plurality of different historical communication templates corresponding to each historical case stage of historical cases in a configured historical period;
[0062] For any historical case stage, the historical communication template corresponding to the maximum communication completion rate in the plurality of communication completion rates of the historical case stage is determined as the target historical communication template of the historical case stage;
[0063] Based on the correspondence between each historical case stage and the corresponding target historical communication template, the correspondence between different case stages and different communication templates is determined.
[0064] It should be noted that the historical cases and the target cases are cases of the same category. That is, the correspondence between different case stages and different communication templates of the target cases of the same category is determined.
[0065] In some embodiments, different case stages can also correspond to multiple communication templates, which are templates for the communication object in different scenarios when answering the phone. For example, when the communication object is driving a vehicle when answering the phone, it can correspond to one communication template. When the communication object is in a meeting when answering the phone, it can correspond to another communication template.
[0066] Further, the communication template includes variable information and fixed information. Variable information can be understood as information that changes with case attributes; fixed information is information that does not change. Variable information can also be information generated in real time during communication with the communication object, such as addresses, mobile phone numbers, moods, and genders mentioned by the communication object.
[0067] Then, according to the correspondence between the different case stages and the different communication templates configured, the target stage corresponding to the target communication template of the target case can be determined.
[0068] It can be understood that the target communication template also includes variable information and fixed information.
[0069] The communication template can be composed of multiple sub-templates, and the combination of the communication template can be determined according to the communication completion rate. The completion rate includes the sub-template completion rate and the communication template completion rate, and the new combination with the highest overall completion rate is selected as the new communication template.
[0070] Template update mechanism:
[0071] Real-time data collection: During the communication between the intelligent robot and the communication object, data is collected in real time, including feedback information of the communication object, communication time, and communication success times.
[0072] Data preprocessing: The collected data is cleaned and labeled to remove invalid data and noise data, and text data is processed by word segmentation, part-of-speech tagging, etc.
[0073] Model update: Establish a model based on machine learning, such as a reinforcement learning model. The model will dynamically adjust the correspondence between different case stages and communication templates based on real-time collected data. The goal of the model is to optimize communication effectiveness and improve communication completion rate and success rate. Using incremental learning technology, the model can be updated online in real time as new data is continuously collected, rather than waiting for a large amount of data to accumulate before offline training.
[0074] Template evaluation mechanism:
[0075] Communication completion rate calculation: For each communication template, calculate its communication completion rate, including sub-template completion rate and overall communication template completion rate. The communication completion rate can be calculated by the following formula:
[0076] Communication completion rate = number of successful communications / total number of communications;
[0077] Sub-template evaluation: For each sub-template, calculate its sub-template completion rate and evaluate and rank the sub-templates based on the sub-template completion rate. The sub-template completion rate can be calculated by the following formula:
[0078] Sub-template completion rate = number of successful communications of the sub-template / total number of communications of the sub-template.
[0079] Template combination evaluation: For different template combination methods, calculate their overall communication completion rate and evaluate and rank the template combination methods based on the overall communication completion rate. The overall communication completion rate can be calculated by the following formula:
[0080] Overall communication completion rate = number of successful communications / total number of communications;
[0081] Template optimization mechanism:
[0082] Select the optimal combination method: According to the overall communication completion rate, select the new combination method with the highest overall completion rate as the new communication template. If the overall communication completion rates of multiple combination methods are similar, other factors such as communication efficiency, communication object satisfaction, etc. can be considered for comprehensive evaluation and selection.
[0083] Sub-template optimization: According to the sub-template completion rate, optimize the sub-templates. For sub-templates with low completion rates, modifications and adjustments can be made, such as adjusting the content, tone, order, etc. of the sub-templates to improve their completion rates.
[0084] Template combination optimization: According to the template combination evaluation results, optimize the template combination method. Different combination methods can be tried, such as adjusting the order of sub-templates, adding or reducing sub-templates, etc. to improve the overall communication completion rate.
[0085] Template diversity maintenance:
[0086] Diversity evaluation: Regularly evaluate the diversity of the communication templates to ensure the diversity of the template set. Diversity metrics such as the similarity between templates, the range of scenarios covered by templates, etc. can be used for evaluation.
[0087] Diversity optimization: If the diversity of the template set is insufficient, the following measures can be taken to optimize:
[0088] Introduce new templates: Introduce new communication templates from other case categories or scenarios to increase the diversity of the template set.
[0089] Template variation: Perform variation operations on existing templates, such as randomly adjusting the content, tone, order, etc. of the templates to generate new templates.
[0090] Template combination variation: Perform variation operations on the template combination method, such as randomly adjusting the order of sub-templates, adding or reducing sub-templates, etc., to generate new template combination methods.
[0091] To increase the richness of the sample, when performing template combination, a random parameter can be introduced to randomly select sub-templates with high scores or newly optimized sub-templates. Here are the specific solutions:
[0092] Sub-template score:
[0093] Each sub-template is scored, which can be based on the completion rate, communication efficiency, and communication object satisfaction of the sub-template.
[0094] The sub-template score can be calculated by the following formula:
[0095] Sub-template score = w1 × sub-template completion rate + w2 × communication efficiency + w3 × communication object satisfaction
[0096] Where w1, w2, and w3 are weight coefficients.
[0097] Random selection:
[0098] When performing template combination, introduce a random parameter α, and the value of α ranges from 0 to 1.
[0099] According to the value of α, randomly select a sub-template:
[0100] When α is small (e.g., α < 0.3), it tends to select sub-templates with high scores.
[0101] When α is large (e.g., α > 0.7), it tends to select newly optimized sub-templates.
[0102] When α is in the middle range (e.g., 0.3 ≤ α ≤ 0.7), randomly select sub-templates with high scores or newly optimized sub-templates.
[0103] Dynamic adjustment of random parameters:
[0104] The random parameter α can be dynamically adjusted according to time or the number of communications. For example, as the number of communications increases, gradually increase the value of α to increase the probability of selecting newly optimized sub-templates.
[0105] Dynamic adjustment of α can be achieved by the following formula:
[0106] α = α0 + β * t
[0107] Where α0 is the initial value, β is the adjustment coefficient, and t is the time or the number of communications.
[0108] A score is calculated for each sub-template, and the score formula is as described above.
[0109] A random parameter a is generated, and a ranges from 0 to 1.
[0110] According to the value of a, a sub-template is selected:
[0111] When a < 0.3, the sub-template with the highest score is selected.
[0112] When a > 0.7, the newly optimized sub-template is selected.
[0113] When 0.3 ≤ a ≤ 0.7, a sub-template with a higher score or a newly optimized sub-template is randomly selected.
[0114] According to the selected sub-template, a new communication template is combined.
[0115] According to the communication completion rate of the new communication template, the template combination method is updated. If the communication completion rate of the new combination method is higher, it is used as the new communication template.
[0116] Step S230, based on the case attribute, determine the actual information corresponding to the variable information, and determine the dialogue instance according to the actual information and the fixed information.
[0117] Specifically, according to the information of the communication object in the case attribute of the target case, the actual information corresponding to the variable information in the target communication template is determined. For example: the work unit of the communication object in the case attribute of the target case is a school, and the communication object is Li; At this time, the actual information corresponding to the salutation variable in the target communication template is determined as: Li teacher.
[0118] In some embodiments, a greeting corresponding to the time of the outbound call can also be added before the salutation variable according to the time of the outbound call; for example: Good morning, Li teacher.
[0119] Then, according to the determined actual information and fixed information, the dialogue instance for this time with the corresponding communication object is determined.
[0120] It should be noted that the current determination is not one dialogue instance, and the foregoing has shown that one case stage can correspond to multiple communication templates, therefore, the determined dialogue instance is also multiple dialogue instances in different scenarios.
[0121] Step S240, according to the case attribute, processing the dialogue instance to generate the voice information of the intelligent robot, so that the intelligent robot communicates with the communication object in the case attribute according to the voice information and the dialogue instance.
[0122] Among them, the voice information includes tone and volume.
[0123] Before step S240 is performed, the method can further include:
[0124] Different timbres of sound are generated by using TTS (Text-to-Speech) technology. Each timbre is configured with a corresponding volume interval.
[0125] When the case category in the case attribute is product promotion, the timbre in the voice information of the intelligent robot can be determined as mature and stable.
[0126] When the case category in the case attribute is customer service, the timbre in the voice information of the intelligent robot can be determined as friendly and friendly.
[0127] When the case category in the case attribute is business communication or professional consulting service, the timbre in the voice information of the intelligent robot can be determined as professional and formal.
[0128] After the corresponding timbre is determined, the current volume interval corresponding to the current timbre is determined according to the correspondence between the different timbres and the different volume intervals configured.
[0129] It should be noted that when the outbound call is connected, the volume of the communication between the intelligent robot and the communication object is determined based on the background volume of the communication object.
[0130] In some embodiments, the background sound is environmental noise, and when the background volume reaches the configured preset background volume, the starting volume is adjusted according to the configured adjustment amplitude to determine the volume of the current communication, so that the communication object can clearly hear the voice of the intelligent robot. The starting volume can be the middle value of the volume interval.
[0131] When the background volume does not reach the configured preset background volume during quiet periods such as early morning and late evening, the lowest value of the volume interval is used to communicate with the communication object; when the communication object feedbacks that it cannot clearly hear, the lowest value of the volume interval is adjusted according to the configured adjustment amplitude to determine the volume of the current communication, so that the communication object can clearly hear the voice of the intelligent robot.
[0132] It should be noted that the current volume is adjusted at any time according to the actual situation during the whole process of communication between the intelligent robot and the communication object.
[0133] After determining the voice information of the intelligent robot, the method further includes:
[0134] The dialogue technique instance is divided to determine a plurality of sub-contents and a communication order of the plurality of sub-contents;
[0135] When the intelligent robot communicates with the communication object in the case attribute according to the current sub-content, the first feedback information of the communication object is obtained;
[0136] According to the first feedback information, a target sub-content in the plurality of sub-contents is determined. Specifically, the first feedback information is processed by a preset ASR model configured to determine a feedback type corresponding to the first feedback information: if the feedback type is interesting and the current sub-content is not communicated, the current sub-content is determined as the target sub-content; that is, the current sub-content is continued to be communicated.
[0137] If the feedback type is interesting and the current sub-content has been communicated, the next sub-content of the current sub-content is determined as the target sub-content according to the communication order of the plurality of sub-contents.
[0138] If the feedback type is not interesting, a target level of not interesting is determined; specifically, a target text corresponding to the first feedback information of the communication object is determined; according to a preset corresponding relationship between different not interesting texts and different levels, a target level corresponding to the target text is determined.
[0139] According to a preset corresponding relationship between different levels and different appeasing information, target appeasing information corresponding to the target level is determined, so that the intelligent robot communicates with the communication object according to the target appeasing information.
[0140] Then, second feedback information of the intelligent robot communicating with the communication object according to the target appeasing information is obtained.
[0141] According to the second feedback information, a communication willingness of the communication object is determined.
[0142] For example: when the intelligent robot asks some questions related to personal information, the feedback type of the first feedback information of the communication object is not interesting, or unwilling to answer such questions, in some embodiments, the target level can be determined according to the tone in the first feedback information, so as to determine the corresponding appeasing information; the intelligent robot communicates with the communication object according to the appeasing information, and obtains the second feedback information, that is, according to the second feedback information, the communication willingness of the communication object is determined.
[0143] If the communication willingness is to continue to communicate, a sub-content in the plurality of sub-contents that is irrelevant to the current sub-content is determined as the target sub-content.
[0144] If the communication willingness is to refuse to communicate, an end communication template corresponding to the refusal to communicate is determined, so that the intelligent robot ends the current communication according to the end communication template.
[0145] In some embodiments, since the communication willingness or the feedback type is determined after the natural language is analyzed by using the model, the result of the determination may be inaccurate due to the difference of languages in different regions, at this time, different languages (dialects) in different regions are involved in the preset ASR model, so that the language of the communication object belongs to the region, and the first feedback information and the second feedback information can be analyzed according to the region.
[0146] That is, a data set containing different cultural customs and etiquette is established in the preset ASR model, so that the intelligent robot makes appropriate responses in different cultural backgrounds. In short, the first feedback information and / or the second feedback information of the communication object is text converted according to different local dialects by the preset ASR model to obtain accurate text information.
[0147] Similarly, cultural markers can be added to the script instance according to the case attribute to enable the preset ASR model to better recognize and process the script instance.
[0148] In some embodiments, in order to ensure the security of all voice data, that is, the first feedback information and / or the second feedback information of the communication object in the process of the preset ASR model, an encryption mechanism is added to the preset ASR model to prevent data leakage.
[0149] The application provides a voice information generation method of an intelligent robot, which comprises the following steps: acquiring a case list to be processed by the intelligent robot, wherein the case list comprises a plurality of target cases, target stages in which the corresponding target cases are located, and case attributes; for any target case, determining a target communication template corresponding to a target stage in which the target case is located based on a correspondence relationship between different case stages and different communication templates configured; the target communication template comprises variable information and fixed information; determining actual information corresponding to the variable information based on the case attribute, and determining a script instance according to the actual information and the fixed information; processing the script instance according to the case attribute, and generating voice information of the intelligent robot to enable the intelligent robot to communicate with a communication object in the case attribute according to the voice information and the script instance; wherein the voice information comprises tone and volume. The application can improve the communication quality of the intelligent robot and reduce the hang-up rate of the communication object.
[0150] Corresponding to the above method, the application also provides an intelligent robot voice information generation device, as shown in the following Figure 3 The device comprises:
[0151] An acquisition unit 310 is configured to acquire a case list to be processed by the intelligent robot, wherein the case list comprises a plurality of target cases, target stages in which the corresponding target cases are located, and case attributes;
[0152] A determination unit 320 is configured to determine a target communication template corresponding to a target stage in which any target case is located based on a correspondence relationship between different case stages and different communication templates configured; the target communication template comprises variable information and fixed information;
[0153] Furthermore, based on the case attributes, the actual information corresponding to the variable information is determined, and based on the actual information and the fixed information, a script instance is determined;
[0154] The generation unit 330 is used to process the dialogue instance according to the case attributes to generate voice information for the intelligent robot, so that the intelligent robot can communicate with the communication object in the case attributes according to the voice information and the dialogue instance; wherein, the voice information includes timbre and volume.
[0155] The functions of each functional unit in the voice information generation device for an intelligent robot provided in the above embodiments of this application can be implemented through the above method steps. Therefore, the specific working process and beneficial effects of each unit in the voice information generation device for an intelligent robot provided in the embodiments of this application will not be repeated here.
[0156] This application also provides an electronic device, such as... Figure 4 As shown, it includes a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440.
[0157] Memory 430 is used to store computer programs;
[0158] When the processor 410 executes the program stored in the memory 430, it performs the following steps:
[0159] Obtain a list of cases to be processed by the intelligent robot, the list of cases including multiple target cases and the target stage and case attributes of the corresponding target cases;
[0160] For any given target case, based on the correspondence between different case stages and different communication templates configured, the target communication template corresponding to the target stage of the target case is determined; the target communication template includes variable information and fixed information.
[0161] Based on the case attributes, determine the actual information corresponding to the variable information, and determine the script instance based on the actual information and the fixed information;
[0162] Based on the case attributes, the script instance is processed to generate voice information for the intelligent robot, enabling the intelligent robot to communicate with the communication object in the case attributes based on the voice information and the script instance; wherein, the voice information includes timbre and volume.
[0163] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0164] The communication interface is used for communication between the electronic device and other devices.
[0165] The memory can include a Random Access Memory (RAM) and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0166] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0167] Since the implementation manners and beneficial effects of the electronic device in the above-mentioned embodiments can be achieved by the steps in the embodiments shown in the above-mentioned embodiments, the specific working process and beneficial effects of the electronic device provided by the embodiments of the present application will not be repeated here. Figure 2 The specific working process and beneficial effects of the electronic device provided by the embodiments of the present application will not be repeated here.
[0168] In another embodiment provided by the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores instructions, when the instructions run on a computer, the computer executes the voice information generation method of the intelligent robot in any one of the above-mentioned embodiments.
[0169] In a further example provided in the present application, a computer program product containing instructions, which when executed on a computer, causes the computer to perform the voice information generation method of the intelligent robot in any of the above examples.
[0170] Those skilled in the art should understand that the examples in the present application can be provided as a method, a system, or a computer program product. Therefore, the examples in the present application can be in the form of an entirely hardware example, an entirely software example, or an example combining software and hardware aspects. Moreover, the examples in the present application can be in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0171] The examples in the present application are described with reference to the flowcharts and / or block diagrams according to the methods, devices (systems), and computer program products of the examples in the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks
[0172] These computer program instructions can also be stored in a computer-readable memory capable of directing the computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including instruction devices that implement the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks
[0173] These computer program instructions can also be loaded into a computer or other programmable data processing device, so that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide functions for implementing the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in a flow or multiple flows and / or blocks Figure 1 The functions specified in a flow or multiple flows and / or blocks
[0174] Unless otherwise defined, technical terms or scientific terms used in the present application shall have the same meaning as commonly understood by one of ordinary skill in the art to which the present application pertains. Unless otherwise defined, the terms "first", "second", and similar terms are used to distinguish one element from another, and are not necessarily used to describe a sequential or chronological order. The terms "comprises", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. The terms "connected", "coupled", or any variant thereof are intended to cover a connection or coupling between or among two or more elements, and can encompass a direct connection or coupling or an indirect connection or coupling through one or more additional elements. The terms "upper", "lower", "left", "right", and the like are used to denote relative positions only and thus are used for convenience and illustrative purposes only and do not require a particular orientation with the described implementation.
[0175] While the preferred embodiments in the present application have been described, those skilled in the art will be able to make additional modifications and variations to these embodiments without departing from the spirit and scope of the preferred embodiments in the present application. Accordingly, it is intended that the preferred embodiments in the present application embrace all such modifications and variations as fall within the scope of the preferred embodiments in the present application.
[0176] Obviously, numerous modifications and variations of the preferred embodiments in the present application are possible in light of the above teachings. It is therefore to be understood that within the scope of the preferred embodiments in the present application, the preferred embodiments can be practiced otherwise than as specifically described.
Claims
1. A voice information generation method of an intelligent robot, the method comprising: The method comprises: obtaining a case list to be processed by an intelligent robot, the case list comprising a plurality of target cases and target stages and case attributes of the corresponding target cases; for any target case, determining a target communication template corresponding to the target stage of the target case based on the correspondence between different case stages and different communication templates configured; the target communication template comprises variable information and fixed information, and the correspondence is configured by: obtaining the communication completion rate of a plurality of different historical communication templates corresponding to each historical case stage in a historical period, determining the historical communication template corresponding to the maximum communication completion rate as the target historical communication template of the historical case stage, and determining the correspondence between different case stages and different communication templates based on the correspondence between each historical case stage and the corresponding target historical communication template; based on the case attributes, determining the actual information corresponding to the variable information, and determining the dialogue instance according to the actual information and the fixed information; processing the dialogue instance according to the case attributes to generate voice information of the intelligent robot, so that the intelligent robot communicates with the communication object in the case attributes according to the voice information and the dialogue instance; wherein the voice information comprises tone and volume; the tone is determined according to the case category in the case attributes, and the volume is determined by: determining the volume interval corresponding to the current tone according to the correspondence between different tones and different volume intervals, and dynamically adjusting the volume interval based on the background volume of the communication object after the call is connected; after determining the voice information of the intelligent robot, the method further comprises: dividing the dialogue instance to determine a plurality of sub-contents and a communication order of the plurality of sub-contents; obtaining first feedback information of the communication object when the intelligent robot communicates with the communication object according to the current sub-content; if the feedback type of the first feedback information is not interested, determining a target level of not interested, and determining target soothing information corresponding to the target level according to the correspondence between different levels and different soothing information, so that the intelligent robot communicates according to the target soothing information; the maintenance method of the communication template comprises diversity evaluation and diversification optimization, the diversification optimization comprises template variation, and the template variation comprises randomly adjusting the order of each sub-template in the communication template; when combining templates, a sub-template is selected based on a random parameter, the random parameter a is dynamically adjusted by the formula a = a0 + β * t, wherein a0 is an initial value, β is an adjustment coefficient, and t is time or the number of communications; and selecting a sub-template based on a random parameter comprises: selecting a sub-template with the highest score when a < 0.3, selecting a newly optimized sub-template when a > 0.7, and randomly selecting a sub-template with a higher score or a newly optimized sub-template when 0.3 ≤ a ≤ 0.7; the sub-template score is calculated by the formula sub-template score = w1 * sub-template completion rate + w2 * communication efficiency + w3 * communication object satisfaction, and w1, w2 and w3 are weight coefficients.
2. The method of claim 1, wherein, In the process that the intelligent robot communicates with the communication object according to the current sub-content, after obtaining the first feedback information of the communication object, the method further comprises: processing the first feedback information through the configured preset ASR model to determine the feedback type corresponding to the first feedback information: if the feedback type is interesting and the current sub-content has not been communicated, the current sub-content is determined as the target sub-content; if the feedback type is interesting and the current sub-content has been communicated, the next sub-content of the current sub-content is determined as the target sub-content according to the communication order of the plurality of sub-contents.
3. The method of claim 1, wherein, After determining the target level corresponding to the target pacification information, the method further comprises: obtaining second feedback information of the communication object after the intelligent robot communicates with the communication object according to the target pacification information; determining the communication intention of the communication object according to the second feedback information; if the communication intention is to continue communication, the sub-content unrelated to the current sub-content in the plurality of sub-contents is determined as the target sub-content.
4. The method of claim 3, wherein, Determining the communication intention of the communication object further comprises: if the communication intention is to refuse communication, an end communication template corresponding to the refusal of communication is determined to make the intelligent robot end this communication according to the end communication template.
5. A voice information generating apparatus of an intelligent robot, characterized by comprising: The device comprises: an acquisition unit configured to acquire a case list to be processed by an intelligent robot, wherein the case list comprises a plurality of target cases, target stages in which the target cases are located, and case attributes; a determination unit configured to determine, for any target case, a target communication template corresponding to a target stage in which the target case is located based on a correspondence relationship between different case stages and different communication templates, wherein the target communication template comprises variable information and fixed information, and the correspondence relationship is configured by: acquiring a communication completion rate of a plurality of different historical communication templates corresponding to each historical case stage in a historical period, determining a historical communication template corresponding to the maximum communication completion rate as a target historical communication template of the historical case stage, and determining the correspondence relationship between the different case stages and the different communication templates based on the correspondence relationship between each historical case stage and the corresponding target historical communication template; determining actual information corresponding to the variable information based on the case attributes, and determining a dialogue instance based on the actual information and the fixed information; a generation unit configured to process the dialogue instance based on the case attributes to generate voice information of the intelligent robot, so that the intelligent robot communicates with a communication object in the case attributes based on the voice information and the dialogue instance, wherein the voice information comprises tone and volume. The tone is determined according to a case category in the case attributes, and the volume is determined by: determining a volume interval corresponding to the current tone according to a correspondence relationship between different tones and different volume intervals, and dynamically adjusting the volume interval based on a background volume of the communication object after the outgoing call is connected. After determining the voice information of the intelligent robot, the method further includes: dividing the dialogue instance, determining a plurality of sub-contents and a communication order of the plurality of sub-contents; obtaining first feedback information of the communication object when the intelligent robot communicates with the communication object according to a current sub-content; if a feedback type of the first feedback information is not interested, determining a target level of not interested, determining target appeasement information corresponding to the target level according to a corresponding relationship between different levels and different appeasement information, and enabling the intelligent robot to communicate according to the target appeasement information; The maintenance mode of the communication template includes diversity evaluation and diversification optimization, the diversification optimization includes template variation, and the template variation includes random adjustment of the order of each sub-template in the communication template; when the template is combined, a sub-template is selected based on a random parameter, the random parameter α is dynamically adjusted through a formula α = α 0 + β * t, wherein α 0 is an initial value, β is an adjustment coefficient, and t is time or the number of communications; and the selection of the sub-template based on the random parameter includes: when α < 0.3, selecting a sub-template with the highest sub-template score, when α > 0.7, selecting a newly optimized sub-template, and when 0.3 ≤ α ≤ 0.7, randomly selecting a sub-template with a relatively high score or a newly optimized sub-template, the sub-template score is calculated through a formula sub-template score = w1 * sub-template completion rate + w2 * communication efficiency + w3 * communication object satisfaction, and w1, w2 and w3 are weight coefficients.
6. An electronic device, comprising: The electronic device includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor is used to execute the program stored on the memory, and realizes the method steps of any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the method steps of any one of claims 1-4. The computer readable storage medium stores a computer program, and the computer program is executed by the processor to realize the method steps of any one of claims 1-4.
Citation Information
Patent Citations
Intelligent outbound system and method, computer system and storage medium
CN112202978A