Script generation method and device based on intelligent outbound call, equipment, storage medium

By analyzing the target audience's historical outbound call voice and intent, personalized outbound call scripts are generated, solving the problem of repetitive scripts in intelligent outbound calling and improving information reach.

CN119402590BActive Publication Date: 2025-10-24CHINA PING AN LIFE INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411522023.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-10-24
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

In existing intelligent outbound calling systems, pre-set script combinations can cause recipients to receive the same type of question multiple times, resulting in low information reach and negatively impacting customer experience.

Method used

By acquiring the target object's historical scenario tasks and outbound call voices, and using voice extraction and intent recognition models, the node confirmation status of historical task nodes is identified, generating target outbound call scripts that are more suitable for the current scenario.

Benefits of technology

It improved the accuracy of matching target intent with the script, reduced repetitive scripts, and increased the information reach rate of outbound calling tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119402590B_ABST
    Figure CN119402590B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a dialogue generation method and device based on intelligent outbound call, equipment and storage medium, and belongs to the technical field of financial technology. The method comprises the following steps: obtaining a target scene task of a target object and historical outbound voice of the target object in a historical scene task; performing voice extraction on the historical outbound voice and historical node dialogue text based on a voice extraction model to obtain historical object voice; performing intention recognition on the historical object voice based on an intention recognition model to obtain a node reply intention category of the historical task node and category matching data; determining a node confirmation state of the target object to the historical task node based on the node reply intention category and the category matching data; and generating a target outbound dialogue based on the node confirmation state and the scene task node. The embodiment of the application can improve the accuracy of matching the object intention with the dialogue for intelligent outbound call, and generate more accurate outbound dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of financial technology, and in particular to a script generation method and device based on intelligent outbound call, equipment and storage medium. BACKGROUND

[0002] Intelligent outbound call is a service that uses artificial intelligence technology to automatically dial and communicate. For example, in the insurance scenario of financial technology, the intelligent outbound call can be used to remind the target object of the key node, and the process of reminding the target object of the key node can include identity confirmation, specific matter reminder, next visit mode reservation, etc. Moreover, the intelligent outbound call usually has a pre-set script combination, and the corresponding script is played by recognizing the answers or questions of the target object, so as to replace the work of manual customer service outbound call to a certain extent and reduce labor costs.

[0003] However, since the process of intelligent outbound call for key node reminder is generally long, the target object may be reminded multiple times before the node task expires. At present, the pre-set script combination is easy to cause the target object to receive multiple outbound calls containing the same question type script, so that the information reach rate of the outbound task is low, and the customer experience is affected. Therefore, how to improve the accuracy of matching the object intention with the script for intelligent outbound call, generate more accurate outbound script, and improve the information reach rate of the outbound task, has become a technical problem to be solved. SUMMARY

[0004] The main purpose of the embodiments of the present application is to provide a script generation method and device based on intelligent outbound call, equipment and storage medium, which aims to improve the accuracy of matching the object intention with the script for intelligent outbound call, generate more accurate outbound script, and improve the information reach rate of the outbound task.

[0005] To achieve the above purpose, a first aspect of the embodiments of the present application provides a script generation method based on intelligent outbound call, which comprises:

[0006] obtaining a target scene task of a target object, the target scene task comprising a scene task node;

[0007] obtaining a historical outbound voice of the target object in a historical scene task, the historical scene task comprising a historical task node and a historical node script text of the historical task node;

[0008] performing voice extraction on the historical outbound voice and the historical node script text based on a pre-set voice extraction model to obtain a historical object voice, the historical object voice being the voice of the target object answering the historical node script text in the historical outbound dialogue;

[0009] performing intent recognition on the historical object voice based on an intent recognition model to obtain a node reply intent category of the historical task node and category matching data, the category matching data being used to indicate a distance between the node reply intent category and speech text of the historical object voice;

[0010] determining a node confirmation state of the target object to the historical task node based on the node reply intent category and the category matching data;

[0011] generating a target outbound call script based on the node confirmation state and the scene task node.

[0012] In some embodiments, the generating a target outbound call script based on the node confirmation state and the scene task node comprises:

[0013] performing node division on the scene task node based on the node confirmation state to obtain a target confirmed node and a target unconfirmed node;

[0014] obtaining node confirmation data of the target object to the target confirmed node;

[0015] performing script generation on the target confirmed node and the node confirmation data based on a script generation model to obtain an outbound opening script;

[0016] performing script generation on the target unconfirmed node based on the script generation model to obtain an outbound unconfirmed node script;

[0017] performing script combination based on the outbound opening script and the outbound unconfirmed node script to obtain the target outbound call script.

[0018] In some embodiments, the performing node division on the scene task node based on the node confirmation state to obtain a target confirmed node and a target unconfirmed node comprises:

[0019] selecting a historical confirmed node from the historical task node based on the node confirmation state;

[0020] performing node matching on the historical confirmed node and the scene task node to obtain a node matching result;

[0021] determining the target confirmed node and the target unconfirmed node from the scene task node based on the node matching result.

[0022] In some embodiments, the method further comprises:

[0023] obtaining a task priority of the target scene task;

[0024] match the target scene task with an interaction mode based on the task priority, to determine a target interaction mode, which is used to indicate a communication mode adopted by the target object to confirm the target task node;

[0025] intelligently call the target object based on the target outbound call script and the target interaction mode.

[0026] In some embodiments, the matching the target scene task with an interaction mode based on the task priority, to determine a target interaction mode, comprises:

[0027] obtaining a pre-allocated interaction mode of the target scene task;

[0028] dividing a preset interaction reference number of the pre-allocated interaction mode based on a preset saturation degree division function, to obtain a saturation degree division interval, which comprises a first saturation interval and a second saturation interval, and the saturation degree of the first saturation interval is higher than that of the second saturation interval;

[0029] interval matching the number of to-be-allocated interaction objects of the pre-allocated interaction mode based on the saturation degree sub-interval, to determine a matching interval of the pre-allocated interaction mode;

[0030] if the matching interval is the first saturation interval, adjusting the pre-allocated interaction mode based on the task priority, to obtain the target interaction mode.

[0031] In some embodiments, the voice extraction based on the preset voice extraction model on the historical outbound call voice and the historical node script text, to obtain a historical object voice, comprises:

[0032] performing noise reduction processing on the historical outbound call voice, to obtain a noise-reduced outbound call voice;

[0033] performing voice coding on the noise-reduced outbound call voice, to obtain an outbound call voice feature;

[0034] performing voice feature separation on the outbound call voice feature, to obtain an outbound object voice feature, which is used to indicate different speaking objects in the historical outbound call voice;

[0035] performing feature decoding on the outbound object voice feature, to obtain an outbound object voice segment;

[0036] selecting a reply object voice segment from the outbound object voice segment based on the historical node script text, the reply object voice segment being a voice segment in which the target object replies to the historical node script text in the historical outbound call dialogue;

[0037] fragment integration is performed on the reply object voice segment to obtain the historical object voice.

[0038] In some embodiments, the intention recognition model performs intention recognition on the historical object voice to obtain a node reply intention category of the historical task node and category matching data, including:

[0039] obtaining a node verification category of the historical task node corresponding to the reply object voice segment, the node verification category including an intention recognition category and a labeling category;

[0040] If the node verification category is the intention recognition category, intention recognition is performed on the reply object voice segment based on an intention recognition model to obtain the node reply intention category and the category matching data;

[0041] If the node verification category is the labeling category, a preset category is determined as the node reply intention category of the reply object voice segment, and a preset matching data is determined as the category matching data of the reply object voice segment.

[0042] To achieve the above object, a second aspect of the embodiment of the present application provides a dialogue generation device based on intelligent outbound call, the device comprising:

[0043] a first obtaining module configured to obtain a target scene task of a target object, the target scene task including a scene task node;

[0044] a second obtaining module configured to obtain historical outbound voice of the target object in a historical scene task, the historical scene task including a historical task node and a historical node dialogue text of the historical task node;

[0045] an extraction module configured to perform voice extraction on the historical outbound voice and the historical node dialogue text based on a preset voice extraction model to obtain a historical object voice, the historical object voice being voice of the target object in reply to the historical node dialogue text in the historical outbound dialogue;

[0046] an identification module configured to perform intention recognition on the historical object voice based on an intention recognition model to obtain a node reply intention category of the historical task node and category matching data, the category matching data being used to indicate a distance between the node reply intention category and voice text of the historical object voice;

[0047] a determination module configured to determine a node confirmation state of the target object to the historical task node based on the node reply intention category and the category matching data;

[0048] The generating module is configured to generate a target outbound call script based on the node confirmation state and the scene task node.

[0049] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.

[0050] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.

[0051] The method and device for generating a script based on intelligent outbound calls, the equipment and the storage medium provided by the present application, by obtaining a target scene task of a target object, the target scene task comprises a scene task node; obtaining the historical outbound voice of the target object in the historical scene task, the historical scene task comprises a historical task node and a historical node script text of the historical task node; further, based on the preset voice extraction model, the historical outbound voice and the historical node script text are subjected to voice extraction to obtain a historical object voice, the historical object voice is the voice of the target object in the historical outbound dialogue in reply to the historical node script text; further, based on the intention recognition model, the historical object voice is subjected to intention recognition to obtain a node reply intention category of the historical task node and category matching data, the category matching data is used to indicate the distance between the node reply intention category and the voice text of the historical object voice; further, based on the node reply intention category and the category matching data, the node confirmation state of the target object to the historical task node is determined; further, based on the node confirmation state and the scene task node, the target outbound call script is generated. Compared with the related art which uses the pre-set fixed script combination, it is easy to cause the reminding object to receive the outbound call containing the same question type script for many times, so that the information reach rate of the outbound task is low. The embodiments of the present application can determine the node confirmation state of the historical task node from the historical object voice by combining voice extraction and intention recognition, so as to generate a target outbound call script more suitable for the current scene task node. In this way, it can be avoided that the target object receives the outbound call containing the same question type script for many times, the accuracy of matching the object intention with the script for intelligent outbound call is improved, a more accurate outbound call script is generated, and thus the information reach rate of the outbound task is improved. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 is a flowchart of the method for generating a script based on intelligent outbound calls provided by the embodiments of the present application;

[0053] Figure 2 is Figure 1one flowchart of step S130 in FIG. 13;

[0054] Figure 3 is Figure 1 one flowchart of step S140 in FIG. 14;

[0055] Figure 4 is Figure 1 one flowchart of step S160 in FIG. 16;

[0056] Figure 5 is Figure 4 one flowchart of step S410 in FIG. 41;

[0057] Figure 6 is another flowchart of the script generation method based on intelligent outbound call provided by the embodiments of the present application;

[0058] Figure 7 is Figure 6 one flowchart of step S620 in FIG. 62;

[0059] Figure 8 is a structural schematic diagram of the script generation apparatus based on intelligent outbound call provided by the embodiments of the present application;

[0060] Figure 9 is a hardware structural schematic diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0061] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0062] It should be noted that although the functional modules are divided in the apparatus schematic diagram, and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a manner different from the module division in the apparatus or the order in the flowchart. The terms "first", "second", etc. in the specification and claims and the above-described drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0064] First, the several terms involved in the present application are analyzed:

[0065] Artificial intelligence (AI): is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science, and artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The research in this field includes robots, language recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0066] Natural language processing (NLP): NLP uses computers to process, understand and use human language (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and is an interdisciplinary subject of computer science and linguistics, also commonly known as computational linguistics. Natural language processing includes syntax analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis and opinion mining, etc. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and language computing related linguistic research.

[0067] Intent Recognition: refers to the process of identifying the intent or purpose behind user inputted statements, text, images and other data through the system.

[0068] Intelligent outbound call is a service that uses artificial intelligence technology to automatically dial and communicate. For example, in the insurance scenario of financial technology, the intelligent outbound call method can be used to remind the reminder object of the key node (such as insurance expiration payment, claim information submission, etc.), and the identity confirmation of the reminder object, the specific matter reminder, the next return way reservation, etc. can be included in the process of reminding the key node. And this intelligent outbound call method usually sets the speech combination in advance, and plays the corresponding speech by recognizing the answers or questions of the reminder object, so as to replace the work of artificial customer service outbound call to a certain extent, thereby reducing the labor cost.

[0069] However, since the process of intelligent outbound call reminding the key node is generally longer, the reminding object may be reminded multiple times before the node task expiration date. At present, this method of setting the speech combination in advance easily causes the reminding object to receive the outbound call containing the same question type speech multiple times, so that the information reach rate of the outbound task is low, and the customer experience is affected. Therefore, how to improve the accuracy of matching the object intention with the speech for intelligent outbound call, generate more accurate outbound speech, and thereby improve the information reach rate of the outbound task, has become a technical problem to be solved.

[0070] Based on this, the embodiments of the present application provide a speech generation method and device based on intelligent outbound call, equipment and storage medium, aiming to improve the accuracy of matching the object intention with the speech for intelligent outbound call, generate more accurate outbound speech, and thereby improve the information reach rate of the outbound task.

[0071] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (Artificial Intelligence, AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0072] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.

[0073] The speech generation method based on intelligent outbound call provided by the embodiments of the present application relates to the field of artificial intelligence technology. The speech generation method based on intelligent outbound call provided by the embodiments of the present application can be applied in a terminal, can also be applied in a server end, and can also be software running in a terminal or a server end. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server end can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, can also be configured as a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and basic cloud computing services such as big data and artificial intelligence platforms; the software can be an application that implements the speech generation method based on intelligent outbound call, etc., but is not limited to the above forms.

[0074] The application is operable in a multitude of generic or specific computer system environments or configurations. Examples of well known computing systems, environments, and / or configurations that can be suitable for use with the application include personal computers, server computers, handheld or laptop devices, tablet devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like. The application can be described in the general context of computer-executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like, that perform particular tasks or implement particular abstract data types. The application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0075] It should be noted that in each of the specific embodiments of the present application, when it is necessary to perform relevant processing according to the object out-bound voice, the object historical data, and the object insured data, and other data related to the object identity or characteristics, the permission or consent of the object is obtained first, and the collection, use, and processing of the data comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the object, the separate permission or separate consent of the object is obtained through a pop-up window or a jump to a confirmation page, and after the separate permission or separate consent of the object is obtained, the necessary user-related data for enabling the embodiments of the present application to normally operate is obtained.

[0076] Please refer to Figure 1 , Figure 1 is an optional flowchart of the script generation method based on intelligent out-bound provided by the embodiments of the present application. In some embodiments of the present application, the method in Figure 1 may specifically include but is not limited to steps S110 to S160.

[0077] Step S110, obtaining a target scene task of a target object;

[0078] Step S120, obtaining a historical out-bound voice of the target object in a historical scene task;

[0079] Step S130, performing voice extraction on the historical out-bound voice and historical node script text based on a preset voice extraction model to obtain historical object voice;

[0080] Step S140, performing intent recognition on the historical object voice based on an intent recognition model to obtain a node reply intent category of the historical task node and category matching data;

[0081] Step S150, determining the node confirmation state of the target object to the historical task node based on the node reply intention category and the category matching data;

[0082] Step S160, generating the target outbound call script based on the node confirmation state and the scene task node.

[0083] In step S110 of some embodiments, the target object refers to an object that needs to be called, i.e., an object that needs to be confirmed by a set of interactive modes. The target scene task refers to a task of a business scene that the outbound call system needs to interact with the target object to complete, for example, a renewal payment reminder task, a renewal business recommendation task, etc. in the insurance field. The target scene task can include multiple scene task nodes, and the outbound call system can complete the confirmation of the target scene task through a series of pre-defined scene task nodes. For example, the target scene task is a renewal payment reminder task in the insurance field, and the scene task nodes can include identity confirmation, payment reminder, appointment of next visit mode, etc. In this way, there can be a sequence relationship, a priority relationship, etc. between the scene task nodes.

[0084] It should be noted that in intelligent outbound call, each scene task node can correspond to a specific dialogue target or question, i.e., a configured outbound call script, for generating corresponding voice to confirm to the target object. Among them, the target scene task and the scene task node can be pre-configured to realize the differentiated management of the outbound call tasks of different business scenes. And since each scene task node can be different, the scope of intention recognition for different target scene tasks will be different.

[0085] In step S120 of some embodiments, the historical scene task refers to a scene task that has been communicated with the target object through a set of interactive modes. The historical scene task includes historical task nodes and historical node script texts of the historical task nodes. The historical task node refers to a task node that has been transferred in the historical outbound call process, and the historical task node can be the same as or different from the scene task node, depending on whether the historical task node has been confirmed or whether the node needs to be confirmed multiple times as configured by the target scene task. The historical node script text refers to a pre-set script text for confirming the historical task node to the target object. For example, if the historical task node is identity confirmation, the corresponding historical node script text can be "Are you xxx yourself?". The historical outbound call voice is the voice generated by the outbound call system based on the historical node script text, and the communication voice generated after the intelligent communication with the target object.

[0086] In step S130 of some embodiments, the speech extraction model refers to a model for extracting the speech of different objects from the input speech. The speech extraction model can be trained based on a machine learning model, a neural network model, or the like. For example, the speech extraction model can be a model constructed based on a Time-Domain Audio Separation Network (TasNet) structure, which can achieve human voice separation while taking into account efficiency and performance to obtain a human voice speech segment containing only the reply of the target object.

[0087] It should be noted that the historical object speech is the speech of the target object replying to the historical node script text in the historical outbound call dialogue. In the historical object speech, the target object can reply to all speech questions corresponding to the historical node script text, or can reply to part of the speech questions corresponding to the historical node script text, which is not limited in particular. In this way, by obtaining the reply speech of the target object in the historical outbound call dialogue, it can be analyzed how the target object responds to a specific script, so as to generate a script more suitable for the target object, which is convenient for use in the next outbound interaction.

[0088] Please refer to Figure 2 , Figure 2 is a specific flowchart of step S130 provided by the embodiments of the present application. In some embodiments of the present application, step S130 can specifically include but is not limited to steps S210 to S260.

[0089] Step S210, performing noise reduction processing on the historical outbound call speech to obtain a noise-reduced outbound call speech;

[0090] Step S220, performing speech coding on the noise-reduced outbound call speech to obtain outbound call speech features;

[0091] Step S230, performing speech feature separation on the outbound call speech features to obtain outbound object speech features;

[0092] Step S240, performing feature decoding on the outbound object speech features to obtain outbound object speech segments;

[0093] Step S250, selecting a reply object speech segment from the outbound object speech segments based on the historical node script text;

[0094] Step S260, performing segment integration on the reply object speech segment to obtain a historical object speech.

[0095] In step S210 of some embodiments, considering that there can be some noise or background sound in the obtained historical outbound call voice, in order to reduce the interference on the subsequent key intent recognition, the application can first preprocess the historical outbound call voice, and the preprocessing at this time can include voice cleaning, noise reduction processing, voice enhancement, etc., without limitation. Noise reduction processing (such as filter and spectral subtraction) can effectively remove background noise and interference sound, thereby obtaining clear outbound call voice.

[0096] In step S220 of some embodiments, after noise reduction processing, the clear noise-reduced outbound call voice can be encoded using a speech coding algorithm to convert the noise-reduced outbound call voice into a series of features that can represent the voice signal. The obtained outbound call voice features can include Mel-frequency cepstral coefficients (MFCCs), Mel-spectral energy features, etc., which can capture key information in the voice signal while reducing the dimensionality of the data.

[0097] In step S230 of some embodiments, further, speech separation techniques can be implemented based on the differences in different object voice features to distinguish the features of different speaking objects. In a multi-speaker environment, this can be achieved by using machine learning techniques such as clustering or classification algorithms to separate features, with the goal of extracting the features of each independent speaker from the mixed voice. Outbound object voice features refer to the different speaking objects in the historical outbound call voice. For example, in the intelligent outbound call scenario of financial technology, the historical outbound call voice can contain the voice of the intelligent customer service and the voice of the target object, and the outbound object voice features can contain the voice features of the intelligent customer service in the outbound call voice features and the voice features of the target object in the outbound call voice features.

[0098] In step S240 of some embodiments, further, the separated outbound object voice features can be decoded to obtain at least one voice segment of the intelligent customer service in the historical outbound call voice and at least one voice segment of the intelligent customer service in the historical outbound call voice. Feature decoding here refers to converting the extracted features back to understandable human voice, and can specifically use text-to-speech synthesis (TTS) technology or other synthesis methods.

[0099] In step S250 of some embodiments, since the voice script of the intelligent customer service in the historical outbound call voice is pre-configured, at least one reply object voice segment of the target object can be selected from the outbound object voice segment based on the historical node script text. The reply object voice segment is the voice segment of the target object in the historical outbound call dialogue in response to the historical node script text.

[0100] It should be noted that the selection of the at least one reply object voice segment of the target object from the outbound object voice segment based on the historical node dialogue text can specifically include: performing voice text conversion on the outbound object voice segment to obtain an outbound object voice segment text; performing text matching on the outbound object voice segment text and the historical node dialogue text to obtain a matching state of the outbound object voice segment text; if the matching state indicates that the outbound object voice segment text matches the historical node dialogue text, it is determined that the corresponding outbound object voice segment belongs to intelligent customer service, i.e., a non-target object; and if the matching state indicates that the outbound object voice segment text does not match the historical node dialogue text, it is determined that the corresponding outbound object voice segment belongs to a non-target object. Moreover, this process can rely on content matching and semantic analysis to ensure that the selected segment matches the target object. In this way, all voice segments in which the target object replies to the historical node dialogue text in the historical outbound call dialogue can be determined.

[0101] In step S260 of some embodiments, further, the obtained at least one reply object voice segment can be integrated to obtain historical object voice generated by the target object in the historical outbound call dialogue. For example, the historical outbound voice is "Please ask you are XXX person? Hmm, yes. Your policy expires on XX month XX day, please remind you to complete the payment before the expiration date to avoid the policy invalid. Good, I know", after steps S210 to S260 of the present application, it can be determined that the reply object voice segment includes "hmm, yes" and "good, I know". Further, after the integration of the voice segment, a complete historical object voice "hmm, yes / good, I know" is formed, wherein " / " is a mark used to separate the reply voice of different historical node dialogue texts.

[0102] It should be noted that the voice segment integration process can use overlap-add, weighted superposition, etc. to ensure that the final reply content of the historical outbound call dialogue is perfectly presented.

[0103] In the above embodiments, because the historical outbound voice obtained in one call interaction can contain the confirmation results of multiple nodes, it is necessary to first perform voice extraction on the historical outbound voice, so as to facilitate more rapid and accurate identification of the intention of the target object.

[0104] In step S140 of some embodiments, after obtaining the historical object voice, the intent recognition model can be used to perform intent recognition on each reply object voice segment in the historical object voice to obtain a node reply intent category of the historical task node and category matching data. Since a historical task node can also include multiple dialogue texts, i.e., through continuous inquiries to determine the confirmation of the target object to the historical task node. The node reply intent category is used to indicate the confirmation degree of the target object to the historical task node, and the node reply intent category can include a confirmation category, an unconfirmed category, an uncertain category, etc. The confirmation category indicates that the target object explicitly expresses a positive intent (such as good, yes, etc.) or provides explicit inquiry data (such as providing an explicit certificate number, address information, etc.). The unconfirmed category indicates that the target object explicitly expresses a negative intent (such as no, don't know, etc.). The uncertain category indicates that the target object does not explicitly express a positive or negative intent (such as unclear, seems, etc.).

[0105] The category matching data is used to indicate the distance between the node reply intent category and the voice text of the historical object voice. The category matching data can be directly determined based on a text similarity value between the node reply intent category and the voice text of the historical object voice, and the text similarity value can be calculated based on a Euclidean distance function, a cosine similarity function, etc. The category matching data can also be a confidence degree obtained by normalizing the obtained text similarity value (such as a value range [0, 1]), and a higher confidence value indicates a higher credibility of the node reply intent category.

[0106] Please refer to Figure 3 , Figure 3 is a specific flowchart of step S140 provided by the embodiments of the present application. In some embodiments of the present application, step S140 can specifically include but is not limited to steps S310 to S320.

[0107] Step S310: obtaining a node verification category of the historical task node corresponding to the reply object voice segment;

[0108] Step S320: performing intent recognition on the reply object voice segment based on the intent recognition model and the node verification category to obtain a node reply intent category and category matching data.

[0109] In step S310 of some embodiments, the node verification category is used to indicate a manner in which the historical task node is verified. The node verification category includes an intent recognition category and a labeling category, where the intent recognition category indicates that the corresponding historical task node needs to be determined by performing intent recognition on the corresponding reply object voice segment. The labeling category indicates that the corresponding historical task node is determined by manual labeling. For example, when the historical scenario task is a renewal payment reminder task in the insurance field, the historical task nodes contained therein include identity verification, payment reminder, and whether to transfer to a human, the node verification category corresponding to identity verification can be set as the intent recognition category, the node verification category corresponding to payment reminder can be set as the intent recognition category, and the node verification category corresponding to whether to transfer to a human can be set as the intent recognition category. For another example, when the historical scenario task contains historical task nodes including identity verification, payment reminder, and appointment for next visit, the node verification category corresponding to identity verification can be set as the intent recognition category, the node verification category corresponding to payment reminder can be set as the intent recognition category, and the node verification category corresponding to appointment for next visit can be set as the labeling category.

[0110] In step S320 of some embodiments, further, the intent recognition model and the node verification category can be used to perform intent recognition on the reply object voice segment to obtain the node reply intent category and the category matching data. That is, the node verification category of the historical task node is considered, and thus a manner for recognizing the corresponding reply object voice segment is determined to determine the node reply intent category and the category matching data.

[0111] It should be noted that, considering that the nodes to be verified in each scenario are a definite and limited set of intents, for example, the “historical task node | node verification category” contained in the historical scenario task 1 can include: {identity confirmation | intent recognition, payment reminder | intent recognition}, therefore, the intent recognition model of the present application can use a model based on the FastText model structure to achieve fast intent recognition, in addition, the intent recognition model can also use other model structures, such as a convolutional neural network model, without specific limitation.

[0112] In some specific embodiments, the process of performing intent recognition on the reply object voice segment based on the intent recognition model and the node verification category to obtain the node reply intent category and the category matching data can include the following two cases:

[0113] If the node verification category is the intent recognition category, the intent recognition model is used to perform intent recognition on the reply object voice segment to obtain the node reply intent category and the category matching data;

[0114] If the node verification category is the labeling category, the preset category is determined as the node reply intent category of the reply object voice segment, and the preset matching data is determined as the category matching data of the reply object voice segment.

[0115] In the above embodiment, if the node verification category is the intent recognition category, the intent recognition model can be used to perform intent recognition on the reply object voice segment to obtain the node reply intent category and the category matching data. If the node verification category is the labeling category, i.e., the artificial labeling method is used, the preset category can be directly determined as the node reply intent category of the reply object voice segment, and the preset matching data can be directly determined as the category matching data of the reply object voice segment. For example, if the historical task node is the reservation of the next visit mode, the preset category can be the telephone confirmation category, and the category matching data is 0, indicating that the target object has not completed the confirmation. That is, if the node verification category is the labeling category, the target scene task is not identified.

[0116] It should be noted that the application can configure one or more node verification categories for a node, and when multiple node verification categories are configured, the node verification categories can be verified based on the priority of the node verification categories.

[0117] In step S150 of some embodiments, the confirmation state of the target object to the historical task node can be further determined according to the node reply intent category and the category matching data. For example, when the node reply intent category is the confirmation category, and the category matching data is the confidence, the higher the confidence corresponding to the category, the higher the credibility of the corresponding category. At this time, a confidence threshold (such as 0.6, 0.7, etc.) can be set, the confidence of different node reply intent categories can be compared with the confidence threshold, the intent recognition result (i.e., the corresponding intent category) lower than the confidence threshold can be removed as an untrusted intent, only the intent category higher than the confidence threshold is retained, and the node confirmation state of the target object is determined based on the intent category corresponding to the maximum confidence higher than the confidence threshold. If the intent category corresponding to the maximum confidence is the confirmation category, the node confirmation state is the confirmed state, i.e., it is considered that the customer has completed the confirmation of this node; if the intent category corresponding to the maximum confidence is the unconfirmed category, the node confirmation state is the unconfirmed state, i.e., it is considered that the customer has not completed the confirmation of this node. For example, the historical task node includes an identity confirmation node, the reply voice of the target object to the historical node dialogue text “May I ask if you are XXX yourself?” is “Hmm, yes”, the node reply intent category is determined as the confirmation category through recognition, and the confidence is 0.8. If the confidence threshold is 0.6, the node confirmation state of the identity confirmation node can be determined as the confirmed state.

[0118] In step S160 of some embodiments, further, the application can generate a target outbound call script based on the node confirmation state and the scene task node. The target scene task is a task after the historical scene task, that is, the application can adjust the script for the next time based on the reply of the target object in the historical outbound call to optimize the opening script of the next outbound call. For example, if the node confirmation state corresponding to the identity confirmation node is the confirmed state, and the node confirmation state corresponding to the payment reminder node is the confirmed state, the opening script of the next time can be "XXX, hello, I am XX who contacted you last time, your identity information is XXXX confirmed last time, and you are also notified that the policy expiration date is XXXX, do you remember?". In this way, the information that has been confirmed can be expressed in the opening script, avoiding the problem of repeated reply of the target object, generating a more accurate outbound call script, thereby improving the information reach rate of the outbound task.

[0119] Referring to Figure 4 , Figure 4 is a specific flowchart of step S160 provided by the embodiments of the application. In some embodiments of the application, step S160 can specifically include but is not limited to steps S410 to S450.

[0120] In step S410, the scene task nodes are divided based on the node confirmation state, to obtain target confirmed nodes and target unconfirmed nodes;

[0121] In step S420, the node confirmation data of the target object to the target confirmed nodes is obtained;

[0122] In step S430, the target confirmed nodes and the node confirmation data are subjected to script generation based on the script generation model, to obtain an outbound opening script;

[0123] In step S440, the target unconfirmed nodes are subjected to script generation based on the script generation model, to obtain an outbound unconfirmed node script;

[0124] In step S450, the outbound opening script and the outbound unconfirmed node script are combined to obtain a target outbound call script.

[0125] In step S410 of some embodiments, the target confirmed nodes refer to the nodes in the scene task nodes with the confirmed state of the node confirmation state. The target unconfirmed nodes refer to the nodes in the scene task nodes with the unconfirmed state of the node confirmation state.

[0126] Referring to Figure 5 , Figure 5is a specific flowchart of step S410 provided in the embodiments of the present application. In some embodiments of the present application, step S410 can specifically include but is not limited to steps S510 to S530.

[0127] In step S510, a historical confirmed node is selected from the historical task nodes based on the node confirmation state.

[0128] In step S520, the historical confirmed node and the scene task node are matched to obtain a node matching result.

[0129] In step S530, a target confirmed node and a target unconfirmed node are determined from the scene task node based on the node matching result.

[0130] In steps S510 to S530 of some embodiments, since the historical scene task and the target scene task can be the same or different, the historical task node and the scene task node can also be the same or different. Thus, the historical confirmed node (i.e., the node corresponding to the confirmed state) can be selected from the historical task nodes based on the node confirmation state first, and then the historical confirmed node and the scene task node are matched to obtain a node matching result, which is used to indicate whether the scene task node has a historical task node that has been confirmed to be completed.

[0131] In step S420 of some embodiments, the node confirmation data refers to the related object data of the corresponding target confirmed node. For example, the target confirmed node is an identity confirmation node, and the node confirmation data is the related identity information previously stored by the target object, which is used to generate a complete target outbound call script.

[0132] In steps S430 to S450 of some embodiments, further, the opening speech of the target outbound call can be generated using the speech generation model in combination with the confirmed nodes, the node confirmation data, and the target unconfirmed node. This speech can reflect the information confirmed by the target object before, so that the conversation is more smooth and personalized. The speech generation model of the present application can be constructed based on a large model, and a prompt text can be generated based on the confirmed nodes, the node confirmation data, and the target unconfirmed node, and the prompt text can be input into the speech generation model to obtain the target outbound call speech. For example, the prompt text is "I am AI phone customer service XX, there is a customer XXX, I need to remind him of

task 1

task 2

node 1

node 2

node 1

node 1

[0133] In the above embodiments, the present application can combine the opening speech, the confirmed nodes, and the unconfirmed node speech to form a complete target outbound call speech. This combined speech will be used in the actual outbound call conversation, which can flexibly guide the conversation process according to the different responses of the user, ensure that the conversation can effectively cover all key nodes, whether confirmed or unconfirmed, avoid the customer answering the same question multiple times, and at the same time, the outbound call speech can be customized for each customer, thereby further improving the customer satisfaction.

[0134] Please refer to Figure 6 , Figure 6 is another specific flowchart of the speech generation method based on intelligent outbound call provided by the embodiments of the present application. In some embodiments of the present application, after step S160, the speech generation method based on intelligent outbound call of the present application can further include but is not limited to steps S610 to S630.

[0135] Step S610, obtaining the task priority of the target scene task;

[0136] Step S620, matching the target scene task based on the task priority to determine the target interaction mode;

[0137] Step S630, intelligent outbound call to the target object based on the target outbound call speech and the target interaction mode.

[0138] In step S610 of some embodiments, the application can also intelligently adjust the touch channel of each outbound call, that is, the interaction mode. Among them, the interaction mode with the target object can include AI outbound call, APP reminder, manual outbound call, etc. The task priority is the importance of the current target scene task pre-configured, and the task priority can be determined based on various factors, such as the urgency, importance of the task, historical response mode of the user, business rules, etc.

[0139] In step S620 of some embodiments, the target interaction mode of the current target scene task and the target object can be determined based on the task priority and the interaction mode matching of the target scene task. The target interaction mode is used to indicate the communication mode adopted to confirm the target task node to the target object. Among them, the determination of the target interaction mode can depend on various factors, such as the nature of the task, the user's preference, the expected response rate, cost-benefit analysis, etc. For example, for high-priority tasks, the system may choose a more direct and immediate communication mode such as telephone call; while for low-priority tasks, it may choose a lower-cost communication mode such as email or SMS.

[0140] Please refer to Figure 7 , Figure 7 is a specific flowchart of step S620 provided by the embodiments of the application. In some embodiments of the application, step S620 can specifically include but is not limited to steps S710 to S740.

[0141] Step S710, obtaining a pre-allocated interaction mode of a target scene task;

[0142] Step S720, performing saturation degree division on a preset interaction reference number of the pre-allocated interaction mode based on a preset saturation degree division function, to obtain a saturation degree division interval;

[0143] Step S730, performing interval matching on the number of to-be-allocated interaction objects of the pre-allocated interaction mode based on the saturation degree sub-interval, to determine a matching interval of the pre-allocated interaction mode;

[0144] Step S740, if the matching interval is the first saturation interval, performing interaction mode adjustment on the pre-allocated interaction mode based on the task priority, to obtain a target interaction mode.

[0145] In step S710 of some embodiments, considering that the current tasks of different interaction modes may be unbalanced, the application can calculate the saturation degree of different interaction modes to intelligently schedule according to the saturation degree, thereby improving the efficiency of intelligent outbound call. The pre-allocated interaction mode refers to a pre-configured interaction mode for the target scene task.

[0146] In step S720 of some embodiments, further, the application can evaluate the efficiency and capacity of the pre-allocated interactive mode using a preset saturation degree division function. This function can divide different saturation degree division intervals according to the preset interactive benchmark number M (such as the monthly average number of processed policies, etc.) of the pre-allocated interactive mode. Saturation degree can be understood as the busy degree or availability of the interactive mode. Based on the preset saturation degree division function, a first saturation interval and a second saturation interval can be divided, where the saturation degree of the first saturation interval is higher than that of the second saturation interval, meaning that the interactive mode in the first saturation interval is more busy or close to its capacity upper limit. For example, the first saturation interval corresponds to a medium-high saturation interval, and the number of to-be-allocated interactive objects N in the medium-high saturation interval is greater than or equal to M*30% (where 30% is a first interval adjustment parameter, which can be flexibly adjusted according to actual needs), and the first saturation interval corresponds to a low saturation interval, and the number of to-be-allocated interactive objects N in the low saturation interval is less than M*30%.

[0147] In step S730 of some embodiments, the number of to-be-allocated interactive objects refers to data that needs to be communicated using the pre-allocated interactive mode, but has not been allocated yet, such as the number of to-be-allocated policies. Further, the number of to-be-allocated interactive objects of the pre-allocated interactive mode can be interval-matched based on the saturation degree sub-interval to determine the saturation interval to which the number of to-be-allocated interactive objects of the pre-allocated interactive mode belongs, i.e., the matching interval of the pre-allocated interactive mode.

[0148] In step S740 of some embodiments, if the matching interval is the first saturation interval, i.e., the pre-allocated interactive mode of the target scenario task belongs to the medium-high saturation interval, then the pre-allocated interactive mode corresponding to the to-be-allocated interactive object of the target scenario task needs to be adjusted based on the task priority of the target scenario task to obtain the target interactive mode of the target scenario task. Specifically, if the pre-allocated interactive mode of the target scenario task belongs to the medium-high saturation interval, and the task priority of the target scenario task is greater than the preset priority, then the pre-allocated interactive mode needs to be adjusted. If the pre-allocated interactive mode of the target scenario task belongs to the medium-high saturation interval, and the task priority of the target scenario task is less than or equal to the preset priority, then the pre-allocated interactive mode does not need to be adjusted.

[0149] For example, the target scene task is "renewal payment reminder", the corresponding task priority is 8, the pre-allocated interaction mode is "manual outbound call", the preset priority is 7, according to steps S720 and S730, it is determined that the matching interval of the pre-allocated interaction mode of the target scene task belongs to the first saturated interval, and the task priority of the pre-allocated interaction mode is greater than 7, at this time, it is not necessary to adjust the pre-allocated interaction mode of the target scene task, and "manual outbound call" is taken as the target interaction mode. If the task priority corresponding to the target scene task is 7, it is determined that the matching interval of the pre-allocated interaction mode of the target scene task belongs to the first saturated interval, and the task priority of the pre-allocated interaction mode is equal to 7, at this time, it is necessary to update and adjust the pre-allocated interaction mode of the target scene task, and the updated target interaction mode can be "AI outbound call".

[0150] Further, the first saturated interval includes a first saturated sub-interval and a second saturated sub-interval, the saturation degree of the first saturated sub-interval is higher than the saturation degree of the second saturated sub-interval. At this time, the first saturated sub-interval corresponds to the high saturated interval, and the second saturated sub-interval corresponds to the low saturated interval. For example, the first saturated interval corresponds to the medium-high saturated interval, and the number of to-be-allocated interaction objects N in the medium-high saturated interval is greater than or equal to M*30% (wherein 30% is the first interval adjustment parameter, which can be flexibly adjusted according to actual needs). Further, the second interval adjustment parameter can be 80%, at this time, the number of to-be-allocated interaction objects N in the first saturated sub-interval is greater than or equal to M*80% (wherein 80% is the second interval adjustment parameter, which can be flexibly adjusted according to actual needs), and the number of to-be-allocated interaction objects N in the second saturated sub-interval is greater than or equal to M*30% and less than M*80%.

[0151] Based on this, step S740 can specifically include the following cases:

[0152] If the matching interval is the first saturated sub-interval, the pre-allocated interaction mode is adjusted to obtain the target interaction mode, and the target interaction mode is different from the pre-allocated interaction mode;

[0153] If the matching interval is the second saturated sub-interval and the task priority is greater than the preset priority, the pre-allocated interaction mode is taken as the target interaction mode;

[0154] If the matching interval is the second saturated sub-interval and the task priority is less than or equal to the preset priority, the pre-allocated interaction mode is adjusted to obtain the target interaction mode, and the target interaction mode is different from the pre-allocated interaction mode.

[0155] It should be noted that the preset saturation degree partition function can be used to divide the saturation degree partition interval into three intervals, a first saturation sub-interval (a high saturation interval), a second saturation sub-interval (a medium saturation interval), and a second saturation interval (a low saturation interval). In addition, the first interval adjustment parameter and the second interval adjustment parameter can be adjusted to update the number of intervals and the specific partition of the intervals, which is not limited in particular.

[0156] In step S630 of some embodiments, further, after determining the target interaction mode, the target object can be called in the target interaction mode for communication in combination with the target call script generated in advance. In this way, the intelligent call of the present application can be performed according to the preset target call script and communication strategy, while the response of the target object can be detected in real time, and the script or interaction strategy can be adjusted as needed. For example, if the target interaction mode is "AI call", if the target object does not answer the phone, the system can automatically send a short message reminder.

[0157] In the above embodiments, the intelligent call system of the present application can flexibly select the most suitable interaction mode according to the task priority and the characteristics of the target object, improve the efficiency and effect of communication, and make the call activity more personalized and target-oriented.

[0158] It should be noted that the non-company software tools or components appearing in the embodiments of the present application are only examples and do not represent actual use.

[0159] The method for generating a call script based on intelligent call provided by the embodiments of the present application proposes a user behavior analysis based on multi-interaction mode human-computer cooperation call strategy. On the one hand, the opening script of the next call can be dynamically generated according to the node information confirmed by the customer (i.e., the node confirmation state of the target object to the historical task node is determined based on the node reply intention category and category matching data), which avoids the customer's multiple replies to the same problem, and at the same time, a personalized call script is customized for each customer to improve customer satisfaction. On the other hand, the saturation degree of each interaction mode is dynamically predicted to adjust the interaction mode with the target object in real time, realize the cooperation of artificial and AI, and improve the information reach rate and efficiency of the call task.

[0160] Please refer to Figure 8 The embodiments of the present application also provide a call script generation device based on intelligent call, which can implement the above-mentioned call script generation method based on intelligent call. The device comprises:

[0161] The first acquisition module 810 is configured to acquire a target scene task of a target object, and the target scene task comprises a scene task node.

[0162] The second acquisition module 820 is configured to acquire historical outbound voice of the target object in a historical scene task, the historical scene task including a historical task node and a historical node dialogue text of the historical task node;

[0163] The extraction module 830 is configured to perform voice extraction on the historical outbound voice and the historical node dialogue text based on a preset voice extraction model to obtain historical object voice, the historical object voice being voice of the target object in reply to the historical node dialogue text in a historical outbound dialogue;

[0164] The recognition module 840 is configured to perform intent recognition on the historical object voice based on an intent recognition model to obtain a node reply intent category of the historical task node and category matching data, the category matching data being used to indicate a distance between the node reply intent category and voice text of the historical object voice.

[0165] The determination module 850 is configured to determine a node confirmation state of the target object to the historical task node based on the node reply intent category and the category matching data.

[0166] The generation module 860 is configured to generate target outbound dialogue based on the node confirmation state and the scene task node.

[0167] The specific implementation of the dialogue generation apparatus based on intelligent outbound of the embodiments of the present application is basically the same as that of the above-mentioned dialogue generation method based on intelligent outbound, and will not be repeated here.

[0168] The embodiments of the present application further provide an electronic device, which includes a memory and a processor, the memory storing a computer program, and the processor implementing the dialogue generation method based on intelligent outbound when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0169] Please refer to Figure 9 , Figure 9 The hardware structure of the electronic device of another embodiment is illustrated, which includes:

[0170] The processor 910 can be implemented in a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0171] The memory 920 can be implemented in the form of a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 920 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 920 and are called and executed by the processor 910 to implement the script generation method based on intelligent outbound calls according to the embodiments of the present application;

[0172] The input / output interface 930 is configured to realize information input and output.

[0173] The communication interface 940 is configured to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, etc.) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, etc.).

[0174] The bus 950 is configured to transmit information between various components (for example, the processor 910, the memory 920, the input / output interface 930, and the communication interface 940) of the device.

[0175] The processor 910, the memory 920, the input / output interface 930, and the communication interface 940 are connected to each other through the bus 950 to realize the communication connection between the devices.

[0176] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the script generation method based on intelligent outbound calls.

[0177] The memory is a non-transitory computer readable storage medium, which can be used to store a non-transitory software program and a non-transitory computer executable program. In addition, the memory can include a high-speed random access memory and can also include a non-transitory memory, for example, at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged relative to the processor. These remote memories can be connected to the processor through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0178] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0179] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and can include more or fewer steps than the figures, or combine certain steps, or different steps.

[0180] The apparatus embodiments described above are merely illustrative, and units described as separate components can or can not be physically separate, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments.

[0181] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the function modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.

[0182] The terms "first", "second", "third", "fourth" and the like used in the description of the present application and the above-described figures (if any) are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0183] It should be understood that in the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" is used to describe the association between the associated objects, which means that there can be three relationships, for example, "A and / or B" can represent three cases: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can mean a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0184] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other manners. For example, the apparatus embodiments described above are merely illustrative, for example, the division of the above units is merely a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0185] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0186] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0187] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that makes a contribution or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program storage media.

[0188] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, but this does not limit the scope of the embodiments of the present application. Any modification, equivalent replacement and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.

Claims

1. A script generation method based on intelligent outbound call, characterized in that, The method comprises: obtaining a target scene task of a target object, the target scene task comprising a scene task node; obtaining historical outbound call voice of the target object in a historical scene task, the historical scene task comprising a historical task node and a historical node dialogue text of the historical task node; performing voice extraction on the historical outbound call voice and the historical node dialogue text based on a preset voice extraction model to obtain historical object voice, the historical object voice being voice of the target object in a historical outbound call dialogue in reply to the historical node dialogue text; performing intent recognition on the historical object voice based on an intent recognition model to obtain a node reply intent category of the historical task node and category matching data, the category matching data being used to indicate a distance between the node reply intent category and voice text of the historical object voice; determining a node confirmation state of the target object for the historical task node based on the node reply intent category and the category matching data; generating a target outbound call dialogue based on the node confirmation state and the scene task node.

2. The method of claim 1, wherein, The generating of the target outbound call dialogue based on the node confirmation state and the scene task node comprises: performing node division on the scene task node based on the node confirmation state to obtain a target confirmed node and a target unconfirmed node; obtaining node confirmation data of the target object for the target confirmed node; performing dialogue generation on the target confirmed node and the node confirmation data based on a dialogue generation model to obtain an outbound call opening dialogue; performing dialogue generation on the target unconfirmed node based on the dialogue generation model to obtain an outbound call unconfirmed node dialogue; performing dialogue combination based on the outbound call opening dialogue and the outbound call unconfirmed node dialogue to obtain the target outbound call dialogue.

3. The method of claim 2, wherein, The performing of the node division on the scene task node based on the node confirmation state to obtain the target confirmed node and the target unconfirmed node comprises: selecting a historical confirmed node from the historical task node based on the node confirmation state; performing node matching on the historical confirmed node and the scene task node to obtain a node matching result; determining the target confirmed node and the target unconfirmed node from the scene task node based on the node matching result.

4. The method of claim 1, wherein, The method further comprises: obtaining a task priority of the target scene task; performing interaction mode matching on the target scene task based on the task priority to determine a target interaction mode, the target interaction mode being used to indicate a communication mode adopted for confirming the target task node to a target object; performing intelligent outbound call on the target object based on the target outbound call dialogue and the target interaction mode.

5. The method of claim 4, wherein, The performing of the interaction mode matching on the target scene task based on the task priority to determine the target interaction mode comprises: obtaining a pre-allocated interaction mode of the target scene task; The preset interaction reference number of the pre-allocated interaction mode is divided based on a preset saturation degree division function to obtain a saturation degree division interval, the saturation degree division interval including a first saturation interval and a second saturation interval, and a saturation degree of the first saturation interval is higher than a saturation degree of the second saturation interval; An interval matching is performed on the to-be-allocated interaction object number of the pre-allocated interaction mode based on the saturation degree sub-interval to determine a matching interval of the pre-allocated interaction mode; If the matching interval is the first saturation interval, an interaction mode adjustment is performed on the pre-allocated interaction mode based on the task priority to obtain the target interaction mode.

6. The method according to any one of claims 1 to 5, characterized in that, The historical object voice is obtained by performing voice extraction on the historical outbound call voice and the historical node script text based on a preset voice extraction model, and the historical object voice is a voice of the target object in a historical outbound call conversation in which the target object replies to the historical node script text. The historical outbound call voice is subjected to noise reduction processing to obtain a noise-reduced outbound call voice; The noise-reduced outbound call voice is subjected to voice coding to obtain outbound call voice features; The outbound call voice features are subjected to voice feature separation to obtain outbound object voice features, the outbound object voice features being used to indicate different speaking objects in the historical outbound call voice; The outbound object voice features are subjected to feature decoding to obtain outbound object voice segments; A reply object voice segment is selected from the outbound object voice segments based on the historical node script text, the reply object voice segment being a voice segment in which the target object replies to the historical node script text in the historical outbound call conversation; The reply object voice segment is subjected to segment integration to obtain the historical object voice.

7. The method of claim 6, wherein, The historical task node is subjected to intent recognition based on an intent recognition model to obtain a node reply intent category of the historical task node and category matching data, and the method includes the following steps: A node verification category of the historical task node corresponding to the reply object voice segment is obtained, the node verification category including an intent recognition category and a labeled category; If the node verification category is the intent recognition category, the reply object voice segment is subjected to intent recognition based on an intent recognition model to obtain the node reply intent category and the category matching data; If the node verification category is the labeled category, a preset category is determined to be the node reply intent category of the reply object voice segment, and a preset matching data is determined to be the category matching data of the reply object voice segment.

8. An apparatus for generating a script based on intelligent outbound call, characterized by, The device includes: A first obtaining module is configured to obtain a target scene task of a target object, the target scene task including a scene task node; A second obtaining module is configured to obtain a historical outbound call voice of the target object in a historical scene task, the historical scene task including a historical task node and a historical node script text of the historical task node; An extraction module is configured to perform voice extraction on the historical outbound call voice and the historical node script text based on a preset voice extraction model to obtain a historical object voice, the historical object voice being a voice of the target object in a historical outbound call conversation in which the target object replies to the historical node script text. An identification module is configured to perform intent identification on the historical object voice based on an intent identification model to obtain a node reply intent category of the historical task node and category matching data, the category matching data being used to indicate a distance between the node reply intent category and speech text of the historical object voice; A determination module is configured to determine a node confirmation state of the target object to the historical task node based on the node reply intent category and the category matching data; A generation module is configured to generate a target outbound call script based on the node confirmation state and the scene task node.

9. An electronic device, comprising: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the method in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Outbound method and device based on intention classification, equipment and medium

    CN117278675A

  • Intelligent dialogue method and device, equipment and medium

    CN117633191A