Training method of call prompt tone classification model, intelligent outbound method and electronic equipment
By identifying and retraining samples whose differences in the call prompt tone classification model are greater than or equal to the preset threshold, the classification accuracy of the existing Chinese and foreign call prompt tone classification model in complex audio environments is solved, and more efficient intelligent outgoing calls are achieved.
Patent Information
- Application Number
- CN202510270981.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing outgoing call prompt classification model is susceptible to interference from factors such as background noise, muted clips and audio quality differences when facing complex audio environments, resulting in a decrease in classification accuracy and thus reducing the efficiency of intelligent outgoing call.
By determining the model category text and regular category text of multiple call prompt tones, the call prompt tones with a difference of greater than or equal to the preset threshold are identified as retraining samples, and the call prompt tones classification model is retrained to improve anti-interference ability and classification accuracy.
The call prompt classification model has improved its anti-interference ability to factors such as background noise, silent clips and audio quality differences, significantly improved the classification ability, thereby improving the efficiency of intelligent outgoing calls.
Smart Images

Figure CN120199271A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent outbound calling, and particularly relates to a method for training a call prompt tone classification model, an intelligent outbound calling method, and an electronic device. Background Art
[0002] In the prior art, intelligent outbound calling refers to the automated outbound calling by telephone implemented by AI (Artificial Intelligence) electronic devices, which is used to achieve more efficient, accurate, and personalized marketing promotion or customer communication. For example, intelligent outbound calling can notify target users of the company's promotion activities, etc., and for another example, intelligent outbound calling can also conduct follow-up visits to target users through voice calls. Therefore, intelligent outbound calling has very broad application prospects.
[0003] During the process of intelligent outbound calling by an AI electronic device, it is necessary to identify the call prompt tones emitted by the called telephone device; the call prompt tone refers to the audio information generated by the communication network or the called telephone device during the process of making an active call through the telephone device to inform the current call status, such as "the number you dialed is temporarily unavailable", etc. The AI electronic device needs to identify and classify the above call prompt tones to determine subsequent outbound calling strategies. For example, when the call prompt tones are respectively "the number you dialed is an invalid number" and "the number you dialed is temporarily unavailable", the subsequent outbound calling strategies of the AI electronic device will obviously make differences.
[0004] In the prior art, an AI electronic device identifies and classifies call prompt tones through an outbound call prompt tone classification model. However, according to actual experience, the existing outbound call prompt tone classification models are easily interfered by factors such as background noise, silent segments, and audio quality differences in the face of complex audio environments, resulting in a decrease in classification accuracy and further reducing the efficiency of intelligent outbound calling. Summary of the Invention
[0005] In view of this, the purpose of the present application is to provide a method for training a call prompt tone classification model, an intelligent outbound calling method, and an electronic device to solve the technical problem that the existing outbound call prompt tone classification models are easily interfered by factors such as background noise, silent segments, and audio quality differences in the face of complex audio environments, resulting in a decrease in classification accuracy.
[0006] In a first aspect, the present application provides a method for training a call prompt tone classification model, the method comprising:
[0007] Determine the model category text and regular category text of multiple call prompt tones;
[0008] Among them, the call prompt tone is collected during the active call; the model category text is determined by identifying and classifying through a trained call prompt tone classification model; the regular category text is determined according to a regular expression;
[0009] Determine the call prompt tones with the difference degree between the model category text and the regular category text in multiple call prompt tones greater than or equal to a preset threshold as retraining samples;
[0010] Retrain the call prompt tone classification model according to the retraining samples.
[0011] In a second aspect, the present application provides an intelligent outbound call method, which is applied to an intelligent outbound call device, and a call prompt tone classification model is deployed in the intelligent outbound call device. The call prompt tone classification model is obtained by training through the training method of the call prompt tone classification model described above; the method includes:
[0012] During the process of making an active call through the intelligent outbound call device, collect the call prompt tone of the target telephone device to be called to obtain a target call prompt tone;
[0013] Input the target call prompt tone into the call prompt tone classification model to obtain the target model category text corresponding to the target call prompt tone;
[0014] Among them, the model category text indicates the category of the call prompt tone;
[0015] Determine whether it is necessary to make an active call to the target telephone device again according to the target model category text.
[0016] In a third aspect, the present application provides a training device for a call prompt tone classification model. The device includes: a text determination module, a sample screening module, and a training module;
[0017] The text determination module is used to determine the model category text and the regular category text of multiple call prompt tones;
[0018] Among them, the call prompt tone is collected during the active call; the model category text is determined by identifying and classifying through a trained call prompt tone classification model; the regular category text is determined according to a regular expression;
[0019] The sample screening module is used to determine the call prompt tones with the difference degree between the model category text and the regular category text in multiple call prompt tones greater than or equal to a preset threshold as retraining samples;
[0020] The training module is used to retrain the call prompt tone classification model according to the retraining samples.
[0021] In a fourth aspect, the present application provides an intelligent outbound calling device, which is applied to an intelligent outbound calling device, and a call prompt tone classification model is deployed in the intelligent outbound calling device. The call prompt tone classification model is obtained by training through the above-mentioned call prompt tone classification model training method; the device includes: a prompt tone acquisition module, a classification module, and a call determination module;
[0022] The prompt tone acquisition module is used to collect a target call prompt tone for the call prompt tone of the target telephone device being called during the process of making an active call through the intelligent outbound calling device;
[0023] The classification module is used to input the target call prompt tone into the call prompt tone classification model to obtain a target model category text corresponding to the target call prompt tone;
[0024] Among them, the model category text indicates the category of the call prompt tone;
[0025] The call determination module is used to determine whether it is necessary to make an active call to the target telephone device again according to the target model category text.
[0026] In a fifth aspect, the present application provides an electronic device, which includes a processor and a memory. The memory is used to store application programs, and the processor runs or executes the software programs stored in the memory to enable the electronic device to implement the above-mentioned call prompt tone classification model training method or intelligent outbound calling method.
[0027] Beneficial effects:
[0028] The present application provides a training method for a call prompt tone classification model. The method includes: determining model category texts and regular category texts of multiple call prompt tones; among them, the call prompt tones are collected during the active call process; the model category texts are determined by identifying and classifying through a trained call prompt tone classification model; the regular texts are determined according to regular expressions; determining call prompt tones with a difference degree greater than or equal to a preset threshold between the model category texts and the regular category texts among the multiple call prompt tones as retraining samples; and retraining the call prompt tone classification model according to the retraining samples;
[0029] In summary, in the present application, regular category texts are introduced to proofread model category texts, and the call prompt sounds corresponding to the model category texts and regular category texts with a difference degree greater than or equal to a preset threshold in the proofreading results are determined as retraining texts, which are used to retrain the call prompt sound classification model, so as to improve the anti-interference ability of the call prompt sound classification model against factors such as background noise, silent segments, and audio quality differences, greatly improve the classification ability of the call prompt sound classification model, and further improve the efficiency of intelligent outbound calls. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for use in the embodiments of the present application will be briefly introduced below. The following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0031] Figure 1 It is a schematic flowchart of a method for training an outbound call prompt sound classification model provided by an embodiment of the present application;
[0032] Figure 2 It is a comparative example diagram of the prediction probabilities of a first model and a VAD algorithm for non-speech segments provided by an embodiment of the present application;
[0033] Figure 3 It is a schematic flowchart of an intelligent outbound call method provided by an embodiment of the present application;
[0034] Figure 4 It is a schematic structural diagram of a training device for a call prompt sound classification model provided by an embodiment of the present application;
[0035] Figure 5 It is a schematic structural diagram of an intelligent outbound call device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In the current era of the rapid development of AI technology, intelligent outbound calls relying on AI robots (or other intelligent outbound call devices) have application advantages such as high efficiency and cost savings, and thus have a very broad application prospect. For example, they have corresponding applications in aspects such as marketing, sales promotion, customer service, and survey feedback.
[0037] During the intelligent outbound call process, the AI robot making the outbound call needs to identify and classify the outbound call prompt sounds collected during the current outbound call process, so as to convert the outbound call prompt sounds in the form of "natural language" into a form that can be understood by the AI robot, facilitating the determination of the subsequent call strategy for the called user during the current outbound call process. For example, if the identification and classification result of the outbound call prompt sound collected during the current outbound call process is "busy", then the AI robot can determine the called user as a "user who needs to be called again", and wait for a period of time before making an intelligent outbound call again; if the identification and classification result of the outbound call prompt sound collected during the current outbound call process is "invalid number", then the AI robot can determine the called user as a "user who does not need to be called again".
[0038] In practical applications, the AI electronic device identifies and classifies the call prompt sounds through an outbound call prompt sound classification model. However, according to practical experience, the existing outbound call prompt sound classification models are easily interfered by factors such as background noise, silent segments, and audio quality differences in the face of complex audio environments, resulting in a decrease in classification accuracy and further reducing the efficiency of intelligent outbound calls.
[0039] It should be emphasized that in actual operation, although the outbound call prompt sound classification model is used to identify and classify call prompt sounds, the data directly input into the outbound call prompt sound classification model are the audio features obtained by extracting features from the call prompt sounds, such as Mel-Frequency Cepstral Coefficients (MFCC), etc.
[0040] To solve the above technical problems, this application provides a training technical solution for an outbound call prompt sound classification model, which is used to solve the problem that the existing outbound call prompt sound classification models are easily interfered by factors such as background noise, silent segments, and audio quality differences in the face of complex audio environments, and further improve the generalization of the outbound call prompt sound classification model.
[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0042] First, this application provides a training method for an outbound call prompt sound classification model, as Figure 1 shown Figure 1 is the flowchart of the training method for the outbound call prompt sound classification model provided by the embodiments of this application. The method includes: S110 to S130, details are as follows:
[0043] S110: Determine the model category text and regular category text of multiple call prompt tones;
[0044] Among them, the call prompt tone is collected during an active call; the model category text is determined by identification and classification through a trained call prompt tone classification model; the regular category text is determined according to a regular expression.
[0045] Specifically, the call prompt tone is collected during the process of the AI electronic device making an active outbound call; in the embodiment of the present application, the regular category text of the call prompt tone is used as a comparison for the model category text of the call prompt tone to determine whether the model category text of the actual meaning relative to the call prompt tone is incorrect; among them, the model category text is determined by identification and classification through a trained call prompt tone classification model; the regular category text is determined according to a regular expression.
[0046] In one implementation, before S110, the method further includes: step (1), the details are as follows:
[0047] Step (1): Process the call prompt tones with two channels among multiple original call prompt tones to obtain call prompt tones with a single channel.
[0048] Specifically, the original call prompt tone is the original version of the call prompt tone, that is, the call prompt tone without any technical processing; in order to improve the accuracy of the obtained model category text and regular category text, the embodiment of the present application detects the number of channels of the audio of the original call prompt tone, processes the call prompt tones with two channels among multiple original call prompt tones to obtain call prompt tones with a single channel, and then uses the call prompt tones with a single channel to determine the model category text and regular category text.
[0049] In actual operation, the number of channels of the audio of the original call prompt tone is detected by running the ffmpeg command in the linux operating system. For example, for the audio file of the original call prompt tone with the file name input_audio.mp3, the ffmpeg command is run in the linux operating system to detect the number of channels of the original call prompt tone. Among them, the ffmpeg command is as follows:
[0050] ffmpeg -i input_audio.mp3 2>&1 | grep 'Stream.*Audio';
[0051] Among them, "-i input_audio.mp3" represents the audio file of the original call prompt tone;
[0052] "2>&1" means redirecting the standard error output to the standard output so that all information can be captured and displayed in the Linux operating system;
[0053] "grep 'Stream.*Audio'" means running the grep command to filter out the lines containing the audio stream information of the original call prompt tone;
[0054] In actual operation, the detection result of the number of channels can be obtained by running the above ffmpeg command, and the detection result is as follows:
[0055] Stream#0:0:Audio:mp3,44100Hz,stereo,fltp,128kb / s;
[0056] Among them, "stereo" means two channels; if "mono" is displayed at the corresponding position, it indicates that the original call prompt tone is mono.
[0057] In one implementation, between step (1) and S110, the method further includes: step (2), the details are as follows:
[0058] Step (2): Removing the silent segments in the call prompt tone through the first model;
[0059] Among them, the first model is obtained by training a convolutional-long short-term memory-fully connected deep neural network.
[0060] Specifically, in the embodiment of the present application, the first model is used to identify the silent segments and / or useless segments in the call prompt tone, and then the identified silent segments and / or useless segments are removed; the first model is obtained by training a convolutional-long short-term memory-fully connected deep neural network (Convolutional, Long Short-Term Memory, FullyConnected Deep Neural Networks; CLDNN); the CLDNN model combines the advantages of a convolutional neural network (Convolutional Neural Network, CNN), a long short-term memory network (Long Short-Term Memory, LSTM), and a fully connected layer (Deep Neural Networks, DNN).
[0061] The bottom layer in the network structure of the first model is a two-layer CNN network. The CNN can effectively extract the time-domain features of audio and can effectively extract the audio features in the current time domain and adjacent time domains. After the CNN network, two layers of LSTM networks are connected. The LSTM network can filter out long-term background noise through the memory and forgetting mechanisms. Finally, a single layer of DNN maps the audio features to speech or non-speech nodes. Thus, the first model has multiple advantages, which are as follows:
[0062] 1) Time-domain feature extraction ability:
[0063] The first model uses the CNN network to extract the time-domain features of speech signals, captures the time-series dependencies of speech signals through the LSTM network, and performs non-linear mapping of complex features through the DNN network. Therefore, the first model can comprehensively capture the spatio-temporal features of speech signals and improve the accuracy of feature extraction;
[0064] 2) Robustness:
[0065] The CNN network in the first model effectively extracts local features, improving the anti-interference ability of the first model against local noise. The LSTM network filters background noise through the memory and forgetting mechanisms, improving the robustness of the first model in a noisy environment;
[0066] 3) Efficient feature learning:
[0067] The CNN network in the first model is used to reduce the input feature dimension, the LSTM network is used for time-series modeling, and the DNN network is used for deep learning of high-dimensional features, enabling the first model to efficiently learn and recognize speech activities;
[0068] 4) Real-time performance:
[0069] The CNN network and LSTM network in the first model can process and analyze speech signals in a short time, meeting the real-time requirements;
[0070] 5) Strong generalization ability:
[0071] The first model integrates multiple neural network structures and thus has good generalization ability, being able to adapt to different speech data and application scenarios. During the training process, the first model learns general features and can perform well on different test data; as Figure 2 shown, Figure 2This is a comparison example diagram of the prediction probabilities of the first model provided by the embodiments of this application and the VAD (Voice Activity Detection) algorithm for non-speech segments. Non-speech segments refer to silent segments and / or useless segments; the VAD algorithm can separate the speech signal and background silence in the call prompt tone, retain the meaningful part in the call prompt tone, and reduce the processing burden for subsequent text recognition.
[0072] Figure 2 In Series 1 in, it represents the curve graph of the prediction probability obtained by the VAD algorithm for non-speech segments, and Series 2 represents the curve graph of the prediction probability obtained by the first model for non-speech segments. According to Figure 2 it can be seen that Series 2 can identify speech segments and non-speech segments more clearly compared to Series 1. For example, it shows a square area as shown in Figure 2 indicating that compared with the VAD algorithm, the first model has stronger prediction ability for non-speech segments.
[0073] In the embodiments of this application, the process of training the first model includes: Step (2.1) to Step (2.4), the details are as follows:
[0074] Step (2.1): Obtain multiple speech data for training the CLDNN model;
[0075] Step (2.2): Extract features for each of the multiple speech data to obtain corresponding audio features, such as MFCC, etc.;
[0076] Step (2.3): Determine the silent segments in each of the multiple speech data, and label the silent segments to obtain labels corresponding to each speech data;
[0077] Among them, the label of the speech data includes the labels of all time frames included in the speech data. The label of a time frame refers to the annotation data indicating whether the time frame is a time frame in the silent segment. For example, if the time frame belongs to the time frame in the silent segment, the annotation data of the time frame can be 1, and if the time frame does not belong to the time frame in the silent segment, the annotation data of the time frame can be 0;
[0078] In actual operation, the labels of all time frames included in the speech data can be organized into a numerical sequence or a two-dimensional vector; it should also be emphasized that the labels of each time frame in the label of the speech data need to be time-aligned with the audio features corresponding to the speech data;
[0079] Step (2.4): Use the audio features of each voice data as training samples and input them together with the corresponding labels into the CLDNN model for iterative training of the CLDNN model until the training stop condition is reached;
[0080] Among them, the training stop condition can be reaching a preset prediction accuracy rate or reaching the maximum number of iterations, etc., and the present application does not make specific limitations in this regard.
[0081] In actual operation, when the iterative training of the CLDNN model reaches the training stop condition, the first model can be obtained. In one implementation manner, S110 includes: Step (3) to Step (4), and the details are as follows:
[0082] Step (3): Perform speech recognition processing on the call prompt tone to obtain a speech recognition text;
[0083] Among them, the speech recognition processing is implemented based on the ASR (Automatic Speech Recognition) technology; the ASR technology is responsible for accurately converting the input audio file of the call prompt tone into a text file; during the speech recognition processing, the timestamp information of the call prompt tone needs to be emphasized. The timestamp can record the start time and end time of each call prompt tone, which is used to ensure the logical coherence of the text file of the call prompt tone, that is, the purpose of retaining the timestamp is to ensure that multiple sentences in the text file of the call prompt tone can be correctly spliced and the specific speaking time of each sentence can be understood.
[0084] In one implementation manner, Step (3) includes: Step (3.1), and the details are as follows:
[0085] Step (3.1): Perform speech recognition processing on the call prompt tone through the second model to obtain a speech recognition text;
[0086] Among them, the second model is obtained by training the Zipformer network structure.
[0087] Specifically, in the embodiment of the present application, when applying the ASR technology for speech recognition processing, it is based on the Zipformer network structure to achieve real-time, high-precision and low-resource-consumption speech recognition; among them, the speech recognition processing based on the Zipformer network structure has multiple advantages, and the details are as follows:
[0088] 1) High-precision recognition:
[0089] The Zipformer network structure can retain the key features of the speech signal during the compression process, so that speech recognition processing can still be efficiently performed even in a bandwidth-limited environment;
[0090] 2) Real-time processing:
[0091] The Zipformer network structure has the ability to respond quickly when compressing and processing speech data; real-time processing enables instant speech recognition processing;
[0092] 3) Resource saving:
[0093] The Zipformer network structure has the special effect of efficient compression, so it can significantly reduce the storage and transmission requirements for speech recognition processing, reduce the occupation of system resources, and enable this speech recognition processing technology to be deployed on devices with limited resources;
[0094] 4) Strong robustness:
[0095] The Zipformer network structure performs excellently in processing different noise and interference environments. Its powerful feature extraction and compression capabilities enable the Zipformer network structure to maintain stable speech recognition processing under various noise conditions; In summary, through the speech recognition processing that retains key timestamp information, the embodiments of this application improve the processing ability and processing efficiency compared with the existing speech recognition processing.
[0096] In the embodiments of this application, the process of training the second model includes: steps (3.2) to (3.5), and the details are as follows:
[0097] Step (3.2): Obtain multiple speech data for training the Zipformer network structure;
[0098] Step (3.3): Extract features for each of the multiple speech data to obtain corresponding audio features, such as MFCC, etc.;
[0099] Step (3.4): Determine the actual text of each of the multiple speech data, and use the actual text as the label of the corresponding speech data;
[0100] Among them, the label of the speech data should be a character sequence, and the characters included in the character sequence can represent the corresponding text through a preset index table. For example, if the sixth character in the character sequence is 4, and in the preset index table, "4" corresponds to "not", then the sixth character in the character sequence represents "not";
[0101] It should also be emphasized that each character in the label of the speech data needs to be time-aligned with the audio feature corresponding to the speech data;
[0102] Step (3.5): Use the audio features of each voice data as training samples and input them together with the corresponding labels into the Zipformer network structure for iterative training of the Zipformer network structure until the training stop condition is reached;
[0103] Among them, the training stop condition can be reaching a preset prediction accuracy rate or reaching the maximum number of iterations, etc. This application does not make specific limitations on this.
[0104] In actual operation, when the iterative training for the Zipformer network structure reaches the training stop condition, the second model can be obtained. Step (4): Determine the prompt tone category corresponding to the speech recognition text according to the regular expression to obtain the regular category text;
[0105] Among them, the number of prompt tone categories that the regular category text can indicate is the same as the number of prompt tone categories that the model category information can indicate.
[0106] Specifically, in the embodiments of this application, the regular category text corresponding to the speech recognition text is determined through the regular expression to improve the accuracy of proofreading for the speech classification text.
[0107] In actual operation, applying the regular expression can automatically extract key information related to the prompt tone category from the speech recognition text and obtain the regular category text according to the key information. The key information can be specific words, patterns, numbers, etc. Therefore, the information related to the prompt tone classification in the speech recognition text can be quickly identified; as shown in Table 1, Table 1 is an example table of regular matching provided by the embodiments of this application.
[0108] Table 1
[0109]
[0110]
[0111] S120: Determine the call prompt tones with the difference degree between the model category text and the regular category text in multiple call prompt tones being greater than or equal to the preset threshold as the retraining samples.
[0112] Specifically, the difference between the model category text and the regular category text may be caused by noise, weak voice signals, or other factors. If the difference degree between the model category text and the regular category text is large, it is generally considered that the recognition and classification result of the model category text is incorrect. Therefore, it is necessary to select the call prompt sounds with a large difference degree between the texts, that is, the call prompt sounds corresponding to the model category texts with a difference degree between the model category text and the regular category text greater than or equal to a preset threshold, as samples for retraining the current call prompt sound classification model; the preset threshold can be determined according to actual needs, and the present application does not make specific limitations on this.
[0113] In actual operation, the "difference degree between the model category text and the regular category text" can be determined according to the semantic similarity degree between the model category text and the regular category text, or can be determined by a pre-trained natural language processing (NLP) model. The present application does not make specific limitations on this.
[0114] In actual operation, the regular category text is derived from the speech recognition text, and the speech recognition text is obtained through the second model; among them, the purpose of performing speech recognition processing on the call prompt sound through the second model to obtain the speech recognition text is to efficiently obtain the actual speech text of the call prompt sound, avoiding wasting a large amount of manpower and material resources by performing manual recognition on each call prompt sound. However, in actual operation, there may be a situation where "due to the poor clarity of the call prompt sound, the speech recognition result obtained through the second model is not very accurate". Therefore, in order to obtain accurate retraining text, after selecting the call prompt sounds corresponding to the model category texts with a difference degree between the model category text and the regular category text greater than or equal to the preset threshold, the call prompt sound and the speech recognition text of the call prompt sound can be verified through manual review to determine whether the difference between the model category text and the regular category text corresponding to the call prompt sound is caused by the model category text or the speech recognition text.
[0115] If it is considered after manual review that the difference between the model category text and the regular category text corresponding to the call prompt sound is caused by the model category text, it means that the model category text and the regular category text corresponding to the call prompt sound are exactly the retraining samples required by the present application for retraining the call prompt sound classification model.
[0116] If it is considered after manual review that the difference between the model category text and the regular category text corresponding to the call prompt tone is caused by the speech recognition text, it means that the model category text and the regular category text corresponding to the call prompt tone are not the retraining samples required for retraining the call prompt tone classification model in this application, and the model category text and the regular category text corresponding to the call prompt tone need to be discarded.
[0117] S130: Retrain the call prompt tone classification model according to the retraining samples.
[0118] Specifically, retraining the call prompt tone classification model can effectively improve the ability of the prompt tone classification model to better adapt to call prompt tones with noise, etc., and greatly improve the classification ability.
[0119] In actual operation, for each call prompt tone determined to be a retraining sample, it is necessary to extract features for each call prompt tone to obtain the audio features of each call prompt tone, and it is also necessary to label each call prompt tone to obtain the label of each call prompt tone. Then, the audio features and labels corresponding to each call prompt tone are input into the call prompt tone classification model. The call prompt tone classification model outputs the model classification text corresponding to each retraining sample, and then, according to the model classification text and the corresponding label, the parameters of the call prompt tone classification model are updated; among them, the labeling method of the label can be manual labeling or other implementable technical methods, and this application does not make specific limitations on this.
[0120] In actual application, there are various ways to update the parameters of the call prompt tone classification model, which can be determined according to the actual situation in the specific training process, and this application does not make specific limitations on this.
[0121] In one implementation, S130 includes: step (5), details are as follows:
[0122] Step (5): Through the retraining samples, perform multiple rounds of iterative training on the call prompt tone classification model until the training stop condition is reached.
[0123] Specifically, the training stop condition can be to reach the preset prediction accuracy rate or reach the maximum number of iterations, etc., and this application does not make specific limitations on this. In actual operation, the training process of any round before the training stop condition is reached includes:
[0124] Step (5.1): Input the retraining samples into the call prompt tone classification model to obtain the model category text corresponding to the retraining samples.
[0125] Specifically, in actual operation, after the retraining samples are input into the call prompt tone classification model, the call prompt tone classification model outputs the model type text corresponding to the retraining samples; according to the model category text of the retraining samples, the loss function value is calculated.
[0126] Step (5.2): Update the model parameters of the call prompt tone classification model according to the model category text of the retraining samples.
[0127] Specifically, if the model parameters of the call prompt tone classification model are updated according to the predicted category text and / or the loss function value, the model parameters of the call prompt tone classification model can be updated according to the model category text and / or the loss function value.
[0128] In detail, as the iterative training of the call prompt tone classification model progresses, the labels of the retraining samples are continuously fed back into the call prompt tone classification model, enabling the call prompt tone classification model to have a clear learning objective; among them, the call prompt tone classification model will adjust the training status and adjustment strategy of the current round according to the training results of the previous round, and update and adjust the model parameters accordingly; among them, the training results include the model category text, the loss function value, and / or other data.
[0129] In the above training process, the call prompt tone classification model will continuously try, verify, and correct until the optimal configuration of the model parameters is found, so that the call prompt tone classification model can become more accurate and intelligent when processing similar scenarios. In summary, in this application, by introducing the regular category text to proofread the model category text, the call prompt tones corresponding to the model category text and the regular category text with the difference degree greater than or equal to the preset threshold are determined as the retraining texts for retraining the call prompt tone classification model, so as to improve the anti-interference ability of the call prompt tone classification model to factors such as background noise, silent segments, and audio quality differences, greatly improve the classification ability of the call prompt tone classification model, and further improve the efficiency of intelligent outbound calls. Second, this application provides an intelligent outbound call method applied to an intelligent outbound call device. The call prompt tone classification model is deployed in the intelligent outbound call device, and the call prompt tone classification model is obtained by training through the training method of the call prompt tone classification model detailed in S110~S130; as Figure 3 shown, Figure 3 is a schematic flowchart of the intelligent outbound call method provided by the embodiment of this application. The method includes: S210~S230, details are as follows:
[0130] S210: During the process of making an active call through the intelligent outbound call device, collect the target call prompt tone for the call prompt tone of the target telephone device being called.
[0131] Specifically, with the support of the outbound call list, the intelligent outbound call device will continuously make active calls to the target phone devices of the users in the outbound call list based on the outbound call list. The target phone device is the phone device that the intelligent outbound call device is currently making an outbound call to.
[0132] During the active call process, if the target phone device emits a call prompt tone, the intelligent outbound call device will record the call prompt tone to obtain the target call prompt tone of the target phone device.
[0133] S220: Input the target call prompt tone into the call prompt tone classification model to obtain the target model category text corresponding to the target call prompt tone;
[0134] Among them, the model category text indicates the category of the call prompt tone.
[0135] Specifically, the intelligent outbound call device extracts features from the target call prompt tone to obtain the target audio features, and then inputs the target audio features into the retrained call prompt tone classification model to obtain the target model category text of the target call prompt tone.
[0136] S230: Determine whether it is necessary to make an active call to the target phone device again according to the target model category text.
[0137] Specifically, since the target model category text can clearly indicate the category of the call prompt tone, it is possible to determine whether to make a call again according to the target model category text. For example, if the target model category text indicates that the target phone number device is in the state of "temporarily unanswered", it can be determined that "it is necessary to make an active call to the target phone device again".
[0138] In actual operation, if it is determined that it is necessary to make an active call to the target phone device again, the intelligent outbound call device can also correspondingly determine the time or interval of the active call again. For example, the time of the active outbound call again can be determined as an expression that can clearly indicate the time of the active call again, such as "8:00 tomorrow morning" or "8:00 in the morning after 1 working day".
[0139] In actual operation, if there is a situation where an active call is made to a target phone device multiple times but the target phone device is not connected, the duration between two adjacent active calls can be continuously lengthened until the target phone device is connected or the number of active outbound calls (for a target phone device) reaches the upper limit.
[0140] Third, the present application provides a training device for a call prompt tone classification model, as Figure 3 shown Figure 4Schematic structural diagram of a training device for a call prompt tone classification model provided by an embodiment of the present application. The device includes: a text determination module 310, a sample screening module 320, and a training module 330;
[0141] The text determination module 310 is configured to determine the model category text and the regular category text of multiple call prompt tones;
[0142] Among them, the call prompt tone is collected during an active call; the model category text is determined by identifying and classifying through a trained call prompt tone classification model; the regular category text is determined according to a regular expression;
[0143] The sample screening module 320 is configured to determine the call prompt tones with a difference degree greater than or equal to a preset threshold between the model category text and the regular category text among the multiple call prompt tones as retraining samples;
[0144] The training module 330 is configured to perform retraining on the call prompt tone classification model according to the retraining samples. In one implementation, the text determination module 310 is further configured to process the original call prompt tones that are stereo among the multiple original call prompt tones to obtain mono call prompt tones.
[0145] In one implementation, the text determination module 310 is further configured to remove the silent segments in the call prompt tone through a first model;
[0146] Among them, the first model is trained for a convolutional-long short-term memory-fully connected deep neural network.
[0147] In one implementation, the text determination module 310 is further configured to perform speech recognition processing on the call prompt tone to obtain speech recognition text;
[0148] The text determination module 310 is further configured to determine the prompt tone category corresponding to the speech recognition text according to the regular expression to obtain the regular category text;
[0149] Among them, the number of prompt tone categories indicated by the regular category text is the same as the number of prompt tone categories indicated by the model category information.
[0150] In one implementation, the text determination module 310 is further configured to perform speech recognition processing on the call prompt tone through a second model to obtain speech recognition text; among them, the second model is trained for a Zipformer network structure.
[0151] In one implementation, the training module 330 is further configured to perform multiple rounds of iterative training on the call prompt tone classification model through the retraining samples until a training stop condition is reached;
[0152] Among them, the training process in any round before the training stop condition is reached includes:
[0153] The training module 330 is further configured to input the retraining samples into the call prompt tone classification model to obtain the model category text corresponding to the retraining samples;
[0154] The training module 330 is further configured to update the model parameters of the call prompt tone classification model according to the model category text of the retraining samples.
[0155] Fourth, the present application provides an intelligent outbound calling device, which is applied to an intelligent outbound calling device, and a call prompt tone classification model is deployed in the intelligent outbound calling device. The call prompt tone classification model is obtained by training through the training method of the call prompt tone classification model detailed in S110 - S130; Figure 5 FIG. is a schematic structural diagram of the intelligent outbound calling device provided by the embodiment of the present application. The device includes: a prompt tone acquisition module 410, a classification module 420, and a call determination module 430;
[0156] The prompt tone acquisition module 410 is configured to collect a target call prompt tone for the call prompt tone of the target telephone device being called during the process of making an active call through the intelligent outbound calling device;
[0157] The classification module 420 is configured to input the target call prompt tone into the call prompt tone classification model to obtain the target model category text corresponding to the target call prompt tone;
[0158] Among them, the model category text indicates the category of the call prompt tone;
[0159] The call determination module 430 is configured to determine whether it is necessary to make an active call to the target telephone device again according to the target model category text.
[0160] Fifth, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of S110 - S130 or S210 - S230 provided in the above embodiments are implemented.
[0161] Fourth, the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of S110 - S130 or S210 - S230 in the above embodiments are executed.
[0162] Fifth, the computer program product provided by this application includes a computer-readable storage medium storing program codes. The instructions included in the program codes can be used to execute the methods in the foregoing method embodiments. For specific implementation, reference can be made to the steps of S110 - S130 or S210 - S230 in the method embodiments, which will not be elaborated herein.
[0163] In the embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0164] In addition, the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0165] Furthermore, in each embodiment of this application, the various functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0166] It should be noted that if a function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0167] In this document, relational terms such as first and second are used solely to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0168] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for training a call prompt tone classification model, characterized in that: The method comprises: Determine model category texts and regular category texts of multiple call prompt tones; The call prompt tone is collected during the active call process; the model category text is determined by identification and classification through the trained call prompt tone classification model; the regular category text is determined according to a regular expression; Determine the call prompt tones, among the plurality of call prompt tones, for which the difference between the model category text and the regular category text is greater than or equal to a preset threshold, as retraining samples; The call prompt tone classification model is retrained according to the retraining samples.
2. The method according to claim 1, characterized in that Before determining the model category texts and regular category texts of the plurality of call prompt tones, the method further includes: The original call prompt tones in dual channels among the multiple original call prompt tones are processed to obtain the call prompt tones in mono channels.
3. The method according to claim 1, characterized in that Before determining the model category texts and regular category texts of the plurality of call prompt tones, the method further includes: Removing the silent segment in the call prompt tone by using the first model; Among them, the first model is obtained by training a convolutional-long short-term memory-fully connected deep neural network.
4. The method according to claim 1, characterized in that: The step of determining the model category texts and the regular category texts of the plurality of call prompt tones includes: Performing speech recognition processing on the call prompt tone to obtain a speech recognition text; According to the regular expression, determining the prompt tone category corresponding to the speech recognition text to obtain a regular category text; The number of prompt sound categories that can be indicated by the regular category text is the same as the number of prompt sound categories that can be indicated by the model category information.
5. The method according to claim 4, characterized in that The performing speech recognition processing on the call prompt tone to obtain a speech recognition text includes: The call prompt tone is processed by speech recognition through a second model to obtain speech recognition text; wherein the second model is obtained by training the Zipformer network structure.
6. The method according to claim 1, characterized in that The retraining of the call prompt tone classification model according to the retraining sample includes: Performing multiple rounds of iterative training on the call prompt tone classification model using the retraining samples until a training stop condition is reached; The training process of any round before the training stop condition is reached includes: Inputting the retraining sample into the call prompt tone classification model to obtain a model category text corresponding to the retraining sample; The model parameters of the call prompt tone classification model are updated according to the model category text of the retraining sample.
7. An intelligent outbound calling method, characterized in that: Applied to an intelligent outbound call device, wherein a call prompt tone classification model is deployed in the intelligent outbound call device, and the call prompt tone classification model is obtained by training the call prompt tone classification model training method according to any one of claims 1 to 6; the method comprises: In the process of actively calling through the intelligent outbound calling device, the call prompt tone of the target telephone device being called is collected to obtain the target call prompt tone; Inputting the target call prompt tone into the call prompt tone classification model to obtain a target model category text corresponding to the target call prompt tone; Wherein, the model category text indicates the category of the call prompt tone; According to the target model category text, it is determined whether it is necessary to actively call the target telephone device again.
8. A training device for a call prompt tone classification model, characterized in that: The device comprises: a text determination module, a sample screening module and a training module; The text determination module is used to determine the model category texts and regular category texts of multiple call prompt tones; The call prompt tone is collected during the active call process; the model category text is determined by identification and classification through the trained call prompt tone classification model; the regular category text is determined according to a regular expression; The sample screening module is used to determine the call prompt tones whose difference between the model category text and the regular category text among the plurality of call prompt tones is greater than or equal to a preset threshold as retraining samples; The training module is used to retrain the call prompt tone classification model according to the retraining samples.
9. An intelligent outbound calling device, characterized in that: Applied to an intelligent outbound call device, wherein a call prompt tone classification model is deployed in the intelligent outbound call device, and the call prompt tone classification model is obtained by training the call prompt tone classification model training method described in any one of claims 1 to 6; the device comprises: a prompt tone collection module, a classification module and a call determination module; The prompt tone collection module is used to collect the call prompt tone of the called target telephone device to obtain the target call prompt tone during the process of actively calling through the intelligent outbound calling device; The classification module is used to input the target call prompt tone into the call prompt tone classification model to obtain a target model category text corresponding to the target call prompt tone; Wherein, the model category text indicates the category of the call prompt tone; The call determination module is used to determine whether it is necessary to actively call the target telephone device again according to the target model category text.
10. An electronic device, characterized in that: The electronic device includes a processor and a memory, the memory is used to store an application, and the processor runs or executes a software program stored in the memory so that the electronic device implements the training method of the call prompt tone classification model described in any one of claims 1 to 6 or the intelligent outbound calling method described in claim 7.