Method and System for Determining Reply Speech of Outbound Call Robot

By generating tone, volume and speech activity rate characteristics combined with text semantic vectors, the RoBERTa language model is used to calculate the correlation score and priority of intent and context, the lack of outgoing call robots in speech comprehension and intention recognition is solved, and more accurate user emotions and intention capture is achieved, and user experience and service efficiency is improved.

CN119621917BActive Publication Date: 2025-07-08NANTONG ZHIDATONG INFORMATION SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510149281.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-08
Estimated Expiration
2045-02-11

Smart Images

  • Figure CN119621917B_ABST
    Figure CN119621917B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for determining the reply words of an outbound robot, which relates to the technical field of dialogue robots. The present invention extracts the pitch, volume, and speech activity rate features of the user's voice signal to generate a voice feature vector; and converts the voice signal into text to generate a text semantic vector, and combines to generate a user state vector; generates a context semantic vector and an intention semantic vector based on the RoBERTa language model, calculates the intention correlation score, calculates the intention completion rate through a prediction model and combines the urgency to generate a priority score, selects the user intention according to the priority scoring function, and trains a reply model based on historical conversation data to generate the most appropriate recovery words in real time, improving the intelligence and adaptability of the robot's reply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of dialogue robots, and specifically to a method and system for determining the reply speech of an outbound robot. Background Art

[0002] In the context of the rapid development of current voice interaction technology, outbound robots, as an intelligent service tool, are widely used in fields such as customer service, marketing, and user support. However, existing outbound robots often face problems such as inaccurate speech understanding, inefficient intent recognition, and lack of personalized reply speech. This is mainly because traditional outbound robots usually rely only on text semantic analysis and ignore the emotional features contained in voice signals, making it difficult for outbound robots to accurately capture the user's emotional state and communication needs. In addition, traditional methods lack an effective comprehensive evaluation mechanism for determining the priority of user intents, resulting in inaccurate intent priority ranking in complex scenarios and affecting the customer experience and service efficiency.

[0003] In the prior art, the publication number CN116975223A discloses a method, device, and storage medium for determining the reply speech of an outbound robot, which includes obtaining a dialogue request text corresponding to a customer who communicates with the outbound robot, and obtaining a plurality of clustering clusters, where each of the clustering clusters includes a question text and its corresponding reply speech; inputting the dialogue request text into a preset semantic prediction model for semantic prediction to obtain a first semantic information vector corresponding to the dialogue request text; based on the first semantic information vector, determining the target clustering cluster to which the dialogue request text belongs in each of the clustering clusters, and determining the reply speech corresponding to the cluster center question text of the target clustering cluster as the reply speech corresponding to the dialogue request text.

[0004] The main problems of the above method are: mainly relying on text semantic analysis to map the user's dialogue request text into a semantic vector, and lacking in extracting the features of the user's voice signal, resulting in limitations in identifying the user's emotional state and hidden intent; classifying question texts and matching them to reply speech through a clustering method, where the division of clustering clusters is based on the static classification results of training data, making it difficult to adapt to dynamic user intents and complex and diverse dialogue scenarios, and the clustering method only focuses on the current dialogue request text and directly matches the question text and reply speech, while ignoring the in-depth analysis of context semantics, thus affecting the reply accuracy.

[0005] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The object of the present invention is to provide a method and system for determining the reply speech of an outbound robot, so as to solve the problems raised in the above-mentioned background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method for determining the reply speech of an outbound robot, the specific steps include:

[0009] Step 1: Collect the voice signal input by the user, judge the pitch change of the voice signal through the YIN algorithm to generate a pitch feature, generate a volume feature by calculating the energy of the voice signal, generate a voice activity rate feature by identifying the time period of voice generation in the voice signal, and generate a voice feature vector based on the pitch feature, volume feature and voice activity rate feature;

[0010] Step 2: Convert the voice signal into text in an automatic speech recognition model, input the text into the BERT model to generate a text semantic vector, and generate a user state vector by combining the text semantic vector and the voice feature vector;

[0011] Step 3: Input the text into a pre-trained language model based on RoBERTa to generate a context semantic vector. For each possible intention, generate an intention semantic vector through the language model based on RoBERTa, and generate a relevance score between the intention and the context based on the context semantic vector and the intention semantic vector;

[0012] Step 4: Input the user state vector and the intention into a pre-trained prediction model to generate an intention completion rate, assign a priority weight based on the urgency of the intention, and generate an intention completion priority score;

[0013] Step 5: Generate a priority scoring function for the user intention based on the relevance score between the intention and the context and the intention completion priority score, calculate the priority of each intention, and output the intention with the highest priority;

[0014] Step 6: Collect historical reply speech data, use the known intention and user state vector as input, and the reply speech as a label to train the reply model, and input the intention with the highest priority score and the user state vector into the reply model to generate the reply speech.

[0015] Furthermore, the principle for generating the voice feature vector is:

[0016] The principle for generating the pitch feature is:

[0017] For each frame of voice signal, calculate the difference function at different delays, and the formula is:

[0018]

[0019] Among them, \(l\) represents the delay, \(d(l)\) represents the difference function corresponding to the delay of \(l\), \(x(t)\) represents the speech signal at time \(t\), \(t\) represents the time index of the speech signal, and \(N\) represents the total number of time indexes;

[0020] The formula for generating the normalized cumulative difference function is:

[0021]

[0022] Among them, \(d\) nor (l j ) represents the normalized difference function corresponding to the delay of \(l\), \(l\) j represents the \(j\)-th delay, j represents the cumulative sum of the difference functions from the 1st to the \(j\)-th delay, \(J\) represents the total number of different delays, and \(j\) represents the delay index; Search for the local minimum in the normalized cumulative difference function, and calculate the pitch frequency based on the delay corresponding to the local minimum. The formula is:

[0023]

[0024]

[0025] Among them, represents the pitch frequency corresponding to the pitch period \(l\) j , \(T\) s represents the speech sampling period;

[0026] Analyze the pitch frequency trajectory according to the time series, and generate the average pitch change rate. The formula is:

[0027]

[0028] Among them, represents the average pitch change rate between the pitch periods \(l\) j and \(l\) j-1 ;

[0029] Combine the average pitch change rates of all periods to generate the pitch feature vector \(F\);

[0030] The principle for generating the volume feature is:

[0031]

[0032] Among them, \(E\) represents the volume feature of the speech signal, and \(N\) represents the total number of sample points in the signal;

[0033] The principle for generating the speech activity rate feature is:

[0034] ​

[0035] Among them, R represents the speech activity rate feature, and T sound represents the duration of voice production in continuous speech, and T toatal represents the total duration of continuous speech;

[0036] The generated speech feature vector is:

[0037] V = [F, E, R]

[0038] Among them, V represents the speech feature vector.

[0039] Furthermore, the principle for generating the user state vector is:

[0040] Input the text into the BERT model, and the generated text semantic vector is:

[0041] Y = BERT(W)

[0042] Among them, Y represents the text semantic vector, and W represents the text of the current speech signal;

[0043] Concatenate the speech feature vector V and the text semantic vector Y to generate the user state vector, denoted as:

[0044] U = [V, Y]

[0045] Among them, U represents the user state vector.

[0046] Furthermore, the principle for generating the relevance score of the intention and the context is:

[0047] The principle for generating the context semantic vector is:

[0048] Input the context of the current conversation into the RoBERTa-based language model, and the generated context semantic vector is:

[0049] H = RoBERTa(W context )

[0050] Among them, H represents the context semantic vector, and W context represents the text context of the current conversation;

[0051] The principle for generating the intention semantic vector is:

[0052] For each intention, generate an intention text description, where the intention represents the purpose of the current conversation with the user. Input the intention text description into the RoBERTa-based language model to generate the intention semantic vector, denoted as:

[0053] K i = RoBERTa(X i )

[0054] Among them, K i represents the intention semantic vector of the i-th intention, and X i represents the i-th intention, where i represents the index of the intention;

[0055] The formula for generating the relevance score between the intention and the context is:

[0056]

[0057] Among them, S i represents the relevance score between the i-th intention and the context, |H| represents the modulus of the context semantic vector, and |X i | represents the modulus of the i-th intention semantic vector.

[0058] Furthermore, the principle for generating the intention completion priority score is:

[0059] The principle for generating the intention completion rate is:

[0060] C i = PRED(X i , U)

[0061] Among them, C i represents the completion rate of the i-th intention, PRED represents prediction by the prediction model, X i represents the i-th intention, and U represents the user state vector;

[0062] The formula for generating the intention completion priority score is:

[0063] B i = C i ·w task,i

[0064] Among them, B i represents the intention completion priority score of the i-th intention, and w task,i represents the priority weight of the i-th intention.

[0065] Furthermore, the formula for generating the priority scoring function is:

[0066] P i = α·S i + β·B i

[0067] Among them, P i represents the priority scoring function, S i represents the relevance score between the intention and the context, and B iIt represents the intention to complete the priority score. α and β respectively represent the weight coefficients of the relevance score between the intention and the context and the intention to complete the priority score.

[0068] The formula for generating the intention with the highest priority is:

[0069] i′ = argmax i P i

[0070] Among them, i′ represents the intention with the highest priority, and argmax i P i represents selecting the intention with the highest priority score from the intentions.

[0071] The present invention also provides a system for determining the reply words of an outbound robot. The system is used to implement the method for determining the reply words of the outbound robot, and specifically includes:

[0072] A voice processing module, which is used to collect the voice signal input by the user, judge the pitch change of the voice signal through the YIN algorithm to generate a pitch feature, generate a volume feature by calculating the energy of the voice signal, generate a speech activity rate feature by identifying the continuous sounding time period in the voice signal, and generate a voice feature vector based on the pitch feature, volume feature and speech activity rate feature;

[0073] A signal conversion module, which is used to convert the voice signal into text in an automatic speech recognition model, input the text into a BERT model to generate a text semantic vector, and generate a user state vector by combining the text semantic vector and the voice feature vector;

[0074] A semantic calculation module, which is used to input the text into a pre-trained language model based on RoBERTa to generate a context semantic vector. For each possible intention, generate an intention semantic vector through the language model based on RoBERTa, and generate the relevance score between the intention and the context based on the context semantic vector and the intention semantic vector;

[0075] An intention processing module, which is used to input the user state vector and the intention into a pre-trained prediction model to generate an intention completion rate, assign a priority weight based on the urgency of the intention, and generate an intention completion priority score;

[0076] An intention scoring module, which is used to generate a priority scoring function for the user intention based on the relevance score between the intention and the context and the intention completion priority score, calculate the priority of each intention, and output the intention with the highest priority;

[0077] The comprehensive output module is used to collect historical reply script data, take the known intent and user state vector as inputs, and the reply script as a label to train the reply model. The intent with the highest priority score and the user state vector are input into the reply model to generate a reply script.

[0078] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0079] The present invention extracts speech feature vectors generated through pitch features, volume features, and speech activity rate features, fully considering the emotions and expression intents hidden in the speech, capturing the user's emotional changes and urgency, making up for the deficiency of the prior art in emotion perception ability, and thus making a more accurate judgment on the priority of user needs. It also converts the speech signal into text information, generates a text semantic vector, combines the text semantic vector and the speech feature vector to generate a user state vector. Through multi-modal fusion, it can better capture the deep intents and emotions in the user's expression and adapt to more complex dialogue scenarios.

[0080] The present invention also generates a context semantic vector and an intent semantic vector, which can more deeply understand the context meaning and potential intent of the user's expression, achieve more accurate semantic matching. Calculating the correlation between the context semantic vector and the intent semantic vector can quantify the matching degree of each intent with the current context. This semantic-based correlation scoring method goes beyond ordinary rule-based or keyword-based matching methods, can better adapt to the diversity and ambiguity in natural language expression. Through the deep matching of the context semantic vector and the intent semantic vector, the generated correlation score can dynamically and accurately match the true intent of the user's expression. The combination of the user state vector and the intent makes the prediction of the intent completion rate not only depend on static semantic information, but also dynamically perceive information such as the user's emotion, attitude, and urgency, effectively improving the accuracy of the intent completion rate prediction. It enables the system to more scientifically evaluate the possibility of achieving a certain intent in complex situations, and through the priority scoring function, it can dynamically sort according to the current interaction status, always ensuring that the system gives priority to responding to the user's most important needs, avoiding resource waste or response delay, and greatly improving the user experience. Description of the Drawings

[0081] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present invention;

[0082] Figure 2 It is a schematic diagram of the system module according to the embodiment of the present invention. Detailed Embodiments

[0083] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further describes the present invention in detail with reference to specific embodiments.

[0084] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "comprising" or "including" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0085] Embodiment:

[0086] Please refer to Figure 1 , the present invention provides a technical solution:

[0087] A method for determining the reply speech of an outbound robot, the specific steps include:

[0088] Step 1: Collect the voice signal input by the user, judge the pitch change of the voice signal through the YIN algorithm to generate a pitch feature, generate a volume feature by calculating the energy of the voice signal, generate a speech activity rate feature by identifying the continuous sounding time period in the voice signal, and generate a voice feature vector based on the pitch feature, volume feature and speech activity rate feature;

[0089] In this embodiment, the principle for generating the voice feature vector is:

[0090] The principle for generating the pitch feature is:

[0091] For each frame of voice signal, calculate the difference function at different delays, and the formula is:

[0092]

[0093] where, l represents the delay, d(l) represents the difference function corresponding to the delay of l, x(t) represents the voice signal at time t, t represents the time index of the voice signal, and N represents the total number of time indices;

[0094] The difference function reflects the similarity degree of the voice signal at different delays. When the delay is exactly one period of the signal, x(t) and x(t + l) will be very similar, and at this time the difference function reaches a local minimum.

[0095] The formula for generating the normalized cumulative difference function is:

[0096]

[0097] Among them, d nor (l j ) represents the normalized difference function corresponding to a delay of l j , where l j represents the jth delay, represents the cumulative sum of the difference functions from the 1st to the jth delay, J represents the total number of different delays, and j represents the index of the delay;

[0098] In the difference function, the value of d(l j ) increases as the delay l j increases. The normalized cumulative difference function emphasizes the characteristics of the low-delay region. When the delay l j increases, d nor (l j ) can avoid the disorderly expansion of the difference function to more accurately detect the pitch; the value of d nor (l j ) ranges from 0 to 1. The smaller d nor (l j ), the higher the similarity of the signal at the delay l j and the stronger the periodicity.

[0099] Find the local minimum in the normalized cumulative difference function, and calculate the pitch frequency based on the delay corresponding to the local minimum. The formula used is:

[0100]

[0101] Among them, represents the pitch frequency corresponding to the pitch period l j , and T s represents the speech sampling period;

[0102] Analyze the pitch frequency trajectory according to the time series to generate the average pitch change rate. The formula used is:

[0103]

[0104] Among them, represents the average pitch change rate between the pitch periods l j and l j-1 ;

[0105] Combine the average pitch change rates of all periods to generate the pitch feature vector F;

[0106] The principle for generating the volume feature is:

[0107]

[0108] Among them, E represents the volume feature of the voice signal, and N represents the total number of sample points in the signal;

[0109] The principle for generating the voice activity rate feature is as follows:

[0110]

[0111] Among them, R represents the voice activity rate feature, and T sound represents the voice duration in continuous speech, and T total represents the total duration of continuous speech;

[0112] The generated voice feature vector is:

[0113] V = [F, E, R]

[0114] Among them, V represents the voice feature vector.

[0115] Step 2: Convert the voice signal into text in the automatic speech recognition model, input the text into the BERT model to generate the text semantic vector, and combine the text semantic vector and the voice feature vector to generate the user state vector;

[0116] The voice signal is converted into text through an automatic speech recognition (ASR) model based on Seq2Seq, which consists of an encoder and a decoder. The encoder encodes the input voice signal into a high-dimensional context vector, and the decoder gradually decodes to generate a text sequence according to the context vector generated by the encoder. At each step, the decoder receives the output and hidden state of the previous time step, generates the hidden state of the current time step, selects the word with the highest probability as the output of the current time step, and uses the output of the current time step as the input of the next time step until the end symbol is generated.

[0117] In this embodiment, the principle for generating the user state vector is as follows:

[0118] Input the text into the BERT model, and the generated text semantic vector is:

[0119] Y = BERT(W)

[0120] Among them, Y represents the text semantic vector, and W represents the text of the current voice signal;

[0121] Concatenate the voice feature vector V and the text semantic vector Y to generate the user state vector, which is expressed as:

[0122] U = [V, Y]

[0123] Among them, U represents the user state vector.

[0124] The user status vector reflects the comprehensive status of speech features and text semantics, including the user's emotional state, the user's semantic intention, and the comprehensive relationship between semantics and emotion.

[0125] Step 3: Input the text into a pre-trained RoBERTa-based language model to generate a context semantic vector. For each possible intention, generate an intention semantic vector through the RoBERTa-based language model, and generate a relevance score between the intention and the context based on the context semantic vector and the intention semantic vector;

[0126] In this embodiment, the principle on which the RoBERTa-based language model is trained is as follows:

[0127] Based on the 24-layer Transformer Encoder, prepare multi-domain corpus data including news, social media, Q&A conversations, product reviews, etc. Use the input speech text in the model, as well as the context information of the text, and the possible intentions, with the context semantic vector as the label to train the language model;

[0128] In this embodiment, the principle on which the relevance score between the intention and the context is generated is as follows:

[0129] The principle on which the context semantic vector is generated is as follows:

[0130] Input the context of the current conversation into the RoBERTa-based language model, and the generated context semantic vector is:

[0131] H = RoBERTa(W context )

[0132] where H represents the context semantic vector, and W context represents the text context of the current conversation;

[0133] The principle on which the intention semantic vector is generated is as follows:

[0134] For each intention, generate an intention text description, where the intention represents the purpose of the current conversation with the user. Input the intention text description into the RoBERTa-based language model to generate an intention semantic vector, denoted as:

[0135] K i = RoBERTa(X i )

[0136] where K i represents the intention semantic vector of the i-th intention, X i represents the i-th intention, and i represents the index of the intention;

[0137] The formula for generating the relevance score between the intention and the context is:

[0138]

[0139] Among them, S i represents the relevance score of the i-th intention to the context, |H| represents the norm of the context semantic vector, and |X i | represents the norm of the i-th intention semantic vector.

[0140] Taking the cosine similarity as the relevance score of the intention to the context, in the pre-trained language model, the semantic embeddings learned from large-scale data make the vector directions of texts that are semantically similar in the semantic space more consistent, and the calculation result is directly standardized between 0 and 1.

[0141] Step 4: Input the user state vector and the intention into the pre-trained prediction model to generate the intention completion rate, assign priority weights based on the urgency of the intention, and generate the intention completion priority score;

[0142] In this embodiment, the principle on which the prediction model is trained is:

[0143] Collect the historical data of the outbound robot's interaction with the user, including the user's voice data, the recognized text content, the user's intention, the success rate of the outbound robot in completing the intention, and the user feedback. Mark the key information in the collected data, including the user intention, the urgency of the intention, and the completion situation;

[0144] Extract the current state of the user based on the user voice feature vector and the text semantic vector, extract the completion rate of similar intentions in the historical data, and combine the historical data to mark the urgency and completion rate of each intention. Using a multi-layer perceptron as the architecture, input the user state vector and the intention, and use the intention completion rate as the label to train the prediction model.

[0145] In this embodiment, the principle on which the intention completion priority score is generated is:

[0146] The principle on which the intention completion rate is generated is:

[0147] C i = PRED(X i , U)

[0148] Among them, C i represents the completion rate of the i-th intention, PRED represents prediction by the prediction model, X i represents the i-th intention, and U represents the user state vector;

[0149] The formula on which the intention completion priority score is generated is:

[0150] B i = Ci ·w task,i

[0151] Among them, B i represents the completion priority score of the i-th intention, and w task,i represents the priority weight of the i-th intention.

[0152] According to the business scenario, set the urgency of different intentions. The longer the waiting time of the intention and the shorter the execution time, the higher the assigned priority; because tasks that can be completed quickly can respond to user needs more efficiently and avoid the decline of user experience caused by long waiting times.

[0153] Step 5: Generate a priority scoring function for the user intention based on the relevance score between the intention and the context and the intention completion priority score, calculate the priority of each intention, and output the intention with the highest priority;

[0154] In this embodiment, the formula for generating the priority scoring function is:

[0155] P i =α·S i +β·B i

[0156] Among them, P i represents the priority scoring function, S i represents the relevance score between the intention and the context, B i represents the intention completion priority score, and α and β respectively represent the weight coefficients of the relevance score between the intention and the context and the intention completion priority score, and α + β = 1;

[0157] The priority scoring function reflects the comprehensive score of the intention. The higher the score of the intention, the more preferentially it is executed. Determine the weight coefficients based on the business focus. For scenarios with higher requirements for the semantic relevance of the context, to avoid answering off-topic, set α > β, α = 0.6, β = 0.4; for scenarios that emphasize the priority of the intention, such as medical calls, fault repairs, etc., set α < β, α = 0.4, β = 0.6

[0158] The formula for generating the intention with the highest priority is:

[0159] i′ = argmax i P i

[0160] Among them, i′ represents the intention with the highest priority, and argmax i P i represents selecting the intention with the highest priority score from the intentions.

[0161] argmax iIndicates finding the one that makes P i Take the maximum value of i, that is, among all intents, find the intent i with the highest priority score, and select it as i′, so as to determine the intent to be executed first.

[0162] Step 6: Collect historical reply speech data, use the known intent and user state vector as input, and the reply speech as a label to train the reply model. Input the intent with the highest priority score and the user state vector into the reply model to generate the reply speech.

[0163] In this embodiment, the reply model is trained based on a deep learning network, and the structure of the deep learning network is as follows:

[0164] Input layer: It contains 2 neurons and is used to input the intent with the highest priority and the user state vector;

[0165] The first hidden layer: It contains 128 neurons and is activated using the ReLU function;

[0166] The second hidden layer: It contains 64 neurons and is activated using the ReLU function;

[0167] Output layer: It contains one neuron and is used to output the reply speech.

[0168] Please refer to Figure 2 , the present invention also provides a system for determining the reply speech of an outbound robot. The system is used to implement the method for determining the reply speech of the outbound robot as described above, and specifically includes:

[0169] A voice processing module, which is used to collect the voice signal input by the user, judge the pitch change of the voice signal through the YIN algorithm to generate a pitch feature, generate a volume feature by calculating the energy of the voice signal, generate a voice activity rate feature by identifying the continuous sounding time period in the voice signal, and generate a voice feature vector based on the pitch feature, volume feature and voice activity rate feature;

[0170] A signal conversion module, which is used to convert the voice signal into text in an automatic speech recognition model, input the text into the BERT model to generate a text semantic vector, and combine the text semantic vector and the voice feature vector to generate a user state vector;

[0171] A semantic calculation module, which is used to input the text into a pre-trained language model based on RoBERTa to generate a context semantic vector. For each possible intent, generate an intent semantic vector through the language model based on RoBERTa, and generate a relevance score between the intent and the context based on the context semantic vector and the intent semantic vector;

[0172] An intent processing module, which is used to input the user status vector and the intent into a pre-trained prediction model, generate an intent completion rate, assign a priority weight based on the urgency of the intent, and generate an intent completion priority score;

[0173] An intent scoring module, which is used to generate a priority scoring function for the user intent based on the relevance score between the intent and the context and the intent completion priority score, calculate the priority of each intent, and output the intent with the highest priority;

[0174] A comprehensive output module, which is used to collect historical reply script data, use the known intent and the user status vector as inputs, and the reply script as a label to train a reply model, and input the intent with the highest priority score and the user status vector into the reply model to generate a reply script.

[0175] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0176] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.

[0177] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered within the protection scope of this application.

Claims

1. A method for determining the reply script of an outbound robot, characterized in that The specific steps include: Step 1: Collect the voice signal input by the user, judge the pitch change of the voice signal through the YIN algorithm to generate pitch features, generate volume features by calculating the energy of the voice signal, generate voice activity rate features by identifying the continuous sounding time period in the voice signal, and generate a voice feature vector based on the pitch features, volume features, and voice activity rate features; Step 2: Convert the voice signal into text in the automatic speech recognition model, input the text into the BERT model to generate a text semantic vector, and generate a user state vector by combining the text semantic vector and the voice feature vector; Step 3: Input the text into the pre-trained RoBERTa-based language model to generate a context semantic vector. For each possible intention, generate an intention semantic vector through the RoBERTa-based language model, and generate a relevance score between the intention and the context based on the context semantic vector and the intention semantic vector; Step 4: Input the user state vector and the intention into the pre-trained prediction model to generate an intention completion rate, assign a priority weight based on the urgency of the intention, and generate an intention completion priority score; Step 5: Generate a priority scoring function for the user intention based on the relevance score between the intention and the context and the intention completion priority score, calculate the priority of each intention, and output the intention with the highest priority; Step 6: Collect historical reply script data, use the known intention and user state vector as input, and the reply script as the label to train the reply model. Input the intention with the highest priority score and the user state vector into the reply model to generate a reply script; The principle for generating the voice feature vector is: The principle for generating the pitch feature is: For each frame of the voice signal, calculate the difference function at different delays, and the formula is: where l represents the delay, d(l) represents the difference function corresponding to the delay of l, x(t) represents the voice signal at time t, t represents the time index of the voice signal, and N represents the total number of time indexes; The formula for generating the normalized cumulative difference function is: where d nor (l j ) represents the normalized difference function corresponding to a delay of l j where l j represents the j-th delay, represents the cumulative sum of the difference functions from the 1st to the j-th delay, J represents the total number of different delays, and j represents the index of the delay; Find the local minimum in the normalized cumulative difference function, and calculate the pitch frequency based on the delay corresponding to the local minimum as the pitch period. The formula is: Among them, represents the pitch period l j the corresponding pitch frequency, T s represents the voice sampling period; Analyze the pitch frequency trajectory according to the time series to generate the average pitch change rate. The formula is: Among them, represents the average pitch change rate between the pitch periods l j and l j-1 ; Combine the average pitch change rates of all periods to generate the pitch feature vector F; The principle for generating the volume feature is: where E represents the voice signal volume feature, and N represents the total number of sample points in the signal; The principle for generating the voice activity rate feature is: Among them, R represents the speech activity rate feature, and T sound represents the vocalization duration in continuous speech, and T total represents the total duration of continuous speech; The generated voice feature vector is: V = [F, E, R] where V represents the voice feature vector; Set the urgency levels of different intents according to business scenarios.

2. The method for determining the reply script of an outbound call robot according to claim 1, wherein: The principle for generating the user state vector in Step 2 is: Input the text into the BERT model, and the generated text semantic vector is: Y = BERT(W) where Y represents the text semantic vector, and W represents the text of the current voice signal; Concatenate the voice feature vector V and the text semantic vector Y to generate the user state vector, denoted as: U = [V, Y] where U represents the user state vector.

3. The method for determining the reply words of an outbound call robot according to claim 1, wherein: The principle for generating the relevance score between the intent and the context in step 3 is as follows: The principle for generating contextual semantic vectors is: The context of the current conversation is input into the language model based on RoBERTa, and the generated context semantic vector is: H = RoBERTa(W context ) Among them, H represents the context semantic vector, and W context represents the text context of the current conversation; The principle for generating the intent semantic vector is: For each intent, generate an intent text description, which indicates the purpose of this conversation with the user. Input the intent text description into the RoBERTa-based language model to generate an intent semantic vector, which is expressed as: K i = RoBERTa(X i ) Among them, K i represents the intention semantic vector of the i-th intention, and X i represents the i-th intention, where i represents the index of the intention; The formula used to generate the intent-context relevance score is: Among them, S i represents the relevance score of the i-th intention to the context, |H| represents the modulus of the context semantic vector, and |X i | represents the modulus of the i-th intention semantic vector.

4. The method for determining the reply words of an outbound robot according to claim 1, wherein: The principle for generating the intention completion priority score in step 4 is: The intent completion rate is generated based on the following principles: C i = PRED(X i , U) Among them, C i represents the completion rate of the i-th intention, PRED represents the prediction by the prediction model, and X i represents the i-th intention, and U represents the user state vector; The formula used to generate the intent completion priority score is: B i = C i · w task,i Among them, B i represents the completion priority score of the i-th intention, and w task,i represents the priority weight of the i-th intention.

5. The method for determining the reply script of an outbound robot according to claim 1, wherein: The formula for generating the priority scoring function in step 5 is: P i = α·S i + β·B i Among them, P i represents the priority scoring function, S i represents the relevance score of the intention and the context, B i represents the intention completion priority score, and α and β respectively represent the weight coefficients of the relevance score of the intention and the context and the intention completion priority score; The formula for generating the highest priority intent is: i′ = argmax i P i Among them, i′ represents the intention with the highest priority, argmax i P i represents selecting the intention with the highest priority score from the intentions.

6. A system for determining the replying words of an outbound call robot, characterized in that: The system is used to implement the method for determining the reply speech of the outbound call robot according to any one of claims 1 to 5, specifically comprising: The speech processing module is used to collect the speech signal input by the user, determine the pitch change of the speech signal through the YIN algorithm, generate the tone feature, generate the volume feature by calculating the energy of the speech signal, generate the speech activity rate feature by identifying the time period of continuous sound in the speech signal, and generate the speech feature vector based on the tone feature, volume feature and speech activity rate feature; The signal conversion module is used to convert the speech signal into text in the automatic speech recognition model, input the text into the BERT model to generate a text semantic vector, and combine the text semantic vector and the speech feature vector to generate a user state vector; The semantic calculation module is used to input the text into the pre-trained RoBERTa-based language model to generate a context semantic vector. For each possible intent, an intent semantic vector is generated through the RoBERTa-based language model, and a relevance score between the intent and the context is generated based on the context semantic vector and the intent semantic vector. The intent processing module is used to input the user state vector and intent into the pre-trained prediction model, generate the intent completion rate, assign priority weights based on the urgency of the intent, and generate the intent completion priority score; The intent scoring module is used to generate a priority scoring function for user intent based on the relevance score between intent and context and the intent completion priority score, calculate the priority of each intent, and output the intent with the highest priority; The comprehensive output module is used to collect historical reply speech data, take known intent and user state vectors as input, reply speech as labels, train the reply model, and input the intent and user state vector with the highest priority score into the reply model to generate reply speech.

Citation Information

Patent Citations

  • Characterization learning-based Chinese automatic speech recognition text restoration method and system

    CN115438154A

  • Method and device for determining reply verbal skill of outbound robot and storage medium

    CN116975223A