Processing method and device for dialogue type audio, equipment and storage medium

The dialogue audio is processed through deep neural networks and large models to generate emotional tags, solving the problem of inefficient manual monitoring and improving the quality inspection efficiency of dialogue audio.

CN119943049APending Publication Date: 2025-05-06银联数据服务有限公司
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411977783.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Manual monitoring of customer service conversations is inefficient, unable to effectively cover a large number of conversations, and there are problems of inconsistency testing.

Method used

Dialogue audio is recognized through deep neural network, first text information is generated, and a large model and emotion model is combined with emotion labels to determine whether the conversation meets the quality inspection requirements.

Benefits of technology

It improves the quality inspection efficiency of conversational audio, enhances the accuracy of emotional labels, and solves the problem of inefficient manual monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943049A_ABST
    Figure CN119943049A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dialogue type audio processing method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: carrying out the recognition of the dialogue type audio through a deep neural network, and obtaining the first text information of the dialogue type audio; inputting the first text information and a large model cue word into a large model to obtain second text information with an emotion label; the large model cue word is used for setting a model role and a text recognition task; and according to the emotion label in the second text information, determining whether the dialogue type audio meets quality inspection requirements or not. In the embodiment of the invention, the dialogue audio is converted into the first text information through the deep neural network, and the processing efficiency is improved by processing the text information; by inputting not only the first text information but also the large model cue word into the large model, the emotional label of the output second text information better meets the requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a processing method, device, equipment and storage medium for conversational audio. Background Art

[0002] When users conduct online business, customer service needs to have good service quality, such as service attitude and business ability. In order to evaluate the service quality of customer service, manual monitoring is usually used to monitor the conversation between customer service and users and score the service quality of customer service.

[0003] However, manual monitoring often has the problem of low efficiency, so a large number of conversations cannot be covered, and manual monitoring also has the problem of inconsistent verification. Summary of the invention

[0004] The embodiments of the present application provide a method, apparatus, device and storage medium for processing conversational audio, which are used to improve the quality inspection efficiency of conversational speech.

[0005] In a first aspect, an embodiment of the present application provides a method for processing conversational audio, including:

[0006] Recognize the conversational audio by using a deep neural network to obtain first text information of the conversational audio;

[0007] Inputting the first text information and the large model prompt words into the large model to obtain the second text information with the emotion label; the large model prompt words are used to set the model role and the text recognition task;

[0008] Determine whether the conversational audio meets quality inspection requirements based on the emotion tag in the second text information.

[0009] In the embodiment of the present application, the conversational audio is converted into the first text information through a deep neural network, and the processing efficiency is improved by processing the text information; by inputting not only the first text information but also the big model prompt words into the big model, the emotional label of the output second text information is made to better meet the needs.

[0010] Optionally, the recognizing the conversational audio by using a deep neural network to obtain the first text information of the conversational audio includes:

[0011] Inputting the conversational audio and the customized vocabulary into a first feedforward neural network to obtain a word feature vector corresponding to each sentence;

[0012] The word feature vectors corresponding to the continuous multiple sentences are processed by the second feedforward neural network to obtain a word feature matrix with semantic information;

[0013] The word feature matrix is ​​input into a neural network model of a self-attention mechanism to obtain first text information of the conversational audio.

[0014] In the embodiment of the present application, by inputting the conversational audio and the customized word library into the first layer of the feedforward neural network, a word feature vector in a set scenario can be obtained, and by inputting the word feature vector into the second layer of the feedforward neural network, multiple consecutive sentences can be combined to obtain more accurate first text information according to the context. Through the two-layer feedforward neural network, the first text information is more in line with the application scenario and the text information obtained according to the context is also more accurate.

[0015] Optionally, the word feature matrix includes the dialogue role corresponding to each sentence; the large model prompt word also includes the dialogue role corresponding to each sentence.

[0016] Optionally, before obtaining the second text information with the emotion tag, the method further includes:

[0017] Inputting the conversational audio into an emotion model to obtain third text information with an auxiliary emotion label;

[0018] The step of inputting the first text information and the large model prompt word into the large model to obtain the second text information with the emotion label includes:

[0019] Adding the auxiliary emotion tag correspondingly to the first text information;

[0020] The first text information with the auxiliary emotion label and the large model prompt word are input into the large model to obtain the second text information with the emotion label.

[0021] In an embodiment of the present application, auxiliary emotion labels of conversational audio are obtained through the emotion model, which can provide a reference for the large model to determine the emotion labels of the conversational audio. The large model is trained through the auxiliary emotion labels, thereby improving the accuracy of the large model in determining the emotion labels.

[0022] Optionally, the large model is trained in the following manner, including:

[0023] Determine the instruction compliance result, emotion label result and event detection result corresponding to the sample text information through the sample text information with emotion label output by the large model; wherein the text recognition task defines the instruction execution sequence, including the output result of emotion label and event detection;

[0024] According to the sample labels, linear similarity calculations are performed on the instruction following results, the emotion label results, and the event detection results to obtain deviation results of the instruction following results, the emotion label results, and the event detection results;

[0025] Determining a loss value based on deviation results of the instruction following result, the emotion label result, and the event detection result;

[0026] The large model is adjusted according to the loss value until the training is completed.

[0027] In the embodiment of the present application, by judging the instruction following results, emotion labeling results, and event detection results in the text recognition task, it is possible to determine whether the training of the large model meets the requirements, calculate the respective deviation results of the above, determine the loss value, and adjust the large model through the loss value, so that the learning ability of the large model is continuously enhanced, and the training results of the large model are also made more accurate.

[0028] Optionally, the recognizing the conversational audio by using a deep neural network includes:

[0029] Removing noise from the conversational audio using a background noise detection model;

[0030] The conversational audio after noise removal is recognized through a deep neural network.

[0031] In an embodiment of the present application, before using a deep neural network to recognize conversational audio, the background noise of the conversational audio is removed, so that the deep neural network can recognize the conversational audio more accurately, avoiding large errors in recognition caused by the noise of the conversational audio.

[0032] Optionally, the method further comprises:

[0033] Determining, by a background noise detection model, whether the conversational audio contains a first fraudulent behavior; and / or

[0034] Determining, by a voiceprint recognition model, whether the conversational audio contains a second fraudulent behavior; and / or

[0035] Determining whether the conversational audio contains a third fraudulent behavior through a voice forgery model; and / or

[0036] By using an intermediary behavior detection model, it is determined whether the conversational audio contains a fourth fraudulent behavior.

[0037] In the embodiments of the present application, the background noise model, voiceprint recognition model, speech forgery model and intermediary behavior detection model can be used to broaden the processing method of conversational audio. Not only can the emotion of conversational audio be identified, but also whether there is fraud or other behavior in the conversational audio, thereby improving the processing field of conversational audio.

[0038] In a second aspect, an embodiment of the present application provides a processing device for conversational audio, including:

[0039] A recognition module, configured to recognize the conversational audio through a deep neural network to obtain first text information of the conversational audio;

[0040] An input module, used to input the first text information and the large model prompt words into the large model to obtain the second text information with the emotion label; the large model prompt words are used to set the model role and the text recognition task;

[0041] A determination module is used to determine whether the conversational audio meets the quality inspection requirements based on the emotion tag in the second text information.

[0042] Optionally, the identification module is specifically used for:

[0043] Inputting the conversational audio and the customized vocabulary into a first feedforward neural network to obtain a word feature vector corresponding to each sentence;

[0044] The word feature vectors corresponding to the continuous multiple sentences are processed by the second feedforward neural network to obtain a word feature matrix with semantic information;

[0045] The word feature matrix is ​​input into a neural network model of a self-attention mechanism to obtain first text information of the conversational audio.

[0046] Optionally, the word feature matrix includes the dialogue role corresponding to each sentence; the large model prompt word also includes the dialogue role corresponding to each sentence.

[0047] Optionally, the input module is specifically used for:

[0048] Inputting the conversational audio into an emotion model to obtain third text information with an auxiliary emotion label;

[0049] The step of inputting the first text information and the large model prompt word into the large model to obtain the second text information with the emotion label includes:

[0050] Adding the auxiliary emotion tag correspondingly to the first text information;

[0051] The first text information with the auxiliary emotion label and the large model prompt word are input into the large model to obtain the second text information with the emotion label.

[0052] Optionally, the input module is specifically used for:

[0053] Determine the instruction compliance result, emotion label result and event detection result corresponding to the sample text information through the sample text information with emotion label output by the large model; wherein the text recognition task defines the instruction execution sequence, including the output result of emotion label and event detection;

[0054] According to the sample labels, linear similarity calculations are performed on the instruction following results, the emotion label results, and the event detection results to obtain deviation results of the instruction following results, the emotion label results, and the event detection results;

[0055] Determining a loss value based on deviation results of the instruction following result, the emotion label result, and the event detection result;

[0056] The large model is adjusted according to the loss value until the training is completed.

[0057] Optionally, the identification module is specifically used for:

[0058] Removing noise from the conversational audio using a background noise detection model;

[0059] The conversational audio after noise removal is recognized through a deep neural network.

[0060] Optionally, the determining module is further used for:

[0061] Determining, by a background noise detection model, whether the conversational audio contains a first fraudulent behavior; and / or

[0062] Determining, by a voiceprint recognition model, whether the conversational audio contains a second fraudulent behavior; and / or

[0063] Determining whether the conversational audio contains a third fraudulent behavior through a voice forgery model; and / or

[0064] By using an intermediary behavior detection model, it is determined whether the conversational audio contains a fourth fraudulent behavior.

[0065] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the processing methods for conversational audio described in the first aspect.

[0066] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device, wherein when the program is run on the computer device, the computer device executes any of the processing methods for conversational audio described in the first aspect above.

[0067] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer device, the computer device executes any step of the processing method for conversational audio described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0069] Figure 1 A system architecture diagram provided for an embodiment of the present application;

[0070] Figure 2 A flowchart of a method for processing conversational audio provided in an embodiment of the present application;

[0071] Figure 3 A flowchart of a method for determining first text information provided in an embodiment of the present application;

[0072] Figure 4 A flowchart of a method for training a large model provided in an embodiment of the present application;

[0073] Figure 5 A schematic diagram of the structure of a processing device for conversational audio provided in an embodiment of the present application;

[0074] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0076] To facilitate understanding of this solution, the application scenario of this solution is introduced below.

[0077] When banks evaluate the service quality of customer service, they usually conduct a comprehensive evaluation based on the chat conversations between users and customer service from the aspects of customer service attitude and business capabilities. At present, the following methods are used to evaluate the service quality of customer service: First, the following modules are needed to process the conversational audio: voice acquisition module, which collects the conversational audio between users and customer service; voice preprocessing module, which performs preprocessing operations such as noise reduction and reverberation on the collected conversational audio; voice recognition module, which uses deep learning and other technologies to convert the conversational audio into text; keyword extraction module, which uses natural language processing algorithms and predefined keyword dictionaries to extract key information related to business handling, such as applicant's name, ID number, application amount, income status, etc., and also extracts keywords related to service quality, such as enthusiasm, patience, etc. Then, a quality inspection rule library is constructed, including: compliance of speech, which stipulates that customer service must use standard and accurate speech to answer applicants' questions and guide the application process; information integrity, which regulates the tone and attitude of customer service; risk prevention and control, which identifies possible fraud or abnormal situations. The quality inspection and analysis module then compares and analyzes the extracted keywords and text of the conversational audio with the rule library to determine whether they meet the quality inspection requirements. Finally, the quality inspection results are output, including the problems in the conversational audio, the severity of the problems, the corresponding conversational audio clips, etc.

[0078] The above method can only make a simple assessment of the service quality of the customer service, but cannot accurately determine the service attitude of the customer service. The following is a detailed introduction to the operation steps of this application:

[0079] First, after taking the conversational audio, this application performs noise reduction processing on the conversational audio to obtain a clearer conversational audio, and inputs the audio into a two-layer feedforward neural network. Through the two-layer feedforward neural network, a relatively accurate first text information of the conversational audio can be obtained. The first text information is then input into the large model to obtain the second text information, which contains an emotional label. Finally, the conversational audio is quality inspected through the second text information to judge the service quality of the customer service, such as the customer service attitude and business capabilities. Among them, the large model needs to be trained before use so that the large model meets the capabilities required by this application. The information used to standardize event scenarios, event detection results, and sample text is input into the large model, and the large model is adjusted through a preset loss function so that the output results of the large model are consistent with the emotional labels of the sample text information.

[0080] See also Figure 1 , is a system architecture diagram provided in an embodiment of the present application, the system architecture includes a terminal device 101 and a server 102.

[0081] The terminal device 101 is pre-installed with a service application for text matching, wherein the service application is a client application, a web application, a small program application, etc. The terminal device 101 can be a smart phone, a POS machine, a desktop notebook, a computer, etc., but is not limited thereto.

[0082] Server 102 is the backend server for business applications. Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0083] See also Figure 2 , is a flow chart of a method for processing conversational audio provided in an embodiment of the present application, comprising the following steps:

[0084] Step 201: Recognize the conversational audio through a deep neural network to obtain first text information of the conversational audio.

[0085] Specifically, conversational audio is a chat mode of one question and one answer between two people. It is commonly seen between users and customer service, such as when user A applies for a credit card online, the chat between user A and customer service B of the bank is conversational audio. Conversational audio includes at least two people, that is, dual-track audio. The deep neural network model can recognize the conversational audio, and the first text information is the text content of the conversational audio obtained by the deep neural network model.

[0086] Step 202: input the first text information and the large model prompt words into the large model to obtain the second text information with the emotion label; the large model prompt words are used to set the model role and the text recognition task.

[0087] Specifically, the large model is used to process the first text information and determine the emotion in the first text information. When the first text information is input into the large model, the model role and text recognition task are also input into the large model. The model role is to inform the large model of the identity information corresponding to each sentence in the first text information. For example, in the first text information, the first sentence corresponds to user A, and the second sentence corresponds to customer service B. The text recognition task is to inform the execution task of the large model, such as "please add corresponding emotion tags to all sentences in the text". After the first text information and the large model prompt word are input into the large model, the second text information obtained has an emotion tag. The second text information can be the first text information and the emotion tag of each sentence in the first text information, or the second text information can be obtained by splitting the first text information into multiple sentences, each of which contains a corresponding emotion tag. Moreover, in the second text information, each sentence also contains the keyword corresponding to the generated emotion tag. For example, in the sentence "I feel good now", the emotion tag is "happy", and the keyword "very good" is marked.

[0088] Step 203: Determine whether the conversational audio meets quality inspection requirements based on the emotion tag in the second text information.

[0089] Specifically, the emotion label in the second text information is judged. If the emotion label meets the quality inspection requirements, it is considered that the service quality of the customer service in the conversational audio meets the quality inspection requirements; if the emotion label does not meet the quality inspection requirements, the service quality of the customer service in the conversational audio does not meet the quality inspection requirements.

[0090] In the embodiment of the present application, the conversational audio is converted into the first text information through a deep neural network, and the processing efficiency is improved by processing the text information; by inputting not only the first text information but also the big model prompt words into the big model, the emotional label of the output second text information is made to better meet the needs.

[0091] In some embodiments, the conversational audio is recognized by a deep neural network to obtain first text information of the conversational audio, such as Figure 3 As shown, the following steps are included:

[0092] Step 301: Input the conversational audio and the customized vocabulary into a first feedforward neural network to obtain a word feature vector corresponding to each sentence.

[0093] Specifically, the customized vocabulary is set according to the application scenario and belongs to the exclusive words of the application scenario. The conversational audio and the customized vocabulary are both input into the first feedforward neural network to obtain the word feature vector corresponding to each sentence. The word feature vector is to divide the sentence into multiple words, each word has a corresponding feature vector, and the word feature vector is the unique identifier of the word.

[0094] For example, user A wants to apply for a credit card. During the chat with customer service B, the chat conversation between user A and customer service B contains multiple exclusive words related to the credit card application scenario, such as "credit card", "maximum limit", "proof of income" and other exclusive words. The word library is customized according to the application scenario of the credit card application. When there are such exclusive words, the exclusive words will not be processed separately. If there is no exclusive word library, the above exclusive words will be split into "credit", "card", "income", "proof", etc. The continuous conversation between user A and customer service B and the customized word library in the credit card application scenario are input into the first feedforward neural network to obtain the word feature vector of each sentence, such as the word feature vector of "credit card" is (0,1,1), and the word feature vector of "proof of income" is "(1,0,1)".

[0095] Step 302: Process the word feature vectors corresponding to the plurality of consecutive sentences through a second feedforward neural network to obtain a word feature matrix with semantic information.

[0096] Specifically, the word feature vector is input into the second feedforward neural network to obtain a word feature matrix, which is composed of word feature vectors. The word feature matrix contains multiple word feature vectors, and the word feature vectors in the word feature matrix are contextual relationships, so the word feature matrix has semantic information.

[0097] In some embodiments, the word feature matrix includes the dialogue role corresponding to each sentence; the large model prompt word also includes the dialogue role corresponding to each sentence.

[0098] Specifically, the word feature matrix integrates dialogue role recognition, sentence segmentation, and punctuation mark special identification functions. The large model prompt words also include the dialogue role corresponding to each sentence.

[0099] Step 303: input the word feature matrix into the neural network model of the self-attention mechanism to obtain the first text information of the conversational audio.

[0100] Specifically, the neural network model of the self-attention mechanism can be a Transformer architecture, which has achieved great success in the field of natural language processing, especially in tasks such as machine translation, text generation, and semantic understanding. It can not only process long sequences, but also capture global dependencies, improving the accuracy and semantic coherence of the model. The Transformer architecture is mainly composed of encoding components and decoding components, each of which consists of a multi-layer encoder (Encoder) and a multi-layer decoder (Decoder). The main function of the encoder is to convert the input sequence into a series of vector representations, while the decoder generates an output sequence based on these vectors.

[0101] Inputting the word feature matrix obtained in the above method into the Transformer architecture enhances the model's ability to understand contextual semantics compared to directly inputting conversational audio into the Transformer architecture.

[0102] In the embodiment of the present application, by inputting the conversational audio and the customized word library into the first layer of the feedforward neural network, a word feature vector in a set scenario can be obtained, and by inputting the word feature vector into the second layer of the feedforward neural network, multiple consecutive sentences can be combined to obtain more accurate first text information according to the context. Through the two-layer feedforward neural network, the first text information is more in line with the application scenario and the text information obtained according to the context is also more accurate.

[0103] In some embodiments, before obtaining the second text information with the emotion tag, the method further includes:

[0104] The conversational audio is input into the emotion model to obtain third text information with auxiliary emotion labels.

[0105] Specifically, the conversational audio is input into the emotion model, and the emotion label of the conversational audio can be obtained through the emotion model. The emotion model is used to identify conversations with large emotional fluctuations in the conversational audio.

[0106] The emotion model may be an Emotion LSTM model, which extracts acoustic feature vectors in conversational audio through Continuous integrate-and-fire (CIF), builds an emotion2vec general speech emotion representation model, records the emotional states of both parties in the conversational audio for real-time tracking, and obtains emotion labels in the conversational audio, and uses the labels as third text information to assist the emotion labels.

[0107] In some embodiments, the first text information and the large model prompt word are input into the large model to obtain the second text information with the emotion tag, including:

[0108] Auxiliary emotion tags are added to the first text information accordingly. The first text information with the auxiliary emotion tags and the large model prompt words are input into the large model to obtain the second text information with the emotion tags.

[0109] Specifically, after the above method, the conversational audio now has the auxiliary emotion tag and the first text information. The auxiliary emotion tag is added to the first text information and input into the large model, and the prompt word is input into the large model at the same time to obtain the second text information with the emotion tag. At this time, the second text information is combined with the auxiliary emotion tag provided by the emotion model.

[0110] In an embodiment of the present application, auxiliary emotion labels of conversational audio are obtained through the emotion model, which can provide a reference for the large model to determine the emotion labels of the conversational audio. The large model is trained through the auxiliary emotion labels, thereby improving the accuracy of the large model in determining the emotion labels.

[0111] In some embodiments, the large model is trained as follows: Figure 4 As shown, the following steps are included:

[0112] Step 401, determine the instruction compliance results, emotion label results and event detection results corresponding to the sample text information through the sample text information with emotion labels output by the large model; wherein the text recognition task defines the output results of instruction execution sequence, emotion labels and event detection.

[0113] Specifically, the sample text information is input into the big model to obtain the sample text information with emotion labels. The corresponding compliance results, emotion label results and event detection results are determined according to the sample text information. The prompt word setting of the big model includes a text recognition task, which defines the instruction execution order, emotion label and event detection. Through the sample text information output by the big model, the instruction compliance results corresponding to the text instruction execution order, the emotion label results corresponding to the emotion label, and the event detection results corresponding to the event detection are determined.

[0114] Even if the results of the large model follow the instruction format, the judgment results on the content are still illusory, that is, the output results of the model deviate from the expected results, and certain manual corrections are required to ensure accuracy. In order to better align the output of the large model with human preferences, the KTO (Kahneman-Tversky Optimisation) reinforcement learning training algorithm is adopted.

[0115] Step 402: According to the sample labels, linear similarity calculations are performed on the instruction following results, the emotion label results, and the event detection results, respectively, to obtain deviation results of the instruction following results, the emotion label results, and the event detection results.

[0116] Specifically, according to the sample labels, the linear similarities corresponding to the instruction following results, emotion label results and event detection results are calculated respectively, and the softmax function is used to complete the normalized output, as shown in formula (1):

[0117] e i =softmax(Linear D →V i )...Formula (1)

[0118] Among them, e i For deviation results, i = Ltd or ser or aed, eLtd indicates that the instruction follows the deviation result, c indicates that the emotion label deviates from the result, and e aed Indicates the event detection deviation result. V i Indicates the weight of the newly added vocabulary.

[0119] Step 403: determine a loss value based on the deviation results of the instruction following result, the emotion label result, and the event detection result.

[0120] Specifically, a loss function is constructed based on the deviation results of the instruction following results, the emotion labeling results, and the event detection results, so as to determine the loss value. The loss function is shown in formula (2):

[0121] f(x) value =Union(e Ltd ,e ser ,e aed |X speech )...Formula (2)

[0122] Among them, X speech is the first text message, f(x) value is the loss value.

[0123] Step 404: Adjust the large model according to the loss value until the training is completed.

[0124] Specifically, according to the loss value, the large model is trained until the requirements are met.

[0125] For example, in order to better adapt the large model to the tasks in the conversational audio emotion analysis scenario, write text recognition task instructions and introduce additional <spkrole> , <speechtoken> , <aed>Special annotations are used to enhance the training effect of domain-specific tasks. The specific training instructions are composed of the following:

[0126] Inputs= <system> SystemPrompt< / system> <user> <spkrole>< / spkrole> {Instruct Prompt}+{{Speech Prompt}}+< / user> .

[0127] Among them, SystemPrompt is the preset instruction code, which uses lisp syntax and is used for model role setting. InstructPrompt is a specific task prompt project, which is used for task specification and output rule content control, and is used to guide the large model to perform special analysis on sample text information. Speech Prompt is the input after the conversational audio is converted into text information. The text information also includes emotion tags and auxiliary emotion tags. The output format not only includes emotion tags, but also includes which word or word in the text generates the emotion tag. The judgment result of each paragraph is output in json format. The cross entropy loss is used as the optimizer during training to obtain a model trained by SFT, so that the fine-tuned model can better follow the specific instruction output. Spkrole is a dialogue role, which is divided into user and customer service.

[0128] The instruction encoding is as follows:

[0129] In the embodiment of the present application, by judging the instruction following results, emotion labeling results, and event detection results in the text recognition task, it is possible to determine whether the training of the large model meets the requirements, calculate the respective deviation results of the above, determine the loss value, and adjust the large model through the loss value, so that the learning ability of the large model is continuously enhanced, and the training results of the large model are also made more accurate.

[0130] In some embodiments, recognizing conversational audio using a deep neural network includes:

[0131] The background noise detection model is used to remove noise from the conversational audio. The conversational audio after noise removal is recognized through a deep neural network.

[0132] Specifically, when actually processing conversational audio, conversational audio with background noise is often encountered. In order to avoid large errors in subsequent processing, the noise in the conversational audio is removed through a background noise detection model, and then recognized through a deep neural network. This can improve the recognition efficiency and accuracy of the conversational audio.

[0133] In an embodiment of the present application, before using a deep neural network to recognize conversational audio, the background noise of the conversational audio is removed, so that the deep neural network can recognize the conversational audio more accurately, avoiding large errors in recognition caused by the noise of the conversational audio.

[0134] In some embodiments, the method for processing conversational audio further includes:

[0135] Determine whether the conversational audio contains a first fraudulent behavior through a background noise detection model; and / or determine whether the conversational audio contains a second fraudulent behavior through a voiceprint recognition model; and / or determine whether the conversational audio contains a third fraudulent behavior through a voice forgery model; and / or determine whether the conversational audio contains a fourth fraudulent behavior through an intermediary behavior detection model.

[0136] Specifically, a background noise detection model is introduced, such as a CNN-RNN network. The background noise is classified through the CNN-RNN network to identify and analyze noise in abnormal environments, such as noisy public places and multi-person conversations. The residual convolutional ResNet network is used as the front end, and the time-delay neural network structure is selected as the backbone. By embedding a context-aware mask module in each layer, a densely connected time-delay neural network architecture is constructed. Combined with sound field analysis technology, multi-granularity pooling tools are used to extract contextual features of complex scales. The generated mask is used to remove irrelevant noise in the features and retain key speaker information, such as the voices of user A and customer service B.

[0137] The introduction of voiceprint recognition models, such as the One-Class SVM dynamic voiceprint recognition model, can not only match the user's baseline voiceprint, but also adapt to the natural changes in their voice through a continuous learning mechanism, such as changes in the voice caused by a cold, to maintain a high-accuracy match. At the same time, Siamese Networks is used to perform in-depth comparison of voiceprint features to improve the ability to resist noise interference.

[0138] Introducing voice forgery models, such as the adversarial training GANs (Generative Adversarial Networks) model, based on real voice scene elements, using multiple new TTS (text-to-speech) models to generate artificially synthesized voices with the same semantic content, while using the GPT architecture to complete voice cloning to improve robustness, and training the system to recognize and resist deep fake voice attacks. By simulating adversarial sample training, the system can recognize the unique pattern features of synthetic voices, such as waveform distortion and unnatural frequency conversion, significantly enhancing anti-counterfeiting capabilities.

[0139] An intermediary behavior detection model is introduced. For example, in the analysis of audio streams, the RNN-LSTM model based on sequence pattern recognition can identify abnormal conversation patterns by analyzing the sequence features, vocabulary patterns, and speech speed changes of the conversation. It can determine whether there are any customer questions during the electronic verification process, especially any hesitation, ambiguity, hesitancy, or obvious waiting delays in identity information verification, so as to intelligently detect the risk of information concealment or fraud by intermediary agents.

[0140] Based on the intermediary agent dialogue and fraud case identification paradigm constructed through the big model, multiple task target detection tasks are constructed to detect whether there are information inconsistencies, such as the user registration information and the PBOC report are inconsistent with the electronic verification results; detect whether there is information concealment, such as the filled-in address is identified as a non-valid address; detect whether there are suspicious intermediary language hits, such as the answer to the reason for applying for a card or the information application submitted is suspected of being a template; detect whether there are abnormal parts in the dialogue content, such as the distance between the delivery address and the customer's residential address is too large, and identify and detect suspicious intermediary agent fraud.

[0141] The following is a partial code example of a method for processing conversational audio provided in an embodiment of the present application:

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149] In the embodiments of the present application, the background noise model, voiceprint recognition model, speech forgery model and intermediary behavior detection model can be used to broaden the processing method of conversational audio. Not only can the emotion of conversational audio be identified, but also whether there is fraud or other behavior in the conversational audio, thereby improving the processing field of conversational audio.

[0150] Based on the same technical concept, the embodiment of the present application provides a processing device for conversational audio, such as Figure 5 As shown, the device 500 includes:

[0151] A recognition module 501 is used to recognize the conversational audio through a deep neural network to obtain first text information of the conversational audio;

[0152] An input module 502 is used to input the first text information and the large model prompt words into the large model to obtain the second text information with the emotion label; the large model prompt words are used to set the model role and the text recognition task;

[0153] The determination module 503 is used to determine whether the conversational audio meets the quality inspection requirements according to the emotion tag in the second text information.

[0154] Optionally, the identification module 501 is specifically used for:

[0155] Inputting the conversational audio and the customized vocabulary into a first feedforward neural network to obtain a word feature vector corresponding to each sentence;

[0156] The word feature vectors corresponding to the continuous multiple sentences are processed by the second feedforward neural network to obtain a word feature matrix with semantic information;

[0157] The word feature matrix is ​​input into a neural network model of a self-attention mechanism to obtain first text information of the conversational audio.

[0158] Optionally, the word feature matrix includes the dialogue role corresponding to each sentence; the large model prompt word also includes the dialogue role corresponding to each sentence.

[0159] Optionally, the input module 502 is specifically used for:

[0160] Inputting the conversational audio into an emotion model to obtain third text information with an auxiliary emotion label;

[0161] The step of inputting the first text information and the large model prompt word into the large model to obtain the second text information with the emotion label includes:

[0162] Adding the auxiliary emotion tag correspondingly to the first text information;

[0163] The first text information with the auxiliary emotion label and the large model prompt word are input into the large model to obtain the second text information with the emotion label.

[0164] Optionally, the input module 502 is specifically used for:

[0165] Determine the instruction compliance result, emotion label result and event detection result corresponding to the sample text information through the sample text information with emotion label output by the large model; wherein the text recognition task defines the instruction execution sequence, including the output result of emotion label and event detection;

[0166] According to the sample labels, linear similarity calculations are performed on the instruction following results, the emotion label results, and the event detection results to obtain deviation results of the instruction following results, the emotion label results, and the event detection results;

[0167] Determining a loss value based on deviation results of the instruction following result, the emotion label result, and the event detection result;

[0168] The large model is adjusted according to the loss value until the training is completed.

[0169] Optionally, the identification module 501 is specifically used for:

[0170] Removing noise from the conversational audio using a background noise detection model;

[0171] The conversational audio after noise removal is recognized through a deep neural network.

[0172] Optionally, the determining module 503 is further configured to:

[0173] Determining, by a background noise detection model, whether the conversational audio contains a first fraudulent behavior; and / or

[0174] Determining, by a voiceprint recognition model, whether the conversational audio contains a second fraudulent behavior; and / or

[0175] Determining, by a voice forgery model, whether the conversational audio contains a third fraudulent behavior; and / or

[0176] By using an intermediary behavior detection model, it is determined whether the conversational audio contains a fourth fraudulent behavior.

[0177] Based on the same technical concept, the embodiment of the present application provides a computer device, which may be a terminal or a server, such as Figure 6 As shown, it includes at least one processor 601 and a memory 602 connected to the at least one processor. The specific connection medium between the processor 601 and the memory 602 is not limited in the embodiment of the present application. Figure 6 For example, the processor 601 and the memory 602 are connected via a bus. The bus can be divided into an address bus, a data bus, a control bus, and the like.

[0178] In the embodiment of the present application, the memory 602 stores instructions that can be executed by at least one processor 601. The at least one processor 601 can execute the steps included in the above-mentioned processing method for conversational audio by executing the instructions stored in the memory 602.

[0179] The processor 601 is the control center of the computer device, and can use various interfaces and lines to connect various parts of the computer device, by running or executing instructions stored in the memory 602 and calling data stored in the memory 602. Optionally, the processor 601 may include one or more processing units, and the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 601. In some embodiments, the processor 601 and the memory 602 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on independent chips.

[0180] Processor 601 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware processor for execution, or can be executed by a combination of hardware and software modules in the processor.

[0181] The memory 602 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 602 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 602 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 602 in the embodiment of the present application can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.

[0182] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the program runs on the computer device, the computer device executes the steps of the above-mentioned processing method for conversational audio.

[0183] Based on the same inventive concept, an embodiment of the present application provides a computer program product, characterized in that the computer program product includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer device, the computer device executes the steps of the above-mentioned processing method for conversational audio.

[0184] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0185] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0186] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0187] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0188] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.< / aed> < / speechtoken> < / spkrole>

Claims

1. A method for processing conversational audio, characterized in that: include: Inputting the conversational audio and the customized vocabulary into a first feedforward neural network to obtain a word feature vector corresponding to each sentence; The word feature vectors corresponding to the continuous multiple sentences are processed by the second feedforward neural network to obtain a word feature matrix with semantic information; Inputting the word feature matrix into a neural network model of a self-attention mechanism to obtain first text information of the conversational audio; Inputting the first text information and the large model prompt words into the large model to obtain the second text information with the emotion label; the large model prompt words are used to set the model role and the text recognition task; Determine whether the conversational audio meets quality inspection requirements based on the emotion tag in the second text information.

2. The method according to claim 1, characterized in that include: The word feature matrix includes the dialogue role corresponding to each sentence; The large model prompt words also include the dialogue role corresponding to each sentence.

3. The method according to claim 1, characterized in that Before obtaining the second text information with the emotion tag, the method further includes: Inputting the conversational audio into an emotion model to obtain third text information with an auxiliary emotion label; The step of inputting the first text information and the large model prompt word into the large model to obtain the second text information with the emotion label includes: Adding the auxiliary emotion tag correspondingly to the first text information; The first text information with the auxiliary emotion label and the large model prompt word are input into the large model to obtain the second text information with the emotion label.

4. The method according to any one of claims 1 to 3, characterized in that: The large model is trained in the following manner, including: Determine the instruction compliance result, emotion label result and event detection result corresponding to the sample text information through the sample text information with emotion label output by the large model; wherein the output results of instruction execution sequence, emotion label and event detection are defined in the text recognition task; According to the sample labels, linear similarity calculations are performed on the instruction following results, the emotion label results, and the event detection results to obtain deviation results of the instruction following results, the emotion label results, and the event detection results; Determining a loss value based on deviation results of the instruction following result, the emotion label result, and the event detection result; The large model is adjusted according to the loss value until the training is completed.

5. The method according to any one of claims 1 to 3, characterized in that: The recognizing the conversational audio by using a deep neural network includes: Removing noise from the conversational audio using a background noise detection model; The conversational audio after noise removal is recognized through a deep neural network.

6. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Determining, by a background noise detection model, whether the conversational audio contains a first fraudulent behavior; and / or Determining, by a voiceprint recognition model, whether the conversational audio contains a second fraudulent behavior; and / or Determining whether the conversational audio contains a third fraudulent behavior through a voice forgery model; and / or By using an intermediary behavior detection model, it is determined whether the conversational audio contains a fourth fraudulent behavior.

7. A processing device for conversational audio, characterized in that: include: A recognition module, configured to input the conversational audio and the customized vocabulary into a first feedforward neural network to obtain a word feature vector corresponding to each sentence; The word feature vectors corresponding to the continuous multiple sentences are processed by the second feedforward neural network to obtain a word feature matrix with semantic information; Inputting the word feature matrix into a neural network model of a self-attention mechanism to obtain first text information of the conversational audio; An input module, used for inputting the first text information and the large model prompt word into the large model to obtain the second text information with the emotion label; The large model prompt words are used to set the model role and text recognition tasks; A determination module is used to determine whether the conversational audio meets the quality inspection requirements based on the emotion tag in the second text information.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of any one of the methods of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that: It stores a computer program executable by a computer device. When the program is run on the computer device, the computer device executes the steps of any method described in claims 1 to 6.

10. A computer program product, characterized in that The computer program product comprises a computer program stored on a computer-readable storage medium, wherein the computer program comprises program instructions, and when the program instructions are executed by a computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Voice quality inspection method and device, equipment, storage medium and program product

    CN120581040A

  • Call processing method, device and equipment

    CN120600025A

  • Emotion recognition method and device based on large model, medium, equipment and product

    CN121483311A