Dialogue tendency recognition method and system based on text structuring and multi-modal fusion
By adopting text structured and multimodal fusion methods in dialogue tendency recognition, dialogue keywords are extracted and speech speed and intonation classification are combined, the problems of identification result bias and useless information misleading in the single modal recognition method are solved, and higher recognition accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510206281.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-06
AI Technical Summary
The existing dialogue tendency recognition methods for out-of-call dialogue mainly rely on a single mode and cannot effectively utilize the complementary advantages and intrinsic connections between different information, resulting in large deviations in the identification results. Due to the limited data volume, they are easily misled by useless information, which affects the accuracy of identification.
A dialogue tendency recognition method based on text structured and multimodal fusion is proposed. By segmenting the initial corpus data, dialogue keywords are extracted, and structured text with labels is generated. Combining speech speed and intonation classification, a deep learning model based on attention mechanism is used for prediction.
It improves the accuracy and reliability of dialogue tendency recognition, reduces the interference of non-critical information on identification results, improves model processing efficiency, and enhances the comprehensiveness and accuracy of identification results through multimodal fusion.
Smart Images

Figure CN120104795A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and system for identifying conversation tendency based on text structuring and multimodal fusion. Background Art
[0002] Intent recognition is a very important technology in conversation tasks. It aims to analyze the intentions and purposes expressed by users, so as to more accurately understand user needs or provide more appropriate related services. For example, in the operator's after-sales follow-up scenario, analyzing the user's emotions and obtaining the conversation tendency between customer service and customers is conducive to quickly determining the follow-up results.
[0003] At present, the conversation tendency recognition of most outbound calls is still limited to the scope of a single modality, and is mainly implemented by speech-to-text or Transformer. The fundamental flaw of this type of single-modality method is that it cannot take advantage of the complementary advantages and inherent connections between different information in outbound calls, and the resulting recognition results may be relatively biased. In addition, due to the limited amount of single-modality data, it is necessary to use as much information as possible, even useless information, but useless information may mislead the model to produce erroneous results. The relevant technology cannot eliminate the redundant components in single-modality data (such as text data), which may directly affect the accuracy of the recognition results, resulting in a decrease in the overall recognition accuracy. Summary of the invention
[0004] In order to solve at least one of the above problems, the main purpose of the embodiments of the present application is to propose a conversation tendency recognition method and system based on text structuring and multimodal fusion, aiming to improve the accuracy and reliability of conversation tendency recognition, and also to improve recognition efficiency.
[0005] To achieve the above objectives, an embodiment of the present application proposes a method for identifying conversation tendency based on text structuring and multimodal fusion, the method comprising:
[0006] The initial corpus data is segmented to obtain multiple segments of voice data based on dialogue turns and multi-party voice data based on roles; wherein the multi-party voice data includes the voice data of the questioner and the voice data of the answerer;
[0007] Extracting keywords from the voice data of the questioner to obtain conversation keywords;
[0008] Generate a structured text with tags according to the multiple voice data segments and the conversation keywords;
[0009] Inputting the labeled structured text into a deep learning model based on an attention mechanism to predict the probability of positive and negative tendencies of the text of the initial corpus data;
[0010] Classifying the speaking speed of the respondent's voice data to obtain the speaking speed emotion tendency probability;
[0011] Performing intonation classification on the respondent's voice data to obtain intonation emotion tendency probability;
[0012] According to the text positive and negative tendency probability, the speech rate emotional tendency probability and the intonation emotional tendency probability, a dialogue positive and negative tendency recognition result is obtained.
[0013] In some embodiments, extracting keywords from the questioner's voice data to obtain conversation keywords includes the following steps:
[0014] Converting the questioner's voice data into first text data;
[0015] Keyword extraction is performed on the first text data to obtain conversation keywords.
[0016] In some embodiments, generating a tagged structured text according to the multiple voice data segments and the conversation keywords comprises the following steps:
[0017] Converting the plurality of voice data into second text data;
[0018] The second text data and the conversation keywords are concatenated to obtain text data to be processed; wherein the text data to be processed includes a plurality of word units;
[0019] Generating a target input matrix according to the text data to be processed;
[0020] Perform category label prediction on the target input matrix to obtain a label prediction result for each word in the text data to be processed.
[0021] In some embodiments, generating a target input matrix according to the text data to be processed comprises the following steps:
[0022] Performing word segmentation processing on the text data to be processed to obtain word units;
[0023] Performing word embedding processing on the word element to obtain a word embedding vector;
[0024] According to the sentence to which the word-gram belongs, assigning a paragraph tag to the word-gram to obtain a segment embedding vector;
[0025] According to the position of the word element in the text data to be processed, assigning a position tag to the word element to obtain a position embedding vector;
[0026] Adding the word embedding vector, the segment embedding vector and the position embedding vector to generate a target input vector for the word unit;
[0027] All the word units in the text to be processed are combined in order to obtain a target input matrix.
[0028] In some embodiments, the step of classifying the speaking rate of the respondent's voice data to obtain the speaking rate emotion tendency probability includes the following steps:
[0029] According to the voice data of the respondent, obtaining the speaking time of the respondent;
[0030] converting the respondent voice data into third text data;
[0031] Determining the total number of words spoken by the respondent based on the third text data;
[0032] Calculate the ratio of the total number of words spoken by the respondent to the speaking time of the respondent to obtain the respondent's speaking speed feature;
[0033] Encoding the third text data to obtain contextual semantic representation;
[0034] splicing the context semantic representation with the respondent speech rate feature to obtain a first splicing result;
[0035] Performing weighted processing on the first splicing result to obtain a first weighted result;
[0036] The first weighted result is input into an activation function for emotion classification to obtain the probability of speech rate emotion tendency.
[0037] In some embodiments, the step of performing intonation classification on the respondent's voice data to obtain intonation emotion tendency probability comprises the following steps:
[0038] Extracting the respondent's voice features from the respondent's voice data by short-time Fourier transform;
[0039] Performing weighted fusion on the respondent's speech features to obtain a second weighted result;
[0040] The second weighted result is input into a second activation function to perform sentiment classification, and obtain the probability of sentiment tendency of the intonation.
[0041] In some embodiments, obtaining the positive and negative tendency recognition result of the dialogue according to the positive and negative tendency probability of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability comprises the following steps:
[0042] The positive and negative tendency probabilities of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability are weighted and integrated to obtain a dialogue tendency probability;
[0043] Extracting positive probability and negative probability from the conversation tendency probability;
[0044] When the positive probability is greater than the negative probability, the conversation positive or negative tendency is identified as positive;
[0045] When the positive probability is less than or equal to the negative probability, the positive or negative tendency of the conversation is identified as negative.
[0046] To achieve the above objectives, another aspect of the embodiment of the present application proposes a conversation tendency recognition system based on text structuring and multimodal fusion, the system comprising:
[0047] The first module is used to segment the initial corpus data to obtain multiple voice data based on dialogue turns and multi-party voice data based on roles; wherein the multi-party voice data includes the voice data of the questioner and the voice data of the answerer;
[0048] The second module is used to extract keywords from the questioner's voice data to obtain dialogue keywords;
[0049] A third module is used to generate a labeled structured text according to the multiple voice data and the conversation keywords;
[0050] The fourth module is used to input the labeled structured text into a deep learning model based on an attention mechanism to predict the probability of positive and negative tendency of the text of the initial corpus data;
[0051] The fifth module is used to classify the speaking speed of the respondent's voice data to obtain the speaking speed emotion tendency probability;
[0052] The sixth module is used to classify the intonation of the respondent's voice data to obtain the probability of intonation emotion tendency;
[0053] The seventh module is used to obtain the positive and negative tendency recognition result of the dialogue according to the positive and negative tendency probability of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability.
[0054] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.
[0055] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0056] The embodiments of the present application include at least the following beneficial effects: The present application provides a method and system for identifying the tendency of a conversation based on text structuring and multimodal fusion. The scheme divides the initial corpus data into multiple voice data segments based on conversation turns and multi-party voice data based on roles, combines the conversation keywords extracted from the questioner's voice into multiple voice data segments, and generates a labeled structured text. The key information in the initial corpus data can be marked, and the interference of non-key information on the recognition results can be reduced, so that the model is more focused on effective information, and the model processing efficiency can be improved. Such a structured text is input into a deep learning model based on an attention mechanism to predict the text positive and negative tendency probabilities of the initial corpus data, and the obtained positive and negative tendency probabilities are highly accurate. The speech speed and intonation of the respondent's voice data are then classified, and the emotional tendency probability therein is extracted as an auxiliary judgment. A new dimension consideration is added on the basis of text analysis, so that the accuracy and reliability of the obtained conversation positive and negative tendency recognition results are further improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The accompanying drawings are used to provide further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0058] Figure 1 It is a schematic diagram of the implementation environment provided by the embodiment of the present application;
[0059] Figure 2 is a role relationship diagram provided in an embodiment of the present application;
[0060] Figure 3 It is a step diagram of a method for identifying a conversation tendency based on text structuring and multimodal fusion provided in an embodiment of the present application;
[0061] Figure 4 It is a schematic diagram of the overall process of a method for identifying a conversation tendency based on text structuring and multimodal fusion provided in an embodiment of the present application;
[0062] Figure 5 It is a schematic diagram of a process for generating a structured text with a tag provided in an embodiment of the present application;
[0063] Figure 6 This is a visualization example diagram of text structure processing provided by an embodiment of the present application;
[0064] Figure 7 is another visualization example diagram of text structuring processing provided by an embodiment of the present application;
[0065] Figure 8 This is a Transformer-based text positive and negative tendency probability prediction framework diagram provided in an embodiment of the present application;
[0066] Fig. 9 It is a module schematic diagram of a conversation tendency recognition system based on text structuring and multimodal fusion provided in an embodiment of the present application;
[0067] Fig.10 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0068] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of devices and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0069] Although the functional modules are divided in the system schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first / S100", "second / S200", etc. in the specification, claims and the above drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0070] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0071] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0072] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0073] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0074] It is understandable that the method for identifying conversation tendency based on text structuring and multimodal fusion provided in the embodiment of the present invention can be applied to any computer device with data processing and computing capabilities, and this computer device can be various terminals or servers. When the computer device in the embodiment is a server, the server is an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Optionally, the terminal is a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited to this.
[0075] like Figure 1 FIG. 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. Figure 1 The implementation environment includes at least one terminal 102 and a server 101. The terminal 102 and the server 101 can be connected to a network wirelessly or wired to complete data transmission and exchange.
[0076] Server 101 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.
[0077] In addition, the server 101 can also be a node server in the blockchain network. Blockchain is a new application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm.
[0078] The terminal 102 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal 102 and the server 101 may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present invention.
[0079] For example, based on Figure 1 In the implementation environment shown, an embodiment of the present invention provides a lateral speed analysis and mapping method. The following is explained using the lateral speed analysis and mapping method applied to the server 101 as an example. It can be understood that the lateral speed analysis and mapping method can also be applied to the terminal 102.
[0080] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0081] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0082] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0083] (1) Questioner and answerer: can also be understood as interviewer and interviewee. Figure 2 As shown in the figure, the questioner is the role that is mainly responsible for guiding the conversation, obtaining information and understanding the situation by asking questions, and usually asks questions to tap the feedback of the interviewee; the answerer is the role that is mainly responsible for providing information and answering the questions raised by the questioner, such as the customer service and the customer in the sales return conversation. It should be noted that the conversation between the questioner and the answerer is not completely one-way, and the answerer may also ask questions to the questioner for some reasons.
[0084] (2) Conversation tendency: The conversation tendency of this application refers to whether the information tendency conveyed by the content provided by the respondent in the conversation is consistent with the questioner's question expectation. For example, in a sales follow-up conversation, the customer service follow-up question expectation is that the customer is satisfied with a certain service. If the customer's reply conveys satisfaction, the conversation tendency of the conversation is positive. If the customer's reply conveys dissatisfaction, the conversation tendency is negative. In a more complex example, for example, the customer service follow-up question expectation is that the customer is aware of all pre-sales contracts. If the customer's reply conveys a clear understanding of the contracts, the conversation tendency is positive. If the customer's reply conveys an unclear or incorrect understanding of the contracts, the conversation tendency is negative.
[0085] In the related technologies, the conversation tendency recognition of most outbound calls is still limited to the scope of a single modality, and is mainly implemented by speech-to-text or Transformer. The fundamental flaw of this type of single-modality method is that it cannot take advantage of the complementary advantages and inherent connections between different information in outbound calls, and the resulting recognition results may be relatively biased. In addition, due to the limited amount of single-modality data, it is necessary to use as much information as possible, even useless information, but useless information may mislead the model to produce erroneous results. The related technology cannot eliminate the redundant components in single-modality data (such as text data), which may directly affect the accuracy of the recognition results, resulting in a decrease in the overall recognition accuracy.
[0086] In view of this, a method and system for identifying a conversation tendency based on text structuring and multimodal fusion are provided in an embodiment of the present application. The scheme divides the initial corpus data into multiple voice data segments based on conversation turns and multi-party voice data based on roles, and combines the conversation keywords extracted from the questioner's voice into multiple voice data segments to generate a labeled structured text. The key information in the initial corpus data can be marked, and the interference of non-key information on the recognition results can be reduced, so that the model is more focused on effective information, and the model processing efficiency can be improved. Such a structured text is input into a deep learning model based on an attention mechanism to predict the text positive and negative tendency probabilities of the initial corpus data, and the obtained positive and negative tendency probabilities are highly accurate. The speech data of the respondent is then classified by speech speed and intonation, and the emotional tendency probability is extracted as an auxiliary judgment. A new dimension consideration is added on the basis of text analysis, so that the accuracy and reliability of the obtained conversation tendency recognition results are further improved.
[0087] Figure 3 is an optional step diagram of a method for identifying a conversation tendency based on text structuring and multimodal fusion provided in an embodiment of the present application. Figure 4 : is a schematic diagram of the overall process of the conversation tendency identification method based on text structuring and multimodal fusion provided in an embodiment of the present application (the conversation between customer service and the customer is taken as an example in the figure), Figure 3 and Figure 4 The method may include but is not limited to steps S100 to S700.
[0088] Step S100, segmenting the initial corpus data to obtain multiple segments of speech data based on dialogue turns and multi-party speech data based on roles; wherein the multi-party speech data includes the questioner's speech data and the answerer's speech data.
[0089] Furthermore, the initial corpus data is voice data in the form of a dialogue. There may be two or more characters in the dialogue. These characters are divided into questioners and answerers in the dialogue. In the embodiment of the present application, based on the rounds of the dialogue, the initial corpus data is divided into multiple segments of voice data, such as Figure 4 As shown in the dialogue speech segments in , for example, when the questioner asks a question, until the respondent answers the question, this is a dialogue round. All the speech of one round is segmented from the speech of the next round to obtain the speech segments of each dialogue round as segmented speech data.
[0090] Based on timbre, the voice data of the questioner and the voice data of the answerer are segmented according to the roles. In this step, GMM-HMM and deep learning methods (such as i-vector) can be used for segmentation. The segmented multiple voice data will be used to generate structured text, the voice data of the questioner will be used for keyword extraction, and the voice data of the answerer will be used to assist in determining the results of the conversation tendency.
[0091] Step S200, extracting keywords from the voice data of the questioner to obtain dialogue keywords.
[0092] Specifically, the questioner's voice data is first converted into first text data, which can be converted into text by using a voice processing method such as Wav2Vec, and then conversation keywords are extracted from the first text data obtained above, which can be extracted by using RAKE (Rapid Automatic Keyword Extraction). For example, the conversation keywords obtained can be expressed as K = {K 1 ,K 2 ,…,K t ,…,K n}.
[0093] Step S300, generating a structured text with tags based on multiple voice data segments and conversation keywords.
[0094] In this step, if Figure 5 As shown, firstly, the above-mentioned multiple voice data are converted into second text data. The voice data conversion here can be performed by using the voice processing method as in step S200. The second text data can be represented as S={S 1 ,S 2 ,…,S t ,…,S n}, where S t For the tth dialogue, the text is split into subword units through word segmentation. The token sequence after word segmentation is:
[0095] T t =[[CLS],t 1 ,t 2 ,…t n [SEP]]
[0096] Among them, t i is the i-th word segmentation unit, and the total length of the sequence is i+2.
[0097] Each token (ie t i ) will be mapped to a vector space of fixed dimension, and the expression of word embedding vector is:
[0098] Etoken (t i )∈R d
[0099] Where d is the word embedding dimension.
[0100] According to the sentence to which the token belongs, each token is assigned a segment tag seg(t i ), which is used to distinguish which sentence it belongs to, and then mapped to the segment embedding vector. The segment embedding vector expression is:
[0101] E segement (t i )∈R d
[0102] Segment tags can be assigned according to the sentence number. If there is only one sentence, all tokens are marked as seg(t i )=0.
[0103] According to the position of the token in the text data to be processed, a position mark is assigned to the token to indicate the position of the token in the sequence. Assume that pos(t i ) is the sequence position of the token, then the expression of the position embedding vector is:
[0104] E position (t i )∈R d
[0105] For each token, the final input vector (which can be called the target input vector for convenience) is the sum of the above three embeddings. The expression of the target input vector is:
[0106] E(t i )=E token (t i )+E segement (t i )+E position (t i )
[0107] Combine the target input vectors of all tokens in order to form the target input matrix as follows:
[0108] E input =[E(t 0 ),E(t 1 ),…,E(t n+1 )] T
[0109] Among them, E input ∈R (n+2)*d .
[0110] In the solution of the present application, not only the dialogue text is used as the input of text structure, but also the dialogue keywords are used as auxiliary input, so that the keyword information provides explicit prompts for the model, thereby effectively improving the label recognition accuracy of the deep learning model. Specifically, after obtaining the second text data, the second text data and the dialogue keywords are spliced to obtain the text data to be processed as follows:
[0111] Input=[CLS]S[SEP]K[SEP]
[0112] The above-mentioned text data to be processed is processed by word segmentation and embedding as described above, and the token sequence obtained is:
[0113] T′ t =[[CLS],t 1 ,t 2 ,…t n [SEP]K 1 ,K,…,K t ,…,K n [SEP]]
[0114] Construct the target input vector for each token:
[0115] E′(t i )=E′ token (t i )+E′ segement (t i )+E′ position (t i )
[0116] Combine the vectors of all tokens in order to form the target input matrix:
[0117] E′ inupt =[E′(t 0 ),E′(t 1 ),…,E′(t n+1 )] T
[0118] Among them, E′ inupt ∈R N*d ; N is the sequence length; the target input matrix includes:
[0119] Token IDs: by looking up the vocabulary and converting each token into its corresponding ID;
[0120] Segment IDs: used to distinguish different text segments (such as dialogue and keywords here);
[0121] Attention Mask: It is used to indicate which positions are valid and which positions are padding. The value corresponding to the position of the valid token is 1, and the value corresponding to the position of the padding token is 0. Generally speaking, BERT needs to be padded to a fixed length. If the input text is short, the input needs to be padded.
[0122] Further reference Figure 5 As shown, the target input matrix is predicted for category labels to obtain the label prediction results for each word in the text data to be processed. Specifically, the BERT-NER model can be used to predict the category labels of the target input matrix. BERT-NER consists of multiple Transformer encoders, and the Transformer encoder consists of a multi-head self-attention mechanism and a feedforward neural network. The multi-head attention mechanism is used to capture the contextual relationship between tokens, and the nonlinear transformation of the feedforward neural network improves the feature representation capability. For the lth encoder, the input is H l-1 , the output is H l ,but:
[0123] H l =Encoder(H l-1 )
[0124] The detailed working process of the Transformer encoder can be expressed as: the vectorized input E′ inupt Feed it into the encoder, and for each input vector, generate query Q, key K, and value V matrices through three sets of linear transformations:
[0125] Q=H l-1 W Q
[0126] K=H l-1 W K
[0127] V=H l-1 W V
[0128] in, is a learnable parameter, The attention weight matrix A is calculated by scaling the dot product:
[0129]
[0130] Use the attention weights to perform weighted summation on the value vector. The Transformer uses a multi-head attention mechanism, so for the i-th head, there is the following expression:
[0131] head i=Attention(Q i ,K i ,V i )
[0132] After concatenating the results of all the headers and performing a linear transformation, we get:
[0133] MultiHead(Q,K,V)=Concat(head 1 ,head 2 ,…,head h )W o
[0134] Among them, W o is the weight matrix, which is randomly generated when the model is initialized and continuously optimized through the back propagation algorithm during the training process.
[0135] The output result H′ of the multi-head attention mechanism (l) After two layers of fully connected networks and activation function processing:
[0136] FFN(h i )=ReLU(h i W 1 +b 1 )W 2 +b 2
[0137] The characteristics of the entire sequence output are:
[0138] FFN(H′)=ReLU(H′ (l) W 1 +b 1 )W 2 +b 2
[0139] Residual connection and layer normalization are used for each sub-layer, as follows ① and ②:
[0140] ①Multi-head self-attention mechanism sublayer:
[0141]
[0142] ② Feedforward neural network sublayer:
[0143]
[0144] After passing through L Transformer encoders, the final context representation is as follows:
[0145] H=H L ∈R n*d
[0146] The obtained context representation is sent to the fully connected layer, which represents the context of each token h i ∈R d , perform a linear transformation and map it to a dimensional space with c label categories:
[0147] z i =h i W FC +b FC
[0148] Among them, z i ∈R c The unnormalized category score corresponding to the i-th token, W FC ∈R d*c is the weight matrix of the fully connected layer, b FC ∈R c is the bias vector of the fully connected layer.
[0149] For the output Z of the entire sequence, the formula is expressed as:
[0150] Z=HW FC +b FC
[0151] Then, the softmax layer normalizes the output z of the layer i The probability distribution converted to label category is as follows:
[0152]
[0153] The output of the entire token sequence can be expressed as:
[0154] P label =Softmax(Z)
[0155] Finally, the category with the highest probability in the softmax output is selected as the predicted label for each token:
[0156]
[0157] Then the predicted label of the entire sequence is expressed as:
[0158]
[0159] By predicting the label of each token, the text data to be processed is labeled and structured text with labels is generated. Figure 6 and Figure 7These are two visualization example diagrams of text structuring processing provided in this application. In the actual process of converting into structured text, important information in the original speech text will be identified and labeled for subsequent prediction.
[0160] Step S400: input the labeled structured text into a deep learning model based on the attention mechanism to predict the probability of positive and negative tendencies of the text of the initial corpus data.
[0161] Among them, the deep learning model based on the attention mechanism can be a model based on the Transformer structure, such as Figure 8 As shown, the steps for predicting the probability of positive and negative tendency of text include:
[0162] For the structured text input sequence X = [x 1 ,x 2 ,…,x n ], and convert it into a word vector by embedding layer to get:
[0163] E=Embed(X)
[0164] Among them, E∈R n*d , n is the sequence length, d is the dimension of the embedding vector. In order to obtain the position relationship of each token in the sequence, position encoding E is added position (x i ), E position (x i )∈R n*d .
[0165] The final input can be expressed as:
[0166] X input =E+E position (x i ), X input ∈R n*d
[0167] The word vector with positional encoding is sent to the encoder for encoding. The process can refer to the Transformer encoder introduced in step S300. After passing through several encoders, the final context representation H = H (L) ∈R n*d .
[0168] In order to reduce the dimension of the entire sequence into a global representation and capture the semantics of the entire text, a [CLS] tag is inserted at the first position of the input, and the [CLS] vector output by the encoder is used as the global representation, expressed as:
[0169] h CLS =H CLS ,hCLS ∈R d
[0170] Input the global representation h into the linear layer and we get:
[0171]
[0172] in, is the weight matrix, C text It is the number of categories for classification. In the embodiment of the present application, it is divided into positive and negative categories, but in some embodiments, it can be divided into more categories.
[0173] Finally, the output is converted into the probability of judgment through the softmax layer:
[0174] P text =softmax(z text )
[0175]
[0176] Among them, P text Indicates the probability of the text belonging to each category.
[0177] Step S500, classify the speaking speed of the respondent's voice data to obtain the speaking speed emotion tendency probability.
[0178] Furthermore, according to the respondent's voice data, the respondent's speaking time t is obtained. speech ; Convert the respondent's voice data into the third text data, and the conversion method can be as described in step S100; Determine the total number of words n spoken by the respondent based on the third text data word ; Calculate the total number of words spoken by the respondent n word Talking time with the respondent speech The ratio of the respondent's speaking speed characteristic S rate , the calculation formula is:
[0179]
[0180] The third text data is encoded to obtain the contextual semantic representation h text , and compare it with the respondent's speaking speed feature S rate Perform splicing to obtain the first splicing result h combined :
[0181] h combined =Concat(h text ,S rate )
[0182] The first concatenation result is weighted to obtain a first weighted result z:
[0183] z combined =W combined ·h combined
[0184] in, d′ is the dimension of the fusion feature, C speed is the number of categories of speech rate emotional tendency. The first weighted result is input into the activation function for emotional classification, and the probability of speech rate emotional tendency P is obtained. speed :
[0185] P speed =softmax(z combined )
[0186] Step S600: classify the intonation of the respondent's voice data to obtain the probability of intonation emotion tendency.
[0187] Furthermore, the respondent voice data X can be extracted by short-time Fourier transform. audio The voice feature of the respondent in the above example can be extracted by using the F0_mean() function to extract the sound spectrum feature. The voice feature is the speech spectrum feature P pitch :
[0188] P pitch =F0_mean(X audio )
[0189] It should be noted that the F0_mean() function is a function used to characterize the statistical characteristics of the fundamental frequency (Fundamental Frequency) of the sound in speech signal processing.
[0190] Combined with the weight parameter W pitch , weighted fusion of the respondent's speech features is performed to obtain a second weighted result z pitch :
[0191] z pitch =W pitch ·P pitch
[0192] The second weighted result is input into the second activation function for sentiment classification to obtain the intonation sentiment tendency probability P intonation :
[0193] P intonation =softmax(z pitch );
[0194] Step S700, obtaining the positive and negative tendency recognition result of the dialogue according to the positive and negative tendency probability of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability.
[0195] Furthermore, the above-obtained Ptext , P speed and P intonation Perform weighted fusion to obtain the probability P of positive and negative tendency of the conversation conversion , the calculation formula is as follows:
[0196] P conversion =αP text +βP speed +γP intonation ;
[0197] Among them, α, β and γ are the weights corresponding to the three probabilities; P conversion contains the positive probability P positive and negative probability P negative .
[0198] Finally, the positive probability P in the conversation tendency probability positive and negative probability P negative For comparison, if P positove >P negative , then the positive and negative tendency recognition result of the dialogue is judged to be positive, otherwise it is judged to be negative.
[0199] In summary, in steps S100 to S700 above, by dividing the initial corpus data into multiple voice data segments based on dialogue turns and multiple voice data segments based on roles, the dialogue keywords extracted from the questioner's voice are combined into multiple voice data segments to generate labeled structured text, which can mark the key information in the initial corpus data, reduce the interference of non-key information on the recognition results, make the model more focused on effective information, and improve the model processing efficiency. Such structured text is input into a deep learning model based on the attention mechanism to predict the text positive and negative tendency probabilities of the initial corpus data, and the obtained positive and negative tendency probabilities are highly accurate. Then, the speech speed and intonation of the respondent's voice data are classified, and the emotional tendency probability is extracted as an auxiliary judgment. A new dimension is added based on text analysis, so that the accuracy and reliability of the obtained dialogue tendency recognition results are further improved.
[0200] Below, taking the dialogue between the customer service (questioner) and the user (answerer) in the operator sales return visit scenario as an example, the solution of the embodiment of the present application is introduced and explained in detail:
[0201] It should be noted that in this application scenario, the user has already purchased the operator's services in the store, and the follow-up customer service will ask the user through an outbound call whether he is aware of and clear about the purchased services, in order to determine whether the sales behavior of the relevant store is compliant.
[0202] like Figure 4As shown, in an embodiment of the present application, a method for identifying a conversation tendency based on text structuring and multimodal fusion is provided, which can be implemented as the following steps in combination with the above scenario:
[0203] Step 1: Obtain the initial corpus data I of the outbound call conversation, and divide the initial corpus data into multiple voice data segments according to the conversation rounds I = {I 1 ,I 2 ,…,I t ,…,I n};
[0204] Step 2: Convert the segmented speech data into corresponding text data S = {S 1 ,S 2 ,…,S t ,…,S n} as input for text structuring;
[0205] Step 3: Use GMM-HMM and deep learning (such as i-vector) to segment the initial corpus data I into multi-party speech data according to timbre and role;
[0206] Step 4: Convert the customer service voice in the multi-party voice data into corresponding text data;
[0207] Step 5: Use RAKE (Rapid Automatic Keyword Extraction) to extract conversation keywords K from customer service text data. 1 ,K 2 ,…,K t ,…,K n};
[0208] Step 6: Generate labeled structured text based on the multiple voice data segments and the conversation keywords.
[0209] Specifically, the text data S of the conversation is 1 ,S 2 ,…,S t ,…,S n} and keyword K={K 1 ,K 2 ,…,K t ,…,K n} combined as input for text structuring:
[0210] Input=[CLS]S[SEP]K[SEP]
[0211] Segment the combined input to get:
[0212] T′ t=[[CLS],t 1 ,t 2 ,…t n [SEP]K 1 ,K,…,K t ,…,K n [SEP]
[0213] One of the input texts is {Customer Service: Did the customer service inform me that there is a penalty fee for this package? Customer: No.}, and the result is:
[0214] ['[CLS]','Customer Service',':','Please','Ask','Customer Service','Have they','informed','that','package','has','breach','fee','? ','Customer',':','No','informed','. ','[SEP]'].
[0215] The conversation keywords extracted through step 5 above are:
[0216] ['[CLS]','Notice','Package','Breach of Contract Fee','[SEP]'].
[0217] The two are combined to obtain the text data input to be processed:
[0218] ['[CLS]','Customer Service',':','Please','Ask','Customer Service','Have they','informed','that','the','package','exists',''breach','fee','? ','Customer',':','No','informed','. ','[SEP]','informed','package','breach','fee','[SEP]'].
[0219] BERT-NER constructed based on BERT-Large uses a 24-layer encoder to perform structured processing on the text. Before the text is fed into the encoder, an input vector is constructed for each input token, and the target input vector is:
[0220] E′(t i )=E′ token (t i )+E ′ segement (t i )+E′ position (t i )
[0221] Combine the vectors of all tokens in order to form the target input matrix:
[0222] E′ inupt =[E′(t 0 ),E′(t1 ),…,E′(t n+1 )] T
[0223] Among them, E′ inupt ∈R N*d , N is the sequence length;
[0224] E′ inupt Including Token IDs, Segment IDs and Attention Mask, specifically:
[0225] Token IDs: By looking up the vocabulary and converting each token to its corresponding ID.
[0226] Segment IDs: used to distinguish different text segments (such as dialogue and keywords here).
[0227] Attention Mask: It is used to indicate which positions are valid and which positions are padding. For example, the value corresponding to the position of a valid token is 1, and the value corresponding to the position of a padding token is 0. Generally speaking, BERT needs to be padded to a fixed length. If the input text is short, the input needs to be padded.
[0228] Exemplarily, the above text data to be processed is converted into:
[0229] Token IDs:[101,1134,872,2974,...].
[0230] Segment IDs:[0,0,0,0,...].
[0231] Attention Mask:[1,1,1,1,...].
[0232] For the lth encoder, the input is H l-1 , the output is H l ,but:
[0233] H l =Encoder(H l-1 )
[0234] The detailed working process of the Transformer encoder can be expressed as: the vectorized input E′ inupt Feed it into the encoder, and for each input vector, generate query Q, key K, and value V matrices through three sets of linear transformations:
[0235] Q=H l-1 W Q
[0236] K=H l-1 W K
[0237] V=H l-1 W V
[0238] in, is a learnable parameter, The attention weight matrix A is calculated by scaling the dot product:
[0239]
[0240] Use the attention weights to perform weighted summation on the value vector. The Transformer uses a multi-head attention mechanism, so for the i-th head, there is the following expression:
[0241] head i =Attention(Q i ,K i ,V i )
[0242] After concatenating the results of all the headers and performing a linear transformation, we get:
[0243] MultiHead(Q,K,V)=Concat(head 1 ,head 2 ,…,head h )W o
[0244] Among them, W o is the weight matrix, which is randomly generated when the model is initialized and continuously optimized through the back propagation algorithm during the training process.
[0245] The output result H′ of the multi-head attention mechanism (l) After two layers of fully connected networks and activation function processing:
[0246] FFN(h i )=ReLU(h i W 1 +b 1 )W 2 +b 2
[0247] The characteristics of the entire sequence output are:
[0248] FFN(H′)=ReLU(H′ (l) W 1 +b 1 )W 2 +b 2
[0249] Residual connection and layer normalization are used for each sub-layer, as follows ① and ②:
[0250] ①Multi-head self-attention mechanism sublayer:
[0251]
[0252] ② Feedforward neural network sublayer:
[0253]
[0254] After passing through L Transformer encoders, the final context representation is as follows:
[0255] H=H L ∈R n*d
[0256] For example, after being processed by the encoder, the context representation of each token is obtained as shown in Table 1 below:
[0257] Table 1
[0258]
[0259] The obtained context representation is sent to the fully connected layer, which represents the context of each token h i ∈R d , perform a linear transformation and map it to a dimensional space with c label categories:
[0260] z i =h i W FC +b FC
[0261] Among them, z i ∈R c The unnormalized category score corresponding to the i-th token, W FC ∈R d*c is the weight matrix of the fully connected layer, b FC ∈R c is the bias vector of the fully connected layer.
[0262] For the output Z of the entire sequence, the formula is expressed as:
[0263] Z=HW FC +b FC
[0264] A fully connected layer is added after the context representation of each token to predict the label of each token. Assume that there are the following categories in the label space:
[0265] Interlocutor classification: B-CUSTOMER_SERVICE, B-CUSTOMER, O;
[0266] Entity classification: B-KEY, B-QUERY, O;
[0267] Answer positive and negative: B-NEGATIVE, B-POSITIVE, O.
[0268] Among them, O represents non-critical information.
[0269] Then the fully connected layer will map the multi-dimensional vector of each token to a 6-dimensional output vector, where each dimension corresponds to a classification probability. For example, the classification probability corresponding to "customer service" is [0.05, 0.02, 0.10, 0.33, 0.11, 0.21]. Each value corresponds to the original score of a certain label, and these scores indicate the possibility that the token belongs to different labels.
[0270] Then, the softmax layer normalizes the output z of the layer i The probability distribution converted to label category is as follows:
[0271]
[0272] The output of the entire token sequence can be expressed as:
[0273] P label =Softmax(Z)
[0274] Finally, the category with the highest probability in the softmax output is selected as the predicted label for each token:
[0275]
[0276] Then the predicted label of the entire sequence is expressed as:
[0277]
[0278] For each token, choose the label with the highest probability.
[0279] After the original score of "Customer Service" passes through the Softmax layer, the final output probability distribution is: [0.152, 0.147, 0.160, 0.201, 0.162, 0.178].
[0280] Finally, each token is labeled with a label category, as shown in Table 2, which is an example table of label categories:
[0281] Table 2
[0282] Token Predicted Label [CLS] o customer service B-CUSTOMER_SERVICE : o please o ask o customer service B-CUSTOMER_SERVICE whether o have B-POSITIVE inform o Should B-QUERY combo B-QUERY exist o Breach of contract fee B-QUERY ? o customer B-CUSTOMER : o No B-NEGATIVE inform o 。 o [SEP] o inform B-KEY combo B-KEY Breach of contract fee B-KEY [SEP] o
[0283] Step 7: Input the labeled structured text into a deep learning model based on an attention mechanism to predict the positive and negative tendencies of the conversations in the initial corpus data.
[0284] Similar to BERT-NER, the labeled structured text obtained above is input and converted into a word vector through an embedding layer transformation:
[0285] E=Embed(X)
[0286] Among them, E∈R n*d , n is the sequence length, d is the dimension of the embedding vector. In order to obtain the position relationship of each token in the sequence, position encoding E is added position (x i ), E position (x i )∈R n*d .
[0287] The final input can be expressed as:
[0288] X input =E+E position (x i ), X input ∈R n*d
[0289] The word vector with positional encoding is sent to the encoder for encoding. The process can refer to the Transformer encoder introduced in step S300. After passing through several encoders, the final context representation H = H (L) ∈R n*d .
[0290] Input the global representation h into the fully connected layer and get:
[0291]
[0292] in, is the weight matrix, C text is the number of categories for classification.
[0293] Finally, the output is converted into the probability of judgment through the softmax layer:
[0294] P text =softmax(z text )
[0295]
[0296] Among them, P text∈R C It is the probability that the text belongs to each category.
[0297] Get the final predicted probability:
[0298] P Text [p pos_text ,p neg_text ]=[0.35,0.65]
[0299] Step 8: Compare the speech speed and intonation of the customer voices in the multi-party voice data. The speech spectrum characteristics can be obtained by short-time Fourier transform. Finally, the speech speed emotion tendency probability P is obtained. speed [p pos_speed ,p neg_speed ]=[0.45,0.55], probability of emotional tendency of intonation P intonation [p pos_into ,p neg_into ]=[0.40,0.60].
[0300] Step 9: According to the positive and negative tendency probability of the text, the speech rate emotional tendency probability and the tone emotional tendency probability, the positive and negative tendency recognition result of the dialogue is obtained, and then it is determined whether the store sales are compliant.
[0301] P=αP Text +βP speed +γP intonation ;
[0302] Among them, α = 0.6, β = 0.1, γ = 0.3, and the final weighted probability, that is, the probability of dialogue tendency recognition, is obtained:
[0303] P[p pos ,p neg ]=[0.375,0.625];
[0304] Since 0.375<0.625, the conversation has a negative tendency, and it is determined that the store's sales behavior to the user is not compliant.
[0305] In summary, the embodiments of the present application have at least the following beneficial effects:
[0306] 1. By setting a weighted dialogue tendency probability output design in multimodal data fusion, the voice modality and text modality data can be linked, thereby correcting the deviation of the judgment result of a single modality data and improving the accuracy of the overall recognition judgment.
[0307] 2. By setting up dialogue text structured processing and keyword extraction methods in the dialogue tendency determination process with text modal data as input, key information can be marked and the model can be assisted in extracting key information, thereby reducing the interference of non-key information on the determination results. Ultimately, the model is more focused on effective information, thereby improving the efficiency and accuracy of dialogue tendency recognition and judgment.
[0308] 3. By using the audio attributes of the respondent's voice, such as speaking speed and intonation, as one of the auxiliary judgment bases, a new dimension of consideration can be added to the traditional text-based content analysis. This not only enables a more comprehensive understanding of the communication context, but also provides additional information support based on factors such as the speaker's emotional changes, thereby further enhancing the reliability and accuracy of conversation tendency identification.
[0309] See also Fig. 9 The embodiment of the present application also provides a conversation tendency recognition system based on text structuring and multimodal fusion, which can implement the above-mentioned conversation tendency recognition method based on text structuring and multimodal fusion. The system includes:
[0310] The first module 201 is used to segment the initial corpus data to obtain multiple voice data based on dialogue turns and multi-party voice data based on roles; wherein the multi-party voice data includes the voice data of the questioner and the voice data of the answerer;
[0311] The second module 202 is used to extract keywords from the questioner's voice data to obtain dialogue keywords;
[0312] The third module 203 is used to generate a structured text with tags according to the multiple voice data and the conversation keywords;
[0313] The fourth module 204 is used to input the labeled structured text into a deep learning model based on an attention mechanism to predict the probability of positive and negative tendency of the text of the initial corpus data;
[0314] The fifth module 205 is used to classify the speech speed of the respondent's voice data to obtain the probability of speech speed emotion tendency;
[0315] The sixth module 206 is used to classify the intonation of the respondent's voice data to obtain the probability of intonation emotion tendency;
[0316] The seventh module 207 is used to obtain the positive and negative tendency recognition result of the dialogue according to the positive and negative tendency probability of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability.
[0317] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0318] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned method for identifying conversation tendency based on text structuring and multimodal fusion when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0319] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0320] See also Fig.10 , Fig.10 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0321] The processor 301 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0322] The memory 302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 302 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 302, and the processor 301 calls and executes the above method;
[0323] Input / output interface 303, used to implement information input and output;
[0324] The communication interface 304 is used to realize the communication interaction between the device and other devices. The communication can be realized through a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0325] A bus 305 that transmits information between the various components of the device (e.g., the processor 301, the memory 302, the input / output interface 303, and the communication interface 304);
[0326] The processor 301 , the memory 302 , the input / output interface 303 and the communication interface 304 are connected to each other in communication within the device via the bus 305 .
[0327] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program implements the above method when executed by a processor.
[0328] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0329] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0330] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0331] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0332] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0333] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0334] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0335] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0336] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0337] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0338] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0339] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0340] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A conversation tendency recognition method based on text structuring and multimodal fusion, characterized in that: The following steps are involved: The initial corpus data is segmented to obtain multiple segments of voice data based on dialogue turns and multi-party voice data based on roles; wherein the multi-party voice data includes the voice data of the questioner and the voice data of the answerer; Extracting keywords from the voice data of the questioner to obtain conversation keywords; Generate a structured text with tags according to the multiple voice data segments and the conversation keywords; Inputting the labeled structured text into a deep learning model based on an attention mechanism to predict the probability of positive and negative tendencies of the text of the initial corpus data; Classifying the speaking speed of the respondent's voice data to obtain the speaking speed emotion tendency probability; Performing intonation classification on the respondent's voice data to obtain intonation emotion tendency probability; According to the text positive and negative tendency probability, the speech rate emotional tendency probability and the intonation emotional tendency probability, a dialogue positive and negative tendency recognition result is obtained.
2. The method according to claim 1, characterized in that: The step of extracting keywords from the questioner's voice data to obtain conversation keywords includes the following steps: Converting the questioner's voice data into first text data; Keyword extraction is performed on the first text data to obtain conversation keywords.
3. The method according to claim 1, characterized in that The step of generating a structured text with a label according to the plurality of voice data segments and the conversation keywords comprises the following steps: Converting the plurality of voice data into second text data; The second text data and the conversation keywords are concatenated to obtain text data to be processed; wherein the text data to be processed includes a plurality of word units; Generating a target input matrix according to the text data to be processed; Perform category label prediction on the target input matrix to obtain a label prediction result for each word in the text data to be processed.
4. The method according to claim 3, characterized in that: The step of generating a target input matrix according to the text data to be processed comprises the following steps: Performing word segmentation processing on the text data to be processed to obtain word units; Performing word embedding processing on the word element to obtain a word embedding vector; According to the sentence to which the word-gram belongs, assigning a paragraph tag to the word-gram to obtain a segment embedding vector; According to the position of the word element in the text data to be processed, assigning a position tag to the word element to obtain a position embedding vector; Adding the word embedding vector, the segment embedding vector and the position embedding vector to generate a target input vector for the word unit; All the word units in the text to be processed are combined in order to obtain a target input matrix.
5. The method according to claim 1, characterized in that The method of classifying the speech rate of the respondent's voice data to obtain the speech rate emotion tendency probability includes the following steps: According to the voice data of the respondent, obtaining the speaking time of the respondent; converting the respondent voice data into third text data; Determining the total number of words spoken by the respondent based on the third text data; Calculate the ratio of the total number of words spoken by the respondent to the speaking time of the respondent to obtain the respondent's speaking speed feature; Encoding the third text data to obtain contextual semantic representation; splicing the context semantic representation with the respondent speech rate feature to obtain a first splicing result; Performing weighted processing on the first splicing result to obtain a first weighted result; The first weighted result is input into an activation function for emotion classification to obtain the probability of speech rate emotion tendency.
6. The method according to claim 1, characterized in that The step of performing intonation classification on the respondent's voice data to obtain intonation emotion tendency probability comprises the following steps: Extracting the respondent's voice features from the respondent's voice data by short-time Fourier transform; Performing weighted fusion on the respondent's speech features to obtain a second weighted result; The second weighted result is input into a second activation function to perform sentiment classification, and obtain the probability of sentiment tendency of the intonation.
7. The method according to claim 1, characterized in that The method of obtaining a dialogue positive and negative tendency recognition result according to the text positive and negative tendency probability, the speech rate emotional tendency probability and the intonation emotional tendency probability comprises the following steps: The positive and negative tendency probabilities of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability are weighted and integrated to obtain a dialogue tendency probability; Extracting positive probability and negative probability from the conversation tendency probability; When the positive probability is greater than the negative probability, the conversation positive or negative tendency is identified as positive; When the positive probability is less than or equal to the negative probability, the positive or negative tendency of the conversation is identified as negative.
8. A conversation tendency recognition system based on text structuring and multimodal fusion, characterized by: include: The first module is used to segment the initial corpus data to obtain multiple voice data based on dialogue turns and multi-party voice data based on roles; wherein the multi-party voice data includes the voice data of the questioner and the voice data of the answerer; The second module is used to extract keywords from the questioner's voice data to obtain dialogue keywords; A third module is used to generate a labeled structured text according to the multiple voice data and the conversation keywords; The fourth module is used to input the labeled structured text into a deep learning model based on an attention mechanism to predict the probability of positive and negative tendency of the text of the initial corpus data; The fifth module is used to classify the speaking speed of the respondent's voice data to obtain the speaking speed emotion tendency probability; The sixth module is used to classify the intonation of the respondent's voice data to obtain the probability of intonation emotion tendency; The seventh module is used to obtain the positive and negative tendency recognition result of the dialogue according to the positive and negative tendency probability of the text, the speech rate emotional tendency probability and the intonation emotional tendency probability.
9. An electronic device, characterized in that: including a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the method according to any one of claims 1 to 7 when executed by the processor.