Multi-level labeling method and device in dialogue scene
By conducting endpoint detection, feature extraction and cluster analysis on dialogue audio data, combined with text recognition and sentiment analysis, multi-level annotation in dialogue scenarios is achieved, and problems of incomplete dialogue records and low data utilization efficiency in the existing technology are solved, efficient and accurate information management and analysis are achieved, and the quality and efficiency of mental health services are improved.
Patent Information
- Application Number
- CN202411911380.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
The existing technology cannot accurately record the complete dialogue scenarios during psychological counseling, and lacks parallel labeling information for training and improving artificial intelligence systems, resulting in low comprehensive utilization efficiency of data.
By performing endpoint detection, feature extraction and cluster analysis on the conversation audio data, the speaker tag and speech time period are obtained, and combined with text recognition and sentiment analysis, multi-level annotation in the conversation scenario, including automatic annotation of speaker information, speech content, emotional information and inquiry strategies.
It realizes efficient and accurate information management and analysis of dialogue scenarios, improves the efficiency of dialogue records, helps to understand communication modes, monitors user emotional changes, and evaluates consultation effects, and is used to train and improve artificial intelligence systems and improves the quality and efficiency of mental health services.
Smart Images

Figure CN119943097A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech signal processing, and specifically to a multi-level annotation method and device in a dialogue scenario. Background Art
[0002] In a natural conversation environment, conversation strategies and behavioral intent recognition are one of the core tasks in a conversation system. Behavioral intent recognition aims to understand the user's goals or needs in a conversation, and conversation strategies refer to how the system chooses the appropriate response method to achieve the desired communication effect in a specific conversation context. Multi-round conversation management involves how to maintain and update the conversation state in consecutive conversation rounds to ensure that the system's response is consistent with the user's intention. In this process, sentiment analysis is also needed to understand the user's emotional state, and combined with personalized learning to predict the user's preferences and behaviors. This technology can help the system better adapt to user needs and emotional changes.
[0003] Labeling conversation strategies and behavioral intentions can help train and improve artificial intelligence systems, and can play a big role in future consulting services, such as automatic emotion recognition, consulting strategy recommendations, etc., thereby improving the overall quality and efficiency of services. However, traditional methods are basically single labeling, such as only speaker labeling, text labeling, or only emotion labeling. This cannot accurately record the complete conversation scene during psychological counseling, is not conducive to monitoring the emotional changes of patients and the doctor's inquiry strategy, and also lacks parallel labeling information for training and improving artificial intelligence systems, so the comprehensive utilization efficiency of data is low.
[0004] In addition, actual conversation data often contains multiple speakers, a large amount of overlapping speech, and useless "noise" fragments. These fragments will introduce a lot of uncertainty in the modeling and reasoning process, affecting the output accuracy of subsequent speech recognition, emotion recognition and other systems. Summary of the invention
[0005] The present application provides a multi-level annotation method in a conversation scenario to solve the problem in the prior art that the traditional method cannot accurately record the complete conversation scenario in the psychological counseling process, is not conducive to monitoring the patient's emotional changes and the doctor's inquiry strategy, and lacks an automatic annotation method for the doctor's inquiry strategy and patient intention in the psychological counseling scenario.
[0006] Correspondingly, the present application also provides a multi-level annotation device in a dialogue scenario, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.
[0007] In order to solve the above technical problems, the present application discloses a multi-level annotation method in a dialogue scenario, the method comprising:
[0008] Perform endpoint detection on the conversation audio data to obtain valid voice clips containing only human voices;
[0009] Extract features from valid speech segments and perform cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments;
[0010] Determine a corresponding target speech segment from the conversation audio data according to the speech time period, and perform text recognition on the target speech segment to obtain text data;
[0011] Based on the text data and target speech segment corresponding to the same speaker label, the text emotion features and speech emotion features of the corresponding speaker are extracted respectively;
[0012] Emotion classification is performed based on text emotion features and speech emotion features to obtain the emotion recognition result of the corresponding speaker;
[0013] The large language model is fine-tuned through preset prompt words, and the target speech segment, text data and emotion recognition results of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy or intention understanding.
[0014] The present application also discloses a multi-level annotation device in a dialogue scenario, the device comprising:
[0015] An endpoint detection module is used to perform endpoint detection on the conversation audio data to obtain valid voice segments containing only human voices;
[0016] The speaker log module is used to extract features from valid speech segments and perform cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments;
[0017] A speech recognition module is used to determine a corresponding target speech segment from the conversation audio data according to the speech time period, and perform text recognition on the target speech segment to obtain text data;
[0018] The multimodal sentiment analysis submodule is used to extract the text sentiment features and speech sentiment features of the corresponding speaker based on the text data and target speech segment corresponding to the same speaker label;
[0019] The multimodal sentiment analysis submodule is also used to perform sentiment classification based on text sentiment features and speech sentiment features to obtain the corresponding speaker's sentiment recognition results;
[0020] The strategy annotation module is used to fine-tune the large language model through preset prompt words, and input the target speech segment, text data and emotion recognition results of the corresponding speaker into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy and intent understanding.
[0021] The present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, one or more methods described in the present application are implemented.
[0022] The present application also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, one or more methods described in the present application are implemented.
[0023] In this application, by performing endpoint detection and automatic speech recognition methods on the conversation audio data, unreliable parts such as background noise, silence, and noisy speech in the real speech can be removed, and valid speech segments of reliable data can be obtained. Then, cluster analysis is performed on the valid speech segments to obtain the target speech segments and text data of the corresponding speaker for emotional feature extraction and emotion recognition. Finally, the large language model is fine-tuned using the preset prompt words, and all the target speech segments, text data, and emotion recognition results of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence to achieve automatic annotation of inquiry strategies and intention understanding. Therefore, this application can perform multi-level annotation on the recorded conversation speech, including speaker information, speech content, emotional information, and inquiry strategies adopted by the inquirer, etc., to provide an efficient and accurate information management and analysis tool for the conversation scene. Through this automated processing method, the efficiency of organizing conversation records can be greatly improved, helping people better understand the communication mode during the conversation process, monitor user emotional changes, and evaluate the consulting effect. In addition, this labeled information can also be used to train and improve artificial intelligence systems, enabling them to play a greater role in future psychological counseling services, such as automatic emotion recognition, counseling strategy recommendations, etc., thereby improving the overall quality and efficiency of mental health services.
[0024] Additional aspects and advantages of the present application will be given in the following description, which will become apparent from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0026] Figure 1 A flowchart of a multi-level annotation method in a conversation scenario provided by an embodiment of the present application;
[0027] Figure 2 This is a general block diagram of the automatic annotation system in the dialogue scenario provided by the embodiment of the present application;
[0028] Figure 3 An audio endpoint detection flow chart provided for an embodiment of the present application;
[0029] Figure 4 A flow chart of speaker clustering provided in an embodiment of the present application;
[0030] Figure 5 A flow chart of speech recognition provided in an embodiment of the present application;
[0031] Figure 6 ASR model sequence to sequence framework diagram provided in the embodiment of the present application;
[0032] Figure 7 A multimodal sentiment analysis flow chart provided for an embodiment of the present application;
[0033] Figure 8 The automatic annotation process of the doctor's inquiry strategy based on the large model provided in the embodiment of the present application;
[0034] Fig. 9 The automatic annotation process based on the large model patient intention strategy provided in the embodiment of the present application;
[0035] Fig.10 A schematic diagram of the structure of a vehicle speed measuring device based on road monitoring video provided in an embodiment of the present application;
[0036] Fig.11 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0037] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.
[0038] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.
[0039] Those skilled in the art will appreciate that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the field to which the present invention belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0040] The solution provided in the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server, wherein the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application is not limited here. For the technical problems existing in the prior art, the multi-level annotation method and device in the dialogue scenario provided by this application is intended to solve at least one of the technical problems of the prior art.
[0041] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0042] The present application embodiment provides a possible implementation method, such as Figure 1 As shown, a flowchart of a multi-level annotation method in a dialogue scenario is provided, and the scheme can be executed by any electronic device, and optionally, can be executed on a server side or a terminal device.
[0043] like Figure 1 and Figure 2 As shown, the method may include the following steps:
[0044] Step 101, performing endpoint detection on the conversation audio data to obtain a valid voice segment containing only human voice.
[0045] Step 102 , extract features from the valid speech segments, and perform cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments.
[0046] In the embodiment of the present application, the problem of "who spoke and when" is solved based on the speaker log, that is, the speech activities of different speakers in the conversation audio data are distinguished and located. The speaker log includes tasks such as speech endpoint detection and speaker segmentation clustering. Depending on whether there is preset speaker information, it can be divided into supervised clustering and unsupervised clustering.
[0047] Speech endpoint detection specifically involves identifying the silent and spoken parts in the conversation audio data and annotating the information of the spoken part. Through speech endpoint detection, unreliable parts such as background noise, silence, and noisy speech in the real conversation audio data can be removed, and the audio clips of reliable data can be obtained.
[0048] Speaker segmentation and clustering specifically involves extracting features from the audio based on the results of endpoint detection, performing cluster analysis on the features, and obtaining speaker information.
[0049] There are two main research models for speaker logging: (1) The modular model relies on multiple independent optimization modules to first cut long audio into short segments and extract speaker embeddings, and then use unsupervised clustering methods to distinguish segments belonging to different speakers. (2) The end-to-end model treats the speaker logging task as a multi-label classification problem and trains a neural network model to directly predict the speaker probability of each person in the audio frame.
[0050] The modular model can better handle audio data with an unknown number of speakers and a long duration. Currently, the duration of a single psychological consultation is generally about one hour, so the embodiment of the present application chooses a modular model for processing. In the modular method, the performance of the clustering module is crucial. In the embodiment of the present application, clustering methods such as spectral clustering (SC) or agglomerative hierarchical clustering (AHC) can be used to distinguish the speech segments of different speakers. Follow the global clustering process, that is, after extracting the voiceprint embedding of the valid voice segment, calculate the pairwise similarities between all embeddings to construct an N×N affinity matrix, and use this affinity matrix for clustering.
[0051] Step 103, determining a corresponding target speech segment from the conversation audio data according to the speech time period, and performing text recognition on the target speech segment to obtain text data.
[0052] In the embodiment of the present application, text recognition is performed on valid speech segments based on automatic speech recognition (Automatic Speech Recognition, ASR) technology. Automatic speech recognition refers to the process of converting human speech signals into readable text. Deep learning technologies such as convolutional neural networks (CNN) and recurrent neural networks (RNN) can be applied to automatic speech recognition. In particular, the sequence-to-sequence (Seq2Seq) model and the introduction of attention mechanism to automatic speech recognition can enable the automatic speech recognition system to better process long sequence data. By performing text recognition on the valid speech segments obtained after endpoint detection through automatic speech recognition technology, accurate text transcription can be achieved. Text recognition results are such as: "Doctor: What time did you go to bed yesterday?", "Patient: Ten o'clock."
[0053] Step 104 , based on the text data and the target speech segment corresponding to the same speaker label, respectively extract the text emotion features and speech emotion features of the corresponding speaker.
[0054] Step 105, performing emotion classification according to the text emotion features and the speech emotion features to obtain the emotion recognition result of the corresponding speaker.
[0055] Emotion recognition results can be happiness, sadness, anger, fear, surprise, disgust, neutral, etc.
[0056] In the embodiment of the present application, text and audio information are jointly analyzed to determine the emotional state of the target. Through multimodal sentiment analysis, the accuracy of the emotion recognition results can be improved.
[0057] Step 106, fine-tune the large language model through preset prompt words, and uniformly input the target speech segment, text data and emotion recognition results of the corresponding speaker into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy or intention understanding.
[0058] In the embodiments of the present application, a large language model is used to achieve intention understanding and inquiry strategy calibration, and the calibration results include encouragement, questioning, restatement, emotional response, challenge, etc.
[0059] In the embodiment of the present application, by performing endpoint detection and automatic speech recognition method on the conversation audio data, unreliable parts such as background noise, silence, noisy speech, etc. in the real speech can be removed, and the effective speech segments of reliable data can be obtained. Then, the effective speech segments are clustered and analyzed to obtain the target speech segments and text data of the corresponding speaker to extract emotional features and recognize emotions. Finally, the large language model is fine-tuned using the preset prompt words, and all the target speech segments, text data and emotion recognition results of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence to realize the automatic annotation of inquiry strategy and intention understanding. Therefore, the embodiment of the present application can perform multi-level annotation on the recorded conversation speech, including speaker information, speech content, emotional information, and inquiry strategy adopted by the inquirer, etc., to provide an efficient and accurate information management and analysis tool for the conversation scene. Through this automated processing method, the efficiency of organizing conversation records can be greatly improved, helping people to better understand the communication mode in the conversation process, monitor user emotional changes, and evaluate the consulting effect. In addition, this labeled information can also be used to train and improve artificial intelligence systems, enabling them to play a greater role in future psychological counseling services, such as automatic emotion recognition, counseling strategy recommendations, etc., thereby improving the overall quality and efficiency of mental health services.
[0060] In an optional embodiment, endpoint detection is performed on the conversation audio data to obtain a valid voice segment containing only human voice, including:
[0061] The endpoint detection method is used to perform endpoint detection on the conversation audio data, and the silent segments, noise segments and music segments in the conversation audio data are removed to obtain the valid speech segments containing only human voices;
[0062] Among them, the endpoint detection method is a short-time energy, zero-crossing rate detection and spectral entropy hybrid detection method.
[0063] In the present application embodiment, Figure 3 As shown, the input conversation audio data segment is first divided into small segments of speech, and a pre-trained voice endpoint detection (VAD) model is used to detect whether it is speech, and it is annotated (if so, it is marked as speech, if not, it is marked as silence) and then output.
[0064] The speech endpoint detection method can be a signal-based endpoint detection method, a machine learning-based endpoint detection method, or a deep neural network-based endpoint detection method.
[0065] VAD based on signal processing mainly focuses on short-time energy, zero-crossing rate detection and spectral entropy:
[0066] (1) Short-time energy: There is a significant difference in energy between speech segments and non-speech segments. The short-time energy is calculated by the following formula:
[0067]
[0068] Wherein, x is the audio signal sampling point, and N is the number of sampling points detected each time.
[0069] (2) Zero-crossing rate detection: The number of times the audio changes in a short period of time. A low zero-crossing rate indicates a speech segment, whereas a high zero-crossing rate indicates a non-speech segment. The calculation is as follows:
[0070]
[0071] (3) Spectral entropy: The spectral entropy of speech segments is large, while the spectral entropy of non-speech segments is small. Therefore, the spectral entropy of the normalized signal power spectral density P is calculated as follows:
[0072]
[0073] VAD based on machine learning is mainly implemented through Gaussian mixture model (GMM), support vector machine (SVM), random forest (RF), deep belief network (Deep BeliefNetworks) and conditional random field (CRF).
[0074] VAD based on signal processing and VAD based on machine learning are often limited by the model's expressiveness. When some noise signals approximate valid speech in time and frequency, the performance will drop significantly. Thanks to the rapid development of deep learning, some more complex models based on deep neural networks have gradually been used for VAD in recent years, such as long short-term memory networks (LSTM), gated recurrent unit networks (GRU) and temporal convolutional networks (TCN). These models regard VAD as a sequence learning task and have achieved good results. Some network structures with larger parameter scales, such as Transformer, can simultaneously complete VAD, speech enhancement and speech separation tasks through self-supervised learning, so that the segmented speech signal only contains clean human voices.
[0075] Existing speaker segmentation and clustering techniques are less effective when processing large-scale speech data. In the embodiments of the present application, production-scale corpora typically contain tens of thousands of hours of speech, which are composed of a large number of small files. Referring to the TFRecord format used in Tensorflow and AIStore, this format uses GNU tar to package each group of small files into a larger fragment. For large data sets, dynamic decompression will be performed during the training phase to read the fragmented files sequentially into memory. On the other hand, for small data sets, traditional data loading functions are supported to load the original files directly from disk. Very competitive results have been achieved on multiple data sets. Deployment code compatible with CPUs and GPUs is also integrated to bridge the gap between research and production systems.
[0076] In an optional embodiment, feature extraction is performed on the valid speech segment, and cluster analysis is performed on the extracted features to obtain speaker labels and corresponding speech time segments, including:
[0077] Perform pre-emphasis, fast Fourier transform and Mel frequency conversion on the valid speech segments to obtain FBANK features;
[0078] Map the FBANK features to a low-dimensional latent space to obtain embedded features;
[0079] Hierarchical clustering analysis is performed on the embedded features to obtain speaker labels and corresponding speech time segments.
[0080] like Figure 4 As shown, in the embodiment of the present application, feature extraction and embedded feature extraction are performed on the valid speech segment after segmentation and marking, the distance matrix between features is calculated, and then the inter-cluster distance (such as the average inter-cluster distance) is calculated, cluster analysis is performed based on the distance between features and the inter-cluster distance, the clusters with the minimum inter-cluster distance are merged, and the inter-cluster distance matrix is updated. Finally, it is determined whether the preset number of clusters is reached. If it is reached, the marked data is output, otherwise the above calculation and merging steps are continued to be looped until the preset target is reached.
[0081] Perform FBANK feature extraction on valid speech segments to obtain FBANK features, as follows:
[0082] (1) Pre-emphasis: High-frequency emphasis is given to the audio priority part:
[0083] s(t)=x(t)-αx(t-1) (5)
[0084] Wherein, s(t) is the pre-emphasized signal, x(t) is the original signal, and α is the pre-emphasis coefficient.
[0085] (2) Fast Fourier Transform:
[0086]
[0087] Where X(f) is the spectrum, x(n) is the time domain signal, f is the frequency, and N is the length of the frame.
[0088] (3) Mel frequency conversion:
[0089] f Mel =2595log 10 (1+f / 700) (7)
[0090] Where f is the frequency.
[0091] The obtained Mel band energy (logarithm value) is the FBANK feature. The FBANK feature of each frame can be represented as a vector, and the dimension depends on the number of Mel filters (usually 20 to 40). The FBANK feature vector of each frame can form a sequence to form a feature representation of the entire speech segment.
[0092] After FBANK feature extraction, the Mel energy of each frame is a relatively low-level feature representation, which describes the short-time spectrum structure of the speech signal. In order to capture more semantic information and high-level features in the speech signal, it is usually necessary to further map these features to a lower-dimensional, more compact, and more expressive space. This process is called embedded feature extraction, which aims to map high-dimensional features (such as FBANK features) to a low-dimensional latent space to obtain embedded features. This is usually achieved through a trained neural network model, the structure of which can be a feedforward network such as a convolutional neural network or a time-delay neural network.
[0093] Perform hierarchical clustering analysis on the embedded features to obtain preliminary speaker classification results. The specific process is as follows:
[0094] One of the core of hierarchical clustering is to calculate the distance between data points, usually calculated by Euclidean distance, cosine similarity, and Manhattan distance.
[0095] (1) Euclidean distance: Euclidean distance is the most common distance metric and is applicable to numerical data. Given two points p = (p1, p2, ..., p n ) and q=(q1,q2,…,q n ), their Euclidean distance is defined as:
[0096]
[0097] Among them, p i and q i are the coordinates of point p and point q in the i-th dimension, and n is the dimension of the data.
[0098] (2) Cosine similarity:
[0099]
[0100] Where p·q is the dot product of vectors p and q, and ||p||, ||q|| are the moduli of the vectors.
[0101] (3) Manhattan distance: It is the sum of the absolute differences between two points in each dimension. Given two points p = (p1, p2, ..., p n ) and q=(q1,q2,…,q n ), the calculation formula is:
[0102]
[0103] Among them, p i and q i are the coordinates of point p and point q in the i-th dimension, and n is the dimension of the data.
[0104] The distance between two clusters also needs to be calculated during the clustering process. Common methods for calculating inter-cluster distance are as follows:
[0105] (1) Minimum distance: The distance between two clusters is defined as the distance between the two closest points in the cluster. That is, for cluster C1 and cluster C2, the distance calculation formula is:
[0106]
[0107] (2) Maximum distance: The distance between two clusters is defined as the distance between the two farthest points in the cluster. That is, for cluster C1 and cluster C2, the distance calculation formula is:
[0108]
[0109] (3) Average distance: The distance between two clusters is defined as the average distance between all point pairs. That is, for clusters C1 and C2, the distance calculation formula is:
[0110]
[0111] The merging process of hierarchical clustering is performed based on the distance metric between clusters. The most similar clusters are continuously merged until a large cluster is obtained or the predetermined number of clusters is reached. The steps of merging are:
[0112] (i) Calculate the distance between clusters;
[0113] (ii) Select the two clusters with the smallest distance to merge;
[0114] (iii) Update the distance matrix between clusters;
[0115] (iv) Repeat steps (i) to (iii) until all data points are merged into one cluster or the set number of clusters is reached.
[0116] Density clustering can also be used in the embodiment of the present application. A preliminary clustering result is obtained through hierarchical clustering or density clustering, and the preliminary clustering result is then used as a basis for further clustering to enhance the classification effect.
[0117] In an optional embodiment, determining a corresponding target speech segment from the conversation audio data according to the speech time period, and performing text recognition on the target speech segment to obtain text data includes:
[0118] Determining a corresponding target speech segment from the conversation audio data according to the speech time segment;
[0119] Extract features of the target speech segment and generate a Mel-spectrogram based on the extracted features;
[0120] The mel-spectrogram is processed using a pre-trained automatic speech recognition model to predict text data.
[0121] In the present application embodiment, Figure 5 As shown, after the dialogue audio data is framed and feature extracted in the above steps 103 and 104, the previously extracted features can be directly used to generate a mel-spectrogram, and a pre-trained automatic speech recognition (ASR) model can be used for training to obtain predicted text data.
[0122] In an optional embodiment, the mel-spectrogram is processed using a pre-trained automatic speech recognition model to predict and obtain text data, including:
[0123] The Mel-spectrogram is input into two one-dimensional convolutional layers and processed using a preset activation function to obtain a spectrum sequence graph;
[0124] Position encoding of the spectrum sequence graph;
[0125] Input the position-encoded spectrum sequence graph into the Transformer-based encoder to obtain the encoded output sequence;
[0126] The encoded output sequence is decoded by the decoder using a cross-attention mechanism, and the output is the text data;
[0127] Among them, the Transformer-based encoder includes multiple Transformer encoder blocks, and the Transformer encoder block includes a multi-layer perceptron and a self-attention module. The self-attention module captures the correlation between elements in the spectrum sequence diagram through an autonomous mechanism.
[0128] In the embodiment of the present application, the ASR model training framework is as follows Figure 6 As shown:
[0129] The Mel spectrum graph is passed into two layers of one-dimensional convolutional layers, and GELU is used as the activation function. After processing, the spectrum sequence graph will be sinusoidally encoded to introduce the position information in the sequence. Sine position encoding calculation method:
[0130]
[0131] Among them, pos is the position in the sequence, i represents the dimension of the feature, and d model is the dimension of the embedding.
[0132] The spectral sequence graph with position encoding is further input into the Transformer-based encoder, which consists of multiple Transformer encoder blocks, each of which contains a multi-layer perceptron (MLP) and a self-attention module. The self-attention mechanism can help the encoder capture the correlation between elements in the sequence. The output of the encoder module is a series of hidden layer representations, which are used to generate the final encoded output sequence. The self-attention mechanism is expressed as:
[0133]
[0134] Among them, Q is the query matrix (Query), K is the key matrix (Key), V is the value matrix (Value), d k It is the dimension of the key, used for scaling to avoid the dot product value being too large.
[0135] The encoder output is decoded using the Transformer decoder with a cross-attention mechanism. The formula for cross-attention is similar to self-attention, with the only difference being that the query comes from the hidden state of the decoder, and the key and value come from the output representation of the encoder. The cross-attention mechanism is expressed as:
[0136]
[0137] Among them, Q d is the invisible query matrix generated by the encoder, K e is the key matrix, generated by the output of the encoder, V e It is a value matrix, also generated by the output of the encoder, representing the information of the input sequence, which will be passed to the decoder after attention calculation.
[0138] In each layer of the encoder and decoder, in addition to the attention mechanism, there is also a multi-layer perceptron (MLP) layer, which usually contains two fully connected layers and a nonlinear activation function (such as ReLU or GELU). MLP helps the model extract higher-dimensional features.
[0139] MLP(x)=GELU(xW1+b1)W2+b2 (17)
[0140] Among them, W1 and W2 are weight matrices, and b1 and b2 are bias terms.
[0141] During the training phase, the ASR model formats the input into a specific task tag sequence, namely Figure 6 The tags in the multi-task training format in . For example, SOT is the start tag; EN is the language tag; TRANSCRIBE: task type tag, specifying the transcription operation performed by the model. 0.0: timestamp information.
[0142] The vector directly learned through ASR model training adaptively generates the information representation of each position in the input sequence, that is, Figure 6 The learned position encoding in improves the expressive power of the model.
[0143] Transformer uses the attention mechanism to predict the next task tag (i.e. Figure 6 The output of the decoder passes through a linear projection layer and softmax to obtain the probability distribution of each possible tag:
[0144] P(y t |y<t,X)=softmax(W o H t ) (18)
[0145] Among them, H t is the output of the decoder at time step t, W o is the projection matrix, y t Generate a tag for the tth time.
[0146] In the embodiment of the present application, a decoder of the Transformer structure is used to process features and generate a sequence of phonemes and words. The corresponding text sequence is gradually output, and a language model is applied to correct and improve the generated text to ensure semantic consistency and fluency.
[0147] In terms of speech recognition, traditional machine learning models can often achieve good performance on the data sets it has learned, but the generalization ability is not good. The training model in the embodiment of the present application uses multiple spliced supervised data sets and pre-training across multiple data sets / fields. Compared with the previous method of relying on a single data set, the robustness of the model has been greatly improved. Reducing data preprocessing operations and using end-to-end training reduces the errors generated between modules and improves the accuracy of the system.
[0148] In an optional embodiment, based on the text data and the target speech segment corresponding to the same speaker label, the text emotion features and the speech emotion features of the corresponding speaker are respectively extracted, including:
[0149] like Figure 7 As shown, text sentiment features are extracted based on the predicted text data:
[0150] Commonly used methods for extracting text sentiment features include word embedding, BERT, and sentiment dictionaries. Word embedding models (such as Word2Vec, GloVe, BERT, etc.) can be used to convert text data into vector representations. Sentiment dictionaries, syntactic features (such as sentence length, vocabulary diversity), and contextual features are used to enhance the sentiment representation of text. For example, in a sentiment dictionary, each word has a corresponding sentiment score. Assuming the dictionary score is s(w i ), then the sentiment feature of the sentence can be obtained by summing or averaging:
[0151]
[0152] Extract speech emotion features based on the target speech segment:
[0153] Commonly used speech emotion feature extraction methods include traditional audio features (such as MFCC, pitch) and deep learning models (such as convolutional neural networks). Pitch is an important feature that reflects emotion. For example, angry or excited speech usually has a higher pitch, while sad speech usually has a lower pitch. Pitch can be calculated by the following formula:
[0154]
[0155] Where τ is the time delay. Find the τ corresponding to the first peak of R(τ), and the tone f0 can be expressed as:
[0156]
[0157] Among them, τ peak is the main cycle peak position.
[0158] Sound intensity (or energy) indicates the loudness of a speech signal and can also reflect emotion. High-energy signals usually indicate emotional excitement (such as anger or excitement), while low-energy signals may indicate calm emotions (such as peace or sadness). The energy of each frame is usually calculated as the sound intensity feature:
[0159]
[0160] In an optional embodiment, emotion classification is performed according to text emotion features and speech emotion features to obtain an emotion recognition result corresponding to the speaker, including:
[0161] Align the text sentiment features with the speech sentiment features to obtain aligned sentiment features;
[0162] The aligned sentiment features are input into the transformer module, and the contextual information in the aligned sentiment features is extracted through the self-attention mechanism to obtain the fusion features;
[0163] The fused features are input into the multi-layer perceptron (MLP) classifier to obtain the emotion recognition result of the corresponding speaker.
[0164] In the present application embodiment, Figure 7 As shown in the figure, the extracted text emotion features and speech emotion features are aligned to align the speech and text data to the same time axis to ensure the timing consistency of emotion analysis. Common feature alignment methods include linear projection, shared space mapping, alignment loss, contrastive learning and other methods. For example, contrastive learning is a method for feature alignment based on similarities and differences between samples. Use contrastive loss (such as InfoNCE or Triplet Loss) to bring multimodal features with the same emotion label closer and features with different emotion labels farther apart. Suppose there is a pair of text and speech features F with the same emotion label T and F S , and negative samples with different labels The contrast loss can be defined as:
[0165]
[0166] Where sim(·,·) represents a similarity function, such as cosine similarity. By minimizing this loss, multimodal features of the same emotion are aligned to similar spatial locations.
[0167] The aligned speech emotion features and text emotion features are concatenated into a multimodal input vector and input into the transformer. The transformer uses the self-attention mechanism to calculate the relationship between different positions in the input sequence and extract the context information in the sequence. After multiple layers of transformer encoding, the encoded fusion features are finally obtained, which contain rich emotion and context information. The encoded features are then pooled to obtain H 全局 , and finally input it into the MLP classifier for judgment.
[0168] MLP contains several fully connected layers, usually using activation functions (such as ReLU):
[0169] h1=ReLU(W1H 全局 +b1) (24)
[0170] After several layers, a softmax output layer is finally connected to map the feature vector to the probability distribution of different emotion categories:
[0171] y=softmax(W out h n +b out ) (25)
[0172] Among them, W out and b out are the weights and biases of the output layer, h n Represents the output vector of the last layer of MLP. The probability of each emotion category is output through the softmax layer, and the category with the highest probability is finally selected as the result of emotion recognition. Commonly used classification labels for emotion classification are happiness, sadness, anger, fear, surprise, disgust, neutral, etc.
[0173] In an optional embodiment, the large language model is fine-tuned by using preset prompt words, and the target speech segment, text data and emotion recognition results of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence to obtain the query strategy or intention understanding, including:
[0174] Set the prompt words and fine-tune the large language model through the LoRA model;
[0175] Pre-train the fine-tuned large language model using a preset dataset;
[0176] The target speech segment, text data, and emotion recognition results are unified into a target file, and the data in the target file is input into the pre-trained large language model sentence by sentence to obtain the query strategy or intent understanding.
[0177] The large language model has shown excellent performance in tasks such as natural language generation, question answering and dialogue understanding. Its powerful language modeling ability can help capture contextual information and generate coherent text. In the embodiment of the present application, a ready-made open source large language model is used and improved, and fine-tuned by adding prompt words. Optionally, professional psychological knowledge and the name and sample of the inquiry strategy for a given psychological consultation and the user's intention category and example can be input to guide the model to generate an inquiry content with an answer or strategy with intention understanding, and help the model identify inquiry strategies with types such as interrogation tone, confirmation information, clarification, and patient intention understanding with emotional components. The data in the target file (generated according to the target speech segment, text data and emotion recognition results) is passed sentence by sentence into the large language model (LLM). By setting different weights in the attention matrix, the model can focus on the part related to the historical target file in the current target file, and the degree of attention to historical information can be dynamically adjusted, which is more suitable for processing complex contexts. After the large language model is fine-tuned by the prompt word, it is pre-trained according to the data set that has been manually annotated to strengthen the understanding ability of the large language model and enhance the analysis and annotation effect. The inquiry strategy or intention understanding with the closest score is selected as the label for automatic annotation to supplement the lack of real content records and training data in the psychological counseling environment.
[0178] The relevant inquiry strategies mainly include: encouragement, questioning, restatement, emotional response, challenge, etc. The relevant intention understanding mainly includes: seeking emotional support, solving specific problems, understanding oneself, improving mental health, overcoming trauma, etc.
[0179] In the embodiments of the present application, the pre-trained knowledge of the large language model is used to fine-tune it through a small amount of accurately labeled data to adapt to the context of specific inquiry strategies and intent understanding, and then the inquiry strategies and intent understanding are converted into semantic hierarchical representations, including intent classification, strategy decomposition, and dependency modeling. The labeled information obtained through the above training can further fill the exploration of the large language model in the aspects of inquiry strategy labeling and intent understanding labeling, can more accurately interpret the patient's potential behavioral intentions, and provide rich information support for clinical practice and research.
[0180] The following fine-tuning techniques can be used in the embodiments of the present application:
[0181] LORA (Low-Rank Adaptation) fine-tuning technology is an innovative and highly parameter-efficient model optimization strategy, which is particularly suitable for the adaptive adjustment of large pre-trained models in specific fields or tasks. Therefore, in the embodiments of this application, LORA fine-tuning technology can be used to fine-tune large language models. The core concept of this technology is to use low-rank matrix decomposition theory to perform fine and localized update operations on the weight matrix in the pre-trained deep learning model, rather than globally retraining all parameters. Specifically, for any weight matrix in the pre-trained model W , whose dimension may be very large. LoRA technology converts this high-dimensional weight matrix W into two low-rank matrices A and B T The product of AB T An approximate representation is made, where the ranks of A and B are much smaller than the rank of the original matrix W, which greatly reduces the number of required parameters. The modified weight matrix can be expressed as:
[0182] W′=W+α·AB T (26)
[0183] Among them, α is a learnable scaling factor used to control the influence of the low-rank part on the original weight. A is a matrix of dimension d×r, and B is a matrix of dimension r×d, r < d, that is, the rank of A and B is much smaller than the number of columns of W. Through this method, LoRA can reduce the computational and storage costs in the process of model fine-tuning without significantly increasing the amount of additional parameters, and can maintain the original generalization ability of the model while enhancing its performance on new tasks, achieving effective fine-tuning and optimization of key parts of the model.
[0184] P-tuningv2 is an innovative local fine-tuning technique that optimizes pre-trained language models. It aims to preserve the inherent knowledge structure of the model while enhancing its ability to handle a variety of downstream NLP tasks. Unlike traditional global parameter tuning methods, this method does not involve major changes to the overall parameters of the model. Instead, it focuses on learning and applying a set of dynamic and continuous vector representations, namely learnable prompt word embeddings, which play a key guiding role in the input sequence and help the model understand and generate outputs that match specific tasks. In practical applications, P-tuning v2 creates a series of prompt words representing different task descriptions or templates to achieve an effective transition from pre-training goals to specific task perception patterns. For example, in text classification scenarios, carefully designed prompt words can guide the model to keenly capture the core category features in the text; in question-answering scenarios, they can help the model accurately locate the answer range or construct the appropriate question context. Since only a small number of prompt-related parameters need to be optimized and adjusted, P-tuning v2 successfully reduces the demand for computing resources, shortens the training cycle, and effectively suppresses the occurrence of overfitting. At the same time, this method enhances the generalization performance of the model when facing unseen tasks by maintaining its basic language understanding ability, so that the model can quickly adapt to new task requirements while maintaining flexibility. In short, through the careful design and optimization of prompts, P-tuning v2 proposes an efficient and universal lightweight fine-tuning strategy, which enables pre-trained language models to quickly and accurately respond to various complex natural language processing challenges without undergoing large-scale retraining.
[0185] Optionally, the following open source large language model can be selected in the embodiments of the present application:
[0186] (1) ChatGLM: One of the most effective open source base models in the Chinese field, optimized for Chinese question-answering and dialogue. After bilingual training of about 1T identifiers in Chinese and English, it is supplemented by technologies such as supervised fine-tuning, feedback self-help, and human feedback reinforcement learning. It focuses on Chinese question-answering and dialogue tasks, especially excellent optimization for Chinese, suitable for Chinese query strategy labeling and intent detection, and compared with larger parameter models, it requires fewer hardware resources, suitable for rapid iteration with limited resources, and provides relatively friendly open source support and fine-tuning interfaces for quick start.
[0187] (2) Qwen / Qwen1.5 / Qwen2 / Qwen2.5: This is a series of Tongyi Qianwen large language models developed by Alibaba Cloud, including models with parameter scales of 1.8 billion (1.8B), 7 billion (7B), 14 billion (14B), 72 billion (72B), and 110 billion (110B). Models of various scales include the basic model Qwen and the dialogue model. The dataset includes a variety of data types such as text and code, covering general and professional fields, and can support context lengths of 8 to 32K. Specific optimizations have been made for alignment data related to plug-in calls. The current model can effectively call plug-ins and be upgraded to an agent.
[0188] (3) InternLM: SenseTime, Shanghai AI Lab, Chinese University of Hong Kong, Fudan University and Shanghai Jiao Tong University jointly released the "Shusheng·Puyu" (InternLM), a large language model with hundreds of billions of parameters. It is reported that "Shusheng·Puyu" has 104 billion parameters and is trained based on a "multilingual high-quality dataset containing 1.6 trillion tokens".
[0189] For example, the doctor-patient dialogue is analyzed, and the doctor-patient dialogue audio is segmented into multiple valid voice segments through step 101. For each valid voice segment, the target voice segment, text data and emotion recognition result of the corresponding speaker (doctor or patient) are obtained through steps 102 to 105. The doctor's target voice segment, text data and emotion recognition result are unified to generate the doctor text; the patient's target voice segment, text data and emotion recognition result are unified to generate the patient text.
[0190] like Figure 8 As shown, in terms of inquiry strategy annotation based on a large language model, in the embodiment of the present application, the current doctor's text (example: How did you sleep last night?) is input into the preset large language model, and the large language model is fine-tuned through the prompt word to guide the large language model to perform inquiry strategy annotation. LoRA technology is used for fine-tuning. By setting different weights in the attention matrix, the model can focus on the part of the current doctor's text related to the historical doctor's text. The degree of attention to historical information can be dynamically adjusted, and the model's understanding of the context can be strengthened. It is convenient to identify the doctor's inquiry strategy and optimize the inquiry strategy annotation results. Inquiry strategy annotations based on large language models or such as encouragement, questioning, restatement, emotional reflection, challenge, etc.
[0191] like Fig. 9As shown, in terms of intent strategy annotation based on a large language model, in an embodiment of the present application, the current patient text (example: I slept well last night.) is input into a preset large language model, and the large language model is fine-tuned through the prompt word to guide the large language model to perform intent strategy annotation. LoRA technology is used for fine-tuning. By setting different weights in the attention matrix, the model can focus on the parts of the current patient text related to the historical patient text. The degree of attention to historical information can be dynamically adjusted, and the model's understanding of the context can be strengthened. It is convenient to identify the patient's intent strategy and optimize the intent strategy annotation results. Intention strategy annotation based on a large language model or such as seeking emotional support, solving specific problems, understanding oneself, improving mental health, overcoming trauma, etc.
[0192] In terms of doctor-patient dialogue, the embodiment of this application aims to identify the inquiry content between doctor-patient interactions and automatically annotate the doctor-patient relationship, doctor-patient emotions, patient intentions, and inquiry strategies. The large language model is used to identify and automatically annotate the inquiry strategies of doctors' psychological consultations in natural environments, and to correct erroneous texts, thereby improving the accuracy and robustness of the system.
[0193] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application also provides a multi-level annotation device in a dialogue scenario, such as Fig.10 As shown, the device comprises:
[0194] The endpoint detection module 1001 is used to perform endpoint detection on the conversation audio data to obtain a valid voice segment containing only human voice;
[0195] The speaker log module 1002 is used to extract features from valid speech segments and perform cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments;
[0196] The speech recognition module 1003 is used to determine the corresponding target speech segment from the conversation audio data according to the speech time period, and perform text recognition on the target speech segment to obtain text data;
[0197] The multimodal sentiment analysis submodule 1004 is used to extract text sentiment features and speech sentiment features of the corresponding speaker based on the text data and target speech segment corresponding to the same speaker label;
[0198] The multimodal sentiment analysis submodule 1004 is further used to perform sentiment classification according to text sentiment features and speech sentiment features to obtain a corresponding speaker's sentiment recognition result;
[0199] The strategy labeling module 1005 is used to fine-tune the large language model through preset prompt words, and uniformly input the target speech segment, text data and emotion recognition results of the corresponding speaker into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy or intention understanding.
[0200] In the embodiment of the present application, by performing endpoint detection and automatic speech recognition method on the conversation audio data, unreliable parts such as background noise, silence, noisy speech, etc. in the real speech can be removed, and the effective speech segments of reliable data can be obtained. Then, the effective speech segments are clustered and analyzed to obtain the target speech segments and text data of the corresponding speaker to extract emotional features and recognize emotions. Finally, the large language model is fine-tuned using the preset prompt words, and all the target speech segments, text data and emotion recognition results of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence to realize the automatic annotation of inquiry strategy and intention understanding. Therefore, the embodiment of the present application can perform multi-level annotation on the recorded conversation speech, including speaker information, speech content, emotional information, and inquiry strategy adopted by the inquirer, etc., to provide an efficient and accurate information management and analysis tool for the conversation scene. Through this automated processing method, the efficiency of organizing conversation records can be greatly improved, helping people to better understand the communication mode in the conversation process, monitor user emotional changes, and evaluate the consulting effect. In addition, this labeled information can also be used to train and improve artificial intelligence systems, enabling them to play a greater role in future psychological counseling services, such as automatic emotion recognition, counseling strategy recommendations, etc., thereby improving the overall quality and efficiency of mental health services.
[0201] The multi-level annotation device in the dialogue scenario provided by the embodiment of the present application can achieve Figures 1 to 9 To avoid repetition, the various processes implemented in the method embodiment are not described here.
[0202] The multi-level annotation device for a conversation scenario in the embodiment of the present application can execute the multi-level annotation method for a conversation scenario provided in the embodiment of the present application, and the implementation principle thereof is similar. The actions performed by each module and unit in the multi-level annotation device for a conversation scenario in each embodiment of the present application correspond to the steps in the multi-level annotation method for a conversation scenario in each embodiment of the present application. For the detailed functional description of each module of the multi-level annotation device for a conversation scenario, please refer to the description in the corresponding multi-level annotation method for a conversation scenario shown in the previous text, which will not be repeated here.
[0203] Based on the same principle as the method shown in the embodiment of the present application, the embodiment of the present application also provides an electronic device, which may include but is not limited to: a processor and a memory; a memory for storing a computer program; a processor for executing the multi-level annotation method in the dialogue scene shown in any optional embodiment of the present application by calling a computer program. Compared with the prior art, the multi-level annotation method in the dialogue scene provided by the present application can remove unreliable parts such as background noise, silence, and noise speech in the real voice by performing endpoint detection and automatic speech recognition methods on the dialogue audio data, and obtain the valid voice segments of reliable data therein. Then, the valid voice segments are clustered and analyzed to obtain the target voice segments and text data of the corresponding speaker to extract emotional features and recognize emotions. Finally, the large language model is fine-tuned using the preset prompt words, and all the target voice segments, text data and emotion recognition results of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence, so as to realize the automatic annotation of inquiry strategy and intention understanding. Therefore, the embodiment of the present application can perform multi-level annotation on the recorded dialogue voice, including speaker information, speech content, emotional information, and inquiry strategy adopted by the inquirer, etc., to provide an efficient and accurate information management and analysis tool for the dialogue scene. This automated processing method can greatly improve the efficiency of organizing conversation records, help people better understand the communication patterns during the conversation, monitor user emotional changes, and evaluate the effectiveness of consultation. In addition, this annotation information can also be used to train and improve artificial intelligence systems, so that they can play a greater role in future psychological counseling services, such as automatic emotion recognition, counseling strategy recommendations, etc., thereby improving the overall quality and efficiency of mental health services.
[0204] In an optional embodiment, an electronic device is also provided, such as Fig.11 As shown, Fig.11 The electronic device 1100 shown may be a server, including: a processor 1101 and a memory 1103. The processor 1101 and the memory 1103 are connected, such as through a bus 1102. Optionally, the electronic device 1100 may further include a transceiver 1104. It should be noted that in actual applications, the transceiver 1104 is not limited to one, and the structure of the electronic device 1100 does not constitute a limitation on the embodiments of the present application.
[0205] Processor 1101 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1101 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0206] The bus 1102 may include a path to transmit information between the above components. The bus 1102 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 1102 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.11 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0207] The memory 1103 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.
[0208] The memory 1103 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 1101. The processor 1101 is used to execute the application code stored in the memory 1103 to implement the contents shown in the above method embodiment.
[0209] Among them, electronic devices include but are not limited to: mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Fig.11 The electronic device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0210] The server provided in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected by wired or wireless communication, and this application does not limit this.
[0211] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content in the aforementioned method embodiment.
[0212] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.
[0213] It should be noted that the computer-readable storage medium mentioned above in the present application can also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0214] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0215] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0216] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the multi-level annotation method and device in the dialogue scenario provided in the above-mentioned various optional implementations.
[0217] Computer program code for performing the operations of the present application may be written in one or more programming languages or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0218] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0219] The modules involved in the embodiments described in this application can be implemented by software or hardware. The name of the module does not limit the module itself in some cases. For example, the endpoint detection module can also be described as "an endpoint detection module for performing endpoint detection on the conversation audio data to obtain valid voice segments containing only human voices".
[0220] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A multi-level annotation method in a dialogue scenario, characterized in that: The method comprises: Perform endpoint detection on the conversation audio data to obtain valid voice clips containing only human voices; Extracting features from the valid speech segments, and performing cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments; Determining a corresponding target speech segment from the conversation audio data according to the speech time period, and performing text recognition on the target speech segment to obtain text data; Based on the text data and the target speech segment corresponding to the same speaker label, respectively extracting text emotion features and speech emotion features of the corresponding speaker; Performing emotion classification according to the text emotion features and the speech emotion features to obtain an emotion recognition result corresponding to the speaker; The large language model is fine-tuned through preset prompt words, and the target speech segment of the corresponding speaker, the text data and the emotion recognition result are uniformly input into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy or intention understanding.
2. The multi-level annotation method in a dialogue scenario according to claim 1, characterized in that: The endpoint detection of the conversation audio data to obtain a valid voice segment containing only human voice includes: The endpoint detection method is used to perform endpoint detection on the conversation audio data, and the silent segments, noise segments and music segments in the conversation audio data are removed to obtain the valid speech segments containing only human voices; The endpoint detection method is a hybrid detection method of short-time energy, zero-crossing rate detection and spectral entropy.
3. The multi-level annotation method in a dialogue scenario according to claim 1, characterized in that: The extracting features of the valid speech segments and performing cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments includes: Pre-emphasis, fast Fourier transform and Mel frequency conversion are performed on the effective speech segments to obtain FBANK features; Mapping the FBANK features to a low-dimensional latent space to obtain embedded features; A hierarchical clustering analysis is performed on the embedded features to obtain speaker labels and corresponding speech time segments.
4. The multi-level annotation method in a dialogue scenario according to claim 1, characterized in that: Determining a corresponding target speech segment from the conversation audio data according to the speech time period, and performing text recognition on the target speech segment to obtain text data, including: determining a corresponding target speech segment from the conversation audio data according to the speech time segment; Extracting features of the target speech segment and generating a mel-spectrogram according to the extracted features; The mel-spectrogram is processed using a pre-trained automatic speech recognition model to predict and obtain the text data.
5. The multi-level annotation method in a dialogue scenario according to claim 4, characterized in that: The using of a pre-trained automatic speech recognition model to process the mel-spectrogram to predict and obtain the text data includes: Input the Mel-spectrogram into two one-dimensional convolutional layers and process it using a preset activation function to obtain a spectrum sequence graph; Position-encoding the spectrum sequence diagram; Input the position-encoded spectrum sequence graph into the Transformer-based encoder to obtain the encoded output sequence; The encoded output sequence is decoded by a decoder through a cross-attention mechanism, and the text data is obtained as output; The Transformer-based encoder includes multiple Transformer encoder blocks, and the Transformer encoder blocks include a multi-layer perceptron and a self-attention module. The self-attention module captures the correlation between elements in the spectrum sequence diagram through an autonomous mechanism.
6. The multi-level annotation method in a dialogue scenario according to claim 1, characterized in that: The performing emotion classification according to the text emotion feature and the speech emotion feature to obtain the emotion recognition result of the corresponding speaker includes: Aligning the text emotion feature with the speech emotion feature to obtain an aligned emotion feature; Inputting the aligned sentiment features into a transformer module, extracting context information from the aligned sentiment features through a self-attention mechanism, and obtaining fused features; The fused features are input into a multi-layer perceptron classifier to obtain the emotion recognition result corresponding to the speaker.
7. The multi-level annotation method in a dialogue scenario according to claim 1, characterized in that: The large language model is fine-tuned by using preset prompt words, and the target speech segment, the text data and the emotion recognition result of the corresponding speaker are uniformly input into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy or intention understanding, including: Set the prompt words and fine-tune the large language model through the LoRA model; Pre-train the fine-tuned large language model using a preset dataset; The target speech segment, the text data and the emotion recognition result are unified into a target file, and the data in the target file is input into the pre-trained large language model sentence by sentence to obtain the inquiry strategy or the intention understanding.
8. A multi-level annotation device in a dialogue scene, characterized in that: The device comprises: An endpoint detection module is used to perform endpoint detection on the conversation audio data to obtain valid voice segments containing only human voices; A speaker log module is used to extract features from the valid speech segments and perform cluster analysis on the extracted features to obtain speaker labels and corresponding speech time segments; A speech recognition module, used to determine a corresponding target speech segment from the conversation audio data according to the speech time period, and perform text recognition on the target speech segment to obtain text data; A multimodal sentiment analysis submodule, for extracting text sentiment features and speech sentiment features of the corresponding speaker based on the text data and the target speech segment corresponding to the same speaker label; The multimodal sentiment analysis submodule is further used to perform sentiment classification according to the text sentiment features and the speech sentiment features to obtain the sentiment recognition result of the corresponding speaker; The strategy annotation module is used to fine-tune the large language model through preset prompt words, and uniformly input the target speech segment, the text data and the emotion recognition result of the corresponding speaker into the fine-tuned large language model sentence by sentence to obtain the inquiry strategy or intention understanding.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Vehicle head-up display method and system for hearing-disabled driver
CN120156307A
Black broadcast semantic automatic identification system and method based on artificial intelligence
CN120932677A
An artificial intelligence-based black broadcast semantic automatic identification system and method
CN120932677B
Dialogue analysis method for multiple speakers
CN121393427A
Demand information identification method and device based on voice information and large model
CN121963739A