Voiceprint recognition, identity confirmation and dialogue implementation method applied to sentiment analysis

Through the combination of voiceprint recognition, mute detection and large language model, the problem of insufficient real-time and recognition capabilities of emotion analysis in the prior art is solved, and a more efficient and natural emotional interaction experience is achieved.

CN119993168APending Publication Date: 2025-05-13南京理工大学紫金学院
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202411979535.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing multimodal emotion analysis technology has problems such as insufficient real-time, insufficient emotion recognition ability, insufficient real-time processing ability, insufficient diversity and naturalness of emotional expression, limited data dependence and generalization ability, and low matching degree with user experience.

Method used

A method for realizing voiceprint recognition, identity confirmation and dialogue applied to sentiment analysis is proposed. The source of recorded audio is confirmed through voiceprint recognition technology, and combined with mute detection and recording control, the recording and audio data are automatically terminated. At the same time, the Whisper model is used to convert speech to text, and emotional recognition and reply generation are performed through the large language model.

Benefits of technology

It improves the real-time nature of emotion analysis and the delicateness of emotion recognition, enhances the diversity and nature of emotional expression, reduces dependence on data, improves user experience, and achieves more efficient voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993168A_ABST
    Figure CN119993168A_ABST
Patent Text Reader

Abstract

The invention discloses a voiceprint recognition, identity confirmation and dialogue implementation method applied to sentiment analysis, which can well realize voiceprint recognition and identity confirmation, and can realize mute detection, recording control and recording end judgment and release. Comprising voiceprint recognition and identity confirmation: firstly, determining whether the source of recorded audio is the sound of a specific person by adopting a voiceprint recognition technology; the voiceprint recognition generally performs identity verification by extracting audio features and comparing the extracted audio features with a pre-stored template. The mathematical formula is as follows: assuming that the feature of the input audio is X = {x1, x2,..., xn} and the voiceprint template is T = {t1, t2,..., tm}, the recognition process can be represented as # imgabs0 #, and the # imgabs1 # represents the recognized user tag xj-ti2 represents the Euclidean distance between feature vectors. According to the method, sound cloning and emotion adjustment are realized through deep learning models (GANs and VAEs). In the whole process, through training and optimization of the deep neural network, high-quality emotional voice is generated, and more natural and personalized voice interaction experience is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for realizing voiceprint recognition, identity confirmation and dialogue applied to sentiment analysis, and belongs to the technical field of artificial intelligence. Background Art

[0002] With the rapid development of artificial intelligence and big data technology, multimodal sentiment analysis systems have gradually become the core application in the field of intelligent interaction. This technology integrates voice, text and visual signals to accurately perceive user emotions and is widely used in many fields, bringing disruptive changes to improving user experience and optimizing service models. However, although multimodal sentiment analysis technology has been widely used in intelligent interaction, mental health, education and other fields, it still has the following significant defects and deficiencies, which restrict its full implementation and more efficient application:

[0003] 1. The real-time nature of message transmission is slow

[0004] Multimodal sentiment analysis systems usually require transferring large amounts of data between multiple modules, including audio streams, video frames, and text content. Current technologies exhibit bottlenecks in the following aspects:

[0005] Large data volume: Real-time transmission of audio, video, and text processing results, especially high-definition video data, brings pressure on bandwidth and processing delay.

[0006] Communication delay: Even with an efficient communication framework such as ROS2, the real-time performance of the system is still difficult to guarantee under high-frequency message transmission and complex scenarios.

[0007] Difficulty of synchronization: The time synchronization requirements for multimodal data are extremely high. Any slight deviation will lead to inaccurate sentiment analysis results, especially in scenarios that require sub-second accuracy.

[0008] 2. Emotion recognition ability still needs to be improved

[0009] Although current technologies can classify common emotions (such as happiness, anger, and sadness), they are obviously insufficient in terms of the subtlety of emotion understanding and the ability to generalize to complex scenarios:

[0010] Lack of refinement: The recognition of mixed emotions (such as "sadness mixed with joy") and emotional intensity is weak, and the generated emotional expression tends to appear mechanical and lacks human depth.

[0011] Insufficient handling of modal conflicts: When the emotional results of speech, vision, and text modalities are inconsistent, the system lacks an effective conflict resolution strategy in decision-making and is prone to making judgments that deviate from reality.

[0012] Cultural and language limitations: Existing technologies are mostly trained based on standardized data sets and have poor adaptability to differences in emotional expression across cultural backgrounds and languages, especially limited effectiveness in processing dialects and accents.

[0013] 3. Insufficient real-time processing capabilities

[0014] Sentiment analysis systems need to process speech, text, and visual signals simultaneously, which places extremely high demands on computing resources. However, existing technologies have shortcomings in the following aspects:

[0015] High computational complexity: Multimodal fusion requires simultaneous processing of audio feature extraction, text semantic analysis, and visual feature analysis, and each step involves a lot of calculations, which limits the system response speed.

[0016] Large hardware resource requirements: In embedded devices or resource-constrained environments, such as robots or in-vehicle systems, the model's operating performance is difficult to meet real-time requirements.

[0017] Insufficient model inference speed: Deep learning models (such as Transformer and YOLO) have slow inference speed in complex scenarios, further dragging down the efficiency of sentiment analysis.

[0018] 4. Lack of diversity and naturalness in emotional expression

[0019] The voice or text responses generated by the current system still lack the richness and naturalness of emotional expression:

[0020] Monotonous expression: The system’s expression of the same emotion is too fixed. For example, happy voice responses often repeat the same tone, lacking richness and variation.

[0021] Emotional fuzziness: The generated voice intonation tends to be blurred when expressing complex emotions (such as excitement but with anxiety), making it difficult for users to perceive specific emotions.

[0022] Insufficient emotional depth in human-computer interaction: The ability to resonate with users emotionally is limited, making it difficult to maintain emotional consistency during long-term interactions.

[0023] 5. Data dependence and limited generalization ability

[0024] The training and reasoning of sentiment analysis technology are highly dependent on data, but data deficiencies directly affect system performance:

[0025] Training data bias: Existing emotion datasets are mostly collected in laboratory environments, lack diversity, and cannot fully represent the complex emotional expressions in the real world.

[0026] Poor domain adaptability: The system's performance in non-training scenarios often drops significantly. For example, in noisy environments or multilingual conversations, the accuracy and robustness of sentiment classification are insufficient.

[0027] High cost of data annotation: High-quality multimodal sentiment data requires precise annotation, but this process is extremely time-consuming and costly, limiting the rapid iteration of technology.

[0028] 6. Poor match with user experience

[0029] Although emotion recognition and interaction can be accomplished technically, there are still the following deficiencies in user experience:

[0030] Response delay affects the smoothness of interaction: a long processing chain increases the system response time, and users may feel that they have to wait too long.

[0031] Emotion recognition results are unreliable: When the system misjudges the user's emotions, the generated responses may cause user dissatisfaction or misunderstanding.

[0032] Unable to adapt to users’ personalized needs: Different users have different expectations for emotional services, but the system currently finds it difficult to achieve truly personalized emotional services. Summary of the invention

[0033] The purpose of the present invention is to address the defects and shortcomings of the above-mentioned prior art and propose a method for realizing voiceprint recognition, identity confirmation and conversation applied to sentiment analysis, which method can well realize voiceprint recognition and identity confirmation, and can also realize silence detection, recording control and recording end judgment and release.

[0034] The technical solution adopted by the present invention to solve the technical problem is: a method for realizing voiceprint recognition, identity confirmation and dialogue applied to sentiment analysis, the method comprising the following steps:

[0035] Step 1: Get user audio information

[0036] 1. Voiceprint recognition and identity confirmation

[0037] First, voiceprint recognition technology is used to confirm whether the source of the recorded audio is the voice of a specific person. Voiceprint recognition usually extracts audio features (such as MFCC, Mel spectrum, etc.) and compares them with pre-stored templates to authenticate the identity.

[0038] Mathematical formula: Assume that the input audio feature is X = {x1, x2, ..., x n}, the voiceprint template is T = {t1, t2, ..., t m}Then the recognition process can be expressed as:

[0039]

[0040] in, Indicates the recognized user tag||x j -t i || 2 Represents the Euclidean distance between feature vectors.

[0041] 2.Silence detection and recording control

[0042] In order to determine when to stop recording based on the duration of silence, a silence threshold needs to be set and the silence state is determined by energy detection of the audio signal.

[0043] Silence detection formula:

[0044] Assume that the instantaneous energy of the audio signal is where x k (t) is the sampling point of the audio signal. By setting a silence threshold ∈, when E(t) is less than ∈ and the duration exceeds the set silence threshold δ, it is considered to enter the silence state.

[0045] 3. Recording end determination and release

[0046] Combining the results of voiceprint recognition and silence detection, when the voice of a specific person is recognized and the continuous silence exceeds the set threshold, the system can automatically end the recording and send the audio data to the specified topic.

[0047] Silence determination and recording end formula: If the silence signal E(t) <∈ is detected for k consecutive times and the silence duration exceeds the threshold δ, the recording ends and the release mechanism is triggered:

[0048]

[0049] 4. Control of recording and publishing process

[0050] In the system, if the recording stops, the system will automatically publish the audio data. This process can be achieved through the publish-subscribe mechanism in ROS2. A buffer can be set, and when the number of audio files reaches a predetermined value, it will be published.

[0051] Release condition formula: Assume that the audio buffer is B = {b1, bx2, ..., b n}When n reaches the set number of voice segments, publish the cached data to the topic:

[0052] Publish = {1, if n ≥ N max

[0053] Where N max The maximum number of caches set triggers data publishing.

[0054] Step 2: Use whisper to convert speech to text and get text results

[0055] 1. Audio input and preprocessing

[0056] First, the audio input needs to be preprocessed and converted into a format suitable for processing by the Whisper model. Audio data is usually stored in the original PCM format and needs to be converted into a feature representation that the Whisper model can recognize.

[0057] Audio preprocessing steps:

[0058] Unify the sampling rate of audio files (Whisper's default sampling rate is 16kHz).

[0059] Convert the audio signal into a Mel Spectrogram, which is a common method for representing audio features.

[0060] Set the audio signal x(t), where t is the time point, first perform windowing and fast Fourier transform (FFT) processing: X(k) = FFT(Window(x(t)))

[0061] Then, the FFT result is mapped to the Mel frequency axis to obtain the Mel spectrum M(t,f):

[0062] M(t,f)=Mel(X(k))

[0063] 2. Whisper model for speech-to-text conversion

[0064] After the audio is preprocessed, the Whisper model is used for speech-to-text conversion. Whisper is an end-to-end speech recognition model whose input is a Mel-spectrogram and output is text.

[0065] Set the input feature to X mel (Mel spectrum), the Whisper model uses a deep neural network (Transformer) to process and decode text from audio features. The model reasoning process is as follows:

[0066]

[0067] in is the output of the Whisper model, representing the transcribed text.

[0068] 3. Output and result publication

[0069] Final recognition result ^Published to other modules in the system through ROS2 topics. Assuming that the text result is published through the topic result_say, the result publishing process is as follows:

[0070]

[0071] Step 3: The text is transmitted to the trained large language model in the server through the API, and the large model response is obtained

[0072] Specific steps 1

[0073] The client sends a request (user voice-to-text conversion result)

[0074] After the client obtains text input from the user, it sends the text input to the server through the API. The client collects the language-to-text input in this step and transmits it to the server through an HTTP request.

[0075] Operation process:

[0076] Subscribe to the text string of the speech-to-text conversion on the topic.

[0077] The client sends the input text to the server via an API request.

[0078] Specific Step 2

[0079] The server receives the request (parses the request data)

[0080] After receiving the client's request, the server first parses the request data and extracts the text input by the user. Then, the text input is passed to the large language model for processing and generates a corresponding reply.

[0081] Operation process:

[0082] The server receives the HTTP POST request and extracts the text entered by the user.

[0083] The server passes the input to the large language model for inference processing.

[0084] Specific Step 3

[0085] The principle and fine-tuning process of emotion recognition and response generation of a large language model on the server side

[0086] 1. Basic principles of large language models

[0087] The large language model of this project uses a combination of pre-training and fine-tuning to train a language model using large-scale text data. The large language model is based on the Transformer architecture. Its core idea is to generate the next word through an autoregressive generative model and capture semantic and contextual information in the process. Its basic training process is as follows: Pre-training stage: Use unsupervised learning methods to train large-scale corpus, with the goal of maximizing the probability distribution of predicting the next word under a given context. Through unsupervised objectives, the model learns the structure and statistical characteristics of the language.

[0088] The objective function of pre-training is to maximize the conditional probability:

[0089]

[0090] Among them, y t is the target word at time step t, x1,...,x t-1 is the previous context word, and T is the length of the text sequence.

[0091] Fine-tuning stage: In the fine-tuning stage, we use annotated data from a specific field (such as sentiment analysis datasets, conversation datasets, etc.) to optimize model parameters through supervised learning. We also add the objective function of a specific task (sentiment analysis) in this stage so that the model can perform better on specific tasks.

[0092] 2. Sentiment Analysis Task

[0093] The task of sentiment analysis aims to identify the emotional categories contained in the text. Sentiments can usually be divided into multiple categories such as positive, negative, neutral, etc. In some situations, the expression of emotions is more detailed, such as "sad", "happy", etc.

[0094] In the large language model, sentiment analysis is handled as a text classification task. The input text is passed into the model after preprocessing, and the model outputs the classification label corresponding to the sentiment (such as "sad", "happy").

[0095] Sentiment classification model: Sentiment analysis model is usually a multi-classification problem. By training on a large amount of labeled data, the model learns to extract features from text and classify them. Assume that the sentiment label is C = {c1, c2, ..., c m}, the goal of the model is to calculate the probability that a given input text x belongs to each sentiment category:

[0096] p(y=c k |x) where y∈C is the probability distribution of the input text x corresponding to the sentiment category y through training data, and the most likely category is selected as the prediction result.

[0097] 3. Fine-tuning process

[0098] 3.1 Fine-tuning objectives

[0099] In order to enable the self-trained large language model to recognize the sentiment of the input text and generate emotional responses, the model needs to be fine-tuned so that it can complete the sentiment analysis and sentiment response generation tasks at the same time. The goals of fine-tuning include two aspects:

[0100] 1. Sentiment classification task: Perform sentiment analysis on the input text and classify the sentiment category (such as "sad").

[0101] 2. Emotional reply generation task: Generate reply text with corresponding emotions based on the input text and emotion labels.

[0102] 3.2 Dataset Design

[0103] For fine-tuning, the required training data should include the following:

[0104] 1. Input text: Contains the user's original input text (such as "I am sad").

[0105] 2. Emotion labeling: Label the emotion category of each input text, such as "sad" or "happy".

[0106] 3. Target response: Generate a suitable response based on the input text and the sentiment tag. For example, for the sentiment tag "sad", the generated response may be "I know you are sad. @sad".

[0107] The design of the dataset can be based on an existing sentiment annotation dataset, or training data can be constructed through manual annotation. Common sentiment classification labels include:

[0108] Positive emotions: happiness, surprise

[0109] Negative emotions: fear, anger, sadness

[0110] Neutral sentiment: Neutral

[0111] 3.3 Loss Function for Fine-tuning

[0112] During fine-tuning, the model needs to minimize the loss of two tasks: the loss of the sentiment classification task and the loss of the text generation task.

[0113] Sentiment classification loss: The loss of sentiment classification usually uses cross-entropy loss. If there are C types of sentiment categories, the model predicts category c. k The probability of p(y=c k |x)

[0114] The true label is y true Then the sentiment classification loss is:

[0115]

[0116] in Is the indicator function, if the true label is c i , then it is 1, otherwise it is 0. Text generation loss: Text generation loss uses language model loss, usually maximizing the likelihood estimate (MLE) of the generated text. For input x and target text y, the generation loss is:

[0117]

[0118] Among them, y t is the target word at time step t, x is the input text, and T is the length of the target text.

[0119] Total loss function: By weighted merging of sentiment classification loss and text generation loss, the total loss is obtained:

[0120] L total =α·L emotion +β·L gen

[0121] Among them, α and β are hyperparameters used to control the relative weights between sentiment classification and text generation tasks.

[0122] 3.4 Multi-task learning and training

[0123] Through multi-task learning, the model will optimize both sentiment classification and text generation tasks during training. Specifically, the loss of sentiment analysis and generation tasks are jointly optimized, and the model will simultaneously learn how to identify the sentiment in the text and generate corresponding responses based on the sentiment.

[0124] During training, the model's parameter updates take into account the losses of both tasks, thereby improving both sentiment recognition and text generation. Ultimately, the model is able to generate natural language responses based on the input sentiment label.

[0125] 4. Emotion recognition and response generation reasoning process

[0126] The fine-tuned self-trained large language model can simultaneously complete the tasks of sentiment analysis and emotional response generation during inference. The inference process includes the following steps:

[0127] Emotion recognition: Given a user input text (such as "I am sad"), the model uses the sentiment classification module to determine the sentiment category of the text (such as "sad"). Emotion label generation: Based on the identified sentiment label, the model uses the sentiment label as context information when generating a reply to ensure that the tone of the generated reply matches the sentiment of the input text. For example, when the sentiment label is "sad", the generated reply may be: "I know you are sad. @sad".

[0128] Generate response: Based on the input text and sentiment label, the model generates response text. For example, "I know you are sad. @sad" or "I understand how you feel. @sad".

[0129] Specific Step 4

[0130] The text reply after the server's large language model inference is returned to the client through the API response. After receiving the reply, the client parses and displays it to the user.

[0131] The final output example is as follows

[0132] Positive emotions

[0133] happy

[0134] People asked: My project today won first place!

[0135] AI answer: Great! You must be very happy, congratulations! @happy (confidence: 95%)

[0136] surprise

[0137] People asked: I actually won the lottery!

[0138] AI answer: Wow, this is a big surprise! So lucky! @Surprise (Confidence: 90%)

[0139] Negative emotions

[0140] fear

[0141] Person asks: I’m afraid I’m going to mess up my presentation tomorrow.

[0142] AI answer: I understand your concerns, but don't be afraid, you will do well! @fear (confidence: 85%)

[0143] anger

[0144] Someone asked: My colleagues kept interrupting me today, it’s so annoying!

[0145] AI answer: I know you are angry, and being interrupted is annoying. @Angry (Confidence: 90%)

[0146] sad

[0147] Someone asked: I have been feeling very depressed recently and always feel unhappy.

[0148] AI answer: I know you are upset. If you need to talk, I am here. @Sad (confidence: 92%)

[0149] Neutral Emotion

[0150] neutral

[0151] Someone asked: The weather today is a bit gloomy, but also quite cool.

[0152] AI answer: Sounds like a good day for relaxation. How should we plan our day? @Neutral (Confidence: 88%)

[0153] Step 4: Voice emotion analysis

[0154] 1. Data preparation and preprocessing

[0155] Dataset construction and selection:

[0156] To ensure the generalization ability of the model, first select a diverse speech emotion dataset. For example, the RAVDESS dataset, TESS dataset, or Emo-DB, which contain multiple emotion categories (such as happiness, anger, sadness, surprise, etc.) and cover speech samples from multiple speakers.

[0157] Each audio file in the dataset contains a labeled emotion category for the training of supervised learning models.

[0158] Data preprocessing:

[0159] Denoising and augmentation: Use voice activity detection (VAD) to remove silent segments and remove background noise through noise suppression techniques such as spectral subtraction or Wiener filtering.

[0160] Framing and windowing: The speech signal is divided into several short time frames (for example, each frame is 20ms long and the overlap between frames is 50%), and each frame is windowed (such as Hamming window) to ensure time-frequency locality.

[0161] Normalization: Normalize the features of each audio clip to remove the amplitude differences between different audio samples. The normalization formula is as follows:

[0162]

[0163] Among them, μ is the mean, σ is the standard deviation, x is the original feature, are the standardized features.

[0164] 2. Feature extraction

[0165] In speech emotion analysis, the time domain and frequency domain features of audio signals can effectively describe the emotional state. We extract the following acoustic features as model input:

[0166] Mel Frequency Cepstral Coefficient (MFCC): MFCC is one of the most commonly used features in speech analysis. It converts the speech signal into a frequency domain signal and then performs a Mel filter to extract the spectrum information of the audio signal. The calculation formula of MFCC is as follows:

[0167] MFCC = DCT(log(Mel(STFT(x))))

[0168] Among them, STFT is short-time Fourier transform, Mel represents Mel-scale filter bank, and DCT is discrete cosine transform.

[0169] Zero Crossing Rate (ZCR): The zero crossing rate reflects the frequency of signal changes and can effectively distinguish speech with different emotions. For example, angry speech usually has a higher zero crossing rate. High zero crossing rate: usually appears in speech with high-frequency components, such as excitement, happiness and other emotions. At this time, the frequency of the speech signal changes quickly, resulting in an increase in the number of times it passes through the zero point. Low zero crossing rate: appears in low-frequency speech or speech with smoother emotions, such as sadness, fatigue, etc.

[0170]

[0171] in, is the indicator function, x i is the i-th speech signal sample.

[0172] Pitch: Pitch reflects the fundamental frequency of the audio signal and is usually expressed in Hertz (Hz). The fundamental frequency can be estimated by methods such as autocorrelation and harmonic analysis. High pitch: Emotions such as happiness and surprise are usually accompanied by higher pitches because the speaker raises his voice to express emotions. Low pitch: Emotions such as sadness, fatigue, or fear will cause the fundamental frequency to decrease, and the voice will sound more stable or low.

[0173] Rate of change of pitch: In excitement or surprise, the pitch may fluctuate greatly over time; while in neutral or sad emotions, the pitch changes more slowly.

[0174] Short-time energy: Short-time energy is the sum of the energy of an audio signal within a unit time window and is used to measure the strength of a speech signal.

[0175]

[0176] Where x(n) is the speech signal. w(n) is the time window function. N is the frame length. High energy: Emotions such as excitement and anger are usually accompanied by greater speech intensity, which is manifested as high short-term energy. Low energy: Emotions such as sadness and fatigue correspond to lower speech intensity and lower short-term energy. Energy fluctuations: Speech signals expressing intense emotions (such as anger or excitement) have larger short-term energy fluctuations.

[0177] Combining these three for a comprehensive analysis

[0178] Combining features such as zero crossing rate, pitch, and short-time energy can form a multidimensional feature vector for emotion classification. Suppose the feature vector of a speech signal is:

[0179] f=[ZCR,Pitch,STE]

[0180] Happy emotions may be manifested as [high ZCR, high Pitch, high STE]

[0181] Sadness may be manifested as [low ZCR, low Pitch, low STE]

[0182] Anger may manifest as [moderate or low ZCR, high Pitch, high STE]

[0183] 3. Model design

[0184] LSTM and Bi-LSTM models:

[0185] LSTM (Long Short-Term Memory) is a neural network model that is particularly suitable for processing time series data. Through its internal "gate mechanism"

[0186] (input gate, forget gate and output gate), LSTM can remember and forget important historical information in speech signals, and is particularly suitable for capturing emotional features in speech.

[0187] The mathematical formula of LSTM is as follows:

[0188] f t =σ(W f ·[h t-1 ,x t ]+b f ) (Forget Gate)

[0189] i t =σ(W i ·[h t-1 ,x t ]+b i ) (Input Gate)

[0190] (Candidate Memory)

[0191] (Memory Update)

[0192] o t =σ(W o ·[h t-1 ,x t ]+b o ) (Output Gate)

[0193] h t =o t tanh(C t ) (Hidden state)

[0194] Among them, f t ,i t , o t are the activation values ​​of the forget gate, input gate, and output gate, respectively, and C t is the cell state, h t is the hidden state, x t is the current input.

[0195] Bi-LSTM (bidirectional LSTM) is an extension of LSTM, allowing the network to process both forward and reverse time series information. In sentiment analysis, speech sentiment often has contextual dependencies, so Bi-LSTM can improve the accuracy of sentiment classification.

[0196] CNN and Transformer models:

[0197] CNN is mainly used to extract local information from features such as spectrograms. It can effectively capture local changes in emotional features through convolution operations. The convolution operation formula of CNN is as follows:

[0198]

[0199] Among them, x is the input feature, w is the convolution kernel, and * represents the convolution operation.

[0200] Transformer is based on the self-attention mechanism, which can handle long-term dependencies and capture complex emotional patterns. The self-attention mechanism enhances the performance of the model by calculating the correlation between each position in the input sequence:

[0201]

[0202] Among them, Q, K, and V are query, key, and value matrices respectively, and d k is the dimension of the key vector, and the softmax function is used to normalize the correlation.

[0203] 4. Model Training

[0204] The loss function used in model training is the cross entropy loss function, which is used to evaluate the difference between the predicted results and the actual labels in the classification task:

[0205]

[0206] Where C is the number of sentiment categories, y i is the true label, p i is the class probability predicted by the model.

[0207] The optimization algorithm uses the Adam optimizer, and its update formula is as follows:

[0208]

[0209] Among them, η is the learning rate, m t and v t are the mean and variance of the gradient, respectively, and ∈ is a constant to prevent division by zero errors.

[0210] The process of voice emotion analysis starts with data preparation and preprocessing. By selecting a variety of speech emotion datasets (RAVDESS, TESS, Emo-DB), the generalization ability of the model is ensured, and the audio data is cleaned and standardized, including denoising (using spectral subtraction or Wiener filtering), frame windowing (such as 20ms frame length, 50% overlapping Hamming window), and normalization of each frame feature to eliminate amplitude differences. In the feature extraction stage, three core acoustic features are selected: zero crossing rate, pitch, and short-term energy, which reflect signal frequency changes, pitch and speech intensity, respectively. Among them, high zero crossing rate and high pitch usually indicate happy or excited emotions, while low zero crossing rate and low pitch are related to sad or tired emotions. These features are combined into a multidimensional feature vector and input into the model. The model uses Bi-LSTM (bidirectional long short-term memory network) to capture the temporal dependency of speech, and combines CNN to extract local spectral features, and uses the self-attention mechanism of Transformer to handle long-term dependencies and complex emotional patterns. During the training process, the cross entropy loss function is used to optimize the accuracy of model predictions, and the Adam optimizer dynamically adjusts the learning rate to accelerate convergence. After the training is completed, the model outputs the emotion category and its confidence based on the input speech features. For example, for the input speech "I have a really happy day today!", the predicted emotion category may be "happy" with a confidence of 0.96. Finally, the classification ability and practical application effect of the model are verified by evaluating its performance.

[0211] Step 5: Get facial information through the camera and analyze the emotion results through yolov8 facial expression

[0212] 1. ROS2 camera data collection

[0213] In ROS2, camera nodes transmit real-time image data by publishing image messages. Image data is published to the / camera / color / image_raw topic in messages of type sensor_msgs / Image.

[0214] Image message structure:

[0215] Image data is published in RGB format, containing color information for each pixel (8-bit per channel).

[0216] Images contain metadata, such as image width, height, encoding format, timestamp, etc.

[0217] 2. YOLOv8 facial expression detection

[0218] YOLOv8 is a target detection model based on convolutional neural network (CNN) that can detect targets in images in real time. The model divides the image into multiple grid cells, predicts the bounding box and category probability within each grid, and then locates the target. In facial expression analysis, YOLOv8 is used for face detection and extracts facial regions for further expression analysis.

[0219] YOLOv8 working principle:

[0220] Input Image: YOLOv8 takes in fixed size images (640x640).

[0221] Convolution operation: Extract image features through multiple convolution layers, gradually reduce the spatial resolution and increase the number of channels to extract high-level features.

[0222] Prediction process:

[0223] YOLOv8 divides the image into S×S grids, each of which is responsible for predicting a bounding box, including coordinates, width, height, confidence, and category probability.

[0224] The coordinates and category probabilities of each grid are predicted, and the sigmoid activation function is used for classification and regression.

[0225] The face detection model identifies the face region in the image and extracts the region for subsequent processing.

[0226] 3. Facial Expression Recognition

[0227] The goal of facial expression recognition is to infer emotional states, such as "happy", "sad", and "angry", from detected face images. Facial expression recognition can be achieved through classical feature extraction methods such as local binary patterns (LBP) and convolutional neural network (CNN) methods based on deep learning.

[0228] Local Binary Pattern (LBP)

[0229] Local Binary Pattern (LBP) is a commonly used facial expression feature extraction method. It generates features that describe the local texture of an image by comparing the grayscale value relationship between the central pixel and the surrounding pixels.

[0230] Each pixel is compared with its neighbors. If the neighboring pixel value is greater than the central pixel value, it is marked as 1, otherwise it is marked as 0.

[0231] These binary values ​​are concatenated into a binary number to form the LBP feature of the pixel.

[0232] By counting the LBP values ​​of all pixels in the image, the LBP histogram of the entire face image is generated as input for subsequent classification.

[0233] LBP formula:

[0234]

[0235] Among them, I i is the neighborhood pixel, I c is the center pixel, and s(x) is the sign function:

[0236]

[0237] P is the number of neighborhood pixels.

[0238] Deep Learning Methods (CNN)

[0239] Deep learning methods automatically learn the complex features of facial expressions through convolutional neural networks (CNNs). The CNN model extracts the spatial features of the image through multiple convolutional layers and pooling layers, and performs classification through fully connected layers.

[0240] Training process: Use a dataset containing labeled facial expression images for training and use the cross-entropy loss function for optimization.

[0241] Emotion Classification

[0242] Emotion classification is to map the extracted facial features to specific emotion categories. Common emotion categories include "happy", "sad", "angry", etc. Classification uses machine learning methods such as support vector machine (SVM), decision tree, neural network, etc.

[0243] Emotion classification formula: The classification model function f learned through the training set, given the input feature x (LBP histogram and CNN feature vector), the model outputs the predicted emotion label y: y = f(x)

[0244] 4. Combination of YOLOv8 and facial expression analysis

[0245] 1. Use YOLOv8 for face detection, detect the face area in the image, and extract the face frame.

[0246] 2. Crop the face area from the image and convert it to a grayscale image in preparation for expression analysis.

[0247] 3. Use LBP or CNN to extract features from the face area to obtain facial expression features.

[0248] 4. Input the extracted features into the emotion classification model to predict the emotion label of the facial expression.

[0249] 5. Data Flow and Workflow

[0250] 1. Camera node obtains data: ROS2 camera node publishes image data through the / camera / color / image_raw topic.

[0251] 2.YOLOv8 face detection: The YOLOv8 model detects faces in the image and returns the bounding box coordinates.

[0252] 3. Expression analysis: extract the facial area and convert it into grayscale, and apply LBP or CNN method to extract expression features.

[0253] 4. Emotion classification: Input facial features into the emotion classification model, predict the emotion label and output the result.

[0254] Step 6: Split the sentences of the large model response Text parsing: Parse the AI ​​output results through regular expressions to separate the sentiment analysis part (such as "Neutral (confidence: 88%)") and the generated dialogue part (such as "It sounds like a relaxing weather, how do you plan your day?").

[0255] Emotional data processing: extract the confidence of the emotional part as a floating point value (such as 0.88) and combine it with the corresponding emotional label (such as neutral). ROS2 message definition and release:

[0256] For the conversation content, publish it to the chat_yuan topic of ROS2.

[0257] For sentiment confidence and labels, publish them to the wenben_em topic in the specified format (such as 0.88 neutral).

[0258] Step 6: Integrate the results of text, visual, and audio sentiment analysis to get the final sentiment

[0259] 1. Modeling of the credibility of each modality

[0260] In the multimodal fusion process of sentiment analysis, the design of credibility is crucial. The output of each modality (text, vision, sound) is usually a probability distribution, that is, the predicted probability of each sentiment category. We can use the confidence of this probability distribution to weight the sentiment prediction results of different modalities.

[0261] Representation of modal output

[0262] The output of each modality is a probability distribution containing sentiment labels and their corresponding confidences, usually written as:

[0263]

[0264] Among them, y i represents the emotion category, p i Represents the confidence of the modality in this category. For example, for text analysis, suppose the model outputs the sentiment label "happy" and its confidence is 0.92, and the probability of other labels is lower.

[0265] Modeling and Normalization of Modal Assurance Criteria

[0266] The output confidence of each modality pip_ipi can be further modeled and normalized through a series of functions, making the final multimodal fusion process more reasonable. Commonly used normalization methods include Softmax normalization

[0267]

[0268] The Softmax function uses exponential mapping and normalization operations to make the output result present a probability distribution. In this way, the prediction results of each modality can be standardized into probability values.

[0269] Credibility Calibration

[0270] In order to further improve the accuracy of multimodal fusion, we can introduce the "Confidence Calibration" technology. Usually, during the training process of deep learning models, the output probability of the model does not necessarily correspond to the actual probability value. Therefore, the following method can be used to calibrate the confidence of each modality:

[0271] Temperature Scaling

[0272]

[0273] Where T is the temperature hyperparameter, which controls the smoothness of the output probability. A lower temperature will make the output of the model sharper (higher confidence), while a higher temperature will smooth the output, making the probabilities of each category closer.

[0274] Through the above normalization and calibration process, the confidence of each modality can be standardized and made more consistent with the actual probability distribution, thereby providing reliable input for the subsequent fusion step.

[0275] 2. Theoretical Framework of Modal Fusion

[0276] Weighted fusion and probability distribution fusion

[0277] Weighted fusion is one of the most commonly used fusion methods for multimodal sentiment analysis. The core idea of ​​weighted fusion is to obtain the final sentiment classification result by weighted summing the output confidence of each modality. Assume that the output of the three modalities of text, vision and sound is and Respectively expressed as the emotion category and corresponding probability distribution of each modality, we define the weighted fusion method as follows:

[0278]

[0279] Among them, w text , w visual , w audio is the weight of each modality, and these weights are determined through training data or cross-validation. In sentiment analysis, we not only need to obtain the sentiment label of each modality, but also need to consider its confidence value. We can fuse the output probabilities of different modalities in a weighted manner to obtain the final fusion probability value of each sentiment category.

[0280] Bayesian Networks and Decision Fusion

[0281] In a more accurate model, we can consider using Bayesian networks for multimodal sentiment analysis. Bayesian networks can infer the optimal sentiment labels and confidence values ​​in the context of multimodal data. Through the principle of Bayesian reasoning, we can derive the final sentiment label and its confidence by constructing a joint probability model that includes the dependencies between text, vision, and sound.

[0282] Assume that the sentiment prediction results for each modality are are independent of each other, we can use Bayes' theorem for modal fusion:

[0283]

[0284] Among them, P(y) is the prior probability of sentiment label y, is the conditional probability of the modality under the sentiment label y, is the marginal probability of the modality. Through the Bayesian network model, the system can more accurately combine the information of different modalities and calculate the final sentiment category and the corresponding confidence.

[0285] 3. Further processing and optimization of confidence

[0286] In order to further optimize the confidence of the final sentiment classification, we can use the uncertainty quantification method of the model. Common uncertainty quantification methods include:

[0287] Monte Carlo Dropout: By enabling Dropout during inference, different network structures are simulated to calculate the uncertainty of the model output. This can provide confidence intervals for each sentiment category, rather than just a single confidence value. Gaussian Process: Gaussian Process can model the uncertainty of the model and help the system evaluate the credibility of sentiment analysis in different contexts.

[0288] Through these methods, the robustness of the sentiment analysis system can be further improved, especially when facing noisy data or complex emotional states, and more reliable sentiment recognition results can be provided.

[0289] Running the example: Workflow for multimodal sentiment analysis

[0290] Input Data

[0291] The system receives data from three modalities:

[0292] 1. Text input: The user said, "I am very happy today!"

[0293] 2. Visual input: The user’s facial expression is recognized as a smile,

[0294] 3. Voice input: Voice waveform shows the emotion is "happy"

[0295] Unimodal Sentiment Analysis

[0296] 1. Text modal analysis: The natural language processing model analyzes the text and predicts the sentiment as "happy" with a confidence level of 0.88

[0297] 2. Visual modality analysis: The computer vision model based on facial expressions predicts the emotion as "happy" with a confidence level of 0.85

[0298] 3. Voice modal analysis: Through voice emotion recognition, the predicted emotion is "happy" with a confidence level of 0.92

[0299] Modal Assurance Normalization and Calibration

[0300] Use temperature scaling to adjust the confidence:

[0301] 4. Text modality confidence is adjusted to 0.90

[0302] 5. Visual modality confidence is adjusted to 0.87

[0303] 6. Sound modal confidence is adjusted to 0.94

[0304] Weighted Fusion

[0305] The modal weights are determined based on historical data as αt=0.4, αv=0.3, and αa=0.3.

[0306] Fusion formula:

[0307] P fusion =α t P t +α v P v +α a P a

[0308] The final fusion confidence is calculated as:

[0309] P fusion =0.4×0.90+0.3×0.87+0.3×0.94=0.903

[0310] Final Output

[0311] The system judges the user's comprehensive emotion as "happy" with a confidence level of 90.3%.

[0312] 7. Output emotion category: happy

[0313] 8. Output confidence: 90.3%

[0314] Practical Application

[0315] The system pushes the output results to the corresponding ROS2 topic:

[0316] 9. Emotion classification result: "happy" is pushed to topic / emotion_result.

[0317] 10. Fusion confidence: 0.903 is pushed to topic / confidence_score.

[0318] Step 7: Voice Emotion Cloning Module-Topic 1: Reading Data from ROS2 Topic

[0319] 1. System Architecture Overview

[0320] In the voice emotion cloning module of this project, the system interacts with data through the ROS2 (Robot Operating System 2) framework. ROS 2 uses topics to transfer information between different nodes. The core task of this module is to read data from two ROS2 topics, one for receiving the text data to be cloned (chat_yuan) and the other for obtaining the user's current emotional state (emotion_result).

[0321] 2. Read text from the chat_yuan topic

[0322] The chat_yuan topic transmits the text data entered by the user. This text data will serve as the basic information for speech synthesis. Through the subscription mechanism of the ROS2 node, the system will obtain the text data on the topic in real time and prepare for subsequent speech cloning processing. In terms of implementation, the system uses the rclpy Python client library of ROS2 to subscribe to the topic. Whenever a topic publishes new data, the subscribed node will receive the data and pass it to the speech cloning model.

[0323] 3. Read the emotional state from the emotion_result topic

[0324] The emotion_result topic contains the user's current emotional state, such as happiness, anger, sadness, etc. The emotion analysis module obtains the user's emotional characteristics through this topic and maps them to the tone parameters of the audio output. Emotional data is generated by real-time monitoring of the user's physiological or psychological state and transmitted to the system in real time through ROS2. Based on this data, the system adjusts the voice features such as tone and intonation in the voice cloning model to make the generated voice match the user's current emotional state.

[0325] 4. Data processing and conversion

[0326] After reading the text and emotional data, the system will process the data and pass the text data to the speech synthesis module, while the emotional data will be combined with the emotional adjustment mechanism in the speech cloning process. Specifically, the emotional data will affect the pitch, speed, emotional color and other characteristics of speech synthesis to achieve the goal of emotional cloning.

[0327] Beneficial effects: The present invention realizes voice cloning and emotion adjustment through deep learning models (GANs, VAEs). Voice cloning first learns the speech features of the target voice through feature extraction, and generates natural and realistic voices through training. Subsequently, the emotion adjustment module uses reference audio to adjust the voice according to the user's emotional state, so that the output voice not only has the timbre of the target person, but also reflects the user's emotional state. The entire process generates high-quality emotional voice through the training and optimization of deep neural networks, providing a more natural and personalized voice interaction experience.

[0328] Based on the above implementations, we can have emotional AI conversations with users. BRIEF DESCRIPTION OF THE DRAWINGS

[0329] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0330] The invention will be further described in detail below in conjunction with the accompanying drawings.

[0331] like Figure 1 As shown, the present invention provides a method for implementing voiceprint recognition, identity confirmation and dialogue applied to sentiment analysis, the method comprising the following steps:

[0332] Step 1: Get user audio information

[0333] Step 1-1: Voiceprint recognition and identity confirmation

[0334] First, voiceprint recognition technology is used to confirm whether the source of the recorded audio is the voice of a specific person. Voiceprint recognition usually extracts audio features (such as MFCC, Mel spectrum, etc.) and compares them with pre-stored templates to authenticate the identity.

[0335] Mathematical formula: Assume that the input audio feature is X = {x1, x2, ..., x n}, the voiceprint template is T = {t1, t2, ..., t m}Then the recognition process can be expressed as:

[0336]

[0337] in, Indicates the identified user tag||x j -t i || 2 Represents the Euclidean distance between feature vectors.

[0338] 2.Silence detection and recording control

[0339] In order to determine when to stop recording based on the duration of silence, a silence threshold needs to be set and the silence state is determined by energy detection of the audio signal.

[0340] Silence detection formula:

[0341] Assume that the instantaneous energy of the audio signal is where x k (t) is the sampling point of the audio signal. By setting a silence threshold ∈, when E(t) is less than ∈ and the duration exceeds the set silence threshold δ, it is considered to enter the silence state.

[0342] 3. Recording end determination and release

[0343] Combining the results of voiceprint recognition and silence detection, when the voice of a specific person is recognized and the continuous silence exceeds the set threshold, the system can automatically end the recording and send the audio data to the specified topic.

[0344] Silence determination and recording end formula: If the silence signal E(t) <∈ is detected for k consecutive times and the silence duration exceeds the threshold δ, the recording ends and the release mechanism is triggered:

[0345]

[0346] 4. Control of recording and publishing process

[0347] In the system, if the recording stops, the system will automatically publish the audio data. This process can be achieved through the publish-subscribe mechanism in ROS2. A buffer can be set, and when the number of audio files reaches a predetermined value, it will be published.

[0348] Release condition formula: Assume that the audio buffer is B = {b1, b2, ..., b n}, when n reaches the set number of voice segments, the cached data is published to the topic:

[0349] Publish = {1, if n ≥ N max

[0350] Where N max The maximum number of caches set triggers data publishing.

[0351] The second step is to use whisper to convert speech to text and obtain text results

[0352] 1. Audio input and preprocessing

[0353] First, the audio input needs to be preprocessed and converted into a format suitable for processing by the Whisper model. Audio data is usually stored in the original PCM format and needs to be converted into a feature representation that the Whisper model can recognize.

[0354] Audio preprocessing steps:

[0355] Unify the sampling rate of audio files (Whisper's default sampling rate is 16kHz).

[0356] Convert the audio signal into a Mel Spectrogram, which is a common method for representing audio features.

[0357] Set the audio signal x(t), where t is the time point, first perform windowing and fast Fourier transform (FFT) processing: X(k) = FFT(Window(x(t)))

[0358] Then, the FFT result is mapped to the Mel frequency axis to obtain the Mel spectrum M(t,f):

[0359] M(t,f)=Mel(X(k))

[0360] 2. Whisper model for speech-to-text conversion

[0361] After the audio is preprocessed, the Whisper model is used for speech-to-text conversion. Whisper is an end-to-end speech recognition model whose input is a Mel-spectrogram and output is text.

[0362] Set the input feature to X mel (Mel spectrum), the Whisper model uses a deep neural network (Transformer) to process and decode text from audio features. The model reasoning process is as follows:

[0363]

[0364] in is the output of the Whisper model, representing the transcribed text.

[0365] 3. Output and result publication

[0366] Final recognition result ^Published to other modules in the system through ROS2 topics. Assuming that the text result is published through the topic result_say, the result publishing process is as follows:

[0367]

[0368] The third step is to transfer the text to the trained large language model in the server through the API and obtain the large model response

[0369] Specific steps 1

[0370] The client sends a request (user voice-to-text conversion result)

[0371] After the client obtains text input from the user, it sends the text input to the server through the API. The client collects the language-to-text input in this step and transmits it to the server through an HTTP request.

[0372] Operation process:

[0373] Subscribe to the text string of the speech-to-text conversion on the topic.

[0374] The client sends the input text to the server via an API request.

[0375] Specific Step 2

[0376] The server receives the request (parses the request data)

[0377] After receiving the client's request, the server first parses the request data and extracts the text input by the user. Then, the text input is passed to the large language model for processing and generates a corresponding reply.

[0378] Operation process:

[0379] The server receives the HTTP POST request and extracts the text entered by the user.

[0380] The server passes the input to the large language model for inference processing.

[0381] Specific Step 3

[0382] The principle and fine-tuning process of emotion recognition and response generation of a large language model on the server side

[0383] 1. Basic principles of large language models

[0384] The large language model of this project uses a combination of pre-training and fine-tuning to train a language model using large-scale text data. The large language model is based on the Transformer architecture. Its core idea is to generate the next word through an autoregressive generative model and capture semantic and contextual information in the process. Its basic training process is as follows: Pre-training stage: Use unsupervised learning methods to train large-scale corpus, with the goal of maximizing the probability distribution of predicting the next word under a given context. Through unsupervised objectives, the model learns the structure and statistical characteristics of the language.

[0385] The objective function of pre-training is to maximize the conditional probability:

[0386]

[0387] Among them, y t is the target word at time step t, x1,...,x t-1 is the previous context word, and T is the length of the text sequence.

[0388] Fine-tuning stage: In the fine-tuning stage, we use annotated data from a specific field (such as sentiment analysis datasets, conversation datasets, etc.) to optimize model parameters through supervised learning. We also add the objective function of a specific task (sentiment analysis) in this stage so that the model can perform better on specific tasks.

[0389] 2. Sentiment Analysis Task

[0390] The task of sentiment analysis aims to identify the emotional categories contained in the text. Sentiments can usually be divided into multiple categories such as positive, negative, neutral, etc. In some situations, the expression of emotions is more detailed, such as "sad", "happy", etc.

[0391] In the large language model, sentiment analysis is handled as a text classification task. The input text is passed into the model after preprocessing, and the model outputs the classification label corresponding to the sentiment (such as "sad", "happy").

[0392] Sentiment classification model: Sentiment analysis model is usually a multi-classification problem. By training on a large amount of labeled data, the model learns to extract features from text and classify them. Assume that the sentiment label is C = {c1, c2, ..., c m}, the goal of the model is to calculate the probability that a given input text x belongs to each sentiment category:

[0393] p(y=c k |x) where y∈C

[0394] Through training data, the model learns the probability distribution of the input text x corresponding to the sentiment category y, and selects the most likely category as the prediction result.

[0395] 3. Fine-tuning process

[0396] 3.1 Fine-tuning objectives

[0397] In order to enable the self-trained large language model to recognize the sentiment of the input text and generate emotional responses, the model needs to be fine-tuned so that it can complete the sentiment analysis and sentiment response generation tasks at the same time. The goals of fine-tuning include two aspects:

[0398] 3. Sentiment classification task: Perform sentiment analysis on the input text and classify the sentiment category (such as "sad").

[0399] 4. Emotional reply generation task: Generate reply text with corresponding emotions based on the input text and emotion labels.

[0400] 3.2 Dataset Design

[0401] For fine-tuning, the required training data should include the following:

[0402] 4. Input text: Contains the user's original input text (such as "I am sad").

[0403] 5. Emotion labeling: Label the emotion category of each input text, such as "sad" or "happy".

[0404] 6. Target response: Generate a suitable response based on the input text and the sentiment tag. For example, for the sentiment tag "sad", the generated response may be "I know you are sad. @sad".

[0405] The design of the dataset can be based on an existing sentiment annotation dataset, or training data can be constructed through manual annotation. Common sentiment classification labels include:

[0406] Positive emotions: happiness, surprise

[0407] Negative emotions: fear, anger, sadness

[0408] Neutral sentiment: Neutral

[0409] 3.3 Loss Function for Fine-tuning

[0410] During fine-tuning, the model needs to minimize the loss of two tasks: the loss of the sentiment classification task and the loss of the text generation task.

[0411] Sentiment classification loss: The loss of sentiment classification usually uses cross-entropy loss. If there are C types of sentiment categories, the model predicts category c. k The probability of p(y=c k |x)

[0412] The true label is y true Then the sentiment classification loss is:

[0413]

[0414] in Is the indicator function, if the true label is c i , then it is 1, otherwise it is 0. Text generation loss: Text generation loss uses language model loss, usually maximizing the likelihood estimate (MLE) of the generated text. For input x and target text y, the generation loss is:

[0415]

[0416] Among them, y t is the target word at time step t, x is the input text, and T is the length of the target text.

[0417] Total loss function: By weighted merging of sentiment classification loss and text generation loss, the total loss is obtained:

[0418] L total =α·L emotion +β·L gen

[0419] Among them, α and β are hyperparameters used to control the relative weights between sentiment classification and text generation tasks.

[0420] 3.4 Multi-task learning and training

[0421] Through multi-task learning, the model will optimize both sentiment classification and text generation tasks during training. Specifically, the loss of sentiment analysis and generation tasks are jointly optimized, and the model will simultaneously learn how to identify the sentiment in the text and generate corresponding responses based on the sentiment.

[0422] During training, the model's parameter updates take into account the losses of both tasks, thereby improving both sentiment recognition and text generation. Ultimately, the model is able to generate natural language responses based on the input sentiment label.

[0423] 4. Emotion recognition and response generation reasoning process

[0424] The fine-tuned self-trained large language model can simultaneously complete the tasks of sentiment analysis and emotional response generation during inference. The inference process includes the following steps:

[0425] Emotion recognition: Given a user input text (such as "I am sad"), the model uses the sentiment classification module to determine the sentiment category of the text (such as "sad"). Emotion label generation: Based on the identified sentiment label, the model uses the sentiment label as context information when generating a reply to ensure that the tone of the generated reply matches the sentiment of the input text. For example, when the sentiment label is "sad", the generated reply may be: "I know you are sad. @sad".

[0426] Generate response: Based on the input text and sentiment label, the model generates response text. For example, "I know you are sad. @sad" or "I understand how you feel. @sad". Specifically include:

[0427] The text reply after the server's large language model inference is returned to the client through the API response. After receiving the reply, the client parses and displays it to the user.

[0428] Operation process:

[0429] The client receives the response data from the server.

[0430] The client parses the JSON response and displays the reply text to the user.

[0431] The final output example is as follows

[0432] Positive emotions

[0433] happy

[0434] People asked: My project today won first place!

[0435] AI answer: Great! You must be very happy, congratulations! @happy (confidence: 95%)

[0436] surprise

[0437] People asked: I actually won the lottery!

[0438] AI answer: Wow, this is a big surprise! So lucky! @Surprise (Confidence: 90%)

[0439] Negative emotions

[0440] fear

[0441] Person asks: I’m afraid I’m going to mess up my presentation tomorrow.

[0442] AI answer: I understand your concerns, but don't be afraid, you will do well! @fear (confidence: 85%)

[0443] anger

[0444] Someone asked: My colleagues kept interrupting me today, it’s so annoying!

[0445] AI answer: I know you are angry, and being interrupted is annoying. @Angry (Confidence: 90%)

[0446] sad

[0447] Someone asked: I have been feeling very depressed recently and always feel unhappy.

[0448] AI answer: I know you are upset. If you need to talk, I am here. @Sad (confidence: 92%)

[0449] Neutral Emotion

[0450] neutral

[0451] Someone asked: The weather today is a bit gloomy, but also quite cool.

[0452] AI answer: Sounds like a good day for relaxation. How should we plan our day? @Neutral (Confidence: 88%)

[0453] Step 3: Voice emotion analysis

[0454] 1. Data preparation and preprocessing

[0455] Dataset construction and selection:

[0456] To ensure the generalization ability of the model, first select a diverse speech emotion dataset. For example, the RAVDESS dataset, TESS dataset, or Emo-DB, which contain multiple emotion categories (such as happiness, anger, sadness, surprise, etc.) and cover speech samples from multiple speakers.

[0457] Each audio file in the dataset contains a labeled emotion category for the training of supervised learning models.

[0458] Data preprocessing:

[0459] Denoising and augmentation: Use voice activity detection (VAD) to remove silent segments and remove background noise through noise suppression techniques such as spectral subtraction or Wiener filtering.

[0460] Framing and windowing: The speech signal is divided into several short time frames (for example, each frame is 20ms long and the overlap between frames is 50%), and each frame is windowed (such as Hamming window) to ensure time-frequency locality.

[0461] Normalization: Normalize the features of each audio clip to remove the amplitude differences between different audio samples. The normalization formula is as follows:

[0462]

[0463] Among them, μ is the mean, σ is the standard deviation, x is the original feature, are the standardized features.

[0464] 2. Feature extraction

[0465] In speech emotion analysis, the time domain and frequency domain features of audio signals can effectively describe the emotional state. We extract the following acoustic features as model input:

[0466] Mel Frequency Cepstral Coefficient (MFCC): MFCC is one of the most commonly used features in speech analysis. It converts the speech signal into a frequency domain signal and then performs a Mel filter to extract the spectrum information of the audio signal. The calculation formula of MFCC is as follows:

[0467] MFCC = DCT(log(Mel(STFT(x))))

[0468] Among them, STFT is short-time Fourier transform, Mel represents Mel-scale filter bank, and DCT is discrete cosine transform.

[0469] Zero Crossing Rate (ZCR): The zero crossing rate reflects the frequency of signal changes and can effectively distinguish speech with different emotions. For example, angry speech usually has a higher zero crossing rate. High zero crossing rate: usually appears in speech with high-frequency components, such as excitement, happiness and other emotions. At this time, the frequency of the speech signal changes quickly, resulting in an increase in the number of times it passes through the zero point. Low zero crossing rate: appears in low-frequency speech or speech with smoother emotions, such as sadness, fatigue, etc.

[0470]

[0471] in, is the indicator function, x i is the i-th speech signal sample.

[0472] Pitch: Pitch reflects the fundamental frequency of the audio signal and is usually expressed in Hertz (Hz). The fundamental frequency can be estimated by methods such as autocorrelation and harmonic analysis. High pitch: Emotions such as happiness and surprise are usually accompanied by higher pitches because the speaker raises his voice to express emotions. Low pitch: Emotions such as sadness, fatigue, or fear will cause the fundamental frequency to decrease, and the voice will sound more stable or low.

[0473] Rate of change of pitch: In excitement or surprise, the pitch may fluctuate greatly over time; while in neutral or sad emotions, the pitch changes more slowly.

[0474] Short-time energy: Short-time energy is the sum of the energy of an audio signal within a unit time window and is used to measure the strength of a speech signal.

[0475]

[0476] Where x(n) is the speech signal. w(n) is the time window function. N is the frame length. High energy: Emotions such as excitement and anger are usually accompanied by greater speech intensity, which is manifested as high short-term energy. Low energy: Emotions such as sadness and fatigue correspond to lower speech intensity and lower short-term energy. Energy fluctuations: Speech signals expressing intense emotions (such as anger or excitement) have larger short-term energy fluctuations.

[0477] Combining these three for a comprehensive analysis

[0478] Combining features such as zero crossing rate, pitch, and short-time energy can form a multidimensional feature vector for emotion classification. Suppose the feature vector of a speech signal is:

[0479] f = [ZCR, Pitch, STE]

[0480] Happy emotions may be manifested as [high ZCR, high Pitch, high STE]

[0481] Sadness may be manifested as [low ZCR, low Pitch, low STE]

[0482] Anger may manifest as [moderate or low ZCR, high Pitch, high STE]

[0483] 3. Model design

[0484] LSTM and Bi-LSTM models:

[0485] LSTM (Long Short-Term Memory) is a neural network model that is particularly suitable for processing time series data. Through its internal "gate mechanism"

[0486] (input gate, forget gate and output gate), LSTM can remember and forget important historical information in speech signals, and is particularly suitable for capturing emotional features in speech.

[0487] The mathematical formula of LSTM is as follows:

[0488] f t =σ(W f ·[h t-1 ,x t ]+b f ) (Forget Gate)

[0489] i t =σ(Wi·[h t-1 ,x t ]+b i ) (Input Gate)

[0490] (Candidate Memory)

[0491] (Memory Update)

[0492] o t =σ(W o ·[h t-1 ,x t ]+b o ) (Output Gate)

[0493] h t =o t tanh(Ct ) (Hidden state)

[0494] Among them, f t, i t , o t are the activation values ​​of the forget gate, input gate, and output gate, respectively, and C t is the cell state, h t is the hidden state, x t

[0495] is the current input.

[0496] Bi-LSTM (bidirectional LSTM) is an extension of LSTM, allowing the network to process both forward and reverse time series information. In sentiment analysis, speech sentiment often has contextual dependencies, so Bi-LSTM can improve the accuracy of sentiment classification.

[0497] CNN and Transformer models:

[0498] CNN is mainly used to extract local information from features such as spectrograms. It can effectively capture local changes in emotional features through convolution operations. The convolution operation formula of CNN is as follows:

[0499]

[0500] Among them, x is the input feature, w is the convolution kernel, and * represents the convolution operation.

[0501] Transformer is based on the self-attention mechanism, which can handle long-term dependencies and capture complex emotional patterns. The self-attention mechanism enhances the performance of the model by calculating the correlation between each position in the input sequence:

[0502]

[0503] Among them, Q, K, and V are query, key, and value matrices respectively, and d k is the dimension of the key vector, and the softmax function is used to normalize the correlation.

[0504] 4. Model Training

[0505] The loss function used in model training is the cross entropy loss function, which is used to evaluate the difference between the predicted results and the actual labels in the classification task:

[0506]

[0507] Where C is the number of sentiment categories, y i is the true label, p iis the class probability predicted by the model.

[0508] The optimization algorithm uses the Adam optimizer, and its update formula is as follows:

[0509]

[0510] Among them, η is the learning rate, m t and v t are the mean and variance of the gradient, respectively, and ∈ is a constant to prevent division by zero errors.

[0511] Summarize

[0512] The process of voice emotion analysis starts with data preparation and preprocessing. By selecting a variety of speech emotion datasets (RAVDESS, TESS, Emo-DB), the generalization ability of the model is ensured, and the audio data is cleaned and standardized, including denoising (using spectral subtraction or Wiener filtering), frame windowing (such as 20ms frame length, 50% overlapping Hamming window), and normalization of each frame feature to eliminate amplitude differences. In the feature extraction stage, three core acoustic features are selected: zero crossing rate, pitch, and short-term energy, which reflect signal frequency changes, pitch and speech intensity, respectively. Among them, high zero crossing rate and high pitch usually indicate happy or excited emotions, while low zero crossing rate and low pitch are related to sad or tired emotions. These features are combined into a multidimensional feature vector and input into the model. The model uses Bi-LSTM (bidirectional long short-term memory network) to capture the temporal dependency of speech, and combines CNN to extract local spectral features, and uses the self-attention mechanism of Transformer to handle long-term dependencies and complex emotional patterns. During the training process, the cross entropy loss function is used to optimize the accuracy of model predictions, and the Adam optimizer dynamically adjusts the learning rate to accelerate convergence. After the training is completed, the model outputs the emotion category and its confidence based on the input speech features. For example, for the input speech "I have a really happy day today!", the predicted emotion category may be "happy" with a confidence of 0.96. Finally, the classification ability and practical application effect of the model are verified by evaluating its performance.

[0513] Step 4: Get facial information through the camera and analyze the emotion results through yolov8 facial expression

[0514] 1. ROS2 camera data collection

[0515] In ROS2, camera nodes transmit real-time image data by publishing image messages. Image data is published to the / camera / color / image_raw topic in messages of type sensor_msgs / Image.

[0516] Image message structure:

[0517] Image data is published in RGB format, containing color information for each pixel (8-bit per channel).

[0518] Images contain metadata, such as image width, height, encoding format, timestamp, etc.

[0519] 2. YOLOv8 facial expression detection

[0520] YOLOv8 is a target detection model based on convolutional neural network (CNN) that can detect targets in images in real time. The model divides the image into multiple grid cells, predicts the bounding box and category probability within each grid, and then locates the target. In facial expression analysis, YOLOv8 is used for face detection and extracts facial regions for further expression analysis.

[0521] YOLOv8 working principle:

[0522] Input Image: YOLOv8 takes in fixed size images (640x640).

[0523] Convolution operation: Extract image features through multiple convolution layers, gradually reduce the spatial resolution and increase the number of channels to extract high-level features.

[0524] Prediction process:

[0525] YOLOv8 divides the image into S×S grids, each of which is responsible for predicting a bounding box, including coordinates, width, height, confidence, and category probability.

[0526] The coordinates and category probabilities of each grid are predicted, and the sigmoid activation function is used for classification and regression.

[0527] The face detection model identifies the face region in the image and extracts the region for subsequent processing.

[0528] 3. Facial Expression Recognition

[0529] The goal of facial expression recognition is to infer emotional states, such as "happy", "sad", and "angry", from detected face images. Facial expression recognition can be achieved through classical feature extraction methods such as local binary patterns (LBP) and convolutional neural network (CNN) methods based on deep learning.

[0530] Local Binary Pattern (LBP)

[0531] Local Binary Pattern (LBP) is a commonly used facial expression feature extraction method. It generates features that describe the local texture of an image by comparing the grayscale value relationship between the central pixel and the surrounding pixels.

[0532] Steps: Compare each pixel with its neighborhood. If the neighborhood pixel value is greater than the central pixel value, it is marked as 1, otherwise it is marked as 0.

[0533] These binary values ​​are concatenated into a binary number to form the LBP feature of the pixel.

[0534] By counting the LBP values ​​of all pixels in the image, the LBP histogram of the entire face image is generated as input for subsequent classification.

[0535] LBP formula:

[0536]

[0537] Among them, I i is the neighborhood pixel, I c is the center pixel, and s(x) is the sign function:

[0538]

[0539] P is the number of neighborhood pixels.

[0540] Deep Learning Methods (CNN)

[0541] Deep learning methods automatically learn the complex features of facial expressions through convolutional neural networks (CNNs). The CNN model extracts the spatial features of the image through multiple convolutional layers and pooling layers, and performs classification through fully connected layers.

[0542] Training process: Use a dataset containing labeled facial expression images for training and use the cross-entropy loss function for optimization.

[0543] Emotion Classification

[0544] Emotion classification is to map the extracted facial features to specific emotion categories. Common emotion categories include "happy", "sad", "angry", etc. Classification uses machine learning methods such as support vector machine (SVM), decision tree, neural network, etc.

[0545] Emotion classification formula: The classification model function f learned through the training set, given the input feature x (LBP histogram and CNN feature vector), the model outputs the predicted emotion label y: y = f(x)

[0546] 4. Combination of YOLOv8 and facial expression analysis

[0547] 5. Use YOLOv8 for face detection, detect the face area in the image, and extract the face frame.

[0548] 6. Crop the face area from the image and convert it to a grayscale image in preparation for expression analysis.

[0549] 7. Use LBP or CNN to extract features from the face area to obtain facial expression features.

[0550] 8. Input the extracted features into the emotion classification model to predict the emotion label of the facial expression.

[0551] 5. Data Flow and Workflow

[0552] 5. Camera node obtains data: ROS2 camera node publishes image data through the / camera / color / image_raw topic.

[0553] 6.YOLOv8 Face Detection: The YOLOv8 model detects faces in the image and returns the bounding box coordinates.

[0554] 7. Expression analysis: extract the facial area and grayscale it, and apply LBP or CNN method to extract expression features.

[0555] 8. Emotion classification: Input facial features into the emotion classification model, predict the emotion label and output the result.

[0556] Step 5: Segment the sentences that the large model responds to

[0557] Text parsing: Parse the AI ​​output using regular expressions to separate the sentiment analysis part (such as “Neutral (confidence: 88%)”) from the generated dialogue part (such as “Sounds like a relaxing day, how do you plan your day?”).

[0558] Emotional data processing: The confidence of the emotional part is extracted as a floating point value (such as 0.88) and combined with the corresponding emotional label (such as neutral).

[0559] ROS2 message definition and publishing:

[0560] For the conversation content, publish it to the chat_yuan topic of ROS2.

[0561] For sentiment confidence and labels, publish them to the wenben_em topic in the specified format (such as 0.88 neutral).

[0562] Step 6: Integrate the results of text, visual, and audio sentiment analysis to get the final sentiment

[0563] 1. Modeling of the credibility of each modality

[0564] In the multimodal fusion process of sentiment analysis, the design of credibility is crucial. The output of each modality (text, vision, sound) is usually a probability distribution, that is, the predicted probability of each sentiment category. We can use the confidence of this probability distribution to weight the sentiment prediction results of different modalities.

[0565] Representation of modal output

[0566] The output of each modality is a probability distribution containing sentiment labels and their corresponding confidences, usually written as:

[0567]

[0568] Among them, y i represents the emotion category, p i Represents the confidence of the modality in this category. For example, for text analysis, suppose the model outputs the sentiment label "happy" and its confidence is 0.92, and the probability of other labels is lower.

[0569] Modeling and Normalization of Modal Assurance Criteria

[0570] The output confidence of each modality pip_ipi can be further modeled and normalized through a series of functions, making the final multimodal fusion process more reasonable. Commonly used normalization methods include Softmax normalization

[0571]

[0572] The Softmax function uses exponential mapping and normalization operations to make the output result present a probability distribution. In this way, the prediction results of each modality can be standardized into probability values.

[0573] Credibility Calibration

[0574] In order to further improve the accuracy of multimodal fusion, we can introduce the "Confidence Calibration" technology. Usually, during the training process of deep learning models, the output probability of the model does not necessarily correspond to the actual probability value. Therefore, the following method can be used to calibrate the confidence of each modality:

[0575] Temperature Scaling

[0576]

[0577] Where T is the temperature hyperparameter, which controls the smoothness of the output probability. A lower temperature will make the output of the model sharper (higher confidence), while a higher temperature will smooth the output, making the probabilities of each category closer.

[0578] Through the above normalization and calibration process, the confidence of each modality can be standardized and made more consistent with the actual probability distribution, thereby providing reliable input for the subsequent fusion step.

[0579] 2. Theoretical Framework of Modal Fusion

[0580] Weighted fusion and probability distribution fusion

[0581] Weighted fusion is one of the most commonly used fusion methods for multimodal sentiment analysis. The core idea of ​​weighted fusion is to obtain the final sentiment classification result by weighted summing the output confidence of each modality. Assume that the output of the three modalities of text, vision and sound is and Respectively expressed as the emotion category and corresponding probability distribution of each modality, we define the weighted fusion method as follows:

[0582]

[0583] Among them, w text , w visual , w audio is the weight of each modality, and these weights are determined through training data or cross-validation. In sentiment analysis, we not only need to obtain the sentiment label of each modality, but also need to consider its confidence value. We can fuse the output probabilities of different modalities in a weighted manner to obtain the final fusion probability value of each sentiment category.

[0584] Bayesian Networks and Decision Fusion

[0585] In a more accurate model, we can consider using Bayesian networks for multimodal sentiment analysis. Bayesian networks can infer the optimal sentiment labels and confidence values ​​in the context of multimodal data. Through the principle of Bayesian reasoning, we can derive the final sentiment label and its confidence by constructing a joint probability model that includes the dependencies between text, vision, and sound.

[0586] Assume that the sentiment prediction results for each modality are are independent of each other, we can use Bayes' theorem for modal fusion:

[0587]

[0588] Among them, P(y) is the prior probability of sentiment label y, is the conditional probability of the modality under the sentiment label y,

[0589] is the marginal probability of the modality. Through the Bayesian network model, the system can more accurately combine the information of different modalities and calculate the final sentiment category and the corresponding confidence.

[0590] 3. Further processing and optimization of confidence

[0591] In order to further optimize the confidence of the final sentiment classification, we can use the uncertainty quantification method of the model. Common uncertainty quantification methods include:

[0592] Monte Carlo Dropout: By enabling Dropout during inference, different network structures are simulated to calculate the uncertainty of the model output. This can provide confidence intervals for each sentiment category, rather than just a single confidence value. Gaussian Process: Gaussian Process can model the uncertainty of the model and help the system evaluate the credibility of sentiment analysis in different contexts.

[0593] Through these methods, the robustness of the sentiment analysis system can be further improved, especially when facing noisy data or complex emotional states, and more reliable sentiment recognition results can be provided.

[0594] Running the example: Workflow for multimodal sentiment analysis

[0595] Input Data

[0596] The system receives data from three modalities:

[0597] 11. Text input: The user said, "I am very happy today!"

[0598] 12. Visual input: The user’s facial expression is recognized as a smile,

[0599] 13. Voice input: Voice waveform shows the emotion is "happy"

[0600] Unimodal Sentiment Analysis

[0601] 14. Text modal analysis: The natural language processing model analyzes the text and predicts the sentiment as "happy" with a confidence level of 0.88

[0602] 15. Visual modal analysis: A computer vision model based on facial expressions predicts the emotion as "happy" with a confidence level of 0.85

[0603] 16. Voice modal analysis: Through voice emotion recognition, the predicted emotion is "happy" with a confidence level of 0.92

[0604] Modal Assurance Normalization and Calibration

[0605] Use temperature scaling to adjust the confidence:

[0606] 1. The text modality confidence is adjusted to 0.90

[0607] 2. Visual modality confidence is adjusted to 0.87

[0608] 3. The sound modal confidence is adjusted to 0.94

[0609] Weighted Fusion

[0610] The modal weights are determined based on historical data as αt=0.4, αv=0.3, and αa=0.3.

[0611] Fusion formula:

[0612] P fusion =α t P t +α v P v +α a P a

[0613] The final fusion confidence is calculated as:

[0614] P fusion =0.4×0.90+0.3×0.87+0.3×0.94=0.903

[0615] Final Output

[0616] The system judges the user's comprehensive emotion as "happy" with a confidence level of 90.3%.

[0617] 1. Output emotion category: happy

[0618] 2. Output confidence: 90.3%

[0619] Practical Application

[0620] The system pushes the output results to the corresponding ROS2 topic:

[0621] 1. Emotion classification result: "happy" is pushed to topic / emotion_result.

[0622] 2. Fusion confidence: 0.903 is pushed to topic / confidence_score.

[0623] Step 7: Voice Emotion Cloning Module-Topic 1: Reading Data from ROS2 Topic

[0624] 1. System Architecture Overview

[0625] In the voice emotion cloning module of this project, the system interacts with data through the ROS2 (Robot Operating System 2) framework. ROS 2 uses topics to transfer information between different nodes. The core task of this module is to read data from two ROS2 topics, one for receiving the text data to be cloned (chat_yuan) and the other for obtaining the user's current emotional state (emotion_result).

[0626] 2. Read text from the chat_yuan topic

[0627] The chat_yuan topic transmits the text data entered by the user. This text data will serve as the basic information for speech synthesis. Through the subscription mechanism of the ROS2 node, the system will obtain the text data on the topic in real time and prepare for subsequent speech cloning processing. In terms of implementation, the system uses the rclpy Python client library of ROS2 to subscribe to the topic. Whenever a topic publishes new data, the subscribed node will receive the data and pass it to the speech cloning model.

[0628] 3. Read the emotional state from the emotion_result topic

[0629] The emotion_result topic contains the user's current emotional state, such as happiness, anger, sadness, etc. The emotion analysis module obtains the user's emotional characteristics through this topic and maps them to the tone parameters of the audio output. Emotional data is generated by real-time monitoring of the user's physiological or psychological state and transmitted to the system in real time through ROS2. Based on this data, the system adjusts the voice features such as tone and intonation in the voice cloning model to make the generated voice match the user's current emotional state.

[0630] 4. Data processing and conversion

[0631] After reading the text and emotional data, the system will process the data and pass the text data to the speech synthesis module, while the emotional data will be combined with the emotional adjustment mechanism in the speech cloning process. Specifically, the emotional data will affect the pitch, speed, emotional color and other characteristics of speech synthesis to achieve the goal of emotional cloning.

[0632] 5. Summary

[0633] This topic lays the foundation for data input of the voice emotion cloning module by describing in detail how to read the text and emotional state data to be cloned from the ROS2 topic. Through the efficient topic communication mechanism of ROS2, the system can obtain and process the text and emotional information input by the user in real time, providing accurate input data for subsequent voice cloning and emotional adjustment.

[0634] Topic 2: Emotion-driven dialogue system - Correspondence table between emotion and tone

[0635] In order to implement an emotional dialogue system, we need to adjust the tone according to the user's emotional state, so that the dialogue is both humane and meets emotional needs. The following is a table showing how the system should adjust the tone in different emotional states to achieve a voice interaction effect that is closer to the user's emotions.

[0636] The tone of the user's emotional state system reply is saved in the tone description of the topic emotion_kelong

[0637]

[0638] explain:

[0639] 1. Sadness: When a user expresses sadness, the system’s tone should be gentle and sincere, showing care and comfort, helping the user feel understood and comforted.

[0640] 2. Surprise: When users are surprised or happy, the tone should be more excited and positive, expressing pleasant and relaxed emotions to enhance the affinity and encouragement of the interaction.

[0641] 3. Fear: When facing the user’s fear, your tone should be calm and soothing, and you should give encouragement and support to make the user feel safe and comfortable.

[0642] 4. Anger: When dealing with angry users, your tone should be calm and rational, while showing understanding and respect, avoiding intensifying emotions and helping to ease conflicts.

[0643] 5. Sadness: When the user is in a sad state, the system should show sympathy and care, and the tone should be comforting to help the user feel warm and supported.

[0644] 6. Neutral: For the neutral emotional state, the system's tone should remain normal and peaceful, neither biased towards joy, anger, sorrow, or happiness, nor indifferent, suitable for daily conversations.

[0645] In this system, the emotion recognition module analyzes the user's emotional state (obtained from the topic emotion_result), then converts these emotional features into corresponding tone descriptions through the emotion conversion module, and finally publishes the adjusted tone features through the topic emotion_kelong. This ensures that the voice output is highly consistent with the user's emotional state, thereby enhancing the emotional resonance of human-computer interaction and improving the user experience. The core innovation of this module is to dynamically adjust the tone of voice replies based on the user's emotional state, ensuring that the conversation content matches the emotional state, and providing a more natural and humanized interactive experience.

[0646] Topic 3: Voice Emotion Cloning Module-Implementation Principle and Process

[0647] The core goal of the voice emotion cloning module is to generate emotional speech output through deep learning methods based on the user input text (chat_yuan) and the user's emotional state (emotion_result). The implementation process of this module includes voice cloning, emotion adjustment, and modification of speech features through reference audio. The following is the implementation principle, process and technical details of this module.

[0648] 1. Basic process of sound cloning

[0649] The implementation of the sound cloning module mainly includes the following steps:

[0650] Text input processing: The system obtains the text data to be cloned from the chat_yuan topic. These texts need to be converted into voice features that can generate sound, usually voice signals generated by speech synthesis technology.

[0651] Feature extraction: In the process of voice cloning, we first need to extract features from the original voice sample. Commonly used features include:

[0652] o Speech Spectrum: Converts an audio signal into a spectrogram, characterizing the frequency components of the sound and how they vary over time.

[0653] o MFCC (Mel-Frequency Cepstral Coefficients): A commonly used speech feature that can effectively capture the timbre information of speech.

[0654] These features act as the "fingerprint" of the sound, providing the necessary information for subsequent models to generate speech.

[0655] Model training: Use deep learning models to train voice cloning. Commonly used models include:

[0656] oGenerative Adversarial Networks (GANs): Generative Adversarial Networks consist of a generator and a discriminator. Through adversarial training, the generator learns how to generate sound samples that are as realistic as possible, while the discriminator determines whether the generated sound is realistic. GANs are a powerful tool for voice cloning and can generate high-quality, natural speech.

[0657] o Variational Autoencoders (VAEs): Variational autoencoders are generative models that can learn the latent distribution of input features and generate new speech features through sampling.

[0658] During training, the model requires a large amount of speech data to learn how to generate natural, realistic sounds from speech features.

[0659] Voice generation: Once the model is trained, new speech features can be input, and the model will generate new sound samples based on the knowledge learned during training. The generated voice should sound like the target voice, that is, have similar timbre, intonation, speed, etc.

[0660] 2. Emotional Adjustment and Emotional Cloning

[0661] Although the cloned voice imitates the target voice in terms of timbre and intonation, it lacks emotional expression. In order to achieve emotional voice output, the system needs to adjust the cloned voice through reference audio to make it consistent with the user's current emotional state. The specific process is as follows:

[0662] Emotional feature extraction: First, obtain the user's emotional state (such as sad, angry, happy, etc.) through the emotion_result topic. Each emotion has its own specific voice features, such as speaking speed, pitch, volume, etc. According to these emotional features, select the corresponding reference audio.

[0663] Select reference audio: For each emotional state, select reference audio that matches the current emotion from the pre-prepared emotional audio library. Each reference audio contains the speech characteristics of a specific emotion. For example, the audio of sad emotion may have a low tone and a slow speech speed, while the audio of angry emotion may have a fast speech speed and a strong tone.

[0664] Emotion adjustment: Use the voice conversion model to adjust the cloned voice’s emotion. The model adjusts the cloned voice’s pitch, speech rate, and emotional color based on the emotional characteristics of the reference audio, making the generated voice sound consistent with the user’s emotional state.

[0665] The technical implementation of this process generally adopts:

[0666] o Transfer learning of acoustic features: Adjust the timbre and emotional performance by transferring the emotional features in the reference audio to the cloned voice.

[0667] oEmotional Intonation Generation: Modify the intonation and emotion of the cloned voice to generate the final speech output with emotion.

[0668] 3. Technical details and model implementation

[0669] Generative Adversarial Networks (GANs): The core of this method is the adversarial training of the generator and the discriminator. The generator generates more and more realistic sounds through continuous optimization, while the discriminator forces the generator to learn more detailed speech details by judging whether the generated audio is natural and realistic. In implementation, the generator network and the discriminator network are usually based on **convolutional neural networks (CNN) and recurrent neural networks (RNN)** to process time series data and high-dimensional features.

[0670] Variational Autoencoders (VAEs): In voice cloning tasks, VAEs are used to capture the latent representation of speech data. By learning the latent distribution of input speech features, VAEs can generate speech features that are similar but not identical to the input data, thereby achieving high-quality voice synthesis.

[0671] Emotion Regulation Network: In order to achieve emotional speech output, the emotion regulation model will fuse these features with the cloned voice according to the characteristics of different emotions (such as pitch, speaking speed, intensity, etc.) to adjust the emotional expression of the speech.

[0672] In summary, the present invention realizes voice cloning and emotion adjustment through deep learning models (GANs, VAEs). Voice cloning first learns the speech features of the target voice through feature extraction, and generates natural and realistic voices through training. Subsequently, the emotion adjustment module uses reference audio to adjust the voice according to the user's emotional state, so that the output voice not only has the timbre of the target person, but also reflects the user's emotional state. The entire process generates high-quality emotional voice through the training and optimization of deep neural networks, providing a more natural and personalized voice interaction experience.

[0673] Based on the above implementations, we can have emotional AI conversations with users.

Claims

1. A method for implementing voiceprint recognition, identity confirmation and dialogue for sentiment analysis, characterized in that the method comprises the following steps: Step 1: Get user audio information 1. Voiceprint recognition and identity confirmation First, voiceprint recognition technology is used to confirm whether the source of the recorded audio is the voice of a specific person. Voiceprint recognition usually authenticates the identity by extracting audio features (such as MFCC, Mel spectrum, etc.) and comparing them with pre-stored templates. Mathematical formula: Assume that the input audio feature is X = {x1, x2, ..., x n}, the voiceprint template is T = {t1, t2, ..., t m}Then the recognition process can be expressed as: in, Indicates the recognized user tag||x j -t i || 2 represents the Euclidean distance between feature vectors, 2.Silence detection and recording control In order to determine when to stop recording based on the length of silence, a silence threshold needs to be set and the silence state is determined by energy detection of the audio signal. Silence detection formula: Assume that the instantaneous energy of the audio signal is where x k (t) is the sampling point of the audio signal. By setting a silence threshold ∈, when E(t) is less than ∈ and the duration exceeds the set silence threshold δ, it is considered to enter the silence state.

3. Recording end determination and release Combining the results of voiceprint recognition and silence detection, when the voice of a specific person is recognized and the continuous silence exceeds the set threshold, the system can automatically end the recording and send the audio data to the specified topic. Silence determination and recording end formula: If the silence signal E(t) <∈ is detected for k consecutive times and the silence duration exceeds the threshold δ, the recording ends and the release mechanism is triggered:

4. Control of recording and publishing process In the system, if the recording stops, the system will automatically publish the audio data. This process can be achieved through the publish-subscribe mechanism in ROS2. A buffer can be set, and when the number of audio files reaches a predetermined value, it will be published. Release condition formula: Assume that the audio buffer is B = {b1, b2, ..., b n}, when n reaches the set number of voice segments, the cached data is published to the topic: Publish={1,if n≥N max Where N max The maximum number of caches set triggers data publishing. Step 2: Use whisper to convert speech to text and get text results 1. Audio input and preprocessing First, the audio input needs to be preprocessed and converted into a format suitable for processing by the Whisper model. Audio data is usually stored in the original PCM format and needs to be converted into a feature representation that the Whisper model can recognize. Audio preprocessing steps: Unify the sampling rate of audio files (Whisper's default sampling rate is 16kHz), Convert the audio signal into a Mel Spectrogram, which is a common method for representing audio features. Set the audio signal x(t), where t is the time point, and first perform windowing and fast Fourier transform (FFT) processing: X(k)=FFT(Window(x(t))) Then, the FFT result is mapped to the Mel frequency axis to obtain the Mel spectrum M(t,f): M(t,f)=Mel(X(k)) 2. Whisper model for speech-to-text conversion After the audio is preprocessed, the Whisper model is used to convert speech to text. Whisper is an end-to-end speech recognition model whose input is a Mel-spectrogram and output is text. Set the input feature to x mel (Mel spectrum), the Whisper model uses a deep neural network (Transformer) to process and decode text from audio features. The model reasoning process is as follows: in is the output of the Whisper model, representing the transcribed text, 3. Output and result publication Final recognition result Publish to other modules in the system through ROS2 topics. Assuming that the text result is published through the topic result_say, the result publishing process is as follows: Step 3: The text is transmitted to the trained large language model in the server through the API, and the large model response is obtained Step 4: Voice emotion analysis 1. Data preparation and preprocessing Dataset construction and selection: To ensure the generalization ability of the model, first select a diverse speech emotion dataset, such as the RAVDESS dataset, TESS dataset, or Emo-DB, which contains multiple emotion categories (such as happiness, anger, sadness, surprise, etc.) and covers speech samples from multiple speakers. Each audio file in the dataset contains a labeled emotion category for the training of supervised learning models, Data preprocessing: Denoising and augmentation: Use voice activity detection (VAD) to remove silent segments and use noise suppression techniques (such as spectral subtraction or Wiener filtering) to remove background noise. Framing and windowing: The speech signal is divided into several short time frames (for example, each frame is 20ms long and the overlap between frames is 50%), and each frame is windowed (such as Hamming window) to ensure time-frequency locality. Normalization: Normalize the features of each audio clip to remove the amplitude differences between different audio samples. The normalization formula is as follows: Among them, μ is the mean, σ is the standard deviation, x is the original feature, is the standardized feature, 2. Feature extraction In speech emotion analysis, the time domain and frequency domain features of audio signals can effectively describe the emotional state. We extract the following acoustic features as model input: Mel Frequency Cepstral Coefficient (MFCC): MFCC is one of the most commonly used features in speech analysis. It converts speech signals into frequency domain signals and then performs Mel filtering to extract the spectrum information of audio signals. The calculation formula of MFCC is as follows: MFCC = DCT(log(Mel(STFT(x)))) Among them, STFT is short-time Fourier transform, Mel represents Mel-scale filter bank, DCT is discrete cosine transform, Zero Crossing Rate (ZCR): The zero crossing rate reflects the frequency of signal changes and can effectively distinguish speech with different emotions. For example, angry speech usually has a higher zero crossing rate. High zero crossing rate: usually appears in speech with high-frequency components, such as excitement, happiness and other emotions. At this time, the frequency of the speech signal changes quickly, resulting in an increase in the number of times it passes through the zero point. Low zero crossing rate: appears in low-frequency speech or speech with smoother emotions, such as sadness, fatigue, etc. in, is the indicator function, x i is the i-th speech signal sample, Pitch: Pitch reflects the fundamental frequency of the audio signal, usually expressed in Hertz (Hz). The fundamental frequency can be estimated by autocorrelation method, harmonic analysis and other methods. High pitch: Emotions such as happiness and surprise are usually accompanied by higher pitches because the speaker will raise his voice to express emotions. Low pitch: Emotions such as sadness, fatigue or fear can cause the fundamental frequency to decrease, making the voice sound more stable or low. Rate of change of pitch: In excitement or surprise, the pitch may fluctuate greatly over time; in neutral or sad emotions, the pitch changes more slowly. Short-time energy: Short-time energy is the sum of the energy of the audio signal in a unit time window, which is used to measure the strength of the speech signal. Where x(n) is the speech signal, w(n) is the time window function, N is the frame length, high energy: emotions such as excitement and anger are usually accompanied by greater speech intensity, which is manifested as high short-term energy, low energy: emotions such as sadness and fatigue correspond to lower speech intensity and lower short-term energy. Energy fluctuations: Voice signals expressing intense emotions (such as anger or excitement) have large short-term energy fluctuations. Combining these three for a comprehensive analysis Combining features such as zero crossing rate, pitch, and short-time energy can form a multidimensional feature vector for emotion classification. Suppose the feature vector of a speech signal is: f=[ZCR,Pitch,STE] Happy emotions may be manifested as [high ZCR, high Pitch, high STE] Sadness may be manifested as [low ZCR, low Pitch, low STE] Anger may manifest as [moderate or low ZCR, high Pitch, high STE] 3. Model design LSTM and Bi-LSTM models: LSTM (Long Short-Term Memory) is a neural network model that is particularly suitable for processing time series data. Through its internal "gate mechanism" (input gate, forget gate, and output gate), LSTM can remember and forget important historical information in speech signals, and is particularly suitable for capturing emotional features in speech. The mathematical formula of LSTM is as follows: f t =σ(W f ·[h t-1 ,x t ]+b f )(Forget Gate) i t =σ(W i ·[h t-1 ,x t ]+b i )(Input gate) o t =σ(W o ·[h t-1 x t ]+b o )(Output gate) h t =o t tanh(C t )(Hidden state) Among them, f t ,i t , o t are the activation values ​​of the forget gate, input gate, and output gate, respectively, and C t is the cell state, h t is the hidden state, x t is the current input, Bi-LSTM (bidirectional LSTM) is an extension of LSTM, allowing the network to process both forward and reverse time series information. In sentiment analysis, speech sentiment often has contextual dependencies, so Bi-LSTM can improve the accuracy of sentiment classification. CNN and Transformer models: CNN is mainly used to extract local information from features such as spectrograms. It can effectively capture local changes in emotional features through convolution operations. The convolution operation formula of CNN is as follows: Among them, x is the input feature, w is the convolution kernel, and * represents the convolution operation. Transformer is based on the self-attention mechanism, which can handle long-term dependencies and capture complex emotional patterns. The self-attention mechanism enhances the performance of the model by calculating the correlation between each position in the input sequence: Among them, Q, K, and V are query, key, and value matrices respectively, and d k is the dimension of the key vector, and the softmax function is used to normalize the correlation.

4. Model Training The loss function used in model training is the cross entropy loss function, which is used to evaluate the difference between the predicted results and the actual labels in the classification task: Where C is the number of sentiment categories, y i is the true label, p i is the class probability predicted by the model, The optimization algorithm uses the Adam optimizer, and its update formula is as follows: Among them, η is the learning rate, m t and v t are the mean and variance of the gradient, ∈ is a constant to prevent division by zero errors, Summarize The process of voice emotion analysis starts with data preparation and preprocessing. By selecting a variety of speech emotion datasets (RAVDESS, TESS, Emo-DB), the generalization ability of the model is ensured, and the audio data is cleaned and standardized, including denoising (using spectral subtraction or Wiener filtering), frame windowing (such as 20ms frame length, 50% overlapping Hamming window), and normalization of each frame feature to eliminate amplitude differences. In the feature extraction stage, three core acoustic features are selected: zero crossing rate, pitch, and short-time energy, which reflect signal frequency changes, pitch, and speech intensity, respectively. A high zero crossing rate and high pitch usually indicate happiness or excitement, while a low zero crossing rate and low pitch are associated with sadness or fatigue. The model uses Bi-LSTM (bidirectional long short-term memory network) to capture the temporal dependency of speech, and combines CNN to extract local spectral features. It also uses the self-attention mechanism of Transformer to process long-term dependencies and complex emotional patterns. During the training process, the cross entropy loss function is used to optimize the accuracy of model prediction. The Adam optimizer dynamically adjusts the learning rate to accelerate convergence. After the training is completed, the model outputs the emotion category and its confidence level based on the input speech features. For example, for the input speech "I have a really happy day today!", the predicted emotion category may be "happy" with a confidence level of 0.

96. Finally, the model performance is evaluated to verify its classification ability and practical application effect. Step 5: Get facial information through the camera and analyze the emotion results through yolov8 facial expression 1. ROS2 camera data collection In ROS2, the camera node transmits real-time image data by publishing image messages. The image data is published to the / camera / color / image_raw topic in the form of messages of the sensor_msgs / Image type. Image message structure: Image data is published in RGB format, including color information for each pixel (8-bit per channel). Images contain metadata, such as image width, height, encoding format, timestamp, etc.

2. YOLOv8 facial expression detection YOLOv8 is a target detection model based on convolutional neural network (CNN). It can detect targets in images in real time. The model divides the image into multiple grid cells, predicts the bounding box and category probability within each grid, and then locates the target. In facial expression analysis, YOLOv8 is used for face detection and extracts facial areas for further expression analysis. The working principle of YOLOv8 is as follows: Input image: YOLOv8 receives fixed size images (640x640), Convolution operation: extract image features through multiple convolution layers, gradually reduce the spatial resolution and increase the number of channels, extract high-level features, Prediction process: YOLOv8 divides the image into S×S grids, each of which is responsible for predicting a bounding box, including coordinates, width, height, confidence, and category probability. Predict the coordinates and category probabilities of each grid, and use the sigmoid activation function for classification and regression. The face detection model identifies the face area in the image and extracts the area for subsequent processing.

3. Facial Expression Recognition The goal of facial expression recognition is to infer emotional states, such as "happy", "sad", and "angry", based on detected face images. Facial expression recognition can be achieved through classical feature extraction methods (such as local binary patterns (LBP)) and convolutional neural network (CNN) methods based on deep learning. Local Binary Pattern (LBP) Local Binary Pattern (LBP) is a commonly used facial expression feature extraction method. It generates features describing the local texture of an image by comparing the grayscale value relationship between the central pixel and the surrounding pixels. Each pixel is compared with its neighbors. If the neighboring pixel value is greater than the central pixel value, it is marked as 1, otherwise it is marked as 0. These binary values ​​are concatenated into a binary number to form the LBP feature of the pixel. By counting the LBP values ​​of all pixels in the image, the LBP histogram of the entire face image is generated as the input for subsequent classification. LBP formula: Among them, I i is the neighborhood pixel, I c is the center pixel, and s(x) is the sign function: P is the number of neighborhood pixels, Deep Learning Methods (CNN) The deep learning method automatically learns the complex features of facial expressions through convolutional neural networks (CNN). The CNN model extracts the spatial features of the image through multiple convolutional layers and pooling layers, and classifies it through fully connected layers. Training process: Use a dataset containing labeled facial expression images for training and use the cross-entropy loss function for optimization. Emotion Classification Emotion classification is to map the extracted facial features to specific emotion categories. Common emotion categories include "happy", "sad", "angry", etc. Classification uses machine learning methods such as support vector machine (SVM), decision tree, neural network, etc. Emotion classification formula: The classification model function f learned through the training set, given the input feature x (LBP histogram and CNN feature vector), the model outputs the predicted emotion label y: y = f(x) 4. Combination of YOLOv8 and facial expression analysis 1. Use YOLOv8 for face detection, detect the face area in the image, and extract the face frame.

2. Crop the face area from the image and convert it into a grayscale image to prepare for expression analysis.

3. Use LBP or CNN to extract features from the face area to obtain facial expression features.

4. Input the extracted features into the emotion classification model to predict the emotion label of the facial expression, 5. Data Flow and Workflow 1. Camera node obtains data: ROS2 camera node publishes image data through the / camera / color / image_raw topic, 2.YOLOv8 face detection: The YOLOv8 model detects faces in the image and returns the bounding box coordinates.

3. Expression analysis: extract the face area and grayscale it, and use LBP or CNN method to extract expression features.

4. Emotion classification: Input facial features into the emotion classification model, predict the emotion label and output the result. Step 5: Segment the sentences that the large model responds to Text parsing: parse the AI ​​output through regular expressions, separate the sentiment analysis part (such as "neutral (confidence: 88%)") and the generated dialogue part (such as "It sounds like a relaxing day, how do you plan your day?"), sentiment data processing: extract the confidence of the sentiment part as a floating point value (such as 0.88) and combine it with the corresponding sentiment label (such as neutral), ROS2 message definition and publishing: For the conversation content, publish it to the chat_yuan topic of ROS2. For sentiment confidence and labels, publish them to the wenben_em topic in the specified format (e.g. 0.88 neutral). Step 6: Integrate the results of text, visual, and audio sentiment analysis to get the final sentiment 1. Modeling of the credibility of each modality In the multimodal fusion process of sentiment analysis, the design of credibility is crucial. The output of each modality (text, vision, sound) is usually a probability distribution, that is, the prediction probability of each sentiment category. We can use the confidence of this probability distribution to weightedly fuse the sentiment prediction results of different modalities. Representation of modal output The output of each modality is a probability distribution containing sentiment labels and their corresponding confidences, usually written as: Among them, y i represents the emotion category, p i Represents the confidence of the modality in this category. For example, for text analysis, suppose the model outputs the sentiment label "happy" and its confidence is 0.92, and the probability of other labels is lower. Modeling and Normalization of Modal Assurance Criteria The output confidence of each modality pip_ipi can be further modeled and normalized through a series of functions to make the final multimodal fusion process more reasonable. Commonly used normalization methods include Softmax normalization The Softmax function uses exponential mapping and normalization operations to make the output results present a probability distribution. In this way, the prediction results of each modality can be standardized into probability values. Credibility Calibration In order to further improve the accuracy of multimodal fusion, we can introduce "Confidence Calibration" Usually, during the training process of deep learning models, the output probability of the model does not necessarily correspond to the actual probability value. Therefore, the following method can be used to calibrate the confidence of each modality: Temperature Scaling Among them, T is the temperature hyperparameter, which controls the smoothness of the output probability. A lower temperature will make the output of the model sharper (higher confidence), while a higher temperature will smooth the output, making the probabilities of each category closer. Through the above normalization and calibration process, the confidence of each modality can be standardized and made more consistent with the actual probability distribution, thereby providing reliable input for the subsequent fusion step.

2. Theoretical Framework of Modal Fusion Weighted fusion and probability distribution fusion Weighted fusion is one of the most commonly used fusion methods for multimodal sentiment analysis. The core idea of ​​weighted fusion is to obtain the final sentiment classification result by weighted summing the output confidence of each modality. Assume that the output of the three modalities of text, vision and sound is and Respectively expressed as the emotion category and corresponding probability distribution of each modality, we define the weighted fusion method as follows: Among them, w text , w visual , w audio are the weights of each modality, and these weights are determined through training data or cross-validation. In sentiment analysis, we not only need to obtain the sentiment label of each modality, but also need to consider its confidence value. We can fuse the output probabilities of different modalities in a weighted manner to obtain the final fusion probability value of each sentiment category. Bayesian Networks and Decision Fusion In a more accurate model, we can consider using Bayesian networks for multimodal sentiment analysis. Bayesian networks can infer the optimal sentiment labels and confidence values ​​in the context of multimodal data. Through the principle of Bayesian reasoning, we can derive the final sentiment labels and their confidence by constructing a joint probability model that includes the dependencies between text, vision, and sound. Assume that the sentiment prediction results for each modality are are independent of each other, we can use Bayes' theorem for modal fusion: Among them, P(y) is the prior probability of sentiment label y, is the conditional probability of the modality under the sentiment label y, is the marginal probability of the modality. Through the Bayesian network model, the system can more accurately combine the information of different modalities and calculate the final emotion category and the corresponding confidence.

3. Further processing and optimization of confidence In order to further optimize the confidence of the final sentiment classification, we can use the uncertainty quantification method of the model. Common uncertainty quantification methods include: Monte Carlo Dropout: By enabling Dropout during inference, different network structures are simulated to calculate the uncertainty of the model output, which can provide confidence intervals for each emotion category instead of just a single confidence value. Gaussian Process: Gaussian Process can model the uncertainty of the model and help the system evaluate the credibility of sentiment analysis in different situations. Through these methods, the robustness of the sentiment analysis system can be further improved, especially when facing noisy data or complex emotional states, and more reliable sentiment recognition results can be provided. Running the example: Workflow for multimodal sentiment analysis Input Data The system receives data from three modalities:

1. Text input: The user said, "I am very happy today!" 2. Visual input: The user’s facial expression is recognized as a smile, 3. Voice input: Voice waveform shows the emotion is "happy" Unimodal Sentiment Analysis 1. Text modal analysis: The natural language processing model analyzes the text and predicts the sentiment as "happy" with a confidence level of 0.88 2. Visual modality analysis: The computer vision model based on facial expressions predicts the emotion as "happy" with a confidence level of 0.85 3. Voice modal analysis: Through speech emotion recognition, the predicted emotion is "happy" with a confidence level of 0.

92. Modal confidence normalization and calibration Use temperature scaling to adjust the confidence:

4. Text modality confidence is adjusted to 0.90 5. Visual modality confidence is adjusted to 0.87 6. Sound modal confidence is adjusted to 0.94 Weighted Fusion The modal weights are determined based on historical data as αt=0.4, αv=0.3, αa=0.3, Fusion formula: P fusion =α t P t +α v P v +α a P a The final fusion confidence is calculated as: P fusion 0.4×0.90+0.3×0.87+0.3×0.94=0.903 Final Output The system judges the user's comprehensive emotion as "happy" with a confidence level of 90.3%.

7. Output emotion category: happy 8. Output confidence: 90.3% Practical Application The system pushes the output results to the corresponding ROS2 topic:

9. Emotion classification result: "happy" is pushed to topic / emotion_result, 10. Fusion confidence: 0.903 is pushed to topic / confidence_score, Step 7: Voice Emotion Cloning Module-Topic 1: Reading Data from ROS2 Topic 1. System Architecture Overview In the voice emotion cloning module of this project, the system interacts with data through the ROS2 (Robot Operating System 2) framework. ROS2 uses topics to realize information transmission between different nodes. The core task of this module is to read data from two ROS2 topics, one for receiving the text data to be cloned (chat_yuan) and the other for obtaining the user's current emotional state (emotion_result).

2. Read text from the chat_yuan topic The chat_yuan topic transmits the text data entered by the user, which will be used as the basic information for speech synthesis. Through the subscription mechanism of the ROS2 node, the system will obtain the text data on the topic in real time and prepare for subsequent speech cloning processing. In terms of implementation, the system uses the rclpy Python client library of ROS2 to subscribe to the topic. Whenever the topic publishes new data, the subscription node will receive the data and pass it to the speech cloning model.

3. Read the emotional state from the emotion_result topic The emotion_result topic contains the user's current emotional state, such as happiness, anger, sadness, etc. The emotion analysis module obtains the user's emotional characteristics through this topic and maps it to the tone parameters of the audio output. The emotional data is generated by real-time monitoring of the user's physiological or psychological state and transmitted to the system in real time through ROS2. Based on this data, the system adjusts the voice characteristics such as tone and intonation in the voice cloning model to make the generated voice conform to the user's current emotional state.

4. Data processing and conversion After reading the text and emotional data, the system will process the data and pass the text data to the speech synthesis module, while the emotional data will be combined with the emotional regulation mechanism in the speech cloning process. Specifically, the emotional data will affect the pitch, speed, emotional color and other characteristics of speech synthesis to achieve the goal of emotional cloning.

2. According to claim 1, a method for implementing voiceprint recognition, identity confirmation and dialogue for emotion analysis is characterized in that: The step 3 comprises: Specific steps 1 The client sends a request (user voice-to-text conversion result) 2 After the client obtains text input from the user, it sends the text input to the server through the API. In this step, the client collects the language-to-text input and transmits it to the server through an HTTP request. Operation process: Subscribe to the text string of the speech-to-text conversion on the topic, The client sends the input text to the server via an API request. Specific Step 2 The server receives the request (parses the request data) After receiving the client's request, the server first parses the request data and extracts the text input by the user. Then, it passes the text input to the large language model for processing and generates a corresponding reply. Operation process: The server receives the HTTP POST request and extracts the text entered by the user. The server passes the input to the large language model for inference processing. Specific Step 3 The principle and fine-tuning process of emotion recognition and response generation of a large language model on the server side 1. Basic principles of large language models The large language model of this project adopts a combination of pre-training and fine-tuning, and uses large-scale text data for training. The large language model is based on the Transformer architecture. Its core idea is to generate the next word through an autoregressive generative model, and capture semantic and contextual information in the process. Its basic training process is as follows: Pre-training stage: Use unsupervised learning methods to train large-scale corpus. The goal is to maximize the probability distribution of predicting the next word under a given context. Through unsupervised goals, the model learns the structure and statistical characteristics of the language. The objective function of pre-training is to maximize the conditional probability: Among them, g t is the target word at time step t, x1,...,x t-1 is the previous context word, T is the length of the text sequence, Fine-tuning stage: In the fine-tuning stage, we use annotated data from a specific field (such as sentiment analysis datasets, conversation datasets, etc.) to optimize model parameters through supervised learning. We also add the objective function of a specific task (sentiment analysis) in this stage to enable the model to perform better on specific tasks.

2. Sentiment Analysis Task The task of sentiment analysis aims to identify the sentiment categories contained in the text. Sentiments can usually be divided into multiple categories such as positive, negative, and neutral. In some situations, the expression of emotions is more detailed, such as "sad" and "happy". In the large language model, sentiment analysis is processed as a text classification task. The input text is passed into the model after preprocessing, and the model outputs the classification label corresponding to the sentiment (such as "sad", "happy"). Sentiment classification model: Sentiment analysis model is usually a multi-classification problem. By training a large amount of labeled data, the model learns to extract features from text and classify them. Assume that the sentiment label is C = {c1, c2, ..., c m }, the goal of the model is to calculate the probability that a given input text x belongs to each sentiment category: p(y=c k |x) where y∈C Through training data, the model learns the probability distribution of the input text x corresponding to the sentiment category y, and selects the most likely category as the prediction result.

3. Fine-tuning process 3.1 Fine-tuning objectives In order to enable the self-trained large language model to recognize the sentiment of the input text and generate emotional responses, the model needs to be fine-tuned so that it can complete the tasks of sentiment analysis and emotional response generation at the same time. The goals of fine-tuning include two aspects:

1. Sentiment classification task: perform sentiment analysis on the input text and classify the sentiment category (such as "sad").

2. Emotional response generation task: Generate a response text with corresponding emotions based on the input text and emotional labels. 3.2 Dataset Design For fine-tuning, the required training data should include the following:

1. Input text: contains the user's original input text (such as "I am sad"), 2. Emotion labeling: label each input text with an emotional category, such as "sad", "happy", 3. Target response: Generate a suitable response based on the input text and the emotional label. For example, for the emotional label "sad", the generated response may be "I know you are sad, @sad", The design of the dataset can be based on an existing sentiment annotation dataset, or training data can be constructed through manual annotation. Common sentiment classification labels include: Positive emotions: happiness, surprise 2 Negative emotions: fear, anger, sadness Neutral sentiment: Neutral 3.3 Loss Function for Fine-tuning During fine-tuning, the model needs to minimize the loss of two tasks: the loss of the sentiment classification task and the loss of the text generation task. Sentiment classification loss: The loss of sentiment classification usually uses cross-entropy loss. If there are C types of sentiment categories, the model predicts category c. k The probability of p(y=c k |x) The true label is y true Then the sentiment classification loss is: in Is the indicator function, if the true label is c i , then it is 1, otherwise it is 0. Text generation loss: The text generation loss uses the language model loss, usually to maximize the likelihood estimate (MLE) of the generated text. For the input x and the target text y, the generation loss is: Among them, y t is the target word at time step t, x is the input text, T is the length of the target text, Total loss function: By weighted merging of sentiment classification loss and text generation loss, the total loss is obtained: L total =α·L emotion +β·L gen Among them, α and β are hyperparameters used to control the relative weights between sentiment classification and text generation tasks. 3.4 Multi-task learning and training Through multi-task learning, the model will optimize both sentiment classification and text generation tasks during training. Specifically, the loss of sentiment analysis and generation tasks are jointly optimized. The model will simultaneously learn how to identify the sentiment in the text and generate corresponding responses based on the sentiment. During training, the model’s parameter updates take into account the losses of both tasks, thus improving both the ability to recognize and generate text. Ultimately, the model is able to generate natural language responses based on the input sentiment labels.

4. Emotion recognition and response generation reasoning process The fine-tuned self-trained large language model can simultaneously complete the tasks of sentiment analysis and emotional response generation during inference. The inference process includes the following steps: Emotion recognition: Given a user input text (such as "I am sad"), the model uses the sentiment classification module to determine the sentiment category of the text (such as "sad"). Emotional label generation: Based on the identified emotional labels, the model uses the emotional labels as contextual information when generating replies to ensure that the tone of the generated replies matches the sentiment of the input text. For example, when the emotional label is "sad", the generated reply may be: "I know you are feeling bad, @sad", Generate response: Based on the input text and sentiment label, the model generates response text, for example, "I know you are sad, @sad" or "I understand how you feel, @sad", Specific Step 4 The text reply after the server's large language model inference is returned to the client through the API response. After receiving the reply, the client parses and displays it to the user.

Citation Information

Cited By

  • Real-time interaction 3D digital holographic cabin method based on deep learning and sound cloning

    CN120318437A

  • Real-time interactive 3D digital holographic cabin method based on deep learning and sound cloning

    CN120318437B

  • Cross-skin color spectrum calibration method and system based on transfer learning, and storage medium

    CN120655937A

  • Sound stimulation adaptive regulation and control system based on abnormal state recognition

    CN120960583A

  • A sound stimulation adaptive regulation system based on abnormal state identification

    CN120960583B