Personalized voice interaction method and related equipment
By selecting a suitable speech recognition model through streaming processing and object feature analysis, the speech recognition of intelligent voice interaction assistants in different groups of people is solved, thus addressing the speech recognition problems in existing technologies, improving the accuracy of speech recognition and user satisfaction, and reducing the demand for computing resources.
Patent Information
- Application Number
- CN202511264011.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-12-02
AI Technical Summary
Existing intelligent voice interaction assistants struggle to accurately recognize audio questions from different groups of people, leading to incorrect answers and impacting user experience, especially in resource-constrained environments where the models are complex and computational resources are high.
The system receives audio streams of questions from users using streaming technology, segments and preprocesses them, performs object feature analysis to obtain service object tags, selects appropriate speech recognition models based on tag classification, determines the user's group by combining the user's audio, and uses personalized speech recognition models for text analysis to improve recognition accuracy.
It improves the accuracy of speech recognition and user satisfaction with responses, reduces the requirements for system computing resources, and enhances the user experience of intelligent voice interaction services.
Smart Images

Figure CN121053983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent voice interaction technology, and in particular to a personalized voice interaction method and related equipment. Background Technology
[0002] In related technologies, the implementation process of intelligent voice interaction services typically involves a user asking a question to an intelligent voice interaction assistant. Upon receiving the user's audio question, the assistant uses a speech recognition model to convert the audio into text, then performs information retrieval based on the text content to obtain the corresponding answer, which is then played back to the user via an audio module. Currently, intelligent voice interaction assistants need to provide services to different groups of people. Due to differences in speaking styles (such as accents) among different groups, speech recognition models struggle to accurately identify the audio questions from different groups, leading to incorrect answers and negatively impacting the user experience. Summary of the Invention
[0003] The main objective of this application is to propose a personalized voice interaction method and related equipment, which aims to accurately identify the audio of the question and improve the user experience of intelligent voice interaction services.
[0004] To achieve the above objectives, one aspect of this application proposes a personalized voice interaction method, comprising the following steps: In response to a voice interaction service request, streaming processing technology is used to receive the audio stream of the question asked by the user. The audio stream of the object question is segmented and preprocessed to obtain the audio segments of the object question. Perform object feature analysis on the audio clips of the questions asked by the object to obtain service object tags; The service object tags are classified into object groups, and the corresponding speech recognition model is selected based on the classification results of the object groups; The speech recognition model is used to sequentially perform speech recognition on the audio segments of the object's question to obtain the object's question text; The question text of the object is analyzed to obtain the answer response.
[0005] In some embodiments, performing object feature analysis on the audio clip of the object to obtain the service object tag includes the following steps: Audio features are extracted from the audio segment posed to the object to obtain audio representation features; The audio representation features are classified and mapped using an audio classification model to obtain service object labels.
[0006] In some embodiments, the step of performing label classification mapping on the audio representation features using an audio classification model to obtain service object labels includes the following steps: Convolution and pooling operations are performed on the audio representation features to obtain the voiceprint features; The voiceprint features are mapped using a fully connected layer to obtain the service object label.
[0007] In some embodiments, the audio classification model is trained through the following steps: Obtain a training dataset, which includes multiple training samples, and the training samples include audio representation features and corresponding object real labels; The training dataset is input into the initialized audio classification model for forward propagation to obtain the object prediction labels for the training samples. The model loss is determined by the total loss function based on the predicted labels and the true labels of the objects; the total loss function includes loss function terms and regularization terms for multiple different types of service object labels. The audio classification model is backpropagated based on the model loss to update the parameters of the audio classification model until a preset condition is met, thus obtaining a trained audio classification model.
[0008] In some embodiments, the step of sequentially performing speech recognition on the audio segments of the object's question using the speech recognition model to obtain the object's question text includes the following steps: The speech recognition model is used to perform speech recognition on the audio segments of the questions asked by the object, and the resulting text segments are obtained. According to the acquisition time sequence of the audio segments of the question asked by the object, the text sequences of multiple segments are concatenated to obtain the question text sequence; The question text sequence is post-processed to obtain the object's question text.
[0009] In some embodiments, the step of performing text analysis on the question text of the object to obtain an answer response includes the following steps: The service object tag is added as a prompt word to the object's question text to obtain the text to be analyzed; The text to be analyzed is input into a large language model for text analysis to obtain the response text. Select the appropriate dialogue prompt configuration based on the service object tag, and determine the response based on the dialogue prompt configuration and the response text.
[0010] In some embodiments, the personalized voice interaction method further includes the following steps: Obtain feedback on the object's satisfaction with the response; The audio stream of the question asked by the object, the text of the question asked by the object, the text of the answer, the satisfaction feedback, and the object information are stored in the database as a single interaction data. The speech recognition model is optimized based on the interaction data in the database.
[0011] To achieve the above objectives, another aspect of the embodiments of this application proposes a personalized voice interaction system, including a data acquisition module, an audio processing module, a voiceprint module, and a response module. The audio processing module includes an audio stream segmentation unit, an audio stream processing unit, a model switching unit, and a fusion unit. The acquisition module is used to respond to voice interaction service requests and uses streaming processing technology to receive the audio stream of the question asked by the user. The audio stream segmentation unit and the audio stream processing unit are used to perform segmentation and preprocessing operations on the object question audio stream to obtain object question audio segments. The voiceprint module is used to perform object feature analysis on the audio segment of the question asked by the object to obtain the service object tag; The model switching unit is used to classify the service object tags into object groups and select the corresponding speech recognition model based on the classification results of the object groups. The fusion unit is used to sequentially perform speech recognition on the audio segments of the object's question using the speech recognition model to obtain the object's question text; The answer module is used to perform text analysis on the question text of the object and obtain an answer response.
[0012] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0014] The embodiments of this application include at least the following beneficial effects: This application provides a personalized voice interaction method, system, electronic device, and program product. In response to a voice interaction service request, this solution uses streaming processing technology to receive an audio stream of a user's question. The audio stream is segmented and preprocessed to obtain audio segments of the user's question. Then, user feature analysis is performed on the audio segments to obtain service object tags. These service object tags are then categorized into different user groups. Based on the categorization results, a corresponding speech recognition model is selected. The speech recognition model is then used to sequentially perform speech recognition on the audio segments of the user's question to obtain the user's question text. By combining the user's audio with the user's group during the voice interaction process, a suitable speech recognition model is adaptively selected to recognize the user's audio, improving the accuracy of audio-to-text recognition. Finally, text analysis is performed on the question text to obtain a response, thereby increasing user satisfaction with the response and improving the user experience of the intelligent voice interaction service. Attached Figure Description
[0015] Figure 1 This is a flowchart of the personalized voice interaction method provided in the embodiments of this application; Figure 2 yes Figure 1 The flowchart of step S103 in the process; Figure 3 yes Figure 2 The flowchart of step S202 in the document; Figure 4 yes Figure 2 The flowchart for training the audio classification model in step SS202; Figure 5 yes Figure 1 The flowchart of step S105 in the process; Figure 6 yes Figure 1 The flowchart of step S106 in the process; Figure 7 This is a flowchart of a personalized voice interaction method provided in another embodiment of this application. Figure 8 This is a schematic diagram of the personalized voice interaction system provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application; Figure 10 This is a schematic diagram of the audio classification model processing procedure provided in the embodiments of this application; Figure 11 This is a schematic diagram of the personalized voice interaction process provided in the embodiments of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0018] Before providing a detailed description of the embodiments of this application, some of the nouns and terms involved in the embodiments of this application will be explained first. The nouns and terms involved in the embodiments of this application are subject to the following interpretations.
[0019] MFCC (Mel-Frequency Cepstral Coefficients) is a core feature extraction technique in speech signal processing. It extracts cepstral feature parameters from speech signals by simulating the human ear's auditory perception mechanism, including Mel frequency and cepstral coefficients. Mel frequency represents the characteristic of the human ear being more sensitive to low frequencies and having reduced high-frequency discrimination, using a nonlinear Mel scale. Cepstral coefficients represent the spectral envelope features extracted by separating the sound source excitation from the vocal tract filtering effect through cepstral analysis.
[0020] CNN (Convolutional Neural Network) is a deep learning model specifically designed for processing grid-like data (such as images and speech spectra). Its core functionality involves local perception, parameter sharing, and hierarchical feature extraction to achieve efficient pattern recognition.
[0021] L1 / L2 regularization is a parameter constraint technique used in machine learning to prevent model overfitting. It limits the size of model weights by introducing a penalty term into the loss function.
[0022] In related technologies, the implementation process of intelligent voice interaction services typically involves a user asking a question to an intelligent voice interaction assistant. Upon receiving the user's audio question, the assistant uses a speech recognition model to convert the audio into text, then performs information retrieval based on the text content to obtain the corresponding answer, which is then played back to the user via an audio module. Existing voice interaction systems primarily offer general services, lacking customized interactions based on individual user characteristics (such as age and gender), and thus failing to effectively adapt to personalized needs. While some intelligent robots exist that can configure voice playback parameters based on the interaction object and use corresponding voice packs (such as a teenager's voice or a licensed celebrity's voice) for voice interaction, allowing for personalized customization of the voice pack, they do not consider the speech recognition aspect. Because different groups of people speak differently (such as accents), speech recognition models struggle to accurately identify the audio questions from different groups, leading to incorrect answers and negatively impacting the user experience. Currently, in order for speech recognition models to accurately identify the speech of different users, the network architecture of the models is required to be complex, and a large amount of training data from different users is needed. In order to achieve real-time personalized voice broadcasting, the system needs powerful computing resources to run the model and adjust the model parameters, which may increase hardware costs, especially in mobile devices or resource-constrained environments.
[0023] In view of this, this application provides a personalized voice interaction method and related device. This solution is based on the speech characteristics of different groups of people and trains corresponding speech recognition models based on audio data of specific groups. During the voice interaction process, the user's group is determined by combining the user's audio, thereby adaptively selecting an appropriate speech recognition model to recognize the user's audio, improving the accuracy of audio recognition into text. Then, text analysis is performed on the question text to obtain the response, thereby improving user satisfaction with the response and enhancing the user experience of the intelligent voice interaction service. In addition, compared with the method of recognizing user audio using a general speech recognition model, this solution's speech recognition model for specific groups can not only accurately recognize audio, but also has a simple model, which can reduce the requirements for system computing resources.
[0024] The personalized voice interaction method provided in this application relates to the field of intelligent voice interaction technology. The personalized voice interaction method provided in this application can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart voice assistant, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the personalized voice interaction method, but is not limited to the above forms.
[0025] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0026] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.
[0027] In one embodiment, the personalized voice interaction method of this solution can be applied to scenarios where mobile terminals interact with customer service systems. Users initiate a voice interaction service request by dialing a voice customer service number on their mobile terminals. The customer service system responds to the request by receiving the audio stream of the user's question using streaming processing technology. The audio stream is segmented and preprocessed to obtain audio segments containing the user's question. These segments are then analyzed for user features to obtain service recipient tags, such as age tags. The user tags are further categorized into different groups, such as classifying individuals aged 65 as elderly. A corresponding speech recognition model is selected based on the group classification. For example, if elderly individuals typically use dialects, a dialect-specific speech recognition model can be chosen for this group. The selected speech recognition model is then used to sequentially perform speech recognition on the audio segments containing the user's question to obtain the question text. The question text is then analyzed to obtain the response, which can be visualized and fed back to the mobile terminal.
[0028] In another embodiment, the personalized voice interaction method of this solution can be applied to scenarios where users interact with a smart voice assistant, which can be integrated into a mobile terminal or a smart speaker. The user wakes up the smart voice assistant with a specific command, thus initiating a voice interaction service request. The smart voice assistant responds to the voice interaction service request by using streaming processing technology to receive the audio stream of the user's question. It then performs segmentation and preprocessing operations on the audio stream to obtain audio fragments of the user's question. Next, it performs user feature analysis on these audio fragments to obtain service object tags, such as age tags. The service object tags are then categorized into different groups, such as classifying 30-year-olds as young adults. Based on the group classification, a corresponding speech recognition model is selected. For example, young adults typically use standard Mandarin Chinese for conversation, so a Mandarin speech recognition model can be selected for this group. The selected speech recognition model is then used to sequentially perform speech recognition on the audio fragments of the user's question to obtain the question text. The text text is then analyzed to obtain the response, which is then relayed back to the user via voice broadcast.
[0029] Figure 1 This is an optional flowchart of the personalized voice interaction method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S106.
[0030] Step S101: In response to the voice interaction service request, the audio stream of the question asked by the object is received using streaming processing technology; Step S102: Perform segmentation and preprocessing operations on the audio stream of the object's question to obtain the audio segment of the object's question; Step S103: Perform object feature analysis on the audio segment of the object's question to obtain the service object tag; Step S104: Classify the service object tags into object groups, and select the corresponding speech recognition model based on the classification results of the object groups; Step S105: Use a speech recognition model to sequentially perform speech recognition on the audio segments of the object's question to obtain the text of the object's question; Step S106: Perform text analysis on the object's question text to obtain the answer response.
[0031] Steps S101 to S106, as illustrated in this embodiment, respond to a voice interaction service request by using streaming processing technology to receive an audio stream of a user's question. The audio stream is then segmented and preprocessed to obtain audio segments of the user's question. These segments are then analyzed for user features to obtain service object tags. The service object tags are then categorized into different user groups. Based on the group categorization, a corresponding speech recognition model is selected. The speech recognition model is then used to sequentially perform speech recognition on the audio segments of the user's question to obtain the user's question text. By combining the user's audio with the user's group identification during the voice interaction process, a suitable speech recognition model is adaptively selected to recognize the user's audio, improving the accuracy of audio-to-text recognition. Finally, text analysis is performed on the question text to obtain a response, thereby increasing user satisfaction with the response and improving the user experience of the intelligent voice interaction service.
[0032] In step S101 of some embodiments, the system responds to a voice interaction service request. In different voice interaction scenarios, this request can be a specific voice wake-up command, a wake-up instruction, or a call to customer service. For a voice interaction service request to call customer service, the request carries information indicating the user's location, such as an IP address. After responding to the voice interaction service request, the audio stream receiving unit in its audio processing module uses streaming processing technology to receive and process the audio stream of the question, achieving low latency while maintaining high precision in audio signal processing. The audio stream can be received using WebRTC or other streaming processing technologies to receive real-time audio streams. WebRTC is an open-source project that provides real-time communication capabilities. Additionally, the Real-Time Messaging Protocol (RTMP) can also be used for real-time audio stream transmission.
[0033] In step S102 of some embodiments, the audio stream segmentation unit in the system's audio processing module performs a segmentation operation on the received object question audio stream to divide it into multiple audio segments, which are then preprocessed subsequently. The audio stream segmentation unit uses a sliding window to segment the continuous object question audio stream, ensuring that the data segments input to the speech recognition model have consistent lengths, facilitating model processing. Specifically, the audio stream segmentation unit sets a sliding window and sets its size and step size, and segments the object question audio stream into multiple audio segments through a segmentation operation. The calculation expression is as follows: ; Where N is the number of segments, S is the step size, T is the window size, and the starting position is t.
[0034] The audio stream processing unit in the system's audio processing module performs preprocessing operations on the audio segments segmented by the audio stream segmentation unit to obtain the corresponding object query audio segments. The preprocessing operations include at least one of the following: noise reduction, volume normalization, and silence removal. Noise reduction can use noise suppression algorithms (such as spectral subtraction and wavelet transform) to remove background noise while preserving the main components of the speech signal. Volume normalization normalizes the audio segments to ensure that the volume of all audio segments is within the same range, avoiding the impact of volume differences on model prediction and training. Silence removal detects and removes silent parts from audio segments using parameters such as audio amplitude, reducing unnecessary computational burden and improving processing efficiency. Furthermore, the preprocessing operations may also include a sampling rate unification operation, converting all audio segments to a uniform sampling rate (e.g., 16kHz) to improve the consistency of the audio data.
[0035] In step S103 of some embodiments, object feature analysis is performed on the object question audio segments to obtain service object tags. Specifically, the system uses streaming processing technology to receive the object question audio stream and simultaneously performs segmentation and preprocessing operations on the object question audio stream, sequentially outputting object question audio segments. The first few object question audio segments are selected for object feature analysis to determine the service object tags. During the object feature analysis, the analysis mainly focuses on voiceprint-related features in the audio, which can determine service object tags such as the user's gender and age.
[0036] In step S104 of some embodiments, this application embodiment trains a corresponding speech recognition model based on the speech characteristics of different groups of people and audio data of specific groups. The speech recognition model is used to convert the input audio data into a text sequence. During voice interaction, the model switching unit in the system's audio processing module determines the user's gender, age, and other service target tags based on the user's audio, and dynamically selects the most suitable speech recognition model according to the group to which the service target tag belongs. For example, for the elderly, a speech recognition model that supports more dialects may be selected, while for children, a speech recognition model that can better handle the pronunciation characteristics of children is selected, thereby improving the accuracy of audio recognition into text.
[0037] For example, the speech recognition model can be switched by conditional judgment. The code representation of the conditional judgment method is as follows: ifuser_age<12:model=children_model elseifuser_age>=60:model=senior_modelelseifuser_age>=60:model=senior_model else:model=general_modelelse:model=general_model In another example, when selecting a speech recognition model, the user's age is determined by analyzing the voiceprint features of the first few audio segments of the object's questions. If the user's age is greater than the age threshold, it indicates that the user may be an elderly person. At this time, the corresponding regional dialect is determined based on the information of the user's location carried in the voice interaction service request. Then, the corresponding dialect language recognition model is selected to perform speech recognition on the object's question audio stream to convert it into a text sequence.
[0038] In step S105 of some embodiments, the audio segments of each object question are sequentially input into the speech recognition model to obtain the corresponding segment text sequence. The fusion unit in the audio processing module of the system fuses the segment text sequences to form a complete recognition output, namely the object question text.
[0039] In step S106 of some embodiments, after obtaining the object's question text, text analysis is performed on the object's question text to obtain an answer response. Text analysis refers to mapping the object's question text into another type of text (i.e., the answer response) through certain mapping rules. In one example, text analysis can be implemented through a large language model, that is, the object's question text is input into a large language model for analysis to obtain the answer response. In another example, text analysis can also be implemented through text retrieval, that is, the object's question text is input into a constructed knowledge graph for information retrieval to obtain the answer response. Furthermore, the answer response can be fed back to the user through visualization or audio playback.
[0040] According to some embodiments of this application, please refer to Figure 2 Step S103 may include, but is not limited to, the following steps S201 to S202: Step S201: Extract audio features from the audio segment of the question asked by the object to obtain audio representation features; Step S202: The audio representation features are classified and mapped using an audio classification model to obtain the service object labels.
[0041] In step S201 of some embodiments, the identification of the service object label can be achieved through a voiceprint module in the system. The voiceprint module includes a feature extraction unit and a feature mapping unit. The feature extraction unit is used to extract audio features from the audio segment of the object question to obtain audio representation features. The audio representation features may include, but are not limited to, spectrograms, Mel frequency cepstral coefficients (MFCCs), zero-crossing rate, etc. MFCC features are used to represent speech feature information, and the zero-crossing rate is used to reflect the rate of change of the audio signal. Specifically, the feature extraction unit is used to convert the preprocessed audio segment of the object question into a spectrogram and display the energy distribution of the audio at different frequencies. By extracting the MFCC features of the displayed audio, the frequency domain characteristics of the audio are captured, the zero-crossing rate of the audio signal is calculated, and the energy features and fundamental frequency features of the audio segment of the object question are extracted. For example, the feature extraction unit can use the librosa library to extract audio features, use the librosa.feature.melspectrogram function to extract the Mel spectrogram, use the librosa.feature.mfcc function to extract MFCC features, and use the librosa.feature.zero_crossing_rate function to calculate the zero-crossing rate. In another example, the feature extraction unit can also use deep learning models (such as ResNet, CNN-LSTM combination, etc.) to extract audio representation features.
[0042] In step S202 of some embodiments, the feature mapping unit is used to perform label classification mapping on the audio representation features through an audio classification model to obtain service object labels (such as gender and age). The audio classification model can be a neural network. In this embodiment, a supervised learning method can be used to train a convolutional neural network model to learn the mapping relationship from audio representation features to service object labels such as gender and age, thereby obtaining the audio classification model. It can be understood that the deep learning model used by the feature extraction unit and the audio classification model used by the feature mapping unit can be trained as one model or separately. In this embodiment, the audio segment of the question asked by the target is represented as an audio representation feature. Then, by performing label classification mapping on the audio representation feature through the audio classification model, user attributes related to voiceprint (i.e., service object labels) can be analyzed, which facilitates the subsequent dynamic adaptive selection of a suitable speech recognition model and improves the accuracy of speech recognition.
[0043] According to some embodiments of this application, please refer to Figure 3 Step S202 may include, but is not limited to, the following steps: Step S301: Perform convolution and pooling operations on the audio representation features to obtain the voiceprint features; Step S302: The voiceprint features are mapped through a fully connected layer to obtain the service object label.
[0044] In this embodiment, please refer to Figure 10 The audio classification model can employ a convolutional neural network (CNN) model, including an input layer, a first convolutional layer, a second convolutional layer, a max-pooling layer, and a fully connected layer. The input layer takes in audio representation features, performs a first convolution operation on these features through the first convolutional layer, then performs a second convolution operation on the output of the first convolutional layer, and finally, a max-pooling layer reduces the dimensionality of the convolutional data and extracts important features. The fully connected layer has H neurons and a corresponding weight matrix W, which transforms the convolutional neural network feature mapping into the vectors required for classification (such as gender or age), achieving end-to-end learning. In one example, the audio classification model of this scheme can also output multiple labels simultaneously, such as mapping gender and age. Based on a convolutional neural network, the audio classification model includes an input layer, a shared feature extraction layer, an age recognition task branch, and a gender recognition task branch. The shared feature extraction layer is used to perform convolution operations to extract shared features for multiple labels. The age recognition task branch is used to perform further convolution on the data output by the shared feature extraction layer to extract age features and then map them to the user's age through a fully connected layer. The gender recognition task branch is used to perform further convolution on the data output by the shared feature extraction layer to extract gender features and then map them to the user's gender through a fully connected layer. This embodiment uses a convolutional neural network to extract and map voiceprint features from the input audio representation features, which can accurately identify the service object labels.
[0045] According to some embodiments of this application, please refer to Figure 4 The audio classification model in step S202 can be trained by the model optimization unit of the speaker module. The model optimization unit can combine the extracted features and corresponding labels into a training dataset, build a convolutional neural network model, train the convolutional neural network model through supervised learning methods, optimize the model parameters through backpropagation algorithm, and evaluate the model performance using methods such as cross-validation to ensure that the model has good generalization ability. The specific training steps are as follows: Step S401: Obtain the training dataset, which includes multiple training samples. The training samples include audio representation features and corresponding object real labels. Step S402: Input the training dataset into the initialized audio classification model for forward propagation to obtain the object prediction labels for the training samples; Step S403: Determine the model loss based on the predicted object label and the true object label using the total loss function; the total loss function includes loss function terms and regularization terms for multiple different types of service object labels. Step S404: Backpropagate the audio classification model based on the model loss to update the parameters of the audio classification model until the preset conditions are met, and a trained audio classification model is obtained.
[0046] In step S401 of some embodiments, in historical voice interaction scenarios, by collecting audio streams from different users, and through the segmentation, preprocessing, and audio feature extraction processes described above, audio representation features of different segments from multiple different users can be obtained. The object's true label is then labeled for each audio segment to form training samples, and multiple training samples form a training dataset. Furthermore, the training dataset can be augmented by adding different types of background noise to the audio, changing the audio speed, etc., to improve the recognition performance of the audio classification model.
[0047] In some embodiments, step S402 involves inputting the training dataset into the initialized audio classification model and performing forward propagation to obtain object prediction labels for multiple training samples. The forward propagation process is the same as steps S301 to S302 described above, and will not be repeated here.
[0048] In some embodiments, step S403 involves determining the model loss based on the predicted object label and the true object label using a total loss function. The total loss function includes losses for the age recognition task and the gender recognition task. The age recognition task loss uses the mean squared error (MSE) loss function, and the gender recognition task loss uses the binary cross-entropy loss function. The total loss function is a weighted sum of the losses from the two tasks, enabling the model to recognize different labels. Furthermore, L1 / L2 regularization parameters can be applied to the loss function to prevent overfitting. L1 / L2 regularization includes L1 regularization and L2 regularization. L1 regularization adds an L1 norm to the loss function, making the convolutional neural network tend towards a sparse solution, while adding an L2 norm reduces the weights of the convolutional neural network model.
[0049] In some embodiments, step S404 involves backpropagating the audio classification model based on the model loss to continuously update the model's parameters. During backpropagation, Dropout technology can be used to prevent overfitting. Dropout randomly discards a portion of neurons during training to prevent overfitting of the convolutional neural network model. Furthermore, during backpropagation, methods such as grid search and random search are used to find the optimal hyperparameter combination. Multiple tasks (such as age and gender recognition) are trained simultaneously, sharing the underlying feature extraction layer to improve the overall model performance. Fine-tuning is performed using a pre-trained model to improve performance with limited data. After training until preset conditions are met, such as reaching a certain number of iterations or achieving a certain level of accuracy, the trained audio classification model is output.
[0050] The model optimization unit in this embodiment continuously iterates and trains the model to optimize its performance under different background noise conditions, thereby improving the model's generalization ability and robustness.
[0051] According to some embodiments of this application, please refer to Figure 5 Step S105 may include, but is not limited to, the following steps: Step S501: Use a speech recognition model to perform speech recognition on the audio segments of the questions asked by the object, and obtain the segment text sequence; Step S502: According to the acquisition time sequence of the audio segments of the question, the multiple segment text sequences are concatenated to obtain the question text sequence; Step S503: Post-process the question text sequence to obtain the object question text.
[0052] In this embodiment, after determining the user tag of the current interaction, a corresponding speech recognition model is selected. The speech recognition model is used to perform speech recognition on the audio segments of the object question in the audio stream to obtain segment text sequences. The fusion unit concatenates multiple segment text sequences according to the acquisition time order of the object question audio segments to obtain the question text sequence. Further post-processing operations are performed on the question text sequence to improve the accuracy and coherence of the final text. Post-processing operations include deduplication, smoothing, and correction. Deduplication is used to remove repeated words or phrases; smoothing is used to smooth the question text sequence using a language model or rules to improve text coherence; correction is used to correct possible errors based on the context information in the question text sequence. Specifically, the fusion unit concatenates the recognition results of each segment (i.e., segment text sequences) according to the time order, and performs deduplication, smoothing, and correction through post-processing functions. The fusion process is represented as follows: ; Where R represents the question text. ( ) represents the post-processing function. ( ) represents the recognition results of splicing the segments in chronological order, r represents the audio segment of the object question, and n represents the total number of audio segments of the object question.
[0053] According to some embodiments of this application, please refer to Figure 6 Step S106 may include, but is not limited to, the following steps: Step S601: Add the service object tag as a prompt word to the object query text to obtain the text to be analyzed; Step S602: Input the text to be analyzed into the large language model for text analysis to obtain the response text; Step S603: Select the corresponding dialogue prompt configuration according to the service object label, and determine the response based on the dialogue prompt configuration and the response text.
[0054] In steps S601 to S602 of some embodiments, service target tags (gender, age) obtained based on voiceprint analysis are added as prompt words to the target's question text to obtain the text to be analyzed. This provides richer information for the large language model text analysis process, thereby improving the accuracy of the answer text. It is understood that, with the user's permission, other personal information pre-stored in the system can also be added as prompt words to the text to be analyzed, further enriching the parameter information of the large language model and thus providing more satisfactory answers to the user.
[0055] In step S603 of some embodiments, the large language model can output corresponding answer text based on the user's question. The answer text can be displayed through a visual interface, which can also display corresponding dialogue prompt configurations for specific user groups. The answer text and dialogue prompts together form the response. The dialogue prompt configuration can be set by operators or developers according to the characteristics of the target user group (such as age, interests, etc.) to set specific dialogue flows and response rules. For example, more concise and clear prompts can be designed for the elderly, and more interesting interactive content can be designed for children, thereby improving the user's voice interaction experience.
[0056] According to some embodiments of this application, please refer to Figure 7 The personalized voice interaction method in this application embodiment may also include, but is not limited to, the following steps: Step S701: Obtain feedback on the subject's satisfaction with the response; Step S702: Store the audio stream of the object's question, the text of the object's question, the text of the answer, the satisfaction feedback, and the object's information as a single interaction data in the database; Step S703: Optimize the speech recognition model based on the interaction data in the database.
[0057] In this embodiment, the system further includes a data callback module, used to optimize the speech recognition model or text analysis algorithm (such as the aforementioned large language model or knowledge graph) based on interaction data in the database. The data callback module employs a data callback mechanism to record every interaction between the user and the system for subsequent analysis and model training. The data callback module includes a data acquisition, data storage, data upload, and data verification process, as detailed below: The data acquisition process includes collecting user input, system response, timestamps, and user information. User input refers to the audio or text content entered by the user, system response refers to the response content generated by the system, timestamps refer to the timestamps of each interaction, and user information refers to the user's identification information (such as user ID) and basic attributes (such as age and gender).
[0058] The data storage process includes database selection, data format conversion, and data security protection. Database selection specifically involves choosing a suitable database system (such as MySQL, MongoDB, etc.) to store interactive data. Data format conversion specifically involves converting the data format into the designed data table structure or document structure to improve data integrity and consistency. Data security protection refers to protecting private data.
[0059] The data upload process is divided into real-time upload and batch upload. Real-time upload is used to upload data to the server immediately after each interaction. Batch upload is used to upload data in batches at regular intervals to reduce network overhead.
[0060] The data verification process includes integrity verification and consistency verification. Integrity verification is used to confirm that each record contains all the necessary fields; consistency verification is used to ensure the consistency and correctness of the data.
[0061] This application embodiment also designs a fallback strategy for the system's response, that is, when the system cannot provide a satisfactory answer, it can guide the user to try other ways to obtain help or directly transfer to human service, without affecting the user experience. Specifically, please refer to... Figure 11 The personalized voice interaction process of the embodiments of this application will be described.
[0062] The user inputs an audio stream of a question into the system, and the system provides a response to the user through the personalized voice interaction method described in the above embodiment. The system determines whether the response is satisfactory (i.e., whether the identification failed) based on the user's interaction with the system. Identification failures include low confidence, user feedback, and multiple questions asked by the user. Low confidence means that when the confidence of the response generated by the system is lower than a certain threshold, it is considered a failure. User feedback means that when the user explicitly expresses dissatisfaction, it is considered a failure. Multiple questions asked by the user means that when the user asks the same question multiple times, it is considered a failure. If a single recognition attempt fails, the user is guided to try other methods for speech recognition. These methods include prompts, links to frequently asked questions (FAQs), and search engine suggestions. Prompts provide helpful information to encourage users to rephrase their questions or provide more context. FFA links help users resolve issues on their own by providing links to answers to common questions. Search engine suggestions recommend that users use a search engine to find relevant information.
[0063] If identification fails multiple times, the user will be transferred to a human agent. Methods for transferring to a human agent include automatic transfer, user request, and priority management. Automatic transfer means that the system automatically transfers the user to a human agent after multiple failures. User request provides the option for users to manually request a human agent, guiding them to connect manually. Priority management sets transfer priorities based on the user's historical interaction records and importance, and then automatically transfers users with higher priorities.
[0064] Please refer to Figure 8 This application also proposes a personalized voice interaction system, including a data acquisition module, an audio processing module, a voiceprint module, and a response module. The audio processing module includes an audio stream segmentation unit, an audio stream processing unit, a model jump unit, and a fusion unit. The acquisition module is used to respond to voice interaction service requests and uses streaming processing technology to receive the audio stream of the question asked by the user. The audio stream segmentation unit and the audio stream processing unit are used to perform segmentation and preprocessing operations on the object question audio stream to obtain object question audio segments; The voiceprint module is used to perform object feature analysis on the audio clips of the object's questions to obtain the service object's tag; The model jump unit is used to classify the service object tags into groups and select the corresponding speech recognition model based on the classification results of the group. The fusion unit is used to sequentially perform speech recognition on the audio segments of the object's question using a speech recognition model to obtain the object's question text; The response module is used to perform text analysis on the question text of the object and obtain the response.
[0065] It is understood that the methods described in the above method embodiments are applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0066] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0067] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0068] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0069] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.
[0070] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0071] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0072] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0073] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0074] The personalized voice interaction method and related devices provided in this application determine the user's group by combining the user's audio during the voice interaction process, thereby adaptively selecting a suitable speech recognition model to recognize the user's audio, improving the accuracy of audio recognition into text, and then performing text analysis on the question text to obtain the answer response, thereby improving the user's satisfaction with the answer response and improving the user experience of intelligent voice interaction services.
[0075] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0076] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0077] The system embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0078] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0079] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0080] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0081] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0082] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0083] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0084] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A personalized voice interaction method, characterized in that, Includes the following steps: In response to a voice interaction service request, streaming processing technology is used to receive the audio stream of the question asked by the user. The audio stream of the object question is segmented and preprocessed to obtain the audio segments of the object question. Perform object feature analysis on the audio clips of the questions asked by the object to obtain service object tags; The service object tags are classified into object groups, and the corresponding speech recognition model is selected based on the classification results of the object groups; The speech recognition model is used to sequentially perform speech recognition on the audio segments of the object's question to obtain the object's question text; The question text of the object is analyzed to obtain the answer response.
2. The method according to claim 1, characterized in that, The process of performing object feature analysis on the audio clip of the question to obtain the service object tag includes the following steps: Audio features are extracted from the audio segment posed to the object to obtain audio representation features; The audio representation features are classified and mapped using an audio classification model to obtain service object labels.
3. The method according to claim 2, characterized in that, The step of performing label classification mapping on the audio representation features using an audio classification model to obtain service object labels includes the following steps: Convolution and pooling operations are performed on the audio representation features to obtain the voiceprint features; The voiceprint features are mapped using a fully connected layer to obtain the service object label.
4. The method according to claim 2, characterized in that, The audio classification model is trained through the following steps: Obtain a training dataset, which includes multiple training samples, and the training samples include audio representation features and corresponding object real labels; The training dataset is input into the initialized audio classification model for forward propagation to obtain the object prediction labels for the training samples. The model loss is determined by the total loss function based on the predicted labels and the true labels of the objects; the total loss function includes loss function terms and regularization terms for multiple different types of service object labels. The audio classification model is backpropagated based on the model loss to update the parameters of the audio classification model until a preset condition is met, thus obtaining a trained audio classification model.
5. The method according to claim 1, characterized in that, The step of using the speech recognition model to sequentially perform speech recognition on the audio segments of the object's question to obtain the object's question text includes the following steps: The speech recognition model is used to perform speech recognition on the audio segments of the questions asked by the object, and the resulting text segments are obtained. According to the acquisition time sequence of the audio segments of the question asked by the object, the text sequences of multiple segments are concatenated to obtain the question text sequence; The question text sequence is post-processed to obtain the object's question text.
6. The method according to any one of claims 1 to 5, characterized in that, The process of analyzing the question text of the object to obtain a response includes the following steps: The service object tag is added as a prompt word to the object's question text to obtain the text to be analyzed; The text to be analyzed is input into a large language model for text analysis to obtain the response text. Select the appropriate dialogue prompt configuration based on the service object tag, and determine the response based on the dialogue prompt configuration and the response text.
7. The method according to claim 6, characterized in that, The personalized voice interaction method further includes the following steps: Obtain feedback on the object's satisfaction with the response; The audio stream of the question asked by the object, the text of the question asked by the object, the text of the answer, the satisfaction feedback, and the object information are stored in the database as a single interaction data. The speech recognition model is optimized based on the interaction data in the database.
8. A personalized voice interaction system, characterized in that, It includes a data acquisition module, an audio processing module, a voiceprint module, and a response module. The audio processing module includes an audio stream segmentation unit, an audio stream processing unit, a model jump unit, and a fusion unit. The acquisition module is used to respond to voice interaction service requests and uses streaming processing technology to receive the audio stream of the question asked by the user. The audio stream segmentation unit and the audio stream processing unit are used to perform segmentation and preprocessing operations on the object question audio stream to obtain object question audio segments. The voiceprint module is used to perform object feature analysis on the audio segment of the question asked by the object to obtain the service object tag; The model switching unit is used to classify the service object tags into object groups and select the corresponding speech recognition model based on the classification results of the object groups. The fusion unit is used to sequentially perform speech recognition on the audio segments of the object's question using the speech recognition model to obtain the object's question text; The answer module is used to perform text analysis on the question text of the object and obtain an answer response.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Voiceprint identification method and apparatus
CN107507612A
Intelligent household electrical appliance control method and device and intelligent household electrical appliance
CN110336723A
Personalized voice interaction method and system
CN115188376A
Voice sign recognition method and device, electronic equipment and storage medium
CN117037846A
Voice recognition-based interaction method, apparatus and device, and storage medium
CN120199252A
Cited By
Dialect speech recognition method, storage medium and electronic device
CN121393420A