Companion robot speech recognition method and system
By performing time-domain and frequency-domain feature processing on speech data, selecting speech features with high wake-up rates and accuracy, and determining speech preference features, the problem of low accuracy in speech recognition frameworks when dealing with regional accents and unclear speech is solved, thereby improving recognition accuracy and user experience.
Patent Information
- Application Number
- CN202411996484.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-12-30
AI Technical Summary
Existing speech recognition frameworks suffer from low accuracy when dealing with speech containing regional accents or unclear speech from young children or the elderly, which negatively impacts user experience.
By acquiring speech data, performing time-domain and frequency-domain feature processing, filtering speech feature vectors with high wake-up rates and accuracy, determining speech preference features, and performing feature alignment and recognition.
It improves the accuracy and response speed of speech recognition, enhancing the user interaction experience.
Smart Images

Figure CN119920237B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and more particularly to a companion robot speech recognition method and system. BACKGROUND
[0002] With the rapid development of artificial intelligence, machine learning, Internet of Things and other technologies, the performance of intelligent companion robots is continuously improving, and the functions are becoming more diverse. Intelligent companion robots can have smooth conversations with users through natural language processing technology and provide customized services according to the personalized needs of users, becoming an important auxiliary tool in modern families, educational institutions and medical service fields. In the application of intelligent companion robots, the accuracy of speech recognition is crucial to user experience. Current speech recognition frameworks often face challenges when processing speech content with local accents, unclear speech of children or the elderly, which directly affects the accuracy of recognition and user experience in the interaction process.
[0003] Therefore, there is an urgent need for a high-accuracy speech recognition method. SUMMARY
[0004] In view of the above problems, the purpose of the present application is to provide a companion robot speech recognition method and system to solve at least one problem existing in the prior art.
[0005] According to one aspect of the present application, a companion robot speech recognition method is provided, applied to an electronic device, comprising:
[0006] Obtaining speech data;
[0007] Processing the speech data based on time domain features and obtaining a time domain feature dataset of first speech feature vectors; inputting the time domain feature dataset of the first speech feature vectors into a companion robot, and obtaining the wake-up rate corresponding to each first speech feature vector in the time domain feature dataset of the first speech feature vectors; screening first speech feature vectors with a wake-up rate higher than a preset first wake-up rate threshold;
[0008] Processing the screened first speech feature vectors based on frequency domain features and obtaining a frequency domain feature dataset of second speech feature vectors; inputting the frequency domain feature dataset of the second speech feature vectors into the companion robot, obtaining the accuracy rate corresponding to each speech feature vector in the frequency domain feature dataset of the second speech feature vectors; screening second speech feature vectors with an accuracy rate higher than a preset second accuracy rate threshold;
[0009] Determining the speech preference features of the companion robot according to the second speech feature vectors.
[0010] In addition, the time domain feature data set of the first speech feature vector comprises a short-time energy feature parameter adjusted first speech feature vector and a zero-crossing rate feature parameter adjusted first speech feature vector.
[0011] In addition, the frequency domain feature data set of the second speech feature vector comprises a mel-frequency cepstrum coefficient feature parameter adjusted second speech feature vector and a spectral centroid feature parameter adjusted second speech feature vector.
[0012] In addition, the short-time energy feature parameter adjusted first speech feature vector comprises a frame size parameter adjusted first speech feature vector, a frame shift parameter adjusted first speech feature vector, a smoothing window size parameter adjusted first speech feature vector and a pre-emphasis coefficient parameter adjusted first speech feature vector.
[0013] The zero-crossing rate feature parameter adjusted first speech feature vector comprises a frame size feature parameter adjusted first speech feature vector, a frame shift feature parameter adjusted first speech feature vector, a window function feature parameter adjusted first speech feature vector and a high-pass filter cutoff frequency feature parameter adjusted first speech feature vector.
[0014] The frame size feature parameter adjusted first speech feature vector in the short-time energy feature parameter adjusted first speech feature vector and the zero-crossing rate feature parameter adjusted first speech feature vector is the same as the frame shift feature parameter adjusted first speech feature vector.
[0015] In addition, the mel-frequency cepstrum coefficient feature parameter adjusted second speech feature vector comprises a mel-frequency range feature parameter adjusted second speech feature vector, a DCT coefficient number feature parameter adjusted second speech feature vector, a pre-emphasis feature parameter adjusted second speech feature vector and a logarithm operation feature parameter adjusted second speech feature vector.
[0016] In addition, the spectral centroid feature parameter adjusted second speech feature vector comprises a spectral resolution feature parameter adjusted second speech feature vector, a spectral smoothing feature parameter adjusted second speech feature vector, a window function feature parameter adjusted second speech feature vector and a self-defined weighting feature parameter adjusted second speech feature vector.
[0017] In addition, the first wake-up threshold and the second wake-up threshold are both 95%.
[0018] In addition, the method further comprises:
[0019] align the voice data to be recognized with the voice preference feature of the companion robot;
[0020] perform voice recognition on the voice data after the feature alignment.
[0021] Further, the optional technical solution is, before the step of processing the voice data based on short-time energy and zero-crossing rate and obtaining a time-domain feature data set, further comprising,
[0022] perform noise reduction preprocessing on the voice data using a Wiener filter model.
[0023] On the other hand, the present application also provides a companion robot voice recognition system, which performs voice recognition using the companion robot voice recognition method as described above; comprising:
[0024] an acquisition unit for acquiring voice data;
[0025] a first screening unit for processing the voice data based on time-domain features and obtaining a time-domain feature data set of first voice feature vectors; inputting the time-domain feature data set of the first voice feature vectors into a companion robot, and acquiring a wake-up rate corresponding to each first voice feature vector in the time-domain feature data set of the first voice feature vectors; screening first voice feature vectors with a wake-up rate higher than a preset first wake-up rate threshold;
[0026] a second screening unit for processing the screened first voice feature vectors based on frequency-domain features and obtaining a frequency-domain feature data set of second voice feature vectors; inputting the frequency-domain feature data set of the second voice feature vectors into a companion robot, acquiring an accuracy rate corresponding to each voice feature vector in the frequency-domain feature data set of the second voice feature vectors; screening second voice feature vectors with an accuracy rate higher than a preset second accuracy rate threshold;
[0027] a feature determination unit for determining a voice preference feature of a companion robot according to the second voice feature vectors.
[0028] The companion robot voice recognition method and system described above, based on time domain features, processes voice data and obtains a time domain feature dataset of first voice feature vectors; the wake-up rate corresponding to each first voice feature vector in the time domain feature dataset of first voice feature vectors is obtained; first voice feature vectors with a wake-up rate higher than a preset first wake-up rate threshold are screened; based on frequency domain features, the screened first voice feature vectors are processed and a frequency domain feature dataset of second voice feature vectors is obtained; the accuracy rate corresponding to each voice feature vector in the frequency domain feature dataset of second voice feature vectors is obtained; second voice feature vectors with an accuracy rate higher than a preset second accuracy rate threshold are screened; and the voice preference features of the companion robot are determined according to the second voice feature vectors. Through the screening mechanism based on the wake-up rate and the accuracy rate, the most effective voice features for the companion robot can be identified, so as to improve the response speed and accuracy of voice recognition according to the most effective voice features, and finally achieve the technical effects of effectively improving the voice recognition accuracy of the companion robot and improving the user interaction experience.
[0029] To the accomplishment of the foregoing and related ends, one or more aspects of the application comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the annexed drawings set forth in detail certain illustrative aspects of the application. These aspects are indicative, however, of but a few of the various ways in which the principles of the application can be employed. Other objects, advantages, and novel features of the application will become apparent from the following detailed description when considered in conjunction with the drawings. BRIEF DESCRIPTION OF DRAWINGS
[0030] Other objects and results of the application will become more fully understood and appreciated only upon consideration of the following detailed description, taken in conjunction with the accompanying drawings. In the drawings:
[0031] Figure 1 A flowchart of a companion robot voice recognition method according to an embodiment of the application;
[0032] Figure 2 A module schematic diagram of a companion robot voice recognition system provided by an embodiment of the application.
[0033] Figure 3 An internal structure schematic diagram of an electronic device for implementing a companion robot voice recognition method provided by an embodiment of the application.
[0034] The same reference numbers in all the drawings indicate similar or corresponding features or functions. DETAILED DESCRIPTION
[0035] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the protection scope of the present application.
[0036] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings. In the description of the embodiments of the present application, "and / or" in the text only represents an association relationship of associated objects, and indicates that there can be three relationships, for example, A and / or B can represent three cases of A existing alone, A and B existing together, and B existing alone.
[0037] Hereinafter, the terms "first" and "second" are only for description purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with "first" and "second" can explicitly or implicitly include one or more features, and in addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.
[0038] In the present specification, the reference "one embodiment" or "some embodiments" and the like means that a particular feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in yet some embodiments" and the like appearing in various places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "include", "contain", "have" and their variants mean "including but not limited to", unless otherwise specifically emphasized.
[0039] In order to describe the companion robot voice recognition method and system of the present application in detail, the specific embodiments of the present application will be described in detail below with reference to the drawings.
[0040] AI is a theory, method, technology and application system for simulating, extending and expanding human intelligence by using a digital computer or a machine controlled by a digital computer, perceiving an environment, acquiring knowledge and using the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that the machine has the functions of perception, reasoning and decision-making.
[0041] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software level technology. Artificial intelligence basic technology generally includes, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning and other major directions.
[0042] Natural language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing involves natural language, i.e. the language used in daily life, and is closely related to linguistics; at the same time, it involves computer science and mathematics, an important technology for model training in the field of artificial intelligence, and a pre-training model, i.e. a large language model (LLM) developed from the field of NLP. After fine-tuning, large language models can be widely used in downstream tasks. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0043] Machine learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It is a specialized study of how computers can simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications span all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.
[0044] Figure 1 A flowchart of a companion robot voice recognition method according to an embodiment of the present application is shown.
[0045] As shown in Figure 1 The companion robot voice recognition method provided in the present embodiment is applied to an electronic device and mainly includes the following steps:
[0046] S110: Acquire voice data.
[0047] The companion robot in the embodiment can be any terminal device capable of voice interaction data. The terminal device can be a smart phone, a tablet computer, a notebook computer, a palm computer, a personal computer, a smart television, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The interaction information between the terminal device and the user can include voice information and video information. It should be noted that the voice data is voice data selected from a language database for user interaction with the robot. The companion robot allows the user to communicate with the system in natural language through its voice interaction function, easily obtains information, and completes various tasks according to user instructions. For example, the companion robot can also provide emotional support services according to the emotions of the user captured by emotion recognition and analysis technology.
[0048] The companion robot involved in the embodiment can include a speaker, a microphone, and / or a camera, and can have, for example, a communication function and be capable of establishing a wired or wireless communication connection with an electronic device such as a mobile phone or a computer. For example, the wireless connection can be a near-field transmission technology such as a wireless fidelity (Wi-Fi) connection or a bluetooth connection. The wired connection can be a universal serial bus (USB) connection, a high definition multimedia interface (HDMI) connection, etc. The type of the communication connection is not limited in the embodiment. The device can share a call function such as answering and making a call with other communication devices through the established communication connection. The device can be built-in with a SIM card to realize answering and making a call, so as to have the function of a mobile communication device. When interacting with the companion robot, the user can conveniently make a call without holding other communication devices.
[0049] In the embodiment, the audio data transmission interface can adopt an I2S interface, a PCM interface, or a PDM interface. The I2S (Inter-IC Sound, Integrated Interchip Sound, or IIS for short) is a digital audio transmission standard for transmitting digital audio data between internal devices of a system, such as a CODEC, a DSP, a digital input / output interface, an ADC, a DAC, and a digital filter. The PCM (pulse code modulation) interface is composed of a clock pulse (BCLK), a frame synchronization signal (FS), and received data (DR) and transmitted data (DX), and accordingly includes four signal lines for transmitting a data clock signal, a frame synchronization clock signal, a received data signal, and a transmitted data signal. The PDM interface transmits PDM encoded data and has only two signal lines for transmitting a clock signal and a data signal.
[0050] Correspondingly, the signal processing module in the embodiment can adopt a DSP (Digital Signal Process) module, or other similar signal processing modules. The signal processing module can be an external signal processing module independent of the electronic device, i.e., a signal processing module arranged outside the electronic device, so as to realize the dynamic music modulation function through an additional device without changing the hardware settings of the electronic device itself on the basis of the existing electronic device; or the signal processing module can be a signal processing module built in the electronic device, so that the user can realize the dynamic music modulation function based on one device.
[0051] Specifically, as an example, the SOC (System On Chip) of the electronic device sends the audio stream data played by the electronic device to the DSP module of the SOC or an external DSP module through the I2S interface of the SOC, so as to modulate and process the audio stream data by the DSP module.
[0052] Before the step of processing the voice data based on the short-time energy and the zero-crossing rate and obtaining the time-domain feature data set, the voice data is preprocessed by a Wiener filtering model. The Wiener filtering is a linear minimum mean square error (LMMSE) estimator, mainly used to recover the original signal from a signal contaminated by noise. Its goal is to minimize the mean square error (MSE) between the recovered signal and the original signal. In speech enhancement, when the voice signal is disturbed by environmental noise to form a noisy voice, the Wiener filter can be used to estimate the pure voice signal. For example, in a mobile phone call or a voice recognition system, the background noise is reduced, and the voice quality and recognition rate are improved.
[0053] The noise reduction processing can also include, for example, removing background noise in the voice data, etc., by excluding the interference of noise, the accuracy of subsequent voice data recognition is improved. In the specific implementation process, the preprocessing can include that after the voice signal to be recognized is divided into multiple audio frames, the feature vector capable of representing the voice signal can be extracted frame by frame. Fourier transform can be performed on each audio frame, and then the frequency domain features are extracted as the feature vector of the audio frame. The preprocessing can also include effective sound detection, speaker separation, speech enhancement, etc.; which is not specifically limited here.
[0054] In the implementation, the speech data can be extracted by principal component analysis, short-time Fourier transform, time-frequency analysis method, mel frequency cepstral coefficients (MFCCs), deep learning feature extraction method, etc. The machine learning model is a mathematical structure or algorithm that can learn useful information from input data and make predictions or decisions on new data. In the field of machine learning, the model is established by training on existing data. The training process involves optimizing the model parameters to best fit the training data and to expect good generalization ability on unknown data. In this embodiment, the preset machine learning model can be a machine learning model suitable for classification and identification tasks, such as logistic regression model, decision tree, random forest, support vector machine or deep learning model.
[0055] Taking the logistic regression model as an example, the following is described. The logistic regression model is a generalized linear model used to handle binary classification problems (which can also be extended to multi-classification problems). Its basic structure is based on linear regression, which assumes that the input speech feature vector is x = [x1, x2, …, xn] (where x is the speech feature obtained by the feature extraction method mentioned above, and n is the feature dimension). The model first calculates the linear combination n ](here x i is the speech feature obtained by the feature extraction method mentioned above, and n is the feature dimension). The model first calculates the linear combination
[0056] z = θ0+ θ1x1+ θ2x2+ … + θ n x n n, where θ [θ0, θ1, …, θ n ] is the parameter (θ0 is the intercept term) that the model needs to learn.
[0057] Then the linear combination result is converted into a probability value by a logistic function (such as a sigmoid function), and the sigmoid function is in the form of The output y represents the probability that the sample belongs to a certain class (for example, the probability that it belongs to the positive class). In the training process, the parameters θ are optimized by maximum likelihood estimation and other methods to minimize the difference between the predicted probability and the true label, so as to build a logistic regression model structure suitable for speech classification and identification tasks.
[0058] S120: processing the speech data based on the time domain feature and obtaining a time domain feature dataset of the first speech feature vector; inputting the time domain feature dataset of the first speech feature vector into the companion robot, and obtaining the wake-up rate corresponding to each first speech feature vector in the time domain feature dataset of the first speech feature vector; and screening the first speech feature vector with a wake-up rate higher than a preset first wake-up rate threshold.
[0059] It should be noted that the time domain features are directly extracted from the waveform of the speech signal without frequency spectrum analysis. They reflect the instantaneous characteristics or changes of the signal. The present application adjusts the speech data from the aspects of short-time energy and zero-crossing rate. That is, the time domain feature data set formed after adjusting the speech data includes the first speech feature vector adjusted by the short-time energy feature parameter and the first speech feature vector adjusted by the zero-crossing rate feature parameter.
[0060] Specifically, the time domain feature data set of the first speech feature vector includes the first speech feature vector adjusted by the short-time energy feature parameter and the first speech feature vector adjusted by the zero-crossing rate feature parameter.
[0061] In the specific implementation process, the first speech feature vector adjusted by the short-time energy feature parameter includes the first speech feature vector adjusted by the frame size parameter, the first speech feature vector adjusted by the frame shift parameter, the first speech feature vector adjusted by the smoothing window size parameter and the first speech feature vector adjusted by the pre-emphasis coefficient parameter. The first speech feature vector adjusted by the zero-crossing rate feature parameter includes the first speech feature vector adjusted by the frame size feature parameter, the first speech feature vector adjusted by the frame shift feature parameter, the first speech feature vector adjusted by the window function feature parameter and the first speech feature vector adjusted by the high-pass filter cutoff frequency feature parameter. Among them, the first speech feature vector adjusted by the frame size feature parameter in the first speech feature vector adjusted by the short-time energy feature parameter and the first speech feature vector adjusted by the zero-crossing rate feature parameter is the same as the first speech feature vector adjusted by the frame shift feature parameter.
[0062] In one specific embodiment, for the short-time energy feature parameter adjustment process, the speech data is respectively adjusted by the frame size parameter according to 10ms, 15ms and 20ms to obtain the first speech feature vector adjusted by the frame size parameter; the speech data is respectively adjusted by the frame shift parameter according to 60%, 50% and 40% to obtain the first speech feature vector adjusted by the frame shift parameter; the speech data is respectively adjusted by the smoothing window size parameter according to 20ms, 25ms and 30ms to obtain the first speech feature vector adjusted by the smoothing window size parameter; and the speech data is respectively adjusted by the pre-emphasis coefficient parameter according to 0.95, 0.96 and 0.97 to obtain the first speech feature vector adjusted by the pre-emphasis coefficient parameter.
[0063] For the zero-crossing rate feature parameter adjustment process, the speech data is respectively adjusted in frame size parameters of 10 ms, 15 ms, and 20 ms to obtain first speech feature vectors adjusted in frame size parameters; the speech data is respectively adjusted in frame shift parameters of 60%, 50%, and 40% to obtain first speech feature vectors adjusted in frame shift parameters; the speech data is respectively adjusted in window function feature parameters of Hamming window, Hanning window, and rectangular window to obtain first speech feature vectors adjusted in window function feature parameters; and the speech data is respectively adjusted in high-pass filter cutoff frequency feature parameters of 80 Hz, 90 Hz, and 100 Hz to obtain first speech feature vectors adjusted in high-pass filter cutoff frequency feature parameters. The first speech feature vectors adjusted in short-time energy feature parameters and the first speech feature vectors adjusted in zero-crossing rate feature parameters obtained after the time-domain parameter adjustment of one speech data are used to form a time-domain feature data set, and the frame size parameters and the frame shift parameters in the short-time energy feature parameters and the zero-crossing rate feature parameters are kept consistent, a orthogonal experiment is designed, and 729 feature vectors are obtained.
[0064] The time-domain feature data set of the first speech feature vectors is input into a companion robot, and the wake-up rates corresponding to each first speech feature vector in the time-domain feature data set are obtained; and the first speech feature vectors with a wake-up rate higher than a preset first wake-up rate threshold are screened. In a specific implementation process, the first wake-up rate threshold is 95%.
[0065] In the implementation process, the accuracy can be obtained by the classification algorithm of support vector machine. Specifically, support vector machine (SVM): SVM is a supervised learning classification algorithm. Its basic idea is to find an optimal hyperplane in the feature space to separate different classes of data points. For the speech frequency domain feature data set, the known accuracy speech feature vector is used as the training sample to train the SVM model. For example, in the binary classification problem (such as accurate and inaccurate), SVM can find a hyperplane according to the frequency domain features of the sample, so that the interval between the two classes of samples is maximum. In the prediction stage, the speech feature vector to be tested is input into the trained SVM model, and the model outputs the class corresponding to the feature vector (and then the accuracy related information can be obtained). The implementation steps of the classification algorithm of support vector machine include collecting labeled data: collecting speech frequency domain feature vectors with known accuracy as training data, including feature vectors and corresponding accuracy labels (for example, accuracy of 80% is marked as 1, and accuracy lower than 80% is marked as 0). At the same time, the training set, the validation set and the test set are divided, usually according to a certain proportion (such as 70% training set, 10% validation set and 20% test set). Data normalization: normalize the speech frequency domain feature vector, so that features of different dimensions have the same scale. Common normalization methods include minimum-maximum normalization, which maps feature values to the [0, 1] interval, and the formula is where x is the original feature value, x min and x max are the minimum and maximum values of the feature dimension. Select kernel function: SVM has multiple kernel functions to choose from, such as linear kernel, polynomial kernel, Gaussian kernel (RBF kernel), etc. For speech frequency domain feature data, select the appropriate kernel function according to the distribution characteristics of the data. If the data may be linearly separable in the feature space, linear kernel may be a good choice; if the data distribution is complex, Gaussian kernel is usually more appropriate. Train the model: use the training set data to determine the parameters of the SVM model, including support vectors and hyperplane parameters, through optimization algorithms (such as sequential minimal optimization algorithm, SMO). In the training process, adjust the model by minimizing the loss function (such as hinge loss function) so that the model can accurately classify the training data. At the same time, use the validation set to adjust the hyperparameters of the model (such as kernel function parameters, penalty coefficient, etc.) to prevent overfitting. Model evaluation: evaluate the trained SVM model using the test set, and common evaluation indicators include accuracy, recall, F1-score, etc. The accuracy calculation formula is where TP is true positive (correctly classified as positive), TN is true negative (correctly classified as negative), FP is false positive (incorrectly classified as positive), and FN is false negative (incorrectly classified as negative).
[0066] Application model: input the voice frequency domain feature data set to be tested into the trained SVM model, the model will output the category corresponding to each voice feature vector (according to the category, the accuracy related information can be inferred). For example, if the output category is 1, it means that the accuracy of the voice feature vector may be high; if the output category is 0, it means that the accuracy may be low.
[0067] S130: processing the screened first voice feature vector based on the frequency domain feature and obtaining the frequency domain feature data set of the second voice feature vector; inputting the frequency domain feature data set of the second voice feature vector into the companion robot, obtaining the accuracy corresponding to each voice feature vector in the frequency domain feature data set of the second voice feature vector; screening the second voice feature vector with accuracy higher than the preset second accuracy threshold.
[0068] In the specific implementation process, the suitable time domain feature can guarantee the large model wake-up rate, and the frequency domain feature can guarantee the accuracy of the large model identification. The frequency domain feature describes the frequency spectrum information of the signal, which is usually extracted by short-time Fourier transform (STFT) or other spectrum analysis methods. The present application will adjust from two directions of mel frequency cepstral coefficient (MFCC) and spectral centroid (Spectral Centroid).
[0069] Specifically, the frequency domain feature data set of the second voice feature vector includes the second voice feature vector after adjusting the mel frequency cepstral coefficient feature parameter and the second voice feature vector after adjusting the spectral centroid feature parameter.
[0070] The second voice feature vector after adjusting the mel frequency cepstral coefficient feature parameter includes the second voice feature vector after adjusting the range feature parameter of mel frequency, the second voice feature vector after adjusting the DCT coefficient number feature parameter, the second voice feature vector after adjusting the pre-emphasis feature parameter and the second voice feature vector after adjusting the logarithmic operation feature parameter. The second voice feature vector after adjusting the spectral centroid feature parameter includes the second voice feature vector after adjusting the spectral resolution feature parameter, the second voice feature vector after adjusting the spectral smoothing feature parameter, the second voice feature vector after adjusting the window function feature parameter and the second voice feature vector after adjusting the self-defined weighting feature parameter.
[0071] In one specific embodiment, for the mel-frequency cepstral coefficient feature parameter adjustment process, the first speech feature vector is adjusted in the mel-frequency range of 20 Hz-8000 Hz, 0 Hz-4000 Hz, and 50 Hz-4000 Hz to obtain the second speech feature vector after mel-frequency range feature parameter adjustment; the first speech feature vector is adjusted in the number of DCT coefficients of 12, 13, and 20 to obtain the second speech feature vector after DCT coefficient number feature parameter adjustment; the first speech feature vector is adjusted in the pre-emphasis of 0.97, 0.9, and 1.0 to obtain the second speech feature vector after pre-emphasis feature parameter adjustment; the first speech feature vector is adjusted in the logarithmic operation of logarithmic operation, logarithmic operation plus smoothing (log+eps), and no logarithmic operation to obtain the second speech feature vector after logarithmic operation feature parameter adjustment.
[0072] For the spectral centroid feature parameter adjustment process, the first speech feature vector is adjusted in the spectral resolution of 512-point FFT, 1024-point FFT, and 2048-point FFT to obtain the second speech feature vector after spectral resolution feature parameter adjustment; the first speech feature vector is adjusted in the frame shift of 60%, 50%, and 40% to obtain the first speech feature vector after frame shift parameter adjustment; the first speech feature vector is adjusted in the spectral smoothing of moving average, Gaussian smoothing, and low-pass filtering to obtain the second speech feature vector after spectral smoothing feature parameter adjustment; the first speech feature vector is adjusted in the self-defined weighting of increasing the weight of 100 Hz to 500 Hz, increasing the weight of 500 Hz to 1500 Hz, and increasing the weight above 2 kHz to obtain the second speech feature vector after self-defined weighting feature parameter adjustment. The first speech feature vector is adjusted in the window function of Hamming window, Hanning window, and rectangular window to obtain the second speech feature vector after window function feature parameter adjustment.
[0073] The second speech feature vector after mel-frequency cepstral coefficient feature parameter adjustment and the second speech feature vector after spectral centroid feature parameter adjustment, which are obtained after the first speech feature vector is adjusted in the frequency domain, are adjusted in the frame size parameter and the frame shift parameter in the short-time energy feature parameter and the zero-crossing rate feature parameter to keep consistent when forming the frequency domain feature data set. Orthogonal experiments are designed to obtain 6561 feature vectors.
[0074] Input the frequency domain feature data set of the second speech feature vector into the companion robot, and obtain the accuracy corresponding to each second speech feature vector in the frequency domain feature data set; and screen the second speech feature vector with an accuracy higher than a preset second accuracy threshold. In a specific implementation process, the second accuracy threshold is 95%. In a specific implementation process, the accuracy can be obtained by using a support vector machine classification algorithm.
[0075] S140: Determine the speech preference feature of the companion robot according to the second speech feature vector.
[0076] In one specific embodiment, the following processing is further included to improve the speech recognition accuracy of the companion robot.
[0077] S150: Perform feature alignment processing on the speech data to be recognized according to the speech preference feature of the companion robot.
[0078] In a specific implementation process, it can be implemented by using a dynamic time warping (DTW) algorithm, which is an algorithm for measuring the similarity of two time series. Its basic principle is to construct a distance matrix, and the elements in the matrix represent the distance of the speech features at two time points. Then find a path with the minimum cumulative distance by dynamic programming algorithm, and this path represents the best alignment of the two speech sequences. The specific implementation process is exemplarily illustrated as follows: distance matrix calculation: calculate the distance matrix between two speech feature sequences (speech feature sequence to be recognized and robot speech preference feature sequence). Assuming that the speech feature sequence to be recognized is X=[x1, x2, …, xn], and the robot speech preference feature sequence is Y=[y1, y2, …, yn], then the element d m in the distance matrix D is calculated as follows: m ij The Euclidean distance or other methods can be used for calculation, such as where l is the dimension of the feature vector. Next, dynamic programming is used to find the optimal path. Initialize a cumulative distance matrix C with a size of (m+1)×(n+1). Set the boundary conditions C[0,0]=0, C[i,0]=∞ (i>0), and C[0,j]=∞ (j>0). Then, the optimal path is found by using the dynamic programming formula C[i,j]=d ij The cumulative distance matrix is filled with `+min(C[i-1, j-1], C[i-1, j], C[i, j-1])`. Finally, starting from the bottom right corner of the matrix, the path with the minimum cumulative distance is found; this path is the alignment path between the two speech feature sequences. Based on the found alignment path, the speech features are aligned. For example, if the alignment path indicates that the i-th feature in the speech feature sequence to be recognized should be aligned with the j-th feature in the robot's speech preference feature sequence, then these two features are matched or further processed, such as by weighted averaging, to obtain the aligned speech features.
[0079] S160: Perform speech recognition on the speech data after feature alignment processing.
[0080] like Figure 2 As shown, this invention provides a companion robot voice recognition system, which utilizes the companion robot voice recognition method described above for voice recognition. Depending on the functions implemented, the companion robot voice recognition system 200 may include an acquisition unit 210, a first filtering unit 220, a second filtering unit 230, and a feature determination unit 240. The unit of this invention can also be referred to as a module, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device.
[0081] In this embodiment, the functions of each module / unit are as follows:
[0082] Acquisition unit 210 is used to acquire voice data.
[0083] The first filtering unit 220 is used to process the speech data based on the time-domain features and obtain the time-domain feature dataset of the first speech feature vector; input the time-domain feature dataset of the first speech feature vector into the companion robot and obtain the wake-up rate corresponding to each first speech feature vector in the time-domain feature dataset of the first speech feature vector; and filter the first speech feature vectors whose wake-up rate is higher than the preset first wake-up rate threshold.
[0084] The second filtering unit 230 is used to process the first speech feature vector selected based on frequency domain features and obtain the frequency domain feature dataset of the second speech feature vector; input the frequency domain feature dataset of the second speech feature vector into the companion robot, obtain the accuracy corresponding to each speech feature vector in the frequency domain feature dataset of the second speech feature vector; and filter the second speech feature vector with an accuracy higher than the preset second accuracy threshold.
[0085] The feature determination unit 240 is used to determine the voice preference features of the companion robot based on the second voice feature vector.
[0086] The companion robot voice recognition system of the present application processes voice data based on short-time energy and zero-crossing rate to obtain a time-domain feature data set, and obtains a wake-up rate corresponding to each first voice feature vector in the time-domain feature data set, screens first voice feature vectors with a wake-up rate higher than a preset first wake-up rate threshold, processes the screened first voice feature vectors based on mel-frequency cepstral coefficients and spectral centroids to obtain a frequency-domain feature data set, obtains an accuracy rate corresponding to each voice feature vector in the frequency-domain feature data set, screens second voice feature vectors with an accuracy rate higher than a preset second accuracy rate threshold, and determines a voice preference feature of the companion robot according to the second voice feature vectors. Through the screening mechanism based on the wake-up rate and the accuracy rate, the most effective voice feature for the companion robot can be identified, so as to improve the response speed and accuracy of voice recognition according to the most effective voice feature, and finally achieve the technical effects of effectively improving the voice recognition accuracy of the companion robot and improving the user interaction experience.
[0087] More specific implementation modes of the above-mentioned companion robot voice recognition system can be referred to the description of the embodiments of the companion robot voice recognition method, which will not be described one by one here.
[0088] As shown in Figure 3 The present application also provides an electronic device 1 for wearing detection.
[0089] The electronic device 1 can include a processor 10, a memory 11 and a bus, and can also include a computer program stored in the memory 11 and executable on the processor 10, such as a companion robot voice recognition program 12. The memory 11 can include not only an internal storage unit of the companion robot voice recognition system but also an external storage device. The memory 11 can be used not only to store installed application software and various data, such as the code of the companion robot voice recognition program, but also to temporarily store data that has been output or will be output.
[0090] The electronic device 1 can include a processor 10, a memory 11 and a bus, and can also include a computer program stored in the memory 11 and executable on the processor 10, such as a companion robot voice recognition program 12. The memory 11 can include not only an internal storage unit of the companion robot voice recognition system but also an external storage device. The memory 11 can be used not only to store installed application software and various data, such as the code of the companion robot voice recognition program, but also to temporarily store data that has been output or will be output.
[0091] The memory 11 includes at least one type of readable storage medium, such as a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as a mobile hard disk of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 11 can include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used to store application software and various data installed in the electronic device 1, such as the code of the companion robot voice recognition program, and can also be used to temporarily store data that has been output or will be output.
[0092] The processor 10 can be composed of an integrated circuit in some embodiments, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more combinations of a central processing unit (CPU), a microprocessor, a digital processing chip, a graphics processor, and various control chips, etc. The processor 10 is the control unit of the electronic device, which connects various components of the entire electronic device through various interfaces and lines, executes programs or modules stored in the memory 11 (such as the companion robot voice recognition program, etc.), and calls data stored in the memory 11 to perform various functions and process data of the electronic device 1.
[0093] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.
[0094] Figure 3 Only the electronic device with components is shown, and those skilled in the art can understand that, Figure 3The illustrated structure does not constitute a limitation on the electronic device 1, and can include fewer or more components than illustrated, or combine certain components, or arrange different components.
[0095] For example, although not shown, the electronic device 1 can also include a power source (such as a battery) to power the various components, and preferably the power source can be logically connected to the at least one processor 10 through a power management system, so that the power management system can implement functions such as charge management, discharge management, and power consumption management. The power source can also include one or more DC or AC power sources, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and any other components. The electronic device 1 can also include various sensors, Bluetooth modules, Wi-Fi modules, and the like, which are not described here.
[0096] Further, the electronic device 1 can also include a network interface, which can optionally include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, and the like), and is typically used to establish a communication connection between the electronic device 1 and other electronic devices.
[0097] Optionally, the electronic device 1 can also include a user interface, which can be a display (Display), an input unit (such as a keyboard (Keyboard)), and optionally can also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch, and the like. The display can also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the electronic device 1 and to display a visualized user interface.
[0098] It should be understood that the embodiments are only for illustration and are not limited in the scope of the patent application by this structure.
[0099] The companion robot voice recognition program 12 stored in the memory 11 in the electronic device 1 is a combination of a plurality of instructions, which, when running in the processor 10, can achieve: obtaining voice data; processing the voice data based on time domain features and obtaining a time domain feature data set of a first voice feature vector; inputting the time domain feature data set of the first voice feature vector into a companion robot and obtaining a wake-up rate corresponding to each first voice feature vector in the time domain feature data set of the first voice feature vector; screening the first voice feature vector with a wake-up rate higher than a preset first wake-up rate threshold; processing the screened first voice feature vector based on frequency domain features and obtaining a frequency domain feature data set of a second voice feature vector; inputting the frequency domain feature data set of the second voice feature vector into the companion robot and obtaining an accuracy rate corresponding to each voice feature vector in the frequency domain feature data set of the second voice feature vector; screening the second voice feature vector with an accuracy rate higher than a preset second accuracy rate threshold; and determining a voice preference feature of the companion robot according to the second voice feature vector.
[0100] Specifically, the processor 10 can refer to the specific implementation method of the above instructions Figure 1 The description of related steps in the corresponding embodiments is not repeated. Further, the modules / units integrated in the electronic device 1 can be stored in a computer readable storage medium if they are realized in the form of software function units and sold or used as independent products. The computer readable medium can include any entity or system capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM).
[0101] The embodiment of the present application further provides a computer readable storage medium, which can be nonvolatile or volatile, and stores a computer program. The computer program is executed by a processor to realize the following steps: obtaining voice data; processing the voice data based on time domain features to obtain a time domain feature dataset of a first voice feature vector; inputting the time domain feature dataset of the first voice feature vector into a companion robot, and obtaining a wake-up rate corresponding to each first voice feature vector in the time domain feature dataset of the first voice feature vector; screening a first voice feature vector with a wake-up rate higher than a preset first wake-up rate threshold; processing the screened first voice feature vector based on frequency domain features to obtain a frequency domain feature dataset of a second voice feature vector; inputting the frequency domain feature dataset of the second voice feature vector into the companion robot, and obtaining an accuracy rate corresponding to each voice feature vector in the frequency domain feature dataset of the second voice feature vector; screening a second voice feature vector with an accuracy rate higher than a preset second accuracy rate threshold; and determining a voice preference feature of the companion robot according to the second voice feature vector.
[0102] Specifically, the computer program executed by the processor specifically implements the method, which can refer to the description of the related steps in the embodiment of the wearing detection method, and details are not described herein.
[0103] In the several embodiments of the present application, it should be understood that the disclosed device, system and method can be implemented in other ways. For example, the above-mentioned system embodiments are merely schematic, and the division of the modules is merely a logical function division, and another division mode can be used in actual implementation.
[0104] The modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.
[0105] In addition, each functional module in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of hardware plus software functional modules.
[0106] It is obvious for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.
[0107] Therefore, embodiments should be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are intended to be embraced therein. No reference is intended to be made to any disclaimer or disclaimer to limit the claims notwithstanding anything to the contrary contained in this patent or any patent accompanying drawing.
[0108] Furthermore, the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural and vice-versa, unless the context clearly requires these exclusions. The composition of a system claimed in a system claim does not preclude the use of other units or systems in addition to or instead of those claimed.
[0109] The accompanying drawings are referred to in illustrating the companion robot voice recognition method and the companion robot voice recognition system according to the present application as described above. However, it should be understood by those skilled in the art that various modifications can be made to the above-mentioned companion robot voice recognition method and the companion robot voice recognition system according to the present application without departing from the scope of the present application. Therefore, the scope of the present application should be determined by the appended claims.
Claims
1. A companion robot voice recognition method applied to an electronic device, the method comprising: The method comprises: acquiring voice data; processing the voice data based on time domain features and obtaining a time domain feature dataset of first voice feature vectors; inputting the time domain feature dataset of the first voice feature vectors into a companion robot, and acquiring a wake-up rate corresponding to each first voice feature vector in the time domain feature dataset of the first voice feature vectors; screening first voice feature vectors with a wake-up rate higher than a preset first wake-up rate threshold; the time domain feature dataset of the first voice feature vectors comprises first voice feature vectors with adjusted short-time energy feature parameters and first voice feature vectors with adjusted zero-crossing rate feature parameters; processing the screened first voice feature vectors based on frequency domain features and obtaining a frequency domain feature dataset of second voice feature vectors; inputting the frequency domain feature dataset of the second voice feature vectors into the companion robot, and acquiring an accuracy rate corresponding to each voice feature vector in the frequency domain feature dataset of the second voice feature vectors; screening second voice feature vectors with an accuracy rate higher than a preset second accuracy rate threshold; the frequency domain feature dataset of the second voice feature vectors comprises second voice feature vectors with adjusted mel-frequency cepstral coefficient feature parameters and second voice feature vectors with adjusted spectral centroid feature parameters; determining voice preference features of the companion robot according to the second voice feature vectors; performing voice recognition according to the voice preference features of the companion robot. 2.The companion robot voice recognition method of claim 1, wherein, The first voice feature vectors with adjusted short-time energy feature parameters comprise first voice feature vectors with adjusted frame size parameters, first voice feature vectors with adjusted frame shift parameters, first voice feature vectors with adjusted smoothing window size parameters, and first voice feature vectors with adjusted pre-emphasis coefficient parameters; The first voice feature vectors with adjusted zero-crossing rate feature parameters comprise first voice feature vectors with adjusted frame size feature parameters, first voice feature vectors with adjusted frame shift feature parameters, first voice feature vectors with adjusted window function feature parameters, and first voice feature vectors with adjusted high-pass filter cutoff frequency feature parameters; The first voice feature vectors with adjusted frame size feature parameters and the first voice feature vectors with adjusted frame shift feature parameters in the first voice feature vectors with adjusted short-time energy feature parameters and the first voice feature vectors with adjusted zero-crossing rate feature parameters are the same. 3.The companion robot voice recognition method of claim 1, wherein, The second voice feature vectors with adjusted mel-frequency cepstral coefficient feature parameters comprise second voice feature vectors with adjusted range feature parameters of mel-frequency, second voice feature vectors with adjusted DCT coefficient number feature parameters, second voice feature vectors with adjusted pre-emphasis feature parameters, and second voice feature vectors with adjusted logarithm operation feature parameters.
4. The companion robot voice recognition method according to claim 1, wherein The second speech feature vector after adjustment of the spectral centroid feature parameter includes a second speech feature vector after adjustment of a spectral resolution feature parameter, a second speech feature vector after adjustment of a spectral smoothing feature parameter, a second speech feature vector after adjustment of a window function feature parameter, and a second speech feature vector after adjustment of a user-defined weighting feature parameter. 5.The companion robot voice recognition method of claim 1, wherein, The first wake-up rate threshold and the second accuracy rate threshold are both 95%. 6.The companion robot voice recognition method of claim 1, wherein, Further comprising, performing feature alignment processing on the voice data to be recognized according to the voice preference feature of the companion robot; Performing voice recognition on the voice data after the feature alignment processing. 7.The companion robot voice recognition method of claim 1, wherein, Before processing the voice data based on time domain features, further comprising, Performing noise reduction preprocessing on the voice data using a Wiener filter model. 8.A companion robot voice recognition system, which performs voice recognition using the companion robot voice recognition method of any one of claims 1-7; comprising: an acquisition unit configured to acquire voice data; a first screening unit configured to process the voice data based on time domain features and obtain a time domain feature dataset of first speech feature vectors; input the time domain feature dataset of the first speech feature vectors into a companion robot, and acquire a wake-up rate corresponding to each first speech feature vector in the time domain feature dataset of the first speech feature vectors; screen first speech feature vectors with a wake-up rate higher than a preset first wake-up rate threshold; the time domain feature dataset of the first speech feature vectors includes first speech feature vectors after adjustment of a short-time energy feature parameter and first speech feature vectors after adjustment of a zero-crossing rate feature parameter; a second screening unit configured to process the screened first speech feature vectors based on frequency domain features and obtain a frequency domain feature dataset of second speech feature vectors; input the frequency domain feature dataset of the second speech feature vectors into the companion robot, and acquire an accuracy rate corresponding to each speech feature vector in the frequency domain feature dataset of the second speech feature vectors; screen second speech feature vectors with an accuracy rate higher than a preset second accuracy rate threshold; the frequency domain feature dataset of the second speech feature vectors includes second speech feature vectors after adjustment of a mel-frequency cepstrum coefficient feature parameter and second speech feature vectors after adjustment of a spectral centroid feature parameter; a feature determination unit configured to determine a voice preference feature of the companion robot according to the second speech feature vectors; The voice recognition system performs voice recognition according to the voice preference feature of the companion robot. 9.An electronic device, comprising: The electronic device includes a memory, a processor, and a companion robot voice recognition program stored on the memory and executable on the processor, and the companion robot voice recognition program implements the steps of the voice recognition method of any one of claims 1-7 when executed by the processor.
10. A computer readable storage medium storing a computer program, characterized in that, The computer program implements the voice recognition method of any one of claims 1-7 when executed by the processor.
Citation Information
Patent Citations
Speech emotion recognition method
CN118866014A
Channel selector arrangement
US5457711A