Voice interaction method and system for virtual reality training platform
By performing frame-to-window processing of voice signals and optimizing the MFCC coefficient weight, the speech recognition accuracy problem of traditional MFCC algorithms under noise interference is solved, and the speech interaction accuracy of the virtual reality training platform is improved.
Patent Information
- Application Number
- CN202510493497.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Traditional MFCC algorithms are disturbed by background noise during speech recognition, resulting in low speech recognition accuracy and affecting the accuracy of speech interaction.
By performing frame-windowing of speech signals, the stationary index and semantic richness are obtained, combined with the resonance noise degree of the formant peak, the weight factor of the MFCC coefficient is optimized using the particle swarm algorithm, and weighted processing is performed to improve speech recognition accuracy.
It reduces the noise error detection phenomenon caused by semantic features, improves the accuracy of speech recognition, and improves the speech interaction accuracy of the virtual reality training platform.
Smart Images

Figure CN120279908A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of voice interaction technology, and in particular, to a voice interaction method and system for a virtual reality training platform. Background Art
[0002] A virtual reality training platform is a training system that generates a highly immersive three-dimensional environment through a computer, combines interactive devices and artificial intelligence technology, and simulates real or fictional scenarios. Its core goal is to provide users with a safe, repeatable, and quantifiable training experience. In the voice interaction of a virtual reality training platform, it is usually divided into steps such as voice capture, preprocessing, recognition, semantic understanding, instruction execution, and feedback. Among them, the accuracy and efficiency of voice recognition directly affect the effect of voice interaction.
[0003] In the field of voice recognition, acoustic feature extraction is a crucial step in a voice recognition system, which will directly affect the performance and accuracy of the voice recognition system. The traditional acoustic feature extraction method is the MFCC algorithm, which has good stability and discrimination and is widely used in voice recognition systems. The core idea of the MFCC algorithm is to simulate the non-linear perception of sound frequency by the human ear and extract the short-time spectral features of speech through cepstrum analysis. However, when the MFCC algorithm performs feature extraction, due to its dependence on spectral energy, background noise in the speech signal will distort the energy distribution of the filter bank, resulting in distorted MFCC coefficient features being extracted, making the accuracy of voice recognition low, and thus affecting the accuracy of voice interaction. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of this application is to provide a voice interaction method and system for a virtual reality training platform, and the specific technical solutions adopted are as follows: In the first aspect, an embodiment of this application provides a voice interaction method for a virtual reality training platform, and the method includes the following steps: Obtain a voice signal input by a user; Perform frame addition and windowing processing on the voice signal; obtain the smoothness index of each frame of voice signal according to the signal intensity difference between adjacent signal points in each frame of voice signal; obtain the pause interval of the entire voice signal according to the position of the sentence turning point in the entire voice signal; obtain the semantic richness of each frame of voice signal according to the signal intensity difference between the signal points equidistant from the nearest pause interval in each frame of voice signal, and the average signal intensity level and length of the signal points belonging to the pause interval; obtain the optimized smoothness index of each frame of voice signal according to the smoothness index and semantic richness of each frame of voice signal; Obtain the formants in the spectrograms of each frame of speech signals; obtain the formant noise degrees of each formant according to the number of formants, semantic richness in the spectrograms of the speech signals where each formant is located, and the signal amplitude differences between each formant and each frequency component within its neighborhood. Obtain multiple MFCC coefficients of each frame of speech signal and the corresponding frequency intervals of each MFCC coefficient in the spectrogram; obtain the weight factors of each MFCC coefficient of a single-frame speech signal according to the formant noise degree of the formant closest to the center point of the frequency interval corresponding to each MFCC coefficient of the single-frame speech signal, the distance between the center point and the closest formant, and the optimized stationary index of this frame of speech signal, so as to weight each MFCC coefficient, and further complete the voice interaction of the virtual reality training platform.
[0005] Preferably, the calculation formula for the stationary index of each frame of speech signal is: ; in the formula, represents the stationary index of the i-th frame of speech signal, represents the total number of signal points in the i-th frame of speech signal, represents the signal intensity of the k-th signal point in the i-th frame of speech signal, represents the signal intensity of the (k + 1)-th signal point in the i-th frame of speech signal, is the first preset constant.
[0006] Preferably, the process of obtaining the pause intervals of the entire speech signal is as follows: Take the ratio of the maximum value of the absolute value of the intensity difference between a single signal point and the signal points at its adjacent moments on the left and right to the intensity of this signal point as the turning degree of this signal point; Cluster all the signal points in the entire speech signal according to the turning degree of each signal point, divide the signal points in the speech signal into several clustering clusters, and mark the signal points in the clustering cluster with the largest average turning degree as the sentence turning points; Mark the intervals formed by the speech signals between each sentence turning point and the closest sentence turning point in the entire speech signal as the pause intervals of the entire speech signal.
[0007] Preferably, the calculation formula for the semantic richness of each frame of speech signal is: ; in the formula, is the semantic richness of the i-th frame of speech signal, represents the total number of signal points except for the pause intervals in the i-th frame of speech signal, represents the signal intensity of the a-th signal point except for the pause intervals in the i-th frame of speech signal, then represents the signal intensity of the corresponding signal point of the a-th signal point except for the pause intervals in the i-th frame of speech signal; represents the average signal strength of the signal points within the pause interval in the i-th frame of the speech signal; represents the length of the signal points within the pause interval in the i-th frame of the speech signal; is the second preset constant; wherein, the corresponding signal point of the a-th signal point refers to the signal point whose temporal distance to the center point of the nearest pause interval is equal to the temporal distance of the a-th signal point to the center point of the nearest pause interval.
[0008] Preferably, the optimized stationary index of each frame of speech signal is the positive fusion result of the stationary index of each frame of speech signal and the semantic richness.
[0009] Preferably, the calculation formula for the resonance noise degree of each formant is: ; in the formula, represents the resonance noise degree of the j-th formant; n represents the total number of formants in the spectrogram where the j-th formant is located; is the semantic richness of the speech signal to which the spectrogram where the j-th formant is located belongs; R represents the number of frequency components between the two nearest minimum value points on both sides of the j-th formant; represents the signal amplitude of the r-th frequency component between the two nearest minimum value points on both sides of the j-th formant; represents the signal amplitude of the j-th formant.
[0010] Preferably, the method for obtaining the frequency interval corresponding to each MFCC coefficient in the spectrogram of the single-frame speech signal is: equally dividing the spectrogram of the single-frame speech signal into z frequency intervals, and obtaining the frequency interval corresponding to each MFCC coefficient through the particle swarm optimization algorithm, where each particle in the algorithm corresponds to z dimensions, each dimension represents each frequency interval, and the fitness function is the sum of the ratios of the sum of the absolute values of the intensity differences between the frequency components in all frequency intervals and the corresponding MFCC coefficients, so as to obtain the frequency interval corresponding to each MFCC coefficient in the spectrogram.
[0011] Preferably, the calculation formula for the weight factor of each MFCC coefficient of the single-frame speech signal is: ; in the formula, is the weight factor of the m-th MFCC coefficient of the i-th frame of the speech signal, is the optimized stationary index of the i-th frame of the speech signal, represents the resonance noise degree of the formant nearest to the center point of the frequency interval corresponding to the m-th MFCC coefficient of the i-th frame of the speech signal, represents the distance between the center point of the frequency interval corresponding to the m-th MFCC coefficient of the i-th frame of the speech signal and the nearest formant.
[0012] Preferably, the specific process of completing the voice interaction of the virtual reality training platform is as follows: taking the weighted MFCC coefficients as inputs, completing voice recognition through a hidden Markov model; converting the result of voice recognition into text information through a language model, and completing semantic understanding through natural language processing technology, and completing the execution and feedback of corresponding instructions through the AI behavior model of the virtual reality training platform.
[0013] In a second aspect, an embodiment of the present application further provides a voice interaction system for a virtual reality training platform, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the voice interaction method for a virtual reality training platform described in any one of the above are implemented.
[0014] The present application has at least the following beneficial effects: In the present application, the voice signal is subjected to frame windowing processing to divide the voice signal into multiple frame voice signals. A stationary index is constructed through the change of the signal intensity in each frame language signal. At the same time, the sentence turning points in the voice signal are obtained through the sentence features when the user inputs the voice signal. Furthermore, the semantic richness is constructed through the change of the voice signal before and after the sentence turning and the voice features in the pause interval to optimize the stationary index, reducing the phenomenon of noise misdetection caused by semantic features; at the same time, a resonance noise degree is constructed according to the local change characteristics of the formants in the spectrogram and the matching condition between the semantic richness and the formants, and the frequency interval corresponding to each MFCC coefficient is obtained through the particle swarm algorithm. The weight factor of the MFCC coefficient is constructed through the optimized stationary index and resonance noise degree, so that the MFCC coefficient affected by interference has a smaller weight in the voice recognition process, and the MFCC coefficient not affected by interference has a larger weight in the voice recognition process, improving the accuracy of voice recognition, and thus improving the voice interaction accuracy of the virtual reality training platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a flowchart of the steps of a voice interaction method for a virtual reality training platform provided by an embodiment of the present application; Figure 2 It is a flowchart for obtaining the weight factor of each MFCC coefficient in a single-frame voice signal provided by an embodiment of the present application. Detailed Implementation Manner
[0017] In order to further elaborate on the technical means and effects adopted by the present application to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, elaborate in detail on a voice interaction method and system for a virtual reality training platform proposed according to the present application, its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0019] The following will specifically describe the specific solution of a voice interaction method and system for a virtual reality training platform provided by the present application in conjunction with the accompanying drawings.
[0020] Please refer to Figure 1 , which shows a flowchart of the steps of a voice interaction method for a virtual reality training platform provided by an embodiment of the present application. The method includes the following steps: Step 1: Obtain the voice signal input by the user.
[0021] The core components of the virtual reality training platform are divided into a hardware layer, a software layer, and a data layer. Among them, the hardware layer includes: a VR headset (providing visual and auditory immersion), a motion capture system (tracking the user's limb movements), a voice input device (supporting voice interaction), and a haptic feedback device (enhancing the sense of physical reality). The software layer includes: a real-time rendering engine (for constructing a dynamic virtual scene), a physics engine (for simulating the interaction logic of objects), and an AI behavior model (for enhancing the complexity of training). The data layer contains: a user behavior database (recording key indicators such as operation paths and reaction times), and a multi-modal data set (for supporting the analysis of training effects).
[0022] When the user uses the virtual reality training platform, the user inputs a voice command through a microphone, and the system collects the voice signal input by the user through the voice input device in the virtual reality training platform, with a sampling frequency of 16 kHz.
[0023] Thus, the collection of the voice signal of the virtual reality training platform can be completed.
[0024] Step 2: Perform frame segmentation and windowing on the speech signal; obtain the stationary index of each frame of speech signal according to the signal intensity difference between adjacent signal points in each frame of speech signal; obtain the pause interval of the entire speech signal according to the position of the sentence turning points in the entire speech signal; obtain the semantic richness of each frame of speech signal according to the intensity difference between the signal points equidistant from the nearest pause interval in each frame of speech signal, as well as the average signal intensity level and length of the signal points belonging to the pause interval; obtain the optimized stationary index of each frame of speech signal according to the stationary index and semantic richness of each frame of speech signal.
[0025] According to the above steps, the speech signals of the virtual reality training platform are collected. When extracting the features of speech signals, the MFCC algorithm is usually used. The MFCC algorithm is the most commonly used feature extraction method in speech signal processing. Its core idea is to simulate the non-linear perception of sound frequency by the human ear and extract the short-time spectrum features of speech through cepstrum analysis. However, when the MFCC algorithm performs feature extraction, due to its dependence on spectral energy, background noise will distort the energy distribution of the filter bank, resulting in feature distortion, and further affecting the accuracy and efficiency of speech recognition.
[0026] First, perform pre-emphasis processing on the speech signal to enhance the energy of high-frequency components and reduce the energy attenuation of the speech signal in the low-frequency part. Then perform frame segmentation and windowing on the speech signal. The generation of the speech signal comes from the vibration of the user's vocal cords, but this vibration pattern is not stable. However, the speech signal has stationarity in a short time interval. Therefore, in this embodiment, the speech signal is segmented into frames, and the frame length range is set to 25 ms. Then the speech signal input by the user can be divided into multiple frames of signals. After the speech signal is divided into frames, in order to reduce the loss of features caused by the discontinuity at both ends of the signal, a window function is applied to each frame of signal. In this embodiment, the Hamming window is used for windowing. Then perform a fast Fourier transform on each frame of speech signal to obtain the corresponding spectrogram. The abscissa of the spectrogram represents the frequency, and the ordinate represents the intensity of the frequency components.
[0027] According to the above steps, the spectrogram of each frame of speech signal is obtained. By analyzing the speech signal in the time domain and frequency domain, the weight factor of the MFCC coefficients of the speech signal is constructed, so as to perform speech recognition through the MFCC coefficients and the corresponding weight factors, improving the accuracy and efficiency of speech recognition. The specific process is as follows: Here, take the i-th frame of speech signal as an example. First, for the time series signal, when the user issues a speech command, that is, inputs a speech signal, the pronunciation pitch of the user in the same sentence is usually relatively stable and the change range is small. Therefore, the stationary index of each frame of speech signal is constructed through the local changes in the time series of the speech signal to characterize the stationarity of each frame of speech signal. In this embodiment, the stationary index of the i-th frame of speech signal is denoted as , and its specific expression is: ; where, represents the smoothness index of the i-th frame of voice signal, represents the total number of signal points within the i-th frame of voice signal, represents the signal strength of the k-th signal point in the i-th frame of voice signal, represents the signal strength of the (k + 1)-th signal point in the i-th frame of voice signal, is the first preset constant, used to prevent the denominator from being zero, and in this embodiment, it is taken as 0.1. The larger it is, the smaller the possibility that the user is affected by environmental noise when inputting the i-th frame of voice signal.
[0028] According to the above steps, the smoothness index of each frame of voice signal can be obtained. The smoothness index of the voice signal is constructed through the intensity change of the voice signal, reflecting the possibility that the voice signal is affected by environmental noise. However, when the voice signal input by the user itself has certain semantic information, the intensity of the voice signal sent by the user may also change greatly, thus showing specific semantic information. Therefore, in order to prevent the phenomenon that the change in the intensity of the voice signal caused by the semantic characteristics of the voice signal itself is misjudged as noise interference, it is necessary to further analyze the voice signals of the upper and lower frames of each frame of voice signal.
[0029] Still taking the i-th frame of voice signal as an example here. For any signal point in the voice signal, the ratio of the maximum value of the absolute value of the intensity difference between the signal point and the signal points at its adjacent left and right moments (the adjacent moment signal points include two signal points: the signal point at the previous moment and the signal point at the next moment) to the intensity of the signal point is used as the turning degree of the signal point. Among them, the larger the absolute value of the signal intensity difference, and the smaller the signal intensity of the signal point itself, it indicates that the sentence input by the user at this time is more likely to pause or end input at this moment, that is, the corresponding signal point at this time is more likely to be the sentence turning point.
[0030] For each signal point in the entire voice signal, the corresponding turning degree can be obtained in the above way. According to the turning degree of each signal point, one-dimensional k-means mean clustering is performed on all signal points in the entire voice signal, where the number of clustering categories is set to 2, the distance metric is the absolute value of the difference in turning degree, and the initial clustering center points are randomly selected, then the signal points in the voice signal can be divided into two clustering clusters. The clustering process is a well-known method and will not be elaborated here. The signal points in the clustering cluster with the largest average turning degree are recorded as the sentence turning points.
[0031] Since the pauses in the user's voice input are often short pauses, usually two sentence turning points can be detected each time. The intervals formed by the voice signals between each sentence turning point and the nearest sentence turning point in the entire voice signal are recorded as the pause intervals of the entire voice signal. Then, the semantic richness of the i-th frame of voice signal can be obtained: ; where is the semantic richness of the i-th frame of voice signal, represents the total number of signal points in the i-th frame of voice signal except for the pause intervals, represents the signal strength of the a-th signal point in the i-th frame of voice signal except for the pause intervals, then represents the signal strength of the corresponding signal point of the a-th signal point in the i-th frame of voice signal except for the pause intervals; represents the average signal strength of the signal points within the pause interval in the i-th frame of voice signal; represents the length of the signal points within the pause interval in the i-th frame of voice signal; is a second preset constant to prevent the denominator from being zero, and any real number less than 1 and greater than 0.1 can be taken. In this embodiment, 0.5 is taken. Wherein, the corresponding signal point of the a-th signal point refers to the signal point whose temporal distance to the center point of the nearest pause interval is equal to the temporal distance of the a-th signal point to the center point of the nearest pause interval.
[0032] The greater the difference between and indicates that the change amplitude of the voice signal before and after the pause is greater, indicating that the possibility of being interfered by the current noise is smaller, and the change is more likely to be brought by the semantic information itself, that is, the semantic richness is greater.
[0033] The semantic richness corresponding to each frame of voice signal is obtained according to the above steps, and then the steady index can be constrained by the semantic richness to improve the detection accuracy of the steady index of the voice signal. In this embodiment, the optimized steady index of the i-th frame of voice signal is: ; where is the optimized steady index of the i-th frame of voice signal, represents the steady index of the i-th frame of voice signal, is the semantic richness of the i-th frame of voice signal.
[0034] Step 3: Obtain the formants in the spectrograms of each frame of speech signals; obtain the formant noise degree of each formant according to the number of formants, semantic richness in the spectrogram of the speech signal where each formant is located, and the signal amplitude difference between each formant and each frequency component in its neighborhood.
[0035] The semantic richness and the smoothness index mainly reflect the instantaneous changes of the speech signal and do not directly reflect the interference brought by noise to the audio features. When extracting speech features through the MFCC algorithm, the auditory characteristics of the human ear will be simulated, and the key features of the speech signal will be compressed in the frequency domain. The spectral information contains important phonemes, timbres, pitches, formants and other features of the speech, and at the same time, the noise will also present certain spectral characteristics. Therefore, in this embodiment, further analysis is performed on the spectrogram corresponding to each frame of speech signal. Taking the spectrogram of the i-th frame of speech signal as an example, the following analysis is carried out. First, local extreme point detection is performed on the signal amplitudes of all frequency components in the spectrogram of the i-th frame of speech signal to obtain the maximum points and minimum points in the spectrogram of the i-th frame of speech signal, and the obtained maximum points are the formants.
[0036] Then, the formant noise degree of the formant can be constructed through the intensity change between the formant and the local frequency components and the matching condition between the formant and the semantic richness. In this embodiment, the formant noise degree of the j-th formant is denoted as , and its specific expression is: ; In the formula, represents the formant noise degree of the j-th formant; n represents the total number of formants in the spectrogram where the j-th formant is located; is the semantic richness of the speech signal to which the spectrogram where the j-th formant is located belongs; R represents the number of frequency components between the two nearest minimum points on both sides of the j-th formant; represents the signal amplitude of the r-th frequency component between the two nearest minimum points on both sides of the j-th formant; represents the signal amplitude of the j-th formant.
[0037] The greater the formant noise degree of a formant, the greater the possibility that the formant is interfered by noise; conversely, the smaller the possibility that the formant is interfered by noise.
[0038] According to the same steps, the formant noise degree of each formant in the spectrogram corresponding to each frame of speech signal can be obtained.
[0039] Step 4: Obtain multiple MFCC coefficients of each frame of speech signal and the corresponding frequency intervals of each MFCC coefficient in the spectrogram; according to the resonance noise degree of the resonance peak closest to the center point of the frequency interval corresponding to each MFCC coefficient of a single-frame speech signal, the distance between the center point and the closest resonance peak, and the optimized smoothness index of this frame of speech signal, obtain the weight factors of each MFCC coefficient of the single-frame speech signal, which are used to weight each MFCC coefficient, and thus complete the voice interaction of the virtual reality training platform.
[0040] According to the spectrogram of each frame of signal, the acquisition of the MFCC coefficients corresponding to each frame of speech signal is completed through the Mel filter bank. The specific process is a well-known method and will not be elaborated here. After obtaining multiple MFCC coefficients corresponding to each frame of speech signal, in this embodiment, the first 12 MFCC coefficients of each frame of speech signal are selected as the MFCC coefficients of the corresponding frame of speech signal. Among them, different MFCC coefficients reflect different information. Low-order MFCC coefficients reflect low-frequency information characteristics, and high-order MFCC coefficients reflect high-frequency information.
[0041] Furthermore, the frequency component range associated with the MFCC coefficients is obtained through the change of the MFCC coefficients of different frames of speech signals and the change of the intensity of the frequency components in the spectrum, and the weights of each MFCC coefficient are adaptively constructed. First, obtain the frequency range of the i-th frame of speech signal, divide it equally into 12 frequency intervals, and use the particle swarm optimization algorithm to obtain the influence frequency intervals of the 12 MFCC coefficients. Then each particle corresponds to 12 dimensions, and each dimension represents each frequency interval. The fitness function is the sum of the ratios of the sum of the absolute values of the intensity differences between the frequency components in all frequency intervals to the corresponding MFCC coefficients, and the frequency intervals corresponding to each MFCC coefficient in the spectrogram are obtained. Among them, the particle swarm optimization algorithm is a well-known technology, and the specific process will not be elaborated.
[0042] The weight factors of each MFCC coefficient are adaptively constructed according to the noise degree of the frequency components in the frequency intervals corresponding to the MFCC coefficients. The flowchart for obtaining the weight factors of each MFCC coefficient in a single-frame speech signal is as Figure 2 shown. In this embodiment, the weight factor of the m-th MFCC coefficient of the i-th frame of speech signal is denoted as , and its specific expression is: ; in the formula, is the weight factor of the m-th MFCC coefficient of the i-th frame of speech signal, is the optimized smoothness index of the i-th frame of speech signal, represents the resonance noise degree of the resonance peak closest to the center point of the frequency interval corresponding to the m-th MFCC coefficient of the i-th frame of speech signal, It represents the distance between the center point of the frequency interval corresponding to the m-th MFCC coefficient of the i-th frame of voice signal and the nearest resonance peak.
[0043] The larger it is, the smaller the interference received by the m-th MFCC coefficient of the i-th frame of voice signal, and the more beneficial this coefficient is for subsequent speech recognition. Therefore, a larger weight should be assigned; the smaller H is, the greater the interference received by the MFCC coefficient, and the more unfavorable this coefficient is for subsequent speech recognition. Therefore, a smaller weight should be assigned.
[0044] According to the above steps, a corresponding weight factor can be obtained for each MFCC coefficient. The corresponding MFCC coefficients are adaptively weighted using the respective weight factors. The adaptively weighted MFCC coefficients are used as input, and speech recognition is completed through a hidden Markov model. The hidden Markov model is a well-known technology, and the specific process will not be elaborated. The result of speech recognition is converted into text information through a language model, and semantic understanding is completed through natural language processing technology. The AI of the virtual reality training platform executes and gives feedback on the corresponding instructions for the model.
[0045] Thus, the voice interaction of the virtual reality training platform is completed.
[0046] Based on the same inventive concept as the above method, an embodiment of the present application also provides a voice interaction system for a virtual reality training platform, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods for a voice interaction method for a virtual reality training platform.
[0047] It should be noted that: the above sequence of embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specifically describes certain embodiments of this specification. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0048] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
[0049] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present application shall be included within the protection scope of the present application.
Claims
1. A voice interaction method for a virtual reality training platform, characterized in that, The method comprises the following steps: Obtain a voice signal input by a user; Perform frame addition and windowing processing on the voice signal; obtain a stationary index of each frame of voice signal according to the signal intensity difference between adjacent signal points in each frame of voice signal; obtain a pause interval of the entire voice signal according to the position of a sentence turning point in the entire voice signal; obtain the semantic richness of each frame of voice signal according to the intensity difference between signal points equidistant from the nearest pause interval in each frame of voice signal, and the average level and length of the signal intensity of the signal points belonging to the pause interval; obtain an optimized stationary index of each frame of voice signal according to the stationary index and semantic richness of each frame of voice signal; Obtain formants in the spectrogram of each frame of voice signal; obtain the formant noise degree of each formant according to the number of formants, semantic richness in the spectrogram of the voice signal of the frame where each formant is located, and the signal amplitude difference between each formant and each frequency component in its neighborhood; Obtain multiple MFCC coefficients of each frame of voice signal and the frequency interval corresponding to each MFCC coefficient in the spectrogram; obtain a weight factor of each MFCC coefficient of a single frame of voice signal according to the formant noise degree of the formant nearest to the center point of the frequency interval corresponding to each MFCC coefficient of the single frame of voice signal, the distance between the center point and the nearest formant, and the optimized stationary index of the frame of voice signal, so as to weight each MFCC coefficient, thereby completing voice interaction of a virtual reality training platform.
2. The voice interaction method for a virtual reality training platform according to claim 1, wherein, The calculation formula for the stationary index of each frame of voice signal is as follows: ; where represents the stationary index of the i-th frame of voice signal, represents the total number of signal points in the i-th frame of voice signal, represents the signal strength of the k-th signal point in the i-th frame of voice signal, represents the signal strength of the (k + 1)-th signal point in the i-th frame of voice signal, is the first preset constant.
3. A voice interaction method for a virtual reality training platform according to claim 1, characterized in that, The process of obtaining the pause interval of the entire voice signal is as follows: Take the ratio of the maximum value of the absolute value of the intensity difference between a single signal point and the signal points at its adjacent moments on the left and right to the intensity of the signal point as the turning degree of the signal point; Cluster all the signal points in the entire voice signal according to the turning degree of each signal point, divide the signal points in the voice signal into several clustering clusters, and record the signal points in the clustering cluster with the largest average turning degree as the sentence turning points; Record each interval formed by the voice signal between each sentence turning point and the nearest sentence turning point in the entire voice signal as the pause interval of the entire voice signal.
4. A voice interaction method for a virtual reality training platform according to claim 1, characterized in that, The calculation formula for the semantic richness of each frame of speech signal is as follows: ; In the formula, is the semantic richness of the i-th frame of speech signal, represents the total number of signal points in the i-th frame of speech signal except for the pause intervals, represents the signal strength of the a-th signal point in the i-th frame of speech signal except for the pause intervals, then represents the signal strength of the corresponding signal point of the a-th signal point in the i-th frame of speech signal except for the pause intervals; represents the average signal strength of the signal points in the i-th frame of speech signal that belong to the pause intervals; represents the length of the signal points in the i-th frame of speech signal that belong to the pause intervals; is the second preset constant; where the corresponding signal point of the a-th signal point refers to the signal point whose temporal distance to the center point of the nearest pause interval is equal to the temporal distance of the a-th signal point to the center point of the nearest pause interval.
5. A voice interaction method for a virtual reality training platform according to claim 1, characterized in that, The optimized stationary index of each frame of voice signal is the positive fusion result of the stationary index and semantic richness of each frame of voice signal.
6. The voice interaction method for a virtual reality training platform according to claim 1, characterized in that, The calculation formula for the resonance noise level of each resonance peak is as follows: ; In the formula, represents the resonance noise level of the j-th resonance peak; n represents the total number of resonance peaks in the spectrogram where the j-th resonance peak is located; is the semantic richness of the speech signal to which the spectrogram where the j-th resonance peak is located belongs; R represents the number of frequency components between the two nearest minimum points on both sides of the j-th resonance peak; represents the signal amplitude of the r-th frequency component between the two nearest minimum points on both sides of the j-th resonance peak; represents the signal amplitude of the j-th resonance peak.
7. A voice interaction method for a virtual reality training platform according to claim 1, characterized in that The method for obtaining the frequency interval corresponding to each MFCC coefficient in the spectrogram is as follows: equally divide the spectrogram of a single frame of voice signal into z frequency intervals, and obtain the frequency interval corresponding to each MFCC coefficient through a particle swarm optimization algorithm, wherein each particle in the algorithm corresponds to z dimensions, each dimension represents each frequency interval, and the fitness function is the sum of the ratios of the sum of the absolute values of the intensity differences between each frequency component in all frequency intervals to the corresponding MFCC coefficient, so as to obtain the frequency interval corresponding to each MFCC coefficient in the spectrogram.
8. A voice interaction method for a virtual reality training platform according to claim 1, characterized in that, The calculation formula for the weight factor of each MFCC coefficient of the single-frame speech signal is as follows: ; In the formula, is the weight factor of the m-th MFCC coefficient of the i-th frame of speech signal, is the optimized stationary index of the i-th frame of speech signal, represents the resonance noise degree of the resonance peak closest to the center point of the frequency interval corresponding to the m-th MFCC coefficient of the i-th frame of speech signal, represents the distance between the center point of the frequency interval corresponding to the m-th MFCC coefficient of the i-th frame of speech signal and the closest resonance peak.
9. A voice interaction method for a virtual reality training platform according to claim 1, characterized in that, The specific process of completing the voice interaction of the virtual reality training platform is as follows: taking the weighted MFCC coefficients as inputs, voice recognition is completed through a hidden Markov model; the result of voice recognition is converted into text information through a language model, and semantic understanding is completed through natural language processing technology, and the corresponding instructions are executed and feedback is provided through the AI behavior model of the virtual reality training platform.
10. A voice interaction system for a virtual reality training platform, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of a voice interaction method for a virtual reality training platform according to any one of claims 1-9 are implemented.
Citation Information
Patent Citations
Human voice highlighting processing method and device in audio
CN104916288A
Intelligent office voice control method and system based on voice recognition
CN117995178A
Intelligent voice interaction method of digital memory intervention system
CN118430542A
Voice signal feature extraction method based on Mel-frequency cepstral coefficient and multi-scale entropy
CN118645088A
Method and apparatus for extracting feature of speech signal by emphasizing speech signal
KR1020060091591A