Voice data processing method based on artificial intelligence
By adopting artificial intelligence-based voice data processing methods in school cafeterias, and using Gaussian mixed model and decision tree to extract order description instructions, the problem of low speech recognition accuracy in the prior art is solved, and a more efficient and accurate ordering process is achieved.
Patent Information
- Application Number
- CN202510443454.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the use scenarios of school cafeterias, the existing method of ordering based on voice recognition has a low recognition accuracy rate, resulting in a high order error rate and affecting the dining experience.
A speech data processing method based on artificial intelligence is adopted to collect students' ordering audio through a microphone, perform pre-processing and denoising processing, and use Gaussian mixed model and decision tree to extract ordering description instructions to improve recognition accuracy.
It improves the accuracy of voice recognition, reduces the error rate of ordering, improves the efficiency and user experience of ordering, and can quickly and accurately determine students' ordering needs.
Smart Images

Figure CN119993125A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech recognition, and in particular relates to a speech data processing method based on artificial intelligence. Background Art
[0002] The traditional way of ordering in school cafeterias usually relies on manual service. Students need to line up and verbally tell the staff what dishes they want. This method is prone to long queues and high order error rates during peak meal times. With the continuous development of speech recognition technology, its application in the catering field can greatly improve ordering efficiency and user experience. The use scenario of speech recognition ordering in schools is usually to use a terminal device installed at the cafeteria window to receive and recognize students' voice ordering instructions. However, the recognition accuracy of existing speech recognition-based ordering methods is low. Summary of the invention
[0003] In order to solve the above problems, the present invention proposes a voice data processing method based on artificial intelligence.
[0004] The technical solution of the present invention is: a voice data processing method based on artificial intelligence comprises the following steps: S1. Collect students' ordering audio through a microphone, pre-process the ordering audio, and generate balanced ordering audio; S2, generating the original order text using a Gaussian mixture model based on the complex feature of each frame of speech signal in the balanced order audio; S3. Use several decision nodes of the decision tree to extract the order description instructions of the original order text to complete the order.
[0005] In S1, the audio of ordering food can be denoised.
[0006] By installing a terminal at the cafeteria window to collect students' ordering audio in advance and determine their ordering instructions, cafeteria staff can understand students' ordering content in advance and improve ordering efficiency.
[0007] Furthermore, S2 includes the following sub-steps: S21, extracting the LPC coefficient of each frame of the speech signal in the balanced ordering audio by using a linear predictive coding method; S22, calculating the complex characteristic of each frame of the speech signal according to the LPC coefficient of each frame of the speech signal and the spectral entropy of the balanced ordering audio; S23, determining the number of Gaussian distributions in the Gaussian mixture model according to the complex characteristic degree of each frame of the speech signal in the balanced ordering audio; S23. Using each Gaussian distribution in the Gaussian mixture model, the balanced ordering audio is converted into the original ordering text.
[0008] The beneficial effect of the above further scheme is: in the present invention, in S21, the speech signal is modeled using an autoregressive model, and the LPC coefficient of each frame of the speech signal is obtained by linear prediction analysis. Combining the LPC coefficient and spectral entropy of each frame of the speech signal, the complex characteristics of the speech signal can be comprehensively evaluated, and the Gaussian mixture model can be used to more effectively identify the key information in the speech signal by comprehensively evaluating the complex characteristics of the speech signal, thereby improving the accuracy of speech recognition.
[0009] Furthermore, in S22, the calculation formula of the complex characteristic s of the speech signal is: ; Where p represents the spectral entropy of the balanced ordering audio, and lpc represents the LPC coefficient of the speech signal.
[0010] Further, S23 includes the following sub-steps: S231, inputting the complex feature of each frame of the speech signal in the balanced ordering audio into a complex discrimination model to obtain a complex discrimination value; S232. The number of complex feature degrees greater than the complex discrimination value is taken as the number of Gaussian distributions.
[0011] The beneficial effect of the above further solution is that in the present invention, the complex discriminant model can be learned and adjusted according to different speech signal characteristics, thereby improving the adaptability of the model. The number of Gaussian distributions is dynamically adjusted according to the complex discriminant value, thereby more accurately simulating the distribution characteristics of the speech signal.
[0012] Furthermore, in S231, the complex discriminant model S * The calculation formula is: ; In the formula, s k represents the complex characteristic of the k-th frame speech signal, K represents the total number of speech signal frames of the balanced ordering audio, and s k_max Represents the maximum value of the complex characteristics of all speech signals, s k_min It represents the minimum value of the complex characteristic of all speech signals, and α represents the learning rate of the complex discriminant model.
[0013] The learning rate can be used to control the step size or speed of updating the parameters of the complex discriminant model, and can also be used to appropriately adjust the complex discriminant value to ensure that the complex discriminant value is reasonable.
[0014] Furthermore, S3 includes the following sub-steps: S31, removing stop words from the original order text to obtain the order text to be mined; S32, taking all the words in the ordering text to be mined as the root nodes of the decision tree; S33, determining leaf nodes and several decision nodes of the decision tree according to the order text to be mined; S34, based on the root node, leaf nodes and several decision nodes of the decision tree, using the decision tree to extract the nouns and quantities of dishes in the ordering text to be mined; S35: Ordering food by using the dish name and quantity as a description instruction.
[0015] The beneficial effect of the above further scheme is: in the present invention, the root node is the starting node of the decision tree, representing the entire data set or text collection. The internal node is also called a decision node, which represents the characteristics or attributes of the text; the leaf node is the end node of the decision tree. Taking all the words of the order text to be mined as the root node of the decision tree means that the decision tree contains all possible information in the text at the initial stage. This comprehensive coverage method helps to ensure that no key information is missed in subsequent steps. According to the content of the order text to be mined, the leaf nodes and several decision nodes of the decision tree are flexibly determined, and the structure is adaptively adjusted according to different situations.
[0016] Further, S33 includes the following sub-steps: S331, generating a split coefficient based on the word embedding vector of each word in the ordering text to be mined; S332, according to the split coefficient, calculate the decision split degree of each word, sort the decision split degrees from large to small, and sort the top ranked words. The decision splitting degree is taken as the decision node, M represents the total number of words in the ordering text to be mined, Indicates rounding up; S333. Take the maximum word embedding vector value as the leaf node.
[0017] Furthermore, in S331, the calculation formula of the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x m represents the word embedding vector of the mth word in the order text to be mined, and M represents the total number of words in the order text to be mined.
[0018] The maximum depth limits the maximum depth of the decision tree to prevent overfitting.
[0019] Furthermore, in S332, the calculation formula of the decision splitting degree D of the vocabulary is: ; In the formula, MIN represents the minimum impurity of node partitioning of the decision tree, and t represents the splitting coefficient.
[0020] The minimum impurity of node splitting is used to limit the growth of the decision tree.
[0021] The beneficial effects of the present invention are as follows: the present invention utilizes a Gaussian mixture model to analyze the complex features of each frame of speech signal in the balanced ordering audio, and the Gaussian mixture model can simulate the distribution characteristics of the speech signal and generate the original ordering text, thereby improving the accuracy of speech recognition; the present invention also utilizes several decision nodes of a decision tree to extract the ordering description instructions of the original ordering text, and can efficiently process and understand complex ordering texts, thereby quickly completing the ordering process, and can quickly determine the students' ordering needs, reducing the possibility of misunderstanding and incorrect ordering. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The present invention is a flowchart of a voice data processing method based on artificial intelligence. DETAILED DESCRIPTION
[0023] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.
[0024] like Figure 1 As shown, the present invention provides a voice data processing method based on artificial intelligence, comprising the following steps: S1. Collect students' ordering audio through a microphone, pre-process the ordering audio, and generate balanced ordering audio; S2, generating the original order text using a Gaussian mixture model based on the complex feature of each frame of speech signal in the balanced order audio; S3. Use several decision nodes of the decision tree to extract the order description instructions of the original order text to complete the order.
[0025] In S1, the audio of ordering food can be denoised.
[0026] By installing a terminal at the cafeteria window to collect students' ordering audio in advance and determine their ordering instructions, cafeteria staff can understand students' ordering content in advance and improve ordering efficiency.
[0027] In this embodiment of the present invention, S2 includes the following sub-steps: S21, extracting the LPC coefficient of each frame of the speech signal in the balanced ordering audio by using a linear predictive coding method; S22, calculating the complex characteristic of each frame of the speech signal according to the LPC coefficient of each frame of the speech signal and the spectral entropy of the balanced ordering audio; S23, determining the number of Gaussian distributions in the Gaussian mixture model according to the complex characteristic degree of each frame of the speech signal in the balanced ordering audio; S23. Using each Gaussian distribution in the Gaussian mixture model, the balanced ordering audio is converted into the original ordering text.
[0028] In the present invention, in S21, the speech signal is modeled using an autoregressive model, and the LPC coefficient of each frame of the speech signal is obtained by linear prediction analysis. Combining the LPC coefficient and spectral entropy of each frame of the speech signal, the complex characteristics of the speech signal can be comprehensively evaluated, and the Gaussian mixture model can be used to more effectively identify the key information in the speech signal by comprehensively evaluating the complex characteristics of the speech signal, thereby improving the accuracy of speech recognition.
[0029] In the embodiment of the present invention, in S22, the calculation formula of the complex characteristic s of the speech signal is: ; Where p represents the spectral entropy of the balanced ordering audio, and lpc represents the LPC coefficient of the speech signal.
[0030] In this embodiment of the present invention, S23 includes the following sub-steps: S231, inputting the complex feature of each frame of the speech signal in the balanced ordering audio into a complex discrimination model to obtain a complex discrimination value; S232. The number of complex feature degrees greater than the complex discrimination value is taken as the number of Gaussian distributions.
[0031] In the present invention, the complex discriminant model can be learned and adjusted according to different speech signal characteristics, thereby improving the adaptability of the model. The number of Gaussian distributions is dynamically adjusted according to the complex discriminant value, thereby more accurately simulating the distribution characteristics of the speech signal.
[0032] In the embodiment of the present invention, in S231, the complex discriminant model S * The calculation formula is: ; In the formula, s k represents the complex characteristic of the k-th frame speech signal, K represents the total number of speech signal frames of the balanced ordering audio, and s k_max Represents the maximum value of the complex characteristics of all speech signals, s k_min It represents the minimum value of the complex characteristic of all speech signals, and α represents the learning rate of the complex discriminant model.
[0033] The learning rate can be used to control the step size or speed of updating the parameters of the complex discriminant model, and can also be used to appropriately adjust the complex discriminant value to ensure that the complex discriminant value is reasonable.
[0034] In this embodiment of the present invention, S3 includes the following sub-steps: S31, removing stop words from the original order text to obtain the order text to be mined; S32, taking all the words in the ordering text to be mined as the root nodes of the decision tree; S33, determining leaf nodes and several decision nodes of the decision tree according to the order text to be mined; S34, based on the root node, leaf nodes and several decision nodes of the decision tree, using the decision tree to extract the nouns and quantities of dishes in the ordering text to be mined; S35: Ordering food by using the dish name and quantity as a description instruction.
[0035] In the present invention, the root node is the starting node of the decision tree, representing the entire data set or text collection. Internal nodes are also called decision nodes, which represent the features or attributes of the text; leaf nodes are the end nodes of the decision tree. Taking all the words in the order text to be mined as the root node of the decision tree means that the decision tree contains all possible information in the text at the initial stage. This comprehensive coverage method helps to ensure that no key information is missed in subsequent steps. According to the content of the order text to be mined, the leaf nodes and several decision nodes of the decision tree are flexibly determined, and the structure is adaptively adjusted according to different situations.
[0036] In this embodiment of the present invention, S33 includes the following sub-steps: S331, generating a split coefficient based on the word embedding vector of each word in the ordering text to be mined; S332, according to the split coefficient, calculate the decision split degree of each word, sort the decision split degrees from large to small, and sort the top ranked words. The decision splitting degree is taken as the decision node, M represents the total number of words in the ordering text to be mined, Indicates rounding up; S333. Take the maximum word embedding vector value as the leaf node.
[0037] In the embodiment of the present invention, in S331, the calculation formula of the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x m represents the word embedding vector of the mth word in the order text to be mined, and M represents the total number of words in the order text to be mined.
[0038] The maximum depth limits the maximum depth of the decision tree to prevent overfitting.
[0039] In the embodiment of the present invention, in S332, the calculation formula of the decision splitting degree D of the vocabulary is: ; In the formula, MIN represents the minimum impurity of node partitioning of the decision tree, and t represents the splitting coefficient.
[0040] The minimum impurity of node splitting is used to limit the growth of the decision tree.
[0041] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.
Claims
1. A voice data processing method based on artificial intelligence, characterized in that: The following steps are involved: S1. Collect students' ordering audio through a microphone, pre-process the ordering audio, and generate balanced ordering audio; S2, generating the original order text using a Gaussian mixture model based on the complex feature of each frame of speech signal in the balanced order audio; S3. Use several decision nodes of the decision tree to extract the order description instructions of the original order text to complete the order.
2. The method for processing speech data based on artificial intelligence according to claim 1, characterized in that: The S2 comprises the following sub-steps: S21, extracting the LPC coefficient of each frame of the speech signal in the balanced ordering audio by using a linear predictive coding method; S22, calculating the complex characteristic of each frame of the speech signal according to the LPC coefficient of each frame of the speech signal and the spectral entropy of the balanced ordering audio; S23, determining the number of Gaussian distributions in the Gaussian mixture model according to the complex characteristic degree of each frame of the speech signal in the balanced ordering audio; S23. Using each Gaussian distribution in the Gaussian mixture model, the balanced ordering audio is converted into the original ordering text.
3. The method for processing speech data based on artificial intelligence according to claim 2, characterized in that: In S22, the calculation formula of the complex characteristic s of the speech signal is: ; Where p represents the spectral entropy of the balanced ordering audio, and lpc represents the LPC coefficient of the speech signal.
4. The method for processing speech data based on artificial intelligence according to claim 2, characterized in that: The S23 comprises the following sub-steps: S231, inputting the complex feature of each frame of the speech signal in the balanced ordering audio into a complex discrimination model to obtain a complex discrimination value; S232. The number of complex feature degrees greater than the complex discrimination value is taken as the number of Gaussian distributions.
5. The method for processing speech data based on artificial intelligence according to claim 4, characterized in that: In S231, the complex discriminant model S * The calculation formula is: ; In the formula, s k represents the complex characteristic of the k-th frame speech signal, K represents the total number of speech signal frames of the balanced ordering audio, and s k_max Represents the maximum value of the complex characteristics of all speech signals, s k_min It represents the minimum value of the complex characteristic of all speech signals, and α represents the learning rate of the complex discriminant model.
6. The method for processing speech data based on artificial intelligence according to claim 1, characterized in that: The S3 comprises the following sub-steps: S31, removing stop words from the original order text to obtain the order text to be mined; S32, taking all the words in the ordering text to be mined as the root nodes of the decision tree; S33, determining leaf nodes and several decision nodes of the decision tree according to the order text to be mined; S34, based on the root node, leaf nodes and several decision nodes of the decision tree, using the decision tree to extract the nouns and quantities of dishes in the ordering text to be mined; S35: Ordering food by using the dish name and quantity as a description instruction.
7. The method for processing speech data based on artificial intelligence according to claim 6, characterized in that: The S33 comprises the following sub-steps: S331, generating a split coefficient based on the word embedding vector of each word in the ordering text to be mined; S332, according to the split coefficient, calculate the decision split degree of each word, sort the decision split degrees from large to small, and sort the top ranked words. The decision splitting degree is taken as the decision node, M represents the total number of words in the ordering text to be mined, Indicates rounding up; S333. Take the maximum word embedding vector value as the leaf node.
8. The method for processing speech data based on artificial intelligence according to claim 7, characterized in that: In S331, the calculation formula of the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x m represents the word embedding vector of the mth word in the order text to be mined, and M represents the total number of words in the order text to be mined.
9. The method for processing speech data based on artificial intelligence according to claim 7, characterized in that: In S332, the calculation formula of the decision splitting degree D of the vocabulary is: ; In the formula, MIN represents the minimum impurity of node partitioning of the decision tree, and t represents the splitting coefficient.
Citation Information
Patent Citations
Patient weak voice endpoint detection method
CN103077728A
Voice navigation system and method of animal robot system
CN103593048A
System and method for food ordering system data mining algorithm
CN106548422A
Speech recognition method, device and equipment, and computer readable storage medium
CN107680597A
Training method of voice endpoint detection model and voice noise reduction method
CN113744725A