A speech data processing method based on artificial intelligence

By pre-processing the ordering audio and Gaussian mixed model analysis combined with decision tree node extraction, the problem of low ordering accuracy of speech recognition is solved, and an efficient and accurate ordering process is achieved.

CN119993125BActive Publication Date: 2025-07-04青岛丹香投资管理有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510443454.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-04
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The existing method of ordering based on voice recognition is low in school cafeterias, resulting in long queues during peak meals and high order error rates.

Method used

Ordering audio is collected through the microphone, and the original ordering text is generated using the Gaussian mixed model. Several decision nodes of the decision tree are used to extract ordering description instructions, including denoising processing, linear prediction coding, complex feature calculation and node determination of the decision tree.

Benefits of technology

It improves the accuracy of voice recognition, reduces the possibility of misunderstandings and wrong ordering, and improves the efficiency of ordering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993125B_ABST
    Figure CN119993125B_ABST
Patent Text Reader

Abstract

The present invention discloses a voice data processing method based on artificial intelligence, belonging to the technical field of speech recognition, which includes the following steps: S1. Collect the ordering audio of students through a microphone, and preprocess the ordering audio to generate balanced ordering audio; S2. Generate the original ordering text using a Gaussian mixture model according to the complexity degree of each frame of voice signal in the balanced ordering audio; S3. Extract the ordering description instructions of the original ordering text using several decision nodes of a decision tree to complete the ordering. The present invention uses several decision nodes of a decision tree to extract the ordering description instructions of the original ordering text, which can efficiently process and understand complex ordering texts, thereby quickly completing the ordering process, quickly determining the ordering needs of students, and reducing the possibility of misunderstanding and incorrect ordering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and particularly relates to a method for processing speech data based on artificial intelligence. Background Art

[0002] Traditional school cafeteria ordering methods usually rely on manual services. Students need to queue up and verbally tell the staff the dishes they want. This method is prone to problems such as long queuing times and high error rates in ordering during peak dining hours. With the continuous development of speech recognition technology, applying it to the catering field can greatly improve the ordering efficiency and user experience. The usage scenario of speech recognition for ordering in schools is usually to use terminal devices installed at the cafeteria windows to receive and recognize students' speech ordering instructions. However, the recognition accuracy of existing ordering methods based on speech recognition is relatively low. Summary of the Invention

[0003] In order to solve the above problems, the present invention proposes a method for processing speech data based on artificial intelligence.

[0004] The technical solution of the present invention is: A method for processing speech data based on artificial intelligence includes the following steps:

[0005] S1. Collect students' ordering audio through a microphone, and preprocess the ordering audio to generate balanced ordering audio;

[0006] S2. Generate an original ordering text using a Gaussian mixture model according to the complexity degree of each frame of speech signal in the balanced ordering audio;

[0007] S3. Extract the ordering description instructions of the original ordering text using several decision nodes of a decision tree to complete the ordering.

[0008] In S1, noise reduction processing can be performed on the ordering audio.

[0009] By installing terminal devices at the cafeteria windows to collect students' ordering audio in advance and determine students' ordering instructions, the cafeteria staff can understand students' ordering content in advance and improve the ordering efficiency.

[0010] Further, S2 includes the following sub-steps:

[0011] S21. Use the linear predictive coding method to extract the LPC coefficients of each frame of speech signal in the balanced ordering audio;

[0012] S22. Calculate the complexity degree of each frame of speech signal according to the LPC coefficients of each frame of speech signal and the spectral entropy of the balanced ordering audio;

[0013] S23. Determine the number of Gaussian distributions in the Gaussian mixture model according to the complexity degree of each frame of speech signal in the balanced ordering audio;

[0014] S23. Convert the balanced order-taking audio into the original order-taking text by using each Gaussian distribution in the Gaussian mixture model.

[0015] The beneficial effect of the above further solution is as follows: In the present invention, in S21, an autoregressive model is used to model the speech signal, and the LPC coefficients of each frame of the speech signal are obtained through linear prediction analysis. Combining the LPC coefficients and spectral entropy of each frame of the speech signal can comprehensively evaluate the complexity degree of the speech signal. By comprehensively evaluating the complexity degree of the speech signal using the Gaussian mixture model, the key information in the speech signal can be more effectively identified, thereby improving the accuracy of speech recognition.

[0016] Furthermore, in S22, the calculation formula for the complexity degree s of the speech signal is: ; where p represents the spectral entropy of the balanced order-taking audio, and lpc represents the LPC coefficient of the speech signal.

[0017] Furthermore, S23 includes the following sub-steps:

[0018] S231. Input the complexity degree of each frame of the speech signal in the balanced order-taking audio into the complexity discrimination model to obtain a complexity discrimination value;

[0019] S232. Take the number of complexity degrees greater than the complexity discrimination value as the number of Gaussian distributions.

[0020] The beneficial effect of the above further solution is as follows: In the present invention, the complexity discrimination model can learn and adjust according to different speech signal characteristics, thereby improving the adaptability of the model. Dynamically adjust the number of Gaussian distributions according to the complexity discrimination value, so as to more accurately simulate the distribution characteristics of the speech signal.

[0021] Furthermore, in S231, the complexity discrimination model S * has the following calculation formula: ; where s k represents the complexity degree of the k-th frame of the speech signal, K represents the total number of frames of the speech signal in the balanced order-taking audio, s k_max represents the maximum value of the complexity degrees of all speech signals, s k_min represents the minimum value of the complexity degrees of all speech signals, and α represents the learning rate of the complexity discrimination model.

[0022] The learning rate can be used to control the step size or speed of the parameter update of the complexity discrimination model, and is also used to appropriately adjust the complexity discrimination value to ensure the rationality of the complexity discrimination value.

[0023] Furthermore, S3 includes the following sub-steps:

[0024] S31. Remove the stop words from the original order-taking text to obtain the order-taking text to be mined;

[0025] S32. Use all the words in the order-taking text to be mined as the root node of the decision tree;

[0026] S33. Determine the leaf nodes and several decision nodes of the decision tree according to the order-taking text to be mined;

[0027] S34. Based on the root node, leaf nodes and several decision nodes of the decision tree, use the decision tree to extract the dish names and quantities in the order-taking text to be mined;

[0028] S35. Use the dish names and quantities as order-taking description instructions to place an order.

[0029] The beneficial effects of the above further solution are as follows: In the present invention, the root node is the starting node of the decision tree, representing the entire data set or text collection. The internal node is also called the decision node, and the decision node represents the features or attributes of the text; the leaf node is the terminal node of the decision tree. Using all the words in the order-taking text to be mined as the root node of the decision tree means that the decision tree contains all possible information in the text at the initial stage. This comprehensive coverage method helps to ensure that no key information is missed in the subsequent steps. Determine the leaf nodes and several decision nodes of the decision tree flexibly according to the content of the order-taking text to be mined, and adaptively adjust the structure according to different situations.

[0030] Further, S33 includes the following sub-steps:

[0031] S331. Generate a splitting coefficient based on the word embedding vectors of each word in the order-taking text to be mined;

[0032] S332. Calculate the decision splitting degree of each word according to the splitting coefficient, sort the decision splitting degrees from large to small, and take the top decision splitting degrees as decision nodes, where M represents the total number of words in the order-taking text to be mined, represents rounding up;

[0033] S333. Take the maximum word embedding vector value as the leaf node.

[0034] Further, in S331, the calculation formula for the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, and x m represents the word embedding vector of the m-th word in the order-taking text to be mined, and M represents the total number of words in the order-taking text to be mined.

[0035] The maximum depth limits the maximum depth of the decision tree and can prevent overfitting.

[0036] Further, in S332, the calculation formula for the decision splitting degree D of the vocabulary is: ; in the formula, MIN represents the minimum impurity of node division of the decision tree, and t represents the splitting coefficient.

[0037] The minimum impurity of node division is used to limit the growth of the decision tree.

[0038] The beneficial effects of the present invention are as follows: The present invention uses a Gaussian mixture model to analyze the complex feature degree of each frame of speech signal in the balanced ordering audio. The Gaussian mixture model can simulate the distribution characteristics of the speech signal and generate the original ordering text, thereby improving the accuracy of speech recognition. The present invention also uses several decision nodes of the decision tree to extract the ordering description instructions of the original ordering text, which can efficiently process and understand complex ordering texts, thereby quickly completing the ordering process, quickly determining the ordering needs of students, and reducing the possibility of misunderstanding and incorrect ordering. Description of the Drawings

[0039] Figure 1 It is a flowchart of a speech data processing method based on artificial intelligence. Detailed Embodiments

[0040] The following further describes the embodiments of the present invention with reference to the drawings.

[0041] As Figure 1 shown, the present invention provides a speech data processing method based on artificial intelligence, including the following steps:

[0042] S1. Collect the ordering audio of students through a microphone, and preprocess the ordering audio to generate balanced ordering audio;

[0043] S2. Generate the original ordering text using a Gaussian mixture model according to the complex feature degree of each frame of speech signal in the balanced ordering audio;

[0044] S3. Use several decision nodes of the decision tree to extract the ordering description instructions of the original ordering text to complete the ordering.

[0045] In S1, the ordering audio can be denoised.

[0046] By collecting the ordering audio of students in advance through the terminal installed at the cafeteria window to determine the ordering instructions of students, the staff in the cafeteria can understand the ordering content of students in advance and improve the ordering efficiency.

[0047] In the embodiment of the present invention, S2 includes the following sub-steps:

[0048] S21. Use the linear predictive coding method to extract the LPC coefficients of each frame of speech signal in the balanced ordering audio;

[0049] S22. Calculate the complexity feature degree of each frame of speech signal according to the LPC coefficients of each frame of speech signal and the spectral entropy of the balanced order-taking audio;

[0050] S23. Determine the number of Gaussian distributions in the Gaussian mixture model according to the complexity feature degree of each frame of speech signal in the balanced order-taking audio;

[0051] S23. Use each Gaussian distribution in the Gaussian mixture model to convert the balanced order-taking audio into the original order-taking text.

[0052] In the present invention, in S21, an autoregressive model is used to model the speech signal, and the LPC coefficients of each frame of speech signal are obtained through linear prediction analysis. Combining the LPC coefficients and spectral entropy of each frame of speech signal can comprehensively evaluate the complexity feature degree of the speech signal. By comprehensively evaluating the complexity feature degree of the speech signal using the Gaussian mixture model, the key information in the speech signal can be more effectively identified, thereby improving the accuracy of speech recognition.

[0053] In the embodiment of the present invention, in S22, the calculation formula for the complexity feature degree s of the speech signal is: ; where p represents the spectral entropy of the balanced order-taking audio, and lpc represents the LPC coefficients of the speech signal.

[0054] In the embodiment of the present invention, S23 includes the following sub-steps:

[0055] S231. Input the complexity feature degree of each frame of speech signal in the balanced order-taking audio into the complexity discrimination model to obtain a complexity discrimination value;

[0056] S232. Use the number of complexity feature degrees greater than the complexity discrimination value as the number of Gaussian distributions.

[0057] In the present invention, the complexity discrimination model can learn and adjust according to different speech signal characteristics, thereby improving the adaptability of the model. Dynamically adjust the number of Gaussian distributions according to the complexity discrimination value, so as to more accurately simulate the distribution characteristics of the speech signal.

[0058] In the embodiment of the present invention, in S231, the complexity discrimination model S * The calculation formula is: ; where s k represents the complexity feature degree of the k-th frame of speech signal, K represents the total number of frames of speech signals in the balanced order-taking audio, s k_max represents the maximum value of the complexity feature degrees of all speech signals, s k_min represents the minimum value of the complexity feature degrees of all speech signals, and α represents the learning rate of the complexity discrimination model.

[0059] The learning rate can be used to control the step size or speed of parameter updates for a complex discriminant model, and is also used to appropriately adjust the complex discriminant value to ensure the rationality of the complex discriminant value.

[0060] In an embodiment of the present invention, S3 includes the following sub-steps:

[0061] S31. Remove the stop words from the original order-taking text to obtain the order-taking text to be mined;

[0062] S32. Use all the words in the order-taking text to be mined as the root node of the decision tree;

[0063] S33. Determine the leaf nodes and several decision nodes of the decision tree according to the order-taking text to be mined;

[0064] S34. Based on the root node, leaf nodes and several decision nodes of the decision tree, use the decision tree to extract the dish names and quantities from the order-taking text to be mined;

[0065] S35. Use the dish names and quantities as order-taking description instructions to place an order.

[0066] In the present invention, the root node is the starting node of the decision tree, representing the entire data set or text collection. The internal nodes are also called decision nodes, and the decision nodes represent the features or attributes of the text; the leaf nodes are the terminal nodes of the decision tree. Using all the words in the order-taking text to be mined as the root node of the decision tree means that the decision tree contains all possible information in the text at the initial stage. This comprehensive coverage method helps to ensure that no key information is missed in the subsequent steps. Determine the leaf nodes and several decision nodes of the decision tree flexibly according to the content of the order-taking text to be mined, and adaptively adjust the structure according to different situations.

[0067] In an embodiment of the present invention, S33 includes the following sub-steps:

[0068] S331. Generate a splitting coefficient based on the word embedding vectors of each word in the order-taking text to be mined;

[0069] S332. Calculate the decision splitting degree of each word according to the splitting coefficient, sort the decision splitting degrees from largest to smallest, and take the top decision splitting degrees as decision nodes, where M represents the total number of words in the order-taking text to be mined, represents rounding up;

[0070] S333. Use the maximum word embedding vector value as the leaf node.

[0071] In an embodiment of the present invention, in S331, the calculation formula for the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x mIt represents the word embedding vector of the m-th word in the to-be-mined meal-ordering text, and M represents the total number of words in the to-be-mined meal-ordering text.

[0072] The maximum depth limits the maximum depth of the decision tree and can prevent overfitting.

[0073] In the embodiment of the present invention, in S332, the calculation formula of the decision splitting degree D of the word is as follows: ; in the formula, MIN represents the minimum impurity of node division of the decision tree, and t represents the splitting coefficient.

[0074] The minimum impurity of node division is used to limit the growth of the decision tree.

[0075] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not deviate from the essence of the present invention according to these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A method for processing voice data based on artificial intelligence, characterized in that, It includes the following steps: S1. Collect the ordering audio of students through a microphone, and preprocess the ordering audio to generate balanced ordering audio; S2. Generate the original ordering text using a Gaussian mixture model according to the complexity feature degree of each frame of voice signal in the balanced ordering audio; S3. Extract the ordering description instructions of the original ordering text using several decision nodes of a decision tree to complete the ordering; The S3 includes the following sub-steps: S31. Remove the stop words of the original ordering text to obtain the ordering text to be mined; S32. Use all the words of the ordering text to be mined as the root node of the decision tree; S33. Determine the leaf nodes and several decision nodes of the decision tree according to the ordering text to be mined; S34. Based on the root node, leaf nodes and several decision nodes of the decision tree, extract the dish names and quantities of the ordering text to be mined using the decision tree; S35. Use the dish names and quantities as ordering description instructions to place an order; The S33 includes the following sub-steps: S331. Generate a splitting coefficient based on the word embedding vectors of each word in the ordering text to be mined; S332. Calculate the decision splitting degree of each word according to the splitting coefficient, sort the decision splitting degrees from largest to smallest, and use the top decision splitting degrees as decision nodes. M represents the total number of words in the meal ordering text to be mined, indicates rounding up; S333. Use the maximum word embedding vector value as the leaf node; In the above S331, the calculation formula of the splitting coefficient t is as follows: ; depth represents the maximum depth of the decision tree, and x m represents the word embedding vector of the m-th word in the to-be-mined ordering text, and M represents the total number of words in the to-be-mined ordering text; In S332, the calculation formula for the decision splitting degree D of the vocabulary is as follows: ; where MIN represents the minimum impurity of node division of the decision tree, and t represents the splitting coefficient.

2. The method for processing voice data based on artificial intelligence according to claim 1, wherein The S2 includes the following sub-steps: S21. Extract the LPC coefficients of each frame of voice signal in the balanced ordering audio using the linear predictive coding method; S22. Calculate the complexity feature degree of each frame of voice signal according to the LPC coefficients of each frame of voice signal and the spectral entropy of the balanced ordering audio; S23. Determine the number of Gaussian distributions in the Gaussian mixture model according to the complexity feature degree of each frame of voice signal in the balanced ordering audio; S23. Use each Gaussian distribution in the Gaussian mixture model to convert the balanced ordering audio into the original ordering text.

3. The method for processing voice data based on artificial intelligence according to claim 2, wherein In S22, the calculation formula for the complexity degree s of the voice signal is as follows: ; where p represents the spectral entropy of the balanced order-taking audio, and lpc represents the LPC coefficient of the voice signal.

4. The method for processing voice data based on artificial intelligence according to claim 2, wherein The S23 includes the following sub-steps: S231. Input the complexity feature degree of each frame of voice signal in the balanced ordering audio into the complexity discrimination model to obtain a complexity discrimination value; S232. Use the number of complexity feature degrees greater than the complexity discrimination value as the number of Gaussian distributions.

5. The method for processing speech data based on artificial intelligence according to claim 4, wherein In S231, the complex discrimination model S * has the following calculation formula: ; where s k represents the complex feature degree of the k-th frame of the voice signal, K represents the total number of frames of the voice signal of the balanced ordering audio, s k_max represents the maximum value of the complex feature degrees of all voice signals, s k_min represents the minimum value of the complex feature degrees of all voice signals, and α represents the learning rate of the complex discrimination model.

Citation Information

Patent Citations

  • Patient weak voice endpoint detection method

    CN103077728A

  • Voice navigation system and method of animal robot system

    CN103593048A