Voice data processing method based on artificial intelligence

By adopting artificial intelligence-based voice data processing methods in school cafeterias, and using Gaussian mixed model and decision tree to extract order description instructions, the problem of low speech recognition accuracy in the prior art is solved, and a more efficient and accurate ordering process is achieved.

CN119993125AActive Publication Date: 2025-05-13青岛丹香投资管理有限公司
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510443454.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-13
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the use scenarios of school cafeterias, the existing method of ordering based on voice recognition has a low recognition accuracy rate, resulting in a high order error rate and affecting the dining experience.

Method used

A speech data processing method based on artificial intelligence is adopted to collect students' ordering audio through a microphone, perform pre-processing and denoising processing, and use Gaussian mixed model and decision tree to extract ordering description instructions to improve recognition accuracy.

Benefits of technology

It improves the accuracy of voice recognition, reduces the error rate of ordering, improves the efficiency and user experience of ordering, and can quickly and accurately determine students' ordering needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993125A_ABST
    Figure CN119993125A_ABST
Patent Text Reader

Abstract

The invention discloses a voice data processing method based on artificial intelligence, and belongs to the technical field of voice recognition, and the method comprises the following steps: S1, collecting the ordering audio of a student through a microphone, carrying out the preprocessing of the ordering audio, and generating a balanced ordering audio; s2, according to the complex feature degree of each frame of voice signal in the balanced ordering audio, generating an original ordering text by using a Gaussian mixture model; and S3, using a plurality of decision nodes of the decision tree to extract an ordering description instruction of the original ordering text to complete ordering. According to the method, the ordering description instruction of the original ordering text is extracted by using the plurality of decision nodes of the decision tree, and the complex ordering text can be efficiently processed and understood, so that the ordering process is rapidly completed, the ordering demand of the student can be rapidly determined, and the possibility of misunderstanding and wrong ordering is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and in particular relates to a speech data processing method based on artificial intelligence. Background Art

[0002] The traditional way of ordering in school cafeterias usually relies on manual service. Students need to line up and verbally tell the staff what dishes they want. This method is prone to long queues and high order error rates during peak meal times. With the continuous development of speech recognition technology, its application in the catering field can greatly improve ordering efficiency and user experience. The use scenario of speech recognition ordering in schools is usually to use a terminal device installed at the cafeteria window to receive and recognize students' voice ordering instructions. However, the recognition accuracy of existing speech recognition-based ordering methods is low. Summary of the invention

[0003] In order to solve the above problems, the present invention proposes a voice data processing method based on artificial intelligence.

[0004] The technical solution of the present invention is: a voice data processing method based on artificial intelligence comprises the following steps: S1. Collect students' ordering audio through a microphone, pre-process the ordering audio, and generate balanced ordering audio; S2, generating the original order text using a Gaussian mixture model based on the complex feature of each frame of speech signal in the balanced order audio; S3. Use several decision nodes of the decision tree to extract the order description instructions of the original order text to complete the order.

[0005] In S1, the audio of ordering food can be denoised.

[0006] By installing a terminal at the cafeteria window to collect students' ordering audio in advance and determine their ordering instructions, cafeteria staff can understand students' ordering content in advance and improve ordering efficiency.

[0007] Furthermore, S2 includes the following sub-steps: S21, extracting the LPC coefficient of each frame of the speech signal in the balanced ordering audio by using a linear predictive coding method; S22, calculating the complex characteristic of each frame of the speech signal according to the LPC coefficient of each frame of the speech signal and the spectral entropy of the balanced ordering audio; S23, determining the number of Gaussian distributions in the Gaussian mixture model according to the complex characteristic degree of each frame of the speech signal in the balanced ordering audio; S23. Using each Gaussian distribution in the Gaussian mixture model, the balanced ordering audio is converted into the original ordering text.

[0008] The beneficial effect of the above further scheme is: in the present invention, in S21, the speech signal is modeled using an autoregressive model, and the LPC coefficient of each frame of the speech signal is obtained by linear prediction analysis. Combining the LPC coefficient and spectral entropy of each frame of the speech signal, the complex characteristics of the speech signal can be comprehensively evaluated, and the Gaussian mixture model can be used to more effectively identify the key information in the speech signal by comprehensively evaluating the complex characteristics of the speech signal, thereby improving the accuracy of speech recognition.

[0009] Furthermore, in S22, the calculation formula of the complex characteristic s of the speech signal is: ; Where p represents the spectral entropy of the balanced ordering audio, and lpc represents the LPC coefficient of the speech signal.

[0010] Further, S23 includes the following sub-steps: S231, inputting the complex feature of each frame of the speech signal in the balanced ordering audio into a complex discrimination model to obtain a complex discrimination value; S232. The number of complex feature degrees greater than the complex discrimination value is taken as the number of Gaussian distributions.

[0011] The beneficial effect of the above further solution is that in the present invention, the complex discriminant model can be learned and adjusted according to different speech signal characteristics, thereby improving the adaptability of the model. The number of Gaussian distributions is dynamically adjusted according to the complex discriminant value, thereby more accurately simulating the distribution characteristics of the speech signal.

[0012] Furthermore, in S231, the complex discriminant model S * The calculation formula is: ; In the formula, s k represents the complex characteristic of the k-th frame speech signal, K represents the total number of speech signal frames of the balanced ordering audio, and s k_max Represents the maximum value of the complex characteristics of all speech signals, s k_min It represents the minimum value of the complex characteristic of all speech signals, and α represents the learning rate of the complex discriminant model.

[0013] The learning rate can be used to control the step size or speed of updating the parameters of the complex discriminant model, and can also be used to appropriately adjust the complex discriminant value to ensure that the complex discriminant value is reasonable.

[0014] Furthermore, S3 includes the following sub-steps: S31, removing stop words from the original order text to obtain the order text to be mined; S32, taking all the words in the ordering text to be mined as the root nodes of the decision tree; S33, determining leaf nodes and several decision nodes of the decision tree according to the order text to be mined; S34, based on the root node, leaf nodes and several decision nodes of the decision tree, using the decision tree to extract the nouns and quantities of dishes in the ordering text to be mined; S35: Ordering food by using the dish name and quantity as a description instruction.

[0015] The beneficial effect of the above further scheme is: in the present invention, the root node is the starting node of the decision tree, representing the entire data set or text collection. The internal node is also called a decision node, which represents the characteristics or attributes of the text; the leaf node is the end node of the decision tree. Taking all the words of the order text to be mined as the root node of the decision tree means that the decision tree contains all possible information in the text at the initial stage. This comprehensive coverage method helps to ensure that no key information is missed in subsequent steps. According to the content of the order text to be mined, the leaf nodes and several decision nodes of the decision tree are flexibly determined, and the structure is adaptively adjusted according to different situations.

[0016] Further, S33 includes the following sub-steps: S331, generating a split coefficient based on the word embedding vector of each word in the ordering text to be mined; S332, according to the split coefficient, calculate the decision split degree of each word, sort the decision split degrees from large to small, and sort the top ranked words. The decision splitting degree is taken as the decision node, M represents the total number of words in the ordering text to be mined, Indicates rounding up; S333. Take the maximum word embedding vector value as the leaf node.

[0017] Furthermore, in S331, the calculation formula of the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x m represents the word embedding vector of the mth word in the order text to be mined, and M represents the total number of words in the order text to be mined.

[0018] The maximum depth limits the maximum depth of the decision tree to prevent overfitting.

[0019] Furthermore, in S332, the calculation formula of the decision splitting degree D of the vocabulary is: ; In the formula, MIN represents the minimum impurity of node partitioning of the decision tree, and t represents the splitting coefficient.

[0020] The minimum impurity of node splitting is used to limit the growth of the decision tree.

[0021] The beneficial effects of the present invention are as follows: the present invention utilizes a Gaussian mixture model to analyze the complex features of each frame of speech signal in the balanced ordering audio, and the Gaussian mixture model can simulate the distribution characteristics of the speech signal and generate the original ordering text, thereby improving the accuracy of speech recognition; the present invention also utilizes several decision nodes of a decision tree to extract the ordering description instructions of the original ordering text, and can efficiently process and understand complex ordering texts, thereby quickly completing the ordering process, and can quickly determine the students' ordering needs, reducing the possibility of misunderstanding and incorrect ordering. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The present invention is a flowchart of a voice data processing method based on artificial intelligence. DETAILED DESCRIPTION

[0023] The embodiments of the present invention are further described below in conjunction with the accompanying drawings.

[0024] like Figure 1 As shown, the present invention provides a voice data processing method based on artificial intelligence, comprising the following steps: S1. Collect students' ordering audio through a microphone, pre-process the ordering audio, and generate balanced ordering audio; S2, generating the original order text using a Gaussian mixture model based on the complex feature of each frame of speech signal in the balanced order audio; S3. Use several decision nodes of the decision tree to extract the order description instructions of the original order text to complete the order.

[0025] In S1, the audio of ordering food can be denoised.

[0026] By installing a terminal at the cafeteria window to collect students' ordering audio in advance and determine their ordering instructions, cafeteria staff can understand students' ordering content in advance and improve ordering efficiency.

[0027] In this embodiment of the present invention, S2 includes the following sub-steps: S21, extracting the LPC coefficient of each frame of the speech signal in the balanced ordering audio by using a linear predictive coding method; S22, calculating the complex characteristic of each frame of the speech signal according to the LPC coefficient of each frame of the speech signal and the spectral entropy of the balanced ordering audio; S23, determining the number of Gaussian distributions in the Gaussian mixture model according to the complex characteristic degree of each frame of the speech signal in the balanced ordering audio; S23. Using each Gaussian distribution in the Gaussian mixture model, the balanced ordering audio is converted into the original ordering text.

[0028] In the present invention, in S21, the speech signal is modeled using an autoregressive model, and the LPC coefficient of each frame of the speech signal is obtained by linear prediction analysis. Combining the LPC coefficient and spectral entropy of each frame of the speech signal, the complex characteristics of the speech signal can be comprehensively evaluated, and the Gaussian mixture model can be used to more effectively identify the key information in the speech signal by comprehensively evaluating the complex characteristics of the speech signal, thereby improving the accuracy of speech recognition.

[0029] In the embodiment of the present invention, in S22, the calculation formula of the complex characteristic s of the speech signal is: ; Where p represents the spectral entropy of the balanced ordering audio, and lpc represents the LPC coefficient of the speech signal.

[0030] In this embodiment of the present invention, S23 includes the following sub-steps: S231, inputting the complex feature of each frame of the speech signal in the balanced ordering audio into a complex discrimination model to obtain a complex discrimination value; S232. The number of complex feature degrees greater than the complex discrimination value is taken as the number of Gaussian distributions.

[0031] In the present invention, the complex discriminant model can be learned and adjusted according to different speech signal characteristics, thereby improving the adaptability of the model. The number of Gaussian distributions is dynamically adjusted according to the complex discriminant value, thereby more accurately simulating the distribution characteristics of the speech signal.

[0032] In the embodiment of the present invention, in S231, the complex discriminant model S * The calculation formula is: ; In the formula, s k represents the complex characteristic of the k-th frame speech signal, K represents the total number of speech signal frames of the balanced ordering audio, and s k_max Represents the maximum value of the complex characteristics of all speech signals, s k_min It represents the minimum value of the complex characteristic of all speech signals, and α represents the learning rate of the complex discriminant model.

[0033] The learning rate can be used to control the step size or speed of updating the parameters of the complex discriminant model, and can also be used to appropriately adjust the complex discriminant value to ensure that the complex discriminant value is reasonable.

[0034] In this embodiment of the present invention, S3 includes the following sub-steps: S31, removing stop words from the original order text to obtain the order text to be mined; S32, taking all the words in the ordering text to be mined as the root nodes of the decision tree; S33, determining leaf nodes and several decision nodes of the decision tree according to the order text to be mined; S34, based on the root node, leaf nodes and several decision nodes of the decision tree, using the decision tree to extract the nouns and quantities of dishes in the ordering text to be mined; S35: Ordering food by using the dish name and quantity as a description instruction.

[0035] In the present invention, the root node is the starting node of the decision tree, representing the entire data set or text collection. Internal nodes are also called decision nodes, which represent the features or attributes of the text; leaf nodes are the end nodes of the decision tree. Taking all the words in the order text to be mined as the root node of the decision tree means that the decision tree contains all possible information in the text at the initial stage. This comprehensive coverage method helps to ensure that no key information is missed in subsequent steps. According to the content of the order text to be mined, the leaf nodes and several decision nodes of the decision tree are flexibly determined, and the structure is adaptively adjusted according to different situations.

[0036] In this embodiment of the present invention, S33 includes the following sub-steps: S331, generating a split coefficient based on the word embedding vector of each word in the ordering text to be mined; S332, according to the split coefficient, calculate the decision split degree of each word, sort the decision split degrees from large to small, and sort the top ranked words. The decision splitting degree is taken as the decision node, M represents the total number of words in the ordering text to be mined, Indicates rounding up; S333. Take the maximum word embedding vector value as the leaf node.

[0037] In the embodiment of the present invention, in S331, the calculation formula of the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x m represents the word embedding vector of the mth word in the order text to be mined, and M represents the total number of words in the order text to be mined.

[0038] The maximum depth limits the maximum depth of the decision tree to prevent overfitting.

[0039] In the embodiment of the present invention, in S332, the calculation formula of the decision splitting degree D of the vocabulary is: ; In the formula, MIN represents the minimum impurity of node partitioning of the decision tree, and t represents the splitting coefficient.

[0040] The minimum impurity of node splitting is used to limit the growth of the decision tree.

[0041] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.

Claims

1. A voice data processing method based on artificial intelligence, characterized in that: The following steps are involved: S1. Collect students' ordering audio through a microphone, pre-process the ordering audio, and generate balanced ordering audio; S2, generating the original order text using a Gaussian mixture model based on the complex feature of each frame of speech signal in the balanced order audio; S3. Use several decision nodes of the decision tree to extract the order description instructions of the original order text to complete the order.

2. The method for processing speech data based on artificial intelligence according to claim 1, characterized in that: The S2 comprises the following sub-steps: S21, extracting the LPC coefficient of each frame of the speech signal in the balanced ordering audio by using a linear predictive coding method; S22, calculating the complex characteristic of each frame of the speech signal according to the LPC coefficient of each frame of the speech signal and the spectral entropy of the balanced ordering audio; S23, determining the number of Gaussian distributions in the Gaussian mixture model according to the complex characteristic degree of each frame of the speech signal in the balanced ordering audio; S23. Using each Gaussian distribution in the Gaussian mixture model, the balanced ordering audio is converted into the original ordering text.

3. The method for processing speech data based on artificial intelligence according to claim 2, characterized in that: In S22, the calculation formula of the complex characteristic s of the speech signal is: ; Where p represents the spectral entropy of the balanced ordering audio, and lpc represents the LPC coefficient of the speech signal.

4. The method for processing speech data based on artificial intelligence according to claim 2, characterized in that: The S23 comprises the following sub-steps: S231, inputting the complex feature of each frame of the speech signal in the balanced ordering audio into a complex discrimination model to obtain a complex discrimination value; S232. The number of complex feature degrees greater than the complex discrimination value is taken as the number of Gaussian distributions.

5. The method for processing speech data based on artificial intelligence according to claim 4, characterized in that: In S231, the complex discriminant model S * The calculation formula is: ; In the formula, s k represents the complex characteristic of the k-th frame speech signal, K represents the total number of speech signal frames of the balanced ordering audio, and s k_max Represents the maximum value of the complex characteristics of all speech signals, s k_min It represents the minimum value of the complex characteristic of all speech signals, and α represents the learning rate of the complex discriminant model.

6. The method for processing speech data based on artificial intelligence according to claim 1, characterized in that: The S3 comprises the following sub-steps: S31, removing stop words from the original order text to obtain the order text to be mined; S32, taking all the words in the ordering text to be mined as the root nodes of the decision tree; S33, determining leaf nodes and several decision nodes of the decision tree according to the order text to be mined; S34, based on the root node, leaf nodes and several decision nodes of the decision tree, using the decision tree to extract the nouns and quantities of dishes in the ordering text to be mined; S35: Ordering food by using the dish name and quantity as a description instruction.

7. The method for processing speech data based on artificial intelligence according to claim 6, characterized in that: The S33 comprises the following sub-steps: S331, generating a split coefficient based on the word embedding vector of each word in the ordering text to be mined; S332, according to the split coefficient, calculate the decision split degree of each word, sort the decision split degrees from large to small, and sort the top ranked words. The decision splitting degree is taken as the decision node, M represents the total number of words in the ordering text to be mined, Indicates rounding up; S333. Take the maximum word embedding vector value as the leaf node.

8. The method for processing speech data based on artificial intelligence according to claim 7, characterized in that: In S331, the calculation formula of the splitting coefficient t is: ; depth represents the maximum depth of the decision tree, x m represents the word embedding vector of the mth word in the order text to be mined, and M represents the total number of words in the order text to be mined.

9. The method for processing speech data based on artificial intelligence according to claim 7, characterized in that: In S332, the calculation formula of the decision splitting degree D of the vocabulary is: ; In the formula, MIN represents the minimum impurity of node partitioning of the decision tree, and t represents the splitting coefficient.

Citation Information

Patent Citations

  • Patient weak voice endpoint detection method

    CN103077728A

  • Voice navigation system and method of animal robot system

    CN103593048A

  • System and method for food ordering system data mining algorithm

    CN106548422A

  • Speech recognition method, device and equipment, and computer readable storage medium

    CN107680597A

  • Training method of voice endpoint detection model and voice noise reduction method

    CN113744725A