Voice processing method of early education robot and early education robot
The early childhood education robot method of voice category recognition and action image consistency verification solves the problems of high computing power consumption and poor early childhood education results of early childhood education robots, achieves more efficient voice processing and interactive diversity, and improves early childhood education results.
Patent Information
- Application Number
- CN202510131366.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Existing early childhood education robots consume a lot of computing power for voice recognition and have poor early childhood education effects, and fail to effectively interact with children's movements and behaviors.
Through speech category recognition, the input speech is divided into adult speech and child speech, and combined with speech content recognition, the speech processing process is simplified. At the same time, the feature consistency between speech and children's action images is used for early childhood education and double verification.
It significantly reduces the computing power consumption of voice recognition, improves the accuracy of command recognition, and enhances the interaction diversity and teaching effect between early childhood education robots and children.
Smart Images

Figure CN119626221B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing, in particular to a speech processing method of an early education robot and the early education robot. BACKGROUND
[0002] Early education robots are gradually becoming an important tool in children's education. These robots usually interact with children through speech recognition technology to realize the presentation of early education content. However, the existing early education robots mainly use speech recognition as the main interaction mode, and the technical complexity and power consumption problems gradually appear.
[0003] The current early education robots need to recognize the speech of adults and children, which consumes a lot of power, and mainly uses speech recognition as the main interaction mode, does not combine the action behavior of children, and the early education effect is poor. SUMMARY
[0004] The present application provides a speech processing method of an early education robot, which is used to solve the technical problems of large speech recognition power consumption and poor early education effect of the existing early education robots.
[0005] In view of the above problems, the present application provides a speech processing method of an early education robot and the early education robot.
[0006] In a first aspect, the present application provides a speech processing method of an early education robot, which comprises: receiving and recognizing a speech instruction input by a user, and when the instruction is recognized as a teaching instruction, showing the user the teaching content in the teaching instruction;
[0007] After the teaching content is shown, receiving a feedback speech of the user and collecting a feedback image of the user;
[0008] The feedback speech and the feedback image are subjected to content consistency verification, and after the consistency verification is passed, it is judged whether it is consistent with the teaching content, an early education processing result is obtained, and the user is shown, wherein the early education processing result includes feedback correct or feedback error.
[0009] In a second aspect, the present application provides an early education robot, which comprises:
[0010] A speech instruction analysis module is configured to receive and recognize a speech instruction input by a user, and when the instruction is recognized as a teaching instruction, show the user the teaching content in the teaching instruction;
[0011] A feedback information collection module is configured to receive a feedback speech of the user and collect a feedback image of the user after the teaching content is shown;
[0012] The feedback information processing module is configured to perform content consistency verification on the feedback voice and the feedback image, determine whether the feedback voice and the feedback image are consistent with the teaching content after the consistency verification passes, obtain early education processing results, and display the early education processing results to the user, wherein the early education processing results include feedback correct or feedback error.
[0013] The one or more technical solutions provided in the present application have at least the following technical effects or advantages:
[0014] The present application provides an early education robot voice processing method and an early education robot. The input voice is divided into adult voice and child voice through voice category recognition, and further combined with voice content recognition, which significantly simplifies the voice processing process and improves the accuracy of instruction recognition. Compared with the traditional voice recognition method, the present application can quickly identify teaching voice, avoid complex calculation of irrelevant voice, and reduce the consumption of computing power. The present application also uses the feature consistency of voice and child action image for early education teaching and double verification, effectively improves the interaction diversity of early education robot in use, and improves the enthusiasm and effect of early education of children. The present application achieves the technical effects of improving the diversity and teaching effect of early education robot and child interaction, and reducing the consumption of voice recognition computing power. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0016] Figure 1 The flowchart of the voice processing method of the early education robot provided by the embodiment of the present application is shown in the figure.
[0017] Figure 2 The logic diagram of the voice processing method of the early education robot provided by the embodiment of the present application is shown in the figure.
[0018] Figure 3 The structure diagram of the early education robot provided by the embodiment of the present application is shown in the figure.
[0019] In the drawings, the components represented by each number are described as follows:
[0020] The voice instruction analysis module 11, the feedback information acquisition module 12, and the feedback information processing module 13. DETAILED DESCRIPTION
[0021] The application provides a voice processing method of an early education robot and the early education robot, and is used for solving the problems of large voice recognition computing power consumption and poor early education effect of the early education robot in the prior art.
[0022] The technical solutions in the embodiments of the application will be clearly and completely described in connection with the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the application.
[0023] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server comprising a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or modules not clearly listed or inherent to the process, method, product or device.
[0024] Embodiment one, as shown in Figure 1 and Figure 2 The application provides a voice processing method of an early education robot, wherein the method is applied to a cooperative charging platform, the cooperative charging platform is in communication connection with a central control unit and a plurality of wireless charging devices, and the method comprises the following steps:
[0025] S100: receiving a voice instruction input by a user, performing identification, and when the voice instruction is identified as a teaching instruction, showing teaching content in the teaching instruction to the user;
[0026] In the process of using the early education robot, the voice input of the user is captured in real time by a voice acquisition device, and is converted into a digital voice signal.
[0027] Then, the collected voice signal is processed, including pre-processing of the signal (such as noise reduction and signal enhancement) and feature extraction (such as extraction of voice spectrum and acoustic features). In the identification process, the voice instruction will be classified into multiple types, such as teaching, reward, power on and off, etc. The "teaching instruction" indicates that the robot enters the education mode, and is used to show specific teaching content.
[0028] For example, the input voice instruction is identified by using a voice category recognition algorithm.
[0029] When the voice instruction is identified as a teaching instruction, the teaching content in the teaching instruction, such as digital teaching content, is shown to the user.
[0030] The step S100 in the method provided in the embodiments of the application comprises:
[0031] Receive voice commands input by the user;
[0032] Performing voice category recognition on the voice instruction to obtain a voice category recognition result, wherein the voice category recognition result includes one of an adult voice and a child voice;
[0033] Performing content recognition on the voice command to obtain a voice content recognition result, wherein the voice content recognition result includes one of power on, power off, teaching, and rewarding;
[0034] Combining the voice category recognition result and the voice content recognition result to obtain a voice command recognition result;
[0035] When the voice instruction recognition result includes adult voice and teaching, the voice instruction is determined to be a teaching instruction, and the teaching content in the teaching instruction is displayed to the user.
[0036] In the embodiment of the present application, a voice acquisition device such as a microphone is used to capture the user's voice input signal in real time. The voice signal is initially processed, including noise removal, volume equalization, and digitization. The processed result is used as the voice command to be recognized, providing high-quality data support for subsequent analysis.
[0037] Furthermore, voice commands are identified and classified to determine whether the source is an adult or a child. Directly identifying the content of voice commands would require significant computational power to perform simultaneous voice recognition, as the speech characteristics of adults and children are significantly different. However, first identifying the source of the voice, then performing specific content recognition, can reduce the computational power consumption of voice recognition.
[0038] Speech category recognition is achieved by analyzing the acoustic characteristics of speech (such as frequency, pitch, and speaking rate). This step aims to distinguish between the primary operator (usually an adult) and the non-operator (usually a child) to prevent children from accidentally triggering certain functions. For example, adult speech has a lower pitch and a more stable speaking rate, while children's speech has a higher pitch and greater fluctuations. Deep learning models or classification algorithms such as support vector machines can accurately make this determination and output speech category recognition results. For example, sample speech data from adults and children is collected, and the actual speech content is labeled as adult or child. Supervised training based on classification algorithms such as deep learning models or support vector machines is then performed to identify the speech category.
[0039] After determining the voice category, the specific content in the voice signal is analyzed, and the semantic information of the voice is extracted through natural language processing (NLP) technology to identify the core instruction of the voice. The identification result includes one of different instruction types such as starting, shutting down, teaching, and rewarding. Teaching, for example, plays digital teaching content, and rewarding, for example, plays reward videos such as cartoon videos. For example, the user's voice "start digital teaching" will be analyzed as a teaching instruction, and "reward the child" will be analyzed as a reward instruction. Through this classification, the operation target of the voice instruction is clear.
[0040] Among them, according to the voice category recognition result, the content of the voice instruction is recognized, for example, if it is recognized as adult voice, an adult voice content recognition model trained based on a natural language processing algorithm is used for recognition, and if it is recognized as child voice, a child voice content recognition model trained based on a natural language processing algorithm is used for recognition. In this way, the voice recognition content can be converted from an 8-classification problem to a 2-classification problem plus a 4-classification problem, which can effectively reduce the size of the voice recognition model and reduce the computing power consumed by voice recognition, making it more lightweight and fast. The voice content recognition model trained based on the natural language processing algorithm is a training method in the prior art, which will not be described here.
[0041] The voice instruction recognition result of the current voice instruction is obtained by combining the voice category recognition result and the voice content recognition result. For example, the voice instruction recognition result can be any one of adult start, adult shutdown, adult teaching, adult reward, child start, child shutdown, child teaching, and child reward.
[0042] Further, when the voice instruction recognition result includes adult voice and teaching, the voice instruction is determined to be a teaching instruction, and the teaching content in the teaching instruction is displayed. For example, the voice instruction is "learn 7", and the early education content of the number "7" such as a video is played for early education teaching.
[0043] Through the combination of voice category recognition and voice content recognition, the response ability of the robot to voice instructions is significantly improved, making the recognition more lightweight and reducing the computing power consumed by voice recognition.
[0044] In the voice instruction recognition result includes adult voice and teaching, the voice instruction is determined to be a teaching instruction, and the teaching content in the teaching instruction is displayed to the user, including:
[0045] When the voice instruction recognition result includes adult voice and teaching, the voice instruction is determined to be a teaching instruction, and the teaching node in the voice instruction is identified.
[0046] The teaching content in the teaching node is displayed to the user.
[0047] In the embodiments of the present application, when the voice instruction input by the user is identified as adult voice and teaching, it is determined that the instruction is a teaching instruction.
[0048] On this basis, the teaching nodes in the voice instruction are identified and processed. For example, after identifying the voice instruction "learn the number 7" as teaching and adult voice, the teaching nodes in the voice instruction are identified and processed.
[0049] The teaching node refers to a knowledge content unit in the teaching content, such as "number 7", "English letters A to D", etc. The identification of the teaching node in the teaching instruction can be performed by a natural language recognition algorithm to obtain the teaching content unit therein as a teaching node. The early education robot can be pre-set with early education courses and arranged with multiple teaching nodes for user selection and voice control playback.
[0050] After successfully parsing the teaching node, the teaching content display is entered, the teaching resource corresponding to the teaching node is called, and the user is displayed. The teaching content can include various forms, such as picture display, animation demonstration and voice explanation, and the specific display mode is optimized according to the node type and the acceptance ability of children.
[0051] The embodiments of the present application utilize voice category recognition to ensure accurate positioning of the instruction, utilize teaching node recognition to realize modular display of the teaching content, and improve the intelligence of early education.
[0052] S200: After the teaching content display, receiving the feedback voice of the user and collecting the feedback image of the user;
[0053] In the embodiments of the present application, after the teaching content display, the learning feedback information of the user is obtained. For example, the feedback information of the child after reading the teaching content is collected.
[0054] Among them, the present application collects the voice and action image of the user feedback through voice interaction and action interaction to identify the target of synchronous voice and action interaction. Then the feedback voice of the user is received and the feedback image of the user is collected.
[0055] Exemplarily, after the teaching content display ends, the child can be prompted to perform voice feedback and action feedback, for example, a prompt voice "say the number you just saw and use your hand to show it out" is played, and then the voice feedback of the child is collected, and the video image of the action feedback of the child is collected. For example, the child can be prompted to stand in a specified position area to perform action feedback, thereby improving the quality of feedback image collection.
[0056] S300: content consistency verification is performed on the feedback voice and the feedback image, after the consistency verification passes, it is judged whether it is consistent with the teaching content, an early education processing result is obtained, and the early education processing result is displayed to the user, wherein the early education processing result includes feedback correct or feedback error.
[0057] As shown in Figure 2 the embodiment of the present application, after receiving the feedback voice and the feedback image of the user, the content of the user feedback in the two kinds of feedback information is compared through consistency verification, to ensure the consistency of the voice feedback and the image feedback in semantics and behavior. That is, whether the number expressed in the voice of the child feedback and the number expressed in the action image of the child feedback is consistent is judged, and then whether the child has mastered the teaching content is judged.
[0058] After the consistency verification passes, whether the feedback content in the feedback voice and the feedback image is the same as the displayed teaching content is compared, the correctness of the user feedback is judged, an early education processing result is obtained, and the early education processing result is displayed to the user. The early education processing result includes feedback correct or feedback error. For example, the teaching content is the number 7, the content expressed by the user in the feedback voice and the feedback image is 7, and the early education processing result of feedback correct is obtained. Otherwise, if the content expressed by the user in the feedback voice and the feedback image is not 7, the early education processing result of feedback error is obtained.
[0059] The early education processing result is displayed to the user, the child is prompted, and the teaching and feedback of a complete teaching content are completed.
[0060] The method provided in the embodiment of the present application includes the following steps S300:
[0061] The feedback voice is converted into an image to obtain a voice image;
[0062] The voice image and the feedback image are subjected to content consistency verification to obtain a consistency verification result;
[0063] When the consistency verification result is inconsistent, feedback error information is displayed;
[0064] When the consistency verification result is consistent, consistent feedback information is obtained, it is judged whether the consistent feedback information is consistent with the teaching content, an early education processing result of feedback correct or feedback error is obtained, and the early education processing result is displayed to the user.
[0065] In the embodiment of the present application, after receiving the feedback voice of the user, the voice signal is processed by a voice image conversion module to convert it into a voice image easy to compare. In this way, the feedback voice is converted into image data of the voice image, the feedback image is also image data, and the consistency of the two image data is verified to improve the verification accuracy and efficiency.
[0066] In an embodiment of the present application, converting the feedback speech into an image to obtain a speech image includes: normalizing the speech signal sequence in the feedback speech to obtain a normalized speech signal sequence;
[0067] Using a Gram angle field, performing polar coordinate conversion on the normalized speech signal sequence to obtain a polar coordinate sequence;
[0068] A GASF graph or a GADF graph is generated according to the polar coordinate sequence as a speech image.
[0069] In the embodiment of the present application, the Gram angular field algorithm is used to perform image conversion processing on the feedback speech.
[0070] In an embodiment of the present application, after receiving the user's feedback voice, the voice signal is first normalized. Normalization refers to adjusting the original voice signal to a standardized range through mathematical methods to eliminate the influence of volume fluctuations and background noise, making the signal more suitable for subsequent feature extraction and analysis. Specifically, the system will scale the amplitude range of the voice signal and map it to [-1,1] or other specified ranges, while filtering out environmental noise through a denoising algorithm to improve the clarity of the signal. This step generates a normalized voice signal sequence.
[0071] The Gramian Angular Field (GAF) algorithm is used to process the normalized speech signal sequence, converting it from a one-dimensional time series to a two-dimensional image representation. Specifically, each data point in the normalized speech signal sequence is mapped to polar coordinates to generate a polar coordinate sequence. For example, the signal amplitude is mapped to the polar radius (i.e., the distance from the origin), and the signal's temporal evolution is mapped to the polar angle (i.e., the angle relative to the polar axis). This representation transforms the time series into intuitive geometric features, making it more suitable for graphical representation and subsequent pattern analysis. Gramian Angular Field plots are then generated, including Gramian Angular Summation Field (GASF) and Gramian Angular Difference Field (GADF). The GASF plot represents the cumulative trend of the time series by calculating the cosine of the polar angle, while the GADF plot reflects the rate of change of the time series by calculating the sine of the polar angle. These two plots capture the dynamic characteristics of the sequence from different dimensions, providing diverse input for subsequent pattern recognition.
[0072] Finally, the generated GASF or GADF graph is used as a two-dimensional image representation of the speech signal, which can intuitively reflect the amplitude and relative change trend of the speech signal over time. Select one of the GASF or GADF graphs as the speech image.
[0073] The voice image not only retains the time correlation and dynamic characteristics of the user feedback voice signal, but also provides the same data type in the form of an image for subsequent consistency verification.
[0074] After the feedback voice image is converted into a voice image, consistency verification is performed on the voice image and the feedback image to determine whether the content expressed by the user in the voice image and the feedback image is consistent, and a consistency verification result is obtained.
[0075] In the embodiments of the present application, the consistency verification of the voice image and the feedback image is performed to obtain a consistency verification result, which includes:
[0076] A sample voice image set and a sample feedback image set are collected, and sample voice image feature vectors and sample feedback image feature vectors are labeled according to the content in the sample voice image and the sample feedback image, to obtain a sample voice image feature vector set and a sample feedback image feature vector set.
[0077] A consistency verification model is constructed using a twin network, wherein the consistency verification model includes a first feature recognition path and a second feature recognition path, and a comparator connected to the output layers of the first feature recognition path and the second feature recognition path.
[0078] The sample voice image set, the sample feedback image set, the sample voice image feature vector set, and the sample feedback image feature vector set are used to supervise the training of the consistency verification model until convergence, wherein the comparator calculates the Euclidean distance between the voice image feature vector and the feedback image feature vector output by the output layers of the first feature recognition path and the second feature recognition path, and performs consistency discrimination supervision training.
[0079] The voice image and the feedback image are input into the trained consistency verification model, and the consistency verification result is obtained by identifying and outputting.
[0080] In the embodiments of the present application, a twin network model is used to perform consistency verification of the voice image and the feedback image.
[0081] Specifically, first, the consistency verification model for consistency verification of the voice image and the feedback image is trained, and the consistency verification model is trained through a twin neural network.
[0082] First, a sample voice image set and a sample feedback image set are collected. The sample voice image set is obtained by converting a standard voice signal into an image (such as a GASF or GADF image), and the sample feedback image set is obtained by collecting and processing the feedback actions or expression images in the user interaction process. Both are used to train the consistency verification model.
[0083] Further, in the sample data preprocessing stage, according to the content of the collected sample speech image and feedback image, the respective feature vectors are extracted and labeled. The feature vector is a high-dimensional data representation form used to describe the features of the image. In the embodiments of the present application, the feature vector reflects the specific content in the speech image and feedback image, such as the number expressed by the user. For example, the feature vector is labeled according to the content of the number 7 in the speech image and feedback image. In this way, the sample speech image feature vector set and the sample feedback image feature vector set are obtained by labeling.
[0084] Further, a consistency verification model is constructed using a Siamese network. The Siamese network is a special neural network structure, which includes two feature recognition paths for processing speech images and feedback images respectively. Among them, the two feature recognition paths are trained using convolutional neural networks and LSTM networks respectively to extract spatial information and temporal information in the speech image and feedback image, and output the respective sample speech image feature vectors and sample feedback image feature vectors. The consistency verification model also includes a comparator connected to the output layers of the first and second feature recognition paths, which is used to calculate the Euclidean distance between the speech image feature vectors and feedback image feature vectors output by the output layers of the first and second feature recognition paths, and then determine whether it is less than the Euclidean distance threshold, and further determine whether the content in the speech image and feedback image is consistent. If it is less than or equal to, it is consistent, otherwise it is not consistent.
[0085] The Euclidean distance threshold can be set according to the Euclidean distance between the sample speech image feature vectors and the sample feedback image feature vectors with the same content, for example, the mean of the Euclidean distances of multiple groups of sample speech image feature vectors and sample feedback image feature vectors with the same content is set as the Euclidean distance threshold.
[0086] In the model training process, a supervised learning method is used to input the sample speech image set, the sample feedback image set, and the corresponding feature vector set into the Siamese network, to identify the speech image feature vector and the feedback image feature vector, to calculate the Euclidean distance between the speech image feature vector and the feedback image feature vector, to determine whether it is less than the Euclidean distance threshold, to generate a consistency discrimination result, and to compare it with the sample labeled consistency label (such as consistent or inconsistent). The training goal is to minimize the error between the output and the labeled value, thereby optimizing the parameters of the Siamese network. After multiple iterations, the model finally converges, that is, the consistency discrimination accuracy of its output reaches the expected standard, for example, the accuracy rate reaches 95%.
[0087] After the consistency verification model training is completed, the currently converted speech image and feedback image are input into the trained consistency verification model, the feature vectors in the speech image and feedback image are identified and extracted, the Euclidean distance is calculated for judgment, and the consistency verification result is output.
[0088] Through the above, a consistency verification model based on twin network training can accurately determine the degree of match between voice feedback and image feedback. This technology not only improves the processing capabilities of multimodal feedback, but also significantly enhances the robot's understanding of user intent. Combined with user action interaction, it can improve the diversity of early childhood education recognition interactions and the effectiveness of early childhood education.
[0089] In an embodiment of the present application, when the consistency check result is inconsistent, an error message is displayed to the user. For example, if the user's voice answer does not match the number expressed by the gesture, a prompt "Please try again" may be displayed.
[0090] When the consistency check result is consistent, consistent feedback information is generated, that is, the content of user feedback in the voice image and the feedback image, such as consistent numbers, and the next step of teaching content comparison is entered to determine whether the consistent feedback information is consistent with the teaching content, obtain the processing results, and display them to the user.
[0091] In the embodiment of the present application, when the consistency check result is consistent, obtaining consistent feedback information, determining whether the consistent feedback information is consistent with the teaching content, obtaining the early education processing result, and displaying it to the user include:
[0092] When the consistency check result is consistent, obtaining consistent feedback information in the voice image and the feedback image;
[0093] Determining whether the consistent feedback information is consistent with the teaching content;
[0094] If so, the correct early childhood education processing result is obtained; if not, the incorrect early childhood education processing result is obtained and displayed to the user.
[0095] In an embodiment of the present application, when the consistency check result is consistent, consistent feedback information contained in the speech image and the feedback image is obtained. For example, a convolutional neural network is used to recognize the feedback image and identify the numbers contained therein as consistent feedback information, or based on the feature vector of the speech image and the feature vector of the feedback image, the content contained in the feature vector, such as the represented number, is extracted as consistent feedback information.
[0096] The consistent feedback information is compared with the displayed teaching content. It is judged whether the content of the consistent feedback information is consistent with the teaching content, for example, whether the numbers in the consistent feedback information are the numbers taught in the teaching content, for example, whether they are both 7. If yes, a correct feedback early education processing result is obtained, and the correct feedback information is displayed to the user to prompt the user that the feedback is correct, for example, by voice broadcast "correct answer!" or display of a reward animation. If no, an incorrect feedback early education processing result is obtained, and the incorrect feedback information is displayed to the user to prompt the user that the feedback is incorrect, for example, by voice prompt "please try again".
[0097] Through the above steps, a complete process from consistency verification of the collected information from voice feedback and action feedback to feedback comparison and then result display is realized. The method not only improves the processing capability of the early education robot for user feedback, but also enhances the interactive experience of the user through accurate correct and incorrect feedback.
[0098] In the embodiments of the present application, after receiving the voice instruction input by the user and performing recognition, the method further includes:
[0099] When the voice instruction recognition result includes adult voice or child voice, and the power-on or power-off, it is judged that the voice instruction is a power-on or power-off instruction, and the power-on or power-off is performed.
[0100] When the voice instruction recognition result includes adult voice, and the teaching or reward, it is judged that the voice instruction is a teaching instruction or a reward instruction, and the teaching or reward is performed.
[0101] When the voice instruction recognition result includes child voice, and the teaching or reward, it is judged that the voice instruction is an invalid instruction, and a prompt information is generated to prompt.
[0102] In the embodiments of the present application, in order to ensure the active use of the early education robot and avoid the use of incorrect instructions by children affecting the display of teaching or reward content, specific voice instructions are set.
[0103] When the voice instruction recognition result shows that the voice category is adult voice or child voice, and the content is power-on or power-off, it is determined that the instruction is a power-on or power-off instruction, and the power-on or power-off is performed according to the instruction. That is, adults and children can control the power-on and power-off of the early education robot.
[0104] When the voice instruction recognition result includes adult voice, and the teaching or reward, it is judged that the voice instruction is a teaching instruction or a reward instruction, and the teaching or reward is performed, for example, playing teaching content or reward content such as an animation. That is, adults can control the teaching and reward content of the early education robot.
[0105] If the voice command recognition results include children's voices, as well as teaching or rewarding content, the voice command is determined to be invalid and a prompt message is generated to remind the user that children cannot control the early childhood education robot's teaching and reward content. This prevents children from interfering with the teaching or reward content and affecting the effectiveness of early childhood education. Prompt messages may include voice prompts such as "Please ask a parent to operate" or "This function requires parental assistance."
[0106] Through the above steps, different voice categories and contents can be intelligently distinguished to ensure accurate triggering of functions and effective protection against misoperation.
[0107] In summary, the embodiments of the present application have at least the following technical effects:
[0108] The embodiment of the present application proposes a method for voice processing of an early childhood education robot, which divides the input voice into adult voice and child voice through voice category recognition, and further combines voice content recognition, which significantly simplifies the voice processing process and improves the accuracy of command recognition. Compared with traditional voice recognition methods, the present invention can quickly identify teaching voices, avoid complex calculations of irrelevant voices, and reduce computing power consumption. The present application also uses the feature consistency between voice and children's action images for early childhood education and double verification, effectively improving the diversity of interaction when using early childhood education robots, and improving children's early childhood education enthusiasm and early childhood education effects. The present application achieves the technical effect of improving the diversity and teaching effect of interactions between early childhood education robots and children, and reducing the consumption of voice recognition computing power.
[0109] Example 2, as Figure 3 As shown, based on the same inventive concept as the speech processing method of an early childhood education robot provided in Example 1, an embodiment of the present invention further provides an early childhood education robot, comprising:
[0110] The voice instruction parsing module 11 is used to receive the voice instruction input by the user, identify it, and when it is identified as a teaching instruction, display the teaching content in the teaching instruction to the user;
[0111] A feedback information collection module 12 is used to receive user feedback voice and collect user feedback images after the teaching content is presented;
[0112] The feedback information processing module 13 is used to perform content consistency verification on the feedback voice and feedback image. After the consistency verification is passed, it is determined whether they are consistent with the teaching content, and the early education processing result is obtained and displayed to the user, wherein the early education processing result includes correct feedback or incorrect feedback.
[0113] In one embodiment, the voice instruction parsing module 11 is further configured to:
[0114] Receive voice commands input by the user;
[0115] performing voice category recognition on the voice instruction to obtain a voice category recognition result, wherein the voice category recognition result comprises one of adult voice and child voice;
[0116] performing content recognition on the voice instruction to obtain a voice content recognition result, wherein the voice content recognition result comprises one of booting up, shutting down, teaching, and rewarding;
[0117] combining the voice category recognition result and the voice content recognition result to obtain a voice instruction recognition result;
[0118] when the voice instruction recognition result comprises adult voice and teaching, determining that the voice instruction is a teaching instruction, and showing teaching content in the teaching instruction to the user.
[0119] When the voice instruction recognition result comprises adult voice and teaching, determining that the voice instruction is a teaching instruction, and showing teaching content in the teaching instruction to the user, comprises:
[0120] when the voice instruction recognition result comprises adult voice and teaching, determining that the voice instruction is a teaching instruction, and identifying a teaching node in the voice instruction;
[0121] showing teaching content in the teaching node to the user.
[0122] In one embodiment, the feedback information processing module 13 is further configured to: perform image conversion on the feedback voice to obtain a voice image;
[0123] performing content consistency verification on the voice image and the feedback image to obtain a consistency verification result;
[0124] when the consistency verification result is inconsistent, showing feedback error information;
[0125] when the consistency verification result is consistent, obtaining consistent feedback information, determining whether the consistent feedback information is consistent with the teaching content, obtaining an early education processing result of feedback being correct or feedback being incorrect, and showing the early education processing result to the user.
[0126] When the feedback voice is converted into a voice image, the method comprises:
[0127] performing normalization processing on a voice signal sequence in the feedback voice to obtain a normalized voice signal sequence;
[0128] performing polar coordinate conversion on the normalized voice signal sequence using a Gram angle field to obtain a polar coordinate sequence;
[0129] According to the polar coordinate sequence, a GASF graph or a GADF graph is generated as a speech image.
[0130] Wherein, the speech image and the feedback image are content consistency checked to obtain a consistency checking result, including:
[0131] A sample speech image set and a sample feedback image set are collected, and sample speech image feature vectors and sample feedback image feature vectors are labeled according to the content in the sample speech image and the sample feedback image, to obtain a sample speech image feature vector set and a sample feedback image feature vector set;
[0132] A consistency checking model is constructed by using a twin network, wherein the consistency checking model includes a first feature recognition path and a second feature recognition path, and a comparator connecting the output layers of the first feature recognition path and the second feature recognition path;
[0133] The sample speech image set, the sample feedback image set, the sample speech image feature vector set and the sample feedback image feature vector set are used to supervise the training of the consistency checking model until convergence, wherein the comparator calculates the Euclidean distance between the speech image feature vectors and the feedback image feature vectors output by the output layers of the first feature recognition path and the second feature recognition path, and performs consistency discrimination supervision training;
[0134] The speech image and the feedback image are input into the trained consistency checking model, and the output is identified to obtain a consistency checking result.
[0135] Wherein, when the consistency checking result is consistent, consistent feedback information is obtained, it is judged whether the consistent feedback information is consistent with the teaching content, an early education processing result is obtained, and the user is displayed, including:
[0136] When the consistency checking result is consistent, the consistent feedback information in the speech image and the feedback image is obtained;
[0137] It is judged whether the consistent feedback information is consistent with the teaching content;
[0138] If yes, a feedback correct early education processing result is obtained, and if no, a feedback error early education processing result is obtained, and the user is displayed.
[0139] Wherein, a speech instruction input by a user is received and recognized, and then the method further includes:
[0140] When the speech instruction recognition result includes adult speech or child speech, and when starting up or shutting down, it is judged that the speech instruction is a start-up or shutdown instruction, and the start-up or shutdown is performed.
[0141] When the voice instruction recognition result includes adult voice, and teaching or rewarding, the voice instruction is determined as a teaching instruction or a rewarding instruction, and teaching or rewarding is performed.
[0142] When the voice instruction recognition result includes child voice, and teaching or rewarding, the voice instruction is determined as an invalid instruction, and a prompt information is generated to prompt.
[0143] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned describes the specific embodiments of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0144] The above-mentioned is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0145] The present application and the drawings are only exemplary descriptions of the present application, and are considered to cover any and all modifications, changes, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art can make various modifications and changes to the present application without departing from the scope of the present application. Thus, if these modifications and changes of the present application belong to the scope of the present application and its equivalent technology, the present application intends to include these modifications and changes.
Claims
1. A speech processing method for an early childhood education robot, characterized in that: The method comprises: Receive a voice command input by a user, perform recognition, and when recognized as a teaching command, display the teaching content in the teaching command; After the teaching content is presented, receiving user feedback voice and collecting user feedback images; Performing consistency check on the feedback voice and feedback image, and after passing the consistency check, determining whether they are consistent with the teaching content, obtaining processing results, and displaying them, including: Converting the feedback voice into an image to obtain a voice image includes: performing normalization processing on the speech signal sequence in the feedback speech to obtain a normalized speech signal sequence; Using a Gram angle field, performing polar coordinate conversion on the normalized speech signal sequence to obtain a polar coordinate sequence; generating a GASF graph or a GADF graph as a speech image according to the polar coordinate sequence; Performing a consistency check on the voice image and the feedback image to obtain a consistency check result includes: Collecting a set of sample speech images and a set of sample feedback images, and labeling the sample speech image feature vectors and the sample feedback image feature vectors according to the contents of the sample speech images and the sample feedback images, to obtain a set of sample speech image feature vectors and a set of sample feedback image feature vectors; A twin network is used to construct a consistency verification model, wherein the consistency verification model includes a first feature recognition path and a second feature recognition path, and a comparator connecting the output layers of the first feature recognition path and the second feature recognition path; Using the sample speech image set, the sample feedback image set, the sample speech image feature vector set, and the sample feedback image feature vector set, supervised training of the consistency verification model until convergence, wherein the comparator calculates the Euclidean distance between the speech image feature vector and the feedback image feature vector output by the output layer of the first feature recognition path and the second feature recognition path to perform consistency discrimination supervised training; Inputting the speech image and the feedback image into the trained consistency verification model, recognizing the output to obtain a consistency verification result; When the consistency check result is inconsistent, display feedback error information; When the consistency check result is consistent, consistent feedback information is obtained, and it is determined whether the consistent feedback information is consistent with the teaching content, and a processing result is obtained and displayed.
2. The speech processing method of the early childhood education robot according to claim 1, characterized in that: Receive the voice command input by the user, identify it, and when it is identified as a teaching voice command, display the teaching content in the teaching command, including: Receive voice commands input by the user; Performing voice category recognition on the voice instruction to obtain a voice category recognition result, wherein the voice category recognition result includes adult voice and child voice; Performing content recognition on the voice command to obtain a voice content recognition result, wherein the voice content recognition result includes power on, power off, teaching, and rewarding; Combining the voice category recognition result and the voice content recognition result to obtain a voice command recognition result; When the voice instruction recognition result includes adult voice and teaching, the voice instruction is determined to be a teaching instruction, and the teaching content in the teaching instruction is displayed.
3. The speech processing method of the early childhood education robot according to claim 2, characterized in that: When the voice instruction recognition result includes adult voice and teaching, determining that the voice instruction is a teaching instruction, and displaying the teaching content in the teaching instruction include: When the voice instruction recognition result includes adult voice and teaching, determining that the voice instruction is a teaching instruction, and identifying a teaching node in the voice instruction; Display the teaching content in the teaching node.
4. The speech processing method of the early childhood education robot according to claim 1, characterized in that: When the consistency check result is consistent, obtaining consistent feedback information, determining whether the consistent feedback information is consistent with the teaching content, obtaining a processing result, and displaying it, including: When the consistency check result is consistent, obtaining consistent feedback information in the voice image and the feedback image; Determining whether the consistent feedback information is consistent with the teaching content; If so, the correct information is displayed; if not, the wrong information is displayed.
5. The speech processing method of the early childhood education robot according to claim 2, characterized in that: After receiving and recognizing a voice command input by a user, the method further includes: When the voice command recognition result includes an adult voice or a child voice, and a power on or power off command, determining that the voice command is a power on or power off command, and turning the device on or off; When the voice instruction recognition result includes an adult voice, and teaching or rewarding, determining that the voice instruction is a teaching instruction or a rewarding instruction, and performing teaching or rewarding; When the voice instruction recognition result includes a child's voice, as well as teaching or rewarding, the voice instruction is determined to be an invalid instruction, and a prompt message is generated for prompting.
6. An early childhood education robot, characterized in that: The early childhood education robot comprises: The voice command parsing module is used to receive the voice command input by the user, identify it, and when it is identified as a teaching command, display the teaching content in the teaching command; A feedback information collection module, configured to receive user feedback voice and collect user feedback images after the teaching content is presented; The feedback information processing module is used to perform consistency verification on the feedback voice and feedback image. After the consistency verification is passed, it is determined whether they are consistent with the teaching content, and the processing results are obtained and displayed.
7. An electronic device, characterized in that: include: Memory for storing computer software programs; A processor is used to read and execute the computer software program, thereby implementing the steps of the speech processing method of the early childhood education robot as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Early education robot speech interaction education system and method
CN108109622A
Children language rehabilitation training method and system based on voice and mouth shape recognition
CN117672024A