Early education terminal interaction system based on voice acquisition and recognition

By employing three-level voice verification and voiceprint feature matching, the problems of authorized user identification and permission management in early education terminal devices have been solved, enabling accurate voice interaction and secure operation of early education terminal devices, thereby improving device stability and user experience.

CN121528211APending Publication Date: 2026-02-13SHANDONG TONGQI WANJIANG EDUCATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511690299.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing early childhood education terminal devices lack a precise voiceprint authorization mechanism in terms of voice interaction, making it impossible to distinguish between authorized and unauthorized users, leading to problems such as misoperation and abuse of permissions. At the same time, the voice data processing is crude, reducing the accuracy and security of interaction, and lacking dynamic monitoring and feedback mechanisms.

Method used

By filtering invalid and expired voice messages through three-level verification, and combining voiceprint feature extraction with matching of authorized voice feature library, we can achieve accurate verification of authorized users, allocate differentiated permissions, build a voiceprint-permission association library, optimize the speech recognition logic for young children, build interactive recognition and capability index, and carry out rational operation and maintenance.

Benefits of technology

It improves the accuracy of recognizing children's voice commands, prevents unauthorized user operations, ensures safety, provides differentiated permission management, and enhances device stability and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528211A_ABST
    Figure CN121528211A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of early education terminal interaction, in particular to an early education terminal interaction system based on voice acquisition and recognition, which comprises a terminal interaction center, a voice acquisition module, a voice processing module, a control authorization module, a voice recognition module, a historical experience module and a content output module, according to the invention, through three-level verification of duration, signal-to-noise ratio and distortion frame ratio, invalid and invalid voices are filtered, the influence of interference signals on interaction is reduced, meanwhile, recognition logic is optimized for infant fuzzy voices, exclusive recognition rate calculation is combined, the recognition accuracy of infant voice instructions is improved, the core demand of an early education scene is adapted, and the early education efficiency is improved. Furthermore, through voiceprint feature extraction and authorized voice feature library matching, accurate verification of authorized users is achieved, operation of unauthorized users is eradicated, use safety is guaranteed, through construction of a voiceprint-permission association library, differentiated permissions are distributed to different users, instructions are bound with the lowest permission level, unauthorized operation is avoided, and content can be played accurately.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of early education terminal interaction, and particularly relates to an early education terminal interaction system based on voice collection and recognition. BACKGROUND

[0002] Currently, as the core tool for early education terminal devices to assist children in early education in families and early education institutions, the interaction mode is upgrading from traditional key and touch screen to voice interaction. However, the conventional voice interaction type early education terminal still has multi-dimensional technical shortcomings, which are difficult to meet the safe, accurate and personalized early education needs. The specific problems are as follows: On the one hand, there is a lack of accurate voiceprint authorization mechanism, which cannot distinguish between authorized users (such as children and parents) and unauthorized users, and is prone to misoperation or abuse of authority. On the other hand, the voice data processing is extensive, and the timeliness and effectiveness of the voice are not checked hierarchically, resulting in invalid voice and distorted voice interfering with the accuracy of interaction, thereby reducing the efficiency of command recognition. At the same time, it is also impossible to allocate differentiated operation permissions according to the user identity (children, parents, temporary users), which poses a security risk and lacks a dynamic monitoring and feedback mechanism for interaction performance to discover low recognition accuracy and long response time and other hidden problems. Therefore, a solution is proposed. SUMMARY

[0003] The purpose of the present application is to provide an early education terminal interaction system based on voice collection and recognition. Through three-level verification, invalid and ineffective voice is filtered to reduce the influence of interference signals on interaction. At the same time, the recognition logic is optimized for children's ambiguous voice to improve the recognition accuracy of children's voice commands. Furthermore, through voiceprint feature extraction and matching with the authorized voice feature library, accurate verification of authorized users is achieved to prevent unauthorized user operation. Differentiated permissions are allocated for different users, and commands are bound to the lowest permission level to avoid unauthorized operation, so as to accurately play content. Based on key indicator analysis feedback, performance risks are warned in advance to ensure stable operation of the device through rational operation and maintenance and improve user experience.

[0004] The purpose of the present application can be achieved by the following technical solution: an early education terminal interaction system based on voice collection and recognition, comprising a terminal interaction center, a voice collection module, a voice processing module, a control authorization module, a voice recognition module, a historical experience module and a content output module; The terminal interaction center is used to store an authorized voice feature library and a candidate voice segment; The voice collection module is used to collect voice samples of a control authorization object (such as a target child or a parent), and to call a pre-set feature extraction algorithm to extract voiceprint features from the voice samples of the control authorization object based on the pre-set feature extraction algorithm to generate an authorized voice feature library; The control authorization module is used for voiceprint-authorization evaluation analysis on the candidate voice segment, matching analysis of the obtained extracted voiceprint features and the authorized voice feature library, and obtaining a check pass or check fail result. When the check passes, the voice recognition module is used for voiceprint feature extraction and instruction-authorization matching analysis on the candidate voice segment, matching of the obtained current authorization level and the minimum authorization level, and obtaining a check fail or pass signal result. The historical experience module is used for historical experience feedback analysis on the collected early education interactive device interaction recognition data and interaction performance data, discriminant processing of the obtained proportion value, and obtaining an operation and maintenance signal or a regular signal result.

[0005] Preferably, the analysis process of the voice processing module is as follows: T1: real-time acquisition of user voice data, acquisition of voice segment duration of the user voice data, and discriminant processing of the voice segment duration; if the voice segment duration is less than a preset voice segment duration threshold, it is determined as invalid voice; if the voice segment duration is greater than or equal to the preset voice segment duration threshold, it is determined as valid voice; T2: when the user voice data is valid voice, further acquire the signal-to-noise ratio and the distortion frame proportion of the voice data, and discriminant process the signal-to-noise ratio and the distortion frame proportion to obtain a failure voice or a sample voice result.

[0006] Preferably, when the user voice data is sample voice, the user voice data is divided into a frame, 10ms frame movement according to a preset duration, the short-time energy intensity of each frame is calculated, and the set limit interval YYmax and YYmin is called; if the short-time energy intensity is higher than YYmax, the corresponding frame is determined as strong voice; if the short-time energy intensity is lower than YYmin, the corresponding frame is determined as non-voice; if the short-time energy intensity is between YYmax and YYmin, the corresponding frame is determined as weak voice. When the continuous 3 frames of signals are strong voice, it is marked as the starting point of the voice segment; after obtaining the starting point of the voice segment, if the frame is non-voice and the duration is greater than or equal to a preset duration, it is marked as the end point of the voice segment, and the signal between the starting point and the end point is extracted as the candidate voice segment.

[0007] Preferably, the analysis process of the control authorization module is as follows: based on a pre-set feature extraction algorithm, voiceprint features are extracted from the candidate voice segment of the user, the extracted voiceprint features are matched with the authorized voice feature library, and a check pass or check fail result is obtained.

[0008] Preferably, the analysis process of the voice recognition module is as follows: Based on the extracted voiceprint features and the voiceprint-privilege association library, the privilege level of the current control authorized object is obtained, if the privilege level is the temporary interaction level L2, the operation time of the current user is obtained, and the operation time is matched with the valid time period to obtain the verification validity or the prompt sound "the authorization time limit has expired, please adjust"; The construction process of the voiceprint-privilege association library: after the parent completes the voiceprint input of the control authorized object, a unique voiceprint feature identifier is assigned to each voiceprint, and the corresponding privilege level is bound.

[0009] Preferably, when the verification is valid or the basic interaction level or the control management level, the candidate voice segment is input into the pre-set instruction recognition model, and the text instruction is output; The text instruction is matched with the pre-set instruction-privilege table, the lowest privilege level corresponding to the input text instruction is input, the current privilege level is matched with the lowest privilege level, and the pass signal or the check failure result is obtained.

[0010] Preferably, the analysis process of the history experience module is as follows: The interaction recognition data of the daily early education interaction device is obtained, the interaction recognition data includes the text instruction recognition accuracy (the proportion of correct matching of the text instruction and the played content) and the infant fuzzy voice recognition rate (the recognition accuracy of the text instruction of the control authorized object for the infant); The sum value of the product of the text instruction recognition accuracy and the infant fuzzy voice recognition rate and the corresponding weight coefficient is set as the interaction recognition index; At the same time, the interaction performance data of the daily early education interaction device is obtained, the interaction performance data includes the total instruction response time delay and the voiceprint check pass rate, and the interaction ability index is calculated based on (1-instruction total response time delay) x corresponding weight coefficient + voiceprint check pass rate x corresponding weight coefficient, and the interaction recognition index and the interaction ability index are discriminated to obtain the performance stability signal or the risk signal; Based on the discrimination results of the daily interaction recognition index and the interaction ability index, a day-result management table is constructed, and the proportion value of the risk signal corresponding to the early education interaction device for 7 consecutive days (if there is no use in the consecutive days, the next day is planned to the previous day until 7 consecutive days are met) is obtained based on the day-result management table; The proportion value is discriminated to obtain the normal signal or the operation and maintenance signal result.

[0011] The beneficial effects of the present application are as follows: (1) The application is verified by three levels of time length, signal-to-noise ratio, and distortion frame proportion, filters invalid and invalid voice, reduces the influence of interference signals on interaction, optimizes the recognition logic for children's fuzzy voice, combines with exclusive recognition rate calculation, improves the recognition accuracy of children's voice instructions, adapts to the core needs of early education scene, and further extracts the voiceprint features and matches the authorized voice feature library to realize accurate verification of authorized users and prevent unauthorized user operation to ensure use safety.

[0012] (2) The application also constructs a voiceprint-permission association library, allocates differentiated permissions for different users, and binds instructions with the lowest permission level to avoid overreach operation, so as to accurately play content, and constructs an interactive recognition index and an interactive ability index to comprehensively quantify key indicators such as recognition accuracy, response time delay, and verification pass rate, analyzes feedback based on key indicators to early warn performance risks, reasonably operates to ensure stable operation of the equipment, and improves user experience. BRIEF DESCRIPTION OF DRAWINGS

[0013] The application will be further described below with reference to the accompanying drawings; Fig. 1 is a reference diagram of the first embodiment of the application; Fig. 2 is a reference diagram of the second embodiment of the application; Fig. 3 is a control authorization module analysis reference diagram of the application. DETAILED DESCRIPTION

[0014] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the application.

[0015] In this document, referring to "embodiments" means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the application. The appearance of this phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive or alternative to other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments; Embodiment one: please refer to Figs. 1 to 3As shown, the application is an early education terminal interaction system based on voice collection and recognition, which comprises a terminal interaction center, a voice collection module, a voice processing module, a control authorization module, a voice recognition module, a historical experience module and a content output module. The terminal interaction center is in bidirectional communication connection with the voice collection module and the historical experience module, in unidirectional communication connection with the voice processing module, in unidirectional communication connection with the control authorization module, in unidirectional communication connection with the content output module, in unidirectional communication connection with the voice recognition module, and in unidirectional communication connection with the content output module. The voice collection module collects voice samples of the control authorization object (such as target children and parents), which include common pronunciation in different scenes (such as calm and fast voice, etc.), and the number is not less than 5, and the length of each is not less than 10 seconds. The pre-set feature extraction algorithm (such as MFCC mel frequency cepstrum coefficient algorithm) is called, and the voiceprint features (including unique features such as timbre, pitch, speed and formant) are extracted from the voice samples of the control authorization object based on the pre-set feature extraction algorithm, an authorized voice feature library is generated, and the obtained authorized voice feature library is sent to the terminal interaction center for storage. The voice processing module is used for time effectiveness and effectiveness analysis of the collected user's voice data, and the specific time effectiveness and effectiveness analysis process is as follows: T1: Real-time acquisition of user's voice data, acquisition of voice segment length of user's voice data, and judgment of voice segment length, if the voice segment length is less than the preset voice segment length threshold, it is determined as invalid voice, the content output module plays the prompt sound "please say it again", if the voice segment length is greater than or equal to the preset voice segment length threshold, it is determined as valid voice; T2: When the user's voice data is valid voice, the signal-to-noise ratio and distortion frame proportion of the voice data are further acquired, and the signal-to-noise ratio and distortion frame proportion are judged, if the signal-to-noise ratio is less than the preset signal-to-noise ratio threshold, or the distortion frame proportion is greater than the preset distortion frame proportion threshold, it is determined as invalid voice, the content output module plays the prompt sound "please say it again", if the signal-to-noise ratio is greater than or equal to the preset signal-to-noise ratio threshold, and the distortion frame proportion is less than or equal to the preset distortion frame proportion threshold, it is determined as sample voice; T3: When the voice data of the user is sample voice, the voice data of the user is divided into a frame according to a preset time length, 10 ms frame movement is performed, the short-time energy intensity (reflecting signal strength) of each frame is calculated, and a set limited interval YYmax and YYmin is called. If the short-time energy intensity is higher than YYmax, the corresponding frame is determined as strong voice. If the short-time energy intensity is lower than YYmin, the corresponding frame is determined as non-voice. If the short-time energy intensity is between YYmax and YYmin, the corresponding frame is determined as weak voice; When the continuous 3 frames of signals are strong voice, the starting point of the voice segment is marked; After the starting point of the voice segment is obtained, if the frame is non-voice and the duration is greater than or equal to a preset duration, the end point of the voice segment is marked; The signal between the starting point and the end point is extracted as a candidate voice segment, and the candidate voice segment is sent to a control authorization module; The control authorization module is used for voiceprint-authorization evaluation analysis on the candidate voice segment. The specific voiceprint-authorization evaluation analysis process is as follows: The voiceprint feature extraction algorithm is used to extract the voiceprint features of the candidate voice segment of the user, and the extracted voiceprint features are matched and analyzed with the authorized voice feature library. If the similarity of the extracted voiceprint features and the voiceprint features in the authorized voice feature library exceeds a preset similarity threshold, it is determined that the voice meets the control authorization object, the verification is passed, and the candidate voice segment is transmitted to the terminal interaction center for storage; If the similarity of the extracted voiceprint features and the voiceprint features in the authorized voice feature library does not exceed the preset similarity threshold, it is determined that the voice does not meet the control authorization object, the verification is not passed, a prompt sound "please use the authorized voice to operate" is played through the content output module, or a text prompt "non-authorized user, unable to execute the instruction" is displayed, and the current interaction is terminated.

[0016] Embodiment two: when the verification is passed, the voice recognition module is used for voiceprint feature extraction and instruction-permission matching analysis on the candidate voice segment. The specific voiceprint feature extraction and instruction-permission matching analysis process is as follows: Based on the extracted voiceprint features and the voiceprint-permission association library, the permission level of the current control authorization object is obtained. If the permission level is temporary interaction level L2, the operation time of the current user is obtained, and the operation time is matched with the valid time period. If the operation time is within the valid time period, it is determined that the verification is valid. If the operation time is not within the valid time period, a prompt sound "authorization expiration, please adjust" is played through the content output module; The construction process of the voiceprint-authority association library: after the parent completes the voiceprint input of the control authorization object (such as a young child, the parent himself, etc.), a unique voiceprint feature identifier is assigned to each voiceprint, and the corresponding authority level is bound (such as a young child bound to a basic interaction level L1, a parent bound to a control management level L3 (L3>L2>L1), and a temporary user bound to a temporary interaction level L2+valid period), and stored in the voiceprint-authority association library; When verifying the valid or basic interaction level or control management level, the candidate voice segment is input to the pre-set instruction recognition model, and the text instruction is output; The text instruction is matched with the pre-set instruction-authority table, and the lowest authority level corresponding to the input text instruction (such as "delete local story" corresponding to L3) is input; The current authority level (basic interaction level L1, control management level L3, and temporary interaction level L2) is matched with the lowest authority level: if the current level ≥ the lowest authority level (such as L3≥L2), a pass signal is generated, and the content output module plays the content corresponding to the text instruction; If the current level < the lowest authority level, the verification fails, and the content output module plays a prompt sound "parent authority is required for operation, please contact the parent"; The historical experience module is used for historical experience feedback analysis of the collected early education interactive device interaction recognition data and interaction performance data. The specific historical experience feedback analysis process is as follows: The daily early education interactive device interaction recognition data is obtained, and the interaction recognition data includes the text instruction recognition accuracy rate (the proportion of correct matching of the text instruction and the played content), such as the young child saying "play the little rabbit obediently" not being misjudged as "play the little cat obediently" and the young child fuzzy voice recognition rate (the recognition accuracy rate of the text instruction of the control authorization object for the young child); The sum of the product of the text instruction recognition accuracy rate and the young child fuzzy voice recognition rate and the corresponding weight coefficient is set as the interaction recognition index, i.e. the product of the text instruction recognition accuracy rate (after standardization) and the corresponding weight coefficient + the product of the young child fuzzy voice recognition rate (after standardization) and the corresponding weight coefficient = the recognition index; The daily early education interactive device interaction performance data is also obtained, and the interaction performance data includes the total instruction response time delay (the sum of the part exceeding the preset time length from the voice text instruction collection to the feedback "whether to execute") and the voiceprint verification pass rate (the proportion of voiceprint feature verification pass when the control authorized user (excluding unauthorized user interference) issues a voice instruction); The interaction ability index is calculated based on (1-the total instruction response time delay (after standardization))×the corresponding weight coefficient + the voiceprint verification pass rate (after standardization)×the corresponding weight coefficient; The interaction recognition index and the interaction capability index are subjected to discrimination processing. If the interaction recognition index is greater than or equal to a preset interaction recognition index threshold value, and the interaction capability index is greater than or equal to a preset interaction capability index threshold value, a performance stability signal is generated. If the interaction recognition index is less than the preset interaction recognition index threshold value, or the interaction capability index is less than the preset interaction capability index threshold value, a risk signal is generated. Based on the discrimination processing results (performance stability signal or risk signal) of the daily interaction recognition index and the daily interaction capability index, a day-result management table is constructed, in which the horizontal axis represents performance stability signals, risk signals, and no use (no use of the early education interactive device on the day), and the vertical axis represents the number of days. Based on the day-result management table, the proportion of risk signals corresponding to the early education interactive device for 7 consecutive days (if there is no use in the consecutive days, the next day is planned to the previous day, for example, if there is no use on the second day, the third day is planned as the second day, until 7 consecutive days are met) is obtained. The proportion is subjected to discrimination processing. If the proportion is less than a preset proportion threshold value, a regular signal is generated. If the proportion is greater than or equal to the preset proportion threshold value, an operation and maintenance signal is generated. The content output module plays a prompt sound corresponding to the regular signal "normal operation and maintenance period management" and a prompt sound corresponding to the operation and maintenance signal "advance operation and maintenance management", and thus performs rational operation and maintenance management on the early education interactive device to improve the use experience of the early education interactive device. As shown above, through three-level verification of time length, signal-to-noise ratio, and distortion frame proportion, invalid and invalid voice is filtered, the influence of interference signals on interaction is reduced, the recognition logic for children's ambiguous voice is optimized, the recognition accuracy of children's voice commands is improved through exclusive recognition rate calculation, the core needs of early education scenes are adapted, and further through voiceprint feature extraction and matching with authorized voice feature library, authorized user accurate verification is realized, non-authorized user operation is prevented, use safety is ensured, a voiceprint-permission association library is constructed, different permissions are allocated to different users (children L1, parents L3, and temporary users L2), commands are bound to the lowest permission level to avoid unauthorized operation, so as to accurately play content, and interaction recognition index and interaction capability index are constructed to comprehensively quantify key indicators such as recognition accuracy, response time delay, and verification pass rate. Based on the analysis and feedback of the key indicators, performance risks are warned in advance, the device is ensured to operate stably through rational operation and maintenance, and the user experience is improved.

[0017] The threshold value is set for result comparison and analysis to determine whether it is good or bad. The size of the threshold value is determined based on large model analysis of sample data and artificial experience, and is also adjusted according to seasonal or rational influence conditions.

[0018] The above merely describes preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art, according to the technical solution and inventive concept of the present application, makes equivalent replacement or change within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.

Claims

1. An early education terminal interactive system based on voice acquisition and recognition, characterized in that, It includes a terminal interaction center, a voice acquisition module, a voice processing module, a control authorization module, a voice recognition module, a history experience module, and a content output module; The terminal interaction center is used to store the authorized speech feature library and candidate speech segments; The voice acquisition module is used to collect voice samples from the authorized control objects (such as target children or parents), and at the same time, it calls the pre-set feature extraction algorithm to extract voiceprint features from the voice samples of the authorized control objects based on the pre-set feature extraction algorithm to generate an authorized voice feature library. The control authorization module is used to perform voiceprint-authorization evaluation and analysis on candidate speech segments, and to match and analyze the extracted voiceprint features with the authorized speech feature library to obtain the result of verification pass or verification fail. When the verification passes, the speech recognition module is used to extract voiceprint features and perform instruction-permission matching analysis on the candidate speech segments, and matches the current permission level with the lowest permission level to obtain a verification failure or success signal. The historical experience module is used to perform historical experience feedback analysis on the collected interaction recognition data and interaction performance data of early childhood education interactive devices, and to process the obtained percentage values ​​to obtain operation and maintenance signals or routine signal results.

2. The early education terminal interactive system based on voice acquisition and recognition according to claim 1, characterized in that, The analysis process of the speech processing module is as follows: T1: Real-time acquisition of user voice data, acquisition of the duration of voice segments, and judgment processing of voice segment duration. If the duration of a voice segment is less than a preset voice segment duration threshold, it is determined to be invalid voice. If the duration of a voice segment is greater than or equal to the preset voice segment duration threshold, it is determined to be valid voice. T2: When the user's voice data is valid, the signal-to-noise ratio and the proportion of distorted frames of the voice data are further obtained, and the signal-to-noise ratio and the proportion of distorted frames are processed to obtain the invalid voice or sample voice results.

3. The early education terminal interactive system based on voice acquisition and recognition according to claim 2, characterized in that, When the user's voice data is sample voice, the user's voice data is divided into frames according to a preset duration and moved in 10ms frames for frame division. The short-time energy intensity of each frame is calculated. At the same time, the set limited intervals YYmax and YYmin are retrieved. If the short-time energy intensity is higher than YYmax, the corresponding frame is determined to be strong voice. If the short-time energy intensity is lower than YYmin, the corresponding frame is determined to be non-voice. If the short-time energy intensity is between YYmax and YYmin, the corresponding frame is determined to be weak voice. If three consecutive frames of signal are strong speech, they are marked as the start of a speech segment. After obtaining the start of a speech segment, if the frame is not speech and the duration is greater than or equal to the preset duration, it is marked as the end of a speech segment. The signal between the start and end points is extracted as a candidate speech segment.

4. The early education terminal interactive system based on voice acquisition and recognition according to claim 1, characterized in that, The analysis process of the control authorization module is as follows: Based on the pre-set feature extraction algorithm, the user's candidate speech segments are extracted for voiceprint features. The extracted voiceprint features are matched and analyzed with the authorized speech feature library to obtain the result of verification pass or verification fail.

5. The early education terminal interactive system based on voice acquisition and recognition according to claim 1, characterized in that, The analysis process of the speech recognition module is as follows: Based on the extracted voiceprint features and the voiceprint-permission association library, the permission level of the current controlled and authorized object is obtained. If the permission level is temporary interactive level L2, the operation time of the current user is obtained, and the operation time is matched with the valid time period to obtain the verification validity or play the prompt sound "Authorization validity expired, please adjust". The process of building the voiceprint-permission association library: After parents complete the voiceprint registration of the control authorization object, a unique voiceprint feature identifier is assigned to each voiceprint and the corresponding permission level is bound.

6. The early education terminal interactive system based on voice acquisition and recognition according to claim 5, characterized in that, When the verification is valid or at the basic interaction level or control management level, the candidate speech segment is input into the pre-set instruction recognition model and the text instruction is output. The text command is matched against a pre-set command-permission table. The lowest permission level corresponding to the input text command is then matched against the current permission level to obtain either a pass signal or a verification failure result.

7. The early education terminal interactive system based on voice acquisition and recognition according to claim 1, characterized in that, The analysis process for the historical experience module is as follows: The system acquires daily interaction recognition data from early childhood education interactive devices. This data includes text command recognition accuracy (the proportion of text commands that correctly match the playback content) and fuzzy speech recognition rate for young children (the accuracy of recognizing text commands that are authorized to children). The sum of the text instruction recognition accuracy and the fuzzy speech recognition rate of young children multiplied by their corresponding weighting coefficients is set as the interactive recognition index; At the same time, the daily interactive performance data of early education interactive devices are obtained. The interactive performance data includes the total response latency of instructions and the voiceprint verification pass rate. The interactive capability index is calculated based on (1-total response latency of instructions) × the corresponding weight coefficient + voiceprint verification pass rate × the corresponding weight coefficient. The interactive recognition index and the interactive capability index are discriminated to obtain a performance stability signal or a risk signal. Based on the daily interaction recognition index and interaction ability index, a daily-result management table is constructed. Based on the daily-result management table, the percentage of risk signals corresponding to the early education interactive device for 7 consecutive days (if there are days without use, the next day will be planned to the previous day until the 7 consecutive days are met) is obtained. The percentage values ​​are processed to obtain either regular signals or maintenance signals.