Multi-line educational interaction method and toy based on AI (artificial intelligence)

By using an AI-based multi-line interactive puzzle method and a bidirectional long short-term memory neural network acoustic model and a speech activity detection model, the problems of delayed response time and crosstalk in multi-person interaction in smart toys are solved, and fast and coherent voice interaction feedback is achieved.

CN121011189APending Publication Date: 2025-11-25深圳宇凡微电子有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511283401.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing smart toys suffer from problems such as delayed response times, crosstalk during multi-user concurrency, and intermittent interaction due to pauses and misjudgments, especially in real-world toy environments where children's pronunciation is unstable and their speech is highly colloquial.

Method used

We employ an AI-based multi-line interactive puzzle method, utilizing a bidirectional long short-term memory neural network acoustic model for speech feature extraction and processing. We combine this with a speech activity detection model to detect brief silence pauses, and improve recognition accuracy and response speed through Viterbi search decoding.

Benefits of technology

Significantly shortens the perceptual waiting time for toy interaction, ensures semantic continuity and feedback consistency in multi-line interaction, and is suitable for multi-line educational interaction scenarios such as home, classroom, and exhibition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121011189A_ABST
    Figure CN121011189A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice recognition, in particular to an AI-based multi-line intelligence development interaction method and toy, and the method comprises the steps: building a voice interaction recognition session, and carrying out the initialization after receiving a start instruction; continuously receiving audio clips and writing the audio clips into a voice feature buffer area; frame-level features are extracted from the segments to be buffered, when the accumulated length exceeds the length of a first numerical value, the sequence with the length is extracted as a batch, and a subsequent second numerical value is obtained to serve as a right-side context; segmenting the batch into sub-blocks according to a third numerical value, splicing left and right contexts for each block to form a CSC sequence, combining into a CSC feature matrix, and sending the CSC feature matrix into an acoustic model to calculate acoustic scores in parallel; detecting transient pause, triggering backtracking, stage output and state reset by using a voice activity detection model based on the score; and searching by combining viterbi with a language model to obtain a final recognition text. According to the invention, the problems of long response time delay in voice interaction of the intelligent toy, easy serial connection caused by multi-person concurrence and incoherent interaction caused by pause misjudgment can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, in particular to a multi-line intelligence interaction method based on AI and a toy. BACKGROUND

[0002] In recent years, smart toys (such as voice story machines, interactive programming robots, voice-controlled building / puzzle toys, etc.) for children and parent-child scenarios have introduced a large amount of voice interaction for guessing, answering, reading by sound, and scenario branch selection, etc. "Intelligence interaction" play. However, the existing solutions mostly rely on whole sentence recognition or single round "one question and one answer". When the children speak fast, interject with laughter, or have a short pause, the system either waits for the end of the whole section before outputting, or misjudges the pause and submits too early, resulting in long response time and fragmented feedback, which easily causes the children's attention to be lost. When multiple people interact around the same toy or multiple toys in parallel, it is also easy to cause mutual interference and context crosstalk among multiple conversations, affecting the entertainment and enlightenment effect. In order to reduce the time delay, some solutions use lightweight or one-way acoustic models, frame skipping decoding, etc., but often at the expense of reduced recognition accuracy, especially in real toy environments where children's pronunciation is unstable and oral language is strong. SUMMARY

[0003] In view of the above technical problems, the present application provides a multi-line intelligence interaction method based on AI and a toy, aiming to solve the problem of long response time in voice interaction of smart toys, and incoherent interaction caused by multi-line interference and pause misjudgment in multi-person concurrent interaction.

[0004] Other characteristics and advantages of the present disclosure will become apparent from the following detailed description, or will be learned by practice of the present disclosure.

[0005] According to an aspect of the present application, a multi-line intelligence interaction method based on AI is provided, the method comprising: establishing a voice interaction recognition session, initializing a recognition process after receiving a voice start instruction, continuously receiving audio segments transmitted by the user client, and writing them into a voice feature buffer; extracting voice features from each received audio segment, converting the audio segment into a corresponding voice feature vector and storing it in the voice feature buffer, when the voice feature vectors accumulated in the voice feature buffer exceed a preset batch first value length, extracting a feature sequence composed of the voice feature vectors with the first value length as the current decoding batch, and simultaneously obtaining a feature sequence with a second value length immediately following it as the right context of the corresponding batch; divide the current feature sequence to be decoded into a plurality of audio sub-blocks according to a preset third numerical length, and attach a left context feature with a length of the second numerical value and a right context feature with a length of the second numerical value to each of the audio sub-blocks in a context-sensitive partitioning manner to generate a corresponding CSC partitioned feature sequence; combine all the CSC partitioned feature sequences to form a CSC feature matrix, and input the CSC feature matrix into a bidirectional long short-term memory neural network acoustic model to perform parallel calculation to obtain an acoustic score sequence corresponding to the CSC partitioned feature sequence; based on the acoustic score sequence, use a deep neural network-based voice activity detection model to detect whether there is a short pause in the acoustic score sequence, if a short pause is detected, backtrack the acoustic score sequence before the time when the pause occurs to obtain an optimal recognition text result before the pause, send the recognition text corresponding to the stage to the user client, and reset the state of subsequent decoding to clear the history information of the completed part; perform a Viterbi search decoding to calculate possible recognition paths according to the acoustic score sequence and in combination with a preset language model to obtain a corresponding voice interaction recognition text result.

[0006] Further, when performing voice feature extraction on the audio segment, the method comprises: extract 40-dimensional log-mel filter bank voice features for each 10-millisecond-long audio segment, and stack the voice features of the first 7 frames and the last 7 frames of the audio segment with the current frame features to obtain a 600-dimensional voice feature vector for each frame.

[0007] Further, the voice activity detection model determines a silent frame by calculating a non-silence probability and a silence probability output by the acoustic model, and specifically comprises: obtain the maximum value of all acoustic state probabilities output by the bidirectional long short-term memory neural network acoustic model for the current frame, which belongs to a non-silence phoneme state, as a non-silence probability, and the maximum value of all acoustic state probabilities output by the bidirectional long short-term memory neural network acoustic model for the current frame, which belongs to a silence phoneme state, as a silence probability, and calculate a log likelihood ratio; when the log likelihood ratio is less than a preset threshold, the corresponding frame is determined as a silent frame.

[0008] Further, the voice activity detection model performs smoothing processing on the short pause determination results of consecutive frames, comprising: calculate the proportion of silent frames in a preset window length, and if the proportion is greater than a preset threshold, the frame corresponding to the center of the window is determined as a silent frame. when the silent determination results of adjacent frames change from silence to non-silence after smoothing processing, it is recognized that a short pause has occurred, and a minimum time interval is set for consecutive twice short pause detection to avoid false triggering output.

[0009] Further, the calculation of the acoustic score sequence is performed synchronously for multiple concurrent audio streams, including: The feature sequences of multiple different audio streams are merged into a batch of the CSC feature matrix, which is input into the bidirectional long short-term memory neural network acoustic model for parallel forward calculation, thereby realizing multi-channel parallel acoustic score calculation to improve the calculation throughput and reduce the data transmission overhead.

[0010] Further, the bidirectional long short-term memory neural network acoustic model decodes the input CSC feature matrix based on a context-sensitive blocking strategy, and each CSC blocking feature sequence is treated as an independent audio sequence for processing.

[0011] Further, the extracted speech feature vectors are stored in a ring-shaped speech feature buffer, and when the length is extracted based on a predetermined value, after the feature sequence of the first numerical length is extracted from the speech feature buffer, the speech feature with the length of the second numerical value is reserved as the context in the speech feature buffer to be used as the left context for the next batch decoding.

[0012] According to a second aspect of the present disclosure, an AI-based multi-line interactive educational toy is provided, the toy comprising: a processor; a communication module; and a memory arranged to store computer executable instructions which, when executed, cause the processor to send a voice start instruction and an audio segment to an external server through the communication module, so that the external server executes the AI-based multi-line interactive method as described above.

[0013] The technical solution of the present application has the following beneficial effects: Compared with the prior art, the speech activity detection model directly reuses the frame-level posterior of the bidirectional long short-term memory neural network acoustic model, calculates the non-silence / silence log-likelihood ratio and performs sliding window smoothing, detects a short silence pause, triggers "backtracking-submission-state reset", and quickly feeds back the correct answer / prompt to the child under the condition of not interrupting the subsequent decoding, significantly shortens the perceived waiting time of the toy interaction, and makes the guessing, answering and reading scoring more coherent and instant reinforcement. Different conversations (different children, different toys or multiple microphone streams of the same toy) can be combined into a batch after being labeled with a conversation identifier, and then parallel forward calculation is performed on the CSC feature matrix and the conversation identifier is split and returned. Combined with the ring buffer and left and right context retention, the natural connection between batches and the context are ensured, which not only reduces repeated calculation and improves throughput, but also ensures the semantic continuity of each interactive and feedback consistency, and is suitable for multi-line interactive scenarios such as family, classroom and exhibition. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 A flow chart of an AI-based multi-line interactive method for an embodiment of the present specification; Figure 2 A structural block diagram of an AI-based multi-line interactive toy for an embodiment of the present specification. DETAILED DESCRIPTION

[0015] Example implementations will now be described more fully with reference to the accompanying drawings. Example implementations can be implemented in any

[0016] In addition, the accompanying drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification. The drawings illustrate exemplary embodiments of the present disclosure and, as such, a change in the drawings can be represented by a change in a reference numeral. In the drawings:

[0017] The present disclosure provides an AI-based multi-line interactive method for a product. Referring to FIG. 1, a flowchart of an AI-based multi-line interactive method provided by an embodiment of the present disclosure is shown. The method can be applied to electronic devices such as personal computers, servers, controllers, etc. The method can be performed by a device, which can be implemented by software and / or hardware. The method can specifically include the following steps S101-S106: Figure 1 In step S101, a voice interactive recognition session is established, and the recognition process is initialized after receiving a voice start instruction. Audio segments transmitted by the user client are continuously received and written into a voice feature buffer.

[0018] ​Among them, first, the voice interaction recognition session is established: when the voice start instruction from the user client arrives, the recognition process is initialized and enters the streaming reception; During the session, continuously receive the audio segments uploaded by the client, and parse them according to the message type. Whenever an audio segment is obtained, immediately enter the feature calculation, and write the obtained frame-level speech features to the speech feature buffer, which is used to queue and stream subsequent processing links, until the session end message is received. The above process ensures that the session establishment, data continuous access and real-time caching mechanism of the feature stream are connected with each other.

[0019] In the speech feature extraction of the audio segment, 40-dimensional log-mel filter bank speech features are extracted for the audio segment of 10 milliseconds, and the speech features of the first 7 frames and the last 7 frames of the audio segment are spliced and stacked with the current frame feature to obtain a speech feature vector of 600 dimensions per frame.

[0020] Specifically, when the audio segment is subjected to speech feature extraction, 40-dimensional log-mel filter banks are extracted from the original audio as basic acoustic features with a frame shift of 10 ms, and the current frame and the features of the past 7 frames and the future 7 frames are spliced and stacked in the feature dimension, thereby forming a stacked feature vector of 600 dimensions per frame; The obtained frame-level features are written to the ring buffer in the order of arrival for subsequent batch decoding and other links to consume. The log-mel filter bank here refers to the logarithmic representation of the filter bank weighting of the power spectrum in the mel frequency scale, which is used to enhance the frequency spectrum components related to speech perception; Frame refers to a short-time analysis unit in a fixed time step (10 ms in this embodiment); Stacking refers to splicing multiple frame features in the time neighborhood to explicitly provide the temporal context, so that the subsequent model can obtain more rich local dynamic information without increasing the real-time delay.

[0021] In step S102, speech feature extraction is performed on each received audio segment, the audio segment is converted into a corresponding speech feature vector and stored in a speech feature buffer, and when the speech feature vectors accumulated in the speech feature buffer exceed a preset batch first value length, a feature sequence composed of the speech feature vectors with the first value length is extracted as the current decoding batch, and a feature sequence with a second value length immediately following it is obtained as the right context of the corresponding batch.

[0022] Among them, the frame-level feature processing is performed on the continuously received audio segment, and the obtained frame vector is sequentially written to the speech feature buffer. When the frame-level speech features in the buffer accumulate to or exceed a preset batch first value length , a segment with a length of characteristic sequence as the current to-be-decoded batch; meanwhile, to ensure that the subsequent model obtains necessary look-ahead information without increasing the overall delay, a characteristic sequence with a length of a second value is taken as the right context of the batch. In the embodiment, the typical values are = 200 frames (about 2s) and = 40 frames; the implementation pops length from the buffer as the content of the batch, while frames immediately following are reserved for subsequent calculation, so as to simulate bidirectional timing dependence at the batch level and reduce the decoding delay.

[0023] To facilitate quantification of the feature size of the batch, the total length of the feature vector can be written as ; wherein is the frame shift, and 10 ms is adopted in the embodiment; is the feature dimension per frame, and 40-dimensional log-mel filter banks are adopted and stacked after 7 past frames and 7 future frames to obtain 600 dimensions. Thus, under the settings of = 200 frames, = 10 ms, = 600, the feature scalar number corresponding to a single to-be-decoded batch is 12000. The speech feature buffer refers to a data structure for sequentially buffering frame-level feature vectors, supporting writing as soon as arriving and batch reading on demand; the first value length of the batch processing is the threshold triggering batch construction, representing the main content length to be sent into the subsequent processing pipeline; the second value is the set of look-ahead frames provided after the main content of the batch, used to improve the robustness of boundary and dynamic judgment in the subsequent link.

[0024] In step S103, the current to-be-decoded characteristic sequence is divided into a plurality of audio sub-blocks according to a preset third value length, and context-sensitive partitioning is adopted to attach a left context feature with a length of the second value and a right context feature with a length of the second value to each audio sub-block, to generate a corresponding CSC partitioned characteristic sequence.

[0025] The current to-be-decoded characteristic sequence is sequentially cut according to a preset third value length, to obtain a group of audio sub-blocks with a fixed length of the third value ; then, a context-sensitive partitioning (CSC) operation is applied to each sub-block: the sub-block is taken as the center, and a left context and a right context with a length of the second value the history and future features of the sub-chunk, thereby extending the original sub-chunk into a chunked feature sequence containing double-sided context. The time window length of a CSC chunked feature sequence thus generated is denoted as , which is jointly constituted by the sub-chunk and the two-sided context, satisfying the formula: The number of CSC chunked sequences available within a batch to be decoded is determined by the number of sub-chunk divisions, satisfying the formula: The above two formulas directly give the key quantitative relationship of this step: the coverage range of each CSC sequence on the time axis, and the number of CSC sequences that can be processed in parallel within the same batch.

[0026] Context-sensitive chunking refers to concatenating a fixed-length central sub-chunk with its left and right context of the same length as a sequence that can be processed independently, to explicitly provide the pre- and post-context information to the model without waiting for the entire audio; the left context and the right context are respectively the frame-level feature sets before and after the time position of the sub-chunk, both of which have a length given by the second value ; the third value is the reference length of the sub-chunk, used to determine the division granularity and parallelism. After using the above process, each sub-chunk in the current batch is extended into a CSC chunked feature sequence with a length of , forming a set of CSC chunked feature sequences for subsequent steps.

[0027] In step S104, all the CSC chunked feature sequences are combined to form a CSC feature matrix, which is input into a bidirectional long short-term memory neural network acoustic model to calculate in parallel, obtaining an acoustic score sequence corresponding to the CSC chunked feature sequence.

[0028] The bidirectional long short-term memory neural network acoustic model decodes and processes the input CSC feature matrix based on the context-sensitive chunking strategy, treating each CSC chunked feature sequence as an independent audio sequence for processing.

[0029] Specifically, the CSC chunked feature sequences obtained in step S103 are organized by batch as a parallelizable feature input: each CSC chunked feature sequence is flattened according to time and kept consistent with the frame stacking representation in the feature dimension, and combined to form a CSC feature matrix. To facilitate quantification, let the sub-chunk length be the third value , and the left and right context lengths be the second value , then the time window length of each CSC chunk is ; let the frame shift be ​​​, the single-frame feature dimension is , the corresponding flattened frame number x feature dimension product is , the number of CSC sequence in the same batch is . Therefore, the size of the CSC feature matrix can be represented as x , which is used for batch processing of acoustic models for forward calculation.

[0030] The CSC feature matrix is input into a bidirectional long short-term memory network (BLSTM) acoustic model, and each column (or each sequence) is separately forward inferred in a parallel manner to obtain a frame-level acoustic score sequence corresponding to each CSC block. Here, bidirectional long short-term memory network refers to a recurrent neural network structure that utilizes past and future information simultaneously in the time dimension; acoustic score refers to the posterior probability or its logarithmic form of each frame belonging to each acoustic state (such as triphone state or subword unit), which is used for subsequent decoding and determination. Based on the context-sensitive blocking strategy, the model treats each CSC block feature sequence as an independent decodable sequence for processing: that is, under the context condition of its left , a stable posterior estimate is generated for the center frame, and all CSC sequences in the batch are calculated in parallel in a batch processing manner, thereby obtaining bidirectional timing information and generating acoustic score output for use without relying on the entire speech.

[0031] In step S105, based on the acoustic score sequence, a deep neural network-based voice activity detection model is used to detect whether there is a short silent pause, and if a short silent pause is detected, the acoustic score sequence before the silent pause occurs is decoded back to obtain the optimal recognition text result before the pause, the corresponding recognition text is sent to the user client, and the state of subsequent decoding is reset to clear the history information of the completed part.

[0032] The voice activity detection model determines the silent frame by calculating the non-silence probability and silence probability output by the acoustic model, specifically including: obtaining the maximum value of all acoustic state probabilities output by the bidirectional long short-term memory neural network acoustic model for the current frame belonging to the non-silence phoneme state as the non-silence probability, and the maximum value belonging to the silence phoneme state as the silence probability, and calculating the log likelihood ratio; when the log likelihood ratio is less than a preset threshold, the corresponding frame is determined as a silent frame.

[0033] And the voice activity detection model performs smoothing processing on the short pause judgment result of the continuous frames, including: calculating the proportion of the preset window length frames in silence, if the proportion is greater than a preset threshold, the frame corresponding to the window center is determined as silence; when the silence judgment result of the adjacent frames changes from silence to non-silence after the smoothing processing, it is identified that a short pause has occurred, and a minimum time interval is set for the continuous two short pause detections to avoid false triggering.

[0034] Specifically, the frame-level acoustic score sequence obtained in the direct reuse step S104 is directly used as the judgment stop basis, and is input into a voice activity detection model (DNN-VAD) based on a deep neural network to detect whether there is a short pause. The voice activity detection model here refers to comparing and evaluating the posterior of each frame on the "non-silence phoneme state set" and the "silence phoneme state set" by using the frame-level output of the deep acoustic model: let the log posterior of the i-th output node be , then the log probability of non-silence and silence at the current time t is defined as: ; ; According to this, the log likelihood ratio is calculated . When is lower than the threshold , the frame is recorded as a silence candidate. In order to suppress jitter, further smoothing is performed on the silence judgment in a sliding window with a length of (2W+1), and the silence ratio R(t)= is calculated. If R(t) , the center frame of the window is determined as silence. Let the smoothed judgment be , when the transition of and occurs, it is considered that a short pause is detected at t; in order to avoid too frequent triggering, a minimum time interval is set between the adjacent two times of triggering, and only when the interval between the current triggering point and the last triggering point is greater than or equal to , the triggering is effective.

[0035] ​Once a short pause is detected, a decoding backtracking is performed on the acoustic score sequence before the pause time: without changing the processing of the following segment, a backward backtracking is performed on the search trajectory of 0…t with the pause time t as the backtracking boundary, the path with the highest probability is selected and the corresponding optimal recognition text is output. The phase text is immediately sent to the user client to support continuous interactive feedback; at the same time, the state of the decoder is reset, that is, the history marks, stacks or survival hypotheses of the confirmed part before t are cleaned up, only the necessary context for continuing decoding after t is retained, so that the search of the subsequent frame continues from a clean and consistent starting point. The above process ensures that detection—submission—reset naturally connects within the same batch: the speech activity detection model provides stable pause positioning based on frame-level posteriori, backtracking guarantees the global consistency of the phase results, and state reset avoids interference of the completed paragraph on subsequent search, thereby realizing low-latency phase return without sacrificing recognition accuracy.

[0036] In step S106, a Viterbi search decoding is performed to calculate possible recognition paths according to the acoustic score sequence and in combination with a preset language model, and a corresponding voice interactive recognition text result is obtained.

[0037] Among them, the frame-level acoustic score sequence obtained in step S104 is input to perform Viterbi search decoding: a group of candidate paths in an active state is maintained on the time axis, and in each frame, the newly arrived acoustic score and the constraint of the preset language model are jointly applied to these candidates to update their path probabilities and perform pruning according to the beam width, so that only the candidates with higher probabilities continue to expand. The language model can adopt an n-gram form to provide prior constraints for legal word (or sub-word / phoneme) sequences when the path is expanded, so as to jointly determine the overall priority of the path with the acoustic score; the Viterbi search estimates the state probability of all possible paths accordingly until the search of the current decoding segment is completed. When the forward search of the segment is completed, the decoder backtracks the accumulated candidates, finds the path with the highest accumulated probability from the end to the front in the state sequence, and maps the optimal path to the corresponding voice interactive recognition text result output. Here, the “Viterbi search decoding” can be understood as a two-stage process: estimating the state probability based on the acoustic score and the language model on all feasible paths; when the current segment search is completed, the state sequence with the highest probability is backtracked to obtain the optimal path; the cooperation of the above two stages enables stable text results to be given while maintaining the consistency of the language model context.

[0038] In an embodiment, the calculation of the acoustic score sequence is performed synchronously for multiple concurrent audio streams, including: The feature sequences of multiple different audio streams are combined into a batch of the CSC feature matrix, which is input into the bidirectional long short-term memory neural network acoustic model for parallel forward calculation, thereby realizing multi-channel parallel acoustic score calculation to improve the calculation throughput and reduce the data transmission overhead.

[0039] In the embodiment, the continuously generated frame-level speech feature vectors are sequentially written into a ring-shaped speech feature buffer. The ring-shaped buffer can be understood as a circular queue with the head connected to the tail, and a write pointer and a read pointer are maintained: the write pointer is sequentially advanced frame by frame to receive new features; when the cumulative number of frames in the buffer reaches a predetermined trigger condition, the read pointer starts to read in batches. To construct the current batch to be decoded, a continuous feature sequence with a first numerical length (i.e., the batch processing length) is extracted from the position of the read pointer as the main content of the batch; after the extraction is completed, the read pointer is only moved forward by the same first numerical length, indicating that this part of the content has been consumed and "clearly separated" from the subsequent processing of the session.

[0040] This batch combination and parallel method brings two effects: one is to change the original calculation of each sequence into batch parallel, which significantly improves the throughput and utilization of the calculation unit of the model; the other is to converge the upload / download between the host and the accelerator from each session to once per batch, thereby reducing the number of round trips and scheduling overhead. After the forward calculation is completed, the acoustic scores corresponding to each sequence are split and sent back to the respective sessions according to the session identifier, and are respectively input into the pause detection and Viterbi search processes of the session, which ensures the independence of multi-path interaction and the context not to be out of line, and maintains the time continuity and boundary stability within the batch.

[0041] In an embodiment, the extracted speech feature vectors are stored in a ring-shaped speech feature buffer, and when the length of the extracted feature sequence reaches a predetermined numerical value, the speech feature with a length of the second numerical value is reserved as the context in the speech feature buffer to be used as the left context for the next batch decoding.

[0042] In the embodiment, the continuously generated frame-level speech feature vectors are sequentially written into a ring-shaped speech feature buffer. The ring-shaped buffer can be understood as a circular queue with the head connected to the tail, and a write pointer and a read pointer are maintained: the write pointer is sequentially advanced frame by frame to receive new features; when the cumulative number of frames in the buffer reaches a predetermined trigger condition, the read pointer starts to read in batches. To construct the current batch to be decoded, a continuous feature sequence with a first numerical length (i.e., the batch processing length) is extracted from the position of the read pointer as the main content of the batch; after the extraction is completed, the read pointer is only moved forward by the same first numerical length, indicating that this part of the content has been consumed and "clearly separated" from the subsequent processing of the session.

[0043] To ensure the time continuity and context consistency between batches, after the above extraction is completed, the feature frame corresponding to the second number of values that immediately follows is not moved, but is continuously retained in the ring buffer, and this segment of un-consumed frame is taken as the left context of the next batch. In this way, when the next batch construction is triggered, the left neighborhood of the position pointed to by the read pointer naturally contains this segment of retained context frame, which can be directly used as the "history information" of the next batch to participate in subsequent processing without repeated calculation or additional copying. This approach ensures the natural connection of adjacent batches on the time axis, avoids the judgment jitter caused by the information gap at the boundary, and maintains the stable operation of the buffer through the read-write strategy of "moving the main content forward and retaining the context". When the write pointer is rewound due to the ring structure, the relative relationship between the read and write pointers remains unchanged, thereby realizing continuous and low-overhead supply of the left context of the next batch without additional steps outside the scope of operation.

[0044] Compared with the prior art, the voice activity detection model of the above embodiment directly reuses the frame-level posterior of the bidirectional long short-term memory neural network acoustic model, calculates the non-silence / silence log likelihood ratio and performs sliding window smoothing, detects a short silence pause, triggers "backtracking-submission-state reset", and rapidly feeds back the correct answer / prompt to the child under the condition of not interrupting subsequent decoding, thereby significantly shortening the perceived waiting time of the toy interaction, making the guessing, answering and reading scoring more coherent and instant reinforcement. Different conversations (different children, different toys or multiple microphone streams of the same toy) can be combined into a batch after being labeled with a conversation identifier, and then parallel forward and return according to the conversation identifier on the CSC feature matrix. In combination with the ring buffer and left and right context retention, the natural connection between batches and the context consistency are ensured, which reduces repeated calculation, improves throughput, ensures the semantic continuity and feedback consistency of each interaction, and adapts to multi-line intelligence interaction scenes such as family, classroom and exhibition.

[0045] Based on the same idea, as shown in Figure 2 , an AI-based multi-line intelligence interactive toy provided by an embodiment of the present application has a structure diagram. The toy comprises: a processor 201; a communication module 202; and a memory 203 arranged to store computer executable instructions, which, when executed, cause the processor to send a voice start instruction and an audio segment to an external server 300 through the communication module, so that the external server executes the AI-based multi-line intelligence interactive method as described above.

[0046] The specific details in the above toy have been described in detail in the method part of the embodiment, and the undisclosed details can be referred to the embodiment content of the method part, and thus will not be described again.

[0047] The above-described diagrams are merely schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not intended for limiting purposes. It is readily understood that the processes shown in the above-described diagrams do not indicate or limit the time sequence of these processes. In addition, it is also readily understood that these processes can be executed, for example, synchronously or asynchronously in a plurality of modules.

[0048] It should be noted that, although several modules are mentioned in the above detailed description, such a division is not mandatory. Indeed, according to the exemplary embodiments of the present disclosure, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into embodied by a plurality of modules or units.

[0049] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practice of the disclosed application. The present application is intended to cover any variations, uses, or adaptations of the present disclosure following, in general, the principles of the present disclosure and including such departures from the present disclosure that come within known or customary practice in the art to which the present disclosure pertains. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

[0050] It is to be understood that the present disclosure is not limited to the precise construction described and shown in the above description and the accompanying drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the claims appended hereto.

Claims

1. An AI-based multi-line interactive puzzle method, characterized in that, The method comprises: establishing a voice interaction recognition session, initializing a recognition process after receiving a voice start instruction, continuously receiving audio segments transmitted by the user client, and writing them into a voice feature buffer; extracting voice features from each received audio segment, converting the audio segment into a corresponding voice feature vector and storing it in the voice feature buffer, when the voice feature vectors accumulated in the voice feature buffer exceed a preset batch first numerical length, extracting a feature sequence composed of the voice feature vectors with a length of the first numerical value as a current decoding batch, and simultaneously obtaining a feature sequence with a length of a second numerical value following it as the right context of the corresponding batch; dividing the current decoding feature sequence into multiple audio sub-blocks according to a preset third numerical length, and attaching a left context feature with a length of the second numerical value and a right context feature with a length of the second numerical value to each audio sub-block in a context-sensitive blocking manner to generate a corresponding CSC blocking feature sequence; combining all the CSC blocking feature sequences to form a CSC feature matrix, and inputting it into a bidirectional long short-term memory neural network acoustic model for parallel calculation to obtain an acoustic score sequence corresponding to the CSC blocking feature sequence; based on the acoustic score sequence, using a voice activity detection model based on a deep neural network to detect whether there is a short pause, if a short pause is detected, decoding backtracking the acoustic score sequence before the time when the pause occurs to obtain the optimal recognition text result before the pause, sending the recognition text of the corresponding stage to the user client, and resetting the state of subsequent decoding to clear the history information of the completed part; performing a Viterbi search decoding to calculate possible recognition paths according to the acoustic score sequence and in combination with a preset language model to obtain a corresponding voice interaction recognition text result. 2.The AI-based multi-line interactive puzzle method according to claim 1, wherein, When extracting voice features from the audio segment, it includes: extracting 40-dimensional log-mel filter bank voice features from the audio segment with a length of 10 milliseconds, and stacking the voice features of the first 7 frames and the last 7 frames of the audio segment with the current frame feature to obtain a voice feature vector with a length of 600 dimensions per frame. 3.The AI-based multi-line interactive puzzle method according to claim 1, wherein, The voice activity detection model determines a silent frame by calculating the non-silence probability and the silence probability output by the acoustic model, specifically including: obtaining the maximum value of all acoustic state probabilities output by the bidirectional long short-term memory neural network acoustic model for the current frame as the non-silence probability, and the maximum value of the silence phoneme state as the silence probability, and calculating the log likelihood ratio; when the log likelihood ratio is less than a preset threshold, the corresponding frame is determined as a silent frame. 4.The AI-based multi-line interactive puzzle method according to claim 1, wherein, The voice activity detection model performs smoothing processing on the short pause determination results of consecutive frames, including: calculating the proportion of silent frames in a preset window length, if the proportion is greater than a preset threshold, the frame corresponding to the center of the window is determined as a silent frame. When the silence determination result of the adjacent frame after smoothing processing changes from silence to non-silence, it is identified that a short silence pause occurs, and a minimum time interval is set for two consecutive short pause detections to avoid false triggering output. 5.The AI-based multi-line interactive puzzle method according to claim 1, wherein, The calculation of the acoustic score sequence is performed synchronously for multiple concurrent audio streams, including: The feature sequences of multiple different audio streams are merged into a batch of CSC feature matrices, which are input into the bidirectional long short-term memory neural network acoustic model for parallel forward calculation, thereby realizing multi-channel parallel acoustic score calculation to improve calculation throughput and reduce data transmission overhead. 6.The AI-based multi-line interactive puzzle method according to claim 1, wherein, The bidirectional long short-term memory neural network acoustic model decodes the input CSC feature matrix based on a context-sensitive blocking strategy, and processes each CSC blocking feature sequence as an independent audio sequence. 7.The AI-based multi-line interactive puzzle method according to claim 1, wherein, The extracted speech feature vectors are stored in a ring-shaped speech feature buffer, and when the length of the extracted feature sequence is based on a predetermined value, after extracting the first value length of the feature sequence from the speech feature buffer, the speech feature with a length of the second value is reserved as the context in the speech feature buffer to be used as the left context for the next batch decoding.

8. An AI-based multi-line interactive educational toy, the toy comprising: a processor; a communication module; and a memory arranged to store computer executable instructions which, when executed, cause the processor to send a voice start instruction and an audio segment to an external server through the communication module, so that the external server performs the AI-based multi-line interactive method according to any one of claims 1-7.