Edge fast response method and system based on adaptive short window sampling
Through adaptive short-window sampling technology and dynamically adjusted noise models, smart toys can respond quickly and effectively deal with noise on low-computing-power platforms, improving the voice interaction experience, solving the problems of slow response and false triggering, and achieving instant feedback and accurate recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU ZHONGDA CHENG TECHNOLOGY DEVELOPMENT CO LTD
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-02
AI Technical Summary
Existing smart toys have significant limitations in response speed, adaptability to low computing power platforms, and handling noisy environments, resulting in sluggish response, false triggering, and misjudgment, which affects the user experience.
Adaptive short-window sampling technology is adopted. By establishing a noise model and setting monitoring parameters during the startup phase, the sampling window is dynamically adjusted. It combines multiple audio features to identify speech events and utilizes a fast feedback link and a background main recognition model for real-time interaction and deep semantic understanding. It learns environmental noise characteristics in real time to adapt to different acoustic scenarios.
It achieves rapid response on low-computing-power devices and effectively copes with noisy environments, enhancing the voice interaction experience of smart toys, ensuring instant feedback and accurate voice recognition, and adapting to complex usage scenarios for children.
Smart Images

Figure CN122135706A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to speech processing methods, and more specifically to an edge-based fast response method and system based on adaptive short-window sampling. Background Technology
[0002] Currently, although many smart toys on the market are trying to enhance the user experience through voice interaction, existing technologies still have significant limitations. These limitations are mainly reflected in three aspects: response speed, ability to adapt to low computing power platforms, and effectiveness in dealing with noisy environments.
[0003] First, regarding response speed, many traditional toys rely on fixed-length frames and a buffering strategy for speech recognition. This means that the recognition process is only triggered when the buffer is full or a prolonged speech duration is detected. This mechanism is not ideal for child users, as children speak quickly and pause frequently, leading to frequent situations where the toy hasn't responded even after the speech has finished. Furthermore, sudden short sounds such as laughter or exclamations are difficult to capture in time, further exacerbating the toy's sluggish response. Second, achieving high-performance speech recognition on resource-constrained embedded platforms is also a significant challenge. Many traditional start-point detection algorithms and speech recognition front-end designs were originally developed for high-computing devices like PCs or mobile phones. Directly applying them to toy-grade chips results in excessive computational load and memory consumption, making them unsuitable for long-term operation. Additionally, high-frequency operation of complex feature extraction and model inference rapidly depletes battery power, shortening usage time. Simultaneously, to maintain system stability, response speed and detection accuracy must be sacrificed, which brings us back to the first problem. Finally, the impact of noisy environments on speech recognition cannot be ignored. In typical scenarios where children use toys, there is often a broad spectrum of noise, such as television sounds, music, and the operation of household appliances, as well as momentary loud noise interference from children shouting and throwing toys. Traditional energy threshold-based starting point detection methods perform poorly in such environments, easily resulting in false triggers or slow responses. Fixed threshold schemes cannot adaptively adjust to environmental changes, potentially leading to oversensitivity in quiet environments and misjudgment in noisy environments.
[0004] In conclusion, these problems are not isolated but rather interconnected and work together. For example, limitations in hardware resources not only directly affect response speed but also limit the ability to handle noise; and slow response speed and poor noise immunity directly impact user experience.
[0005] Therefore, it is necessary to design a new method that can respond quickly, effectively cope with noisy environments, and adapt to low-computing-power edge devices to improve the voice interaction experience of smart toys. Summary of the Invention
[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide an edge fast response method and system based on adaptive short window sampling.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: an edge fast response method based on adaptive short window sampling, comprising:
[0008] Acquire and analyze audio data during the startup phase, establish a noise model and set initial monitoring parameters to obtain a noise baseline;
[0009] Based on the noise benchmark, the ambient sound changes of the input signal are continuously monitored, and the sampling window is dynamically adjusted in combination with the decision threshold to adapt to different acoustic scenarios.
[0010] The input signal is comprehensively evaluated using multiple audio features to identify valid speech events.
[0011] Based on the aforementioned voice events, a dual response of real-time interaction and deep semantic understanding is achieved using a fast feedback loop and a backend main recognition model.
[0012] Its further technical solution is as follows: after the dual response of real-time interaction and deep semantic understanding based on the voice event using a fast feedback link and a background main recognition model, it also includes:
[0013] During operation, it continuously learns the characteristics of environmental noise and automatically adjusts the noise model and decision threshold.
[0014] The further technical solution is as follows: acquiring and analyzing audio data during the startup phase, establishing a noise model, and setting initial monitoring parameters to obtain a noise baseline includes:
[0015] Acquire and statistically analyze audio data during the startup phase, and calculate the energy, spectral distribution characteristics, and long-term signal-to-noise ratio of environmental noise to obtain a noise model;
[0016] The sampling window length and decision threshold are set according to the noise model to obtain the noise benchmark.
[0017] The further technical solution is as follows: based on the noise benchmark, continuously monitoring the environmental sound changes of the input signal, and dynamically adjusting the sampling window to adapt to different acoustic scenarios in conjunction with the decision threshold, including:
[0018] During normal operation, the microphone input signal is continuously monitored and compared with the noise benchmark to adjust the monitoring strategy in real time.
[0019] The further technical solution is as follows: During normal operation, continuously monitoring the input signal of the microphone and comparing it with the noise benchmark to adjust the monitoring strategy in real time includes:
[0020] The energy difference and spectral characteristic difference between the input signal and the noise reference are compared to obtain the characteristic difference results;
[0021] Adjust the sampling window length and sliding step size based on the aforementioned feature difference results;
[0022] When a sound transition that meets the requirements appears in the feature difference results, the analysis time window is shortened to speed up the confirmation of the speech start point.
[0023] The further technical solution is as follows: the feature difference result includes the instantaneous signal-to-noise ratio, using SNR=10log 10 The instantaneous signal-to-noise ratio is calculated using (Ps / Pn), where log10 represents the logarithm to the base 10, Ps represents the energy estimate of the useful speech component within the current analysis window of the input signal, and Pn represents the energy estimate of the current environmental noise floor.
[0024] Its further technical solution is as follows: the method of comprehensively evaluating the input signal using multiple audio features to identify valid voice events includes:
[0025] The instantaneous signal-to-noise ratio, energy mutation index, and spectral flatness characteristics are comprehensively evaluated to determine valid speech events.
[0026] Its further technical solution is as follows: Based on the voice event, the dual response of real-time interaction and deep semantic understanding is achieved using a fast feedback link and a background main recognition model, including:
[0027] Upon detecting the voice event, an immediate feedback is provided through a lightweight and fast response mechanism;
[0028] The background main recognition model is activated to perform deep analysis and semantic understanding on the complete speech segment corresponding to the speech event in order to obtain the analysis results;
[0029] Complex responses are driven by the analysis results.
[0030] Its further technical solution is: the continuous learning of environmental noise characteristics during operation, and the automatic adjustment of the noise model and decision threshold, including:
[0031] During operation, environmental noise statistics are updated using data from speechless segments, and the decision threshold and window parameters are automatically adjusted based on the environmental noise level.
[0032] The present invention also provides an edge fast response system based on adaptive short window sampling, comprising:
[0033] The baseline establishment unit is used to acquire and analyze audio data during the startup phase, establish a noise model, and set initial monitoring parameters to obtain a noise baseline.
[0034] The adjustment unit is used to continuously monitor the changes in ambient sound of the input signal based on the noise benchmark, and dynamically adjust the sampling window to adapt to different acoustic scenarios in combination with the decision threshold.
[0035] The recognition unit is used to comprehensively evaluate the input signal using multiple audio features and identify valid speech events.
[0036] The response unit is used to provide a dual response based on the voice event, employing a fast feedback link and a background master recognition model for real-time interaction and deep semantic understanding.
[0037] The advantages of this invention compared to existing technologies are as follows: This invention establishes a noise benchmark by acquiring and analyzing audio data during the startup phase to build a noise model and set initial monitoring parameters. Subsequently, based on this benchmark, it continuously monitors changes in ambient sound and dynamically adjusts the sampling window according to decision thresholds to adapt to different acoustic scenarios. It comprehensively evaluates the input signal using multiple audio features to accurately identify valid voice events. Once a voice event is detected, the system immediately provides an instant interactive response through a fast feedback link, while the main recognition model in the background performs deep semantic understanding, ensuring rapid response and effective handling of noisy environments. Furthermore, this method is designed with low-computing-power edge devices in mind, employing lightweight algorithms and strategies to reduce computational resource consumption, enabling smart toys to provide a smooth and natural voice interaction experience even under resource constraints.
[0038] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. Attached Figure Description
[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a flowchart illustrating the edge fast response method based on adaptive short window sampling provided in an embodiment of the present invention.
[0041] Figure 2 A schematic block diagram of an edge fast response system based on adaptive short window sampling provided in an embodiment of the present invention;
[0042] Figure 3 A schematic block diagram of a computer device provided for an embodiment of the present invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0045] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0046] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0047] Please see Figure 1 , Figure 1 This is a flowchart illustrating the edge-based fast response method based on adaptive short-window sampling provided in an embodiment of the present invention. This method is applied in a server. Through adaptive short-window sampling technology, the method analyzes environmental noise to establish a noise model during the startup phase and sets initial monitoring parameters to adapt to different acoustic scenarios. During operation, it continuously monitors changes in the ambient sound of the input signal and dynamically adjusts the sampling window. It combines multiple audio features to evaluate and identify valid voice events, and utilizes a fast feedback link and a backend main recognition model to achieve a dual response of real-time interaction and deep semantic understanding. Simultaneously, it continuously learns and updates environmental noise characteristics to automatically adjust relevant parameters. Thus, while ensuring low computational power requirements, it achieves effective handling and rapid response to noisy environments, significantly improving the voice interaction experience of smart toys.
[0048] Figure 1 This is a flowchart illustrating the edge-fast response method based on adaptive short-window sampling provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps S110 to S150.
[0049] S110. Acquire and analyze the audio data during the startup phase, establish a noise model and set initial monitoring parameters to obtain a noise baseline.
[0050] In this embodiment, the noise benchmark refers to a set of parameters determined based on the collected ambient background sound during the initial startup of the device. These parameters include, but are not limited to, the energy level of the ambient noise, its spectral distribution characteristics, and the long-term signal-to-noise ratio (SNR). This benchmark is used in subsequent processing to compare with the real-time input sound signal to determine whether it contains valid speech content.
[0051] In one embodiment, step S110 described above may include steps S111 to S112.
[0052] S111. Acquire and statistically analyze the audio data during the startup phase, and calculate the energy, spectral distribution characteristics, and long-term signal-to-noise ratio of the environmental noise to obtain a noise model.
[0053] In this embodiment, the noise model refers to a set of parameters describing the current environmental noise characteristics, calculated using statistical methods based on the environmental audio data collected during the startup phase. Specifically, this includes, but is not limited to:
[0054] Energy level: measures the overall loudness of ambient noise.
[0055] Spectral distribution characteristics: Describes the distribution of environmental noise at different frequencies, which helps to distinguish different types of sound sources.
[0056] Long-term signal-to-noise ratio: This measures the ratio of background noise to the potential speech signal in the environment, which is crucial for determining when to begin speech recognition.
[0057] This information together constitutes our understanding of current environmental noise, or what is known as a noise model.
[0058] S112. Set the sampling window length and decision threshold according to the noise model to obtain the noise benchmark.
[0059] In this step, the initial monitoring parameters of the system are set based on the noise model constructed in the previous step, mainly including:
[0060] Sampling window length: Determines the length of the audio segment analyzed each time. Adjusting this length according to the characteristics of ambient noise can optimize detection sensitivity and accuracy.
[0061] Decision threshold: A standard used to determine whether an input signal contains valid speech. For example, in low-noise environments, a more lenient threshold can be used to capture more possible speech signals; while in high-noise environments, a stricter threshold is needed to avoid false triggering.
[0062] In this way, the system can adaptively adjust its monitoring strategy for different acoustic scenarios, thereby improving its performance in practical applications, which is especially important in application scenarios such as children's smart toys, where there are high requirements for instant response and anti-interference capabilities.
[0063] S120. Based on the noise benchmark, continuously monitor the changes in ambient sound of the input signal, and dynamically adjust the sampling window to adapt to different acoustic scenarios in conjunction with the decision threshold.
[0064] In this embodiment, during normal operation, the input signal of the microphone is continuously monitored and compared with the noise benchmark in order to adjust the monitoring strategy in real time.
[0065] During normal operation, the system continuously monitors the microphone input signal and compares it to a previously established noise baseline. This process aims to adjust the monitoring strategy in real time to ensure effective recognition of speech signals and triggering of appropriate responses in various acoustic environments.
[0066] In one embodiment, step S120 described above may include steps S121 to S123.
[0067] S121. Compare the energy difference and spectral characteristic difference between the input signal and the noise reference to obtain the characteristic difference result.
[0068] In this embodiment, the feature difference result refers to a set of indicators determined by analyzing the energy difference and spectral distribution difference between the current input audio signal and the previously established noise model. Specifically, these indicators include, but are not limited to:
[0069] Instantaneous signal-to-noise ratio (SNR): Used to measure the proportion of useful speech components in the current input signal relative to background noise.
[0070] Energy difference: Reflects whether the energy level of the current input signal is significantly higher or lower than the noise reference.
[0071] Spectral characteristic difference: describes the difference between the spectral distribution of the current input signal and the noise reference, which helps to distinguish speech signals from other types of noise.
[0072] The feature difference results include instantaneous signal-to-noise ratio (SNR), using SNR=10log 10 The instantaneous signal-to-noise ratio is calculated using (Ps / Pn), where log10 represents the logarithm to the base 10, Ps represents the energy estimate of the useful speech component within the current analysis window of the input signal, and Pn represents the energy estimate of the current environmental noise floor.
[0073] S122. Adjust the sampling window length and sliding step size based on the feature difference results.
[0074] Based on the feature difference results obtained in the previous step, the system dynamically adjusts the sampling window length and sliding step size to capture the speech signal more accurately. Specific adjustment strategies include:
[0075] When the environment is stable: If the feature difference results indicate that the current environment is relatively stable and there is no obvious speech activity, then keep the current sampling window length and sliding step size unchanged, and continue to listen to the environment with a relatively relaxed standard.
[0076] When environmental fluctuations occur: When significant energy transitions or spectral structure changes are detected, the system will correspondingly shorten the sampling window length and reduce the sliding step size, thereby improving the sensitivity to the speech start point and the positioning accuracy.
[0077] S123. When a sound transition that meets the requirements appears in the feature difference results, shorten the analysis time window and speed up the confirmation of the speech start point.
[0078] This step is the core of the entire dynamic adjustment mechanism, specifically designed for rapid confirmation of the speech start point. The specific implementation details are as follows:
[0079] Voice transition detection: The system monitors the energy and spectral changes in the feature difference results. Once a significant transition phenomenon is detected (such as a sudden change from background noise to a voice with speech features), it immediately enters the "voice candidate state".
[0080] Window shrinking mechanism: After detecting a possible speech start point, the system quickly shortens the current analysis time window from the long default setting to a shorter time period, thereby reducing waiting time and speeding up the confirmation process of the speech start point.
[0081] Early judgment: The shortened analysis time window allows the system to determine the starting point of speech in a very short time, rather than waiting for the entire long window to end. This method greatly improves response speed, allowing children to receive feedback from toys immediately after uttering the first syllable, enhancing the realism and immediacy of the interaction.
[0082] Through the above steps, the method of this embodiment can achieve adaptive adjustment to different acoustic environments, which not only improves the accuracy and stability of speech recognition, but also significantly improves the user experience, especially in the application scenario of children's smart voice toys that require fast response.
[0083] S130: Utilizes multiple audio features to comprehensively evaluate the input signal and identify valid speech events.
[0084] In this embodiment, a speech event refers to the actual speech signal determined by analyzing and synthesizing multiple audio features. To achieve this, the system employs a variety of features for comprehensive evaluation to ensure accurate differentiation of the actual speech signal from background noise and other non-speech events.
[0085] The instantaneous signal-to-noise ratio, energy mutation index, and spectral flatness characteristics are comprehensively evaluated to determine valid speech events.
[0086] To identify valid voice events, the system used the following key audio features and evaluated them comprehensively:
[0087] Instantaneous Signal-to-Noise Ratio (SNR): The instantaneous signal-to-noise ratio refers to the ratio of the energy of useful speech components to the energy of ambient noise within the current analysis window. SNR is one of the important indicators for determining whether the current frame contains effective speech. A higher SNR value usually indicates the presence of significant speech components.
[0088] The energy mutation index reflects how the energy level of the current frame changes relative to the previous frame or the baseline noise level. Specifically, it is obtained by calculating the energy difference between adjacent frames. A significant increase in energy is often an important indicator of the start of speech. However, a single energy mutation is not sufficient to confirm a speech event, as other high-energy events (such as a door slam) can also cause energy mutations.
[0089] Spectral flatness is a measure of the uniformity of a frequency spectrum. It is commonly used to distinguish between white noise and speech signals. White noise has high spectral flatness, while speech signals typically have low spectral flatness due to their specific frequency component distribution. Spectral flatness can help a system further differentiate speech from other types of noise. For example, if a signal has high spectral flatness, it is likely not a speech signal.
[0090] To accurately identify valid voice events, the system performs joint analysis and decision-making based on the above features:
[0091] On edge devices, the system estimates the instantaneous signal-to-noise ratio with extremely low computational complexity. This is achieved by simplifying the algorithm, ensuring fast computation even in resource-constrained environments.
[0092] The system does not rely on a single feature to make decisions, but rather combines features such as instantaneous signal-to-noise ratio, energy mutation index, and spectral flatness for comprehensive evaluation. The specific steps are as follows:
[0093] Initial screening: First, the system performs an initial screening of the current frame based on the instantaneous signal-to-noise ratio and energy mutation index. If both of these indicators show potential speech activity, the process proceeds to the next step.
[0094] Spectral Feature Verification: Next, the system calculates the spectral flatness of the current frame and compares it with a preset threshold. If the spectral flatness is below a certain threshold, and the instantaneous signal-to-noise ratio and energy change index both meet the requirements, then the frame is considered to potentially contain valid speech.
[0095] The system will only trigger a speech event when all features collectively indicate that the current frame does indeed contain valid speech. Specific conditions include:
[0096] The sound becomes louder (high energy mutation index);
[0097] The spectral structure changes significantly (low spectral flatness, non-white noise).
[0098] Speech components are dominant (high instantaneous signal-to-noise ratio);
[0099] To ensure real-time responsiveness while maintaining high recognition accuracy, the system employs a tiered processing mechanism. The initial screening phase uses relatively lenient standards to reduce missed triggers, while the final trigger confirmation phase uses a more stringent multi-feature joint decision to avoid false triggers. The system continuously optimizes the threshold settings for various features based on data accumulated during operation to adapt to different environments and application scenarios. For example, in noisy environments, the system may appropriately increase the thresholds for certain features to reduce the false trigger rate.
[0100] By employing this multi-feature comprehensive evaluation method, the method in this embodiment can effectively identify children's speech signals in complex acoustic environments and provide timely feedback, thereby greatly enhancing the interactive experience and practicality of toys.
[0101] S140. Based on the voice event, a dual response of real-time interaction and deep semantic understanding is achieved using a fast feedback link and a background main recognition model.
[0102] In one embodiment, step S140 described above may include steps S141 to S143.
[0103] S141. Upon detecting the voice event, provide immediate feedback through a lightweight and fast response mechanism.
[0104] Once the system confirms the presence of a voice event—that is, that a child has begun to speak—it immediately activates a lightweight, fast-response chain to provide instant feedback. The key to this step is ensuring both rapid feedback and low computational complexity, so that it can operate efficiently even on resource-constrained edge devices.
[0105] Immediate feedback can be simple, such as flashing lights, a toy making a "hmm?" sound, or a motorized facial expression responding. This type of immediate feedback is designed to give children the feeling that "the toy is listening to me."
[0106] Implementation: Since complete semantic understanding is not required at this stage, the system relies solely on the results of front-end processing (such as energy mutations, spectral feature changes, etc.) to trigger immediate feedback. This link is designed to minimize latency, typically completing the process from detecting the start of speech to outputting immediate feedback within milliseconds, thereby greatly shortening the response time of children's subjective perception.
[0107] S142. Start the background main recognition model to perform deep analysis and semantic understanding on the complete speech segment corresponding to the speech event to obtain the analysis result.
[0108] At the same time, the system will also launch a more complex main recognition model in the background to perform deep analysis and semantic understanding of the entire speech segment. This step aims to obtain more accurate information so that a more accurate and comprehensive response can be made subsequently.
[0109] Unlike the immediate feedback provided earlier, complete speech segment analysis here involves the entire speech segment, not just the beginning. By understanding the whole sentence or paragraph, the system can better grasp the child's intentions. Semantic parsing is performed using deep learning algorithms or other advanced speech recognition technologies, which typically require more computing resources and are therefore executed in the background.
[0110] S143. Drive complex responses based on the analysis results.
[0111] Ultimately, based on the analysis results provided by the main recognition model in the background, the system will drive corresponding complex responses. These responses are not limited to simple sounds or actions, but may also include richer content such as storytelling, answering questions, and playing music.
[0112] Compared to immediate feedback, this type of response places greater emphasis on the accuracy and relevance of the content. For example, if a child asks a question, the system not only needs to understand the question itself but also provide an appropriate answer. In addition to verbal responses, it can also incorporate various forms of interaction, such as visual (e.g., displaying animations) and tactile (e.g., moving toys), to enhance the overall experience.
[0113] By simultaneously activating the lightweight, fast-response link and the main backend recognition model, the system can ensure instantaneous feedback speed without sacrificing the accuracy and richness of subsequent responses. Considering the resource limitations of edge devices, the system cleverly allocates tasks—lightweight tasks are handled directly by the frontend, while complex tasks are handled by the backend. This improves efficiency and avoids unnecessary resource waste. Through the above strategies, the method in this embodiment significantly enhances the interactive experience between users (especially children) and smart voice toys, enabling toys to not only respond quickly but also provide meaningful interactive content, thus strengthening their companionship attributes.
[0114] S150: During operation, it continuously learns the characteristics of environmental noise and automatically adjusts the noise model and decision threshold.
[0115] In this embodiment, during operation, the environmental noise statistics are updated using data from speechless segments, and the decision threshold and window parameters are automatically adjusted according to the environmental noise level.
[0116] In this embodiment, in order to ensure that the intelligent voice toy maintains stable and reliable performance in various complex environments, the system is designed with an adaptive mechanism. This mechanism can continuously learn and update the characteristics of environmental noise during operation, and automatically adjust the noise model and decision threshold accordingly.
[0117] Learning during the silent period:
[0118] Data Acquisition: During system operation, especially during periods when no voice activity is detected (the so-called "silent period"), the system continuously records a certain amount of audio data. This data is primarily used to analyze the background noise level of the current environment.
[0119] Statistical calculation: Using this data without speech segments, the system can calculate the latest environmental noise energy, spectral distribution, and other statistical parameters. This step is performed in real time, ensuring that the noise model can reflect environmental changes in a timely manner.
[0120] Noise floor tracking and scene change detection:
[0121] Moving average processing: The energy and spectral distribution of the newly collected silent segments are incorporated into the noise floor statistics and smoothed by moving average to reduce the impact of short-term fluctuations on long-term trends.
[0122] Scene change judgment: When the system detects a significant change in the background noise energy or spectral characteristics (e.g., a continuous increase in background noise energy or a shift to low-frequency mechanical noise as the main component), it is considered that the current environment has changed. In this case, the noise model and decision parameters need to be re-evaluated and adjusted.
[0123] Adaptive to high-noise scenes:
[0124] Increase the stringency of triggering conditions: If the ambient noise level increases (such as when a TV is turned on or a vacuum cleaner is started), the system will automatically raise the decision threshold and shorten the sampling window parameters. This tightens the triggering conditions and avoids frequent false triggering in noisy environments.
[0125] Low-noise scene restoration:
[0126] Lowering the stringency of trigger conditions: Conversely, when the environment becomes quiet again, the system gradually lowers the decision threshold, relaxes the trigger conditions, and extends the sampling window length. This is done to improve sensitivity and prevent missing softer children's voices.
[0127] Parameter update and loop monitoring:
[0128] Parameter Update: The parameters adjusted above will be used as the new baseline to guide subsequent real-time monitoring. This means the system always makes decisions based on the latest environmental information.
[0129] Cyclic monitoring: This process is a continuous cycle in which the system constantly collects data from the environment, analyzes noise characteristics, and adjusts parameters, thereby achieving a high degree of sensitivity and adaptability to environmental changes.
[0130] By implementing step S150, the method of this embodiment effectively solves two major problems existing in the traditional fixed window start point detection algorithm:
[0131] High false trigger rate: In noisy environments, traditional methods are prone to generating too many false triggers due to fixed thresholds. This invention, however, employs an adaptive method that flexibly adjusts the threshold based on the actual ambient noise level, significantly reducing the probability of false triggers.
[0132] High missed trigger rate: In quiet environments, especially when children's voices are faint, the fixed window method may miss some important speech signals. This invention improves the ability to capture faint sounds while maintaining accuracy by dynamically extending the window length.
[0133] Furthermore, this adaptive mechanism boasts excellent scalability and flexibility, enabling children's smart voice toys of different brands and price points to achieve the best user experience based on their own hardware capabilities and usage scenarios. It also provides a solid foundation for subsequent functional expansion; for example, the integration of modules such as emotion recognition and motion capture can achieve better performance with the help of this framework.
[0134] This embodiment relates to the fields of speech signal processing and embedded edge computing, with a particular focus on short-window sampling and real-time streaming speech processing methods, speech start-point detection and endpoint decision algorithms, and edge inference model design deployed on resource-constrained chips. Its primary application scenario is children's smart voice toys, designed to support low-power MCUs and dedicated voice SoCs, and adapt to the needs of high-noise environments such as homes and kindergartens. A key objective of this method is to achieve an instant interactive experience of "speak a sentence, get an immediate response," significantly shortening the overall interaction time from sound generation to toy response by transforming the traditional "complete recording followed by recognition" mode into a streaming response mode that involves simultaneous acquisition and decision-making.
[0135] To enhance user experience, the method in this embodiment aims to provide feedback within a very short time after the first syllable of a child's speech, thereby strengthening the child's feeling of "being responded to" and "having companionship." Furthermore, the system maintains stable triggering even in noisy, fragmented, and irregular speaking conditions, and is compatible with various situations such as rapid shouting, interruptions, and overlapping speech by multiple people, while avoiding false triggers caused by background noise or distant voices. During long-term operation, the system also maintains a low false trigger rate and a low missed trigger rate. For mass-produced toy products, this invention provides an engineered implementation path, ensuring that the algorithm can be deployed in memory conditions ranging from tens to hundreds of KB, and can be adapted to different types of microphones and shell acoustic characteristics through parameter configuration. This not only facilitates subsequent algorithm iterations through firmware upgrades but also allows toys of different price points and brands to adopt this technology.
[0136] The core technical problems addressed by the method in this embodiment include significantly improving voice triggering and response speed, compressing the initial feedback delay to a range perceived as "almost immediate" by children while ensuring recognition accuracy; and supporting immediate feedback such as facial expressions, lighting, and actions based on the judgment results of the previous few frames before the child has finished speaking. Simultaneously, this embodiment proposes a short-window sampling mechanism that adapts to noise fluctuations, dynamically adjusting the sampling window length and judgment threshold based on the real-time environmental signal-to-noise ratio and instantaneous energy changes to reduce false triggers and promptly identify the true speech segment in noisy environments. Finally, this embodiment forms an overall algorithm architecture suitable for low-computing-power edge devices. The computational load of the front-end feature extraction process is controllable, avoiding large-scale matrix operations as much as possible, and implementing multi-level judgments within the limits of chip resources, reducing the number of calls to the main recognition model. This allows the technology to be flexibly deployed on toy hardware platforms of different price points and brands.
[0137] Specifically, the first step is system initialization and environmental baseline construction (startup phase). Raw audio stream data is acquired within the first few seconds after the device powers on. This includes: an environmental analysis module performing statistical analysis on the acquired audio data; establishing a noise model to calculate the energy, spectral distribution characteristics, and long-term signal-to-noise ratio of the initial environmental noise as a "quiet baseline"; and automatically setting the initial sampling window length and decision threshold based on the above results to prepare for subsequent monitoring.
[0138] The second step involves real-time monitoring and adaptive window length adjustment (runtime phase). This step requires acquiring real-time microphone input speech frames and established baseline noise frames. Specific implementation details include: calculating the energy and spectral feature differences between the current speech frame and the baseline noise frame; dynamically adjusting the window length and sliding step based on a simplified instantaneous signal-to-noise ratio (SNR) metric; and triggering "window contraction" and "early decision" mechanisms when a significant energy transition is detected and the spectral structure shifts from noise to speech, enabling the system to confirm the speech initiation point within a very short time after the child begins to speak.
[0139] The third step is the joint triggering decision based on multi-dimensional features (computation and decision-making stage), which requires data such as instantaneous signal-to-noise ratio estimates, energy mutation indices, and spectral flatness. Specific operations include: low-power SNR calculation using the formula SNR=10log 10 (Ps / Pn) Estimate the instantaneous signal-to-noise ratio; compare the SNR with simplified features such as energy mutation index and spectral flatness; the system only determines to enter the response process when the comprehensive conditions meet the criteria for a valid speech event.
[0140] The fourth step is dual-link parallel response processing (interactive feedback stage), which requires the "voice start point" signal determined in step three and the subsequent complete voice audio stream. Specific measures include: activating the pre-processor fast response link to output simple feedback within milliseconds to create a sense of immediacy; simultaneously starting the main recognition model to perform deep decoding and semantic understanding of the complete voice segment; and driving more complex language responses or action responses after accurate semantic results are obtained, ensuring both speed and accuracy.
[0141] The final step is environmental self-calibration and dynamic threshold updating (maintenance and loop phase), which requires acquiring data including non-speech segment data during operation and long-term environmental noise statistics. Specifically, this involves: automatically updating environmental noise statistics using non-speech segment data; automatically adjusting the decision threshold and window parameters based on changes in environmental noise; and using the updated parameters as a new benchmark to return to the second step for continued real-time monitoring, forming a closed-loop management system. These steps work together to ensure the system's efficiency and stability.
[0142] In this embodiment, when the child utters the first syllable, the light on the toy flashes or displays facial expressions, reinforcing the feeling that "the toy is listening to me attentively." This immediate feedback is not limited to the results of speech recognition but is broken down into two levels: "immediate feedback" and "complete response," significantly enhancing the toy's companionship attributes.
[0143] Adaptive noise modeling and threshold update mechanisms enable toys to operate for extended periods in complex environments such as television sounds, background music, and multi-person conversations without frequent false triggers. Furthermore, for sudden high-energy non-speech events (such as door slamming or smashing sounds), the system uses joint spectral features to identify them, avoiding misinterpretation as children speaking.
[0144] The system operates under low load most of the time, only activating the relatively power-intensive speech recognition and semantic understanding modules upon confirmation of a voice event. The core algorithm consists of basic operations such as addition, subtraction, multiplication, division, and table lookup, making it compatible with low-end MCUs without floating-point units, and its memory usage can be controlled to the tens of KB level suitable for toy products.
[0145] The sampling and rapid response framework formed in this embodiment can be seamlessly integrated with other modules such as emotion recognition, motion capture, and touch sensing. Different manufacturers can add their own unique features on top of a unified interface, such as dynamically adjusting the toy's facial expressions and response methods based on speech speed and tone changes.
[0146] The method system in this embodiment can identify the speech initiation point and provide immediate feedback within a very short time after the child begins to speak, enhancing the interactive experience. Adaptive noise modeling and threshold updating ensure accurate differentiation between speech and non-speech events even in complex noise environments, avoiding false triggering. The design takes into account hardware limitations and adopts a low-power strategy, reducing the pressure on edge device computing and energy consumption. The unified interface design makes it easy to integrate more advanced functions, improving the user experience.
[0147] The hardware in this embodiment includes at least a microphone array or a single microphone, an audio front-end amplifier and filter, a battery and power management unit, a low-power processor, or a dedicated voice SoC. It is equipped with output channels such as a speaker, LED strip, and facial expression drive motor for real-time feedback and complete responses. A simple accelerometer can be optionally configured to distinguish between motion and voice interaction.
[0148] The data acquisition and adaptive short-window scheduling module continuously reads data from the sound card or ADC and switches the window length and step size according to the current strategy. The environment analysis and threshold estimation module maintains long-term noise statistics, estimates the instantaneous signal-to-noise ratio, and updates the trigger parameters. The start-point detection and fast response module immediately drives feedback after determining the start of speech, and simultaneously marks the start position of the recognition segment. The speech recognition and semantic understanding module is responsible for decoding the complete segment and driving subsequent action scripts.
[0149] After powering on, the system enters a silent sampling phase, collecting several samples from environments without speech to estimate baseline noise energy, spectral distribution, and initial signal-to-noise ratio (SNR). Based on the estimation results, a set of initial short-window parameters is selected to prevent frequent accidental touches before understanding the environment. New audio data is continuously read, and energy changes, spectral structure changes, and instantaneous SNR estimates are calculated for each window. The system determines whether a speech event is likely at the current moment; when the indicator value continuously exceeds a threshold, a "speech candidate state" is triggered. The quick response module immediately triggers a short feedback script the moment the starting point is confirmed. The complete speech segment is sent to the recognition and semantic understanding module, which analyzes the type of interaction the child wishes to engage in and triggers the corresponding advanced response process. The system automatically collects several silent segments and incorporates their energy and spectral distribution into the background noise statistics. When the statistical results differ significantly from the initial environment, the appropriate window length and threshold are recalculated. The system records the child's average speaking volume, common speech rate, and pause patterns, fine-tuning the starting point detection sensitivity without compromising safety.
[0150] Tested in various typical environments, including a simulated living room, kindergarten classroom, and outdoor playground, the method presented in this embodiment significantly reduces the first-response time and improves the response speed satisfaction rate under the same false trigger rate. In high-noise scenarios, the method maintains stable start-point detection through adaptive threshold and window adjustment. Battery life tests show that the method in this embodiment significantly reduces overall power consumption compared to continuous high-load recognition strategies.
[0151] The aforementioned edge-based fast response method based on adaptive short-window sampling establishes a noise baseline by acquiring and analyzing audio data during the startup phase to build a noise model and set initial monitoring parameters. Subsequently, based on this baseline, it continuously monitors changes in ambient sound and dynamically adjusts the sampling window according to a decision threshold to adapt to different acoustic scenarios. It comprehensively evaluates the input signal using multiple audio features to accurately identify valid voice events. Once a voice event is detected, the system immediately provides an instant interactive response through a fast feedback loop, while the main recognition model in the background performs deep semantic understanding, ensuring rapid response and effective handling of noisy environments. Furthermore, this method is designed with low-computing-power edge devices in mind, employing lightweight algorithms and strategies to reduce computational resource consumption, enabling smart toys to provide a smooth and natural voice interaction experience even under resource constraints.
[0152] Figure 2 This is a schematic block diagram of an edge fast response system 300 based on adaptive short window sampling provided in an embodiment of the present invention. Figure 2As shown, corresponding to the above-described edge fast response method based on adaptive short window sampling, the present invention also provides an edge fast response system 300 based on adaptive short window sampling. This edge fast response system 300 includes a unit for executing the above-described edge fast response method based on adaptive short window sampling, and the system can be configured in a server. Specifically, please refer to... Figure 2 The edge fast response system 300 based on adaptive short window sampling includes a reference establishment unit 301, an adjustment unit 302, an identification unit 303, a response unit 304, and a learning unit 305.
[0153] The baseline establishment unit 301 is used to acquire and analyze audio data during the startup phase, establish a noise model, and set initial monitoring parameters to obtain a noise baseline. The adjustment unit 302 is used to continuously monitor changes in ambient sound of the input signal based on the noise baseline, and dynamically adjust the sampling window to adapt to different acoustic scenarios in conjunction with a decision threshold. The recognition unit 303 is used to comprehensively evaluate the input signal using multiple audio features and identify valid speech events. The response unit 304 is used to provide a dual response based on the speech events, employing a fast feedback link and a background main recognition model for real-time interaction and deep semantic understanding. The learning unit 305 is used to continuously learn the characteristics of ambient noise during operation and automatically adjust the noise model and decision threshold.
[0154] In one embodiment, the reference establishment unit 301 includes: a noise model generation subunit, used to acquire and statistically analyze audio data during the startup phase, and calculate the energy, spectral distribution characteristics and long-term signal-to-noise ratio level of the ambient noise to obtain a noise model; and a setting subunit, used to set the sampling window length and decision threshold according to the noise model to obtain a noise reference.
[0155] In one embodiment, the adjustment unit 302 is configured to continuously monitor the input signal of the microphone during normal operation and compare it with the noise reference in order to adjust the monitoring strategy in real time.
[0156] In one embodiment, the adjustment unit 302 includes:
[0157] The comparison subunit is used to compare the energy difference and spectral characteristic difference between the input signal and the noise reference to obtain the feature difference result; the length adjustment subunit is used to adjust the sampling window length and sliding step size based on the feature difference result; the shortening subunit is used to shorten the analysis time window and speed up the speech start point confirmation speed when a sound transition that meets the requirements appears in the feature difference result.
[0158] In one embodiment, the recognition unit 303 is used to comprehensively evaluate the instantaneous signal-to-noise ratio, energy mutation index, and spectral flatness features to determine valid speech events.
[0159] In one embodiment, the response unit 304 includes:
[0160] The instant feedback subunit is used to provide instant feedback through a lightweight and fast response mechanism after the voice event is detected; the parsing subunit is used to start the background main recognition model to perform deep parsing and semantic understanding of the complete voice segment corresponding to the voice event in order to obtain the parsing result; and the driving subunit is used to drive a complex response based on the parsing result.
[0161] In one embodiment, the learning unit 305 is used to update the environmental noise statistics using data from speechless segments during operation, and to automatically adjust the decision threshold and window parameters according to the environmental noise level.
[0162] It should be noted that those skilled in the art can clearly understand that the specific implementation process of the above-mentioned edge fast response system 300 based on adaptive short window sampling and each unit can be referred to the corresponding description in the foregoing method embodiments. For the sake of convenience and brevity, it will not be repeated here.
[0163] The aforementioned edge fast response system 300 based on adaptive short window sampling can be implemented as a computer program, which can, for example... Figure 3 It runs on the computer device shown.
[0164] Please see Figure 3 , Figure 3 This is a schematic block diagram of a computer device provided in an embodiment of this application. The computer device 500 can be a server, wherein the server can be a standalone server or a server cluster composed of multiple servers.
[0165] See Figure 3 The computer device 500 includes a processor 502, a memory, and a network interface 505 connected via a system bus 501. The memory may include a non-volatile storage medium 503 and internal memory 504.
[0166] The non-volatile storage medium 503 may store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions that, when executed, cause the processor 502 to perform an edge fast response method based on adaptive short-window sampling.
[0167] The processor 502 provides computing and control capabilities to support the operation of the entire computer device 500.
[0168] The internal memory 504 provides an environment for the execution of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute an edge fast response method based on adaptive short window sampling.
[0169] This network interface 505 is used for network communication with other devices. Those skilled in the art will understand that... Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device 500 to which the present application is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0170] The processor 502 is used to run a computer program 5032 stored in the memory to implement all the steps of the edge fast response method based on adaptive short window sampling.
[0171] It should be understood that in the embodiments of this application, the processor 502 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0172] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program includes program instructions and can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0173] Therefore, the present invention also provides a storage medium. This storage medium may be a computer-readable storage medium. The storage medium stores a computer program, wherein when executed by a processor, the computer program causes the processor to perform all steps of the edge-fast response method based on adaptive short-window sampling.
[0174] The storage medium can be any computer-readable storage medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.
[0175] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0176] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of each unit is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0177] The steps in the method of this invention can be adjusted, merged, or reduced in order according to actual needs. The units in the system of this invention can be merged, divided, or reduced according to actual needs. Furthermore, the functional units in the various embodiments of this invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0178] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0179] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. An edge-fast response method based on adaptive short-window sampling, characterized in that, include: Acquire and analyze audio data during the startup phase, establish a noise model and set initial monitoring parameters to obtain a noise baseline; Based on the noise benchmark, the ambient sound changes of the input signal are continuously monitored, and the sampling window is dynamically adjusted in combination with the decision threshold to adapt to different acoustic scenarios. The input signal is comprehensively evaluated using multiple audio features to identify valid speech events. Based on the aforementioned voice events, a dual response of real-time interaction and deep semantic understanding is achieved using a fast feedback loop and a backend main recognition model.
2. The edge fast response method based on adaptive short window sampling according to claim 1, characterized in that, After the dual response of real-time interaction and deep semantic understanding based on the voice event using a fast feedback loop and a background main recognition model, it also includes: During operation, it continuously learns the characteristics of environmental noise and automatically adjusts the noise model and decision threshold.
3. The edge fast response method based on adaptive short window sampling according to claim 1, characterized in that, The process of acquiring and analyzing audio data during the startup phase, establishing a noise model, and setting initial monitoring parameters to obtain a noise baseline includes: Acquire and statistically analyze audio data during the startup phase, and calculate the energy, spectral distribution characteristics, and long-term signal-to-noise ratio of environmental noise to obtain a noise model; The sampling window length and decision threshold are set according to the noise model to obtain the noise benchmark.
4. The edge fast response method based on adaptive short window sampling according to claim 1, characterized in that, The step of continuously monitoring environmental sound changes of the input signal based on the noise benchmark, and dynamically adjusting the sampling window to adapt to different acoustic scenarios in conjunction with a decision threshold, includes: During normal operation, the microphone input signal is continuously monitored and compared with the noise benchmark to adjust the monitoring strategy in real time.
5. The edge fast response method based on adaptive short window sampling according to claim 4, characterized in that, During normal operation, continuously monitoring the microphone's input signal and comparing it with the noise benchmark to adjust the monitoring strategy in real time includes: The energy difference and spectral characteristic difference between the input signal and the noise reference are compared to obtain the characteristic difference results; Adjust the sampling window length and sliding step size based on the aforementioned feature difference results; When a sound transition that meets the requirements appears in the feature difference results, the analysis time window is shortened to speed up the confirmation of the speech start point.
6. The edge fast response method based on adaptive short window sampling according to claim 5, characterized in that, The feature difference results include instantaneous signal-to-noise ratio (SNR), using SNR=10log 10 The instantaneous signal-to-noise ratio is calculated using (Ps / Pn), where log10 represents the logarithm to the base 10, Ps represents the energy estimate of the useful speech component within the current analysis window of the input signal, and Pn represents the energy estimate of the current environmental noise floor.
7. The edge fast response method based on adaptive short window sampling according to claim 6, characterized in that, The method of comprehensively evaluating the input signal using multiple audio features to identify valid speech events includes: The instantaneous signal-to-noise ratio, energy mutation index, and spectral flatness characteristics are comprehensively evaluated to determine valid speech events.
8. The edge fast response method based on adaptive short window sampling according to claim 1, characterized in that, The dual response based on the voice event, employing a fast feedback loop and a backend main recognition model for real-time interaction and deep semantic understanding, includes: Upon detecting the voice event, an immediate feedback is provided through a lightweight and fast response mechanism; The background main recognition model is activated to perform deep analysis and semantic understanding on the complete speech segment corresponding to the speech event in order to obtain the analysis results; Complex responses are driven by the analysis results.
9. The edge fast response method based on adaptive short window sampling according to claim 2, characterized in that, The process of continuously learning environmental noise characteristics and automatically adjusting the noise model and decision threshold during operation includes: During operation, environmental noise statistics are updated using data from speechless segments, and the decision threshold and window parameters are automatically adjusted based on the environmental noise level.
10. An edge-based fast response system based on adaptive short-window sampling, characterized in that, include: The baseline establishment unit is used to acquire and analyze audio data during the startup phase, establish a noise model, and set initial monitoring parameters to obtain a noise baseline. The adjustment unit is used to continuously monitor the changes in ambient sound of the input signal based on the noise benchmark, and dynamically adjust the sampling window to adapt to different acoustic scenarios in combination with the decision threshold. The recognition unit is used to comprehensively evaluate the input signal using multiple audio features and identify valid speech events. The response unit is used to provide a dual response based on the voice event, employing a fast feedback link and a background master recognition model for real-time interaction and deep semantic understanding.