An asynchronous language input mode closed-loop control method based on evolution state machine feedback

By calculating the language stage state and acquiring multi-source signals, the language input mode is dynamically switched, which solves the problem of discontinuous input mode switching in the existing technology. It realizes stable language input under low-interaction or screenless terminals, is suitable for multi-terminal collaboration, and supports the long-term development of language capabilities.

CN122450404APending Publication Date: 2026-07-24SHANGHAI SENING INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI SENING INFORMATION TECH CO LTD
Filing Date
2026-01-09
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing language learning or language interaction systems lack a unified scheduling and dynamic switching control method for multiple target language input modes, resulting in discontinuous input mode switching and affecting the continuity and consistency of language input. This is especially difficult to adapt to scenarios involving young users, screenless terminals, or long-term immersive use.

Method used

By introducing a language stage state calculation and constraint mechanism, collecting multi-source trigger signals, and setting mode switching conditions, dynamic switching and control of multiple target language input modes are realized, including voice input, interactive behavior, environmental signals, etc. The mode switching decision is made in combination with the current language stage state, and a cross-aliasing algorithm is used to ensure continuity.

Benefits of technology

It enables orderly switching between multiple language input modes without relying on explicit teaching or human intervention, adapts to multi-terminal collaborative operation, is suitable for low-interaction or screenless terminals, maintains a long-term continuous and stable language input environment, and supports the evolution of language ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122450404A_ABST
    Figure CN122450404A_ABST
Patent Text Reader

Abstract

The application discloses an asynchronous language input mode closed-loop control method and system based on evolution state machine feedback. The method comprises: continuously maintaining a multi-dimensional evolution state matrix representing a target voice input process; collecting multi-source trigger signals containing voice features, behavior interaction and scene labels; calculating mode switching gain based on the matching degree of the evolution state matrix and the multi-source trigger signals, and driving the system to dynamically migrate between multiple asynchronous audio input modes when the gain meets the standard. The method performs smooth compensation of the audio stream gain during the migration process, and synchronously updates the audio processing path and the output strategy. The system reversely corrects the evolution state matrix according to the feedback features fed back by the second terminal in real time, forms a closed-loop scheduling of the target language auditory load output frequency, complexity and cooperation ratio. The application realizes long-term consistency and logical coherence of the audio input environment under the condition of non-text starting and low interaction intervention, and reserves a synchronous alignment interface for cross-modal semantic mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of language input system control technology, and in particular to a method for dynamically switching and controlling multiple target language input modes in a target language immersion learning scenario. It is applicable to language input organization systems under multi-terminal collaboration, low interactivity, or non-explicit teaching conditions. Background Technology

[0002] In existing language learning or language interaction systems, language input modes typically operate through fixed procedures or manual selection. For example, switching between playback, interaction, or response modes often relies on explicit user operation or preset teaching processes, making it difficult to adapt to younger users, screenless devices, or long-term immersive usage scenarios.

[0003] In language acquisition, especially in immersive environments where auditory input is the primary mode, different language input modes (such as passive auditory input, real-time responsive input, and environment-triggered input) exhibit varying degrees of effectiveness at different stages and in different contexts. However, existing technologies lack a unified control method that can dynamically switch between multiple target language input modes based on user behavior, system status, and environmental factors, while maintaining input continuity. Furthermore, the switching logic often lacks stable constraints, easily leading to language input interruptions and unpredictable system behavior, thus affecting the continuity and consistency of language input.

[0004] Existing technologies generally lack a unified control method that models the language input stage, system operating state, and input mode scheduling, making it difficult to form an accumulative and evolving operational relationship between different input modes.

[0005] Therefore, it is necessary to propose a dynamic switching and control method for unified scheduling of target language input modes at the system operation level, so as to achieve orderly switching between multiple input modes without relying on explicit teaching instructions or frequent manual intervention, thereby constructing a long-term, continuous and low-interference language input environment. Purpose of the invention

[0006] The purpose of this invention is to provide a dynamic switching and control method for target language immersion input modes. By introducing a calculation and constraint mechanism for language stage states, it enables orderly switching between multiple target language input modes, thereby constructing a long-term, continuous and stable target language input environment without relying on explicit teaching processes or manual intervention. Technical solution

[0007] To achieve the above objectives, the present invention adopts the following technical solution.

[0008] A method for dynamically switching and controlling a target language immersive input mode is applied to a language input system including a server, at least one first terminal, and at least one second terminal, wherein the first terminal has voice acquisition and network communication capabilities, and the second terminal has voice acquisition, voice playback, and network communication capabilities. The method includes the following steps:

[0009] S1: Language Phase State Initialization and Maintenance

[0010] When the system starts up or resets, it obtains the target language stage status information of the current user, sets the initial target language input mode and corresponding running parameters, and continuously updates the language stage status during system operation.

[0011] The language stage state is a computable system state quantity, including at least:

[0012] • One of the following: input strength, input complexity, or input response requirements;

[0013] • The cumulative duration interval of continuous target language input;

[0014] • The frequency or proportion of interruptions during target language input;

[0015] • The probability range of the second terminal responding to the target language input within the most recent preset time window.

[0016] S2: Multi-source trigger signal acquisition

[0017] During system operation, trigger signals related to input mode switching are continuously collected, including but not limited to:

[0018] • Voice input signals from the first or second terminal;

[0019] • Interactive behavior signals from the second terminal;

[0020] • Scene triggering signals based on time, location, or environmental identifiers;

[0021] • Time-triggered signals that are related to system runtime or usage frequency.

[0022] S3: Calculation of Mode Switching Conditions

[0023] Based on the collected trigger signals, combined with the current target language input mode and language stage status, it is determined whether the preset mode switching conditions are met.

[0024] The mode switching conditions are used to limit the conversion relationship between different input modes and to avoid frequent switching or input interruption.

[0025] The mode switching conditions include at least one of the following constraints:

[0026] • Minimum dwell time constraint for input patterns;

[0027] • Trigger constraint when the cumulative amount of target language input reaches a preset threshold;

[0028] • One-way or semi-one-way transfer constraints for input mode switching.

[0029] S4: Dynamic switching of input mode and parameter update

[0030] When the mode switching conditions are met, the control system switches from the current target language input mode to the target input mode, as shown in the attached diagram. Figure 2 As shown, the running parameters corresponding to the target input mode are updated synchronously.

[0031] The target language input mode can be any one of a variety of predefined modes, and different modes have at least one difference in the triggering method, output rhythm or interaction requirements of language input.

[0032] S5: Pattern-based input / output path control

[0033] Based on the current target language input pattern, select the corresponding audio processing path or output strategy, including at least one:

[0034] Different audio caching and playback triggering strategies;

[0035] Different voice input processing delays or real-time processing strategies;

[0036] • Whether to enable real-time or delayed processing mechanisms for voice input.

[0037] S6: Status Feedback and Continuous Scheduling

[0038] During the target language input and output process, the system operation status and user behavior changes are continuously monitored, and subsequent mode switching judgments and input / output strategies are dynamically adjusted based on feedback results to maintain the continuity and consistency of target language input. Beneficial effects

[0039] Compared with the prior art, the present invention has at least the following beneficial effects:

[0040] 1. Organize target language input from the system operation level rather than the teaching process level to achieve unified scheduling and control of input modes;

[0041] 2. Through a dynamic switching mechanism, multiple immersive input modes can work together in an orderly manner during long-term use, avoiding input fatigue or efficiency decline caused by a single mode;

[0042] 3. Supports multi-terminal collaborative operation, and does not depend on specific terminal form or explicit operation method, adapting to low-interaction or screenless terminal scenarios;

[0043] 4. It has good scalability and is suitable for different language pairs and usage scenarios;

[0044] 5. To achieve a systematic organization of auditory input of the target language without requiring text presentation as a prerequisite;

[0045] 6. Reserve system interfaces for future language capability evolution. Attached Figure Description

[0046] Figure 1 Flowchart of dynamic switching and control method for target language immersive input mode

[0047] Figure 2 Schematic diagram of the mapping relationship between language stage states and target language input patterns Specific Implementation

[0048] Example 1: A basic operating method for a target language immersive input mode

[0049] This embodiment illustrates the implementation of the method of the present invention in a basic operating state.

[0050] When the system is running, it first acquires the user's language stage status information. The language stage status is used to characterize the user's stage position in the target language acquisition process. This status is a long-term state variable maintained internally by the system and is not calculated in real time based on a single usage behavior.

[0051] The system pre-stores multiple target language immersive input modes. Each input mode defines at least one control parameter for the target language and the first language in auditory output, including but not limited to output ratio, output order, and whether it is accompanied by semantic auxiliary output.

[0052] After obtaining the current language stage status, the system selects the input mode that matches the stage status according to the preset stage-mode correspondence, and outputs the auditory content of the target language to the user in the input mode.

[0053] In this embodiment, the system does not provide the user with target language content in text form during operation, nor does it require the user to perform explicit learning operations, thereby achieving immersive organization of target language auditory input.

[0054] Example 2: A method for dynamic switching of input modes based on language stage states

[0055] Based on Example 1, this example further illustrates the dynamic switching process of the input mode.

[0056] During continuous operation, the system updates the user's language stage status. These stage status updates can be based on usage duration, usage frequency, historical stage paths, or other stage evolution rules, without relying on immediate test results or explicit feedback.

[0057] When the system detects a change in the language stage state, it triggers an input mode switching decision process. Based on the updated stage state, the system re-determines the appropriate input mode from among the various input modes.

[0058] During the switching process, the system performs smooth transition control on the ongoing auditory input, so that the target language input can complete the adjustment of mode parameters while maintaining content continuity, thereby avoiding frequent interruptions or abrupt changes.

[0059] To maintain the continuity of the immersive environment, this method employs a cross-fading algorithm in its dynamic switching process. When switching from "passive hearing mode" to "interactive response mode," the system does not immediately interrupt the current audio stream. Instead, it reduces the gain of the current stream while simultaneously increasing the weight of the interactive guide stream according to a preset envelope curve, ensuring the physical continuity of the acoustic environment.

[0060] The above method enables dynamic switching of the target language input mode as the language stage changes.

[0061] Example 3: An Implementation Method for Asymmetric Collaborative Output of First Language and Target Language

[0062] This embodiment illustrates the collaborative method between the first language and the target language in input mode control.

[0063] In this embodiment, the input pattern defines an asymmetric relationship between the first language and the target language in the system output. The target language is continuously output as the primary auditory input, while the first language appears only in an auxiliary form under specific conditions.

[0064] The asymmetric collaborative output mode features automatic attenuation. Based on the "semantic proficiency" index in the evolutionary state matrix, the system dynamically compresses the output duration and frequency of the first language auxiliary audio track. This switching is sub-second-level, driven by the system's automatically perceived cognitive load, requiring no user intervention.

[0065] The specific conditions include, but are not limited to: the stage state being within a preset range, structural nodes of the target language content, or predefined semantic prompt trigger points of the system.

[0066] In this input mode, the system controls the output frequency, position, and duration of the first language so that it does not form a complete teaching explanation, but only serves as implicit semantic support for the target language input.

[0067] Through the above methods, the system achieves collaborative output of the first language and the target language without explicit instruction.

[0068] Example 4: An Input Mode Control Method for Multi-Terminal Collaborative Operation

[0069] This embodiment illustrates the application of the method of the present invention in a multi-terminal environment.

[0070] The system includes at least one first terminal with voice acquisition and network connectivity capabilities, and one second terminal with voice acquisition, voice playback, and network connectivity capabilities. The second terminal is used by the target user to receive auditory input in a specific usage scenario.

[0071] During system operation, the language stage status and input mode control logic are uniformly maintained by the first terminal or cloud system, and control instructions corresponding to the current input mode are sent to the second terminal.

[0072] The second terminal outputs auditory content in the target language according to the control instructions, and collects the user's voice input information for system status updates when necessary.

[0073] The above method achieves the decoupling of input mode control logic from terminal form.

[0074] Example 5: A method for phase consistency control in long-term continuous use

[0075] This embodiment illustrates the operation of the method of the present invention during long-term use.

[0076] The system continuously saves the user's language stage state during multiple usage cycles, so that the language input starting point does not need to be re-initialized each time it starts.

[0077] When a user restarts the system after a period of inactivity, the system directly restores the corresponding input mode based on the historical state, rather than reverting to the initial state.

[0078] In this way, auditory input of the target language maintains consistency over long-term use, avoiding input fragmentation.

[0079] Example 6: An implementation method suitable for screenless or low-interaction terminals

[0080] This embodiment illustrates the application in terminals with no text display or low interactivity.

[0081] In this embodiment, the second terminal does not have or does not enable text display function, and the system operates only through auditory output and minimal interactive signals.

[0082] The system's input mode switching, stage status maintenance, and language output control are all completed in the background, requiring no explicit operation from the user.

[0083] Therefore, the method of the present invention is applicable to application scenarios such as wearable devices and embedded terminals.

[0084] Example 7: An Implementation Method for Input Mode Combination Control

[0085] In this embodiment, the system is not limited to using only a single input mode when the system is in the same language stage state.

[0086] The system combines multiple input modes according to preset rules, using time segments or content structures as units, and maintains consistency in the overall input strategy during the combination process.

[0087] The combination method is automatically controlled by the system, and the user does not need to be aware of the specific mode changes.

[0088] Example 8: An implementation method for reserving interfaces for subsequent language presentation formats

[0089] This embodiment illustrates the scalability of the method of the present invention.

[0090] The system is designed to decouple the input mode control module from the language presentation format. In this embodiment, the language presentation format is auditory output.

[0091] When the system introduces other language presentation formats, only the corresponding presentation parameters need to be added to the input mode control module, without changing the stage state management logic.

[0092] The mode switching command in this method not only includes audio control parameters but also carries a set of synchronization logic index packets. These index packets inform the system which logical coordinates the background semantic world model should synchronously jump to at the current audio mode switching point. This provides a precise timestamp alignment reference for future text mapping rendering or spatial audio localization.

[0093] In this way, the method of the present invention provides a system-level expansion space for the subsequent evolution of language capabilities.

Claims

1. A closed-loop control method for asynchronous language input patterns based on evolutionary state machine feedback, applied to a system including a server, a first terminal, and a second terminal, characterized in that, The method includes: Maintaining the evolution state matrix: Real-time updating of the evolution state matrix S, which represents the target speech input process, based on the cumulative entropy of the audio payload, the response delay gradient, and the interruption rate vector; Feature signal deconstruction: Continuously acquire multi-source trigger signals and deconstruct the time-domain features and attribution labels of the signals; Evaluation of switching strategy: Based on the matching degree between the evolution state matrix S and the current operating mode, calculate the mode switching gain coefficient; Closed-loop mode migration: When the gain coefficient reaches the preset threshold, the control system migrates from the current asynchronous input mode to the target mode and performs smooth fade-in and fade-out compensation for the audio stream gain. Parameter adaptive update: Based on the migrated mode, the audio output path and voice acquisition sensitivity of the second terminal are adjusted in real time to form a closed-loop scheduling.

2. The method according to claim 1, characterized in that, The language stage state is determined based on at least one or more of the following: the cumulative duration interval of continuous target language input, the frequency of target language input interruption, and the probability interval of the second terminal generating a response within a preset time window.

3. The method according to claim 1, characterized in that, The mode switching conditions include at least the minimum dwell time constraint of the input mode or the cumulative threshold constraint of the target language input.

4. The method according to claim 1, characterized in that, Different target language input modes correspond to different audio processing paths, and the audio processing paths differ in at least one of the following: audio caching strategy, playback triggering conditions, or voice input processing latency.

5. The method according to claim 1, characterized in that, The trigger signal includes at least one of the following: voice acquisition signal, interactive behavior signal, environmental trigger signal, or time trigger signal.

6. The method according to claim 1, characterized in that, The target language input content is not required to be presented as text during system operation.

7. The method according to claim 1, characterized in that, The method supports operation whether the input language and the target language are the same or different, and adjusts the input and output processing paths based on the consistency relationship.

8. The method according to claim 1, characterized in that, The switching of the input mode does not depend on the user's explicit selection of a specific input mode.

9. The method according to claim 1, characterized in that, The second terminal is activated as needed during system operation and participates in the output or interaction of target language input when activated.

10. The method according to claim 1, characterized in that, The method provides an interface for the subsequent evolution of target language capabilities, supporting language presentation formats other than auditory input.