AI-based keyboard error automatic correction method and system

By adopting AI dual-layer asynchronous collaborative architecture in the keyboard input error correction system, combining local lightweight models and cloud-based large models, using user input features to construct an intention matrix for error correction, and through recursive knowledge distillation and user feedback dynamic adjustment strategies, the problems of error correction quality and response delay in the existing technology are solved, achieving efficient and personalized keyboard input error correction effects.

CN120215722AActive Publication Date: 2025-06-27渴创技术(深圳)有限公司

Patent Information

Application Number
CN202510697959.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-06-27
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing keyboard input error correction technology performs poorly in complex contexts and lacks device adaptability and personalized adjustment capabilities, resulting in insufficient response delays and error correction quality.

Method used

Using a dual-layer asynchronous collaboration architecture based on AI, combining local lightweight models and cloud-based large models, the input intention probability matrix is ​​constructed by collecting the contact pressure distribution, sliding trajectory and residence time input by user, error detection and correction are carried out, and error correction strategies are dynamically adjusted through recursive knowledge distillation and user feedback.

Benefits of technology

It realizes the improvement of error correction quality while ensuring response speed, reducing user perception delay, providing a smoother input experience, and adapting to different user input habits and device status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215722A_ABST
    Figure CN120215722A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of keyboard error correction, and discloses an AI-based keyboard error automatic correction method and system, and the method comprises the steps: collecting the contact pressure distribution, sliding track and stay time of keyboard input by a user, and constructing an input intention probability matrix; performing error detection and preliminary correction through the first local lightweight model to obtain a first correction result, and transmitting the first correction result to the cloud large model for deep semantic analysis to obtain a second correction result; performing recursive knowledge distillation on the first local lightweight model according to the second correction result to obtain a second local lightweight model; and collecting behavior feedback data of accepting or rejecting the second correction result by the user, performing task allocation dynamic adjustment on the second local lightweight model and the cloud large model, and outputting a double-layer asynchronous collaborative error correction strategy, thereby improving the error correction quality while ensuring the response speed, effectively reducing the user perception delay, and improving the user experience. And smoother input experience is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of keyboard error correction, and particularly to an AI-based automatic keyboard error correction method and system. Background Art

[0002] Traditional input error correction methods mainly rely on static dictionaries and simple rules, and cannot effectively handle errors in complex contexts, especially performing poorly in non-standard text environments such as social media abbreviations, technical terms, and multilingual mixed scenarios. Although existing cloud-based large model solutions have made breakthroughs in accuracy, they consume huge computing resources, resulting in significant response delays and seriously affecting the user experience; while lightweight local models can achieve fast responses, they have obvious deficiencies in complex semantic understanding and context-related error handling.

[0003] The current keyboard input error correction technologies on the market lack effective device adaptability and personalized adjustment capabilities, and cannot dynamically adjust error correction strategies according to different users' input habits and device states. Especially on resource-constrained mobile devices, factors such as computing power, battery life, and network conditions will significantly affect the performance and usability of the error correction system. Most solutions either rely entirely on cloud processing, resulting in unstable experiences during network fluctuations, or only rely on simple local models, resulting in insufficient error correction quality, lacking a collaborative architecture that combines the advantages of both. Summary of the Invention

[0004] This application provides an AI-based automatic keyboard error correction method and system, thereby improving the error correction quality while ensuring the response speed, effectively reducing the user-perceived latency, and providing a smoother input experience.

[0005] In the first aspect of this application, an AI-based automatic keyboard error correction method is provided. The AI-based automatic keyboard error correction method includes: Collect the contact pressure distribution, sliding trajectory, and dwell time of the user's keyboard input and construct an input intention probability matrix; Input the input intention probability matrix into a first local lightweight model for error detection and preliminary correction to obtain a first correction result; Transmit the first correction result and the input intention probability matrix to a cloud large model for in-depth semantic analysis to obtain a second correction result; Perform recursive knowledge distillation on the first local lightweight model according to the second correction result to obtain a second local lightweight model; Collect the behavioral feedback data of the user's acceptance or rejection of the second correction result, and dynamically adjust the task allocation of the second local lightweight model and the cloud large model based on the acceptance or rejection behavioral feedback data, and output a two-layer asynchronous collaborative error correction strategy.

[0006] Combined with the first aspect, in the first implementation manner of the first aspect of the present invention, collecting the contact pressure distribution, sliding trajectory, and dwell time of the user's keyboard input and constructing an input intention probability matrix includes: Collecting key pressure data through an embedded sensor array in the user device to obtain contact pressure distribution data; Tracking the movement coordinates and corresponding timestamps of the user's finger on the touch screen to obtain sliding trajectory data; Recording the contact start time and release time of the user's key press to obtain key time series feature data; Performing intention mapping calculation based on the contact pressure distribution data, the time series feature data, and the sliding trajectory data to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input.

[0007] Combined with the first aspect, in the second implementation manner of the first aspect of the present invention, performing intention mapping calculation based on the contact pressure distribution data, the time series feature data, and the sliding trajectory data to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input includes: Performing Gaussian distribution calculation on the contact pressure distribution data to obtain a key position heat map, where the central value of the heat map represents the key press intention intensity; Calculating the user's finger movement vector and speed characteristics according to the sliding trajectory data to obtain a direction offset matrix; Calculating the time interval ratio between adjacent characters based on the key time series feature data to obtain a time correlation weight; Performing Bayesian probability fusion on the key position heat map, the direction offset matrix, and the time correlation weight to obtain an initial intention probability distribution; Combined with the standard keyboard layout geometric distance matrix, performing spatial correlation weighted adjustment on the initial intention probability distribution to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input.

[0008] Combined with the first aspect, in the third implementation manner of the first aspect of the present invention, inputting the input intention probability matrix into the first local lightweight model for error detection and preliminary correction to obtain a first correction result includes: Mapping the input character sequence through the character-level embedding layer of the first local lightweight model to obtain a character vector sequence; Inputting the character vector sequence into the lightweight expert mixture layer of the first local lightweight model, and distributing the input to 4 expert sub-networks through the routing network for parallel feature extraction to obtain a multi-perspective feature representation; Selectively activate the multi-view feature representation using a simplified sparse activation module, only retaining the key information path to obtain a sparse feature matrix; Input the sparse feature matrix and the input intention probability matrix into a lightweight multi-level attention structure for cross-modal information fusion to obtain a fused semantic representation. The lightweight multi-level attention structure includes 4 attention heads and 2 layers of Transformer structures; Input the fused semantic representation into the error detection branch and the correction branch of the dual-branch prediction head in the first local lightweight model for parallel computing, and perform probability weighting in combination with a preset n-gram word frequency statistical table to obtain a first correction result.

[0009] Combined with the first aspect, in the fourth implementation manner of the first aspect of the present invention, the transmitting the first correction result and the input intention probability matrix to a cloud large model for in-depth semantic analysis to obtain a second correction result includes: Transmit the first correction result, the input character sequence, and the input intention probability matrix to a cloud server to obtain a cloud input data set; Perform a merging process on the cloud input data set, the user's historical input sequence, and the application context information to obtain a target context data set; Input the target context data set into a cloud large model that adopts a mixture of experts architecture, a sparse activation mechanism, and a multi-level attention structure for processing to obtain in-depth semantic association data; Perform conditional probability maximization calculation based on the in-depth semantic association data to generate an optimal corrected text sequence, and return the optimal corrected text sequence as the second correction result to the user device.

[0010] Combined with the first aspect, in the fifth implementation manner of the first aspect of the present invention, the recursively distilling knowledge of the first local lightweight model according to the second correction result to obtain a second local lightweight model includes: Use a pseudo-label generator to pair the second correction result and the input character sequence to obtain a pseudo-label training sample pair; Calculate the correction consistency, prediction confidence, and uniqueness of the pseudo-label training sample pair through a quality evaluation function, and generate a sample quality score set based on the correction consistency, the prediction confidence, and the uniqueness; Perform threshold screening on the pseudo-label training sample pair based on the sample quality score set, and add the samples in the pseudo-label training sample pair with a sample quality score higher than a preset value to the recursive training cache to obtain a target distillation training set; Input the target distillation training set into the dynamic weighted distillation module, and train the first local lightweight model using the combined distillation loss function to obtain a knowledge compression parameter set. The first local lightweight model includes a character-level embedding layer, a lightweight mixture-of-experts layer, a simplified sparse activation module, a lightweight multi-level attention structure, and a dual-branch prediction head. Update the parameters and optimize the structure of the first local lightweight model according to the knowledge compression parameter set. At the same time, establish a reverse knowledge flow channel to summarize new error patterns to the server, and obtain a second local lightweight model.

[0011] Combined with the first aspect, in the sixth implementation manner of the first aspect of the present invention, collect the behavioral feedback data of the user's acceptance or rejection of the second correction result, and dynamically adjust the task allocation of the second local lightweight model and the cloud large model based on the acceptance or rejection behavioral feedback data, and output a two-layer asynchronous collaborative error correction strategy, including: Monitor and collect the user's selection operation of accepting or rejecting the correction suggestion, the corrected editing behavior, the response time, and the acceptance rate in the application scenario to obtain behavioral feedback data; Perform Bayesian weight optimization calculation on the behavioral feedback data to obtain an error correction satisfaction score; Construct personalized tuning data including the user's commonly used vocabulary, error type distribution, and scenario preferences based on the error correction satisfaction score; Monitor and calculate the device processor load, memory occupancy, battery power, and network latency, and generate a task allocation plan in combination with the personalized tuning data; Configure and adjust the computing tasks, data transmission frequency, and error handling logic between the second local lightweight model and the cloud large model according to the task allocation plan to obtain a two-layer asynchronous collaborative error correction strategy.

[0012] Combined with the first aspect, in the seventh implementation manner of the first aspect of the present invention, configure and adjust the computing tasks, data transmission frequency, and error handling logic between the second local lightweight model and the cloud large model according to the task allocation plan to obtain a two-layer asynchronous collaborative error correction strategy, including: Calculate the semantic complexity of the input character sequence in real time according to the task allocation plan, and use an adaptive threshold function to assign simple spelling errors to the second local lightweight model for processing, and assign semantic dependency errors to the cloud large model for processing, to obtain a binary error type diversion mechanism; Dynamically adjust the computing resource allocation of the second local lightweight model based on the binary error type diversion mechanism, reserve computing resources for keyboard input in high-frequency usage areas, and at the same time execute an upper limit control on the call frequency of the cloud large model to obtain a differential response scheduling strategy; Perform n-gram analysis based on the user's historical input patterns, construct a keyboard input prediction tree, and trigger potential error correction calculations in advance according to the keyboard input prediction tree to obtain a pre-judgment error correction cache; Based on the pre-judgment error correction cache and the differential response scheduling strategy, use the token bucket algorithm to control the smoothness of sending requests to the cloud large model, and implement network condition-aware request batch merging to obtain an anti-interference communication pipeline; Combine the anti-interference communication pipeline with the input delay sensitivity function, obtain results from the second local lightweight model, and allow the results of the cloud large model to be returned asynchronously and replaced without perception, forming a two-layer asynchronous collaborative error correction strategy.

[0013] The second aspect of this application provides an AI-based automatic keyboard error correction system, and the AI-based automatic keyboard error correction system includes: A collection module for collecting the contact pressure distribution, sliding trajectory, and dwell time of the user's keyboard input and constructing an input intention probability matrix; A preliminary correction module for inputting the input intention probability matrix into a first local lightweight model for error detection and preliminary correction to obtain a first correction result; A semantic analysis module for transmitting the first correction result and the input intention probability matrix to a cloud large model for in-depth semantic analysis to obtain a second correction result; A knowledge distillation module for recursively distilling knowledge from the first local lightweight model according to the second correction result to obtain a second local lightweight model; An output module for collecting the user's behavioral feedback data on accepting or rejecting the second correction result, and dynamically adjusting the task allocation of the second local lightweight model and the cloud large model based on the accepting or rejecting behavioral feedback data, and outputting a two-layer asynchronous collaborative error correction strategy.

[0014] Compared with the prior art, the present application has the following beneficial effects: By establishing a two-layer asynchronous collaborative architecture of a lightweight local model and a large cloud model, a balance between real-time response and high accuracy is achieved. The local model is responsible for immediate response, while the large cloud model asynchronously optimizes the results, thereby improving the error correction quality while ensuring the response speed. The physical characteristics of the user's keystrokes are captured by the contact dynamic feature perception network, and an input intention probability matrix is constructed. Compared with the traditional method that only relies on static key position information, it can more accurately understand the user's true input intention, especially in the scenarios of fast typing and one-handed operation. Through the recursive knowledge distillation self-correction mechanism, the knowledge of the large cloud model is continuously compressed into the local lightweight model, enabling the local model to continuously self-optimize, reducing the dependence on a large amount of labeled data, and achieving continuous improvement of the system performance over time. Based on the collection of user feedback signals, a personalized tuning network is established to perform fine-grained adjustment according to the input habits and application scenarios of different users, improve the error correction acceptance rate, and enhance the user experience, especially in terms of specific domain technical terms and personal habitual expressions. The local lightweight model adopts a lightweight expert mixing layer, a sparse activation mechanism, and a multi-level attention structure, which improves the model expression ability while maintaining low computational complexity, and achieves efficient error recognition and correction. By constructing an anti-interference communication pipeline and a differential response scheduling strategy, the system can operate stably under various network conditions and device states, and improves the adaptability of the system in weak network environments and low-resource devices. Based on the user's historical input patterns, a keyboard input prediction tree is constructed to trigger potential error correction calculations in advance, and combined with the unperceived replacement mechanism, effectively reduces the user's perceived latency and provides a smoother input experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] The structures, ratios, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention.

[0017] Figure 1 It is a schematic flowchart of the AI-based automatic keyboard error correction method provided by the embodiment of the present invention; Figure 2 It is a schematic block diagram of the structure of the AI-based keyboard error automatic correction system provided by an embodiment of the present invention. Specific embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0020] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0021] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Please refer to Figure 1 , an embodiment of the AI-based keyboard error automatic correction method in the embodiments of this application includes: Step 100, collect the contact pressure distribution, sliding trajectory, and residence time of the user's keyboard input and construct an input intention probability matrix; It can be understood that the execution subject of this application can be an AI-based keyboard error automatic correction system, or a terminal or a server, and specific limitations are not made here. In the embodiments of this application, the server is used as the execution subject for illustration.

[0022] Specifically, a high-precision input dynamic perception layer is constructed based on the embedded sensor array on the user device side. This perception layer real-time acquires multi-modal physical information in each user key operation to reflect the actual input behavior of the user. At the moment when the user presses a certain key, the sensor array collects the pressure response information of the contact area on the key surface to form a two-dimensional pressure distribution matrix, which includes the maximum pressure value at the center point of the key, the pressure gradient information in the surrounding area, and the spatial position information corresponding to each pressure point, forming high-resolution contact pressure distribution data. At the same time, the trajectory generated when the user's finger moves on the touch screen surface is continuously tracked. This process generates a sliding trajectory data sequence by recording the coordinate positions and timestamp information of the finger at each time point. The timing behavior of each key operation is modeled, that is, the whole process from the finger first contacting the key surface to the final complete release is recorded, and time series feature data is established through continuous time sampling, including the contact start time, the pressure rising edge, the time when the maximum pressure appears, the pressure falling edge, and the final release time, to characterize the persistence and rhythm of the input action. The contact pressure distribution data, time series feature data, and sliding trajectory data are jointly input into a specific contact dynamic feature perception network. This network adopts a multi-channel input structure, performs local convolution processing and global timing modeling on the spatial pressure map, sliding trajectory curve, and time series vector respectively, and performs multi-modal fusion in the middle layer. By introducing an attention mechanism, the response intensity of the intention area is enhanced, so as to realize the deep mapping of the user input intention. The network output is a two-dimensional probability matrix, where each matrix element represents the conditional probability that the true intention is character j when the user presses character i, forming an input intention probability matrix.

[0023] Model the contact pressure distribution data with a Gaussian distribution. Taking the maximum pressure point of the button detected in the sensor array as the center, calculate the pressure attenuation degree in the horizontal and vertical directions, fit it into a two-dimensional Gaussian function, and form a key heat map. The Gaussian peak corresponding to the center of the image represents the intensity of the button press intention. The closer to the center, the clearer the intention. Conversely, it represents an increased possibility of deviating from the target button. The edge area of the heat map captures non-centered pressing behaviors that may be caused by finger tilt, rotation, or accidental touch of the user. Calculate the motion vector sequence of the finger on the touch panel according to the sliding trajectory data, obtain the speed and direction changes through the coordinate difference between consecutive frames in time, and then construct a direction offset matrix. This matrix reflects the inertial offset pattern of the finger path before or during the click by the user, and can prompt the system whether there is a drag-type misoperation or a deviation trend of the target character. At the same time, based on the key time series feature data, extract the time interval between each consecutive input, calculate the relative time ratio of the target character and its surrounding characters on the time axis, and establish a time correlation weight to express the connection between the current input and adjacent inputs in terms of rhythm and continuity. This weight can reveal whether the user continuously clicks on misaligned characters or whether there is an incomplete input beat. Through the Bayesian probability fusion mechanism, jointly model the spatial concentration information reflected by the key heat map, the trajectory offset trend expressed by the direction offset matrix, and the input rhythm dependence characterized by the time correlation weight, construct the coupling relationship between the prior and the likelihood, and obtain the initial intention probability distribution of each character pair for the candidate characters. In order to improve the positioning accuracy of this distribution in the actual layout, introduce the geometric distance matrix in the standard keyboard layout as a spatial correction factor, and use the physical distance between each candidate character and the currently pressed character as a penalty factor to participate in the adjustment of the intention distribution. That is, if a certain candidate character has a high weight in the initial intention probability, but it is far from the current character in the layout, it will be weakened according to the distance weighting function, while adjacent characters will have their probabilities increased in the intention ambiguity scenario. This series of processes constructs an input intention probability matrix that conforms to the spatial layout logic, has the ability to perceive the trajectory movement trend, and integrates the time dynamic characteristics. Each element in the matrix represents the conditional probability that when the user presses a certain character, their true intention is another character.

[0024] Step 200: Input the input intention probability matrix into the first local lightweight model for error detection and preliminary correction to obtain the first correction result; Specifically, the input character sequence is mapped through the character-level embedding layer of the first local lightweight model. A compact vectorization mechanism is adopted to map each input character through a 64-dimensional low-dimensional embedding space, generating a set of continuous character vector sequences. The character vector sequences are input into the lightweight mixture-of-experts layer of the first local lightweight model. In this layer, a routing network is configured, whose role is to dynamically allocate the input vectors according to the current context semantics, and route the data of different segments to 4 independent expert sub-networks respectively. Each sub-network has a convolution-attention hybrid coding module with different structures, so as to achieve multi-perspective parallel feature extraction while keeping the number of parameters below 2MB. The sub-networks complement each other in capturing different levels of language features such as spelling patterns, syntactic structures, and character position dependencies, and output a fused multi-perspective feature representation. To compress the computational amount and focus on the input signals with the most error-correcting value, a simplified sparse activation module is introduced to perform feature selection on the above multi-perspective feature representation. The activation path is controlled by a gating unit, only activating the channels with higher information density while masking redundant or ambiguous features, forming a feature matrix with a sparse structure. The sparse feature matrix is fused with the input intent probability matrix, and cross-modal information interaction is achieved through a lightweight multi-level attention structure. This attention structure consists of a two-layer Transformer composed of 4 attention heads and has low-latency modeling ability. During the fusion process, the system extracts the coupling mapping between the dependencies between characters and the input intent offset signals respectively, and performs context weighting reinforcement on each character to generate a comprehensive semantic expression representing the error-correcting state of the current input in terms of structure, intent, and context. The fused semantic representation is respectively input into the error detection branch and the correction branch of the dual-branch prediction head in the local model. The error detection branch uses the convolutional receptive field combined with the attention mechanism to predict the error probability of each character at each position, while the correction branch constructs a corresponding alternative probability distribution for each character, calculates the optimal set of alternative characters through a classifier. During this process, the system calls the embedded n-gram word frequency statistical table to perform probability-weighted fusion at the language level on the candidate alternative results and the context lexical fluency to improve the overall text semantic coherence and error-correcting rationality, and finally outputs the first correction result.

[0025] Step 300: Transmit the first correction result and the input intent probability matrix to the cloud large model for in-depth semantic analysis to obtain the second correction result; It should be noted that the first correction result, the input character sequence, and the input intent probability matrix are transmitted to the cloud server to obtain a cloud input dataset. This dataset encapsulates the correction judgments made by the local model at the character level and the touchpoint intent information demonstrated by the user at the physical input layer, providing an initial state with both semantic and physical input characteristics for the cloud model. The cloud input dataset received at the cloud is merged with the historical input sequences synchronized on the user device side and the context data of the currently activated application. A target context dataset with greater contextual integrity is constructed. The historical input sequences reflect the user's long-term language usage habits, and the application context provides an information field for the current input's semantic environment, such as the email body, search box, editor, etc. The system unifies and formats this information from different sources through context splicing, positional encoding expansion, and context window scrolling mechanisms to ensure its semantic continuity and reasoning relevance. The target context dataset is input into a large language model deployed in the cloud. This large model is designed with a mixture of experts architecture, that is, each layer contains multiple structurally heterogeneous expert sub-modules. During the processing, through a sparse activation mechanism, only the part of the experts most suitable for the current input context is selected to participate in the calculation, thereby maintaining the reasoning efficiency in the large-scale parameter space. At the same time, the model internally nests multi-level attention structures. In each Transformer layer, there are high-order cross-segment attention heads and local subsequence attention modules, enabling the model to capture deep semantic associations across sentences, tasks, and even languages in long text sequences. In this process, the model performs semantic compression and vector mapping on the input data, and then constructs a high-dimensional relationship graph between characters, words, and between the input intent probability matrix and language expressions through an attention-weighted aggregation mechanism, thereby generating deep semantic association data with long-range dependence modeling ability and multi-modal context interaction ability. Based on the deep semantic association data, a conditional probability maximization operation is performed to construct an error correction output space, and under the premise of the target conditions, a greedy decoding and beam search strategy are adopted to generate an optimal corrected text sequence. This sequence literally corrects the misspelled words, missing words, and incorrect word order in the input, and optimizes semantic coherence, context consistency, and language style unity through a context integration mechanism to form a second correction result. The optimal corrected text sequence is encapsulated as a cloud response result and asynchronously returned to the user device.

[0026] Step 400: Recursively perform knowledge distillation on the first local lightweight model according to the second correction result to obtain a second local lightweight model; Specifically, a pseudo-label generator is used to pair the second correction result with the input character sequence to form a pseudo-label training sample pair consisting of an input-output mapping relationship. Such sample pairs serve as the carrier for knowledge transfer between the teacher model and the student model and are screened to ensure the effectiveness and stability of distillation training. The sample pairs are input into a quality assessment function, which quantitatively evaluates each pair of pseudo-label samples from three dimensions: one is correction consistency, that is, the degree of overlap between the second correction result and the local first correction result at the key character positions; the second is prediction confidence, which is statistically obtained from the maximum value of the softmax probability at the token level in the output of the cloud model; the third is the uniqueness index, which is used to determine whether the sample appears frequently in the historical training set to prevent overfitting of the training to repeated patterns. By weighted integration of these three indicators, a sample quality score is generated for each pair of samples, and a sample quality score set is constructed. Based on the score set, a pseudo-label sample screening operation is performed. A dynamic threshold mechanism is set, and the upper and lower limit ranges of the threshold are determined according to the current device resource status, training frequency, and historical model convergence speed. Only the pseudo-label samples with a sample quality score higher than the preset threshold are added to the local recursive training cache to form the target distillation training set for knowledge distillation. The target training set is input into the local dynamic weighted distillation module, which performs knowledge transfer training based on a combined loss function. It contains two main loss terms: one is the KL divergence term, which is used to measure the difference in the prediction distributions between the student model and the cloud large model, and the other is the cross-entropy term, which is used to evaluate the hard alignment ability of the student model to the pseudo-labels. The weights of the two are adjusted through a time decay function, thus realizing a progressive convergence mechanism from "imitating soft knowledge" to "matching explicit labels". During the actual distillation execution process, based on the current model structure, the system distributes the knowledge distillation process to all modules of the first local lightweight model, including the character-level embedding layer, lightweight mixture-of-experts layer, simplified sparse activation module, lightweight multi-level attention structure, and dual-branch prediction head. The system fine-tunes the subset of parameters of each module to ensure the stability of the overall architecture and gradually replaces the existing parameters to form a new knowledge compression parameter set. After training is completed, the parameters of the model are updated according to this parameter set, and in necessary cases, pruning, activation path reconfiguration, or channel reconstruction operations are performed on the sub-network structure of the mixture-of-experts layer to achieve structural-level optimization and obtain the second local lightweight model. To achieve two-way knowledge flow and co-evolution of the model system, a reverse knowledge flow channel is established to upload the difficult-to-fit samples or new error patterns identified during the distillation training process to the cloud server. Such information is used to update the global error pattern database and the re-training data pool of the large model, thereby enhancing the adaptability of the cloud model to dynamic language evolution.The entire knowledge distillation process relies on being automatically triggered during device idle periods, has progressive, self-cleaning, and resource-awareness capabilities, can continuously compress and transmit language knowledge without affecting the user input experience, achieves the goal of the model evolving itself over time, and ultimately enables the second local lightweight model to significantly improve accuracy and generalization ability while retaining real-time performance.

[0027] Step 500: Collect the behavioral feedback data of the user's acceptance or rejection of the second correction result, and dynamically adjust the task allocation of the second local lightweight model and the cloud large model based on the acceptance or rejection behavioral feedback data, and output a two-layer asynchronous collaborative error correction strategy.

[0028] Specifically, continuously monitor key behavioral events generated during the interaction between the user and the error correction system on the device side, including explicit operation selections of whether the user accepts the error correction suggestions, such as clicking the "Accept" or "Withdraw" operations, the editing behavior of whether the user continues to modify the text after the correction suggestions, the response time from the suggestion pop-up to the final confirmation, and the error correction acceptance rate of the user in specific application scenarios, such as in the input method in chat, office, or search environments. By capturing and timestamping the behavioral characteristics of these dimensions, construct a behavior feedback dataset with timeliness and context relevance. Introduce a Bayesian optimization model to perform fusion calculations on these behavior feedbacks, establish an error correction satisfaction scoring function considering the influence weights between different feedback types, and deduce the overall acceptance willingness of the user for the error correction suggestions in the current context through a dynamically updated prior model and posterior inference mechanism, and output a quantified satisfaction scoring result. Based on this scoring result, generate a personalized tuning dataset, which includes a common vocabulary vector table of the user during the long-term input process to reflect the user's preferred word usage habits, and also includes a statistical distribution matrix of its common error types to capture error-prone character pairs or grammar patterns during input. At the same time, record the tolerance threshold and acceptance intensity of the user for automatic error correction in different application scenarios to construct a personalized context tuning parameter set. On this basis, synchronously start the device status perception module to monitor the processor load, memory occupancy, battery power level of the current device, and the latency and bandwidth of the current network connection in real time, and combine and model them with the above personalized tuning data. Calculate the optimal error correction task allocation plan through a resource-intention coupling mapping function. The plan specifies which types of error correction requests should be preferentially completed using the local model or the cloud model under the current device and network status, and sets the call frequency, execution latency budget, and cache synchronization window size of each model. Adjust the runtime configuration of the second local lightweight model and the cloud large model respectively according to the task allocation plan. For example, when the user's error correction satisfaction score is high and the device's operating resources are sufficient, increase the proportion of local model usage and reduce cloud dependence to save energy consumption; while in scenarios with ambiguous semantics or complex contexts, when the user's error correction acceptance rate is low, the system moderately increases the cloud inference frequency and dynamically pulls a larger window of the user's historical data to participate in in-depth semantic analysis. In a weak network environment, the system activates an asynchronous compensation mechanism, preferentially executes the local fast response strategy and temporarily caches the error correction results in the local prediction buffer, and then the cloud compensates for the unified results after the network stabilizes. Through the above mechanism, form a two-layer asynchronous collaborative error correction strategy that continuously adapts to changes in user habits.

[0029] Perform semantic complexity analysis on each input character sequence in real time according to the task allocation scheme. This process extracts the language structure features of the current input segment through nested context windows, constructs a semantic tensor representation by combining dependency grammar relations and part-of-speech tags, and further discriminates through an adaptive threshold function. Problems with clear spelling error features, low context coupling degrees, and those that can be directly solved by local character replacement are classified as simple errors and handed over to the second local lightweight model for processing. For problems with complex semantic dependencies such as semantic drift, long-distance dependencies, and disambiguation word selection, they are assigned to the cloud large model for processing, thus forming a binary error type diversion mechanism with clear task boundaries. After completing the error type diversion, based on this mechanism, adjust the computing resource usage strategy of the second local lightweight model, and preferentially allocate processor cycles, model weight caches, and memory access channels to the keyboard areas with higher current input frequencies, such as the main key area, word start area, or positions in continuous input concentration, to improve the response speed to high-frequency input patterns. At the same time, set a dynamic frequency upper limit for the invocation of the cloud large model, and control the remote inference frequency under the premise of meeting the latency tolerance, so as to avoid triggering frequent invocations under high load or weak network, resulting in device power consumption or response congestion, and form a differential response scheduling strategy. Construct an n-gram analysis structure based on the user's historical input patterns, extract common input prefix combinations, frequently misspelled spelling paths, and word order dependency relationships, and generate a dynamically extended keyboard input prediction tree on this basis. This prediction tree is not only used for language model completion but also for identifying future possible spelling or semantic error input trends. The system uses this prediction result to pre-invoke the second local lightweight model or cache the inference results of the cloud model in advance, forming a pre-judgment error correction cache mechanism to reduce the online computing burden. To achieve double-layer asynchronous collaboration, introduce the token bucket algorithm to smoothly schedule the cloud request traffic, adjust the token generation rate in combination with the input rhythm distribution and network state prediction model, and dynamically determine the request batch processing merging strategy according to the network conditions. Automatically aggregate multiple error correction requests in a weak network environment to reduce the communication load, and at the same time disperse requests in a high-bandwidth state to improve the response real-time performance, forming an anti-interference communication pipeline with bandwidth adaptation ability and anti-jitter performance. Combine this anti-interference communication pipeline with the input delay sensitivity function. This function evaluates the timeliness requirements of the error correction task based on the current input type, application scenario, and user interaction rate. Under the premise of meeting the subjective smooth experience, preferentially obtain quick preliminary correction results from the second local lightweight model, and allow the cloud model to return more refined semantic error correction outputs asynchronously in the background, and perform seamless replacement on the local results after returning, ensuring that the user interface has no delayed jump and the input fluency is not affected, and construct a double-layer asynchronous collaborative error correction strategy with adaptive task routing, predictive response control, asynchronous communication, and semantic optimization capabilities.

[0030] In the embodiments of the present application, by establishing a two - layer asynchronous collaborative architecture of a lightweight local model and a large cloud model, a balance between real - time response and high accuracy is achieved. The local model is responsible for immediate response, while the large cloud model asynchronously optimizes the results, thereby improving the error - correction quality while ensuring the response speed. The contact dynamic feature perception network is used to capture the physical features of the user's keystrokes and construct an input intention probability matrix. Compared with the traditional method that only relies on static key position information, it can more accurately understand the user's true input intention, especially in scenarios of fast typing and one - hand operation, with significant effects. Through the recursive knowledge distillation self - correction mechanism, the knowledge of the large cloud model is continuously compressed into the local lightweight model, enabling the local model to continuously self - optimize, reducing the dependence on a large amount of labeled data, and achieving continuous improvement of system performance over time. Based on the collection of user feedback signals, a personalized tuning network is established to perform fine - grained adjustment according to different users' input habits and application scenarios, improving the error - correction acceptance rate and enhancing the user experience, especially in terms of specific - field technical terms and personal idiomatic expressions. The local lightweight model adopts a lightweight mixture - of - experts layer, a sparse activation mechanism, and a multi - level attention structure, which improves the model's expressive ability while maintaining low computational complexity, and achieves efficient error recognition and correction. By constructing an anti - interference communication pipeline and a differential response scheduling strategy, the system can operate stably under various network conditions and device states, improving the adaptability of the system in weak - network environments and low - resource devices. Based on the user's historical input patterns, a keyboard input prediction tree is constructed to trigger potential error - correction calculations in advance, combined with an oblivious replacement mechanism, effectively reducing the user - perceived latency and providing a smoother input experience.

[0031] In a specific embodiment, the process of executing step 100 may specifically include the following steps: Collect the key - press pressure data through the embedded sensor array in the user device to obtain the contact pressure distribution data; Track the movement coordinates and corresponding timestamps of the user's finger on the touch screen to obtain the sliding trajectory data; Record the contact start time and release time of the user's keystrokes to obtain the keystroke time - series feature data; Perform intention mapping calculations based on the contact pressure distribution data, time - series feature data, and sliding trajectory data to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input.

[0032] Specifically, a multimodal input behavior capture and modeling mechanism is constructed. This mechanism is based on an embedded sensor array as the underlying support, combines multi-channel data streams and integrates information in three dimensions: time, space and path, so as to fully restore the user's actual operation behavior at the signal source. When the user performs keyboard input operations, the pressure-sensitive array sensor installed under the soft key surface or touch surface is used to collect the physical contact pressure changes between the fingertips and the key surface in real time with a millisecond sampling period, where each key press corresponds to a two-dimensional pressure response image, which numerically reflects the distribution of unit pressure values ​​at different spatial positions. Due to the presence of a certain oblique angle and local eccentricity in fingertip pressure, the pressure map presents an asymmetric Gaussian shape, in which the pressure peak in the central area represents the actual maximum force point, and the gradient structure formed by its decreasing pressure toward the periphery reflects the finger's tilt trend and micro-motion trajectory. At the same time, the sliding trajectory of the finger on the touch area is continuously tracked, relying on a high-resolution capacitive touch positioning module, which detects the continuous movement of the center of gravity of the finger charge distribution through a matrix-distributed sensing unit, thereby obtaining a set of coordinate points under multiple continuous time stamps. Each element in the sequence contains both spatial position and time information, which can effectively describe the sliding path and rhythm of the user before, during or after triggering the key. By calculating the velocity and acceleration of the first-order derivative and the second-order derivative of the sequence, the characteristics of the finger's movement direction, trajectory turning point, sliding amplitude, and input inertia trend are restored, providing key decision-making basis in fuzzy input, accidental touch or non-vertical operation scenarios. While completing the pressure and trajectory information collection, the time dynamic characteristics of the key process are synchronously recorded. The event-driven mechanism is used to record the three key moments of each key operation: the starting moment of the key trigger (contact start time), the maximum pressure peak moment during the continuous contact process, and the release time when the fingertip leaves the key surface, and a time series composed of time nodes is constructed. The sequence is further expanded into multiple time intervals. These time parameters are used to describe the input rhythm, reaction time and pressing habits, and are jointly modeled with data from other dimensions to analyze whether the user has behavioral deviations such as rapid tapping, hesitant input or rhythm drift.The contact pressure distribution data, sliding trajectory data, and key-press time series feature data are uniformly input into a specially constructed intention mapping network (TDF-Net). This network adopts a three-stage structure design. In the first stage, a multi-channel convolutional network is used to extract features from the pressure image and the sliding trajectory image. The extracted local image features can be represented as a multi-dimensional vector tensor, retaining spatial position and gradient change information. In the second stage, bidirectional LSTM units are introduced to process time series data, including the pressure change rate between time points, the dynamic trend of the sliding trajectory, and the fingertip separation and combination frequency, etc., and bidirectional modeling is used to ensure that historical and future information can participate in the reasoning when determining the intention. In the third stage, all the aforementioned features are sent to the intention mapping layer for vector fusion and probability transformation, and finally a two-dimensional intention mapping matrix is output. This matrix performs probability distribution constraints on each row through softmax normalization to ensure that each input has a complete candidate character intention map. To improve the adaptability and generalization ability of this mapping process, a self-supervised learning mechanism is used to train the intention mapping model, where historical input behaviors and user manual modification results are used as pseudo-label signals to guide the parameter update of the intention discrimination model, enabling the system to self-optimize the recognition logic during continuous use. This intention probability matrix can reflect the ideal character intention under the standard input state and can provide highly robust intention correction support in the face of input noise, accidental touch disturbances, or complex input habits, serving as an important input basis for the subsequent error detection and character correction modules.

[0033] In a specific embodiment, the process of performing step of intention mapping calculation based on the contact pressure distribution data, time series feature data, and sliding trajectory data to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input may specifically include the following steps: Perform Gaussian distribution calculation on the contact pressure distribution data to obtain a key position heat map, where the central value of the heat map represents the key-press intention intensity; Calculate the user's finger movement vector and speed characteristics based on the sliding trajectory data to obtain a direction offset matrix; Calculate the time interval ratio between adjacent characters based on the key-press time series feature data to obtain a time correlation weight; Perform Bayesian probability fusion on the key position heat map, direction offset matrix, and time correlation weight to obtain an initial intention probability distribution; Combine the standard keyboard layout geometric distance matrix to perform spatial correlation weighted adjustment on the initial intention probability distribution to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input.

[0034] Specifically, for each keyboard press event, based on the two-dimensional contact pressure distribution data collected by the embedded pressure sensor array, an original pressure map is constructed. This map takes the key contact center as a reference point and organizes the pressure values recorded by the sensing units at each position in the horizontal and vertical directions into a density matrix. To enhance the structural interpretability and central directivity of this matrix, a two-dimensional Gaussian kernel function is constructed with the peak point of the contact pressure as the Gaussian center, and the entire pressure map is fitted to obtain a key position heat map with physical significance. The center value of the heat map represents the intensity of the key press intention, that is, the intensity of the user's concentrated input willingness at this key position, while the attenuation amplitude in its peripheral area reflects the degree of contact blur and the stability of finger touch. At the same time, the sliding trajectory data of the user's finger is obtained from the touch coordinate recording module, and a two-dimensional motion path is constructed through the position information sequence under continuous timestamps. Further, first-order difference and velocity vector calculations are performed on it to obtain a feature matrix reflecting the motion direction and trajectory amplitude. A vector is constructed for the position difference between every two consecutive moments, and the velocity modulus value and direction angle per unit time are calculated to form a direction offset matrix. This matrix is used to represent the path offset trend from the starting contact to the target key. The system identifies whether there are behaviors such as trajectory corners, inertial slips, or direction drifts to adjacent keys through this matrix. This phenomenon is particularly common in rapid tapping, multi-finger typing, or diagonal input. Analyze the key press time series feature data, that is, the contact start time, release time, and duration of each press action. By normalizing the time difference between different character input events, the time interval ratio between adjacent character pairs is calculated, and then a time correlation weight matrix is constructed. This weight is used to measure whether there are rhythmic mistakes during the user's rapid typing, such as key position deviation or finger lift delay caused by too fast input speed, and at the same time, to identify situations such as thinking interruptions caused by short pauses. The key position heat map, direction offset matrix, and time correlation weight matrix are jointly processed, and a Bayesian probability fusion mechanism is used to integrate the information in the three dimensions into a unified initial intention probability distribution. In the Bayesian framework, for each candidate character j given the pressed character i, its posterior probability P(j|i) as the user's true intention is expressed as a product combination of multiple conditional factors and normalization, that is, the system constructs a prior probability model and uses the pressure distribution, motion direction, and time dependence as conditional evidence for reasoning, so as to generate a normalized initial probability distribution among all possible characters. This distribution captures the physical manifestations and temporal dynamics of the user's input behavior and establishes an intention transfer relationship between characters in a statistical sense, enabling the system to still identify the user's true target character with a high confidence even under interference conditions such as accidental touches, fuzzy contacts, or rapid input. A standard keyboard layout geometric distance matrix is introduced to perform spatial correlation weighted adjustment on the initial probability distribution.The geometric distance matrix defines the Euclidean distance between every two characters in the keyboard. Each probability in the initial intention distribution is weighted according to the distance function. That is, for characters with a relatively short distance, their final probability remains the same or moderately increases, while for candidate characters with a long distance, their probability is compressed to an extremely low level to conform to the physical accessibility of user operations. This spatial weighting mechanism makes the intention probability matrix reasonable not only in the dynamic dimension of the input behavior but also constrained in the spatial structure, preventing the non-physical diffusion of incorrect intentions. By jointly modeling the concentration of the pressure distribution, the trend of the sliding trajectory, the continuity of the time rhythm, and the constraint of the geometric structure, an input intention probability matrix is generated.

[0035] In a specific embodiment, the process of executing step 200 may specifically include the following steps: The input character sequence is mapped through the character-level embedding layer of the first local lightweight model to obtain a character vector sequence; The character vector sequence is input into the lightweight mixture-of-experts layer of the first local lightweight model, and the input is assigned to 4 expert sub-networks through the routing network for parallel feature extraction to obtain a multi-perspective feature representation; A simplified sparse activation module is applied to selectively activate the multi-perspective feature representation, and only the key information paths are retained to obtain a sparse feature matrix; The sparse feature matrix and the input intention probability matrix are input into a lightweight multi-level attention structure for cross-modal information fusion to obtain a fused semantic representation. The lightweight multi-level attention structure includes 4 attention heads and 2 layers of Transformer structures; The fused semantic representation is input into the error detection branch and the correction branch of the dual-branch prediction head in the first local lightweight model for parallel calculation, and probability weighting is combined with a preset n-gram word frequency statistical table to obtain a first correction result.

[0036] Specifically, the input character sequence is mapped through the character-level embedding layer of the first local lightweight model. This embedding layer is equivalent to a compact semantic representation module that maps each discrete character into a set of dense vector representations of a fixed dimension based on its language features such as number, part of speech, usage frequency, and contextual performance. These vectors record the abstract positions of the characters in the language space and possess a certain degree of semantic similarity, enabling characters with similar usages or similar structures to be relatively close in the vector space, providing a basic condition for subsequent structure perception. The character vector sequence is input into the lightweight mixture-of-experts layer in the local model. Under the condition of limited computing resources, this module introduces an expert selection mechanism to assign different types of feature extraction tasks to different sub-networks to simulate a "call-on-demand" computing method. Among them, the routing network, as the discrimination center, assigns activation probabilities to each character position based on the context state and semantic deviation features of the current input sequence and divides these positions into 4 expert sub-networks. Each sub-network is designed as a feature processing unit with different structures. For example, some sub-networks are good at identifying local spelling patterns, some are good at judging language rhythm characteristics, and some are good at processing context connection relationships. The multi-path parallel processing method not only improves the depth of semantic extraction but also enhances the model's ability to model diverse error patterns. The results respectively output by the four expert sub-networks are concatenated and merged into a multi-perspective feature representation. A simplified sparse activation module is applied to selectively activate features in the multi-perspective feature representation to ensure that the system only retains the activation paths with higher information density and greater semantic value, thereby saving unnecessary computational overhead and enhancing the controllability of the prediction results. The sparse activation mechanism scores the activation values of each channel through a simplified gating structure and a dynamic channel scoring strategy, and only selects a part of the channels with the highest scores to output for the next-stage processing, while the remaining low-weight channels are suppressed or masked to obtain a sparse feature matrix. The sparse feature matrix is fused with the input intention probability matrix generated by an input behavior dynamic modeling module (such as TDF-Net). The input intention probability matrix is an offset intention estimate derived from the physical behavior expressions (including contact pressure, time distribution, sliding trajectory, etc.) during the user input process, providing another interpretation angle in the character sequence dimension, that is, the probability mapping relationship between the currently pressed character and the possible true intention character. To effectively couple the semantic features and intention features, a lightweight multi-level attention structure is introduced. The structure includes two Transformer encoding layers, and four attention heads are set in each layer, enabling it to build multi-dimensional information association relationships between the sparse features and the intention matrix. This structure concatenates the character features and the intention distribution as a composite input and uses the multi-head attention mechanism to respectively learn the context dependencies between characters, the coupling relationship between characters and intentions, the upstream and downstream trends between intentions, and the importance of the semantic-behavior coupling path.The attention layer strengthens the joint expression of key positions layer by layer by comparing the activation value distribution, the path attention map, and the dynamic weight backpropagation, and ensures gradient smoothness and model stability in the Transformer structure through residual connections and normalization layers, and finally outputs a fused semantic representation. The fused semantic representation is input into the error detection branch and the correction branch of the dual-branch prediction head in the first local lightweight model for parallel computing. Among them, the error detection branch uses a small convolutional module combined with a position attention mechanism to output the error probability for each character in the sequence bit by bit, and marks the positions higher than the set threshold as high-risk areas; while the correction branch constructs a probability distribution of alternative characters based on the semantic fusion representation and outputs a sorted result from all possible candidates. At this stage, to prevent the system from making judgments solely based on model training experience, a preset n-gram word frequency statistical table is introduced. This table records the probability relationships of common character combinations in the language. After the model makes a correction suggestion, the word frequency statistical results are called to probabilistically weight and correct the prediction results, so that the correction suggestions are not only reasonable and fluent, but also conform to conventional language usage habits and context fluency. Word frequency statistics provide a prior enhancement at the language level, enabling the local model to have higher language understanding and generation capabilities. Through the above steps, the first correction result is generated.

[0037] In a specific embodiment, the process of executing step 300 may specifically include the following steps: Transmit the first correction result, the input character sequence, and the input intent probability matrix to the cloud server to obtain a cloud input data set; Perform a merging process on the cloud input data set, the user's historical input sequence, and the application context information to obtain a target context data set; Input the target context data set into a cloud large model that adopts a mixture of experts architecture, a sparse activation mechanism, and a multi-level attention structure for processing to obtain deeply semantically associated data; Perform conditional probability maximization calculation based on the deeply semantically associated data to generate an optimal corrected text sequence, and return the optimal corrected text sequence to the user device as the second correction result.

[0038] Specifically, the first correction result is packaged together with the original input character sequence it is based on, and the corresponding input intention probability matrix is attached to form a basic data unit. To ensure transmission efficiency and data structure consistency, a unified data format encapsulation mechanism is adopted in the packaging stage. At the same time, differential compression and redundancy elimination algorithms are used to compress duplicate information, ensuring a relatively low data volume even under mobile network conditions, and it is sent to the cloud server through an encrypted channel to construct a cloud input dataset. The cloud input dataset is merged with the user's historical input sequences and application context information to construct a target context dataset. The user's historical input sequences are periodically uploaded to the cloud cache service by the local model in the idle state. Its content includes phrases with high recent usage frequency by the user, habitual word order structures, pragmatic features, and personalized misspelling correction records. These historical information play a role in context supplementation and semantic inertia guidance in the current input analysis; while the application context information comes from the specific usage scenario where the input occurs. Different language style weights and error correction tolerance parameters are defined according to each application type, so as to ensure targeted reasoning strategies in terms of language expression style, error correction strictness, and semantic coherence. After the above information integration is completed, a high-dimensional and multi-level target context dataset covering user behavior, language habits, scenario background, and input content is constructed. To effectively process inputs with complex structures and rich information dimensions, the large model deployed in the cloud adopts a mixture-of-experts architecture, which integrates multiple sub-models with different structures. Each sub-model is called an "expert" and is responsible for modeling specific types of language phenomena. For example, some experts are good at dealing with grammatical logical structures, some other experts pay more attention to word order combinations or sentence pattern transformations, and there are also some experts focusing on spelling changes or style features. After receiving the target context dataset, the model activates the sparse activation mechanism. According to the semantic distribution, context features, and target offset direction of the current input, the most suitable part is selected from all experts for activation and participation in the calculation, thus significantly reducing the redundant calculation cost and avoiding invalid parameter calls. After sparse activation, the multi-level attention structure inside the model starts to run, and complex dependencies between each character, word, phrase, and even context paragraph in the input are captured through a layer-by-layer attention weighting mechanism. In this process, the model identifies whether there are problems such as context incoherence, unreasonable grammatical structure, and semantic deviation in the input character sequence, and performs semantic alignment adjustment on the potential contradictions between the first correction result and the input intention, so as to construct deep semantic association data with strong semantic consistency and tight context coupling in the high-dimensional representation space. Based on the semantic association data, a conditional probability maximization inference task is performed. The model re-scores each character in each position, generates a set of candidate words according to the current semantic environment and language rules, and sequentially selects the path with the highest probability at each step to generate the final text sequence.To improve the overall naturalness and context consistency of the text, a beam search algorithm is introduced to optimize the path, avoiding the chain deviation caused by local optimality. At the same time, the language model is combined to globally evaluate the syntactic smoothness of the whole sentence, and the optimal corrected text sequence is generated. The optimal corrected text sequence is packaged as the second correction result and returned to the client, and confidence information and local character replacement identifiers are attached during the return process to guide whether the local model performs display updates, annotation prompts, or imperceptible replacements.

[0039] In a specific embodiment, the process of executing step 400 may specifically include the following steps: Use a pseudo-label generator to pair the second correction result and the input character sequence to obtain a pseudo-label training sample pair; Calculate the correction consistency, prediction confidence, and uniqueness of the pseudo-label training sample pair through a quality evaluation function, and generate a sample quality score set based on the correction consistency, prediction confidence, and uniqueness; Based on the sample quality score set, perform threshold screening on the pseudo-label training sample pair, and add the samples in the pseudo-label training sample pair with sample quality scores higher than the preset value to the recursive training cache to obtain a target distillation training set; Input the target distillation training set into the dynamic weighted distillation module, and use the combined distillation loss function to train the first local lightweight model to obtain a knowledge compression parameter set. The first local lightweight model includes a character-level embedding layer, a lightweight mixture of experts layer, a simplified sparse activation module, a lightweight multi-level attention structure, and a dual-branch prediction head; Update the parameters and optimize the structure of the first local lightweight model according to the knowledge compression parameter set, and at the same time establish a reverse knowledge flow channel to summarize the new error patterns to the server to obtain a second local lightweight model.

[0040] Specifically, the pseudo-label generator is used to pair the second correction result and the input character sequence and construct them into training sample pairs. This pairing process ensures that each original input sequence forms a pair of positive and negative sample mapping relationships with its corresponding high-confidence output from the cloud, and serves as the pseudo-supervised training data in future distillation learning, becoming the reference basis for the local model to imitate the cloud decision-making mechanism. To avoid interference from low-quality pseudo-labels in the local model learning, after the sample pairs are generated, a quality assessment module is launched to perform multi-dimensional quality analysis and evaluation on each group of samples. The difference degree between the second correction result from the cloud and the first correction result from the local is evaluated through the correction consistency index. This index is used to judge whether they are highly consistent or have differences in terms of character position, correction path, and result text, and accordingly, a consistency weight is assigned. Then, the prediction confidence of each pseudo-label sample is calculated. This confidence comes from the character probability distribution density or language model score output by the cloud model during the generation of this correction result, and is a key parameter to measure the stability and credibility of this result. Then, the uniqueness index is introduced. By comparing whether there are similar structures or overlapping contents in this sample and the historical training data, it is judged whether it has novelty, generalization value, and model challenge, so as to evaluate its training value. The scores of the above three dimensions are fused to construct a sample quality score set, and each pseudo-label sample pair is assigned a composite score reflecting its training reference value. A dynamically updated threshold is set to screen high-quality pseudo-label training samples. Only the sample pairs with scores higher than this threshold are retained and put into the recursive training cache to construct the target distillation training set for local distillation optimization. This training set has strong consistency, structural innovation, and high confidence, and effectively avoids the training pollution risk caused by wrong reasoning, boundary interference, or data anomalies through the sample screening mechanism, ensuring the stability and improvement amplitude of the subsequent local model distillation training. The target distillation training set is input into the dynamic weighted distillation module of the local model. This module is the execution unit for compressing and migrating cloud knowledge to the local model. During the distillation process, a combined distillation loss mechanism is introduced, that is, both the soft loss of the local model in aligning with the cloud prediction distribution and the hard loss of its alignment degree with the pseudo-label output are considered. The former is reflected in the local model imitating the prediction probability distribution of the cloud on each candidate character, and the latter is reflected in its ability to consistently output the pseudo-label result text. The weight ratio of the two losses is automatically adjusted through a time-related function. When the model has not converged, the distribution alignment training is weighted more to improve stability. After the model is gradually fitted, the pseudo-label alignment weight is increased to improve accuracy and generalization ability. The dynamic weighting mechanism ensures that the model has different training objectives at different stages, maximizing the training value of each batch of data.Meanwhile, the training scope of the local model is not limited to the output layer, but covers all its key structural units, including the character-level embedding layer, lightweight mixture of experts layer, simplified sparse activation module, lightweight multi-level attention structure, and dual-branch prediction head. The parameters of each layer participate in the training process through the backpropagation mechanism and continuously optimize their parameter states under the drive of the combined loss function. During the training process, the system monitors the parameter convergence speed and gradient change trend of each module, dynamically freezes the converged paths to avoid redundant calculations. At the same time, after the model weights are updated, the structural configuration is fine-tuned according to the convergence effect, such as pruning low-contribution expert sub-networks, reconstructing the sparse activation gating order, or merging attention paths, to optimize the model structure without increasing the parameter scale, forming a second local lightweight model with stable structure, fast response, and enhanced performance. To further improve the co-evolution mechanism of the system, a reverse knowledge flow channel is constructed during the distillation training, and the high-value new error samples, new input patterns, or rare misspelled word structures identified during the distillation process are uploaded to the server anonymously. The server conducts aggregated analysis and incremental training patch updates for the large model. This mechanism ensures that while the local model is continuously optimized, it also indirectly promotes the enrichment of the cloud model's corpus and the improvement of error correction accuracy, thus forming a positive feedback closed loop of end-cloud collaborative knowledge circulation and enhancing the response ability and long-term evolution potential of the entire error correction system to new language changes, rare usages, and user-specific input habits. Through the above steps, the iterative evolution process of the first local lightweight model to the second local lightweight model is completed.

[0041] In a specific embodiment, the process of executing step 500 may specifically include the following steps: Monitor and collect the user's selection operation of accepting or rejecting the correction suggestion, the corrected editing behavior, response time, and acceptance rate in the application scenario to obtain behavior feedback data; Perform Bayesian weight optimization calculation on the behavior feedback data to obtain an error correction satisfaction score; Construct personalized tuning data including the user's commonly used vocabulary, error type distribution, and scenario preferences based on the error correction satisfaction score; Monitor and calculate the device processor load, memory occupancy, battery power, and network latency, and generate a task allocation plan in combination with the personalized tuning data; Adjust the configuration of the second local lightweight model and the cloud large model according to the task allocation plan to obtain a two-layer asynchronous collaborative error correction strategy.

[0042] Specifically, each time the user receives an error correction suggestion, the behavior feedback collection module is activated to monitor their subsequent operations in real time. Four types of key feedback signals are recorded during this process: First, whether the user accepts or rejects the error correction suggestion provided by the system. If the user clicks "Accept" or actively retains the system suggestion, it is recorded as positive feedback. If the user clicks "Cancel" or continues to modify it to other characters, it is recorded as negative feedback. Second, whether the user continues to edit the text after accepting or rejecting the correction. The system determines whether the correction suggestion fully meets the user's expression needs based on whether the editing action covers the corrected area. Third, the system records the response time experienced by the user from the pop-up of the error correction suggestion to the final confirmation, and infers the user's judgment confidence in the corrected content by analyzing the length of this time. Fourth, the system counts the acceptance rate of error correction suggestions in the same type of scenario according to the current application scenario, such as chatting, note-taking, email, web input, etc., to reflect the user's tolerance and dependence on system intervention in different usage contexts. After the above feedback data collection is completed, it is input into the Bayesian weight optimization model for multi-dimensional weight normalization and satisfaction inference processing. The role of the Bayesian optimization model at this stage is to weight and fuse different types of feedback signals to construct a comprehensive error correction satisfaction scoring model. Different signals are scored by setting feedback priorities. For example, the acceptance or rejection operation has the highest weight, followed by the editing behavior, then the response time, and the scenario acceptance rate provides a long-term trend reference. Inside the model, the marginal probability of each type of signal is continuously updated according to the prior distribution of historical feedback, and the weighting strategy is dynamically updated according to the posterior inference of the new input data, outputting an error correction satisfaction score between 0 and 1, which is used to quantify the effectiveness of the current model output in the user's subjective perception. Based on the error correction satisfaction score, a personalized parameter adjustment data model is constructed. This model consists of three parts: the set of the user's frequently used words. By analyzing the words, phrases, and proper nouns with higher frequencies in the historical input records and error correction feedback, a user preference word vector is constructed to affect future word frequency weighting and candidate priority ranking; the user-specific error type distribution matrix. This matrix is based on the frequency statistics of error types that appear in the user's error correction history, such as spelling errors, semantic errors, word order misplacement, format errors, etc., to form a preferred error correction pattern. This information can be used to predict high-risk input structures in advance; the preference vector of the user's acceptance of error correction suggestions in different application scenarios. This vector records whether the user tends to accept system modifications or retain the original expression in a specific scenario, such as being inclined to retain Internet slang in social applications and being inclined to correct formal grammar in document writing. The combination of the three constitutes a user-specific error correction behavior portrait. The device resource monitoring module is activated to perform real-time sampling of the current processor load, memory occupancy, battery power, and network latency.Represent various resource metrics in the form of a state vector, and jointly model them with the aforementioned personalized tuning data to establish a task allocation function under multi-objective constraints. Considering multiple objectives such as performance bottlenecks, power consumption control, and user experience, generate an adaptive error correction task allocation scheme under different conditions. For example, when the processor occupancy is low and the battery is fully charged, the system performs more character-level error correction inference tasks on the local model; when the network condition is good but the local resources are tense, the system preferentially sends complex semantic judgment tasks to the cloud for processing; in a weak network environment but the user is in a high-response scenario (such as real-time conference chat), the system activates a conservative strategy, performs high-confidence fast error correction locally, and the cloud asynchronously compensates for semantic optimization to ensure double guarantees of real-time performance and language accuracy. Configure and adjust the second local lightweight model and the cloud large model according to the task allocation scheme to construct a two-layer asynchronous collaborative error correction strategy. Under this strategy, limit or expand the processing scope, execution frequency, and memory scheduling of the local model, and adapt to the current device state through dynamic channel control and quantization precision switching within the model structure; at the same time, adjust the call rhythm, request content compression ratio, and batch processing merge window of the cloud large model to ensure the optimal balance between the latency and semantic value of the cloud return results. In the execution of the overall strategy, introduce an asynchronous coordination mechanism based on timestamps, so that the preliminary error correction results of the local model are immediately displayed without affecting the input fluency, while the final optimized results returned by the cloud are replaced in the background to ensure that users always experience a continuous and high-quality input process, realizing end-cloud collaborative dynamic evolution and resource optimization response driven by user feedback.

[0043] In a specific embodiment, the process of performing step S103 of configuring and adjusting the second local lightweight model and the cloud large model according to the task allocation scheme to obtain a two-layer asynchronous collaborative error correction strategy may specifically include the following steps: Calculate the semantic complexity of the input character sequence in real time according to the task allocation scheme, and use an adaptive threshold function to allocate simple spelling errors to the second local lightweight model for processing, and allocate semantic dependency errors to the cloud large model for processing to obtain a binary error type splitting mechanism; Dynamically adjust the computing resource allocation of the second local lightweight model based on the binary error type splitting mechanism, reserve computing resources for keyboard input in high-frequency usage areas, and at the same time execute an upper limit control on the call frequency of the cloud large model to obtain a differential response scheduling strategy; Perform N-gram analysis based on the user's historical input pattern, construct a keyboard input prediction tree, and trigger potential error correction calculations in advance according to the keyboard input prediction tree to obtain a pre-judged error correction cache; Based on the pre-judged error correction cache and the differential response scheduling strategy, use the token bucket algorithm to control the smoothness of sending requests to the cloud large model, and implement network condition-aware request batch merging to obtain an anti-interference communication pipeline; Combine an anti-interference communication pipeline with an input delay sensitivity function to obtain results from a second local lightweight model, while allowing the results of the cloud large model to be returned asynchronously and replaced imperceptibly, forming a two-layer asynchronous collaborative error correction strategy.

[0044] Specifically, a lightweight semantic analysis engine is used to perform dynamic semantic complexity modeling on the current input character sequence. This model constructs a dynamic semantic complexity score by analyzing the part-of-speech distribution, syntactic structure, context dependence degree, and phrase-level collocation rules in the current input, combined with the lexical adhesion degree, semantic ambiguity, and ambiguity intensity between characters. On this basis, an adjustable adaptive threshold function is introduced. This function automatically adjusts the discrimination threshold according to multiple variables such as device load, user feedback score, input scenario type, and historical semantic error rate, and accordingly classifies input errors into two categories: If the current input is judged to have obvious character offsets, spelling similarity errors, or simple errors at the independent word level, the system will assign it to the second local lightweight model for processing; conversely, if the input involves context dependence, structural nesting, semantic ambiguity, or multi-sentence spanning features, the error will be marked as a semantic complex error and sent to the cloud large model for processing, forming a "spelling-semantic" binary error type shunt mechanism. Based on the above shunt judgment results, the computing resource allocation strategy of the second local lightweight model is dynamically optimized and adjusted. To ensure the keyboard input response speed in high-frequency usage areas, a heat map is constructed based on the user's keyboard operation behavior trajectory, and combined with the user input frequency, key hit rate, and mis-touch record, local priority scheduling is performed on the local model memory cache and computing core, reserving additional model computing paths and parameter spaces for the commonly used key areas and high-risk character areas, so as to achieve fast response and instant correction. For the cloud model, to prevent bandwidth congestion and soaring latency caused by large-scale complex error upload requests, a cloud call frequency upper limit control strategy is introduced. The call frequency upper limit is set according to factors such as real-time network bandwidth, historical average response time, server processing load, and user acceptable waiting duration, avoiding affecting the overall error correction fluency of the system due to frequent triggering of semantic requests, thus constituting a differential response scheduling strategy for end-cloud division of labor. After completing the basic shunt and resource allocation, an active prediction ability is constructed to improve the overall response preposition and error correction predictability. N-gram analysis is performed based on the user's historical input records to construct a user-specific keyboard input prediction tree. This prediction tree uses characters, phrases, and semantic fragments as nodes, and time sequence, input habits, and language transition probabilities as edge weights, and is used to deduce the character sequence that the user may input next and its corresponding potential error patterns. The system pre-expands the high-probability paths through this prediction tree, activates the inference logic of the local lightweight model in advance, pre-generates a list of error correction candidates for some high-risk character fragments, forms a prediction error correction cache, and outputs error correction suggestions without delay when the user's real input matches the predicted path, thus greatly shortening the response time and reducing the backend inference pressure. To effectively coordinate the data interaction and request control between the local and the cloud, a cloud request scheduling module based on the token bucket algorithm is introduced. This module smoothly controls the speed of sending error correction requests to the cloud by generating a fixed number of tokens per unit time, avoiding cloud response blocking caused by high-concurrency requests in a short time.Meanwhile, this scheduling mechanism is linked with the network condition awareness module. When the network latency increases or the link jitter is severe, the token generation rate is automatically reduced, and the request batch merging strategy is enabled. Multiple adjacent or related error correction requests are merged into a single request packet for transmission to maximize bandwidth utilization and minimize network overhead. This anti-interference communication pipeline mechanism not only optimizes the bandwidth pressure of the request path but also enhances the stable response ability in complex scenarios, and is applicable to the collaborative error correction system operation in mobile scenarios, weak network environments, or conditions where multiple users use simultaneously. With the collaborative cooperation of all the above components, the anti-interference communication pipeline is linked and integrated with the input delay sensitivity function, which calculates the tolerance range of each input action for response delay based on dimensions such as input type, usage scenario, user input speed, and task importance. When it is detected that the input is a high-real-time scenario such as an online conversation, a real-time note, or a quick question-and-answer system, the system forces the local model result to be the preferred output path, and the cloud result is processed asynchronously. After the cloud result is returned, it automatically compares the accuracy of the local result and replaces the displayed content imperceptibly if necessary to ensure the continuity of the user experience and the immediacy of the response. In low-interaction-density or static task scenarios such as document writing, report writing, or email reply, it is allowed to wait moderately for the cloud semantic result before outputting the final error correction text, thus implementing the response quality priority strategy.

[0045] The above describes the AI-based automatic keyboard error correction method in the embodiments of the present application. Next, the AI-based automatic keyboard error correction system in the embodiments of the present application will be described. Please refer to Figure 2 , an embodiment of the AI-based automatic keyboard error correction system in the embodiments of the present application includes: A collection module 11, configured to collect the contact pressure distribution, sliding trajectory, and residence time of the user's keyboard input and construct an input intention probability matrix; A preliminary correction module 12, configured to input the input intention probability matrix into a first local lightweight model for error detection and preliminary correction to obtain a first correction result; A semantic analysis module 13, configured to transmit the first correction result and the input intention probability matrix to a cloud large model for in-depth semantic analysis to obtain a second correction result; A knowledge distillation module 14, configured to perform recursive knowledge distillation on the first local lightweight model according to the second correction result to obtain a second local lightweight model; An output module 15, configured to collect the behavior feedback data of the user's acceptance or rejection of the second correction result, and perform dynamic adjustment of task allocation on the second local lightweight model and the cloud large model based on the acceptance or rejection behavior feedback data, and output a two-layer asynchronous collaborative error correction strategy.

[0046] Through the collaborative cooperation of the above-mentioned various components, by establishing a two-layer asynchronous collaborative architecture of a lightweight local model and a large cloud model, a balance between real-time response and high accuracy is achieved. The local model is responsible for instant response, while the large cloud model asynchronously optimizes the results, thereby improving the error correction quality while ensuring the response speed. The physical characteristics of the user's keystrokes are captured by the contact dynamic feature perception network, and an input intention probability matrix is constructed. Compared with the traditional method that only relies on static key position information, it can more accurately understand the user's true input intention, especially in the scenarios of rapid typing and one-handed operation. Through the recursive knowledge distillation self-correction mechanism, the knowledge of the large cloud model is continuously compressed into the local lightweight model, enabling the local model to continuously self-optimize, reducing the dependence on a large amount of labeled data, and achieving continuous improvement of the system performance over time. Based on the collection of user feedback signals, a personalized tuning network is established to perform fine-grained adjustment according to the input habits and application scenarios of different users, improving the error correction acceptance rate and enhancing the user experience, especially in terms of specific domain technical terms and personal habitual expressions. The local lightweight model adopts a lightweight expert mixture layer, a sparse activation mechanism, and a multi-level attention structure, which improves the model expression ability while maintaining low computational complexity, and achieves efficient error recognition and correction. By constructing an anti-interference communication pipeline and a differential response scheduling strategy, the system can maintain stable operation under various network conditions and device states, improving the adaptability of the system in weak network environments and low-resource devices. Based on the user's historical input patterns, a keyboard input prediction tree is constructed to trigger potential error correction calculations in advance, combined with an oblivious replacement mechanism, effectively reducing the user's perceived latency and providing a smoother input experience.

[0047] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0048] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0049] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An AI-based automatic keyboard error correction method, characterized in that, Including: Collecting the contact pressure distribution, sliding trajectory, and dwell time of the user's keyboard input and constructing an input intention probability matrix; Inputting the input intention probability matrix into the first local lightweight model for error detection and preliminary correction to obtain a first correction result; Transmitting the first correction result and the input intention probability matrix to the cloud large model for in-depth semantic analysis to obtain a second correction result; Performing recursive knowledge distillation on the first local lightweight model according to the second correction result to obtain a second local lightweight model; Collecting the behavioral feedback data of the user's acceptance or rejection of the second correction result, and dynamically adjusting the task allocation of the second local lightweight model and the cloud large model based on the acceptance or rejection behavioral feedback data, and outputting a two-layer asynchronous collaborative error correction strategy.

2. The AI-based automatic keyboard error correction method according to claim 1, wherein The collecting the contact pressure distribution, sliding trajectory, and dwell time of the user's keyboard input and constructing an input intention probability matrix includes: Collecting key pressure data through the embedded sensor array in the user device to obtain contact pressure distribution data; Tracking the moving coordinates and corresponding timestamps of the user's finger on the touch screen to obtain sliding trajectory data; Recording the contact start time and release time of the user's key press to obtain key time series feature data; Performing intention mapping calculation based on the contact pressure distribution data, the time series feature data, and the sliding trajectory data to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input.

3. The AI-based automatic keyboard error correction method according to claim 2, characterized in that, The performing intention mapping calculation based on the contact pressure distribution data, the time series feature data, and the sliding trajectory data to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input includes: Performing Gaussian distribution calculation on the contact pressure distribution data to obtain a key position heat map, where the central value of the heat map represents the key press intention intensity; Calculating the user finger movement vector and speed feature according to the sliding trajectory data to obtain a direction offset matrix; Calculating the time interval ratio between adjacent characters based on the key time series feature data to obtain a time correlation weight; Performing Bayesian probability fusion on the key position heat map, the direction offset matrix, and the time correlation weight to obtain an initial intention probability distribution; Performing spatial correlation weighted adjustment on the initial intention probability distribution in combination with the standard keyboard layout geometric distance matrix to obtain an input intention probability matrix representing the correspondence between the characters pressed by the user and the characters actually intended to be input.

4. The AI-based automatic keyboard error correction method according to claim 1, wherein The inputting the input intention probability matrix into the first local lightweight model for error detection and preliminary correction to obtain a first correction result includes: Performing mapping processing on the input character sequence through the character-level embedding layer of the first local lightweight model to obtain a character vector sequence; Inputting the character vector sequence into the lightweight expert mixture layer of the first local lightweight model, and distributing the input to 4 expert subnets through the routing network for parallel feature extraction to obtain a multi-perspective feature representation; Selectively activate the features of the multi-view feature representation using a simplified sparse activation module, and only retain the key information path to obtain a sparse feature matrix; Input the sparse feature matrix and the input intent probability matrix into a lightweight multi-level attention structure for cross-modal information fusion to obtain a fused semantic representation. The lightweight multi-level attention structure includes 4 attention heads and 2 layers of Transformer structures; Input the fused semantic representation into the error detection branch and the correction branch of the dual-branch prediction head in the first local lightweight model for parallel computing, and perform probability weighting in combination with a preset n-gram word frequency statistical table to obtain a first correction result.

5. The AI-based automatic keyboard error correction method according to claim 4, characterized in that, The transmitting the first correction result and the input intent probability matrix to a cloud large model for in-depth semantic analysis to obtain a second correction result includes: Transmit the first correction result, the input character sequence, and the input intent probability matrix to a cloud server to obtain a cloud input data set; Merge and process the cloud input data set with the user's historical input sequence and application context information to obtain a target context data set; Input the target context data set into a cloud large model that adopts a mixture of experts architecture, a sparse activation mechanism, and a multi-level attention structure for processing to obtain in-depth semantic association data; Perform conditional probability maximization calculation based on the in-depth semantic association data, generate an optimal corrected text sequence, and return the optimal corrected text sequence as the second correction result to the user device.

6. The AI-based automatic keyboard error correction method according to claim 1, wherein The recursively distilling knowledge from the first local lightweight model according to the second correction result to obtain a second local lightweight model includes: Use a pseudo-label generator to pair the second correction result and the input character sequence to obtain a pseudo-label training sample pair; Calculate the correction consistency, prediction confidence, and uniqueness of the pseudo-label training sample pair through a quality evaluation function, and generate a sample quality score set based on the correction consistency, the prediction confidence, and the uniqueness; Perform threshold screening on the pseudo-label training sample pair based on the sample quality score set, and add the samples in the pseudo-label training sample pair with a sample quality score higher than a preset value to the recursive training cache to obtain a target distillation training set; Input the target distillation training set into a dynamic weighted distillation module, and use a combined distillation loss function to train the first local lightweight model to obtain a knowledge compression parameter set. The first local lightweight model includes a character-level embedding layer, a lightweight mixture of experts layer, a simplified sparse activation module, a lightweight multi-level attention structure, and a dual-branch prediction head; Update the parameters and optimize the structure of the first local lightweight model according to the knowledge compression parameter set, and at the same time establish a reverse knowledge flow channel to summarize new error patterns to the server to obtain a second local lightweight model.

7. The AI-based automatic keyboard error correction method according to claim 1, wherein Collect the behavioral feedback data on the user's acceptance or rejection of the second correction result, and dynamically adjust the task allocation of the second local lightweight model and the cloud large model based on the acceptance or rejection behavioral feedback data, and output a two-layer asynchronous collaborative error correction strategy, including: Monitor and collect the user's selection operation of accepting or rejecting the correction suggestion, the corrected editing behavior, the response time, and the acceptance rate in the application scenario to obtain behavioral feedback data; Perform Bayesian weight optimization calculation on the behavioral feedback data to obtain an error correction satisfaction score; Construct personalized tuning data including user's commonly used words, error type distribution, and scenario preferences based on the error correction satisfaction score; Monitor and calculate the device processor load, memory occupancy, battery power, and network latency, and generate a task allocation plan in combination with the personalized tuning data; Adjust the configuration of the second local lightweight model and the cloud large model according to the task allocation plan to obtain a two-layer asynchronous collaborative error correction strategy.

8. The AI-based automatic keyboard error correction method according to claim 7, wherein The adjusting the configuration of the second local lightweight model and the cloud large model according to the task allocation plan to obtain a two-layer asynchronous collaborative error correction strategy includes: Real-time calculate the semantic complexity of the input character sequence according to the task allocation plan, and use an adaptive threshold function to assign simple spelling errors to the second local lightweight model for processing, and assign semantic dependency errors to the cloud large model for processing to obtain a binary error type diversion mechanism; Dynamically adjust the computing resource allocation of the second local lightweight model based on the binary error type diversion mechanism, reserve computing resources for keyboard input in high-frequency usage areas, and at the same time execute an upper limit control on the call frequency of the cloud large model to obtain a differential response scheduling strategy; Perform N-gram analysis based on the user's historical input pattern, construct a keyboard input prediction tree, and trigger potential error correction calculations in advance according to the keyboard input prediction tree to obtain a pre-judged error correction cache; Based on the pre-judged error correction cache and the differential response scheduling strategy, use the token bucket algorithm to control the smoothness of sending requests to the cloud large model, and implement network condition-aware request batch merging to obtain an anti-interference communication pipeline; Combine the anti-interference communication pipeline with the input delay sensitivity function, obtain the result from the second local lightweight model, and allow the result of the cloud large model to be returned asynchronously and replaced without perception to form a two-layer asynchronous collaborative error correction strategy.

9. An AI-based automatic keyboard error correction system, characterized in that, For executing the AI-based keyboard error automatic correction method according to any one of claims 1-8, the AI-based keyboard error automatic correction system includes: A collection module for collecting the contact pressure distribution, sliding trajectory, and residence time of the user's keyboard input and constructing an input intention probability matrix; A preliminary correction module for inputting the input intention probability matrix into a first local lightweight model for error detection and preliminary correction to obtain a first correction result; A semantic analysis module for transmitting the first correction result and the input intention probability matrix to a cloud large model for in-depth semantic analysis to obtain a second correction result; A knowledge distillation module, configured to perform recursive knowledge distillation on the first local lightweight model according to the second correction result, so as to obtain a second local lightweight model; An output module, configured to collect behavioral feedback data on the acceptance or rejection of the second correction result by a user, and dynamically adjust the task allocation of the second local lightweight model and the cloud large model based on the acceptance or rejection behavioral feedback data, and output a two-layer asynchronous collaborative error correction strategy.

Citation Information

Patent Citations

  • Input error correcting method, input error correcting device, automatic error correcting method, automatic error correcting device and mobile terminal

    CN103425412A

  • Key mistaken touch error correction method and device

    CN112015279A

  • Interaction system and method for AI intelligent mouse

    CN117784959A

  • Automatic language selection for text input in messaging context

    EP1727024A1

  • Neural query auto-correction and completion

    US11106690B1

Cited By

  • Integrated metal key matrix anti-interference method, system, equipment and medium

    CN120909161A