An intelligent text correction method based on mouse interaction
By using a mouse-interactive intelligent text correction method, which generates candidate terms by fusing AI models and multi-source data, the problem of low efficiency in intelligent text recognition in existing technologies is solved, and efficient and accurate error correction is achieved.
Patent Information
- Application Number
- CN202510800451.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2026-03-17
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing intelligent speech recognition and image recognition models often make mistakes when recognizing text, causing users to frequently use mouse and keyboard operations, which reduces text input efficiency.
By using a mouse-interactive intelligent text correction method, an AI recognition model is used to identify erroneous text, and a correction pop-up window that follows the cursor appears. Users can select the correct term with the mouse to correct the error. Candidate terms are generated by multi-source data fusion, and the selection cost is reduced by using pinyin index and radical index.
It enables the error correction process to be completed solely through mouse operations, significantly improving the efficiency of text recognition error correction. It supports cross-platform compatibility and multiple application scenarios, ensuring the accuracy and efficiency of error correction.
Smart Images

Figure CN120633637B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent text correction technology, specifically to an intelligent text correction method based on mouse interaction. Background Technology
[0002] With the development of artificial intelligence, AI tools such as speech recognition and image recognition have made great strides and are gradually becoming more widely used.
[0003] In practical applications, speech recognition (ASR) models or image recognition (OCR) models often exhibit text recognition errors. For example, when using a smart voice mouse for voice input, incorrect text is frequently recognized. Users must perform numerous mouse and keyboard operations, requiring frequent alternation between these actions, to correct these errors. This significantly reduces the efficiency of text input.
[0004] Therefore, there is an urgent need for an intelligent text correction method based on mouse interaction, which can achieve efficient correction of erroneous text by using mouse interaction and intelligent models. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent text correction method based on mouse interaction, which solves the problem of low efficiency in correcting text errors identified by intelligent recognition in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent text correction method based on mouse interaction, comprising the following steps:
[0007] S1. The AI recognition model provides the results of the recognized text and displays them in the input box in the APP;
[0008] S2. Determine if there are errors in the recognition results in the APP's input box;
[0009] If the text in the recognition result contains no errors, the error correction process ends directly.
[0010] S3. Use the mouse button to activate the text correction mode and start the text correction program;
[0011] S4. A pop-up error correction window will appear and follow the cursor movement.
[0012] S5. Move the cursor near an incorrect text using the mouse;
[0013] S6. Obtain information such as the APP handle, input box content, cursor context, and cursor position;
[0014] S7. Based on the obtained cursor context and cursor position, determine the current erroneous term and generate confusion candidate terms;
[0015] S8. Select N optimal obfuscation candidate terms from the obfuscation candidate terms and display them in the error correction pop-up window;
[0016] S9. Select the correct term from the error correction pop-up window using the mouse;
[0017] S10: Replace the incorrect text in the APP with the selected correct words, and exit the text correction mode by pressing the button on the mouse.
[0018] Preferably, the APP in step S1 includes, but is not limited to, APPs on Windows and macOS systems such as Word, WPS, Notepad, WeChat, and browsers.
[0019] Preferably, the AI recognition model in step S1 is one of the speech recognition ASR model and the image recognition OCR model.
[0020] Preferably, in step S5, when moving the cursor, the cursor is moved to one of the locations before or after the erroneous text.
[0021] Preferably, the method for determining the current erroneous term based on the cursor position and context content in step S7 includes the following steps:
[0022] S7.1 Provide the cursor position and context content to the AI model at the same time. The AI model will provide the offset of the erroneous term relative to the current cursor position based on this information.
[0023] S7.2, Provide all possible candidate terms for confusion;
[0024] S7.3 merges and removes duplicates from the given candidate terms for obfuscation to obtain a list of candidate terms for obfuscation of erroneous terms.
[0025] Preferably, the AI large model includes, but is not limited to, one of the Qwen large model and the DeepSeek large model.
[0026] Preferably, the candidate words for confusion given in step S7.2 include, but are not limited to, the speech recognition results after extracting the wrong words from historical speech localization, the confusion dictionary of the wrong text, the homonym dictionary of the wrong text, and the confusion words given by the AI big model based on the wrong words and all the content in the entire input box.
[0027] Preferably, in step S8, all candidate obfuscated terms, erroneous terms, and contextual content are simultaneously input into the AI big model. The AI big model sorts the terms in the obfuscated term list from highest to lowest probability of replacement and displays the top N terms in the error correction window.
[0028] Preferably, when the correct term is not displayed in the error correction pop-up window in step S8, all the confusing candidate terms are classified according to the set filtering mechanism and displayed in the error correction pop-up window. Then, the filtering conditions are operated by moving the mouse in the y direction of the error correction pop-up window, and the confusing candidate terms under the filtering conditions are operated in the x direction of the error correction pop-up window to select the correct term.
[0029] Preferably, the filtering mechanism is one of pinyin or radical.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. The present invention relates to an intelligent text correction method based on mouse interaction, which can complete the entire correction process with only mouse operation, avoiding keyboard switching and greatly improving the efficiency of text recognition and correction.
[0032] 2. The present invention relates to an intelligent text correction method based on mouse interaction, which employs intelligent models and multi-source data fusion to ensure the accuracy of error identification and candidate word recommendation.
[0033] 3. The present invention relates to an intelligent text correction method based on mouse interaction that supports cross-platform (Windows, MacOS) and multiple application scenarios (Word, WPS, WeChat, browser, etc.); it adopts commonly used pinyin index and radical index to reduce the user's selection cost and can adapt to different correction scenarios. Attached Figure Description
[0034] Figure 1 This is a flowchart illustrating an intelligent text correction method based on mouse interaction according to the present invention.
[0035] Figure 2 This is a schematic diagram of an error correction pop-up window in an embodiment of an intelligent text error correction method based on mouse interaction according to the present invention. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0037] Example, refer to Figure 1-2 , An intelligent text error correction method based on mouse interaction, comprising the following steps:
[0038] S1. The recognition results of the ASR model are displayed in the APP. The APP includes, but is not limited to, APPs in Windows and MacOS systems such as Word, WPS, Notepad, WeChat, browsers, etc. <00000Merge these confusing candidate entries together and remove duplicates to obtain a comprehensive list of confusing candidate entries for the current incorrect entry. In this embodiment, it is "Mixc City, Wanxiangcheng, Wanxiang City, Wangxiangcheng, Play Thread, Play Formation...".
[0048] S8. Screen out N optimal confusing candidate entries from the confusing candidate entries and display them in the error correction pop-up window.
[0049] Here, by inputting all the confusing candidate entries, the incorrect entry, and the context content into the Qwen or DeepSeek large model at the same time, the large model sorts the confusing candidate entry list according to the decreasing possibility of replacement. Then, the top N entries after sorting are displayed in the error correction window. In this embodiment, it is to display "Mixc City, Wanxiang City, Wangxiangcheng".
[0050] S9. Click and select the correct entry from the error correction pop-up window using the mouse.
[0051] This program classifies all the confusing candidate entries according to the preset screening mechanism (pinyin or radicals) and displays them in the error correction pop-up window, as Figure 2 shown.
[0052] The user operates the error correction pop-up window with the mouse, operates the screening conditions in the y direction, and operates the confusing candidate entries under the screening conditions in the x direction to select the correct entry. [[ID=||]]
[0053] When the user clicks on the correct candidate entry, this is achieved by calling the system-level text message interface and keyboard key events. In this embodiment, it is to click on the entry "Mixc City".
[0054] S10. Replace the incorrect text in the APP with the selected correct entry.
[0055] S11. The user operates the mouse to move the cursor near the position of the next incorrect text, and executes from S4 to S10 until all errors have been corrected.
[0056] S12. Press the error correction mode button on the mouse to exit the error correction mode, and this error correction process ends.
[0057] [[ID=3||]]The system of this error correction method includes: an interaction module for receiving mouse trigger signals, controlling the error correction mode switch and cursor position tracking;
[0058] An information acquisition module that obtains the handle, text content, and cursor position through the operating system API (such as WM_GETTEXT and EM_GETSEL for Windows, and AXUIElement for MacOS).
[0059] The intelligent analysis module generates erroneous terms and offsets based on cursor context and an intelligent model, and integrates candidate terms from multiple sources.
[0060] The sorting and display module sorts candidate words and displays them in a pop-up window, supporting pinyin indexing and filtering.
[0061] The replacement module implements text replacement through system-level interfaces (such as EM_SETSEL, EM_REPLACESEL, or AXUIElementSetAttributeValue).
[0062] When the user presses the error correction mode button, the mouse sends a key press data packet to this program via 2.4G or Bluetooth protocol. This program then obtains the handle of the currently active window through the system interface.
[0063] In Windows, the GetForegroundWindow interface can be used with the following code:
[0064] HWND hwnd=GetForegroundWindow();
[0065] Use the following code in macOS:
[0066]
[0067]
[0068] Get all the text in the input boxes in the APP through the system-level API.
[0069] In Windows, using the WM_GETTEXT message and the SendMessage interface, an application with code like the following can retrieve all the text.
[0070]
[0071]
[0072] Get the input cursor position in the app via system-level API
[0073] In Windows systems, the EM_GETSEL message is obtained using the SendMessage interface, with an application code like the following.
[0074] HWND hwndEdit = / * target window handle * / ;
[0075] DWORD start,end;
[0076] SendMessage(hwndEdit,EM_GETSEL,(WPARAM)&start,(LPARAM)&end);
[0077] / / start == end indicates the insertion point; otherwise, it's a selection area.
[0078] In macOS, this can be obtained using Apple's accessibility framework AXUIElement, with an application like the following code.
[0079] AXUIElementRef systemWide=AXUIElementCreateSystemWide();
[0080] AXUIElementRef focusedApp=NULL;
[0081] AXUIElementRef focusedUIElement=NULL;
[0082] AXUIElementCopyAttributeValue(systemWide,
[0083] kAXFocusedApplicationAttribute,(CFTypeRef*)&focusedApp);
[0084] AXUIElementCopyAttributeValue(systemWide,
[0085] kAXFocusedUIElementAttribute,(CFTypeRef*)&focusedUIElement);
[0086] / / Get the row number or range of the insertion point
[0087] CFTypeRef value = NULL;
[0088] AXUIElementCopyAttributeValue(focusedUIElement,
[0089] CFSTR("AXSelectedTextRange"),&value);
[0090] Error information for calculating the input cursor position
[0091] Organize the text content from step 2 and the input cursor position from step 3 into a JSON format, as shown below.
[0092] {
[0093] "content": "Hefei Fantasy City is very popular".
[0094] “caretpos”: (3,3,0)
[0095] }
[0096] Add the following prompt words
[0097] You will receive a text message and a tuple `(start, end, y)` representing the position of the input cursor. The text may be from speech recognition and therefore may contain recognition errors.
[0098] Please complete the following tasks:
[0099] 1. Based on the cursor's position, analyze the characters before and after it to identify possible speech recognition errors (such as semantic inconsistencies, common sense errors, typos, etc.).
[0100] 2. Identify the incorrect characters (there may be one or more).
[0101] 3. Output the offset of these error characters relative to the cursor starting position (i.e., character subscript - cursor starting subscript).
[0102] Please return in the following JSON format:
[0103]
[0104] Input the JSON content and the hint words into the qwen / deepseek model, and the qwen / deepseek model will provide an answer, in the form of:
[0105] {
[0106] "wrong":"Dream City",
[0107] "offset": [-1, 0, 1]
[0108] }
[0109] Get all obfuscation candidates for the incorrect term
[0110] To cover as many obfuscation candidates as possible, this program uses the following sources of obfuscation candidates.
[0111] Speech recognition results after extracting erroneous words from historical speech localization
[0112] Error text obfuscation dictionary
[0113] Dictionary of words with the same pinyin as the erroneous text
[0114] The Qwen / DeepSeek large model generates obfuscated terms based on incorrect terms and all content in the input box.
[0115] The prompts for the Qwen / DeepSeek large model are as follows:
[0116] You are a speech recognition error analysis assistant. You have received a text file after speech recognition and a word that was incorrectly identified. Based on the context, similarity of pronunciation, and common confusion phenomena, please provide possible correct candidate words (confused words) for the incorrect word.
[0117] Example as follows:
[0118]
[0119] The results are as follows
[0120] {
[0121] "candidates":["Vientiane City","Wangxiang City","Wangxiang City"]
[0122] }
[0123] Calculate the ranking of all confusion candidates
[0124] All obfuscation candidates from the five sources are placed into a single list, and duplicates are removed from the list. This results in a list of obfuscation candidates, which is then sorted using the Qwen / DeepSeek large model.
[0125] The prompt words are as follows
[0126] You are a speech recognition post-processing assistant.
[0127] Now we have the following input:
[0128] content: Complete speech recognition text content
[0129] wrong_word: Identifies incorrect words
[0130] candidates: Multiple candidate terms provided by the user or system, which may be the original words of an incorrect term.
[0131] Please sort the entries in candidates according to the probability that they are the most likely correct entries, taking into account factors such as contextual semantics, similarity of pronunciation, fluency of language, and common collocations, and return the sorted list.
[0132] Example as follows:
[0133]
[0134] Output example is as follows
[0135] {
[0136] "sorted_candidates":["Vientiane City","Dream City","Prosperous City","Hopeful City","Delusional City"]
[0137] }
[0138] A pop-up window appears near the cursor position, displaying confusion candidates that can be indexed by pinyin.
[0139] The open-source tool pypinyin was used to annotate the pinyin of all obfuscation candidates in section 6, and the annotated pinyin results were used to index the obfuscation candidates in section 6. Based on the cursor position, a pop-up window near the cursor displays the candidates and their pinyin index. Furthermore, based on the different pinyin pronunciations, the corresponding obfuscation candidates are displayed.
[0140] After the user selects the correct term, the system interface is called to replace the incorrect term with the correct one.
[0141] In Windows, use code like the following:
[0142] / / Set a new selection area
[0143] SendMessage(hwndEdit,EM_SETSEL,start,end);
[0144] / / Replace with new string
[0145] SendMessage(hwndEdit,EM_REPLACESEL,TRUE,(LPARAM)L"New content");
[0146] In macOS, use code like the following
[0147] AXUIElementRefsystemWide=AXUIElementCreateSystemWide();
[0148] AXUIElementReffocusedElement=NULL;
[0149] / / Get the currently focused control (such as an input box)
[0150] AXUIElementCopyAttributeValue(systemWide,
[0151] kAXFocusedUIElementAttribute,(CFTypeRef*)&focusedElement);
[0152] / / Get the current cursor position (e.g., location=5, length=0)
[0153] CFTypeRefselectedRange=NULL;
[0154] AXUIElementCopyAttributeValue(focusedElement,
[0155] CFSTR("AXSelectedTextRange"),&selectedRange);
[0156] / / Assuming it's a CFRange or similar structure wrapped in NSValue.
[0157] CFRange range = ...; / / Parse the value of selectedRange
[0158] / / Expand forward by 1, expand backward by 1
[0159] range.location=MAX(0,range.location-1);
[0160] range.length=range.length+2;
[0161] / / Set a new selection area
[0162] AXUIElementSetAttributeValue(focusedElement,CFSTR("AXSelectedTextRange"),(__bridge CFTypeRef)([NSValue valueWithRange:range]));
[0163] / / Replace content (Note that AXValue corresponds to the entire text in the input box, but some controls support replacing the currently selected area)
[0164] AXUIElementSetAttributeValue(focusedElement,CFSTR("AXValue"),CFSTR("replacement content"));
[0165] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0166] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent text correction based on mouse interaction, characterized in that, The method comprises the following steps: S1, the AI recognition model gives the recognition result of the text, and displays it in the input box in the APP; S2, judge the error of the recognition result in the APP input box; Wherein, when the text in the recognition result has no error, the error correction program is directly ended; S3, the user presses the error correction mode button on the mouse to wake up the text error correction mode, and the program opens the error correction pop-up window; S4, the error correction pop-up window moves with the cursor; S5, move the cursor to the vicinity of an error text by mouse; S6, get the handle of APP, input box content, cursor context content, cursor position and other information; S7, according to the cursor context content and cursor position, judge the current error word and generate the confusion candidate word; Wherein, the method for judging the current error word by cursor position and context content comprises the following steps: S7.1, provide the cursor position and context content to the AI large model at the same time, and the AI large model gives the offset of the error word relative to the current cursor position according to these information; S7.2, give all possible confusion candidate words; S7.3, merge and deduplicate the given confusion candidate words to get the confusion candidate list of error words; S8, select N optimal confusion candidate words from the confusion candidate words, and display them in the error correction pop-up window; when the correct word is not displayed in the error correction pop-up window, classify all confusion candidate words according to the set screening mechanism, and display them in the error correction pop-up window, then operate the screening condition in the y direction of the error correction pop-up window by mouse, and operate the confusion candidate words under the screening condition in the x direction of the error correction pop-up window, in order to select the correct word S9, select the correct word from the error correction pop-up window by mouse; S10, replace the error text in the APP with the selected correct word, and exit the text error correction mode through the button on the mouse.
2. The method of claim 1, wherein the method is based on mouse interaction. The APP in step S1 includes but is not limited to Word, WPS, Notepad, WeChat, browser and other APPs in Windows and MacOS system.
3. The method of claim 1, wherein the method further comprises: The AI recognition model in step S1 is one of speech recognition ASR model and image recognition OCR model.
4. The method of claim 1, wherein the method further comprises: When moving the cursor in step S5, the cursor moves to one of the front or back of the error text.
5. The method of claim 1, wherein: The AI large model includes but is not limited to one of Qwen large model and DeepSeek large model.
6. The method of claim 1, wherein: The confusion candidate words given in step S7.2 include but are not limited to the speech recognition result after extracting the error word from the historical voice, the confusion dictionary of error text, the homonym word dictionary of error text, and the confusion word given by the AI large model according to the error word and all contents in the input box.
7. The method of claim 1, wherein: In step S8, all confusion word candidates, error words and context contents are input into the AI large model at the same time, the AI large model sorts the word replacement possibilities in the confusion candidate word list from high to low according to the sorting, and displays the first N words in the error correction window.
8. The method of claim 1, wherein: The screening mechanism is one of pinyin and radical.
Citation Information
Patent Citations
License plate number identification result correction method based on full mouse operation
CN105528139A
Text input method and device, equipment, medium and computer program product
CN116466830A
Text error correction method and device, equipment and storage medium
CN119443087A