Language landscape text grammar error correction method and system based on image recognition
By capturing images using a mobile terminal camera and combining OCR and AR technologies, real-time grammatical error diagnosis and correction of linguistic landscape texts has been achieved. This solves the problem of cumbersome grammar checks and the inability to correct errors in real time in existing technologies, and improves the accuracy and interactivity of translation quality assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG PETROCHEMICAL INST
- Filing Date
- 2026-02-04
- Publication Date
- 2026-05-05
AI Technical Summary
Existing methods for grammatical error checking in linguistic landscape text translation suffer from being cumbersome, inefficient, and unable to achieve real-time, in-situ error correction.
Images are captured by a mobile terminal camera, and combined with optical character recognition (OCR) and a grammar error diagnosis model, grammar error diagnosis is performed in real time. Augmented reality (AR) technology is used to overlay error correction information onto the images, enabling instant and automatic grammar correction and translation quality assessment.
It enables real-time and automatic grammatical error diagnosis and correction in physical scenarios, improving the accuracy and interactivity of information delivery, and providing a technological extension from single grammatical error correction to dual grammatical-semantic quality assessment.
Smart Images

Figure CN121981124A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and natural language processing technology, specifically to a method and system for grammatical error correction of linguistic landscape text based on image recognition, which is particularly suitable for intelligent error correction and standardization of linguistic landscape text. Background Technology
[0002] With the acceleration of globalization, the linguistic landscape of public spaces, such as bilingual signage, wayfinding information, and promotional posters, has become a crucial dimension for measuring internationalization, particularly in terms of the grammatical accuracy and idiomatic expression of its foreign language translations. However, in reality, such translated texts frequently contain various grammatical errors and Chinglish, severely impacting international image and the effective transmission of information. Currently, quality control methods for such texts face significant bottlenecks, primarily relying on the following two types of technical solutions, both of which have fundamental flaws: The first type of technical solution is based on optical character recognition (OCR), which is limited to converting text in an image into a character sequence. The system architecture lacks a functional module for performing syntactic and semantic analysis on the sequence, and therefore cannot identify and correct the language errors in the text itself.
[0003] The second type of technical solution is online syntax analysis, which can analyze text, but its application requires users to manually convert physical text into numerical format. This non-automated interaction mode results in a cumbersome and inefficient process, making it impossible to perform real-time, in-situ checks on text in the physical environment.
[0004] Therefore, developing a method and system for grammatical error correction of language landscape text based on image recognition has become an urgent problem to be solved. Summary of the Invention
[0005] The purpose of this invention is to provide a method and system for correcting the grammar of language landscape text based on image recognition that can seamlessly integrate visual perception, intelligent diagnosis and augmented reality feedback, so as to solve the problems in the background technology where the first type of technical solution (based on optical character recognition) lacks deep analysis capabilities, and the second type of technical solution (online grammar analysis) cannot directly interact with the physical scene.
[0006] Another objective of this invention is to enable users to obtain accurate grammatical error diagnosis, correction suggestions, and immersive visual feedback in real time and automatically simply by using a mobile device to photograph the target text.
[0007] To achieve the above objectives, the technical solution provided by this invention is: a method for correcting linguistic landscape text syntax based on image recognition, comprising the following steps: S1: Capture images containing linguistic landscape text using the mobile terminal's camera; S2: Perform optical character recognition on the image to extract text content in at least one target language; S3: Input the identified target language text into the grammar error diagnosis model to obtain grammar diagnosis results; the grammar diagnosis results include at least the error location, error type, and correction suggestions. S4: On the display screen of the mobile terminal, the grammar diagnosis results are spatially aligned with and overlaid with the text area in the image in an augmented reality manner.
[0008] Furthermore, when the text extracted in step S2 contains both the original text and the translated text, step S3 further includes: Semantic correlation analysis is performed on the source and translated texts to generate translation quality comments. These comments are used to assess the naturalness and accuracy of the translation and are incorporated as supplementary information into the grammatical diagnostic results.
[0009] Furthermore, in step S3, the grammatical error diagnosis model is a personalized grammatical diagnosis model based on the mother tongue background, which can perform in-depth attribution analysis based on the mother tongue background information associated with the text and output the diagnosis results.
[0010] Furthermore, step S4 specifically includes: S41: Locate a text area on the camera feed or a captured still image that is being previewed in real time on the display screen; S42: In the text area, highlight words and phrases with grammatical errors in a first-person visual style according to the location of the error; S43: In response to the user's triggering action on the highlighted annotation, a feedback window pops up, which integrates and displays the error type, correction suggestions, and translation quality comments.
[0011] This invention also provides a language landscape text grammar correction system based on image recognition, comprising: The image acquisition module is used to control the camera to acquire images containing linguistic landscape text; The OCR processing module, connected to the image acquisition module, is used to perform text recognition on the image and output text in at least one target language. The grammar diagnosis module, connected to the OCR processing module, is used to receive text, call the grammar error diagnosis model for processing, and output grammar diagnosis results; the grammar diagnosis results include at least the error location, error type, and correction suggestions; The AR rendering and display module is connected to the image acquisition module and the grammar diagnosis module respectively, and is used to align the diagnosis results with the text positions in the image and overlay them for display.
[0012] Furthermore, it also includes a translation quality analysis module, connected to the OCR processing module, used to compare and analyze the semantic consistency between the source text and the translated text, and generate translation quality comments; the translation quality comments are used to evaluate the naturalness and accuracy of the translation, and are included as supplementary information in the grammar diagnosis results.
[0013] Furthermore, the syntax diagnostic module can be configured in two modes: Mode A is the cloud-based intelligent mode, which calls diagnostic services deployed in the cloud via network API; Mode B is a lightweight terminal mode that runs a lightweight syntax error diagnosis model locally on the terminal.
[0014] The advantages of this invention compared to the prior art are: This invention integrates image recognition, AI diagnostics, and AR display modules, enabling users to obtain real-time analysis results in situ overlay feedback within the same field of view that captures physical text. This fundamentally solves the problem of real-time verification caused by operation switching and scene separation.
[0015] Based on grammatical error correction, this invention expands the technical dimension from single grammatical error correction to dual quality assessment of "grammar-semantics" through correlation analysis between the original text and the translated text.
[0016] This invention uses AR technology to precisely spatially anchor diagnostic information and combines it with interactive pop-ups to achieve a seamless transformation from abstract conclusions to concrete guidance, greatly improving the accuracy of information delivery and the directness of interaction.
[0017] This invention provides the necessary real-time visual input and immersive feedback channels for deep syntax diagnostic algorithms. This coupling of the "front-end interactive system" and the "back-end intelligent core" constitutes a complete technical closed loop from the underlying algorithm to the upper-level application.
[0018] This invention systematically applies deep grammatical and semantic diagnosis and augmented reality interaction technology to the quality verification scenario of linguistic landscape texts in public places, providing efficient tools for fields such as international urban governance and standardization of cultural and tourism services, and realizing an application paradigm upgrade from passive error correction to proactive quality control, and from offline review to real-time on-site handling. Attached Figure Description
[0019] Figure 1 This is a flowchart of a language landscape text grammar correction method based on image recognition according to the present invention.
[0020] Figure 2 This is a system block diagram of a language landscape text grammar correction system based on image recognition according to the present invention.
[0021] Figure 3This is a schematic diagram of the augmented reality (AR) display effect.
[0022] Figure 4 This is a diagram illustrating the connection between the front-end diagnostic module and the back-end service. Detailed Implementation
[0023] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0024] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0025] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0026] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0027] The following detailed description, in conjunction with the accompanying drawings, provides a method and system for correcting grammatical errors in language landscape text based on image recognition, in further detail according to the present invention.
[0028] Combined with appendix Figure 1-4 This invention will be described in detail below.
[0029] A linguistic landscape text grammar correction method based on image recognition achieves real-time correction of grammatical errors in linguistic landscape text through a collaborative process encompassing image acquisition, OCR recognition, grammar diagnosis, and AR overlay display. The method includes the following steps: S1: Image Acquisition Images containing linguistic landscape text are captured using the mobile terminal's camera. These images can be real-time previews from the camera or still images captured and stored. The mobile terminal's camera must support autofocus to ensure clear text images, providing high-quality input for subsequent OCR recognition. It should also be adaptable to different lighting environments, optimizing image capture through automatic exposure adjustment to suit various public scenarios, including outdoor and indoor settings.
[0030] S2: OCR Text Extraction The system performs Optical Character Recognition (OCR) processing on the acquired images to extract text content in at least one target language. During OCR processing, the images are first preprocessed, including grayscale conversion, noise reduction, tilt correction, and text region segmentation to improve text recognition accuracy. Then, a multilingual recognition model is used, adaptable to multiple commonly used languages such as Chinese, English, Japanese, and French, supporting simultaneous extraction of monolingual, bilingual, and multilingual text. When the extracted text contains both source and translation text, the language type and correspondence are simultaneously labeled, providing data support for subsequent translation quality analysis.
[0031] S3: Grammar Diagnosis and Translation Quality Analysis The identified target language text is input into the grammar error diagnosis model to obtain grammar diagnosis results, which include at least the error location, error type, and correction suggestions. The grammar error diagnosis model is a personalized grammar diagnosis model based on native language background. It can perform in-depth attribution analysis based on the native language background information associated with the text. For example, for problems such as tense confusion and article misuse that are common among Chinese-speaking users in English text, it provides targeted attribution and correction suggestions based on native language transfer patterns, improving the personalization and practicality of error correction.
[0032] When the text extracted in step S2 includes both the source text and the translated text, this step also includes semantic association analysis of the source and translated texts to generate translation quality comments. Translation quality comments are used to evaluate the naturalness and accuracy of the translation, specifically including assessments of dimensions such as semantic consistency, sentence structure regularity, and vocabulary naturalness. For example, it determines whether the translated text has semantic deviations, whether it conforms to the expression habits of the target language, and whether there are ambiguities. These comments are then incorporated as supplementary information into the grammar diagnosis results, providing users with more comprehensive text optimization suggestions.
[0033] S4: AR Overlay Display On the mobile terminal's display screen, the grammar diagnosis results are spatially aligned with and overlaid with the text area in the image using augmented reality, achieving real-time correlation between the error correction results and the scene. This includes the following sub-steps: S41: Text Region Positioning: Based on the text region coordinates determined during OCR processing, the text is accurately located on the real-time preview of the camera image or the captured static image on the display screen, ensuring that the virtual error correction information is precisely aligned with the real text region.
[0034] S42: Error Highlighting: In the located text area, based on the error location in the grammar diagnosis results, words and phrases with grammatical errors are highlighted in a first-person visual style to help users quickly identify the error location.
[0035] S43: Interactive Feedback Display: In response to the user's triggering action on the highlighted annotation, a feedback window pops up, which integrates and displays the error type, correction suggestions, and translation quality comments, so that the user can understand the error information and optimization solutions in detail.
[0036] A language landscape text grammar correction system based on image recognition is used to implement the above methods. Through modular design, the system enables the coordinated operation of various functions, ensuring efficient progress throughout the grammar correction process. The system includes the following core modules: The image acquisition module connects to the mobile terminal's camera hardware to control camera startup, focusing, exposure adjustment, and image acquisition. It supports both real-time preview acquisition and still image capture modes. The acquired image data is synchronously transmitted to the OCR processing module and the AR rendering display module. This module has environmental adaptability, automatically adjusting acquisition parameters based on light intensity and text clarity to ensure image quality.
[0037] The OCR processing module connects to the image acquisition module to receive and process the acquired image data, extracting the text content. Internally, it integrates an image preprocessing unit, a multilingual recognition unit, and a text association tagging unit: the image preprocessing unit performs grayscale conversion, noise reduction, tilt correction, and text region segmentation to improve recognition accuracy; the multilingual recognition unit supports text recognition in multiple commonly used languages, enabling simultaneous extraction of monolingual, bilingual, and multilingual text; the text association tagging unit, when the original and translated texts are extracted, tags the language type and correspondence, providing data support for the translation quality analysis module, and finally outputs the processed text data to the grammar diagnosis module and the translation quality analysis module.
[0038] The translation quality analysis module connects to the OCR processing module and receives data containing both the source and translated texts from the OCR output. Using a semantic association analysis algorithm, it compares the semantic consistency, sentence structure, and vocabulary idiomaticity of the source and translated texts to generate translation quality comments. This module is adaptable to translation analysis of different language combinations, and the comments are synchronously transmitted to the grammar diagnosis module as supplementary information.
[0039] The grammar diagnosis module is respectively connected to the OCR processing module and the translation quality analysis module, and is used to receive the text data output by the OCR processing module and the translation quality comments output by the translation quality analysis module, call the grammar error diagnosis model for processing and output the grammar diagnosis results. The grammar diagnosis module can be configured in a dual mode to meet the usage requirements in different scenarios: Mode A is the cloud intelligent mode, which calls the diagnosis service deployed on the cloud through the network API, and relies on the powerful computing resources and massive corpus of the cloud to achieve high-precision grammar diagnosis and in-depth attribution analysis; Mode B is the terminal lightweight mode, which runs a lightweight grammar error diagnosis model locally on the terminal, and can achieve the basic grammar correction function without network connection, adapting to the immediate usage requirements in the network-free environment. The grammar diagnosis results integrate the error location, error type, correction suggestions and translation quality comments, and are synchronously transmitted to the AR rendering display module.
[0040] The AR rendering display module is respectively connected to the image acquisition module and the grammar diagnosis module, and is used to receive the real-time picture or static image output by the image acquisition module, and the grammar diagnosis results output by the grammar diagnosis module. Through the spatial positioning algorithm, the diagnosis results are accurately spatially aligned with the text area in the image, and then the virtual error correction information is superimposed and displayed on the screen of the mobile terminal through the rendering engine. This module supports the real-time synchronous update of virtual information and the real picture, ensuring that when the user moves the mobile terminal, the error correction information always corresponds accurately to the text area. At the same time, it provides an interactive response function to receive the trigger operation of the user on the highlighted annotation and pop up a feedback window.
[0041] The specific implementation process of a grammar error correction method and system for language landscape texts based on image recognition according to the present invention is as follows: Embodiment
[0042] Reference Figure 1 and Figure 2 , this embodiment demonstrates the complete end-to-end workflow of the system. The user opens the dedicated application on the mobile terminal and aims the camera at a public warning sign with the Chinese text "小心滑倒" and the English text "Caution: Slip and Fall Down".
[0043] The image acquisition module controls the camera to capture an image. The OCR processing module recognizes the Chinese text "小心滑倒" and the English text "Caution: Slip and Fall Down".
[0044] The grammar diagnosis module sends the English text to a grammar error diagnosis service for processing. As a preferred solution for achieving in-depth, personalized diagnosis, the system calls upon a service interface provided by a personalized grammar diagnosis system based on native language background (i.e., the system protected by the associated patent). This service can perform in-depth analysis based on text features (which can be inferred or preset to be Chinese native language background), identifying that the original expression "Slip and Fall Down" does not conform to English signage conventions and is redundant; the idiomatic expression should be "Caution: Wet Floor," and the error type is marked as UNIDIOMATIC_EXPRESSION. Simultaneously, the translation quality analysis module compares the semantics of the Chinese and English texts, generating a comment: "The translation conveys the meaning, and 'Caution: Wet Floor' is the most standard and idiomatic warning sign for slippery surfaces in English-speaking countries, with high public awareness and a clear warning effect."
[0045] The AR rendering module highlights problematic phrases (such as "Slip and Fall Down") with a red wavy underline on the real-time preview screen of the phone. When the user clicks on the highlighted area, an interactive feedback window pops up, displaying: "Error: Inauthentic expression. Correction: Caution: Wet Floor. Explanation: This is the most standard English warning." Example
[0046] This embodiment focuses on illustrating the modular and configurable architecture of the system of the present invention, such as... Figure 2 and Figure 4 As shown.
[0047] The system can be deployed as a standalone application on a mobile terminal, a built-in service of the operating system, or a lightweight app. Its core syntax diagnostic capability supports multiple configuration modes to adapt to different scenarios' requirements for real-time performance, accuracy, and network dependency. Mode A (Cloud-based Intelligent Mode): such as Figure 4 As shown, the grammar diagnostic module acts as a client, calling a powerful diagnostic service deployed in the cloud via a network API. This model is particularly suitable for accessing specialized grammar diagnostic systems with deep personalized analysis capabilities, such as those in Example 1, to achieve the highest accuracy in diagnosis and attribution.
[0048] Mode B (Terminal Lightweight Mode): The syntax diagnostic module runs a lightweight syntax error diagnostic model locally on the terminal to meet the requirements of no network environment or extremely high real-time requirements.
[0049] Through the configurable architecture described above, the system of the present invention can achieve an optimal balance between resource consumption and diagnostic capabilities.
[0050] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.
Claims
1. A method for correcting grammatical errors in linguistic landscape text based on image recognition, characterized in that, Includes the following steps: S1: Capture images containing linguistic landscape text using the mobile terminal's camera; S2: Perform optical character recognition on the image to extract text content in at least one target language; S3: Input the identified target language text into the grammar error diagnosis model to obtain grammar diagnosis results; the grammar diagnosis results include at least the error location, error type, and correction suggestions. S4: On the display screen of the mobile terminal, the grammar diagnosis results are spatially aligned with and overlaid with the text area in the image in an augmented reality manner.
2. The method for correcting linguistic landscape text syntax based on image recognition according to claim 1, characterized in that: When the text extracted in step S2 contains both the original text and the translated text, step S3 further includes: Semantic correlation analysis is performed on the source and translated texts to generate translation quality comments. These comments are used to assess the naturalness and accuracy of the translation and are incorporated as supplementary information into the grammatical diagnostic results.
3. The method for correcting linguistic landscape text syntax based on image recognition according to claim 2, characterized in that: In step S3, the grammatical error diagnosis model is a personalized grammatical diagnosis model based on the mother tongue background, which can perform in-depth attribution analysis based on the mother tongue background information associated with the text and output the diagnosis results.
4. The method for correcting linguistic landscape text syntax based on image recognition according to claim 3, characterized in that: Step S4 specifically includes: S41: Locate a text area on the camera feed or a captured still image that is being previewed in real time on the display screen; S42: In the text area, highlight words and phrases with grammatical errors in a first-person visual style according to the location of the error; S43: In response to the user's triggering action on the highlighted annotation, a feedback window pops up, which integrates and displays the error type, correction suggestions, and translation quality comments.
5. A linguistic landscape text grammar correction system based on image recognition, characterized in that, include: The image acquisition module is used to control the camera to acquire images containing linguistic landscape text; The OCR processing module is used to perform text recognition on images and output text in at least one target language. The grammar diagnosis module is used to receive text, call the grammar error diagnosis model for processing, and output the grammar diagnosis results; the grammar diagnosis results include at least the error location, error type, and correction suggestions; The AR rendering and display module is used to align and overlay diagnostic results with the text positions in the image.
6. A language landscape text grammar correction system based on image recognition according to claim 5, characterized in that: It also includes a translation quality analysis module, which is used to compare and analyze the semantic consistency between the source text and the translated text, and generate translation quality comments; the translation quality comments are used to evaluate the naturalness and accuracy of the translation, and are included as supplementary information in the grammar diagnosis results.
7. A language landscape text grammar correction system based on image recognition according to claim 6, characterized in that: The syntax diagnostic module can be configured in two modes: Mode A is the cloud-based intelligent mode, which calls diagnostic services deployed in the cloud via network API; Mode B is a lightweight terminal mode that runs a lightweight syntax error diagnosis model locally on the terminal.