Handwritten character recognition method and device, equipment, storage medium and chip

By selecting the target path architecture in the handwriting recognition model and dynamic resource allocation, the problem of insufficient accuracy and robustness in handwriting recognition in the education field is solved, and fast and accurate text recognition is achieved.

CN120496102APending Publication Date: 2025-08-15NEW ORIENTAL EDUCATION & TECH GRP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510422296.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, general-purpose models cannot guarantee the accuracy and robustness of text recognition in handwriting recognition applications in the field of education, and the response time is long.

Method used

The handwriting recognition model is adopted, and the subject text recognition is recognized by selecting the target path architecture in m path architectures, dynamic resource allocation is performed based on the input mode, task type, input complexity and input characteristics, and the self-attention mechanism Transformer architecture is used to process handwritten text information.

Benefits of technology

It improves the accuracy and practicality of handwritten text recognition, is suitable for different disciplines, reduces response time, and enhances the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496102A_ABST
    Figure CN120496102A_ABST
Patent Text Reader

Abstract

The invention relates to a handwritten character recognition method and device, equipment, a storage medium and a chip, and the method comprises the steps: obtaining handwritten character information which comprises the OCR information of a target subject in an education scene; calling a handwriting recognition model to carry out subject character recognition on the handwritten character information to obtain a corresponding recognition result; the handwriting recognition model comprises m path architectures, m is a positive integer, and the target path architecture is used for performing subject character recognition on the handwritten character information; the target path architecture is a path architecture selected from m path architectures at least based on one of an input mode, a task type, input complexity and input characteristics of the handwritten character information. In this way, the technical problems that in the prior art, a universal model is redundant, the response time is long, and the accuracy of character recognition and the robustness of a model algorithm cannot be guaranteed can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a handwritten text recognition method, device, equipment, storage medium, and chip. Background Art

[0002] Optical Character Recognition (OCR) technology is often used in handwriting recognition applications in the education field. It often encounters problems such as blurred handwriting, different handwriting formats between different people, and synonymy of different symbols.

[0003] In related technologies, a general-purpose model is often used to identify different subjects before performing detection and recognition. However, for the aforementioned handwriting recognition application, the use of a general-purpose model cannot guarantee the accuracy of text recognition. Moreover, the redundancy of general-purpose models can lead to long response times for text recognition and cannot guarantee the robustness of the model algorithm. Summary of the Invention

[0004] To overcome the problems existing in the related technologies, the present disclosure provides a handwritten text recognition method, device, equipment, storage medium and chip to solve the technical problems in the above-mentioned related technologies, such as the redundancy of general models, long response time, inability to guarantee the accuracy of text recognition and the robustness of model algorithms.

[0005] According to a first aspect of an embodiment of the present disclosure, a handwritten text recognition method is provided, comprising: Acquiring handwritten text information, wherein the handwritten text information includes OCR information of a target subject in an educational scenario; Calling a handwriting recognition model to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; the handwriting recognition model includes m path architectures, m is a positive integer, wherein a target path architecture is used to perform subject text recognition on the handwritten text information, and the target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity and input features of the handwritten text information.

[0006] In some embodiments, the calling of the handwriting recognition model to perform subject text recognition on the handwritten text information to obtain the corresponding recognition result includes: Calling the target path architecture in the handwriting recognition model, performing subject text recognition on the handwritten text information according to the allocated target resources, and obtaining the recognition result; The target resource is a resource allocated based on the input complexity and / or task type of the handwritten text information, and the input complexity and the target resource are positively correlated.

[0007] In some embodiments, the target path architecture is selected based on the input complexity of the handwritten text information. Before calling the target path architecture in the handwriting recognition model, the method further includes: Determining the inference depth and / or inference width of the handwriting recognition model based on the input complexity of the handwritten text information; Based on the inference depth and / or inference width of the handwriting recognition model, the target path architecture matching the input complexity is selected from m path architectures.

[0008] In some embodiments, the input complexity is positively correlated with the inference depth and / or inference width of the target pathway architecture, respectively.

[0009] In some embodiments, the handwriting recognition model includes a self-attention mechanism Transformer architecture, and the Transformer architecture is used to perform context relevance and key information processing on the handwritten text information.

[0010] In some embodiments, the Transformer architecture is an architecture among the m path architectures.

[0011] In some embodiments, obtaining handwritten text information includes: Obtain handwritten OCR information of the target subject; Preprocessing the handwritten OCR information to obtain the handwritten text information; The preprocessing includes at least one of data cleaning, standardization, word segmentation and tagging.

[0012] In some embodiments, the input complexity is determined based on the amount of information and time complexity corresponding to the handwritten text information, the amount of model parameters, the amount of calculation and the amount of memory consumption corresponding to the handwriting recognition model, and the time complexity is used to indicate the time required to process the handwritten text information.

[0013] In some embodiments, the input modality includes at least one of text, image, and code; and / or the task type includes at least one of translation, code generation, text summarization, and information filtering.

[0014] According to a second aspect of an embodiment of the present disclosure, there is provided a handwritten text recognition device, comprising: an acquisition module configured to acquire handwritten text information, wherein the handwritten text information includes OCR information of a target subject in an educational scenario; The processing module is configured to call a handwriting recognition model to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; the handwriting recognition model includes m path architectures, m is a positive integer, wherein a target path architecture is used to perform subject text recognition on the handwritten text information, and the target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity and input features of the handwritten text information.

[0015] For the contents not introduced or described in the embodiments of the present disclosure, please refer to the relevant introduction in the aforementioned method embodiments, and the embodiments of the present disclosure do not limit them.

[0016] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the executable instructions to implement the steps of the above-mentioned handwritten text recognition method.

[0017] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the handwritten text recognition method provided in the first aspect of the present disclosure are implemented.

[0018] According to a fifth aspect of an embodiment of the present disclosure, a chip is provided, comprising: a processor and an interface; the processor is configured to read instructions to execute the steps of the above-mentioned handwritten text recognition method.

[0019] The technical solution provided by the embodiments of the present disclosure may include the following beneficial effects: an electronic device acquires handwritten text information, the handwritten text information including OCR information of a target subject in an educational scenario; a handwriting recognition model is called to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; wherein the handwriting recognition model includes m path architectures, a target path architecture is used to perform subject text recognition on the handwritten text information, the target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity, and input features of the handwritten text information, and m is a positive integer. The handwriting recognition model provided by the present disclosure can be applied to handwritten text recognition scenarios of different subjects, can quickly and accurately perform handwritten text recognition for different subjects, can improve the accuracy and practicality of handwritten text recognition, and can effectively solve technical problems existing in related technologies such as the inability to guarantee text recognition accuracy using general models, the long response time caused by the redundancy of general models, and the inability to guarantee the robustness of the model algorithm.

[0020] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0022] Figure 1 The figure is a flowchart of a handwritten text recognition method according to an exemplary embodiment.

[0023] Figure 2 The figure is a flowchart of obtaining handwritten text information according to an exemplary embodiment.

[0024] Figure 3 The figure is a schematic diagram of a process of determining a target path architecture according to an exemplary embodiment.

[0025] Figure 4 The figure is a flowchart of handwritten text recognition according to an exemplary embodiment.

[0026] Figure 5 The figure is a flowchart of a model training according to an exemplary embodiment.

[0027] Figure 6 The figure is a schematic structural diagram of a handwritten text recognition device according to an exemplary embodiment.

[0028] Figure 7 The figure is a schematic structural diagram of an electronic device according to an exemplary embodiment.

[0029] Figure 8 The figure is a schematic structural diagram of a chip according to an exemplary embodiment. DETAILED DESCRIPTION

[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0031] It should be noted that all actions of acquiring signals, information or data in the present disclosure are performed with the authorization of the corresponding device owner.

[0032] In some scenarios, such as for senior science subjects involving mathematics, physics, and chemistry, teachers will use the OCR recognition function of the test papers in the process of grading the test papers. In some scenarios, it is difficult to distinguish whether the defining formula belongs to mathematics or physics. For example, students may need to understand and apply complex problems that include physics, mathematics, and engineering formulas. In order to facilitate business docking, reduce communication costs and redundancy of open platforms, it is very necessary to develop a unified version of the handwriting recognition model. In actual applications, OCR technology is often used in handwriting recognition applications in the field of education, such as blurred handwriting, diverse font formats of different people's writing, and synonymy of different symbols. Therefore, a handwriting recognition model that is both universal and accurate is needed. Specifically: (1) Blurred handwriting: Since everyone has different writing habits and skills, some people's handwriting may be blurry, which is a challenge for OCR technology. In addition, if the text is written on paper, the wear and stains of the paper may also make the handwriting blurry.

[0033] (2) Diversity of fonts and formats: In the field of education, a variety of fonts and formats may be involved, including different languages, mathematical symbols, scientific symbols, etc. This requires OCR technology to have sufficient flexibility and adaptability.

[0034] (3) Contextual understanding: Handwritten text often contains rich contextual information, which is very important for understanding the meaning of the text. However, OCR technology may have difficulty understanding this contextual information.

[0035] (4) Error correction: If OCR technology generates errors during the recognition process, these errors need to be corrected promptly and accurately. However, error correction is a complex process that requires OCR technology to have a high degree of accuracy and reliability.

[0036] Existing technologies require differentiating between different subjects before relying solely on a single model for detection and recognition. However, these general models cannot guarantee algorithm robustness and cannot guarantee recognition accuracy for handwritten text in various forms. Furthermore, due to the redundancy of these models, their response time is relatively long, making them incapable of rapid handwritten text recognition.

[0037] To solve the above problems, the present disclosure provides a handwritten text recognition method, device, equipment, storage medium and chip. Figure 1 FIG. 1 is a flow chart of a handwritten text recognition method according to an exemplary embodiment. Figure 1 The method shown is applied to an electronic device, and the method may include the following implementation steps: S101. Acquire handwritten text information, where the handwritten text information includes OCR information of a target subject in an education scenario.

[0038] In this disclosure, the handwritten text information may refer to OCR information of a target subject in an educational setting, which may include, but is not limited to, mathematics, Chinese, English, physics, chemistry, or other subjects. This disclosure does not limit the format of the handwritten text information, which may include, but is not limited to, images, text, or other formats.

[0039] The present disclosure does not limit the manner in which the handwritten text information is obtained. For example, the present disclosure may obtain the handwritten text information from other devices (such as a server or a network platform) through the network. Figure 2 FIG. 1 is a flow chart showing a process of obtaining handwritten text information according to an exemplary embodiment. Figure 2 The process shown may include the following implementation steps: S201: Obtain handwritten OCR information of the target subject.

[0040] In the present disclosure, the above-mentioned handwritten OCR information may refer to the information obtained by scanning and recognizing the target subject after the user inputs the handwritten information using OCR technology. For the introduction of the above-mentioned target subject, please refer to the relevant introduction in the aforementioned embodiment, which will not be repeated here.

[0041] The present disclosure does not limit the implementation method for obtaining the above-mentioned handwritten OCR information. For example, the handwritten content of the target subject can be directly scanned and recognized using OCR technology, or it can be obtained from other devices or network platforms through the Internet.

[0042] S202: Pre-process the handwritten OCR information to obtain the handwritten text information; wherein the pre-processing includes at least one of data cleaning, standardization, word segmentation, and tagging.

[0043] The present disclosure can pre-process the above-mentioned handwritten OCR information to obtain the above-mentioned handwritten text information. The above-mentioned pre-processing may refer to a processing method that the system customizes according to actual needs. For example, it may include but is not limited to any one or more of the following combinations: data cleaning, standardization, word segmentation, annotation processing or other customized processing methods. Among them, the above-mentioned data cleaning may include but is not limited to, for example, stop word removal, punctuation, number and special character removal, removal of irrelevant information and potential noise, etc. The above-mentioned standardization may refer to processing in a pre-configured standardized format to ensure that the data is in a standardized format in preparation for the next step of analysis. The above-mentioned word segmentation may refer to dividing the text into individual words for the convenience of subsequent processing. The above-mentioned annotation processing may refer to breaking down the text into smaller units, such as words or sub-words, and breaking down the text into individual sentences for the convenience of subsequent analysis and processing.

[0044] S102. Call a handwriting recognition model to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; wherein, the handwriting recognition model includes m path architectures, and a target path architecture is used to perform subject text recognition on the handwritten text information. The target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity, and input features of the handwritten text information, and m is a positive integer.

[0045] The present disclosure may employ an active path selection strategy to select the most appropriate target path architecture from m path architectures in a handwriting recognition model based on the relevant information of the handwritten text information, and then use the target path architecture to perform subject text recognition on the handwritten text information, outputting a corresponding recognition result. The relevant information may include, but is not limited to, any one or a combination of the following information: the input modality of the handwritten text information, the task type, the input complexity, the input characteristics, or other information that influences path selection.

[0046] This disclosure does not limit the specific implementation of the handwriting recognition model described above. For example, it can typically be a large language model such as PaLM 2. This disclosure also does not limit the internal structure of the handwriting recognition model described above (such as PaLM 2). For example, it can adopt a self-attention Transformer architecture including an encoder and decoder, adopt deeper network layers and a larger parameter scale, adopt some new attention mechanisms or other innovative modules, etc., to improve the model's performance in specific fields or tasks.

[0047] The following is a brief introduction to the above active path selection strategy: (1) Multimodal processing: The handwriting recognition model is able to process data in multiple modalities, including but not limited to any one or a combination of the following modalities: text, images, code, or other modalities. In practical applications, the model may have some mechanism to identify and determine the input modality of the above handwritten text information and select the (target) path architecture that is most suitable for processing the modality.

[0048] (2) Multi-task adaptability: The model supports multiple tasks and can select the most appropriate (target) path architecture for handling the task type. The task type or tasks may include, but are not limited to, any one or a combination of the following: translation, code generation, text summarization, information filtering, or other tasks or task types customized according to actual needs.

[0049] (3) Adaptive reasoning: The model can choose different reasoning depths and / or widths based on the input complexity (also known as the level of sophistication). For example, for simple questions, the model can choose a shallower reasoning path architecture; for complex questions, which may require a deeper thinking process, the model can choose a deeper reasoning path architecture.

[0050] (4) Attention mechanism: The attention mechanism allows the model to focus on certain key parts when processing sequence data. This can be seen as a path selection mechanism because it determines which information is more important.

[0051] (5) Multi-layer architecture: The model uses a multi-layer architecture, and different layers may focus on processing different types of information. The model can select a specific layer / specific path architecture for processing based on the input features.

[0052] (6) Dynamic routing: There can be some mechanism within the model to dynamically determine the path of information transmission. For example, some inputs may pass directly through certain layers / path architectures, while other inputs may require more processing, etc.

[0053] (7) Mixture of Experts: The model can include a Mixture of Experts (MoE), which is a model structure that can include multiple expert modules, each of which is good at processing different types of inputs. The model can dynamically select the most appropriate expert module to process based on the input type.

[0054] In one specific embodiment, to ensure optimal utilization of computing resources, the disclosed model can employ adaptive computing. During inference, the model can dynamically allocate corresponding target resources (or computing power) based on input complexity and / or task type. For more complex sentences or queries requiring additional processing power, the model can intelligently allocate resources to maintain efficiency while providing accurate and timely responses. In a specific implementation, the disclosed model can invoke the target path architecture in the aforementioned handwriting recognition model to perform subject-specific text recognition on the handwritten text information according to the allocated target resources, thereby outputting corresponding recognition results. The target resources can be allocated based on the input complexity and / or task type of the handwritten text information. The input complexity and target resources are positively correlated, meaning that greater input complexity leads to greater target resources allocated, while smaller input complexity leads to smaller target resources allocated. When the task type is a specified type (e.g., a non-language subject), the disclosed model can allocate a larger target resource. Conversely, when the task type is a non-specified type (e.g., a simple subject like mathematics), the disclosed model can allocate a smaller target resource.

[0055] In another specific embodiment, the above input complexity is used as an example to introduce the specific implementation method of determining the target path architecture. Figure 3 FIG. 1 is a flow chart showing a target path architecture determination process according to an exemplary embodiment. Figure 3 The process shown may include the following implementation steps: S301: Determine the inference depth and / or inference width of the handwriting recognition model based on the input complexity of the handwritten text information.

[0056] The present disclosure can select different inference depths and / or inference widths based on different input complexities. Generally, the above-mentioned input complexity is positively correlated with the model's inference depth and / or inference width, respectively. For example, the greater the input complexity, the greater the model's inference depth and inference width; conversely, the smaller the input complexity, the smaller the model's inference depth and inference width.

[0057] In an optional embodiment, the above-mentioned input complexity of the present disclosure may be determined based on the amount of information and time complexity corresponding to the above-mentioned handwritten text information, the amount of model parameters, the amount of computation and the amount of memory consumption corresponding to the above-mentioned handwritten text recognition model. The above-mentioned amount of information may refer to the size used to measure / describe the above-mentioned handwritten text information. Taking the above-mentioned handwritten text information as image information as an example, the amount of information may refer to the image width W and image height H corresponding to the above-mentioned handwritten text information. The above-mentioned time complexity may refer to the time t required for processing the above-mentioned handwritten text information, etc. The above-mentioned amount of model parameters may refer to the number p of model parameters possessed by the above-mentioned handwriting recognition model, etc. The above-mentioned amount of computation may refer to the amount of computation required for the above-mentioned handwriting recognition model to perform text recognition. For example, the above-mentioned handwriting recognition model includes n convolutional layers, and the amount of computation during the operation of each convolutional layer is c. i , then the above calculation amount C is nc i The memory consumption may refer to the total memory size consumed by the handwriting recognition model. For example, the memory occupied by the model parameters of the handwriting recognition model is m p The memory required for the intermediate calculation results of the above handwriting recognition model is m i , then the memory consumption m is m p With m i sum.

[0058] Referring to the example in the above embodiment, the above input complexity O can be expressed as shown in the following formula (1): Formula (1) in, 、 、 、 、 ,and These are pre-given weight factors that reflect / indicate the importance of each complexity component. These weight factors can be adjusted according to specific application scenarios and personal preferences, and this disclosure does not make too many restrictions or details on this.

[0059] For example, assuming that all the above complexity components are equally important, they can be set to the same weight of 1 / 6. At the same time, assuming that the image width W corresponding to the above handwritten text information is 28 pixels, the image height H is also 28 pixels, the above handwriting recognition model has p=10000 model parameters, the entire model has 3 convolutional layers, and the computational load of each convolutional layer is: c1=1000, c2=2000, c3=1500, then the total computational load of the model C= c1+c2+c3=4500. The memory m occupied by the model parameters of the above handwriting recognition model is p The memory occupied by the intermediate calculation results is 40,000 bytes. iis 20,000 bytes, the total memory consumption of the model is m=m p +m i = 60,000 bytes. The time complexity t corresponding to the above handwritten text information is assumed to be 0.01 seconds. Substituting these values into the above formula (1), it can be calculated that the above input complexity O is approximately equal to 12426.0.

[0060] S302: Based on the inference depth and / or inference width of the handwriting recognition model, select the target path architecture that matches the input complexity from m path architectures.

[0061] The present disclosure selects a matching corresponding path architecture from m path architectures based on the determined inference depth and / or inference width as the final target path architecture, thereby performing corresponding handwritten text recognition.

[0062] It should be noted that the handwriting recognition model (e.g., PaLM 2) disclosed herein may include an innovative m-path architecture, which distinguishes it from traditional language models. Unlike traditional models characterized by a single path of information flow, the handwriting recognition model (PaLM 2) introduces multiple path architectures, each of which can process different types of information, enabling the model to develop subtle, targeted expertise for each aspect of language understanding. Please also refer to Figure 4 FIG. 1 is a flow chart showing a handwritten text recognition process according to an exemplary embodiment. Figure 4 As shown, the handwriting recognition model (PaLM 2) disclosed in the present invention can perform handwriting recognition of different subjects or multiple subjects, thereby outputting multi-subject handwriting recognition results.

[0063] It can be understood that the above-mentioned handwriting recognition model (PaLM 2) includes multi-path architectures, which can be decoupled and operate independently, that is, one path does not interfere with / affect the processing of other paths. For example, one path architecture can focus on syntactic structure, analyzing grammar and word order; while another path architecture can emphasize the semantic meaning of the text. Path decoupling enables the model to focus on various aspects of language understanding, thereby gaining a more comprehensive understanding of the input text. Although these path architectures operate independently, they are not isolated from each other. The present disclosure can adopt a general artificial intelligence architecture (such as the Pathways architecture) to allow these path architectures to interact and exchange relevant information to promote comprehensive language understanding. The interaction between path architectures promotes cross-learning and enhances the overall understanding ability of the model.

[0064] In an optional embodiment, the handwriting recognition model (PaLM 2) includes a Transformer architecture, which can be used to perform contextual relevance and key information processing on the handwritten text information. Optionally, the Transformer architecture can be an architecture in the m-path architecture. In a specific implementation, the handwriting recognition model (PaLM2) can be built on the basis of the Transformer architecture, which has revolutionized natural language processing by introducing a self-attention mechanism, allowing the model to more effectively capture long-range dependencies and context. Among them, the self-attention mechanism enables the model to measure the importance of different words in a sentence based on contextual relevance, thereby achieving more accurate prediction and understanding of the text. The Transformer architecture improves training efficiency and supports parallel processing, making it suitable for large-scale language models such as PaLM 2.

[0065] The following describes relevant embodiments of handwriting recognition model training. The model training involved in this disclosure can be applied to training devices, which may include but are not limited to electronic devices, servers, or other devices with model training capabilities. Figure 5 FIG. 1 is a flow chart showing a model training process according to an exemplary embodiment. Figure 5 Taking the PaLM 2 model as an example, the present disclosure may first obtain a training data set, where the training data set includes at least one set of training samples, and the training samples include handwriting OCR information and handwriting recognition results.

[0066] The present disclosure does not limit the acquisition of the above-mentioned training data set. For example, a large and diverse initial data set can be obtained from various sources, and the initial data set can include texts from books, articles, websites, social media and other language resources. Furthermore, the above-mentioned initial data set can be carefully pre-processed, which can specifically refer to the relevant introduction of the aforementioned pre-processing, for example, the initial text is cleaned to eliminate irrelevant information, special symbols and potential noise. The tagging process can break down the text into smaller units, such as words or sub-words, and break down the text into individual sentences. This pre-processing step can ensure that the data is in a standardized format in preparation for further analysis.

[0067] Next, pre-training is performed on a massive training dataset. After obtaining the pre-processed training dataset, PaLM2 enters the unsupervised pre-training phase. During this phase, the model learns to predict missing words in sentences, understand context, and generate coherent text. The pre-training phase involves iterative training on a large dataset, exposing PaLM2 to a wide range of language patterns, structures, and semantics. During model training, an initial learning rate can be set. After each training session, the model's loss value can be calculated to check whether it continues to decrease. If the loss value no longer decreases significantly, this may indicate that the learning rate is too high or the model is nearing convergence. If the loss value does not change as expected, a rule can be set. For example, if the loss value does not decrease significantly after several consecutive model training sessions, the learning rate can be reduced. This learning rate reduction rule can be pre-set, for example, by multiplying the learning rate by a decay factor, such as 0.1. Model training then continues using the new learning rate, and the change in the loss value is monitored. During the above pre-training phase, the present disclosure may repeat the above steps regularly and adjust the learning rate according to the changes in the model loss value, thereby achieving model pre-training better and more accurately.

[0068] Next, the model is fine-tuned based on a specific domain or task. While pre-training gives PaLM 2 a broad understanding of language, fine-tuning takes it a step further by specializing the model to a specific domain or task. Fine-tuning narrows the model focus by training on a smaller, domain-specific or task-specific training dataset for a specific application. The domain-specific or task-specific training dataset may include, but is not limited to, datasets for specific domains or tasks such as sentiment analysis, question answering, and natural language understanding. Fine-tuning helps the model adjust its knowledge and expertise to meet the specific requirements of different real-world language processing tasks, making it more valuable and practical in various contexts. Figure 5 In the paper, the mapper involved in model fine-tuning can be understood as the multi-path architecture of the model to a certain extent, and model fine-tuning can be understood as fine-tuning the model through training datasets such as instructions and dialogues to adapt to better handwritten text recognition in specific fields or specific tasks.

[0069] After selecting a target path architecture and processing the input handwritten text, the PaLM 2 model can generate and output corresponding recognition results based on the designed fine-tuning task. This disclosure does not limit the recognition results output by the model; they can take various forms, such as predicted words for language completion tasks, sentiment scores for sentiment analysis, or detailed answers to questions in question-answering tasks.

[0070] As can be seen, the handwriting recognition model (PaLM 2) described in this disclosure utilizes a Transformer-like encoder and decoder architecture, which has demonstrated excellent performance on various natural speech processing tasks and effectively captures the semantic information of the input text. Compared to traditional models, the handwriting recognition model (PaLM 2) can utilize deeper network layers, enabling it to learn more abstract and complex feature representations, thereby improving the model's understanding and generation capabilities. To further expand the model's learning capabilities, the handwriting recognition model (PaLM 2) can have a larger number of parameters. This larger model size allows it to cover a wider range of knowledge domains and enhance its generalization performance across various tasks. The disclosed model can incorporate innovations and optimizations to its attention mechanism to better capture key information and contextual relevance / dependencies in the input text, thereby enhancing the model's semantic understanding capabilities. The training dataset for the disclosed model can be further expanded and optimized to cover a wider range of knowledge domains. This will provide the model with more comprehensive background knowledge, leading to even better performance on tasks such as question answering and text generation. For specific application scenarios, the present disclosure can perform targeted model fine-tuning and optimization on the handwriting recognition model (PaLM 2) to further improve the model's performance on these tasks.

[0071] By adopting the above-mentioned embodiments of the present disclosure, the robustness of the handwriting recognition model algorithms of different subjects can be improved, and the calls of different business units can be conveniently supported, thereby reducing communication costs. Moreover, the above-mentioned model algorithm has good compatibility and is easy to expand and integrate. In a specific implementation, the electronic device obtains handwritten text information, and the handwritten text information includes OCR information of the target subject in the educational scenario; the handwriting recognition model is called to perform subject text recognition on the handwritten text information to obtain the corresponding recognition result; wherein, the handwriting recognition model includes m path architectures, and the target path architecture is used to perform subject text recognition on the handwritten text information, and the target path architecture is a path architecture selected from m path architectures based on at least one of the input modality, task type, input complexity and input features of the handwritten text information, and m is a positive integer. The handwriting recognition model provided by the present invention can be applied to handwriting recognition scenarios in different disciplines, and can quickly and accurately perform handwriting recognition for different disciplines, thereby improving the accuracy and practicality of handwriting recognition. It can effectively solve technical problems existing in related technologies such as the inability to guarantee the accuracy of text recognition, the redundancy of general models resulting in long response times, and the inability to guarantee the robustness of model algorithms.

[0072] The electronic devices involved in the present disclosure may also be referred to as terminal devices, which may include but are not limited to, for example, smart phones, wearable devices (personal hubs), personal or mobile multimedia players, personal digital assistants, laptop computers, tablet computers, smart books, handheld computers or other devices with communication functions.

[0073] Based on the above examples, see Figure 6 FIG. 1 is a schematic diagram showing the structure of a handwritten text recognition device according to an exemplary embodiment. Figure 6 The device shown can be applied to an electronic device, and the device may include an acquisition module 601 and a processing module 602. The acquisition module 601 is configured to acquire handwritten text information, wherein the handwritten text information includes OCR information of a target subject in an educational scenario; The processing module 602 is configured to call a handwriting recognition model to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; In which, the handwriting recognition model includes m path architectures, and the target path architecture is used to perform subject text recognition on the handwritten text information. The target path architecture is a path architecture selected from m path architectures based on at least one of the input modality, task type, input complexity and input features of the handwritten text information, and m is a positive integer.

[0074] In some embodiments, the processing module 602 is configured to: Calling the target path architecture in the handwriting recognition model, performing subject text recognition on the handwritten text information according to the allocated target resources, and obtaining the recognition result; The target resource is a resource allocated based on the input complexity and / or task type of the handwritten text information, and the input complexity and the target resource are positively correlated.

[0075] In some embodiments, the target path architecture is selected based on the input complexity of the handwritten text information. Before calling the target path architecture in the handwriting recognition model, the processing module 602 is further configured to: Determining the inference depth and / or inference width of the handwriting recognition model based on the input complexity of the handwritten text information; Based on the inference depth and / or inference width of the handwriting recognition model, the target path architecture matching the input complexity is selected from m path architectures.

[0076] In some embodiments, the input complexity is positively correlated with the inference depth and / or inference width of the target pathway architecture, respectively.

[0077] In some embodiments, the handwriting recognition model includes a self-attention mechanism Transformer architecture, and the Transformer architecture is used to perform context relevance and key information processing on the handwritten text information.

[0078] In some embodiments, the Transformer architecture is an architecture among the m path architectures.

[0079] In some embodiments, the acquisition module 601 is configured to: Obtain handwritten OCR information of the target subject; Preprocessing the handwritten OCR information to obtain the handwritten text information; The preprocessing includes at least one of data cleaning, standardization, word segmentation and tagging.

[0080] In some embodiments, the input complexity is determined based on the amount of information and time complexity corresponding to the handwritten text information, the amount of model parameters, the amount of calculation and the amount of memory consumption corresponding to the handwriting recognition model, and the time complexity is used to indicate the time required to process the handwritten text information.

[0081] In some embodiments, the input modality includes at least one of text, image, and code; and / or the task type includes at least one of translation, code generation, text summarization, and information filtering.

[0082] The above-mentioned device can obtain handwritten text information, and the handwritten text information includes OCR information of the target subject in the educational scenario; call the handwriting recognition model to perform subject text recognition on the handwritten text information to obtain the corresponding recognition result; wherein, the handwriting recognition model includes m path architectures, and the target path architecture is used to perform subject text recognition on the handwritten text information, and the target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity and input characteristics of the handwritten text information, and m is a positive integer. The handwriting recognition model provided by the present disclosure can be applied to handwritten text recognition scenarios of different subjects, and can quickly and accurately perform handwritten text recognition for different subjects, which can improve the accuracy and practicality of handwritten text recognition, and can effectively solve technical problems existing in related technologies such as the inability to guarantee text recognition accuracy using general models, the long response time caused by the redundancy of general models, and the inability to guarantee the robustness of model algorithms.

[0083] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the handwritten text recognition method provided by the present disclosure are implemented.

[0084] Figure 7 5 is a schematic diagram showing the structure of an electronic device according to an exemplary embodiment. For example, the electronic device 500 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0085] Reference Figure 7 , the electronic device 500 may include one or more of the following components: a processing component 502 , a memory 504 , a power component 506 , a multimedia component 508 , an audio component 510 , an input / output interface 512 , a sensor component 514 , and a communication component 516 .

[0086] Processing component 502 generally controls the overall operation of device 500, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. Processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the handwritten text recognition method described above. In addition, processing component 502 may include one or more modules to facilitate interaction between processing component 502 and other components. For example, processing component 502 may include a multimedia module to facilitate interaction between multimedia component 508 and processing component 502.

[0087] The memory 504 is configured to store various types of data to support operations on the device 500. Examples of such data include instructions for any application or method operating on the device 500, contact data, phone book data, messages, pictures, videos, etc. The memory 504 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0088] Power supply component 506 provides power to the various components of device 500. Power supply component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 500.

[0089] The multimedia component 508 includes a screen that provides an output interface between the device 500 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 508 includes a front-facing camera and / or a rear-facing camera. When the electronic device 500 is in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.

[0090] The audio component 510 is configured to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is configured to receive external audio signals when the device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 also includes a speaker for outputting audio signals.

[0091] The input / output interface 512 provides an interface between the processing component 502 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0092] The sensor assembly 514 includes one or more sensors for providing various aspects of the status assessment of the device 500. For example, the sensor assembly 514 can detect the open / closed state of the device 500, the relative positioning of components, such as the display and keypad of the device 500. The sensor assembly 514 can also detect changes in the position of the device 500 or a component of the device 500, the presence or absence of user contact with the device 500, the orientation or acceleration / deceleration of the device 500, and temperature changes of the device 500. The sensor assembly 514 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 514 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 514 may also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0093] The communication component 516 is configured to facilitate wired or wireless communication between the device 500 and other devices. The device 500 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 516 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0094] In an exemplary embodiment, the device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described handwriting recognition method.

[0095] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions. The instructions are executable by the processor 520 of the device 500 to perform the above-described superordinate handwritten character recognition method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0096] In addition to being an independent electronic device, the above-mentioned device can also be part of an independent electronic device. For example, in one embodiment, the device can be an integrated circuit (IC) or a chip, where the integrated circuit can be a single IC or a collection of multiple ICs. The chip can include, but is not limited to, the following types: GPU (Graphics Processing Unit), CPU (Central Processing Unit), FPGA (Field Programmable Gate Array), DSP (Digital Signal Processor), ASIC (Application Specific Integrated Circuit), SOC (System on Chip, SoC), etc. The above-mentioned integrated circuit or chip can be used to execute executable instructions (or code) to implement the above-mentioned handwritten text recognition method. The executable instructions can be stored in the integrated circuit or chip, or obtained from other devices or equipment, such as the integrated circuit or chip including a processor, memory, and an interface for communicating with other devices. The executable instruction can be stored in the memory, and when the executable instruction is executed by the processor, the above-mentioned handwritten text recognition method is implemented; alternatively, the integrated circuit or chip can receive the executable instruction through the interface and transmit it to the processor for execution to implement the above-mentioned handwritten text recognition method.

[0097] In another exemplary embodiment, a computer program product is provided. The computer program product includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing the above-mentioned handwritten text recognition method when executed by the programmable device.

[0098] See Figure 8 FIG. 1 is a schematic diagram showing the structure of a chip according to an exemplary embodiment. Figure 8 The chip 600 shown includes a processor 601 and an interface 602. Optionally, it may also include a memory 603. The number of the processor 601 may be one or more, and the number of the interface 602 may be multiple.

[0099] In one embodiment, for a case where a chip is used to implement the method embodiment described in the present disclosure: The interface 602 is used to receive or output signals; The processor 601 is configured to execute part or all of the contents of the handwritten text recognition method embodiment.

[0100] It is understandable that the processor in the embodiments of the present disclosure can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by hardware integrated logic circuits in the processor or software instructions. The above processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0101] It is understood that the memory in the embodiments of the present disclosure may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0102] It should be noted that the descriptions of the above storage medium, device, and chip embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, storage medium, and device embodiments of the present disclosure, please refer to the descriptions of the method embodiments of the present disclosure for understanding.

[0103] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the present disclosure. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0104] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A handwritten character recognition method, characterized in that: include: Acquiring handwritten text information, wherein the handwritten text information includes OCR information of a target subject in an educational scenario; Calling a handwriting recognition model to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; the handwriting recognition model includes m path architectures, m is a positive integer, wherein a target path architecture is used to perform subject text recognition on the handwritten text information, and the target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity and input features of the handwritten text information.

2. The method according to claim 1, characterized in that The calling of the handwriting recognition model to perform subject text recognition on the handwritten text information to obtain the corresponding recognition result includes: Calling the target path architecture in the handwriting recognition model, performing subject text recognition on the handwritten text information according to the allocated target resources, and obtaining the recognition result; The target resource is a resource allocated based on the input complexity and / or task type of the handwritten text information, and the input complexity and the target resource are positively correlated.

3. The method according to claim 2, characterized in that The target path architecture is selected based on the input complexity of the handwritten text information. Before calling the target path architecture in the handwriting recognition model, the method further includes: Determining the inference depth and / or inference width of the handwriting recognition model based on the input complexity of the handwritten text information; Based on the inference depth and / or inference width of the handwriting recognition model, the target path architecture matching the input complexity is selected from m path architectures.

4. The method according to claim 3, characterized in that The input complexity is positively correlated with the inference depth and / or inference width of the target path architecture, respectively.

5. The method according to claim 1, wherein The handwriting recognition model includes a self-attention mechanism Transformer architecture, which is used to process the context relevance and key information of the handwritten text information.

6. The method according to claim 5, characterized in that The Transformer architecture is an architecture among the m path architectures.

7. The method according to claim 1, characterized in that The obtaining of handwritten text information includes: Obtain handwritten OCR information of the target subject; Preprocessing the handwritten OCR information to obtain the handwritten text information; The preprocessing includes at least one of data cleaning, standardization, word segmentation and tagging.

8. The method according to any one of claims 1 to 7, characterized in that The input complexity is determined based on the information volume and time complexity corresponding to the handwritten text information, the model parameter volume, calculation volume and memory consumption corresponding to the handwriting recognition model. The time complexity is used to indicate the time required to process the handwritten text information.

9. The method according to any one of claims 1 to 7, characterized in that The input modality includes at least one of text, image, and code; and / or the task type includes at least one of translation, code generation, text summarization, and information filtering.

10. A handwritten character recognition device, characterized in that: include: an acquisition module configured to acquire handwritten text information, wherein the handwritten text information includes OCR information of a target subject in an educational scenario; A processing module is configured to call a handwriting recognition model to perform subject text recognition on the handwritten text information to obtain a corresponding recognition result; The handwriting recognition model includes m path architectures, where m is a positive integer, wherein a target path architecture is used to perform subject text recognition on the handwritten text information, and the target path architecture is a path architecture selected from the m path architectures based on at least one of the input modality, task type, input complexity, and input features of the handwritten text information.

11. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the executable instructions to implement the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

13. A chip, characterized in that: The method comprises a processor and an interface; the processor is used to read instructions to execute the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Handwritten text recognition method and system based on image-structure multi-mode mutual learning

    CN121259846A