Object recognition processing method, processing device, electronic device and storage medium

By identifying and processing different types of text objects, and using the converter learning model to realize the conversion and recognition of text objects, the problems of low recognition accuracy and batching efficiency in the prior art are solved, and the efficiency of natural language processing is improved.

CN112001152BActive Publication Date: 2025-05-30HANGZHOU DANA TECH INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010864625.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-25
Publication Date
2025-05-30
Estimated Expiration
2041-02-13

AI Technical Summary

Technical Problem

In the field of natural language processing, especially text-to-text conversion, it is difficult to effectively identify and process different types of text objects, such as calculation and knowledge questions, resulting in inaccuracy of recognition and incorrect correction efficiency.

Method used

By obtaining the object to be identified, identifying the type of the object based on the type recognition model, determining the processing rules based on the type, and using the converter learning model to identify and process the objects, realizing the conversion and recognition of different types of text objects.

Benefits of technology

The accuracy and efficiency of text object recognition are improved, and different types of text problems can be automatically corrected, solving the problems of difficulty in identification and low correction efficiency in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112001152B_ABST
    Figure CN112001152B_ABST
Patent Text Reader

Abstract

An object recognition processing method, a processing device, an electronic device, and a non-transitory computer-readable storage medium. The object recognition processing method includes: obtaining an object to be recognized; recognizing the type of the object to be recognized based on a type recognition model; determining a processing rule corresponding to the object to be recognized according to the type of the object to be recognized; in response to the type of the object to be recognized being a basic type, taking the object to be recognized as a target object to be recognized according to the processing rule, and in response to the type of the object to be recognized being a non-basic type, converting the object to be recognized through a converter learning model according to the processing rule to convert the object to be recognized into a target object to be recognized; performing recognition processing on the target object to be recognized through the converter learning model to obtain a target result corresponding to the object to be recognized. The type of the target object to be recognized is a basic type.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to an object recognition processing method, a processing device, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] A Transformer learning model (Transformer model) consists of an encoder and a decoder, and can implement the conversion from one text to another text, thus being widely applied in the field of natural language processing (NLP), for example, machine translation, question-and-answer systems, text and speech recognition, etc. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides an object recognition processing method, including: obtaining an object to be recognized; recognizing the type of the object to be recognized based on a type recognition model; determining a processing rule corresponding to the object to be recognized according to the type of the object to be recognized; and recognizing and processing the object to be recognized through a Transformer learning model according to the processing rule to obtain a target result corresponding to the object to be recognized, where, according to the processing rule, recognizing and processing the object to be recognized through the Transformer learning model to obtain a target result corresponding to the object to be recognized includes: in response to the type of the object to be recognized being a basic type, taking the object to be recognized as a target object to be recognized according to the processing rule, and in response to the type of the object to be recognized being a non-basic type, converting the object to be recognized through the Transformer learning model according to the processing rule to convert the object to be recognized into a target object to be recognized; and recognizing and processing the target object to be recognized through the Transformer learning model to obtain a target result corresponding to the object to be recognized, where the type of the target object to be recognized is the basic type.

[0004] For example, in the object recognition processing method provided in an embodiment of the present disclosure, recognizing and processing the target object to be recognized through the Transformer learning model to obtain a target result corresponding to the object to be recognized includes: recognizing and processing the target object to be recognized through the Transformer learning model to obtain a processing result corresponding to the target object to be recognized; and processing the processing result according to the processing rule to obtain a target result corresponding to the object to be recognized.

[0005] For example, in the object recognition processing method provided in an embodiment of the present disclosure, the basic type includes computational solution questions or knowledge-based short answer questions, and the non-basic type includes computational fill-in-the-blank questions, computational true-or-false questions, computational multiple-choice questions, knowledge-based fill-in-the-blank questions, knowledge-based true-or-false questions, or knowledge-based multiple-choice questions.

[0006] For example, in the object recognition processing method provided in an embodiment of the present disclosure, the to-be-recognized object is transformed by the transformer learning model to transform the to-be-recognized object into a target to-be-recognized object, including: in response to the type of the to-be-recognized object being the computational fill-in-the-blank question or the knowledge fill-in-the-blank question, directly transforming the to-be-recognized object into the target to-be-recognized object by the transformer learning model; in response to the type of the to-be-recognized object being the computational true-or-false question, deleting the judgment result in the to-be-recognized object by the transformer learning model to transform the to-be-recognized object into a first intermediate to-be-recognized object, and transforming the first intermediate to-be-recognized object into the target to-be-recognized object, wherein the type of the first intermediate to-be-recognized object is the computational fill-in-the-blank question; in response to the type of the to-be-recognized object being the computational multiple-choice question or the knowledge multiple-choice question, deleting each option in the to-be-recognized object by the transformer learning model and transforming the stem in the to-be-recognized object into the target to-be-recognized object; in response to the type of the to-be-recognized object being the knowledge true-or-false question, transforming the to-be-recognized object into a third intermediate to-be-recognized object by the transformer learning model, wherein the type of the third intermediate to-be-recognized object is the knowledge multiple-choice question, deleting each option in the third intermediate to-be-recognized object and transforming the stem in the third intermediate to-be-recognized object into a second intermediate to-be-recognized object, and transforming the second intermediate to-be-recognized object into the target to-be-recognized object, wherein the type of the second intermediate to-be-recognized object is the knowledge fill-in-the-blank question.

[0007] For example, in the object recognition processing method provided in an embodiment of the present disclosure, according to the type of the to-be-recognized object, a processing rule corresponding to the to-be-recognized object is determined, including: in response to the type of the to-be-recognized object being the computational fill-in-the-blank question or the knowledge fill-in-the-blank question, the processing rule corresponding to the to-be-recognized object includes adding keywords to the target to-be-recognized object; in response to the type of the to-be-recognized object being the computational multiple-choice question or the knowledge multiple-choice question, the processing rule corresponding to the to-be-recognized object includes selecting an option that is the same as the processing result from each option in the to-be-recognized object; in response to the type of the to-be-recognized object being the computational true-or-false question, the processing rule corresponding to the to-be-recognized object includes comparing the processing result with the judgment result in the to-be-recognized object; in response to the type of the to-be-recognized object being the knowledge true-or-false question, the processing rule corresponding to the to-be-recognized object includes selecting an option that is the same as the processing result from each option in the third intermediate to-be-recognized object.

[0008] For example, in the object recognition processing method provided in an embodiment of the present disclosure, according to the processing rule, the processing result is processed to obtain a target result corresponding to the object to be recognized, including: in response to the type of the object to be recognized being the calculation type solution question or the knowledge type short answer question, directly outputting the processing result as the target result; in response to the type of the object to be recognized being the calculation type fill-in-the-blank question or the knowledge type fill-in-the-blank question, taking the part corresponding to the keyword in the processing result as the target result; in response to the type of the object to be recognized being the calculation type multiple-choice question or the knowledge type multiple-choice question, selecting the option identical to the processing result from the various options in the object to be recognized, and taking the serial number of the option identical to the processing result as the target result; in response to the type of the object to be recognized being the calculation type true / false question, comparing the processing result with the judgment result in the object to be recognized to obtain a comparison result, and taking the comparison result as the target result, where the comparison result includes correct, wrong or pending; in response to the type of the object to be recognized being the knowledge type true / false question, selecting the option identical to the processing result from the various options in the third intermediate object to be recognized, and taking the option identical to the processing result as the target result.

[0009] For example, in the object recognition processing method provided in an embodiment of the present disclosure, when the type of the object to be recognized is the calculation type fill-in-the-blank question, the keyword is "fill in the blank", and when the type of the object to be recognized is the knowledge type fill-in-the-blank question, the keyword is "fill in Chinese characters".

[0010] For example, in the object recognition processing method provided in an embodiment of the present disclosure, obtaining the object to be recognized includes: obtaining an input image, where the input image includes the object to be recognized; processing the input image through a region recognition model to obtain a region to be recognized including the object to be recognized; and processing the region to be recognized through a character recognition model to obtain the object to be recognized.

[0011] For example, in the object recognition processing method provided in an embodiment of the present disclosure, the input image is an image of a test paper, and the object to be recognized is a test question on the test paper.

[0012] For example, in the object recognition processing method provided in an embodiment of the present disclosure, the converter learning model is a model based on a neural network.

[0013] At least one embodiment of the present disclosure provides a processing device, including: an acquisition module configured to acquire an object to be recognized; a type recognition module configured to recognize the type of the object to be recognized based on a type recognition model; a rule determination module configured to determine a processing rule corresponding to the object to be recognized according to the type of the object to be recognized; a converter module configured to recognize and process the object to be recognized through a converter learning model according to the processing rule to obtain a target result corresponding to the object to be recognized, wherein when the converter module executes the step of recognizing and processing the object to be recognized through the converter learning model according to the processing rule to obtain a target result corresponding to the object to be recognized, the following steps are included: in response to the type of the object to be recognized being a basic type, taking the object to be recognized as a target object to be recognized according to the processing rule; in response to the type of the object to be recognized being a non-basic type, converting the object to be recognized through the converter learning model according to the processing rule to convert the object to be recognized into a target object to be recognized; and recognizing and processing the target object to be recognized through the converter learning model to obtain a target result corresponding to the object to be recognized, wherein the type of the target object to be recognized is the basic type.

[0014] At least one embodiment of the present disclosure provides an electronic device, including: a memory configured to non-transiently store computer-readable instructions; a processor configured to run the computer-readable instructions, and when the computer-readable instructions are run by the processor, the object recognition processing method according to any one of the above embodiments is implemented.

[0015] At least one embodiment of the present disclosure provides a non-transient computer-readable storage medium, wherein the non-transient computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the object recognition processing method according to any one of the above embodiments is implemented. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0017] Figure 1 It is a schematic flowchart of an object recognition processing method provided by at least one embodiment of the present disclosure;

[0018] Figure 2 It is a schematic block diagram of a processing device provided by at least one embodiment of the present disclosure;

[0019] Figure 3 It is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure;

[0020] Figure 4 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure; and

[0021] Figure 5 Schematic diagram of a hardware environment provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0022] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0023] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms "including" or "comprising" and the like mean that the elements or items appearing before the term cover the elements or items listed after the term and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0024] In order to keep the following description of the embodiments of the present disclosure clear and concise, the detailed descriptions of some known functions and known components are omitted in the present disclosure.

[0025] At least one embodiment of the present disclosure provides an object recognition processing method, a processing device, an electronic device, and a non-transitory computer-readable storage medium. The object recognition processing method includes: obtaining an object to be recognized; recognizing the type of the object to be recognized based on a type recognition model; determining a processing rule corresponding to the object to be recognized according to the type of the object to be recognized; and recognizing and processing the object to be recognized through a converter learning model according to the processing rule to obtain a target result corresponding to the object to be recognized.

[0026] For example, according to the processing rules, the converter learning model performs recognition processing on the object to be recognized to obtain the target result corresponding to the object to be recognized, including: in response to the type of the object to be recognized being a basic type, according to the processing rules, the object to be recognized is used as the target object to be recognized; in response to the type of the object to be recognized being a non-basic type, according to the processing rules, the converter learning model performs conversion on the object to be recognized to convert the object to be recognized into the target object to be recognized; the converter learning model performs recognition processing on the target object to be recognized to obtain the target result corresponding to the object to be recognized. The type of the target object to be recognized is a basic type.

[0027] The object recognition processing method uses different processing rules to recognize different types of objects to be recognized (such as word problems, etc.) to obtain the target result corresponding to the object to be recognized. The object recognition processing method can reduce the amount of calculation and improve the parallel efficiency without compromising the final recognition result, and the recognition accuracy is relatively high.

[0028] For example, the object recognition processing method can implement automatically grading word problems in a test paper, automatically obtaining the answers to word problems in a sample test paper, etc. Thus, the technical solution provided by the present disclosure can solve the problems of difficult recognition of word problems and low efficiency of grading test papers in the prior art.

[0029] It should be noted that the object recognition processing method provided by the embodiments of the present disclosure can be applied to the processing device provided by the embodiments of the present disclosure, and the processing device can be configured on an electronic device. The electronic device can be a personal computer, a mobile terminal, etc., and the mobile terminal can be a hardware device such as a mobile phone or a tablet computer with various operating systems.

[0030] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0031] Figure 1 It is a schematic flowchart of an object recognition processing method provided by at least one embodiment of the present disclosure.

[0032] For example, as Figure 1 shown, the object recognition processing method provided by the embodiments of the present disclosure includes steps S10 to S13.

[0033] As Figure 1 shown, first, the object recognition processing method includes step S10: obtaining the object to be recognized.

[0034] For example, the object to be recognized may include printed or machine - input characters, symbols, graphics, etc., and may also include handwritten characters, symbols, and graphics, etc. The characters may include Chinese (e.g., Chinese characters or pinyin), English, Japanese, French, Korean, Latin, numbers (e.g., dates, weights, dimensions, etc.), etc. The symbols may include greater - than symbol, less - than symbol, percent sign, etc. The graphics may include circles, rectangles, etc.

[0035] For example, in some embodiments, step S10 includes: obtaining an input image, where the input image includes the object to be recognized; processing the input image through a region recognition model to obtain a region to be recognized including the object to be recognized; and processing the region to be recognized through a character recognition model to obtain the object to be recognized.

[0036] For example, in some embodiments, the input image may be an image of a test paper, an image of an exercise set, etc. For example, the test paper may be a sample test paper (e.g., the sample test paper may include a test paper filled with reference answers or standard answers, etc.) or a student's exam test paper, etc. The object to be recognized may be a question on the test paper or an exercise in the exercise set, etc. For example, the object to be recognized may be a word problem, etc. The test paper may be a test paper for various subjects, such as Chinese, Mathematics, Foreign Languages (e.g., English, etc.). Similarly, the exercise set may also be an exercise set for various subjects, etc. For example, the shape of the input image may be a rectangle, etc. The shape, size, etc. of the input image may be set by the user according to the actual situation.

[0037] For example, the input image may be an image captured by an image acquisition device (e.g., a digital camera or a mobile phone, etc.). The input image may be a grayscale image or a color image. It should be noted that the input image refers to a form that visually presents the object to be recognized, such as a picture of the object to be recognized, etc. Also, for example, the input image may also be obtained through scanning, etc. For example, the input image may be an image directly captured by the image acquisition device or an image obtained after pre - processing the captured image. For example, in order to avoid the influence of the data quality, data imbalance, etc. of the image directly captured by the image acquisition device on the object recognition processing, before processing the input image, the object recognition processing method may also include an operation of pre - processing the image directly captured by the image acquisition device. The pre - processing may include, for example, operations such as cropping, Gamma correction, or noise reduction filtering on the image directly captured by the image acquisition device. The pre - processing can eliminate the irrelevant information or noise information in the input image to facilitate better subsequent processing of the input image.

[0038] For example, the region recognition model can be implemented using machine learning techniques and run on, for example, a general-purpose computing device or a dedicated computing device, and the region recognition model is a pre-trained model. The region recognition model can be implemented using neural networks such as a deep convolutional neural network (DEEP-CNN) or a deep residual network (ResNet).

[0039] For example, the shape of the region to be recognized can be a regular shape or an irregular shape. The shape of the region to be recognized is determined by the object to be recognized. For example, if the characters, symbols, and graphics in the object to be recognized are arranged in a row, the shape of the region to be recognized can be a rectangle. The present disclosure includes but is not limited to the above cases, as long as the region to be recognized can cover the object to be recognized.

[0040] For example, the character recognition model can be implemented based on technologies such as optical character recognition (OCR) and run on, for example, a general-purpose computing device or a dedicated computing device. For example, the character recognition model can also be a pre-trained model, and the character recognition model can be a neural network-based model.

[0041] Next, as Figure 1 shown, in step S11, the type of the object to be recognized is recognized based on the type recognition model.

[0042] For example, the type recognition model can be implemented based on machine learning techniques, and the type recognition model is a pre-trained model. The type recognition model can be implemented using neural networks such as a convolutional neural network (CNN) and run on, for example, a general-purpose computing device or a dedicated computing device.

[0043] For example, the types of the object to be recognized include basic types and non-basic types. In some embodiments, the basic types can include computational solution questions, knowledge-based question-and-answer questions, etc., and the non-basic types include computational fill-in-the-blank questions, computational true-or-false questions, computational multiple-choice questions, knowledge-based fill-in-the-blank questions, knowledge-based true-or-false questions, knowledge-based multiple-choice questions, etc. It should be noted that the types of the object to be recognized include but are not limited to the above examples, and can also be other suitable types, such as application questions, etc.

[0044] For example, in some embodiments, an example of a computational solution question is expressed as "The school bought 10 stationery boxes and 8 notebooks, and paid a total of 85.6 yuan. Given that each stationery box is 6 yuan, how much is each notebook?". An example of a knowledge-based question-and-answer question is expressed as "What kind of phenomenon is the rotation of the optical disc in the optical disc drive?".

[0045] For example, a calculation fill-in-the-blank question refers to a fill-in-the-blank question in which the question stem contains multiple numbers and the answer filled in by the user is also a number. In some embodiments, an example of a calculation fill-in-the-blank question is expressed as: "The number of Jia is 19.6, which is 1.25 less than the number of Yi. What is the sum of the numbers of Jia and Yi? ()". The content within the parentheses "()" represents the answer that the user needs to fill in, and this answer is a number.

[0046] For example, a calculation multiple-choice question refers to a multiple-choice question in which the question stem contains multiple numbers and each option is a number. In some embodiments, an example of a calculation multiple-choice question is expressed as: "When the side length of a square is expanded by 3 times, the area is also expanded by () times. A. 3 B. 6 C. 9 D. 27". The question stem of this calculation multiple-choice question is "When the side length of a square is expanded by 3 times, the area is also expanded by () times.", and each option of this calculation multiple-choice question is "A. 3 B. 6 C. 9 D. 27".

[0047] For example, a calculation true / false question refers to a true / false question in which the question stem contains multiple numbers. In some embodiments, an example of a calculation true / false question is expressed as: "The base of an isosceles triangle is 5 cm and the waist is 3 cm. Its perimeter is 11 cm. ()". The content within the parentheses "()" represents the answer that the user needs to fill in, and this answer can be "correct" or "false" (or "right" or "wrong").

[0048] For example, a knowledge fill-in-the-blank question refers to a fill-in-the-blank question in which there are no or very few numbers in the question stem, and the answer filled in by the user has no numbers, purely for testing knowledge points. In some embodiments, an example of a knowledge fill-in-the-blank question is expressed as: "The rotation of an optical disc in an optical disc drive belongs to () phenomenon?". The content within the parentheses "()" represents the answer that the user needs to fill in, and this answer is text, not a number.

[0049] For example, a knowledge multiple-choice question refers to a multiple-choice question in which there are no or very few numbers in the question stem, and there are no numbers in each option, purely for testing knowledge points. In some embodiments, an example of a knowledge multiple-choice question is expressed as: "The surface of the desk in the classroom is (). A. rectangle B. square C. circle". The question stem of this knowledge multiple-choice question is "The surface of the desk in the classroom is ().", and each option of this knowledge multiple-choice question is "A. rectangle B. square C. circle".

[0050] For example, a knowledge true / false question refers to a true / false question in which there are no or very few numbers in the question stem, purely for testing knowledge points. In some embodiments, an example of a knowledge multiple-choice question is expressed as: "In an acute triangle, the sum of any two interior angles is greater than a right angle. ()". The content within the parentheses "()" represents the answer that the user needs to fill in, and this answer can be "correct" or "false" (or "right" or "wrong").

[0051] Next, as Figure 1As shown, in step S12, according to the type of the object to be recognized, a processing rule corresponding to the object to be recognized is determined.

[0052] For example, different types of objects to be recognized correspond to different processing rules.

[0053] Finally, as Figure 1 shown, in step S13, according to the processing rule, the object to be recognized is recognized and processed by the transducer learning model to obtain the target result corresponding to the object to be recognized.

[0054] For example, the transducer learning model is a neural network-based model. The transducer learning model is used to implement the conversion of one text to another text. The transducer learning model uses the multi-head self-attention mechanism. The multi-head self-attention mechanism can comprehensively utilize various aspects of information / features. Different heads will pay attention to position information, grammatical information, semantic information, rare words, etc. differently, so that the transducer learning model can learn relevant information in different representation subspaces. For example, when a user browses a web page, the user may pay more attention to dark text in terms of color, and pay attention to large and bold text in terms of font. Here, color and font are two different representation subspaces. Paying attention to both color and font (i.e., the multi-head self-attention mechanism) can effectively locate the emphasized content on the web page.

[0055] For example, the transducer learning model can process all words or symbols in a sequence in parallel, and at the same time use the self-attention mechanism to combine the context with more distant words. By processing all words in parallel and allowing each word to notice other words in the sentence in multiple processing steps, the training speed of the transducer learning model is relatively fast. When the transducer learning model is applied to the field of machine translation, the translation results obtained by the transducer learning model are relatively accurate.

[0056] For example, in some embodiments, step S13 may include: in response to the type of the object to be recognized being a basic type, according to the processing rule, taking the object to be recognized as the target object to be recognized; in response to the type of the object to be recognized being a non-basic type, according to the processing rule, converting the object to be recognized by the transducer learning model to convert the object to be recognized into the target object to be recognized; and performing recognition processing on the target object to be recognized by the transducer learning model to obtain the target result corresponding to the object to be recognized.

[0057] For example, the type of the target object to be recognized is a basic type.

[0058] For example, in some embodiments, in step S13, the target object to be recognized is recognized by a transformer learning model to obtain a target result corresponding to the object to be recognized, including: the target object to be recognized is recognized by the transformer learning model to obtain a processing result corresponding to the target object to be recognized; according to the processing rule, the processing result is processed to obtain a target result corresponding to the object to be recognized.

[0059] For example, the object to be recognized is transformed by a transformer learning model to transform the object to be recognized into a target object to be recognized, including: in response to the type of the object to be recognized being a computational fill-in-the-blank question or a knowledge fill-in-the-blank question, the object to be recognized is directly transformed into a target object to be recognized by the transformer learning model; in response to the type of the object to be recognized being a computational true / false question, the judgment result in the object to be recognized is deleted by the transformer learning model to transform the object to be recognized into a first intermediate object to be recognized, and the first intermediate object to be recognized is transformed into a target object to be recognized, where the type of the first intermediate object to be recognized is a computational fill-in-the-blank question; in response to the type of the object to be recognized being a computational multiple-choice question or a knowledge multiple-choice question, each option in the object to be recognized is deleted by the transformer learning model, and the stem in the object to be recognized is transformed into a target object to be recognized; in response to the type of the object to be recognized being a knowledge true / false question, the object to be recognized is transformed into a third intermediate object to be recognized by the transformer learning model, where the type of the third intermediate object to be recognized is a knowledge multiple-choice question, each option in the third intermediate object to be recognized is deleted, and the stem in the third intermediate object to be recognized is transformed into a second intermediate object to be recognized, and the second intermediate object to be recognized is transformed into a target object to be recognized, where the type of the second intermediate object to be recognized is a knowledge fill-in-the-blank question.

[0060] For example, when the type of the object to be recognized is a computational fill-in-the-blank question, a computational true / false question, or a computational multiple-choice question, the type of the target object to be recognized is a computational solution question; when the type of the object to be recognized is a knowledge fill-in-the-blank question, a knowledge true / false question, or a knowledge multiple-choice question, the type of the target object to be recognized is a knowledge essay question.

[0061] For example, according to the processing rules, the processing result is processed to obtain a target result corresponding to the object to be recognized, including: in response to the type of the object to be recognized being a calculation-based solution question or a knowledge-based short-answer question, directly outputting the processing result as the target result; in response to the type of the object to be recognized being a calculation-based fill-in-the-blank question or a knowledge-based fill-in-the-blank question, taking the part corresponding to the keyword in the processing result as the target result; in response to the type of the object to be recognized being a calculation-based multiple-choice question or a knowledge-based multiple-choice question, selecting the option that is the same as the processing result from each option in the object to be recognized, and taking the serial number of the option that is the same as the processing result as the target result; in response to the type of the object to be recognized being a calculation-based true-or-false question, comparing the processing result with the judgment result in the object to be recognized to obtain a comparison result, and taking the comparison result as the target result, where the comparison result includes correct, wrong, or pending; in response to the type of the object to be recognized being a knowledge-based true-or-false question, selecting the option that is the same as the processing result from each option in the third intermediate object to be recognized, and taking the option that is the same as the processing result as the target result.

[0062] For example, in some embodiments, step S12 includes: in response to the type of the object to be recognized being a calculation-based fill-in-the-blank question or a knowledge-based fill-in-the-blank question, the processing rule corresponding to the object to be recognized includes adding a keyword to the target object to be recognized; in response to the type of the object to be recognized being a calculation-based multiple-choice question or a knowledge-based multiple-choice question, the processing rule corresponding to the object to be recognized includes selecting the option that is the same as the processing result from each option in the object to be recognized; in response to the type of the object to be recognized being a calculation-based true-or-false question, the processing rule corresponding to the object to be recognized includes comparing the processing result with the judgment result in the object to be recognized; in response to the type of the object to be recognized being a knowledge-based true-or-false question, the processing rule corresponding to the object to be recognized includes selecting the option that is the same as the processing result from each option in the third intermediate object to be recognized.

[0063] For example, in the case where the type of the object to be recognized is a calculation-based fill-in-the-blank question, the keyword is "fill in the blank", and in the case where the type of the object to be recognized is a knowledge-based fill-in-the-blank question, the keyword is "fill in Chinese". For example, for a calculation-based fill-in-the-blank question, since the answer to a calculation-based fill-in-the-blank question does not need to output the calculation steps, only the number needs to be filled in (this number is the target result corresponding to the calculation-based fill-in-the-blank question), therefore, the keyword "fill in the blank" is added in front of the stem of the calculation-based fill-in-the-blank question to indicate that the converter learning model does not need to output the calculation steps, but only outputs the number. For a knowledge-based fill-in-the-blank question, since the answer corresponding to a knowledge-based fill-in-the-blank question does not require a complete answer, but only the key words need to be output (these words are the target result corresponding to the knowledge-based fill-in-the-blank question), therefore, the keyword "fill in Chinese" is added in front of the stem of the knowledge-based fill-in-the-blank question to indicate that the converter learning model does not need to output a complete answer, but only outputs the key words as the answer.

[0064] For example, for calculation multiple-choice questions or knowledge multiple-choice questions, the calculation multiple-choice question or knowledge multiple-choice question includes a question stem and each option, and the target result corresponding to the calculation multiple-choice question or knowledge multiple-choice question is the serial number of one or more options. For calculation multiple-choice questions, each option is a number, and for knowledge multiple-choice questions, each option is text. It should be noted that for calculation multiple-choice questions, in some examples, at least some of the options may also include text and / or symbols, graphics, etc.; similarly, for knowledge multiple-choice questions, in some examples, at least some of the options may also include numbers and / or symbols, graphics, etc.

[0065] For example, for calculation true / false questions, the calculation true / false question includes a judgment result, and the purpose of the calculation true / false question is to judge whether the judgment result is correct. Thus, the target result corresponding to the calculation true / false question can be correct, false, or undetermined.

[0066] For example, for knowledge true / false questions, the purpose of the knowledge true / false question is to judge whether the knowledge point reflected by the test question is correct. Thus, the target result corresponding to the calculation true / false question can be correct, false, or undetermined.

[0067] For example, when the type of the object to be recognized is a calculation short-answer question or a knowledge essay question, the object to be recognized is used as the target object to be recognized, and the target object to be recognized is subjected to conversion and recognition processing through a converter learning model, so as to obtain the processing result corresponding to the target object to be recognized, and finally the processing result is directly output as the target result.

[0068] For example, the converter learning model can be pre-trained through a large number of training samples (various types of test questions such as calculation short-answer questions, knowledge essay questions, calculation fill-in-the-blank questions, knowledge fill-in-the-blank questions, calculation true / false questions, knowledge true / false questions, calculation multiple-choice questions, knowledge multiple-choice questions, etc. and their corresponding target results).

[0069] For example, in one embodiment, a training sample includes a knowledge-based question-and-answer problem and its corresponding target result. The knowledge-based question-and-answer problem is expressed as: What kind of phenomenon is the rotation of an optical disc in an optical disc drive? The target result corresponding to this knowledge-based question-and-answer problem is expressed as: The rotation of an optical disc in an optical disc drive is a rotational phenomenon. After training a converter learning model with a large number of similar knowledge-based question-and-answer problems, the converter learning model can learn that: A rotating in B is a rotational phenomenon. Thus, when the knowledge-based question-and-answer problem "What kind of phenomenon is the rotation of an optical disc in an optical disc drive?" is input into the converter learning model, the converter learning model can obtain the processing result: "The rotation of an optical disc in an optical disc drive is a rotational phenomenon", and this processing result is directly output as the target result. That is to say, the converter learning model can output "The rotation of an optical disc in an optical disc drive is a rotational phenomenon"; for another example, when the knowledge-based question-and-answer problem "What kind of phenomenon is the rotation of the clock hands on a clock face?" is input into the converter learning model, the converter learning model can obtain the processing result: "The rotation of the clock hands on a clock face is a rotational phenomenon", and this processing result is directly output as the target result. That is to say, the converter learning model can output "The rotation of the clock hands on a clock face is a rotational phenomenon".

[0070] For example, in one embodiment, a training sample includes a calculation-based problem-solving question and its corresponding target result. The calculation-based problem-solving question is expressed as: The school bought 10 stationery boxes and 8 notebooks, and paid a total of 85.6 yuan. Given that each stationery box costs 6 yuan, how much does each notebook cost? The target result corresponding to the calculation-based problem-solving question is expressed as: (85.6 - 10 * 6) / 8 = 3.2 yuan Answer: Each notebook costs 3.2 yuan. After training a converter learning model with a large number of similar calculation-based problem-solving questions, when the calculation-based problem-solving question "The school bought 10 stationery boxes and 8 notebooks, and paid a total of 85.6 yuan. Given that each stationery box costs 6 yuan, how much does each notebook cost?" is input into the converter learning model, the converter learning model can obtain the processing result: "(85.6 - 10 * 6) / 8 = 3.2 yuan Answer: Each notebook costs 3.2 yuan". This processing result is directly output as the target result. That is to say, the converter learning model can output "(85.6 - 10 * 6) / 8 = 3.2 yuan Answer: Each notebook costs 3.2 yuan".

[0071] For example, when the type of the object to be identified is a calculation-type fill-in-the-blank question or a knowledge-type fill-in-the-blank question, the converter learning model can directly convert the object to be identified into a target object to be identified, that is, convert the calculation-type fill-in-the-blank question into a calculation-type answer question, and convert the knowledge-type fill-in-the-blank question into a knowledge-type question-answer question; then, the converter learning model performs identification processing on the target object to be identified, thereby obtaining the processing result corresponding to the target object to be identified, and finally outputs the part of the processing result corresponding to the keyword as the target result.

[0072] For example, in some embodiments, for a calculation-type fill-in-the-blank question "Number A is 19.6, which is 1.25 less than number B. What is the sum of numbers A and B?", first, the calculation-type fill-in-the-blank question is converted into a calculation-type answer question "Number A is 19.6, which is 1.25 less than number B. What is the sum of numbers A and B?" In addition, according to the processing rules corresponding to the calculation-type fill-in-the-blank question, it is necessary to add the keyword "fill-in-the-blank" before the stem of the calculation-type fill-in-the-blank question to indicate that no calculation steps need to be output, thereby converting the calculation-type fill-in-the-blank question into a target object to be recognized. The target object to be identified is represented as "Fill in the blank: Number A is 19.6, which is 1.25 less than Number B. What is the sum of numbers A and B?"; then, the target object to be identified is input into the converter learning model, and the converter learning model can obtain the processing result "19.6+1.25+19.6=40.45\nAnswer: 40.45"; finally, according to the processing rules, the part of the processing result corresponding to the keyword is used as the target result, that is, "40.45" is used as the target result, and finally, the converter learning model outputs "40.45". It should be noted that for calculation-type fill-in-the-blank questions, the part of the processing result corresponding to the keyword can be the number after "Answer:" in the processing result.

[0073] For example, in some embodiments, for the knowledge-based fill-in-the-blank question "The rotation of an optical disc in an optical disc drive belongs to () phenomenon?", first, this knowledge-based fill-in-the-blank question is converted into a knowledge-based question-and-answer question "What phenomenon does the rotation of an optical disc in an optical disc drive belong to?". In addition, according to the processing rules corresponding to the knowledge-based fill-in-the-blank question, the keyword "fill in Chinese" needs to be added before the stem of the knowledge-based fill-in-the-blank question to indicate that a complete answer does not need to be output, thereby converting this knowledge-based fill-in-the-blank question into a target object to be recognized, and the target object to be recognized is expressed as "Fill in Chinese: What phenomenon does the rotation of an optical disc in an optical disc drive belong to?"; then, this target object to be recognized is input into the converter learning model, and the converter learning model can obtain the processing result "The rotation of an optical disc in an optical disc drive belongs to the rotation phenomenon"; finally, according to the processing rules, the part corresponding to the keyword in the processing result is used as the target result, that is, "rotation" is used as the target result. Finally, the converter learning model outputs "rotation". It should be noted that for the knowledge-based fill-in-the-blank question, the part corresponding to the keyword in the processing result can be the text corresponding to the position of "what" in the target recognition object in the processing result.

[0074] For example, when the type of the object to be recognized is a calculation-based true / false question, the converter learning model can delete the judgment result in the object to be recognized to convert the object to be recognized into a first intermediate object to be recognized, and the type of the first intermediate object to be recognized is a calculation-based fill-in-the-blank question. Then, the first intermediate object to be recognized is converted into a target object to be recognized; then, the converter learning model performs recognition processing on the target object to be recognized, thereby obtaining the processing result corresponding to the target object to be recognized. Then, the processing result is compared with the judgment result in the object to be recognized to obtain a comparison result, and finally the comparison result is output as the target result.

[0075] For example, for a calculation-based true / false question, the judgment result can be a number.

[0076] For example, the comparison result can include correct, wrong, or pending. For example, when the processing result and the judgment result are the same, the comparison result is correct; when the processing result and the judgment result are different, the comparison result is wrong or pending. It should be noted that since the correct rate output by the converter learning model is not 100%, if the processing result is different from the deleted judgment result (i.e., the number), the conclusion that the calculation-based true / false question is wrong cannot be obtained. At this time, the converter learning model can output wrong or pending, and can output this calculation-based true / false question and the processing result for the user to judge. In addition, in some other embodiments, if the processing result is the same as the deleted judgment result, at this time, the converter learning model can output correct, and can output this calculation-based true / false question and the processing result for the user to verify.

[0077] For example, in some embodiments, for the calculation-based true / false question "The base of an isosceles triangle is 5 cm and the waist is 3 cm. Its perimeter is 11 cm. ( )", first, determine the judgment result in this calculation-based true / false question, that is, the number "11". Then, the converter learning model can delete this judgment result (i.e., "11") and regard the position corresponding to this judgment result as the part to be filled in, so as to convert the calculation-based true / false question into a calculation-based fill-in-the-blank question (i.e., the first intermediate object to be recognized) "The base of an isosceles triangle is 5 cm and the waist is 3 cm. Its perimeter is ( ) cm"; Next, convert the converted calculation-based fill-in-the-blank question into a calculation-based solution question "The base of an isosceles triangle is 5 cm and the waist is 3 cm. What is its perimeter in centimeters?" In addition, according to the processing rules corresponding to the calculation-based fill-in-the-blank question, the keyword "fill in the blank" needs to be added before the stem of the calculation-based fill-in-the-blank question to indicate that no calculation steps need to be output. Finally, the calculation-based true / false question is converted into the target object to be recognized, and the target object to be recognized is expressed as "Fill in the blank: The base of an isosceles triangle is 5 cm and the waist is 3 cm. What is its perimeter in centimeters?"; Then, input this target object to be recognized into the converter learning model, and the converter learning model can obtain the intermediate processing result "5 + 3 + 3 = 11 cm\nAnswer: 11"; Next, take the part corresponding to the keyword in the intermediate processing result (i.e., the number after "Answer:" in the intermediate processing result) as the processing result, that is, take "11" as the processing result; Finally, according to the processing rules corresponding to the calculation-based true / false question, compare the processing result with the judgment result in the object to be recognized to obtain the comparison result, and take the comparison result as the target result. In this example, the processing result "11" is the same as the judgment result "11", so the comparison result is "correct", and finally the converter learning model outputs "correct".

[0078] For example, when the type of the object to be recognized is a calculation-based multiple-choice question, the converter learning model deletes each option in the object to be recognized and converts the stem in the object to be recognized into the target object to be recognized; Then, the converter learning model performs recognition processing on the target object to be recognized to obtain the processing result corresponding to the target object to be recognized. Next, select the option that is the same as the processing result from each option in the object to be recognized. Finally, output the serial number of the option that is the same as the processing result as the target result.

[0079] For example, the serial numbers of each option can include A, B, C, D or 1, 2, 3, 4, etc. The number of each option can be determined according to the actual situation. For example, 3 or 4, etc.

[0080] For example, in some embodiments, for the computational multiple-choice question "If the side length of a square is enlarged by 3 times, the area is enlarged by () times. A. 3 B. 6 C. 9 D. 27", where the question stem is "If the side length of a square is enlarged by 3 times, the area is enlarged by () times.", and each option is "A. 3 B. 6 C. 9 D. 27". The question stem in the computational multiple-choice question is a computational fill-in-the-blank question. First, determine each option in the computational multiple-choice question, that is, "A. 3 B. 6 C. 9 D. 27". Then, the converter learning model can delete each option, and then convert the question stem (i.e., the computational fill-in-the-blank question) in the computational multiple-choice question into a computational answer question "If the side length of a square is enlarged by 3 times, how many times is the area enlarged?". In addition, according to the corresponding processing rules for the computational fill-in-the-blank question, the keyword "Fill in the blank" needs to be added before the question stem of the computational fill-in-the-blank question. Finally, the computational judgment question is converted into a target object to be recognized, and the target object to be recognized is expressed as "Fill in the blank: If the side length of a square is enlarged by 3 times, how many times is the area enlarged?". Then, input the target object to be recognized into the converter learning model, and the converter learning model can obtain an intermediate processing result "3 * 3 = 9\nAnswer: 9". Then, take the part corresponding to the keyword in the intermediate processing result (i.e., the number after "Answer:" in the intermediate processing result) as the processing result, that is, take "9" as the processing result. Finally, according to the corresponding processing rules for the computational multiple-choice question, select the option that is the same as the processing result from each option in the object to be recognized. In this example, the processing result "9" is the same as the option "9", so the serial number "C" corresponding to this option "9" is used as the target result, and finally the converter learning model outputs "C".

[0081] For example, when the type of the object to be recognized is a knowledge multiple-choice question, the converter learning model deletes each option in the object to be recognized and converts the question stem in the object to be recognized into a target object to be recognized. Then, the converter learning model performs recognition processing on the target object to be recognized to obtain a processing result corresponding to the target object to be recognized. Then, select the option that is the same as the processing result from each option in the object to be recognized. Finally, output the serial number of the option that is the same as the processing result as the target result.

[0082] For example, the serial numbers of each option can include A, B, C, D or 1, 2, 3, 4, etc. The number of each option can be determined according to the actual situation, for example, 3 or 4, etc.

[0083] For example, in some embodiments, for the knowledge multiple-choice question "The surface of the desk in the classroom is (). A. Rectangle B. Square C. Circle", where the question stem is "The surface of the desk in the classroom is ().", and each option is "A. Rectangle B. Square C. Circle". The question stem in the knowledge multiple-choice question is a knowledge fill-in-the-blank question. First, determine each option in the knowledge multiple-choice question, that is, "A. Rectangle B. Square C. Circle". The converter learning model can delete each option and convert the question stem (i.e., the knowledge fill-in-the-blank question) in the knowledge multiple-choice question into a knowledge short-answer question "What is the surface of the desk in the classroom". In addition, according to the processing rules corresponding to the knowledge fill-in-the-blank question, the keyword "Fill in Chinese" needs to be added before the question stem of the knowledge fill-in-the-blank question to indicate that a complete answer does not need to be output. Since the converted knowledge short-answer question may be relatively broad, there will be multiple answers to the converted knowledge short-answer question (for example, for the above knowledge multiple-choice question, the part in the parentheses in the question stem can be filled with rectangle, or flat, plane, etc.). Therefore, each option of the knowledge multiple-choice question can be supplemented after the converted knowledge short-answer question to limit the output of the converter learning model. Finally, the knowledge multiple-choice question is converted into a target object to be recognized, and the target object to be recognized is expressed as "Fill in Chinese: What is the surface of the desk in the classroom?\nChoose\nRectangle\nSquare\nCircle"; then, input the target object to be recognized into the converter learning model, and the converter learning model can obtain the processing result "Answer: Rectangle"; finally, according to the processing rules corresponding to the knowledge multiple-choice question, select the option that is the same as the processing result from each option in the object to be recognized. In this example, the processing result "Rectangle" is the same as the option "Rectangle", so the serial number "A" corresponding to the option "Rectangle" is used as the target result, and finally the converter learning model outputs "A".

[0084] For example, when the type of the object to be recognized is a knowledge true / false question, the converter learning model converts the object to be recognized into a third intermediate object to be recognized, then deletes each option in the third intermediate object to be recognized, and converts the question stem in the third intermediate object to be recognized into a second intermediate object to be recognized, and finally converts the second intermediate object to be recognized into a target object to be recognized. For example, the type of the third intermediate object to be recognized is a knowledge multiple-choice question, and the type of the second intermediate object to be recognized is a knowledge fill-in-the-blank question. Then, the converter learning model performs recognition processing on the target object to be recognized to obtain the processing result corresponding to the target object to be recognized, then selects the option that is the same as the processing result from each option in the third intermediate object to be recognized, and finally outputs the option that is the same as the processing result as the target result.

[0085] For example, since knowledge-based true / false questions actually represent a choice between right and wrong, each option in the third intermediate object to be recognized can include correct and incorrect (or, right and wrong).

[0086] For example, in some embodiments, for the knowledge-based true / false question "In an acute triangle, the sum of any two interior angles is greater than a right angle. ()", where the stem is "In an acute triangle, the sum of any two interior angles is greater than a right angle". First, convert the stem in this knowledge-based true / false question into a knowledge-based multiple-choice question (i.e., the third intermediate object to be recognized) "In an acute triangle, the sum of any two interior angles being greater than a right angle is (). A. Correct B. Incorrect", where the stem of the converted knowledge-based multiple-choice question is expressed as "In an acute triangle, the sum of any two interior angles being greater than a right angle is ().", which is a knowledge-based fill-in-the-blank question, and each option is "A. Correct B. Incorrect"; then, the converter learning model can delete each option in the converted knowledge-based multiple-choice question and convert the stem in the converted knowledge-based multiple-choice question into a knowledge-based short-answer question "What is the sum of any two interior angles in an acute triangle greater than a right angle?", in addition, according to the processing rules corresponding to the knowledge-based fill-in-the-blank question, the keyword "Fill in Chinese" needs to be added before the stem of the knowledge-based fill-in-the-blank question to indicate that a complete answer does not need to be output. Since the converted knowledge-based short-answer question may be relatively broad and there are multiple answers to the converted knowledge-based short-answer question, the options of the above-converted knowledge-based multiple-choice question can be supplemented after the converted knowledge-based short-answer question to limit the output of the converter learning model. Finally, the knowledge-based true / false question is converted into a target object to be recognized, and the target object to be recognized is expressed as "Fill in Chinese: What is the sum of any two interior angles in an acute triangle greater than a right angle? \nSelect \nCorrect \nIncorrect"; then, input this target object to be recognized into the converter learning model, and the converter learning model can obtain the processing result "Answer: Correct"; finally, according to the processing rules corresponding to the knowledge-based true / false question, select the option that is the same as the processing result from each option in the object to be recognized. In this example, the processing result "Correct" is the same as the option "Correct", so the option "Correct" is used as the target result, and finally the converter learning model outputs "Correct".

[0087] It should be noted that the specific embodiments described above are all illustrative and not used to limit the embodiments of the present disclosure.

[0088] It should be noted that in the embodiments of the present disclosure, each model (for example, a converter learning model, a region recognition model, a character recognition model, a type recognition model, etc.) is not just a mathematical model, but a module that can receive input data, perform data processing, and output processing results. This module can be a software module, a hardware module (such as a hardware neural network), or implemented in a combination of software and hardware. In some embodiments, each model includes code and programs stored in a memory; a processor can execute the code and programs to implement some or all of the functions of each model as described above. In still other embodiments, each model can include a circuit board or a combination of multiple circuit boards for implementing the functions as described above. In some embodiments, the combination of the one circuit board or multiple circuit boards can include: (1) one or more processors; (2) one or more non-transitory computer-readable memories connected to the processors; and (3) firmware stored in the memory that can be executed by the processors.

[0089] It should be understood that in the embodiments of the present disclosure, before obtaining the object to be recognized, the object recognition processing method further includes: a training phase. The training phase includes the process of training a converter learning model, a type recognition model, a region recognition model, a character recognition model, etc. Regarding the training phase, reference can be made to the conventional training process in the art, which will not be elaborated here.

[0090] Corresponding to the above object recognition processing method, at least one embodiment of the present disclosure further provides a processing device, Figure 2 which is a schematic block diagram of a processing device provided by at least one embodiment of the present disclosure.

[0091] For example, as Figure 2 shown, the processing device 200 includes an acquisition module 201, a type recognition module 202, a rule determination module 203, and a converter module 204.

[0092] For example, the acquisition module 201 is used to acquire the object to be recognized.

[0093] The type recognition module 202 is used to recognize the type of the object to be recognized based on the type recognition model.

[0094] The rule determination module 203 is used to determine the processing rule corresponding to the object to be recognized according to the type of the object to be recognized.

[0095] The converter module 204 is used to recognize and process the object to be recognized through the converter learning model according to the processing rule to obtain the target result corresponding to the object to be recognized.

[0096] For example, when the converter module 204 executes the step of identifying and processing the object to be recognized through the converter learning model according to the processing rules to obtain the target result corresponding to the object to be recognized, the following steps are included: in response to the type of the object to be recognized being the basic type, according to the processing rules, the object to be recognized is used as the target object to be recognized; in response to the type of the object to be recognized being a non-basic type, according to the processing rules, the object to be recognized is converted through the converter learning model to convert the object to be recognized into the target object to be recognized; the target object to be recognized is recognized and processed through the converter learning model to obtain the target result corresponding to the object to be recognized. The type of the target object to be recognized is the basic type.

[0097] For example, the acquisition module 201, the type recognition module 202, the rule determination module 203, and / or the converter module 204 include codes and programs stored in the memory; the processor can execute the codes and programs to implement some or all of the functions of the acquisition module 201, the type recognition module 202, the rule determination module 203, and / or the converter module 204 as described above. For example, the acquisition module 201, the type recognition module 202, the rule determination module 203, and / or the converter module 204 can be dedicated hardware devices used to implement some or all of the functions of the acquisition module 201, the type recognition module 202, the rule determination module 203, and / or the converter module 204 as described above. For example, the acquisition module 201, the type recognition module 202, the rule determination module 203, and / or the converter module 204 can be a circuit board or a combination of multiple circuit boards for implementing the functions as described above. In the embodiments of the present application, the combination of the one circuit board or multiple circuit boards may include: (1) one or more processors; (2) one or more non-transitory memories connected to the processor; and (3) firmware stored in the memory and executable by the processor.

[0098] It should be noted that the acquisition module 201 is used to implement Figure 1 the step S10 shown, the type recognition module 202 is used to implement Figure 1 the step S11 shown, the rule determination module 203 is used to implement Figure 1 the step S12 shown, and the converter module 204 is used to implement Figure 1 the step S13 shown. Therefore, the specific description of the acquisition module 201 can refer to the relevant description of the step S10 in the embodiments of the above object recognition processing method, the specific description of the type recognition module 202 can refer to the relevant description of the step S11 in the embodiments of the above object recognition processing method, and the specific description of the rule determination module 203 can refer to the relevant description of the step S12 in the embodiments of the above object recognition processing method Figure 1 shown, and the specific description of the type recognition module 202 can refer to the relevant description of the step S11 in the embodiments of the above object recognition processing method. The specific description of the rule determination module 203 can refer to the relevant description of the step S12 in the embodiments of the above object recognition processing method Figure 1 shown, and the specific description of the rule determination module 203 can refer to the relevant description of the step S12 in the embodiments of the above object recognition processing method Figure 1Regarding the relevant description of step S12 shown, for the specific description of the converter module 204, reference can be made to the embodiments of the above object recognition processing method Figure 1 Regarding the relevant description of step S13 shown. In addition, the processing device can achieve technical effects similar to those of the foregoing object recognition processing method, which will not be elaborated here.

[0099] At least one embodiment of the present disclosure further provides an electronic device Figure 3 It is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0100] For example, as Figure 3 shown, the electronic device includes a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. The processor 1001, the communication interface 1002, and the memory 1003 communicate with each other through the communication bus 1004, and components such as the processor 1001, the communication interface 1002, and the memory 1003 can also communicate through a network connection. The present disclosure does not limit the type and function of the network here.

[0101] For example, the memory 1003 is used to non-transiently store computer-readable instructions. When the processor 1001 is used to run the computer-readable instructions, the object recognition processing method described in any of the above embodiments is implemented when the computer-readable instructions are run by the processor 1001. For the specific implementation of each step of the object recognition processing method and the relevant explanatory content, reference can be made to the embodiments of the object recognition processing method above, which will not be elaborated here.

[0102] For example, the implementation manner in which the processor 1001 executes the program stored on the memory 1003 to implement the object recognition processing method is the same as the implementation manner mentioned in the embodiment part of the foregoing object recognition processing method, which will not be elaborated here either.

[0103] For example, the communication bus 1004 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0104] For example, the communication interface 1002 is used to implement communication between the electronic device and other devices.

[0105] For example, the processor 1001 and the memory 1003 can be set on the server side (or cloud).

[0106] For example, the processor 1001 may control other components in the electronic device to perform desired functions. The processor 1001 may be a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be of the X86 or ARM architecture, etc.

[0107] For example, the memory 1003 may include any combination of one or more computer program products. The computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium. The processor 1001 may run the computer-readable instructions to implement various functions of the electronic device. Various application programs and various data may also be stored in the storage medium.

[0108] For example, for a detailed description of the process of the electronic device performing object recognition processing, reference may be made to the relevant descriptions in the embodiments of the object recognition processing method, and repeated parts will not be elaborated.

[0109] Figure 4 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 4 shown, one or more computer-readable instructions 1101 may be non-temporarily stored on the storage medium 1100. For example, when the computer-readable instructions 1101 are executed by the processor, one or more steps in the object recognition processing method described above may be executed.

[0110] For example, the storage medium 1100 may be applied to the above-mentioned electronic device and / or the processing device 200. For example, the storage medium 1100 may include the memory 1003 in the electronic device.

[0111] For example, for the description of the storage medium 1100, reference may be made to the description of the memory in the embodiments of the electronic device, and repeated parts will not be elaborated.

[0112] Figure 5 Schematic diagram of a hardware environment provided by at least one embodiment of the present disclosure. The electronic device provided by the present disclosure may be applied to an Internet system.

[0113] Utilize Figure 5 The computer system provided in Figure 5 can implement the functions of the image processing device and / or the electronic device involved in the present disclosure. Such a computer system may include a personal computer, a laptop computer, a tablet computer, a mobile phone, a personal digital assistant, smart glasses, a smart watch, a smart ring, a smart helmet, and any smart portable device or wearable device. The specific system in this embodiment explains a hardware platform including a user interface using a functional block diagram. Such a computer device may be a general-purpose computer device or a special-purpose computer device. Both types of computer devices can be used to implement the image processing device and / or the electronic device in this embodiment. The computer system may include any components required to implement the information for image processing described currently. For example, the computer system can be implemented by a computer device through its hardware devices, software programs, firmware, and their combinations. For convenience, Figure 5 In Figure 5 , only one computer device is drawn, but the relevant computer functions for implementing the information required for image processing described in this embodiment can be implemented in a distributed manner by a group of similar platforms, distributing the processing load of the computer system.

[0114] As Figure 5 shown, the computer system may include a communication port 250, connected to which is a network for data communication. For example, the computer system can send and receive information and data through the communication port 250, that is, the communication port 250 can enable the computer system to communicate wirelessly or wireline with other electronic devices to exchange data. The computer system may also include a processor group 220 (i.e., the processor described above) for executing program instructions. The processor group 220 may be composed of at least one processor (e.g., a CPU). The computer system may include an internal communication bus 210. The computer system may include different forms of program storage units and data storage units (i.e., the memory or storage medium described above), such as a hard disk 270, a read-only memory (ROM) 230, a random access memory (RAM) 240, which can be used to store various data files for computer processing and / or communication, and the possible program instructions executed by the processor group 220. The computer system may also include an input / output component 260, and the input / output component 260 is used to implement the input / output data stream between the computer system and other components (e.g., the user interface 280, etc.).

[0115] Generally, the following devices may be connected to the input / output component 260: input devices including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices including, for example, magnetic tapes, hard disks, etc.; and communication interfaces.

[0116] Although Figure 5 a computer system with various devices is shown, it should be understood that it is not required for the computer system to have all the shown devices. Instead, the computer system may have more or fewer devices.

[0117] For the present disclosure, the following points also need to be noted:

[0118] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0119] (2) For clarity, in the drawings used to describe the embodiments of the present invention, the thickness and dimensions of layers or structures are enlarged. It can be understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under the other element, or there may be intermediate elements.

[0120] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0121] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claimed rights.

Claims

1. An object recognition and processing method, comprising: obtaining an object to be recognized; recognizing the type of the object to be recognized based on a type recognition model; determining a processing rule corresponding to the object to be recognized according to the type of the object to be recognized; recognizing and processing the object to be recognized through a converter learning model according to the processing rule to obtain a target result corresponding to the object to be recognized, wherein the object to be recognized is a word problem and the target result is the answer to the word problem, wherein, according to the processing rule, recognizing and processing the object to be recognized through the converter learning model to obtain a target result corresponding to the object to be recognized includes: in response to the type of the object to be recognized being a basic type, taking the object to be recognized as a target object to be recognized according to the processing rule, and in response to the type of the object to be recognized being a non-basic type, converting the object to be recognized through the converter learning model according to the processing rule to convert the object to be recognized into a target object to be recognized, wherein the basic type includes computational solution questions or knowledge-based question-and-answer questions, and the non-basic type includes computational fill-in-the-blank questions, computational true-or-false questions, computational multiple-choice questions, knowledge-based fill-in-the-blank questions, knowledge-based true-or-false questions, or knowledge-based multiple-choice questions; recognizing and processing the target object to be recognized through the converter learning model to obtain a target result corresponding to the object to be recognized, wherein the type of the target object to be recognized is the basic type, wherein obtaining the object to be recognized includes: obtaining an input image, wherein the input image includes the object to be recognized; processing the input image through a region recognition model to obtain a region to be recognized including the object to be recognized; processing the region to be recognized through a character recognition model to obtain the object to be recognized.

2. The object recognition and processing method according to claim 1, wherein, recognizing and processing the target object to be recognized through the converter learning model to obtain a target result corresponding to the object to be recognized includes: recognizing and processing the target object to be recognized through the converter learning model to obtain a processing result corresponding to the target object to be recognized; processing the processing result according to the processing rule to obtain a target result corresponding to the object to be recognized.

3. The object recognition and processing method according to claim 1, wherein, converting the object to be recognized through the converter learning model to convert the object to be recognized into a target object to be recognized includes: in response to the type of the object to be recognized being the computational fill-in-the-blank question or the knowledge-based fill-in-the-blank question, directly converting the object to be recognized into the target object to be recognized through the converter learning model; In response to the type of the object to be recognized being the computational judgment question, the converter learning model deletes the judgment result in the object to be recognized, so as to convert the object to be recognized into a first intermediate object to be recognized, and convert the first intermediate object to be recognized into the target object to be recognized, wherein the type of the first intermediate object to be recognized is the computational fill-in-the-blank question; In response to the type of the object to be recognized being the computational multiple-choice question or the knowledge multiple-choice question, the converter learning model deletes each option in the object to be recognized, and converts the stem in the object to be recognized into the target object to be recognized; In response to the type of the object to be recognized being the knowledge judgment question, the converter learning model converts the object to be recognized into a third intermediate object to be recognized, wherein the type of the third intermediate object to be recognized is the knowledge multiple-choice question, deletes each option in the third intermediate object to be recognized, and converts the stem in the third intermediate object to be recognized into a second intermediate object to be recognized, and converts the second intermediate object to be recognized into the target object to be recognized, wherein the type of the second intermediate object to be recognized is the knowledge fill-in-the-blank question.

4. The object recognition processing method according to claim 3, wherein, determining a processing rule corresponding to the object to be recognized according to the type of the object to be recognized, includes: In response to the type of the object to be recognized being the computational fill-in-the-blank question or the knowledge fill-in-the-blank question, the processing rule corresponding to the object to be recognized includes adding keywords to the target object to be recognized; In response to the type of the object to be recognized being the computational multiple-choice question or the knowledge multiple-choice question, the processing rule corresponding to the object to be recognized includes selecting an option that is the same as the processing result from each option in the object to be recognized; In response to the type of the object to be recognized being the computational judgment question, the processing rule corresponding to the object to be recognized includes comparing the processing result with the judgment result in the object to be recognized; In response to the type of the object to be recognized being the knowledge judgment question, the processing rule corresponding to the object to be recognized includes selecting an option that is the same as the processing result from each option in the third intermediate object to be recognized.

5. The object recognition processing method according to claim 4, wherein, processing the processing result according to the processing rule to obtain a target result corresponding to the object to be recognized, includes: In response to the type of the object to be recognized being the computational solution question or the knowledge essay question, directly outputting the processing result as the target result; In response to the type of the object to be recognized being the computational fill-in-the-blank question or the knowledge fill-in-the-blank question, using the part corresponding to the keyword in the processing result as the target result; In response to the type of the object to be recognized being the computational multiple-choice question or the knowledge multiple-choice question, select the option that is the same as the processing result from each of the options in the object to be recognized, and use the serial number of the option that is the same as the processing result as the target result; In response to the type of the object to be recognized being the computational true / false question, compare the processing result with the judgment result in the object to be recognized to obtain a comparison result, and use the comparison result as the target result, where the comparison result includes correct, incorrect, or pending; In response to the type of the object to be recognized being the knowledge true / false question, select the option that is the same as the processing result from each of the options in the third intermediate object to be recognized, and use the option that is the same as the processing result as the target result.

6. The object recognition processing method according to claim 4, wherein, when the type of the object to be recognized is the computational fill-in-the-blank question, the keyword is "fill in the blank", when the type of the object to be recognized is the knowledge fill-in-the-blank question, the keyword is "fill in Chinese".

7. The object recognition processing method according to claim 1, wherein, the input image is an image of a test paper, and the object to be recognized is a test question on the test paper.

8. The object recognition processing method according to any one of claims 1 to 7, wherein, the converter learning model is a model based on a neural network.

9. A processing device, comprising: an acquisition module for acquiring an object to be recognized; a type recognition module for recognizing the type of the object to be recognized based on a type recognition model; a rule determination module for determining a processing rule corresponding to the object to be recognized according to the type of the object to be recognized; a converter module for recognizing and processing the object to be recognized through a converter learning model according to the processing rule to obtain a target result corresponding to the object to be recognized, where the object to be recognized is a text question and the target result is the answer to the text question, where when the converter module executes the step of recognizing and processing the object to be recognized through the converter learning model according to the processing rule to obtain a target result corresponding to the object to be recognized, it includes executing the following steps: In response to the type of the object to be recognized being a basic type, according to the processing rule, use the object to be recognized as the target object to be recognized. In response to the type of the object to be recognized being a non-basic type, according to the processing rule, convert the object to be recognized through the converter learning model to convert the object to be recognized into a target object to be recognized, where the basic type includes computational solution questions or knowledge essay questions, and the non-basic type includes computational fill-in-the-blank questions, computational true / false questions, computational multiple-choice questions, knowledge fill-in-the-blank questions, knowledge true / false questions, or knowledge multiple-choice questions; Recognize and process the target object to be recognized through the converter learning model to obtain a target result corresponding to the object to be recognized, where the type of the target object to be recognized is the basic type Among them, the obtaining module is further configured to: Obtain an input image, where the input image includes the object to be recognized; Process the input image through a region recognition model to obtain a region to be recognized including the object to be recognized; Process the region to be recognized through a character recognition model to obtain the object to be recognized.

10. An electronic device comprising: A memory for non-transiently storing computer-readable instructions; A processor for running the computer-readable instructions, and when the computer-readable instructions are run by the processor, implementing the object recognition processing method according to any one of claims 1 to 8.

11. A non-transient computer-readable storage medium wherein the non-transient computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, implementing the object recognition processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and system for automatically generating data acquisition module

    CN111369290A