Image scoring method, device, equipment and storage medium based on artificial intelligence
Through the feature extraction and stitching layer of the image scoring model, the problem of failure to effectively pay attention to the cleanliness and details of the entire test paper in the prior art is solved, and the accurate scoring of the text content is achieved.
Patent Information
- Application Number
- CN202110620463.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-03
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2041-06-03
AI Technical Summary
In the intelligent scoring method in the field of education, the method based on font template matching only uses the similarity between the font library and the test paper text, and fails to effectively pay attention to the cleanliness of the entire test paper, resulting in inaccurate evaluation; it is difficult to pay attention to the essential details of the image based on computer vision image classification technology, resulting in inaccurate scoring.
The overall dimensional features of the text image are extracted through the first feature extraction layer of the image scoring model. After chunking processing, the detail dimensional features are extracted by the second feature extraction layer, and the overall and detailed features are fused through the feature stitching layer. Finally, the score prediction layer is used to score and predict, achieving accurate scoring of the text content.
Effectively integrate the overall and detailed features in text images to provide more accurate text content ratings, improving the accuracy of ratings.
Smart Images

Figure CN113822847B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and in particular to an artificial intelligence-based image scoring method, apparatus, device, and computer-readable storage medium. Background Art
[0002] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0003] The introduction of artificial intelligence (AI) technology in education has led to rapid development of intelligent teaching. Intelligent grading can be achieved by recognizing text in scanned images of test papers. However, intelligent grading often only considers the recognized text, requiring the participation of teaching and research personnel in the grading of the test papers.
[0004] Among related test score assessment methods, font template matching, while a simple model, only utilizes the similarity between the font library and the test text, without considering the neatness of the entire test paper. Computer vision image classification technology primarily relies on annotated results to classify and score the entire image, failing to focus on essential image details and resulting in inaccurate test score assessments. Summary of the Invention
[0005] The embodiments of the present application provide an artificial intelligence-based image scoring method, device, and computer-readable storage medium, which can effectively integrate the features of the text content in the text image in the detail dimension and the features of the text content in the text image in the overall dimension, and score the text content more accurately and effectively.
[0006] The technical solution of the embodiment of the present application is implemented as follows:
[0007] The present invention provides an artificial intelligence-based image scoring method, including:
[0008] Performing feature extraction on a text image including text content through a first feature extraction layer of an image scoring model to obtain features of the text content in the text image in an overall dimension;
[0009] Performing block processing on the text image to obtain an image block sequence including at least two image blocks;
[0010] Performing feature extraction on the image block sequence through a second feature extraction layer of an image scoring model to obtain features of the text content in the text image in a detail dimension;
[0011] The features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain splicing features;
[0012] The score prediction layer of the image scoring model is used to predict the score of the splicing feature to obtain a first score corresponding to the text content.
[0013] The present invention provides an artificial intelligence-based image scoring device, comprising:
[0014] A first feature extraction module is configured to perform feature extraction on a text image including text content through a first feature extraction layer of an image scoring model to obtain features of the text content in the text image in an overall dimension;
[0015] An image block module, configured to perform block processing on the text image to obtain an image block sequence including at least two image blocks;
[0016] A second feature extraction module is configured to perform feature extraction on the image block sequence through a second feature extraction layer of an image scoring model to obtain features of the text content in the text image in a detail dimension;
[0017] A feature splicing module is used to perform feature splicing on the features of the text content in the overall dimension and the features of the text content in the detail dimension through a feature splicing layer of an image scoring model to obtain splicing features;
[0018] The rating prediction module is used to perform rating prediction on the splicing feature through the rating prediction layer of the image rating model to obtain a first rating corresponding to the text content.
[0019] In the above solution, the device further includes an image preprocessing module, which is used to perform region positioning on the text image to be rated to determine a target region corresponding to the text content in the text image to be rated;
[0020] Based on the target area, the text image to be rated is segmented to obtain the text image.
[0021] In the above solution, the image preprocessing module is further used to perform region segmentation on the text image to be rated based on the target region to obtain a segmented image including the text content;
[0022] The segmented image is subjected to interference removal processing to obtain the text image.
[0023] In the above solution, the second feature extraction module includes a mapping layer, a context feature extraction layer, and a detail feature extraction layer, and the second feature extraction module is further configured to perform feature mapping on the image block sequence through the mapping layer to obtain an image representation sequence corresponding to the image block sequence;
[0024] Performing context feature extraction on the image representation sequence through the context feature extraction layer to obtain context features;
[0025] The detail feature extraction layer performs detail feature extraction on the context feature to obtain features of the text content in the text image in a detail dimension.
[0026] In the above solution, the mapping layer in the second feature extraction module includes a conversion layer and a position encoding layer, and the second feature extraction module is further used to perform vector conversion on the image block sequence through the conversion layer to obtain corresponding feature vectors;
[0027] Performing position coding on the image block sequence through the position coding layer to obtain a corresponding position coding vector;
[0028] The feature vector and the position encoding vector are added in a position-by-position manner to obtain an image representation sequence corresponding to the image block sequence.
[0029] In the above solution, the context feature extraction layer in the second feature extraction module includes at least two sub-feature extraction layers, and the second feature extraction module is further configured to perform context feature extraction on the image representation sequence through the at least two sub-feature extraction layers to obtain at least two context features;
[0030] Correspondingly, the second feature extraction module is further configured to combine the at least two context features to obtain a combined context feature corresponding to the image representation sequence;
[0031] Detail features are extracted from the combined context features to obtain features of the text content in the text image in a detail dimension.
[0032] In the above solution, the device further includes a model training module, which is used to obtain text image samples including text content and standard scores corresponding to the text image samples;
[0033] Performing feature extraction on the text image sample through the first feature extraction layer of the image scoring model to obtain features of the text content in the text image sample in an overall dimension;
[0034] Performing block processing on the text image sample to obtain an image block sequence including at least two image blocks;
[0035] Performing feature extraction on the image block sequence through a second feature extraction layer of an image scoring model to obtain features of the text content in the text image in a detail dimension;
[0036] The features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain splicing features;
[0037] Using the score prediction layer of the image scoring model, the splicing features are scored and predicted to obtain a predicted score corresponding to the text content;
[0038] Based on the difference between the predicted score and the standard score, model parameters of the image scoring model are updated.
[0039] In the above solution, the rating prediction module is further used to extract the text content of the text image to obtain the target text content;
[0040] Matching the target text content with the standard text content to obtain a matching result, and determining a second score for the target text content based on the matching result;
[0041] Based on the first score and the second score, a comprehensive score corresponding to the text content is determined.
[0042] An embodiment of the present application provides an electronic device, including:
[0043] a memory for storing executable instructions;
[0044] The processor is used to implement the artificial intelligence-based image scoring method provided in the embodiment of the present application when executing the executable instructions stored in the memory.
[0045] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute instructions to implement the artificial intelligence-based image scoring method provided in the embodiment of the present application.
[0046] The embodiments of the present application have the following beneficial effects:
[0047] In the embodiment of the present application, a first feature extraction layer of an image scoring model is used to extract features from a text image including text content, thereby obtaining features of the text content in the text image in the overall dimension, based on which the overall features of the text content can be effectively focused on; the text image is then segmented to obtain an image block sequence including at least two image blocks; and a second feature extraction layer of the image scoring model is used to extract features from the image block sequence, thereby obtaining features of the text content in the text image in the detail dimension, based on which the detail features corresponding to the text content can be effectively extracted; and then a feature splicing layer of the image scoring model is used to splice the features of the text content in the overall dimension and the features of the text content in the detail dimension to obtain splicing features, based on which the overall features and the detail features of the text content can be effectively integrated to provide richer feature information; and a score prediction layer of the image scoring model is used to predict the splicing features to obtain a first score corresponding to the text content. In this way, the features of the text content in the text image in the detail dimension and the features of the text content in the text image in the overall dimension can be effectively integrated to more accurately and effectively score the text content. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 This is a schematic diagram of an optional architecture of an artificial intelligence-based image scoring system provided in an embodiment of the present application;
[0049] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0050] Figure 3 This is an optional flowchart of the artificial intelligence-based image scoring method provided in an embodiment of the present application;
[0051] Figure 4 This is an optional structural diagram of the image scoring model provided in the embodiment of the present application;
[0052] Figure 5 This is an optional schematic diagram of a text image to be rated provided in an embodiment of the present application;
[0053] Figure 6 This is an optional schematic diagram of image segmentation provided in an embodiment of the present application;
[0054] Figure 7 This is an optional schematic diagram of text image segmentation provided in an embodiment of the present application;
[0055] Figure 8 This is an optional structural diagram of the second feature extraction layer provided in an embodiment of the present application;
[0056] Figure 9This is an optional structural diagram of the second feature extraction layer provided in an embodiment of the present application;
[0057] Figure 10 This is an optional structural diagram of the attention mechanism model provided in the embodiment of the present application;
[0058] Figure 11 This is an optional flowchart of the training method of the artificial intelligence-based image scoring model provided in an embodiment of the present application;
[0059] Figure 12 This is an optional flowchart of the artificial intelligence-based image scoring method provided in an embodiment of the present application;
[0060] Figure 13 This is an optional structural diagram of an artificial intelligence-based image scoring system provided in an embodiment of the present application;
[0061] Figure 14 This is an optional flowchart of the artificial intelligence-based image scoring method provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0063] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0064] If similar descriptions of "first / second" appear in the application documents, the following explanation is added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0066] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0067] 1) Transformer Encoder: The Transformer Encoder module uses a multi-attention mechanism instead of the sequential structure of recurrent neural networks, enabling parallel training and global information.
[0068] 2) Vision Transformer Encoder: This combines knowledge from computer vision and natural language processing to segment the original image into blocks, flatten the blocks into sequences, and then feeds the original Transformer Encoder. Finally, a fully connected layer is used to classify the image.
[0069] 3) Visual Geometry Group Network (VGG): The VGG model is a traditional image classification model. It outperforms most other model frameworks in multiple transfer learning tasks. There are multiple versions of the VGG model, each with slightly different network structures. The most common one is VGG16.
[0070] In related technologies, there are two main methods for evaluating the test paper score corresponding to the text content in a text image: one is a technology based on font template matching. Its basic idea is to establish a library of different font templates, and calculate the similarity between the handwritten image on the test paper and the template image one by one by matching the image with the template, so as to score the test paper. Although the technology based on font template matching has a simple model, it only uses the similarity between the font library and the test paper text, and does not refer to the neatness of the entire test paper. In addition, the technology based on template matching also has problems such as slow calculation speed and inaccurate template matching.
[0071] The other method is based on computer vision image classification technology, which uses traditional convolutional neural networks to classify and score the entire scanned image of the test paper. Computer vision image classification technology mainly relies on the results of annotation to classify and score the entire image, making it difficult to focus on the essential details of the image.
[0072] Based on this, the embodiments of the present application provide an artificial intelligence-based image scoring method, device, electronic device and computer-readable storage medium, which can effectively integrate the features of the text content in the text image in the detail dimension and the features of the text content in the text image in the overall dimension, and score the text content more accurately and effectively.
[0073] First, the artificial intelligence-based image scoring system provided in the embodiment of the present application is described. Figure 1 , Figure 1 This is an optional architectural diagram of an AI-based image scoring system provided in an embodiment of the present application. To support an AI-based image scoring application, in the AI-based image scoring system 100, terminals (terminal 400-1 and terminal 400-2 are shown as examples) are connected to a server 200 via a network 300.
[0074] In some embodiments, the terminal can be a laptop computer, a tablet computer, a desktop computer, a smart phone, a dedicated messaging device, a portable gaming device, a smart speaker, a smart watch, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms. The network 300 can be a wide area network or a local area network, or a combination of the two. The terminal (terminal 400-1 and terminal 400-2 are shown as examples) and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0075] In some embodiments, the terminal can be a laptop computer, a tablet computer, a desktop computer, a smart phone, a dedicated messaging device, a portable gaming device, a smart speaker, a smart watch, etc., but is not limited thereto.
[0076] Server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDNs), and big data and artificial intelligence platforms. Network 300 can be a wide area network or a local area network, or a combination of the two. The terminal and the server can be connected directly or indirectly via wired or wireless communication, which is not limited in the embodiments of the present application.
[0077] The terminal (such as terminal 400 - 1 ) is used to send a scoring request to the server 200 to request the server 200 to score the text content in the text image carried in the scoring request.
[0078] Server 200 is configured to parse a text image including text content from a scoring request, perform feature extraction on the text image including text content through a first feature extraction layer of an image scoring model, and obtain features of the text content in the text image in an overall dimension; perform block processing on the text image to obtain an image block sequence including at least two image blocks; perform feature extraction on the image block sequence through a second feature extraction layer of the image scoring model to obtain features of the text content in the text image in a detail dimension; perform feature splicing on the features of the text content in the overall dimension and the features of the text content in the detail dimension through a feature splicing layer of the image scoring model to obtain a spliced feature; perform score prediction on the spliced feature through a score prediction layer of the image scoring model to obtain a first score corresponding to the text content;
[0079] The server 200 is further configured to return a first score for the text content to the terminal;
[0080] The terminal is further configured to receive and present the first score for the text content sent by the server 200 .
[0081] In some embodiments, an image scoring client is provided on the terminal (an image scoring client 410-1 and an image scoring client 410-2 are shown as examples). The user scores the text content in the text image based on the image scoring client. Based on the selection operation of the text image containing text content, a scoring instruction for the text content in the text image is triggered. The image scoring client responds to the scoring instruction and sends a scoring request to the server. After the server parses the text image to be scored containing text content from the scoring request, it extracts features of the text image containing text content through the first feature extraction layer of the image scoring model to obtain the overall dimension of the text content in the text image. features; the text image is divided into blocks to obtain an image block sequence including at least two image blocks; the image block sequence is subjected to feature extraction by the second feature extraction layer of the image scoring model to obtain features of the text content in the text image in the detail dimension; the features of the text content in the overall dimension and the features of the text content in the detail dimension are subjected to feature splicing by the feature splicing layer of the image scoring model to obtain splicing features; the score prediction layer of the image scoring model is subjected to score prediction on the splicing features to obtain a first score corresponding to the text content, and the first score is returned to the image scoring client, which presents the first score for the text content in the text image.
[0082] Next, the electronic device for implementing the above-mentioned artificial intelligence-based image scoring method provided in the embodiment of the present application is described. Figure 2 , Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. In practical applications, the electronic device 500 can be implemented as Figure 1 The terminal or server in the Figure 1 Taking the server 200 shown as an example, an electronic device that implements the artificial intelligence-based image scoring method according to an embodiment of the present application is described. Figure 2 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not shown in FIG. Figure 2 Various buses are labeled as bus system 540 .
[0083] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0084] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0085] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 510.
[0086] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0087] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0088] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0089] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, WiFi, and USB;
[0090] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0091] The input processing module 554 is configured to detect one or more user inputs or interactions from one of the one or more input devices 532 and to translate the detected inputs or interactions.
[0092] In some embodiments, the artificial intelligence-based image scoring device provided in the embodiments of the present application can be implemented in software. Figure 2 An AI-based image scoring device 555 stored in memory 550 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a first feature extraction module 5551, an image segmentation module 5552, a second feature extraction module 5553, a feature splicing module 5554, and a score prediction module 5555. These modules are logical and can be arbitrarily combined or further separated according to the functions they implement. The functions of each module will be described below.
[0093] In other embodiments, the artificial intelligence-based image scoring device provided in the embodiments of the present application can be implemented in hardware. As an example, the artificial intelligence-based image scoring device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the artificial intelligence-based image scoring method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0094] Next, the image scoring method based on artificial intelligence provided by the embodiment of the present application is described. In some embodiments, the image scoring method based on artificial intelligence provided by the embodiment of the present application can be implemented by the terminal or the server alone, or by the terminal and the server in collaboration. Taking the server implementation as an example, see Figure 3 , Figure 3 This is an optional flow chart of the artificial intelligence-based image scoring method provided in the embodiment of the present application, which will be combined with Figure 3 The steps shown are explained.
[0095] In step 101, the server performs feature extraction on a text image including text content through the first feature extraction layer of the image scoring model to obtain features of the text content in the text image in an overall dimension.
[0096] Before using the image scoring model to score text images containing text content, the structure of the image scoring model is first explained. Figure 4 , Figure 4It is an optional structural diagram of the image scoring model provided by the embodiment of the present application. Since the image scoring model provided by the embodiment of the present application needs to pay attention to the features of the text content in the overall dimension and the features of the text content in the detail dimension at the same time, the image scoring model includes a first feature extraction layer (number 1), a second feature extraction layer (number 2), a feature splicing layer (number 3) and a score prediction layer (number 4). Among them, the first feature extraction layer is used to extract the features of the text content in the text image in the overall dimension; the second feature extraction layer is used to extract the features of the text content in the text image in the detail dimension; the feature splicing layer is used to perform feature splicing on the overall features output by the first feature extraction layer and the detail features output by the second feature extraction layer, and obtain splicing features that simultaneously integrate the overall features and detail features; the score prediction layer is used to perform score prediction on the splicing features output by the feature splicing layer to obtain a first score for the text content.
[0097] Based on the aforementioned image scoring model structure, the first feature extraction layer is described. The first feature extraction layer can be a common convolutional neural network model. The first feature extraction layer processes the text image to obtain overall features (global features), including paragraph layout, degree of erasure, line spacing, paragraph spacing, neatness, and text content layout style.
[0098] In actual implementation, the first feature extraction layer can use the VGG16 convolutional neural network model. The VG G16 convolutional neural network model has a total of 16 layers. The specific layer structure can be referred to in existing related technologies and will not be repeated here. The VGG16 convolutional neural network model is used to extract features of the text content in the text image to obtain the overall dimensional features of the text content in the text image.
[0099] In some embodiments, the text image received by the server containing the text content to be rated is often a scanned image containing multiple text content areas. If a target area containing a large amount of text content needs to be rated, the target area containing the text content to be rated must first be located. Specifically, the text image to be rated is region-localized to determine the target area in the text image to be rated corresponding to the text content; based on the target area, the text image to be rated is region-segmented to obtain the text image.
[0100] In actual implementation, the region positioning of the text image to be scored is completed based on the unique collective characteristics of the target area. At the same time, the original scanned image with tilt and perspective distortion can be corrected according to the collective information to obtain a text image that only retains the content of the text to be scored.
[0101] For example, see Figure 5 , Figure 5 This is an optional schematic diagram of a text image to be scored provided in an embodiment of the present application. The text image to be scored is a scanned English composition test paper. The composition area (handwriting area) is located, and the four-point coordinates corresponding to the upper left corner (numbered 1), the lower left corner (numbered 2), the lower right corner (numbered 3), and the upper right corner (numbered 4) are obtained. The composition area is cut according to the four-point coordinate information to obtain a text image corresponding to the composition area to be scored.
[0102] In some embodiments, the text image containing the text content to be scored, obtained through the aforementioned region positioning, often contains some interference information and image scanning noise. Because the first feature extraction layer, i.e., the convolutional neural network model, has certain requirements for the text images to be processed, it is necessary to perform image preprocessing on the text image containing the text content. Specifically, based on the target region, the text image to be scored is segmented to obtain a segmented image containing the text content; and the segmented image is subjected to interference removal processing to obtain the text image.
[0103] In actual implementation, since text images containing text content are often scanned images, in order to reduce the impact of scanning noise and other interference factors, it is necessary to perform image denoising. Methods for image denoising include, but are not limited to, image processing such as denoising and binarization. The embodiments of this application do not limit the specific image denoising methods.
[0104] In step 102, the text image is divided into blocks to obtain an image block sequence including at least two image blocks.
[0105] The reason for dividing the text image into blocks is explained. Figure 4 The second feature extraction layer of the image scoring model is used to extract detailed features of the text content in the text image (referred to as detail features). Detail features primarily refer to the handwriting of the text content, including whether words break, word continuity, repetition, and distance between words. To extract these detail features, a sequence-to-sequence attention mechanism model, such as the Transformer Encoder model, is generally used. This means that the second feature extraction layer typically incorporates a sequence-to-sequence attention mechanism.
[0106] However, in practical applications, sequence-to-sequence models are mostly used to solve natural language processing problems, requiring the input information to be a sequence, and text images including text content (whether color mode images or grayscale mode images) are represented by matrices. Therefore, when using sequence-to-sequence models for image processing, the matrix corresponding to the image needs to be converted into a sequence representation, that is, the text image needs to be divided into blocks to obtain a series of image blocks, and based on the image blocks, the matrix representation of the image is converted into a sequence representation.
[0107] For a description of the conversion method from the matrix representation of an image to a sequence representation, see Figure 6 , Figure 6 This is an optional schematic diagram of image segmentation provided in an embodiment of the present application. After obtaining a complete scanned text image, the text image is first divided into a series of continuous image blocks of the same size (number 1), and then the image sequence blocks (number 2) are formed in order from left to right and from top to bottom. The feature format after mapping is as follows:
[0108]
[0109] In the above feature format, we first transform the image X∈H×W×C into an X P ∈N×(P 2 C) Flattened image block sequence, there are N = HW / P in this image block sequence 2 flattened image blocks, each block has a dimension of (P 2 C), where P is the size of each image block and C is the number of channels. For color images, C is equal to 3, and for grayscale images, C is equal to 1.
[0110] For example, see Figure 6 Taking a scanned image of an English essay as an example, let's assume the text image is a 48*48 grayscale image with 1 channel. It is divided into multiple 16*16 image blocks (numbered 1) and organized into an image block sequence from left to right and top to bottom. The image block sequence (numbered 2) contains nine flattened image blocks, each of which is flattened into a vector with a dimension D of 256 (16*16*1). Finally, we obtain an image representation sequence for the text image, which consists of nine consecutive vectors of dimension 256.
[0111] In some embodiments, the input text image containing text content is usually a text image corresponding to a subjective question on a test paper, such as an English composition, a Chinese composition, or the like. In order to facilitate block processing, the blocks can be directly divided into rows, see Figure 7 , Figure 7This is an optional schematic diagram of text image segmentation provided in an embodiment of the present application. Figure 7 In this example, a grayscale text image containing text content is divided into rows, with each row being divided into an image block. Numbers 1-9 indicate that the text image is divided into nine blocks of the same size. For example, if the block size is 24*16, each block is flattened into a vector of dimension (24*16*1), resulting in an image representation sequence of the text image containing nine consecutive vectors of dimension (24*16*1).
[0112] In step 103, feature extraction is performed on the image block sequence through the second feature extraction layer of the image scoring model to obtain features of the text content in the text image in the detail dimension.
[0113] From the above description, it can be seen that the second feature extraction layer requires the input information to be the image sequence representation corresponding to the text image, and the second feature extraction layer extracts the features of the text content in the detail dimension (detail features), where the detail features mainly refer to the writing situation of the text content, including whether words have line breaks, continuity between words, degree of repetition, distance between words, etc.
[0114] Before extracting the features of the text content in the detail dimension through the second feature extraction layer, the structure of the second feature extraction layer is first described. In some embodiments, see Figure 8 , Figure 8 This is an optional structural diagram of the second feature extraction layer provided in an embodiment of the present application. The second feature extraction layer (No. 1) can be divided into a mapping layer (No. 2), a context feature extraction layer (No. 3) and a detail feature extraction layer (No. 4). The mapping layer is used to perform feature mapping on the image block sequence to obtain an image representation sequence corresponding to the image block sequence; the context feature extraction layer is used to perform context feature extraction on the image representation sequence to obtain context features; the detail feature extraction layer is used to perform detail feature extraction on the context features to obtain features of the text content in the text image in the detail dimension. Among them, the mapping layer (No. 2) can be further divided into a conversion layer and a position coding layer. The conversion layer is used to perform vector conversion on the image block sequence to obtain the corresponding feature vector; the position coding layer is used to perform position coding on the image block sequence to obtain the corresponding position coding vector; then the sub-feature vectors corresponding to each image block in the feature vector and the position coding vector corresponding to each image block are added in a position-by-position manner to obtain the image representation sequence corresponding to the image block sequence.
[0115] In actual implementation, the mapping layer performs a linear transformation on the vector corresponding to the image block; the context feature layer can be a common sequence to sequence attention mechanism model; the detail feature extraction layer can be a common multi-layer feedforward neural network model, which is used to classify or regress the context features output by the attention mechanism model to obtain the corresponding detail features.
[0116] In actual implementation, the context feature extraction layer usually uses an attention mechanism model. For example, it can be an attention mechanism model based on Transformer Encoder; the detail feature extraction layer can be a multi-layer feedforward neural network, also known as a multi-layer perceptron (MLP). Among them, the detail feature extraction layer can use a discrete interval classification method or a continuous regression numerical method to calculate the context features for the continuous context feature sequence output by the context feature extraction layer to obtain the corresponding detail features. The detail feature extraction layer can be a multi-layer perceptron network (MLP) model. The MLP model can be calculated using a continuous regression numerical method or a discrete interval classification method. However, due to the review of the paper score, the final score obtained is discrete data and the overall score is not high, so a discrete interval splitting method can be used for calculation.
[0117] The method of obtaining the image representation sequence is described. In some embodiments, see Figure 9 , Figure 9 This is an optional structural diagram of the second feature extraction layer provided in an embodiment of the present application. For the image block sequence (number 1) input into the second feature extraction layer, since the text image containing text content is directly split into a series of continuous image blocks when the image is segmented, the position information of each image block relative to the original text image is lost. In order to compensate for the lost position information, a mapping layer is set to perform position encoding on the original position information of each image block in the text image, obtain a position encoding vector (Position Embedding) corresponding to each image block, and perform positional addition with the representation vector (PatchEmbedding) of the image block to obtain a feature vector containing position information. For example, number 3 in the figure represents a feature vector with position information obtained after the 9th image block is feature mapped by the mapping layer. According to this method, the above operation is performed on each image block to obtain an image representation sequence containing position information of the text image.
[0118] In actual implementation, a vector e representing position information is usually specified for each position. i , i is a positive integer, vector e iAfter adding the corresponding image block’s Patch Embedding, we get the encoding vector with position information, which will be used in the following calculation process. i It is set manually, not learned by neural network. Each position has a different e i , and e i The dimension of is consistent with that of the Patch Embedding, and is added to the Patch Embedding corresponding to the image block (that is, two vectors with the same dimension and elements at the same position are added).
[0119] For example, Figure 9 In [1], a 48×48×1 target text image (grayscale image) is input. The image block size P is set (e.g., to 16). This divides the image into N P×P×C (16×16×1) blocks, where N = H×W / (P×P) (N = 9). After obtaining the N image blocks, a linear transformation is used to convert them into D-dimensional feature vectors (D = P×P×C, D = 16×16×1 = 256), which are then added with the position encoding vector. The input target text image has three patches per row and three patches per column, each of which is 16*16 in size (containing 256 value vectors). Each image block (Patch) is numbered in row order, and the numbers {1, 2, 3, 4, 5, 6, 7, 8, 9} are used as the input of the position coding layer, and a D-dimensional position coding vector is output. That is, the dimension of the position coding vector corresponding to the image block is the same as the dimension of the feature vector corresponding to the image block. Then the position coding vector and the feature vector of the corresponding Patch are added in place to obtain the feature vector of each image block with position information.
[0120] The context feature extraction layer is described. In some embodiments, see Figure 9 In
[15] , the context feature extraction layer can be a multi-head attention mechanism model based on Transformer Encoder (No. 5). The input information of this layer is the image representation sequence composed of the feature vectors of each image block with position information.
[0121] The essence of the attention mechanism is a means of filtering out high-value information from a large amount of information. Among a large amount of information, the importance of different information to the result is different. This importance can be reflected by assigning attention weights of different sizes. In other words, the attention mechanism can be understood as a mechanism for assigning weights when synthesizing multiple inputs.
[0122] See also Figure 10 , Figure 10This is an optional structural diagram of the attention mechanism model provided in an embodiment of the present application. The Transformer Encoder attention mechanism model is mainly composed of multi-head self-attention (Multi-Head Attention) and a multi-layer perception network or a multi-layer feedforward neural network (MLP) (two layers of fully connected networks using GE LU activation function). Layer Norm (normalization or standardization layer, Norm layer) and residual connection (Add layer) are added before Multi-Head Attention and MLP. Add means residual connection (Residual Connection) is used to prevent network degradation, and Norm means Layer Normalization, which is used to normalize the activation values of each layer.
[0123] In some embodiments, the context feature extraction layer includes multiple sub-feature extraction layers, and the structure of each sub-feature extraction layer is as follows: Figure 10 As shown, the multi-head attention mechanism model includes multiple attention sub-networks with different network parameters. The network parameters in each attention sub-network are used to characterize the importance (or influence) of each context feature element on the score of the text content from different angles. The image representation sequence is input into each attention sub-network respectively, and the outputs of all attention sub-networks are spliced to obtain the feature vector corresponding to the text content in the text image. Specifically, the image representation sequence is subjected to context feature extraction by the at least two sub-feature extraction layers to obtain at least two context features; correspondingly, the context feature is subjected to detail feature extraction by the detail feature extraction layer to obtain the features of the text content in the text image in the detail dimension, including: combining the at least two context features to obtain the combined context feature corresponding to the image representation sequence; and extracting the detail feature of the combined context feature to obtain the features of the text content in the text image in the detail dimension.
[0124] By adopting a multi-head attention mechanism model, weights can be set for the feature elements of the text content in different detail dimensions from multiple different angles. For example, taking the text image corresponding to an English composition as an example, the detail feature corresponding to the degree of adhesion between words can be used to set the weight of the detail feature when scoring the text content (i.e., evaluating the English composition paper score) (reflecting the influence of the detail feature of the degree of adhesion between words in evaluating the English composition paper score). Or, weights can be set based on the detail feature of inconsistent handwriting size. Or, weights can be set based on the relationship between a segmented word and other segmented words in the context. From different angles, comprehensively considering the influence of the corresponding weights of various contextual features in multiple different dimensions can improve the accuracy of text content scoring (i.e., the accuracy of evaluating the English composition paper score).
[0125] In step 104, the features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain spliced features.
[0126] Before performing feature splicing, it is necessary to receive a first notification message and a second notification message, wherein the first notification message is used to indicate the features of the text content in the overall dimension (overall features) received in the text image output by the first feature extraction layer, and the second notification message is used to indicate the features of the text content in the detail dimension (detail features) received in the text image output by the second feature extraction layer; after receiving the first notification message and the second notification message, the overall features and detail features obtained by splicing are obtained to obtain the splicing features.
[0127] In actual implementation, the obtained splicing features combine the detailed features of the text content in the text image with the overall features of the text content in the text image. Figure 5 The splicing feature combines the overall neatness of the text image and the writing of the text content in the text image with the text image corresponding to the English composition (or the text image corresponding to other common subjective questions). It provides richer feature information for evaluating the paper score of the text image and can more effectively improve the accuracy of the paper score evaluation.
[0128] In step 105, a score prediction layer of the image scoring model is used to perform score prediction on the splicing features to obtain a first score corresponding to the text content.
[0129] Here, the first score is derived from a fusion of detailed features of the text content in the text image and overall features of the text content in the text image. If the text image corresponds to a subjective question in a common test paper, the first score is the score for the subjective question.
[0130] In some embodiments, the rating prediction layer can be one or more fully connected layers, which are used to perform dimensionality reduction processing on the splicing features, and input the splicing features after dimensionality reduction into a classifier, which outputs the rating level to which the rating corresponding to the text content in the text image belongs or the probability of matching a certain rating level. Among them, the rating level is pre-divided into different score segments before the rating prediction is performed, and each score segment corresponds to a rating level. The embodiment of the present application does not limit the form of division of the rating levels.
[0131] In actual implementation, a softmax classifier can be used to give a score corresponding to the text content in the text image. The score is based on the features of the text content in the text image in the detail dimension and the features of the text content in the text image in the overall dimension, and is more accurate.
[0132] In some embodiments, see Figure 11 , Figure 11 This is an optional flow chart of the training method of the artificial intelligence-based image scoring model provided in the embodiment of the present application, based on Figure 3 Before step 101, it is necessary to train the image scoring model to obtain the trained image scoring model. Figure 11 The steps shown are explained.
[0133] Step 201: The server obtains a text image sample including text content and a standard score corresponding to the text image sample.
[0134] Step 202 : Perform feature extraction on the text image sample through the first feature extraction layer of the image scoring model to obtain features of the text content in the text image sample in the overall dimension.
[0135] Step 203 : performing block processing on the text image sample to obtain an image block sequence including at least two image blocks.
[0136] Step 204 : Perform feature extraction on the image block sequence through the second feature extraction layer of the image scoring model to obtain features of the text content in the text image in the detail dimension.
[0137] In step 205 , the features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain spliced features.
[0138] Step 206 : Using the score prediction layer of the image scoring model, score prediction is performed on the splicing features to obtain a predicted score for the corresponding text content.
[0139] Step 207 : updating the model parameters of the image scoring model based on the difference between the predicted score and the standard score.
[0140] Repeat steps 202 to 207 until the image scoring model reaches a convergence condition, thereby obtaining a trained image scoring model. Here, the convergence condition may be the convergence of the model loss function, the convergence of the model parameters, the reaching of the maximum number of iterations, or the reaching of the maximum training time.
[0141] In some embodiments, a text image including text content is input into the aforementioned trained image scoring model to obtain a first score corresponding to the text content. If the text image is a text image corresponding to a subjective question in the test paper, the first score is the score of the subjective question. In order to obtain a comprehensive score corresponding to the text content, it is necessary to obtain a second score corresponding to the text content determined based on the difference between the text content and the standard answer information. Here, the second score refers to the score of whether the answer to the subjective question is correct. The final score corresponding to the subjective question is then determined based on the first score (the score) and the second score (the score of whether the answer is correct). Specifically, the text content of the text image is extracted to obtain the target text content; the target text content is matched with the standard text content to obtain a matching result, and based on the matching result, the second score of the target text content is determined; based on the first score and the second score, the comprehensive score corresponding to the text content is determined.
[0142] First, the method for extracting text content from a text image is explained. It should be noted that there is not a single method for extracting text content. There are multiple methods, including but not limited to the following: extracting text content from a text image using technologies such as optical character recognition (OCR); or extracting text content from a text image using semantic segmentation methods based on deep learning. The text content extracted from the text image is used as the target text content. For example, if the text image is a scanned image of subjective questions on an exam paper, the target text content generally refers to handwritten text content.
[0143] Then, a second score is determined based on the acquired target text content. The second score is the score obtained by matching the text content in the text image with the standard text content (also known as standard answer information). In some embodiments, a specific method for matching the target text content with the standard text content may include, through a semantic feature extraction network, extracting a first semantic vector corresponding to the target text content and a second semantic vector corresponding to the standard text content, respectively, and performing a similarity calculation on the first semantic vector and the second semantic vector. Based on the similarity calculation result, a second score for the target text content is completed (i.e., the higher the similarity, the more similar the text content is to the standard text content, and the higher the second score for the corresponding target text content).
[0144] In other embodiments, the second score corresponding to the target text content obtained can also be obtained by obtaining and calculating the dimension scores of the target text content under each scoring dimension, wherein the scoring dimension can include one or more combinations of the topic-related dimension, the sentence fluency dimension, the literary dimension, and the intention dimension, and then determining the second score of the target text content based on the calculated dimension scores. Among them, the specific method of determining the second score based on the dimension scores can include accumulating the obtained dimension scores and using the obtained accumulated score as the second score of the target text content; or setting corresponding weights for different scoring dimensions, and determining the second score of the target text content based on the obtained dimension scores and the weights of the scoring dimensions corresponding to the dimension scores. In the embodiments of the present application, there is no limitation on the method of determining the second score of the target text content.
[0145] After obtaining the first score and the second score of the target text content, determine the comprehensive score of the target text content. It should be noted that the method for determining the comprehensive score of the target text content can be selected and set according to different application scenarios or different needs. In some embodiments, the first score and the second score obtained can be directly summed, and the sum score can be used as the comprehensive score of the text content. In other embodiments, the weight corresponding to the first score and the weight corresponding to the second score can also be pre-set, and then the comprehensive score of the text content is determined based on the first score, the weight corresponding to the first score, the second score, and the weight corresponding to the second score. For example, the first score is A, the weight corresponding to the first score is wa, the second score is B, and the weight corresponding to the second score is w b , the comprehensive score is w a ×A+w b × B. This embodiment of the present application does not limit the method for determining the comprehensive score of the target text content.
[0146] In summary, the embodiment of the present application performs feature extraction on a text image including text content through the first feature extraction layer of the image scoring model to obtain the features of the text content in the text image in the overall dimension, based on which the overall features of the text image can be effectively focused on; then the text image is segmented to obtain an image block sequence including at least two image blocks; and the image block sequence is subjected to feature extraction through the second feature extraction layer of the image scoring model to obtain the features of the text content in the text image in the detail dimension, based on which the detail features in the text image can be effectively extracted; then, through the feature splicing layer of the image scoring model, the features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced to obtain spliced features, based on which the contextual information in the text image and the overall features of the text image can be simultaneously integrated and richer feature information can be provided; through the score prediction layer of the image scoring model, the spliced features are scored and predicted to obtain the first score of the corresponding text content. In this way, the features of the text content in the text image in the detail dimension and the features of the text content in the text image in the overall dimension can be effectively integrated, and the text content can be scored more accurately and effectively.
[0147] Next, we will continue to introduce the image scoring method based on artificial intelligence provided by the embodiment of this application. Figure 12 This is an optional flow chart of the artificial intelligence-based image scoring method provided in the embodiment of the present application, see Figure 12 The artificial intelligence-based image scoring method provided in the embodiment of the present application is implemented collaboratively by the client and the server. Figure 12 The steps shown are explained.
[0148] Step 301: The terminal sends a rating request carrying a text image to be rated to a server.
[0149] In step 302 , the server parses the received rating request and obtains a text image to be rated containing text content.
[0150] In step 303 , the server performs region positioning on the obtained text image to be rated, and determines a target region corresponding to the text content in the text image to be rated.
[0151] Step 304 : Based on the determined target region, the text image to be rated is segmented to obtain a target text image.
[0152] Here, the target text image contains only the text content that needs to be scored.
[0153] In step 305 , the server obtains the trained image scoring model and inputs the target text image into the first feature extraction layer of the image scoring model.
[0154] Step 306 : Perform feature extraction on the target text image through the first feature extraction layer of the image scoring model to obtain overall features of the text content in the target text image.
[0155] Step 307 : performing block processing on the target text image to obtain an image block sequence including at least two image blocks.
[0156] Here, the target text image is output through step 304 .
[0157] Step 308 : Perform feature mapping on the image block sequence through the mapping layer of the image scoring model to obtain an image representation sequence corresponding to the image block sequence.
[0158] Step 309 : extract context features from the image representation sequence through the context feature extraction layer of the image scoring model to obtain context features.
[0159] In step 310 , detail feature extraction is performed on the context features through the detail feature extraction layer of the image scoring model to obtain features of the text content in the text image in the detail dimension.
[0160] Step 311 , through the feature splicing layer of the image scoring model, the features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced to obtain spliced features.
[0161] Step 312: Use the score prediction layer of the image scoring model to perform score prediction on the splicing features to obtain a first score corresponding to the text content.
[0162] In step 313 , the server sends the first score corresponding to the text content to the terminal.
[0163] In step 314 , the terminal receives a first score corresponding to the text content and displays the first score.
[0164] It should be noted that step 306 extracts features of the text content in the text image in the overall dimension, while steps 307 to 310 extract features of the text content in the text image in the detail dimension. Step 306 and steps 307-310 are two parallel processes and are independent of each other. There is no strict order in which steps 306 and steps 307-310 must be executed.
[0165] The embodiment of the present application obtains a splicing feature with richer features by fusing the overall features of the text content in the target text image in the overall dimension and the detail features of the text content in the target text image in the detail dimension. The text content is scored based on the splicing features, which can accurately evaluate the writing condition of the text content and the neatness of the paper, and thus more accurately and effectively evaluate the paper score corresponding to the text content, and can greatly reduce the manpower cost of teaching and research personnel and accelerate the process of intelligent paper score evaluation.
[0166] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0167] At present, the intelligent scoring of scanned images of test papers in the education field often only refers to the text results after recognition, and teaching and research personnel are required to participate in the scoring process of the paper scores.
[0168] Among related exam paper scoring technologies, font template matching, while a simple model, only utilizes the similarity between the font library and the exam text, without considering the neatness of the entire exam paper. Furthermore, template matching-based technologies suffer from slow computational speeds and inaccurate template matching. Computer vision-based image classification technologies primarily rely on annotated results to classify and score the entire image, making it difficult to focus on essential image details.
[0169] Based on this, the embodiment of the present application proposes a Transformer-based image scoring algorithm, which integrates the Transformer and convolutional neural network to comprehensively score the paper score, and comprehensively considers the scoring dimensions of the paper score, including the overall dimension and the detail dimension. Among them, the overall dimension includes paragraph layout, degree of correction, line spacing, word size, etc.; the detail dimension includes whether a word has a line break, the continuity between words, the degree of repetition, and the distance between words. Obtain reliable results and save labor costs.
[0170] The embodiment of the present application is optimized based on computer vision image classification. By integrating the Transformer module and the Resnet Block module (the residual network corresponding to the Resnet Block is used in the Transformer module) to build a model, it can not only effectively extract the context information (detail features) in the image, but also pay attention to the overall features of the image. The embodiment of the present application can more accurately predict the student's test score through the improved artificial intelligence-based image scoring model. Figure 13 , Figure 13 This is an optional structural diagram of an artificial intelligence-based image scoring system provided in an embodiment of the present application, combined with Figure 13It can be seen that the entire process of using the image scoring model to score the scanned images of students' answers includes the image preprocessing process and the image scoring model processing process. For the specific processing process, see Figure 14 , Figure 14 This is an optional flow chart of the artificial intelligence-based image scoring method provided in the embodiment of the present application. Figure 14 The steps shown are explained.
[0171] In step 401, the client collects a scanned image of the student's answer and sends the scanned image of the student's answer as input information to the server.
[0172] In step 402, the server receives a scanned image of the student's answer and pre-processes the image to obtain a target text image.
[0173] The image preprocessing method is explained below. First, template matching is performed on the scanned image of the student's answer sheet based on the test paper template. The image is then segmented based on the student's subjective answer area marked on the template. The coordinates of the top left, bottom left, bottom right, and top right corners of the subjective answer area are extracted. Based on these four coordinates, the text image of the answer area is determined. De-noising and binarization are then performed on the text image to reduce the impact of scanning noise, resulting in the target text image.
[0174] Step 403: Input the target text image into the first feature extraction layer to perform feature extraction to obtain full image features.
[0175] Here, the first feature extraction layer performs image segmentation on the target text image to obtain the overall features of the target text image. The first feature extraction layer can be a common convolutional neural network structure. In actual implementation, the basic VGG16 model network can be used to extract features from the target text image to obtain the full-image features (overall features) corresponding to the target text image. Full-image features (overall features) generally refer to the neatness of the entire surface of the target text image.
[0176] Step 404 : performing image block processing on the target text image to obtain an image block sequence including at least two image blocks.
[0177] When scoring the target text image, the writing conditions of the target text also need to be considered, such as the adhesion between words in the target text, whether a single word spans across lines in the case of English words, and whether the characters or letters in a word vary in size. Based on this, the second feature extraction layer based on the Transformer Encoder model is used to extract the contextual features of the target text image to obtain the corresponding contextual features. However, in actual applications, the Transformer Encoder is used to solve natural language processing (NLP) problems, requiring the input information to be a sequence of words, while the target text image (whether a color image or a grayscale image) is represented by a matrix. When using the Transformer for image processing, the matrix needs to be converted into a sequence.
[0178] For a detailed description of the matrix-to-sequence conversion method, see Figure 6 To obtain a complete target text image, first divide the target text image into a series of continuous image blocks of the same size (number 1), and then form image sequence blocks (number 2) in order from left to right and from top to bottom. The feature format after mapping is as follows:
[0179]
[0180] In the above feature format, we first transform the image X∈H×W×C into an X P ∈N×(P 2 C) Flattened image block sequence, there are N = HW / P in this image block sequence 2 flattened image blocks, each block has a dimension of (P 2 C), where P is the size of each image block and C is the number of channels. For color images, C is equal to 3, and for grayscale images, C is equal to 1.
[0181] For example, see Figure 6 , set the target text image to be a grayscale image of size 48*48, with the number of channels being 1, divided into multiple image blocks of size 16*16 (numbered 1), and forming an image block sequence from left to right and from top to bottom. There are a total of 9 flattened image blocks in the image block sequence (numbered 2), and each image block is flattened into a vector with a dimension of 256 (16*16*1).
[0182] Step 405: Input the image block sequence into the second feature extraction layer to perform context feature extraction to obtain context features.
[0183] Here, the second feature extraction layer adopts the Transformer Encoder model structure, see Figure 10A Transformer Encoder model can contain multiple encoding layer encoders. As shown in the figure, "Lx" represents the number of encoding layers L. The Transformer Encoder uses a multi-head attention mechanism to replace the original single-head attention, which has a stronger feature fusion effect.
[0184] In step 406 , the obtained context features are input to a detail feature extraction layer to extract detail features and obtain features of the text content in the target text image in the detail dimension.
[0185] Here, the detail feature extraction layer can be a multi-layer perceptron network (MLP) model. The MLP model can perform calculations using a continuous regression numerical method or a discrete interval classification method. However, since the review is based on the paper score, the final score is discrete data and the overall score is not high. Therefore, a discrete interval splitting method can be used for calculation to obtain the corresponding detail features.
[0186] It should be noted that steps 404 to 406 are the detailed process of extracting context features from the target text image based on the Transformer Encoder model, and are performed in parallel with the process of extracting full-image features through a convolutional neural network (VGG) in step 403.
[0187] In step 407 , the server receives the first notification message and the second notification message, and performs feature splicing on the obtained overall features and detail features to obtain spliced features.
[0188] Before performing feature splicing, it is necessary to first receive a first notification message and a second notification message, wherein the first notification message is used to indicate that the overall features of the text content in the target text image have been received (i.e., the output of step 403), and the second notification message is used to indicate that the detailed features of the text content in the target text image have been received (i.e., the output of step 406); after receiving the first notification message and the second notification message, the overall features and detailed features are spliced together to obtain the spliced features.
[0189] Step 408: Output the obtained splicing features to the score prediction layer to perform paper score prediction and obtain the final paper score.
[0190] Step 409: Send the final paper score to the client.
[0191] The automatic scoring of paper scores provided in the embodiment of the present application uses scanned images of students' answer results as training samples to train a constructed image scoring model to obtain a trained image scoring model. The model takes into account both the contextual information in the image and the overall characteristics of the image. The model is used to automatically score the students' answers, which can accurately evaluate the students' writing conditions and the neatness of the paper, and thus more accurately and effectively evaluate the students' paper scores, and can greatly reduce the labor costs of teaching and research personnel and accelerate the intelligent scoring process.
[0192] The following continues to describe the exemplary structure of the artificial intelligence-based image scoring device 555 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the artificial intelligence-based image scoring device 555 of the memory 550 may include:
[0193] A first feature extraction module 5551 is configured to perform feature extraction on a text image including text content through a first feature extraction layer of an image scoring model to obtain features of the text content in the text image in an overall dimension;
[0194] An image block module 5552 is configured to perform block processing on the text image to obtain an image block sequence including at least two image blocks;
[0195] A second feature extraction module 5553 is configured to perform feature extraction on the image block sequence through a second feature extraction layer of an image scoring model to obtain features of the text content in the text image in a detail dimension;
[0196] A feature splicing module 5554 is configured to perform feature splicing on the features of the text content in the overall dimension and the features of the text content in the detail dimension through a feature splicing layer of an image scoring model to obtain spliced features;
[0197] The rating prediction module 5555 is used to perform rating prediction on the splicing feature through the rating prediction layer of the image rating model to obtain a first rating corresponding to the text content.
[0198] In some embodiments, the apparatus further comprises an image preprocessing module, the image preprocessing module being configured to perform region positioning on the text image to be rated, so as to determine a target region in the text image to be rated corresponding to the text content;
[0199] Based on the target area, the text image to be rated is segmented to obtain the text image.
[0200] In some embodiments, the image preprocessing module is further configured to perform region segmentation on the text image to be rated based on the target region to obtain a segmented image including the text content;
[0201] The segmented image is subjected to interference removal processing to obtain the text image.
[0202] In some embodiments, the second feature extraction module includes a mapping layer, a context feature extraction layer, and a detail feature extraction layer, and the second feature extraction module is further configured to perform feature mapping on the image block sequence through the mapping layer to obtain an image representation sequence corresponding to the image block sequence;
[0203] Performing context feature extraction on the image representation sequence through the context feature extraction layer to obtain context features;
[0204] The detail feature extraction layer performs detail feature extraction on the context feature to obtain features of the text content in the text image in a detail dimension.
[0205] In some embodiments, the mapping layer in the second feature extraction module includes a conversion layer and a position encoding layer, and the second feature extraction module is further configured to perform vector conversion on the image block sequence through the conversion layer to obtain a corresponding feature vector;
[0206] Performing position coding on the image block sequence through the position coding layer to obtain a corresponding position coding vector;
[0207] The feature vector and the position encoding vector are added in a position-by-position manner to obtain an image representation sequence corresponding to the image block sequence.
[0208] In some embodiments, the context feature extraction layer in the second feature extraction module includes at least two sub-feature extraction layers, and the second feature extraction module is further configured to perform context feature extraction on the image representation sequence through the at least two sub-feature extraction layers to obtain at least two context features;
[0209] Correspondingly, the second feature extraction module is further configured to combine the at least two context features to obtain a combined context feature corresponding to the image representation sequence;
[0210] Detail features are extracted from the combined context features to obtain features of the text content in the text image in a detail dimension.
[0211] In some embodiments, the apparatus further comprises a model training module, the model training module being configured to obtain a text image sample including text content and a standard score corresponding to the text image sample;
[0212] Performing feature extraction on the text image sample through the first feature extraction layer of the image scoring model to obtain features of the text content in the text image sample in an overall dimension;
[0213] Performing block processing on the text image sample to obtain an image block sequence including at least two image blocks;
[0214] Performing feature extraction on the image block sequence through a second feature extraction layer of an image scoring model to obtain features of the text content in the text image in a detail dimension;
[0215] The features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain splicing features;
[0216] Using the score prediction layer of the image scoring model, the splicing features are scored and predicted to obtain a predicted score corresponding to the text content;
[0217] Based on the difference between the predicted score and the standard score, model parameters of the image scoring model are updated.
[0218] In some embodiments, the rating prediction module is further configured to extract text content from the text image to obtain target text content;
[0219] Matching the target text content with the standard text content to obtain a matching result, and determining a second score for the target text content based on the matching result;
[0220] Based on the first score and the second score, a comprehensive score corresponding to the text content is determined.
[0221] The embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the artificial intelligence-based image scoring method provided by the embodiment of the present application, for example, Figure 3 The method shown.
[0222] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.
[0223] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0224] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (for example, files storing one or more modules, subroutines, or code portions).
[0225] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.
[0226] To sum up, the embodiments of the present application can effectively integrate the features of the text content in the text image in the detail dimension and the features of the text content in the text image in the overall dimension, score the text content more accurately and effectively, and significantly reduce the manpower cost of teaching and research personnel and accelerate the intelligent grading process.
[0227] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. An image scoring method based on artificial intelligence, characterized in that: The method comprises: Performing feature extraction on a text image including text content through a first feature extraction layer of an image scoring model to obtain features of the text content in the text image in an overall dimension, wherein the overall dimension refers to an overall layout style of the text content in the text image; Performing block processing on the text image to obtain an image block sequence including at least two image blocks; Performing feature mapping on the image block sequence through a mapping layer of the image scoring model to obtain an image representation sequence corresponding to the image block sequence, wherein the image representation sequence includes original position information of each image block relative to the text image; Performing context feature extraction on the image representation sequence through a context feature extraction layer of the image scoring model to obtain context features; Performing detail feature extraction on the context features through the detail feature extraction layer of the image scoring model to obtain features of the text content in the text image in a detail dimension, wherein the detail dimension refers to the writing condition of the text content in the text image; The features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain splicing features; The score prediction layer of the image scoring model is used to predict the score of the splicing features to obtain a first score corresponding to the text content, wherein the first score is the paper score of the text content.
2. The method according to claim 1, wherein Before extracting features from the text image including text content through the first feature extraction layer of the image scoring model, the method further includes: Performing region positioning on the text image to be rated to determine a target region corresponding to the text content in the text image to be rated; Based on the target area, the text image to be rated is segmented to obtain the text image.
3. The method according to claim 2, wherein The step of performing region segmentation on the text image to be rated based on the target region to obtain the text image includes: Based on the target area, the text image to be scored is segmented to obtain a segmented image including the text content; The segmented image is subjected to interference removal processing to obtain the text image.
4. The method according to claim 1, wherein The mapping layer includes a conversion layer and a position encoding layer. The mapping layer performs feature mapping on the image block sequence to obtain an image representation sequence corresponding to the image block sequence, including: Performing vector conversion on the image block sequence through the conversion layer to obtain corresponding feature vectors; Performing position coding on the image block sequence through the position coding layer to obtain a corresponding position coding vector; The feature vector and the position encoding vector are added in a position-by-position manner to obtain an image representation sequence corresponding to the image block sequence.
5. The method according to claim 1, wherein The context feature extraction layer includes at least two sub-feature extraction layers, and the context feature extraction is performed on the image representation sequence through the context feature extraction layer to obtain context features, including: Performing context feature extraction on the image representation sequence respectively through the at least two sub-feature extraction layers to obtain at least two context features; The step of extracting detail features from the context features through the detail feature extraction layer to obtain features of the text content in the text image in a detail dimension includes: combining the at least two context features to obtain a combined context feature corresponding to the image representation sequence; Detail features are extracted from the combined context features to obtain features of the text content in the text image in a detail dimension.
6. The method according to claim 1, wherein The method further comprises: Extracting the text content of the text image to obtain target text content, where the target text content is the user's answer; Matching the target text content with the standard text content to obtain a matching result, and determining a second score for the target text content based on the matching result, wherein the standard text content is standard answer information and the second score is a correctness score for the target text content; Based on the first score and the second score, a comprehensive score corresponding to the text content is determined.
7. The method according to claim 1, wherein Before extracting features from the text image including text content through the first feature extraction layer of the image scoring model, the method further includes: Obtaining a text image sample including text content and a standard score corresponding to the text image sample; Performing feature extraction on the text image sample through the first feature extraction layer of the image scoring model to obtain features of the text content in the text image sample in an overall dimension; Performing block processing on the text image sample to obtain an image block sequence including at least two image blocks; Performing feature extraction on the image block sequence through a second feature extraction layer of an image scoring model to obtain features of the text content in the text image in a detail dimension; The features of the text content in the overall dimension and the features of the text content in the detail dimension are spliced together through the feature splicing layer of the image scoring model to obtain splicing features; Using the score prediction layer of the image scoring model, the splicing features are scored and predicted to obtain a predicted score corresponding to the text content; Based on the difference between the predicted score and the standard score, model parameters of the image scoring model are updated.
8. An image scoring device based on artificial intelligence, characterized in that: The device comprises: a first feature extraction module configured to perform feature extraction on a text image including text content using a first feature extraction layer of an image scoring model to obtain features of the text content in the text image in an overall dimension, wherein the overall dimension refers to an overall layout style of the text content in the text image; An image block module, configured to perform block processing on the text image to obtain an image block sequence including at least two image blocks; a second feature extraction module, configured to perform feature mapping on the image block sequence through the mapping layer of the image scoring model to obtain an image representation sequence corresponding to the image block sequence, wherein the image representation sequence includes original position information of each image block relative to the text image; perform context feature extraction on the image representation sequence through the context feature extraction layer of the image scoring model to obtain context features; perform detail feature extraction on the context features through the detail feature extraction layer of the image scoring model to obtain features of the text content in the text image in a detail dimension, wherein the detail dimension refers to the writing condition of the text content in the text image; A feature splicing module, configured to perform feature splicing on the features of the text content in the overall dimension and the features of the text content in the detail dimension through the feature splicing layer of the image scoring model to obtain splicing features; The scoring prediction module is used to predict the score of the splicing features through the scoring prediction layer of the image scoring model to obtain a first score corresponding to the text content, wherein the first score is the paper score of the text content.
9. The device according to claim 8, characterized in that Before extracting features from the text image including text content through the first feature extraction layer of the image scoring model, the method further includes: An image preprocessing module is used to perform region positioning on the text image to be rated, so as to determine a target region corresponding to the text content in the text image to be rated; The image preprocessing module is further configured to perform region segmentation on the text image to be rated based on the target region to obtain the text image.
10. The device according to claim 9, characterized in that The step of performing region segmentation on the text image to be rated based on the target region to obtain the text image includes: The image preprocessing module is further configured to perform region segmentation on the text image to be rated based on the target region to obtain a segmented image including the text content; The image preprocessing module is further used to perform interference removal processing on the segmented image to obtain the text image.
11. The device according to claim 8, characterized in that The mapping layer includes a conversion layer and a position encoding layer. The mapping layer performs feature mapping on the image block sequence to obtain an image representation sequence corresponding to the image block sequence, including: The second feature extraction module is further configured to perform vector conversion on the image block sequence through the conversion layer to obtain corresponding feature vectors; The second feature extraction module is further configured to perform position encoding on the image block sequence through the position encoding layer to obtain a corresponding position encoding vector; The second feature extraction module is further configured to perform position-wise addition of the feature vector and the position encoding vector to obtain an image representation sequence corresponding to the image block sequence.
12. The device according to claim 8, characterized in that The context feature extraction layer includes at least two sub-feature extraction layers, and the context feature extraction is performed on the image representation sequence through the context feature extraction layer to obtain context features, including: The second feature extraction module is further configured to perform context feature extraction on the image representation sequence through the at least two sub-feature extraction layers to obtain at least two context features; The step of extracting detail features from the context features through the detail feature extraction layer to obtain features of the text content in the text image in a detail dimension includes: The second feature extraction module is further configured to combine the at least two context features to obtain a combined context feature corresponding to the image representation sequence; The second feature extraction module is further configured to perform detail feature extraction on the combined context feature to obtain features of the text content in the text image in a detail dimension.
13. A computer-readable storage medium, characterized in that Executable instructions are stored for implementing the method for image scoring based on artificial intelligence as described in any one of claims 1 to 7 when executed by a processor.
14. An electronic device, characterized in that: include: a memory for storing computer-executable instructions; A processor, configured to implement the artificial intelligence-based image scoring method according to any one of claims 1 to 7 when executing the computer-executable instructions stored in the memory.
Citation Information
Patent Citations
Vehicle type identification method and system and storage medium
CN109359696A
Writing rating method for composition area in test paper scanning image
CN110929674A
Character recognition method and device, electronic equipment and storage medium
CN112686263A