Test question tag generation method, device and equipment

By combining visual and linguistic representation extractors, test item tags are automatically generated, solving the problem of high cost and low efficiency in manually determining test item tags, and achieving efficient and accurate test item tag generation.

CN121921783APending Publication Date: 2026-04-24IFLYTEK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2025-12-29
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In test-based learning scenarios, existing technologies require manual identification of test item tags, resulting in high costs and low efficiency.

Method used

Test item labels are automatically generated using a preset visual representation extractor and linguistic representation extractor. First, the visual representation extractor extracts features from the test item image to obtain visual features. Then, the linguistic representation extractor performs inference to generate target classification labels, and finally, the test item labels are obtained.

Benefits of technology

The system enables automated generation of test question tags, reducing generation costs and improving efficiency while ensuring the accuracy and consistency of the tags.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921783A_ABST
    Figure CN121921783A_ABST
Patent Text Reader

Abstract

The invention provides a test question tag generation method, device and equipment. The test question tag generation method comprises the steps of obtaining a test question image of a target test question; performing feature extraction on the test question image through a preset visual representation extractor to obtain visual features; reasoning the visual features through a preset language representation extractor to obtain a target classification label of the target test question under at least one classification dimension; and obtaining the test question tag of the target test question according to the at least one target classification tag, thereby reducing the generation cost of the test question tag and improving the generation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and more specifically, to a method, apparatus, and device for generating test item tags. Background Technology

[0002] With the deepening of educational informatization and the widespread implementation of multimodal learning technologies, various learning scenarios for test questions have emerged, such as video explanations of test questions, interactive answering, and AR-based scenario demonstrations.

[0003] Currently, in test-based learning scenarios, it's necessary to manually determine the labels of test questions before performing corresponding learning operations based on those labels. For example, selecting a suitable model to answer a question based on its label. However, this method of determining test question labels is costly and inefficient. Summary of the Invention

[0004] This application provides a method, apparatus, and device for generating test item tags. First, a visual feature extraction method is used to extract features from the test item image of the target test item using a preset visual representation extractor. Then, a preset linguistic representation extractor is used to infer the visual features to obtain a target classification label for the target test item in at least one classification dimension. Finally, based on the at least one target classification label, a test item tag for the target test item can be obtained. This enables automated generation of test item tags, thereby reducing the generation cost and improving generation efficiency.

[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.

[0006] According to one aspect of this application, a method for generating test item labels is provided, comprising: acquiring a test item image of a target test item; extracting features from the test item image using a preset visual representation extractor to obtain visual features; reasoning about the visual features using a preset linguistic representation extractor to obtain a target classification label for the target test item in at least one classification dimension; and obtaining a test item label for the target test item based on the at least one target classification label.

[0007] According to one aspect of this application, a test item label generation apparatus is provided, comprising: an image acquisition module for: acquiring a test item image of a target test item; a feature extraction module for: extracting features from the test item image using a preset visual representation extractor to obtain visual features; a feature reasoning module for: reasoning from the visual features using a preset linguistic representation extractor to obtain a target classification label for the target test item in at least one classification dimension; and a label generation module for: obtaining a test item label for the target test item based on at least one target classification label.

[0008] According to one aspect of this application, an electronic device is provided, comprising: a processor and a memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the steps of the above-described test item tag generation method.

[0009] According to one aspect of this application, a chip is provided, comprising: a processor for calling and running a computer program from a memory, such that the processor performs the steps of the above-described test item label generation method.

[0010] According to one aspect of this application, a computer-readable storage medium is provided for storing a computer program that causes a computer to perform the steps of the above-described test item label generation method.

[0011] Based on the above technical solution, the visual features of the target test question image can be extracted first by using a preset visual representation extractor; then, the visual features can be inferred by a preset linguistic representation extractor to obtain the target classification label of the target test question in at least one classification dimension; finally, the test question label of the target test question can be obtained based on at least one target classification label, thereby realizing the automatic generation of test question labels, which can reduce the generation cost of test question labels and improve the generation efficiency.

[0012] Other features and advantages of the embodiments of this application will become apparent from the following detailed description, or may be learned in part by practice of this application.

[0013] It should be understood that the above general description and the following detailed description are merely exemplary and do not constitute a limitation on this application. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0015] Figure 1 This diagram illustrates an application scenario of a test item tag generation method provided in one embodiment of this application. Figure 2 A flowchart illustrating a test item tag generation method according to an embodiment of this application is shown; Figure 3 A schematic diagram illustrating a test item tag generation method according to an embodiment of this application is shown; Figure 4A block diagram of a test question tag generation apparatus according to an embodiment of this application is shown; Figure 5 A schematic diagram of the structure of a computer system suitable for implementing the electronic devices of the present application is shown. Detailed Implementation

[0016] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided to make the description of this application more complete and to fully convey the concept of the exemplary embodiments to those skilled in the art. The accompanying drawings are schematic illustrations of this application and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0017] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more exemplary embodiments. Numerous specific details are provided in the following description to give a full understanding of exemplary embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced with one or more specific details omitted, or other methods, components, steps, etc., can be employed. In other instances, well-known structures, methods, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.

[0018] Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different networks, processor devices, or microcontroller devices.

[0019] Figure 1 This is an application scenario diagram of the test question tag generation method provided in one embodiment, such as... Figure 1 As shown, this application scenario includes terminal 110 and server 120.

[0020] In some embodiments, the technical solution of this application is mainly applied to test question tag generation scenarios, such as test question tag generation scenarios in education and teaching scenarios. Specifically, it can be for any subject such as mathematics or Chinese, or test question tag generation scenarios for different educational stages such as primary school and junior high school, but it is not limited to these.

[0021] In some embodiments, the terminal 110 can send the image of the target test question to the server 120. The server 120 can extract features from the test question image using a preset visual representation extractor to obtain visual features; infer the visual features using a preset linguistic representation extractor to obtain the target classification label of the target test question in at least one classification dimension; obtain the test question label of the target test question based on the at least one target classification label, and send the test question label to the terminal 110. The terminal 110 can display, store, etc., the test question label.

[0022] It is understood that the above application scenario is only an example and does not constitute a limitation on the test question tag generation method provided in the embodiments of this application.

[0023] The server 120 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal 110 can be a smartphone, tablet, laptop, desktop computer, smart speaker, or smartwatch, but is not limited to these. The terminal 110 and server 120 can be directly or indirectly connected via wired or wireless communication, and this application does not limit this connection.

[0024] The specific implementation process of the embodiments of this application will be described in detail below.

[0025] Figure 2 A flowchart illustrating a test item tag generation method according to an embodiment of this application is shown. This test item tag generation method can be executed by a device with computing power, such as the aforementioned terminal 110 or server 120. (Refer to...) Figure 2 As shown, the method for generating test item tags can include at least S210 to S240, which are described in detail below: In S210, the image of the target test question is obtained.

[0026] For example, the target test question can be any subject, grade level, image quality, or question type; in addition, the target test question can include the question title, or it can include the question title and the answer content, which is not limited in this application.

[0027] For example, the test question image of the target test question can be a test question image captured by photograph, or a digital test question image obtained by extracting, processing or converting from teaching materials, electronic question banks, PDF documents, etc.

[0028] For example, before performing subsequent embodiments, the test question image can be adjusted to a specified pixel size or a preset resolution to adapt to the input requirements of the subsequent model. For instance, the test question image can be adjusted to a 448×448 pixel image (i.e., the width and height of the image are both 448 pixels), but it is not limited to this.

[0029] In the above embodiments, since the image can preserve all visual elements of the test question as is: text, formulas, charts, layout, color, clarity, whether it is occluded, etc., it can provide original materials for subsequent analysis. Moreover, the image is more suitable for the most frequent use scenario of users taking pictures to search for questions. Therefore, there is no need to perform additional format conversion on the test question, which is convenient for direct use in subsequent downstream scenarios (e.g., test question tutoring system).

[0030] Furthermore, the processed standardized image tensor will be directly passed to the subsequent visual representation extractor for visual feature extraction, providing a unified and high-quality visual input for obtaining visual features and determining classification labels.

[0031] In S220, the visual features are obtained by extracting features from the test question image through a preset visual representation extractor.

[0032] For example, a visual representation extractor can be a pre-trained model that extracts deep semantic features from images, specifically a visual model.

[0033] For example, a visual representation extractor can be based on a Transformer architecture model. For instance, it could be a Vision Transformer (ViT); or it could be a model based on a vision transformer, but it is not limited to these.

[0034] For example, the InternViT-6B model can be pre-trained using knowledge distillation techniques to obtain InternViT-300M, which can then be used as a visual representation extractor. Furthermore, a variant of InternViT-300M, InternViT-300M-448px, supports dynamically input images with a fixed resolution of 448×448 and a tiled size of 448×448. It exhibits strong robustness, efficient optical character recognition (OCR) capabilities, and high-resolution image processing capabilities, making it suitable for image classification and feature extraction tasks. Therefore, it can also be used as a visual representation extractor.

[0035] Alternatively, a lightweight student model, InternViT-300M-448px-Distill, can be obtained by learning from the teacher model, InternViT-6B-448px-V1.5, using knowledge distillation techniques (e.g., training with cosine distillation loss to align the output features of the student model with those of the teacher model in the feature space). This student model has 0.3B parameters and is configured as follows: a 24-layer network structure with 1024 hidden layers, 16 attention heads, and a standard Layer Norm layer as the normalization mechanism. Then, the student model can be integrated with a Large Language Model (LLM) to obtain the aforementioned visual representation extractor. The integrated student model and LLM can be trained using at least one of the following methods: training with dynamic high-resolution techniques (to enable the visual representation extractor to flexibly handle test images of different sizes and resolutions, improving the fine granularity of visual understanding); or training with the loss corresponding to Next Token Prediction (NTP).

[0036] For example, the visual processing capabilities of a visual representation extractor can be used to encode and integrate test question images, extracting visual element information from the test question images to obtain visual features. These visual features can represent the visual element information.

[0037] For example, visual element information includes at least one of the following: text layout and structure information, non-text visual entity information, and image quality and status information.

[0038] The text layout and structure information includes at least: the positions of different functional text blocks such as the question stem and options in the question image; the line order, paragraph structure, and relative positional relationship between characters in the question image; and specific formatting such as "underline" for fill-in-the-blank questions, "brackets" and "numbering sequence" for multiple-choice questions in the question image.

[0039] Non-textual visual entity information includes at least: the outlines and types of geometric figures (e.g., triangles, circles), coordinate systems, function graphs, and statistical charts (e.g., bar charts, line charts) in the test question image; the content and location of mathematical formulas, chemical equations, physical symbols, etc. in the test question image; and illustrations in the test question image, such as photographs of real objects and schematic diagrams.

[0040] Image quality and status information includes at least: image clarity information such as the sharpness of text edges, overall contrast, and noise level in the test question image; test integrity information such as whether the test question image has been cropped or whether there is an excessively small effective content area (e.g., the question stem, attached figures, etc.); and interference information such as whether the test question image contains occlusions (e.g., fingers, stationery, shadows, reflections) and their impact on key content (e.g., the question itself).

[0041] It should be noted that the visual element information can correspond to the classification dimensions and classification labels in subsequent embodiments, and the visual element information can be used to determine the corresponding classification label under the corresponding classification dimension. For example, the image quality and state information in the above-mentioned visual element information can be used to determine the classification dimensions "the applicability dimension of the target test question to the test question learning system" and "the image quality dimension of the test question image"; the non-text visual entity information and text layout and structure information in the above-mentioned visual element information can be used to determine the classification dimension "the test question modality dimension of the target test question", but are not limited thereto.

[0042] In the above embodiments, precise visual element information can be extracted by combining the downstream test question classification requirements, i.e. the classification label division requirements, so as to output effective and accurate visual features and provide standardized input for subsequent processing steps; moreover, the visual features retain information such as the content semantics, spatial layout and quality status of the test question image, which can provide accurate and complete visual basis for subsequent classification label judgment.

[0043] In some embodiments, after obtaining the visual features, at least one of the following processes may be performed before executing S230, but not limited to: The first step is to merge the local features of each preset adjacent region in the visual features to perform inference based on the merged visual features; wherein the number of lexical units corresponding to the merged visual features is less than the number of lexical units corresponding to the visual features.

[0044] For example, suppose a visual feature is a three-dimensional tensor containing spatial dimensions (height H, width W) and channel dimensions C. The visual feature can be represented as (H, W, C), where each spatial location (i, j) corresponds to a C-dimensional feature vector. The height and width directions are divided with a step size of n (a positive integer) to obtain (H / n)*(W / n) non-overlapping n*n neighboring regions. For each neighboring region, the local features (containing n*n feature vectors) in the neighboring region are directly concatenated along the channel direction to obtain a new feature vector. Each new feature vector is arranged according to the position of the local feature corresponding to the new feature vector in the visual feature to obtain the merged visual feature, which can be represented as (H / n, W / n, n*C).

[0045] For example, visual features can be downsampled using a downsampling method (e.g., Pixel Unshuffle) to obtain processed visual features.

[0046] In this process, assuming that each spatial location corresponds to one word, the number of words included in the visual features is H × W; the number of words corresponding to the merged visual features is (H / n) * (W / n) = H × W / (n * n). Thus, without losing information in the visual features, the spatial resolution of the visual features can be reduced and the channel dimension increased, thereby reducing the number of words that the subsequent language model needs to process and improving the overall processing efficiency.

[0047] The second step involves aligning the visual features with word vectors to obtain the corresponding text features; correspondingly, S230 includes: reasoning about the text features through a language representation extractor.

[0048] For example, visual features can be preprocessed, such as through layer normalization, to stabilize their numerical distribution and facilitate subsequent transformations. Then, the visual features can be linearly mapped to text features in the semantic space that match the dimensions supported by the subsequent language representation extractor. Alternatively, non-linear activation functions (such as GELU or ReLU) can be used to perform non-linear activation calculations on intermediate features. Multiple linear mapping and non-linear activation steps can be performed.

[0049] For example, visual features can be converted into text features using a multilayer perceptron projector (MLPProjector).

[0050] In the above process, visual features can be converted into semantic features that can be used by the language representation extractor for reasoning, so that the language representation extractor can directly call its own reasoning ability to determine the classification label of the test question.

[0051] In S230, the visual features are inferred through a preset language representation extractor to obtain the target classification label of the target test item in at least one classification dimension.

[0052] For example, at least one classification dimension includes at least one of the following: the test modality dimension of the target test item, the image quality dimension of the test item image, and the applicability dimension of the target test item to the test item learning system.

[0053] Regarding the question modality dimension, the corresponding classification tags include: unimodal tags and multimodal tags. Quest modality refers to the information presentation format of the question, such as text, charts, images, formulas, etc. A unimodal tag indicates that the question is presented only in one format, such as text. A multimodal tag indicates that the question, in addition to text, also contains structured or semi-structured visual elements such as charts, images, and formulas, requiring the combination of information from multiple presentation formats to fully solve the problem.

[0054] Regarding the image quality dimension, it refers to the clarity or blurriness of the test question image. The corresponding classification labels include at least one level of blurriness. For example, it can include the following five levels of blurriness: A (unblurred label, meaning the test question image is completely clear); B (slightly blurry label, meaning the test question image is slightly blurry but easily identifiable); C (moderately blurry label, meaning the test question image is identifiable upon close inspection); D (severely blurry label, meaning a small portion of the text in the test question image is illegible); E (extremely blurry label, meaning more than 90% of the content in the test question image is illegible). Each level of blurriness corresponds to a classification label and a level of clarity or blurriness.

[0055] Regarding the applicability dimension, it refers to whether the test question image is suitable for use in a test question-assisted learning system. This system can be a downstream system, such as a model that determines whether a question can be answered and executes answering steps, but it is not limited to this. The corresponding classification tags include: Applicable tags (meaning the test question is complete, clear, unobstructed, and belongs to a type of question suitable for assisted learning, such as multiple-choice questions, fill-in-the-blank questions, etc.) and Inapplicable tags. The Inapplicable tags can include the following sub-tags: Test Question Validity Defect Tags (meaning the test question image corresponds to an invalid test question, such as not being a question, missing illustrations, or the question being too large / too small), Test Question Obstruction Tags (meaning the test question in the image is obstructed by an object, such as a hand, stationery, page turning, etc., obscuring text or images), Test Question Incomplete Tags (meaning the test question in the image is horizontally / vertically incomplete, or some answer materials or question requirements are missing), and Unsuitable Question Type Tags (meaning the type of question in the test question image does not belong to a type of question suitable for assisted learning, such as operation questions, listening questions, reading questions, copying questions, etc.). For example, a language representation extractor can be a large language model, which refers to an artificial intelligence model trained on a large amount of data and with strong language understanding and generation capabilities.

[0056] Exemplarily, S230 specifically includes the following steps: Obtain a label generation instruction (Prompt), where the label generation instruction at least includes: classification standard information and classification example information corresponding to classification dimensions; convert the label generation instruction into instruction tokens (Tokens) through a text tokenizer (TextTokenizer); perform inference on the visual features and instruction tokens through a language representation extractor to obtain at least one target classification label.

[0057] For example, the label generation instruction can be converted into instruction tokens through a preset vocabulary (Tokenizer), which facilitates the language representation extractor to understand and analyze the label generation instruction.

[0058] For example, considering the preloading (prefilling) time of the language representation extractor, the model inference efficiency can be improved by streamlining the label generation instruction. Moreover, since determining the classification label involves a classification task and the logic is relatively simple, the label generation instruction can be set as a simple instruction. For example, the label generation instruction can be expressed as: "<Image Here (refers to the test question image placed), complete three tasks according to the test question image: 1. Modal classification (corresponding to determining the target classification label under the test question dimension); 2. Blur classification (refers to determining the target classification label under the image quality dimension); 3. Usability classification (refers to determining the target classification label under the applicable dimension). Only output the result without explanation.>" For example, in order to effectively avoid the semantic ambiguity problem in the inference of the language representation extractor and improve the label parsing efficiency and accuracy, it can be determined that the classification standard information is as shown in Table 1 below.

[0059] Table 1

[0060] Exemplarily, in order for the language representation extractor to output more accurate and structured classification labels, specific tokens corresponding to specific classification labels can be added to the vocabulary corresponding to the language representation extractor; determine the mapping relationship between each classification label and the corresponding token in the vocabulary; according to the mapping relationship, perform inference on the visual features through the language representation extractor to obtain at least one target classification label.

[0061] For example, the specific classification label can be any classification label, and this application does not limit it.

[0062] For example, the following code can be used to set the specific tokens corresponding to specific classification labels in the vocabulary: {"Id (lexicon identifier for a specific lexicon): 151671", content (specific lexicon): <|answer yes|> "lstrip": false (lexicon attribute: do not remove leading whitespace), "normalized": false (lexicon attribute: do not normalize), "rstrip": false (lexicon attribute: do not remove leading whitespace), "single word": false (lexicon attribute: not a single word), "special: false (lexicon attribute: not a special control character);} "Id": 151672,content":<answer no> "lstrip": false, "normalized": false, "rstrip": false, "single word": false, "special": false}.

[0063] For example, two specific lexical units can be introduced into the language representation extractor: <|answer_yes|> and <|answer_no|>. <|answer_yes|> corresponds to dictionary ID (lexical identifier) ​​151671, and <|answer_no|> corresponds to dictionary ID 151672. Regarding the question modality dimension, if the target classification label is a unimodal label, the corresponding lexical unit is <|answer_yes|>, and the target classification label output by the language representation extractor can be <|answer_yes|>; if the target classification label is a multimodal label, the corresponding lexical unit is <|answer_no|>, and the target classification label output by the language representation extractor can be <|answer_no|>. For the applicable dimension, if the target classification label is an applicable label, the corresponding lexical unit is <|answer_yes|>, and the target classification label output by the language representation extractor can be <|answer_yes|>; if the target classification label is an inapplicable label, the corresponding lexical unit is <|answer_no|>, and the target classification label output by the language representation extractor can be <|answer_no|>.

[0064] For image quality labels, the corresponding category labels can be mapped to the existing English letters AE in the vocabulary, without needing to set specific words.

[0065] For example, visual features and label generation instructions can be input together into a language representation extractor. The language representation extractor can understand the target test question based on the visual features and, based on the constraint information corresponding to the label generation instructions (e.g., the classification criteria information shown in Table 1), and in combination with the above mapping relationship, make three classification decisions to generate target classification labels for each classification dimension. The target classification labels are generated based on the constraint information corresponding to the label generation instructions and are labels that correspond to / are bound to the lexical units in the vocabulary.

[0066] Furthermore, the language representation extractor can be trained through static semantic annotation and dynamic fine-tuning, enabling it to form a stable cognitive association between lexical units and classification labels. For example, the correct classification label of each training sample can be forcibly and explicitly replaced with the corresponding lexical unit by manual intervention or pre-set rules; subsequently, through diverse training tasks and feedback mechanisms, the language representation extractor can dynamically adjust the mapping relationship between lexical units and classification labels.

[0067] In the above embodiments, the mapping relationship can be determined by binding lexical units with classification tags, and dynamic reasoning can be performed accordingly, thereby achieving an accurate and structured representation of the target classification tags. Furthermore, the classification of the target can be divided into multiple classification dimensions, with each dimension corresponding to multiple classification tags. Therefore, the problems of coarse segmentation and insufficient adaptability can be solved, ensuring that the determined classification tags can accurately match the differentiated needs of the multimodal test question learning system, improving the adaptability and accuracy for diverse test questions.

[0068] In S240, the test item label of the target test item is obtained based on at least one target classification label.

[0069] For example, at least one target classification label can be combined in a preset order to obtain the question label. For instance, the question label can be determined as: (target classification label under the question modality dimension, target classification label under the image quality dimension, and target classification label under the applicability dimension). For example, the language representation extractor described above can be used to obtain the test item label of the target test item based on at least one target classification label, and this application does not limit this.

[0070] For example, question tags could be:<answer_no> B<answer_yes> Among them, the first "<answer_no> The first "B" indicates that the target classification label under the modality dimension of the test question is a multimodal label; the second "B" indicates that the target classification label under the image quality dimension is a B-level label / slightly blurry label; the third "<|answer_yes|>" indicates that the target classification label under the applicability dimension is an applicable label.

[0071] In the above embodiments, a unified structured label combining "question modality, image quality, and applicability" can be generated. This avoids the problems of data flow fragmentation, inconsistent results, and complex downstream integration caused by using different models to generate independent classification labels in related technologies. It can achieve efficient collaboration and standardized output of the label generation chain, reduce system integration costs, and enhance scalability and large-scale deployment capabilities.

[0072] In some embodiments, the extractor described above can also be trained through the following steps, but are not limited to: Obtain a training task set containing multiple types of test questions; predict at least one training target classification label corresponding to the training task set using a visual representation extractor and a linguistic representation extractor; train the visual representation extractor and the linguistic representation extractor based on the cross-validation results corresponding to at least one training target classification label.

[0073] For example, the training task set may include multiple types of questions that differ in modality (plain text, mixed text and image), image quality (clear to extremely blurry), completeness (complete, incomplete, occluded), and question type (study aid questions, operation questions, listening questions, etc.).

[0074] For example, the cross-validation result corresponding to the training objective classification label can be the result determined based on the label consistency check corresponding to at least one training objective classification label. For example, the consistency check includes at least: whether there is a contradiction in the representation information between at least one training objective classification label (e.g., "extremely ambiguous label" and "applicable label" are simultaneously labels corresponding to the same test item), whether the training objective classification label is inconsistent with the corresponding visual feature (e.g., the predicted training objective classification label is a unimodal label, but the visual feature representation of the test item contains a chart), and whether it is consistent with the classification standard information (e.g., the predicted training objective classification label is an inapplicable label, but the visual feature representation of the test item is a listening comprehension question).

[0075] The technical solution of this application will be further illustrated below with a schematic diagram: In some embodiments, in conjunction with the above embodiments, such as Figure 3As shown, an image of the test question (e.g., containing the question itself) can be acquired first. Then, the image is set to a fixed resolution according to the resolution supported by the visual representation extractor. Next, the visual representation extractor can extract features from the image to obtain visual features. Then, the visual features are sampled by the lexical adjustment module to reduce the number of lexical units while ensuring that the information contained in the visual features is not lost. Then, the visual features are mapped by the feature transformation module to obtain text features. At the same time, the label generation instruction (prompt instruction) can be determined, and the instruction is divided into multiple lexical units by the text segmenter. After that, the lexical units and text features can be input together into the language representation extractor to obtain the test question label.

[0076] It should be noted that the models, networks, modules, or units in this application can be parts of a whole model, network, module, or unit, or they can be independent models, networks, modules, or units. This application does not impose any restrictions on this. For example, the visual representation extractor, the lexical adjustment module, and the feature transformation module can be different parts of the same model, or they can be independent parts. The visual representation extractor and the language representation extractor can be the same model or different models. This application does not impose any restrictions on this.

[0077] In some embodiments, the technical solution of this application can be adapted to the pre-screening and structured preprocessing scenarios of multimodal learning assistance systems / test-based learning assistance systems. For example, it can provide automated and high-precision test question label judgment results for applications such as photo-based question search tools and intelligent teaching assistance platforms. Specifically, it can provide structured and high-quality test question labels for downstream answering models or rejection models (referring to modules used to identify and reject test question inputs that cannot be answered effectively). This facilitates the downstream answering models or rejection models to effectively identify and intercept inapplicable or low-quality test questions (such as extremely vague, incompatible question types, incomplete content, etc.), thereby avoiding ineffective resource consumption and system operation errors, and fundamentally ensuring the efficient, stable, and reliable operation of the multimodal learning assistance system.

[0078] The technical solution of this application can realize integrated and automated collaborative discrimination of test item modality, image quality and applicability to auxiliary learning system, and output three-in-one label. It solves the problems of functional fragmentation between various label determination modules, rough judgment and inefficiency and high cost caused by reliance on manual determination in related technologies for determining multi-dimensional labels. It can improve the accuracy and efficiency of label determination, and provide structured and comprehensive test item labels for downstream answering models, rejection models and other models.

[0079] Figure 4A block diagram of a test item tag generation apparatus according to an embodiment of this application is shown. This test item tag generation apparatus may be a software unit or a hardware unit, or a combination of both, as part of a computer device. Figure 4 As shown, the test question tag generation device 400 provided in this application embodiment may specifically include: Image acquisition module 410 is used to: acquire the image of the target test question; feature extraction module 420 is used to: extract features from the test question image using a preset visual representation extractor to obtain visual features; feature reasoning module 430 is used to: reason about the visual features using a preset linguistic representation extractor to obtain the target classification label of the target test question in at least one classification dimension; label generation module 440 is used to: obtain the test question label of the target test question based on at least one target classification label.

[0080] In some embodiments, at least one classification dimension includes at least one of the following: the test modality dimension of the target test item, the image quality dimension of the test item image, and the applicability dimension of the target test item to the test item learning system.

[0081] In some embodiments, the feature processing module 450 is configured to: merge the local features of each preset adjacent region in the visual features, so as to perform inference based on the merged visual features; wherein the number of lexical units corresponding to the merged visual features is less than the number of lexical units corresponding to the visual features.

[0082] In some embodiments, the feature processing module 450 is used to: perform word vector alignment on visual features to obtain text features corresponding to the visual features; correspondingly, the feature inference module 430 is specifically used to: infer text features through a language representation extractor.

[0083] In some embodiments, the label generation module 440 is specifically used to: combine at least one target classification label in a preset order to obtain test item labels.

[0084] In some embodiments, the feature reasoning module 430 is specifically used to: add specific words corresponding to specific classification labels to the vocabulary of the language representation extractor; determine the mapping relationship between each classification label and the corresponding word in the vocabulary; and, based on the mapping relationship, reason about the visual features through the language representation extractor to obtain at least one target classification label.

[0085] In some embodiments, the feature reasoning module 430 is specifically used to: obtain a label generation instruction, which includes at least: classification standard information and classification example information corresponding to the classification dimension; convert the label generation instruction into instruction lexical units through a text segmenter; and reason about the visual features and instruction lexical units through a language representation extractor to obtain at least one target classification label.

[0086] In some embodiments, the model training module 460 is specifically used to: obtain a training task set containing multiple types of test questions; predict at least one training target classification label corresponding to the training task set through a visual representation extractor and a linguistic representation extractor; and train the visual representation extractor and the linguistic representation extractor based on the cross-validation results corresponding to the at least one training target classification label.

[0087] The specific implementation of each module in the test question tag generation device provided in this application embodiment can refer to the content of the test question tag generation method described above, and will not be repeated here.

[0088] Each module in the aforementioned test question label generation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0089] Figure 5 A schematic diagram of a computer system suitable for implementing the embodiments of this application is shown. It should be noted that... Figure 5 The computer system 500 of the electronic device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0090] like Figure 5 As shown, the computer system 500 includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 502 or programs loaded from storage section 508 into Random Access Memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An Input / Output (I / O) interface 505 is also connected to the bus 504.

[0091] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network (LAN) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 510 as needed so that computer programs read from them can be installed into storage section 508 as needed.

[0092] Specifically, according to embodiments of this application, the processes described in the flowcharts above can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts above. In such embodiments, the computer program can be downloaded and installed from a network via communication section 509, and / or installed from removable medium 511. When the computer program is executed by central processing unit (CPU) 501, it performs the various functions defined in the apparatus of this application.

[0093] In some embodiments, an electronic device is also provided, comprising: Processor; and Memory is used to store the processor's executable instructions; The processor is configured to execute the steps in the above method embodiments by executing executable instructions.

[0094] In some embodiments, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0095] In some embodiments, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0096] It should be noted that the computer-readable storage medium of this application can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, disk storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, the computer-readable signal medium may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, radio frequency, etc., or any suitable combination thereof.

[0097] This embodiment is only used to illustrate this application. The selection of software and hardware platform architecture, development environment, development language, message acquisition source, etc. in this embodiment can be varied. Based on the technical solution of this application, any improvement or equivalent transformation made to a certain part according to the principle of this application should not be excluded from the protection scope of this application.

[0098] It should be noted that the terminology used in the embodiments of this application and the appended claims is for the purpose of describing specific embodiments only, and is not intended to limit the embodiments of this application.

[0099] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of the embodiments of this application.

[0100] If implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application embodiment, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method of this application embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices, apparatuses and methods can be implemented in other ways.

[0103] For example, the division of units, modules, or components in the device embodiments described above is merely a logical functional division. In actual implementation, there may be other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some units, modules, or components may be ignored or not executed.

[0104] For example, the units / modules / components described above as separate / display components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the units / modules / components can be selected to achieve the objectives of the embodiments of this application, depending on actual needs.

[0105] Finally, it should be noted that the mutual coupling or direct coupling or communication connection shown or discussed above can be an indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.

[0106] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be included within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.

Claims

1. A method for generating test item tags, characterized in that, include: Obtain the image of the target test question; The visual features are obtained by extracting features from the test question image using a preset visual representation extractor. The visual features are inferred by a preset language representation extractor to obtain the target classification label of the target test question in at least one classification dimension; The test item label of the target test item is obtained based on at least one of the target classification labels.

2. The method according to claim 1, characterized in that, The at least one classification dimension includes at least one of the following: the test modality dimension of the target test question, the image quality dimension of the test question image, and the applicability dimension of the target test question to the test question learning system.

3. The method according to claim 1, characterized in that, Before reasoning about the visual features using a preset language representation extractor, the method further includes: The local features of each preset adjacent region in the visual features are merged separately, so as to perform the inference based on the merged visual features; The number of lexical units corresponding to the merged visual features is less than the number of lexical units corresponding to the visual features.

4. The method according to claim 1, characterized in that, Before reasoning about the visual features using a preset language representation extractor, the method further includes: Word vector alignment is performed on the visual features to obtain the text features corresponding to the visual features; Accordingly, the reasoning of the visual features using a preset language representation extractor includes: The text features are inferred using the language representation extractor.

5. The method according to claim 1, characterized in that, The step of obtaining the test item label of the target test item based on at least one of the target classification labels includes: The at least one target classification label is combined in a preset order to obtain the test item label.

6. The method according to any one of claims 1-5, characterized in that, The step of reasoning about the visual features using a preset language representation extractor to obtain the target classification label of the target test item in at least one classification dimension includes: Add specific lexical units corresponding to specific category labels to the vocabulary of the language representation extractor; Determine the mapping relationship between each category label and the corresponding word element in the vocabulary; Based on the mapping relationship, the visual features are inferred through the language representation extractor to obtain at least one target classification label.

7. The method according to any one of claims 1-5, characterized in that, The step of reasoning about the visual features using a preset language representation extractor to obtain the target classification label of the target test item in at least one classification dimension includes: Obtain a tag generation instruction, which includes at least: classification standard information and classification example information corresponding to the classification dimension; The tag generation instructions are converted into instruction tokens using a text segmenter. The language representation extractor is used to reason about the visual features and the instruction lexical units to obtain at least one target classification label.

8. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain a training task set containing various types of test questions; Using the visual representation extractor and the language representation extractor, predict at least one training target classification label corresponding to the training task set; The visual representation extractor and the language representation extractor are trained based on the cross-validation results corresponding to the at least one training target classification label.

9. A test item label generation device, characterized in that, include: The image acquisition module is used to: acquire the image of the target test question; The feature extraction module is used to: extract features from the test question image using a preset visual representation extractor to obtain visual features; The feature reasoning module is used to: reason about the visual features through a preset language representation extractor to obtain the target classification label of the target test question in at least one classification dimension; The tag generation module is used to: obtain the question tag of the target question based on at least one of the target classification tags.

10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to perform the method of any one of claims 1 to 8 by executing the executable instructions.