An examination question information extraction method and device, electronic equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 深圳市星桐科技有限公司
- Filing Date
- 2023-06-07
- Publication Date
- 2026-06-02
Smart Images

Figure CN116758554B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and storage medium for extracting test information. Background Technology
[0002] In recent years, with the development of online education, more and more users are using test question banks to search for test questions and obtain problem-solving ideas or answers. The most common method is to take a picture to search for questions.
[0003] In related technologies, a user uploads a test question image, the Q&A system obtains the test question image, and pushes multiple test questions most similar to the user's uploaded test question, along with their stems, answers, and explanations, for the user's reference; and for test questions containing illustrations, image search technology is introduced to obtain the test question image. Summary of the Invention
[0004] According to one aspect of this disclosure, a method for extracting test question information is provided, including:
[0005] Determine the initial question type, initial illustration type, and illustration image for illustrated questions;
[0006] Based on the initial question type, the initial illustration type, and the illustration image, question feature information is obtained, wherein the question feature information includes at least illustration feature information;
[0007] Test question information is obtained based on the feature information of the illustration.
[0008] According to another aspect of this disclosure, a test item recommendation method is provided, comprising:
[0009] The method described in the exemplary embodiments of this disclosure determines the question information of the illustrated test question to be searched;
[0010] Recommended questions are obtained from the question bank based on the information of the question to be searched.
[0011] According to another aspect of this disclosure, a device for extracting test question information is provided, comprising:
[0012] The determination module is used to determine the initial question type, initial illustration type, and illustration image for illustrated questions;
[0013] The module is configured to obtain test question feature information based on the initial test question type, the initial illustration type, and the illustration image, wherein the test question feature information includes at least illustration feature information;
[0014] The obtaining module is also used to obtain test question information based on the illustration feature information.
[0015] According to another aspect of this disclosure, a test item recommendation device is provided, comprising:
[0016] The determining module determines the question information of the illustrated test questions to be searched based on the method described in the exemplary embodiments of this disclosure;
[0017] The acquisition module is used to acquire recommended test questions from the test question bank based on the test question information to be searched.
[0018] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0019] Processor; and memory that stores programs;
[0020] The program includes instructions that, when executed by the processor, cause the processor to perform the method according to an exemplary embodiment of the present disclosure.
[0021] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided, the non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method according to exemplary embodiments of this disclosure.
[0022] One or more technical solutions provided in the exemplary embodiments of this disclosure can obtain illustration images from illustrated test questions, and obtain test question feature information including at least illustration feature information based on the initial test question type, the initial illustration type, and the illustration image. Therefore, the illustration feature information not only contains relevant features of the illustration image, but also integrates the initial test question type and the initial illustration type, thereby analyzing the illustration image more accurately and avoiding interference from other factors in the illustrated test questions. Then, test question information is obtained based on the illustration feature information. It is evident that the exemplary embodiments of this disclosure can extract illustration images from illustrated test questions, allowing for both overall analysis of the illustrated image and individual analysis of the illustration image. By combining these two analyses, interference from other textual information besides the illustration in the illustrated test question image can be avoided, improving the search capability for illustrated test questions and enhancing the accuracy of test question information extraction. Based on this, when test question information for illustrated test questions to be searched is determined based on the method of the exemplary embodiments of this disclosure, and recommended test questions are obtained from the test question database based on the search question information, test question filtering can be performed efficiently, improving the accuracy of test question search, thereby providing users with accurate test question recommendations. Attached Figure Description
[0023] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this disclosure, illustrate exemplary embodiments of the present disclosure and are used to explain the disclosure, but do not constitute an undue limitation of the disclosure. In the drawings:
[0024] Figure 1A schematic diagram of an example system in which the various methods described herein may be implemented according to exemplary embodiments of the present disclosure;
[0025] Figure 2 A schematic flowchart of a method for extracting test question information according to an exemplary embodiment of this disclosure is shown;
[0026] Figure 3 A schematic diagram illustrating the text recognition process of a text recognition model according to an exemplary embodiment of the present disclosure is shown.
[0027] Figure 4 A schematic flowchart of a test question recommendation method according to an exemplary embodiment of this disclosure is shown;
[0028] Figure 5 A flowchart illustrating a test question similarity determination method according to an exemplary embodiment of the present disclosure is shown.
[0029] Figure 6 A schematic block diagram of the functional modules of a test question information extraction device according to an exemplary embodiment of the present disclosure is shown;
[0030] Figure 7 A schematic block diagram of the functional modules of a test item recommendation device according to an exemplary embodiment of the present disclosure is shown;
[0031] Figure 8 A schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown;
[0032] Figure 9 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0034] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0035] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0036] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0037] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0038] Before introducing the embodiments of this disclosure, the relevant terms involved in the embodiments of this disclosure are first defined as follows:
[0039] Image search uses image recognition technology to search for similar images on the internet based on visual features such as color distribution, geometric shape, and texture of the original image.
[0040] Forward reasoning starts with atomic statements in the knowledge base and applies reasoning rules in the forward direction to extract more data until the goal is reached.
[0041] Optical Character Recognition (OCR) refers to the process of analyzing and recognizing textual data in image files to obtain text and layout information. In other words, it involves recognizing the text in an image and returning it as text.
[0042] Differentiable Binarization Network (DBNet) algorithm, also known as Differentiable Binarization Processing, achieves adaptive thresholding at various points on the heatmap by inserting binarization operations into the segmentation network for combined optimization.
[0043] A bounding box is a simple geometric space that, in a 3D point cloud, contains a clustered set of points. Constructing bounding boxes for the target point set allows the extraction of the geometric attributes of obstacles, which are then used as observations by the tracking module. Transforming the scattered target point cloud into regular objects using bounding boxes makes it easier for the decision-making module to plan motion trajectories.
[0044] Convolutional Recurrent Neural Networks (CRNNs) are primarily used for end-to-end recognition of text sequences of variable length. Instead of segmenting individual characters first, they transform text recognition into a time-dependent sequence learning problem, which is image-based sequence recognition.
[0045] Convolutional Neural Networks (CNNs) are a type of feedforward neural network that includes convolutional computations and has a deep structure. They are one of the representative algorithms of deep learning.
[0046] A recurrent neural network (RNN) is a type of recurrent neural network that takes sequential data as input, recursively moves along the direction of the sequence, and connects all nodes (recurrent units) in a chain-like manner.
[0047] The CTCLoss (Connectionist Temporal Classification Loss) loss function is designed to address the misalignment between the labels of neural network data and the network's predicted data output.
[0048] Bidirectional Long Short-Term Memory (LSTM) is a type of time-recurrent neural network suitable for processing and predicting important events with relatively long intervals and delays in time series.
[0049] Word2Vec is an open-source tool developed by Google for calculating word vectors. It is a shallow neural network that converts words in natural language into dense vectors that computers can understand.
[0050] A Huffman tree, also known as an optimal binary tree, is a type of binary tree with the shortest weighted path length.
[0051] The YOLOv5s network model is the most commonly used lightweight object detection model, implemented based on the PyTorch framework. It includes five versions: YOLOv5n, YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x. Among them, YOLOv5s is fast and has a small model size, making it easy to scale to embedded devices for production use.
[0052] CenterNet is a classic object detection algorithm proposed in the 2019 paper "Objects as Points". The algorithm uses an anchor-free approach to implement object detection and other extended tasks. CenterNet treats object detection as a standard keypoint estimation problem, representing the object as a single point at the center of its bounding box. Other attributes, such as object size, dimensions, orientation, and pose, are directly regressed from the image features at this center point.
[0053] EfficientNets was proposed by Google Brain engineers Mingxing Tan and Quoc V. Le in their paper "EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks". The basic network architecture of this model was designed using neural architecture search. Convolutional neural network models are typically trained under known hardware resources. When you have better hardware resources, you can scale the network model to obtain better training results. To scale the research model for the system, Google Brain researchers proposed a novel model scaling method for the basic network model of EfficientNets. This method uses simple and efficient composite coefficients to balance network depth, width, and input image resolution.
[0054] Global average pooling is a common feature in neural networks. There are four common pooling operations: mean-pooling, max-pooling, stochastic-pooling, and global average pooling. Pooling layers have a significant effect: reducing the size of feature maps, which means reducing computation and memory requirements.
[0055] MobileNet is a convolutional neural network proposed by Google in 2017 for use in mobile devices and embedded systems. Its main applications include smartphones, drones, robots, autonomous driving, augmented reality, and more.
[0056] SqueezeNet, proposed by Han et al., is a lightweight and efficient CNN model with 50x fewer parameters than AlexNet, yet its performance is close to that of AlexNet. At an acceptable performance level, smaller models offer many advantages over larger models.
[0057] Elasticsearch is a powerful open-source search engine with numerous robust features that help us quickly find the content we need from massive amounts of data. Elasticsearch, combined with Kibana, Logstash, and Beats, forms the Elastic Stack (ELK), which is widely used in log data analysis, real-time monitoring, and other fields. Elasticsearch is the core of the Elastic Stack, responsible for storing, searching, and analyzing data.
[0058] In recent years, with the increasing prevalence of online learning systems in the education environment and the growing number of online learners, more and more students are using question banks to search for test questions and obtain solutions or answers. For example, students can use electronic devices with camera capabilities to capture images of the test questions they want to search for and upload them to the server corresponding to the Q&A system. The server then identifies the image of the test question, matches it with the most similar question in the question bank, and provides the matched question and its explanation process back to the student.
[0059] However, in practical applications, most Q&A systems primarily rely on text-based searches. This results in poor search results for illustrated questions with minimal text, often leading to significant discrepancies between the actual questions presented to students and the images. To improve the accuracy of searches for illustrated questions, existing technologies largely incorporate image-based search. However, this common technique is a fuzzy search, relying on similarity to the original image. The results are often broad, coarse, or even unsuccessful. Furthermore, interference from text and other factors within illustrated questions, coupled with the vast number of questions in the question bank, makes it difficult and inefficient for students to find matching questions when they take photos of purely image-based or predominantly image-based questions. Even if the original question exists in the question bank, the search may fail to find the correct answer, or the searched question may not be the one the student is looking for. This results in difficulties finding relevant questions when searching for images, leading to inefficiency and difficulty in searching for questions.
[0060] To address the aforementioned issues, this disclosure proposes a method for extracting test question information. When the test question is an illustrated test question, the text and illustrations in the illustrated test question image can be detected separately. The detection results, along with the original image data, are input into the test question image feature extraction module, which outputs the feature vector of the illustrated test question image. Finally, a feature vector search library is used to search for the most similar test question in the test question database based on the feature vector of the illustrated test question image and provide it to the student user.
[0061] Figure 1 A schematic diagram of an example system in which various methods described herein can be implemented according to exemplary embodiments of this disclosure is shown. Figure 1 As shown, the system 100 of the exemplary embodiments of this disclosure may include: a user device 110, a computing device 120, and a data storage system 130.
[0062] like Figure 1 As shown, the user equipment 110 can communicate with the computing device 120 via a communication network. This communication network can be a wired communication network or a wireless communication network. The wired communication network can be a communication network based on power line carrier technology, and the wireless communication network can be a local area network (LAN) or a wide area network (WAN). The LAN can be a Wi-Fi network, a Zigbee network, a mobile communication network, or a satellite communication network, etc.
[0063] like Figure 1 As shown, the user equipment 110 may include a computer, mobile phone, or information processing center, etc., as a smart terminal. The user equipment 110 can act as an image acquisition terminal for the test question image to be searched and initiate a request to the computing device 120. The computing device 120 can be a cloud server, network server, application server, or management server, etc., with data processing capabilities, to implement the extraction and search methods. The server may be configured with processors, which may include text detection processors and image detection processors, to complete the tasks of extracting test question information and searching for test questions.
[0064] like Figure 1 As shown, the data storage system 130 described above can store a database of test question information. The database can be located on the computing device 120 or on another network server. The data storage system 130 can be separate from the computing device 120 or integrated into the computing device 120.
[0065] In practical applications, computer devices can acquire test question information based on test question images. If the test question to be searched is one without illustrations, only the text can be detected. If the test question to be searched is one with illustrations, both the text and the illustrations in the test question image are detected. The detection results are then input together with the original image data into the test question image feature extraction module to obtain the feature vector of the test question image. At this point, a database containing test question feature vectors can be constructed based on the feature vectors. Based on this, when a student user can acquire the test question image to be searched through their user device and upload it to the computing device via a communication network, the computer device can recognize the test question image, obtain its feature vector, and search the database for multiple test questions that match the feature vector of the test question image. The computer then returns the searched test questions and their analysis process to the student user in descending order of similarity for reference.
[0066] The method for extracting test question information according to the exemplary embodiments of this disclosure can be applied to a server or a chip in a server. The method of the exemplary embodiments of this disclosure is described in detail below with reference to the accompanying drawings.
[0067] Figure 2 A schematic flowchart of a method for extracting test question information according to an exemplary embodiment of this disclosure is shown. The method for extracting test question information according to an exemplary embodiment of this disclosure includes:
[0068] Step 201: Determine the initial question type, initial illustration type, and illustration image of the illustrated test questions. It should be understood that, in terms of question content, the illustrated test questions of this exemplary embodiment can be basic subject test questions or various applied subject test questions; in terms of the target audience, they can be student test questions or various possible non-student test questions, such as various vocational qualification test questions, but are not limited thereto.
[0069] In practical applications, the initial question type, initial illustration type, and illustration image of illustrated test questions can be manually labeled directly, or they can be determined through a network model. For example, after acquiring illustrated test question images based on user devices, the initial question type can be determined based on the illustrated test question images. Simultaneously, the illustration location information can be determined based on the illustrated test question images, and the illustration image can be obtained from the illustrated test questions based on the illustration location information; finally, the initial illustration type can be determined based on the illustration image.
[0070] Step 202: Based on the initial question type, initial illustration type, and illustration image, obtain question feature information, which includes at least illustration feature information. Here, in the process of obtaining question feature information based on the initial question type, initial illustration type, and illustration image, the initial question type and initial illustration type can provide a reference for feature information extraction. This ensures that the question feature information extracted from the illustration image not only contains the features of the illustration image but also references the features implicit in the initial question type and initial illustration type, ensuring that the extracted question feature information is more comprehensive and accurate.
[0071] Step 203: Obtain test question information based on illustration feature information. Here, when obtaining test question information based on illustration feature information, to enrich the content of the test question information, it can be determined based on illustration feature information and illustration size information. Furthermore, to ensure that the test question information can be more fully referenced, multiple identical illustration size information can be copied N times, i.e., expanded by an order of magnitude, thereby increasing the data volume and improving the model's generalization ability. Based on this, test question information can be obtained through the above method. When searching for test questions in the test question bank using this information, the accuracy of the search results can be guaranteed. Simultaneously, a test question bank can be constructed based on the test question information, thus ensuring both search efficiency and search accuracy during test question searches.
[0072] As can be seen, the exemplary embodiments of this disclosure can not only perform overall analysis of illustrated images, but also analyze illustrated images separately. By combining the two analyses, interference from other textual information in the illustrated test question images can be avoided, improving the search capability of illustrated test questions, and also improving the accuracy of test question information extraction.
[0073] In one possible implementation, when determining the corresponding initial test question type based on the illustrated test question image in the exemplary embodiment of this disclosure, the test question text information of the illustrated test question image can be identified based on a text recognition model, the word vectors of multiple sub-texts contained in the test question text information can be determined based on the test question text information, and the word vectors of multiple sub-texts can be processed based on a text classification model to obtain the initial test question type.
[0074] In practical applications, the text recognition model of this exemplary embodiment can be an optical character recognition (OCR) model. Before inputting the illustrated test question image into the optical character recognition model, the illustrated test question image can be processed according to at least one of the following processing methods as needed:
[0075] The first method is to scale the illustrated test question image to a size suitable for the input image of the text recognition model. For example, the illustrated test question image can be scaled or combined with blanking to be processed into a square or rectangle, depending on the actual situation.
[0076] The second method involves standardizing the image data: centering the data by removing the mean. Based on convex optimization theory and knowledge of data probability distribution, data centering conforms to the data distribution law and makes it easier to achieve generalization effect after training. Alternatively, the image data can be normalized, transforming the pixel value range from 0 to 255 into 0 to 1, thus facilitating subsequent processing.
[0077] Figure 3 A schematic diagram illustrating the text recognition process of a text recognition model according to an exemplary embodiment of this disclosure is shown. Figure 3 As shown, the text recognition model can include a text detection module and a text line recognition module. The preprocessed illustrated test question image data is input into the text detection module for forward inference. The text detection module outputs the text line coordinate information of each text line in the illustrated test question image. Then, based on the text line coordinate information, text line images are obtained from the illustrated test question, and these text line images are input into the text line recognition module to obtain the test question text information.
[0078] For example, the text detection module of this exemplary embodiment can be a DBNet text detection module based on the DBNet algorithm, wherein the network skeleton of the DBNet text detection module can be a feature pyramid skeleton network. Based on this, an image of a test question with illustrations can be input into the feature pyramid skeleton network. Then, based on the upsampling method of the feature pyramid skeleton network, the output of the feature pyramid is transformed to the same size and cascaded to generate a feature map. Next, a probability map and a threshold map are generated based on the feature map. Then, an approximate binary map is calculated using the probability map and the threshold map. Labels are then expanded based on the approximate binary map to form text boxes. Furthermore, during the training phase, supervision is applied to the threshold map, the probability map, and the approximate binary map, with the latter two sharing the same supervision. During the inference phase, bounding boxes can be easily obtained from the latter two, making the DBNet text detection module perform excellently in both efficiency and detection results in the field of text detection. It should be understood that other text detection modules can also be used in practical applications, such as the EAST text detection module based on the EAST (Efficient and Accuracy Scene Text) network, and the PSENet text detection module based on the Progressive Scale Expansion Network (PSENet).
[0079] In this exemplary embodiment, after determining the position coordinates of each text line in an illustrated test question image, the text line position coordinates and the illustrated test question image are input into a text line recognition module. The text line recognition module extracts multiple lines of text images from the illustrated test question image based on the text line position coordinates and text borders. However, since user-uploaded illustrated test question images may have issues such as text tilting or distortion, the extracted multiple lines of text images need to be corrected to unify their positions. For example, horizontal and perspective corrections can be performed on the multiple lines of text images. Next, an image preprocessing module preprocesses the corrected multiple lines of text images, which may include scaling them to a standard size, filling in missing parts of the corrected images, or standardizing or normalizing the lines of text images. Finally, the corrected and preprocessed multiple lines of text images are input into the text line recognition module for forward inference. The text line recognition module outputs the test question text information corresponding to the multiple lines of text images.
[0080] For example, the text line recognition module of this exemplary embodiment can be a CRNN model, which may include convolutional layers, recurrent layers, and transcription layers. The convolutional layers can be CNN models, primarily used to extract features from the input image of the illustrated test question, obtaining a feature map of the image. The recurrent layers can be RNN models, using a bidirectional long short-term memory network to predict the feature sequence, learning each feature vector in the feature sequence, and outputting a predicted distribution of test question text information. The transcription layer includes a CTCLoss loss function, thereby using the CTC algorithm to transform the series of test question text information distributions obtained from the recurrent layers into the final test question text information sequence, i.e., the test question text information. It should be understood that in practical applications, text line recognition modules incorporating attention mechanisms or other text line recognition modules can also be used.
[0081] In practical applications, after obtaining the test text information output by the text recognition model, the word vectors of multiple sub-texts contained in the test text information can be determined based on the test text information. For example, the word vector tool (Word2Vec) can be used to convert the test text information into word vectors of multiple sub-texts. The Word2Vec tool mainly includes two models: the continuous bag of words (CBOW) module and the skip-gram module. CBOW is trained to predict the target word based on the context to obtain word vectors, while Skip-gram is trained to predict surrounding words based on the target word to obtain word vectors. Based on this, the test text information can be converted into word vectors of multiple sub-texts through CBOW and Skip-gram.
[0082] For example, the test text information is first segmented into words, then the frequency of each character is counted, then a Huffman tree corresponding to all characters is constructed, and the corresponding Huffman code is assigned to each character. Finally, the characters are trained based on the CBOW module and the skip character module to obtain the word vectors of multiple subtexts contained in the test text information.
[0083] After obtaining the word vectors of multiple sub-texts contained in the test question text, these word vectors can be input into a trained text classification model. The text classification model performs forward inference on the word vectors of the multiple sub-texts and outputs the initial test question type. Test question types include fill-in-the-blank questions, multiple-choice questions, and calculation questions, etc.
[0084] In one possible implementation, the aforementioned text classification model can be the fastText text classification model. When a sequence of word vectors from multiple sub-texts is input into the fastText text classification model, the model can output the initial question type corresponding to this sequence of word vectors. It should be understood that different question types can be labeled using numbers or English letters, allowing the fastText text classification model to directly output the initial question type number corresponding to different initial question types. Based on the initial question type number, not only can multiple illustrated questions be classified, but the question category information corresponding to different illustrated questions can also be seen more clearly and concisely.
[0085] An exemplary embodiment of this disclosure can also determine the location information of illustrations based on illustrated test question images. First, the illustrated test questions are preprocessed. Then, the preprocessed test question image data is input into a trained illustration detection model. For example, the illustration detection model can be a YOLOv5 object detection model, which has very high detection efficiency and good detection results in the field of object detection. Forward inference is then performed in the YOLOv5 object detection model based on the preprocessed test question image data. The YOLOv5 object detection model outputs the location information of the illustrations contained in the test question image based on the preprocessed test question image data. It should be understood that the preprocessing method here is consistent with the image preprocessing module method and will not be described in detail here. Furthermore, the illustration detection model can also be a single-stage object detection algorithm or the CenterNet object detection algorithm, etc.
[0086] In practical applications, the aforementioned illustration location information and the illustrated test question image can also be input into the illustration type determination module, and the illustration image can be obtained from the illustrated test question based on the illustration location information. For example, the corresponding illustration area data can be extracted from the illustrated test question image based on the illustration location information. When the illustration detection module does not detect the illustration image in the illustrated test question image, the entire illustrated test question image is used as the illustration image, and then the initial illustration type is determined based on the illustration image.
[0087] In one possible implementation, the initial illustration type of the exemplary embodiments of this disclosure may be determined by an illustration type determination module based on an illustration image, wherein the illustration type determination module includes a first feature extraction module, a moving-flipping-bottom convolution module, and a second feature extraction module, thereby determining the initial illustration type based on the illustration image.
[0088] For example, the inset type module can be an EfficientNet model. When an inset image is input into a trained EfficientNet model, the shallow image features of the inset image are first extracted based on the first feature extraction module; then, the shallow image features are extracted based on the moving-flipping bottleneck convolution module to obtain the depth image features; finally, the depth image features are extracted based on the second feature extraction module and forward inference is performed to obtain the initial inset type. The inset type here can include coordinate graph type, line graph type, or geometric image type, etc. It should be understood that the inset type module can also be a MobileNet model, a SqueezeNet model, etc.
[0089] In one alternative approach, the moving-flipping bottleneck convolution first compresses the feature map corresponding to the shallow image features and then performs global average pooling along the channel dimension to obtain the global image features of the feature map corresponding to the shallow image features along the channel dimension. Next, it activates the global image features to reduce the number of channels and computational cost. Then, it uses a sigmoid activation function to obtain the weights of different channels and multiplies these weights with the feature map corresponding to the input illustration image for feature fusion, thus obtaining the depth image features. The moving-flipping bottleneck convolution module essentially performs attention operations along the channel dimension. This attention mechanism allows the illustration type module to focus more on the channel features containing the most illustration information while suppressing less important channel features, thereby ensuring the completeness and accuracy of extracting the content information contained in the illustration image and guaranteeing the accuracy of the initial illustration type determination. It should be understood that the channel dimension here can be the illustration image height, illustration image width, and the number of illustration image channels, etc. It should also be understood that different illustration types can be labeled with numbers or letters, allowing the illustration type module to directly output the illustration type number corresponding to different illustration types.
[0090] As can be seen, the exemplary embodiment of this disclosure detects the text portion in the illustrated test question image based on an optical character recognition model, and detects the illustration portion in the illustrated test question image based on an illustration detection module and an illustration type determination module. Based on this, the same illustrated test question image is detected from two angles, which ensures accurate detection of both the text portion and the illustration portion in the illustrated test question image, making the judgment of the initial test question type and the initial illustration type more accurate.
[0091] An exemplary embodiment of this disclosure can also stitch together the initial question type, the initial illustration type, and the illustration image to obtain first illustration stitching information; wherein, the first illustration stitching information can be described in matrix form, the matrix having a width of W and a height of H, where W represents the width of the illustration image, and H is determined by the height of the illustration image, the question category number, and the illustration category.
[0092] For example, the initial question type, initial illustration type, and illustration image can be input into the illustration feature extraction model. The illustration image is scaled to a fixed size. The illustration feature extraction model concatenates and converts the initial question type number corresponding to the initial question type, the initial illustration type number corresponding to the illustration type, and the illustration image data corresponding to the illustration image into a W*H matrix data form. Based on this, the concatenated illustration information is input into the illustration feature extraction model to obtain the illustration feature information corresponding to the illustrated question image. This allows the illustration feature extraction model to be trained and to focus on the feature information of the initial question type number and the initial illustration type number. It should be understood that the specific scaling size of the illustration image depends on the actual situation.
[0093] The exemplary embodiments of this disclosure can also standardize the illustration image data corresponding to the illustration image, and input the standardized illustration image data into the trained illustration feature extraction model. The illustration feature extraction model can be a MobileNetV3 model. The MobileNetV3 model has three outputs: a first output, a second output, and a third output. The first output is used to output illustration feature vectors of fixed dimensions, the second output is used to output the initial test question type number, and the third output is used to output the initial illustration type number.
[0094] The MobileNetV3 model described above adds two outputs to the MobileNet model's network structure. Both outputs are fully connected layers, outputting the initial question type number and the initial illustration type number, respectively. Based on this, multi-task training can facilitate better feature extraction for the illustration feature extraction model, while also providing a certain degree of correction to the input initial question type number and initial illustration type number. The question feature information in this exemplary embodiment includes corrected initial question type and corrected initial illustration type, which can also be obtained based on the MobileNetV3 model.
[0095] In practical applications, the illustration feature vector can be further processed. For example, the test question information can include illustration feature information and illustration size information. Based on this, size information can be added to the illustration feature vector, so that the illustrations in the illustrated test questions can be represented more accurately.
[0096] In one optional approach, the illustration feature information and illustration size information can be concatenated to obtain corresponding second illustration concatenation information. This second illustration concatenation information contains multiple identical illustration size parameters, including the illustration aspect ratio parameter and / or the area ratio parameter of the illustration image within the illustrated question. The concatenation rule can be to copy the illustration aspect ratio parameter and the area ratio parameter of the illustration image within the illustrated question N times and concatenate them to the end of the illustration feature vector. Here, N is an empirical value that can be adaptively adjusted based on the dimension of the illustration feature vector in practical applications. Finally, the question information is determined based on each second illustration concatenation information, the corrected initial question type, and the corrected initial illustration type.
[0097] As can be seen, the illustrated images can be separated from the illustrated test questions. This allows for both overall analysis of the illustrated images and analysis of the illustrated images individually. It takes into account not only the initial test question type but also the initial illustration type, and further integrates the illustration size information, thereby more accurately representing the illustration information of the illustrated test questions and ensuring the accuracy and comprehensiveness of illustration detection and illustration feature extraction.
[0098] For example, the illustration feature information corresponding to the illustrated test question image may include the corrected initial test question type, the corrected initial illustration type, and illustration feature information. When the corrected initial test question type, the corrected initial illustration type, and the illustration feature information are correctly combined, a complete test question information can be obtained. One test question information corresponds to one illustrated test question image, that is, one illustrated test question. Based on this, the illustration feature information not only contains the relevant features of the illustration image, but also integrates the initial test question type and the initial illustration type, thereby more accurately analyzing the illustration image, avoiding interference from other factors in the illustrated test question, and thus ensuring the accuracy of the illustration feature information.
[0099] The recommended test questions for the exemplary embodiments of this disclosure can be applied to servers or chips in servers. The methods of the exemplary embodiments of this disclosure are described in detail below with reference to the accompanying drawings.
[0100] Figure 4 A schematic flowchart of a test item recommendation method according to an exemplary embodiment of this disclosure is shown. The test item recommendation method according to an exemplary embodiment of this disclosure includes:
[0101] Step 401: The method for extracting test question information based on the exemplary embodiments of this disclosure determines the test question information of the illustration test question to be searched.
[0102] In practical applications, student users can input images of exam questions to be searched into the server using their user devices. The server first extracts the exam question information using the method described in this exemplary embodiment, determining the exam question information corresponding to the image. Then, it inputs this information into a feature vector search library. The feature vector search library compares the exam question information with the existing exam question information to determine the matching questions. The exam question library includes exam question information for multiple questions determined using the method described in this exemplary embodiment. These multiple questions can include various types of questions, such as multiple-choice questions, application problems, and geometry problems. It should be understood that the feature vector search library used in this exemplary embodiment is the ElasticSearch library. Based on ElasticSearch's distributed file storage, each feature vector can be indexed and used for searching. During queries, feature vectors can be combined to improve exam question search efficiency. In practical applications, other vector search libraries such as Milvus can also be used.
[0103] Step 402: Obtain recommended questions from the question bank based on the question information to be searched. The question information of the question to be searched in this exemplary embodiment can be determined by the second illustration splicing information in the question information extraction method of this exemplary embodiment. It should be understood that the question information to be searched is the question information corresponding to the question to be searched.
[0104] For example, firstly, multiple candidate questions can be obtained from the question bank based on the illustration feature information in the question information to be searched. Then, the similarity between the illustration feature information of each candidate question and the illustration feature information of the question information to be searched is determined. If the similarity between the candidate question and the question information to be searched meets the filtering criteria, the candidate question is determined as a recommended question. If the identity information of the candidate question matches the identity information of the question information to be searched, the similarity between each candidate question and the question information to be searched is updated. It should be understood that the similarity comparison is based on comparing the question information corresponding to the candidate question with the question information corresponding to the question information to be searched.
[0105] In practical applications, the above screening conditions may include the similarity between the illustration feature information of the candidate test question and the illustration feature information of the test question to be searched being greater than or equal to a preset threshold; or the candidate test questions may be sorted in descending order of similarity, with the test question to be searched being the kth candidate test question, where k is the total number of candidate test questions that is greater than 0 and less than or equal to 0; or a combination of both may be used.
[0106] When the filtering criteria include the similarity between the candidate question and the question to be searched being greater than or equal to a preset threshold, for example, when the similarity between the candidate question and the question to be searched is greater than 95%, the candidate question is determined as a recommended question, and the filtered recommended questions are fed back to the user's device.
[0107] When the filtering criteria include sorting candidate questions in descending order of similarity, when the question to be searched is determined to be the kth candidate question in the sorting results, the kth candidate question and the candidate questions in the sorting results before the kth candidate question are determined as recommended questions, and the multiple recommended questions are fed back to the user device, where k is the total number of candidate questions that is greater than 0 and less than or equal to 0.
[0108] When the filtering criteria include both of the above filtering methods, multiple candidate questions can be obtained first based on a preset threshold. Assuming 100 candidate questions are selected, they are ranked according to their similarity to the question to be searched. When the question to be searched is determined to be the 4th candidate question, the 4th candidate question and the three candidate questions ranked before it are designated as recommended questions, and these four recommended questions are fed back to the user's device. It should be understood that in practical applications, the question to be searched may include multiple types of question information. Therefore, when determining similarity, the average similarity calculation method can be directly used to obtain the average similarity of multiple question information for the same question to be searched, and matching can be performed based on the average similarity. Alternatively, the similarity of multiple question information can be accumulated, and the final accumulated result can be used as the similarity result for matching.
[0109] In one optional embodiment, the identity information of the candidate test questions in this disclosure includes at least the initial test question type and the initial illustration type of the candidate test questions. The identity information of the test question information to be searched includes at least the initial test question type and the initial illustration type of the test question information to be searched. Based on this, during the test question search process, after comparing the test question information corresponding to the candidate test questions with the test question information corresponding to the test question information to be searched to obtain multiple recommended test questions, the initial test question type and the initial illustration type corresponding to the candidate test questions and the test question information to be searched can also be compared. When the candidate test question matches the initial test question type corresponding to the test question information to be searched, the similarity between the candidate test question and the test question information to be searched is incremented by t. When the candidate test question does not match the initial test question type corresponding to the test question information to be searched, the similarity between the candidate test question and the test question information to be searched remains unchanged. When the candidate test question matches the initial illustration type corresponding to the test question information to be searched, the similarity between the candidate test question and the test question information to be searched is incremented by t. When the candidate test question does not match the initial illustration type corresponding to the test question information to be searched, the similarity remains unchanged. Based on this, the similarity between each candidate test question and the test question information to be searched is updated. It should be understood that the value of 't' in the similarity increase setting can be set according to the application, and can be 10%, 5%, etc.
[0110] In another optional approach, the identity information of candidate questions also includes the modified initial question type and modified initial illustration type of the candidate questions, and the identity information of the question to be searched also includes the modified initial question type and modified initial illustration type of the question to be searched. Based on this, during the question search process, after comparing the question information corresponding to the candidate questions with the question information corresponding to the question to be searched to obtain multiple recommended questions, the modified initial question type and modified initial illustration type corresponding to the candidate questions and the question to be searched can also be compared. When the candidate question matches the modified initial question type corresponding to the question to be searched, the similarity between the candidate question and the question to be searched is incremented by t; when the candidate question does not match the modified initial question type corresponding to the question to be searched, the similarity remains unchanged. Similarly, when the candidate question matches the modified initial illustration type corresponding to the question to be searched, the similarity between the candidate question and the question to be searched is incremented by t; when the candidate question does not match the modified initial illustration type corresponding to the question to be searched, the similarity remains unchanged. Based on this, the similarity between each candidate question and the question to be searched is updated.
[0111] Figure 5 A flowchart illustrating a test item similarity determination method according to an exemplary embodiment of this disclosure is shown. Figure 5As shown, the information of the test questions to be searched is input into the feature vector search library. The feature vector search library can obtain multiple candidate test questions from the test question library based on the illustration feature information in the test question information. For example, 100 test questions corresponding to multiple illustration feature information with a similarity greater than 90% can be identified as candidate test questions, and the multiple candidate test questions are sorted in descending order of illustration feature information similarity. Then, the initial test question types corresponding to the 100 candidate test questions are compared one by one. When the initial test question type corresponding to the candidate test question is consistent with the initial test question type in the test question information to be searched, the similarity is increased by 10%. When the initial test question type corresponding to the candidate test question is inconsistent with the initial test question type in the test question information to be searched, the similarity remains unchanged. The initial illustration types corresponding to the 100 candidate test questions are compared. When the initial test question type corresponding to the candidate test question is consistent with the initial illustration type in the test question information to be searched, the similarity is increased by 10%. When the initial test question type corresponding to the candidate test question is inconsistent with the initial illustration type in the test question information to be searched, the similarity remains unchanged. The comparison of 10 For the corrected initial question types corresponding to 0 candidate questions, the similarity is increased by 10% when the initial question type of a candidate question matches the corrected initial question type in the search question information; otherwise, the similarity remains unchanged. Similarly, for the corrected initial illustration types corresponding to 100 candidate questions, the similarity is increased by 10% when the initial question type of a candidate question matches the corrected initial illustration type in the search question information; otherwise, the similarity remains unchanged.
[0112] In practical applications, it can simultaneously compare question information, initial question type, initial illustration type, revised initial question type, and revised initial illustration type, or it can compare them sequentially. This updates the similarity between each candidate question and the question to be searched, and then performs a comprehensive ranking. A preset threshold can be set; for example, when 100 candidate questions are selected, the preset similarity threshold is 90%. Therefore, multiple candidate questions with a similarity greater than 90% can be identified as recommended questions and fed back to the user. Furthermore, based on the similarity ranking results of multiple candidate questions, the illustration feature information of multiple candidate questions and the question to be searched can be determined. The 15th candidate question, matched by the similarity of the illustration features in the search query, is recommended as the first candidate question. Multiple candidate questions ranked before the 15th candidate question and those ranked before it in the overall ranking are also recommended to the user. This approach not only recommends the most matching candidate question to the user but also recommends similar questions, enabling efficient question filtering, ensuring accurate search results, and providing related information to broaden the user's problem-solving strategies and improve learning.
[0113] One or more technical solutions provided in the exemplary embodiments of this disclosure can obtain illustration images from illustrated test questions, and obtain test question feature information including at least illustration feature information based on the initial test question type, the initial illustration type, and the illustration image. Therefore, the illustration feature information not only contains relevant features of the illustration image, but also integrates the initial test question type and the initial illustration type, thereby analyzing the illustration image more accurately and avoiding interference from other factors in the illustrated test questions. Then, test question information is obtained based on the illustration feature information. It is evident that the exemplary embodiments of this disclosure can extract illustration images from illustrated test questions, allowing for both overall analysis of the illustrated image and individual analysis of the illustration image. By combining these two analyses, interference from other textual information besides the illustration in the illustrated test question image can be avoided, improving the search capability for illustrated test questions and enhancing the accuracy of test question information extraction. Based on this, when test question information for illustrated test questions to be searched is determined based on the method of the exemplary embodiments of this disclosure, and recommended test questions are obtained from the test question database based on the search question information, test question filtering can be performed efficiently, improving the accuracy of test question search, thereby providing users with accurate test question recommendations.
[0114] The foregoing primarily describes the solutions provided by the embodiments of this disclosure from the perspective of the server. It is understood that, in order to implement the above functions, the server includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0115] This disclosure embodiment can divide the server into functional units according to the above method example. For example, it can divide each function into a separate functional module, or it can integrate two or more functions into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this disclosure embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.
[0116] In the case of dividing each functional module according to its corresponding function, an exemplary embodiment of this disclosure provides a test question information extraction device, which can be a server or a chip applied to a server. Figure 6 A schematic block diagram of the functional modules of a test question information extraction device according to an exemplary embodiment of the present disclosure is shown. Figure 6 As shown, the test question information extraction device 600 includes:
[0117] A device for determining the degree of dynamic knowledge mastery, characterized in that the device comprises:
[0118] Module 601 is used to determine the initial question type, initial illustration type, and illustration image of the illustrated question;
[0119] The module 602 is configured to obtain test question feature information based on the initial test question type, the initial illustration type, and the illustration image, wherein the test question feature information includes at least illustration feature information;
[0120] The obtaining module 602 is also used to obtain test question information based on the illustration feature information.
[0121] In one possible implementation, the determining module 601 is further configured to determine the initial question type, initial illustration type, and illustration image of the illustrated question based on the illustrated question image; determine illustration location information based on the illustrated question image; obtain the illustration image from the illustrated question based on the illustration location information; and determine the initial illustration type based on the illustration image.
[0122] In one possible implementation, the determination module 601 is further configured to identify the test text information of the illustrated test question image based on the test question image using a text recognition model; determine the word vectors of multiple sub-texts contained in the test question text information based on the test question text information; and process the word vectors of the multiple sub-texts based on a text classification model to obtain the initial test question type.
[0123] In one possible implementation, the initial illustration type is determined by an illustration type determination module based on the illustration image. The illustration type determination module includes a first feature extraction module, a move-flip bottleneck convolution module, and a second feature extraction module. The initial illustration type is determined based on the illustration image. The obtaining module 602 is further configured to extract shallow image features of the illustration image based on the first feature extraction module; extract the shallow image features based on the move-flip bottleneck convolution module to obtain depth image features; and extract the depth image features based on the second feature extraction module to obtain the initial illustration type.
[0124] In one possible implementation, the step of obtaining test question feature information based on the initial test question type, the initial illustration type, and the illustration image, and the obtaining module 602 is further configured to stitch the initial test question type, the initial illustration type, and the illustration image together to obtain first illustration stitching information; and input the illustration stitching information into the illustration feature extraction model to obtain the corresponding illustration feature information.
[0125] In one possible implementation, the first illustration stitching information is described in matrix form, the matrix having a width of W and a height of H, where W represents the width of the illustration image, and H is determined by the height of the illustration image, the question category number, and the illustration category.
[0126] In one possible implementation, the question information includes the illustration feature information and illustration size information.
[0127] In one possible implementation, the test question feature information includes a corrected initial test question type and a corrected initial illustration type. The obtaining module 602 is further configured to splice the illustration feature information with the illustration size information to obtain corresponding second illustration splicing information; and to determine test question information based on each second illustration splicing information, the corrected initial test question type, and the corrected initial illustration type.
[0128] In one possible implementation, the second illustration stitching information contains multiple identical illustration size information, including illustration aspect ratio parameters and / or area ratio parameters of the illustration image in the illustrated question.
[0129] In the case of dividing each functional module according to its corresponding function, an exemplary embodiment of this disclosure provides a test question recommendation device, which can be a server or a chip applied to a server. Figure 7 A schematic block diagram of the functional modules of a test item recommendation device according to an exemplary embodiment of the present disclosure is shown. Figure 7 As shown, the test question recommendation device 700 includes:
[0130] The determining module 701 determines the question information of the illustrated test questions to be searched based on the method described in the exemplary embodiments of this disclosure;
[0131] The acquisition module 702 is used to acquire recommended test questions from the test question bank based on the test question information to be searched.
[0132] In one possible implementation, the step of obtaining recommended questions from the question bank based on the question information to be searched, the obtaining module 702 is further configured to obtain multiple candidate questions from the question bank; determine the similarity between each candidate question and the question information to be searched; if the similarity between the candidate question and the question information to be searched meets the filtering conditions, determine the candidate question as a recommended question; wherein, if the identity information of the candidate question matches the identity information of the question information to be searched, update the similarity between each candidate question and the question information to be searched.
[0133] In one possible implementation, the filtering conditions include: the similarity between the candidate question and the information of the question to be searched is greater than or equal to a preset threshold; and / or the candidate questions are sorted in descending order of similarity, where the question to be searched is the kth candidate question, and k is the total number of candidate questions that is greater than 0 and less than or equal to 0.
[0134] In one possible implementation, the question bank includes question information of a plurality of candidate questions determined by the method described in the exemplary embodiments of this disclosure; the identity information of the candidate questions includes at least the initial question type and initial illustration type of the candidate questions, and the identity information of the question information to be searched includes at least the initial question type and initial illustration type of the question information to be searched.
[0135] In one possible implementation, the number of recommended test questions is multiple, and the test question information of the test question to be searched is determined by an exemplary embodiment of this disclosure; the identity information of the candidate test questions also includes the modified initial test question type and the modified initial illustration type of the candidate test questions, and the identity information of the test question to be searched also includes the modified initial test question type and the modified initial illustration type of the test question to be searched.
[0136] Figure 8 A schematic block diagram of a chip according to an exemplary embodiment of the present disclosure is shown. Figure 8 As shown, the chip 800 includes one or more processors 801 and a communication interface 802. The communication interface 802 can support the server in performing the data transmission and reception steps in the above method, and the processor 801 can support the server in performing the data processing steps in the above method.
[0137] Optional, such as Figure 8 As shown, the chip 800 also includes a memory 803, which may include read-only memory and random access memory, and provides operation instructions and data to the processor. A portion of the memory may also include non-volatile random access memory (NVRAM).
[0138] In some implementations, such as Figure 8 As shown, processor 801 executes corresponding operations by calling operation instructions stored in memory (which may be stored in the operating system). Processor 801 controls the processing operations of any terminal device; processor can also be called a central processing unit (CPU). Memory 803 may include read-only memory and random access memory, and provides instructions and data to processor 801. A portion of memory 803 may also include NVRAM. For example, in applications, memory, communication interfaces, and other components are coupled together via a bus system, which may include, in addition to a data bus, a power bus, a control bus, and a status signal bus, etc. However, for clarity, in... Figure 8 The general labeled all buses as Bus System 804.
[0139] The methods disclosed in the embodiments of this disclosure can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above methods can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.
[0140] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.
[0141] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.
[0142] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.
[0143] refer to Figure 9The present invention describes a structural block diagram of an electronic device that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0144] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0145] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, output unit 907, storage unit 908, and communication unit 909. Input unit 906 can be any type of device capable of inputting information to electronic device 900. Input unit 906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 908 may include, but is not limited to, disk and optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0146] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the methods of the exemplary embodiments of this disclosure can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the methods of the exemplary embodiments of this disclosure by any other suitable means (e.g., by means of firmware).
[0147] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0148] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0149] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0150] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0151] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0152] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this disclosure are performed, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
[0153] Although this disclosure has been described in conjunction with specific features and embodiments, it will be apparent that various modifications and combinations can be made therein without departing from the spirit and scope of this disclosure. Accordingly, this specification and drawings are merely exemplary illustrations of the disclosure as defined by the appended claims and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this disclosure. It is obvious that those skilled in the art can make various alterations and modifications to this disclosure without departing from its spirit and scope. Thus, this disclosure is also intended to include any such modifications and modifications that fall within the scope of the claims of this disclosure and their equivalents.
Claims
1. A method for extracting test question information, characterized in that, The method includes: Determine the initial question type, initial illustration type, and illustration image for illustrated questions; Based on the initial question type, the initial illustration type, and the illustration image, question feature information is obtained, wherein the question feature information includes at least illustration feature information; Test question information is obtained based on the illustration feature information; The step of obtaining test question feature information based on the initial test question type, the initial illustration type, and the illustration image includes: The initial question type, the initial illustration type, and the illustration image are stitched together to obtain first illustration stitching information; wherein, the first illustration stitching information is described in matrix form, the matrix has a width of W and a height of H, where W represents the width of the illustration image, and H is determined by the height of the illustration image, the question category number, and the illustration category; The first illustration stitching information is input into the illustration feature extraction model to obtain the corresponding illustration feature information; The test question information includes the illustration feature information and illustration size information. The test question feature information also includes correcting the initial test question type and correcting the initial illustration type. The method further includes: The illustration feature information and illustration size information are combined to obtain the corresponding second illustration combination information; The test question information is determined based on the splicing information of each second illustration, the corrected initial test question type, and the corrected initial illustration type.
2. The method according to claim 1, characterized in that, The process of determining the initial question type, initial illustration type, and illustration image for illustrated questions includes: The initial question type is determined based on the illustrated question image; Determine the location information of the illustrations based on the image of the illustrated test question; The illustration image is obtained from the illustrated test question based on the illustration location information; The initial illustration type is determined based on the illustration image.
3. The method according to claim 2, characterized in that, The step of determining the corresponding initial question type based on the illustrated question image includes: The test text information of the illustrated test question image is identified based on a text recognition model; Based on the test question text information, determine the word vectors of multiple sub-texts contained in the test question text information; The initial question type is obtained by processing the word vectors of multiple sub-texts based on a text classification model.
4. The method according to claim 2, characterized in that, The initial illustration type is determined by an illustration type determination module based on the illustration image. The illustration type determination module includes a first feature extraction module, a shift-flip bottleneck convolution module, and a second feature extraction module. Determining the initial illustration type based on the illustration image includes: The shallow image features of the illustration image are extracted based on the first feature extraction module; The shallow image features are extracted based on the moving flip bottleneck convolution module to obtain the depth image features; The initial illustration type is obtained by extracting the depth image features based on the second feature extraction module.
5. The method according to claim 1, characterized in that, The second illustration splicing information contains multiple identical illustration size information, which includes the illustration aspect ratio parameter and / or the area ratio parameter of the illustration image in the illustrated test question.
6. A test item recommendation method, characterized in that, include: The search question information is determined based on the method described in any one of claims 1 to 5; Recommended questions are obtained from the question bank based on the information of the question to be searched.
7. The method according to claim 6, characterized in that, The step of obtaining recommended test questions from the test question bank based on the test question information to be searched includes: Retrieve multiple candidate questions from the question bank; Determine the similarity between each candidate question and the information of the question to be searched; If the similarity between the candidate question and the information of the question to be searched meets the screening criteria, the candidate question is determined as a recommended question. If the identity information of the candidate question matches the identity information of the question to be searched, the similarity between each candidate question and the question to be searched is updated.
8. The method according to claim 7, characterized in that, The filtering criteria include: the similarity between the candidate questions and the information of the questions to be searched is greater than or equal to a preset threshold; and / or The candidate questions are sorted in descending order of similarity, and the question to be searched is the k-th candidate question, where k is the total number of candidate questions that is greater than 0 and less than or equal to 0.
9. The method according to claim 7, characterized in that, The question bank includes question information of a plurality of candidate questions determined by the method of any one of claims 1 to 5; The identity information of the candidate test questions includes at least the initial test question type and the initial illustration type of the candidate test questions, and the identity information of the test question information to be searched includes at least the initial test question type and the initial illustration type of the test question information to be searched.
10. The method according to claim 9, characterized in that, The question information of the illustration to be searched is determined by the second illustration splicing information in claim 6 or 7; The identity information of the candidate test questions also includes the modified initial test question type and the modified initial illustration type of the candidate test questions, and the identity information of the test questions to be searched also includes the modified initial test question type and the modified initial illustration type of the test questions to be searched.
11. A device for extracting test question information, characterized in that, The device includes: The determination module is used to determine the initial question type, initial illustration type, and illustration image for illustrated questions; The module is configured to obtain test question feature information based on the initial test question type, the initial illustration type, and the illustration image, wherein the test question feature information includes at least illustration feature information; The obtaining module is also used to obtain test question information based on the illustration feature information; The obtaining module is further configured to stitch together the initial question type, the initial illustration type, and the illustration image to obtain first illustration stitching information; wherein, the first illustration stitching information is described in matrix form, the matrix has a width of W and a height of H, where W represents the width of the illustration image, and H is determined by the height of the illustration image, the question category number, and the illustration category; The first illustration stitching information is input into the illustration feature extraction model to obtain the corresponding illustration feature information; The test question information includes the illustration feature information and illustration size information. The test question feature information also includes correcting the initial test question type and correcting the initial illustration type. The obtaining module is further used for: The illustration feature information and illustration size information are combined to obtain the corresponding second illustration combination information; The test question information is determined based on the splicing information of each second illustration, the corrected initial test question type, and the corrected initial illustration type.
12. A test question recommendation device, characterized in that, The device includes: The determining module determines the search question information of the illustration question to be searched based on the method described in any one of claims 1 to 5; The acquisition module is used to acquire recommended test questions from the test question bank based on the test question information to be searched.
13. An electronic device, characterized in that, include: processor; And, the memory for storing programs; The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 10.
14. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing the computer to perform the method according to any one of claims 1 to 10.