Page generation method and device, and exercise book generation method and device
The automated generation of workbooks using a layout generation model solves the problem of high manpower and time costs in educational supplement publishing, achieving efficient and low-cost workbook generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 深圳市星桐科技有限公司
- Filing Date
- 2022-08-11
- Publication Date
- 2026-05-12
AI Technical Summary
In the current technology, publishing educational supplements requires a lot of manpower and time, is costly and has a long cycle, and cannot efficiently generate workbooks under different teaching syllabi.
By using a layout generation model to process layout images, a target layout including background images and target area location information is generated, and questions are arranged in the target area, thus achieving automated workbook generation.
It reduced typesetting costs, improved the efficiency and speed of workbook generation, reduced manual intervention, and shortened the publication cycle.
Smart Images

Figure CN115293104B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method and apparatus for generating page layouts and a method and apparatus for generating workbooks. Background Technology
[0002] Practice constitutes a large proportion of students' learning. Based on different teaching syllabi and different knowledge focuses, a large number of workbook publishing companies produce and publish a large number of different teaching aids. The publication of a teaching aid requires a large number of teaching and research teachers to compile questions based on knowledge points, and graphic designers to perform layout and other operations. Overall, it requires a lot of manpower, the cost is relatively high, and the publication cycle is relatively long and constrained by manpower. Summary of the Invention
[0003] According to one aspect of this disclosure, a method for generating page layouts is provided, comprising:
[0004] Get the layout image, which records one or more typeset objects;
[0005] The layout image is processed using a layout generation model to generate a target layout, which includes a background image and positional information of one or more target regions on the background image, each target region corresponding to a target object.
[0006] According to another aspect of this disclosure, a method for generating workbooks is provided, comprising:
[0007] Retrieve a set of questions, which includes multiple questions;
[0008] Multiple target layouts are generated using the layout generation method disclosed herein, wherein each target layout includes a background image and positional information of one or more target regions on the background image;
[0009] The above questions are arranged in the target areas of multiple target pages to obtain a workbook, where each target area corresponds to one question.
[0010] According to another aspect of this disclosure, a layout generation apparatus is provided, comprising:
[0011] The acquisition module is used to acquire the layout image, which records one or more objects that have been formatted.
[0012] The generation module is used to process the layout image using the layout generation model to generate a target layout, wherein the target layout includes a background image and the position information of one or more target regions on the background image, and each target region corresponds to a target object.
[0013] According to another aspect of this disclosure, a workbook generating apparatus is provided, comprising:
[0014] The acquisition module is used to acquire a set of questions, which includes multiple questions.
[0015] The layout generation module is used to generate multiple target layouts using the layout generation method disclosed herein, wherein each target layout includes a background image and position information of one or more target regions on the background image;
[0016] The layout module is used to arrange the above multiple questions in target areas of multiple target pages to obtain a workbook, where each target area corresponds to one question.
[0017] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0018] Processor; and
[0019] Stored program memory,
[0020] The program includes instructions that, when executed by the processor, cause the processor to perform the methods of this disclosure.
[0021] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the methods of this disclosure.
[0022] One or more technical solutions provided in the embodiments of this application use a layout generation model to process a layout image that records one or more typeset objects, and generate a target layout including a background image and one or more target areas on the background image. Each target area corresponds to a target object, which can realize the typesetting of target objects and reduce typesetting costs. Attached Figure Description
[0023] Further details, features, and advantages of this disclosure are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:
[0024] Figure 1 A flowchart of a layout generation method according to an exemplary embodiment of the present disclosure is shown;
[0025] Figure 2 A schematic block diagram of a layout generation model according to an exemplary embodiment of the present disclosure is shown;
[0026] Figure 3 A schematic block diagram of a layout generation model according to an exemplary embodiment of the present disclosure is shown;
[0027] Figure 4A flowchart illustrating a method for training a first sub-network according to an exemplary embodiment of the present disclosure is shown;
[0028] Figure 5 A schematic block diagram of a first generative model according to an exemplary embodiment of the present disclosure is shown;
[0029] Figure 6 A flowchart illustrating a method for training a second sub-network according to an exemplary embodiment of the present disclosure is shown;
[0030] Figure 7 A schematic block diagram of a second generative model according to an exemplary embodiment of the present disclosure is shown;
[0031] Figure 8 A flowchart of a workbook generation method according to an exemplary embodiment of the present disclosure is shown;
[0032] Figure 9 A schematic block diagram of a layout generation apparatus according to an exemplary embodiment of the present disclosure is shown;
[0033] Figure 10 A schematic block diagram of a workbook generation apparatus according to an exemplary embodiment of the present disclosure is shown;
[0034] Figure 11 A schematic block diagram of a system according to an exemplary embodiment of the present disclosure is shown;
[0035] Figure 12 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0036] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0037] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0038] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0039] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0040] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0041] Before describing the solution disclosed herein, the technical aspects involved in this disclosure are explained as follows.
[0042] Text detection has a wide range of applications and is a prerequisite for many computer vision tasks, such as image search, text recognition, identity authentication, and visual navigation. The main purpose of text detection is to locate the position of text lines or characters in an image. Accurate text localization is both important and challenging because, compared to general object detection, text has characteristics such as multiple directions, irregular shapes, extreme aspect ratios, and diverse fonts, colors, and backgrounds. Therefore, algorithms that are successful in general object detection often cannot be directly transferred to text detection. However, with the resurgence of deep learning in recent years, research on text detection has become a hot topic, resulting in a large number of methods specifically for text detection, all of which have achieved good detection results. Based on the technical characteristics of the methods used in text detection, currently popular text detection methods can be broadly divided into two categories. The first category is sliding window-based text detection methods. These methods are based on the general idea of object detection, setting a large number of anchor boxes with different aspect ratios and sizes. Using these anchor boxes as a sliding window, they traverse and search the image or the feature map obtained from convolutional operations on the image. For each searched location, they classify whether the box contains text. The advantage of this method is that after the text box is identified, no further processing is needed. The disadvantage is the excessive computational cost. Not only does it require a large amount of computing resources, but it also takes a long time. The second type is based on the calculation of connected components, also known as the segmentation-based method. It mainly uses a fully convolutional neural network model to extract image features, then binarizes the feature map and calculates its connected components. Then, depending on the application scenario (i.e., different training datasets), it uses some specific methods to determine the position of text lines. The advantage of this method is that it is fast and has a small amount of computation. The disadvantage is that the post-processing steps are cumbersome and involve a lot of computation and optimization. This not only consumes a lot of time, but also the rationality and effectiveness of the post-processing strategy strictly limits the performance of the algorithm.
[0043] Text recognition is the process of identifying character sequences (for Chinese, one character is a Chinese character; for English, one character is a letter) from images containing text. It is an extremely challenging subject. Besides complex image backgrounds and varying lighting, the complexity of the output space is also a major difficulty. Since text is composed of a variable number of letters, text recognition in natural scenes requires identifying sequences of varying lengths from images. Currently, there are two main approaches: one is a bottom-up strategy, breaking down the recognition problem into character detection, character recognition, and character combination, solving them one by one; the other is a holistic analysis strategy, i.e., a sequence-to-sequence approach, first encoding the image and then decoding the sequence to directly obtain the entire string. While the first method is effective, it requires character-level annotation, meaning the position and information of each character in the input image need to be labeled, which is labor-intensive. The second method, while simpler to annotate and capable of transcribing strings, may result in either over- or under-identified characters.
[0044] The paper "Attention is all you need" proposes a famous classic network structure, the Transformer, which consists of an encoder and a decoder. The decoder comprises multiple stacked basic modules, which mainly consist of multi-head self-attention layers, skip connections, layer normalization, and feedforward neural networks. The decoder also comprises multiple basic modules, which differ from the first module in that they include two multi-head self-attention layers. The Transformer design not only greatly accelerates network training and inference time but also effectively improves the accuracy of various tasks. Originally designed for natural language understanding tasks, it is now widely used in computer vision tasks due to its excellent performance, achieving remarkable results in multiple tasks.
[0045] Residual networks (ResNet) are a well-known type of image classification network for natural scenes. They effectively address the performance degradation that occurs as the number of network layers increases, allowing for more complex feature pattern extraction with increased layer depth. The core of ResNet is a structure called a residual block. The main characteristic of the residual structure is its cross-layer skip connection; a residual block consists of multiple convolutional layers. The output of the input after passing through the residual block is added to the input channel-wise and point-wise. This is equivalent to the input having two branches: one passing through the residual block, and the other directly bypassing it. Finally, the two branches are merged. ResNet has several well-known structures with different numbers of convolutional layers, such as 18, 34, 50, 101, and 152. In addition, there are various variants such as ResNext, all of which achieve good results in natural scene image classification.
[0046] CenterNet is an anchor-free method for general object detection, which can be considered a regression-based method. Its general idea is to first define the total number of object categories N, and then output N+2+2 channels. It only predicts the center point of the object, outputting a score image for each category (each pixel's value is between 0 and 1, representing the probability that the point is the center of a certain category). Therefore, there are N score images. Because the predicted center point cannot be guaranteed to be the true center point during prediction (as offsets often occur), two additional channels are used to predict the center point's offset (one for the x-axis and one for the y-axis). In addition, the remaining two channels are used to predict the distances of the center point from the left and top borders of the bounding box. The actual post-processing involves finding the possible center points of the object in the score images by setting a threshold, then correcting the center points based on their corresponding x and y offsets, and finally obtaining the bounding box directly from the center points and the predicted width and height. For example, the above offset is explained as follows: Suppose the width and height of the original image are W and H, and the size of the predicted feature map is W / 4 and H / 4. Then, a point (10,10) in the original image corresponds to a point (2.5,2.5) in the feature map. However, since the image is discrete and its coordinates are integer values, we round up. (10,10) corresponds to (2,2). Therefore, the offset of the center point in the feature map relative to the original image is (0.5,0.5).
[0047] Visual Image Processing (VAE) is an important generative model consisting of an encoder and a decoder. It typically uses the lower bound of the log-likelihood as its optimization objective, so the loss function of a VAE model generally consists of reconstruction loss and cross-entropy loss. A VAE encodes the input through the encoder and then inputs the encoded data into the decoder to reconstruct the original input. In most cases, the reconstructed image is very close to the original image. VAE models are stable to train and faster. The encoding that a VAE transforms into the input may be parameters of a certain distribution or a feature map, etc.
[0048] The present disclosure is described below with reference to the accompanying drawings.
[0049] This disclosure provides an exemplary embodiment of a layout generation method that can be applied to electronic devices such as servers and / or clients.
[0050] Figure 1 A flowchart of a layout generation method according to an exemplary embodiment of the present disclosure is shown, such as... Figure 1 As shown, the layout generation method includes steps S101 to S102.
[0051] Step S101: Obtain a layout image, which records one or more typeset objects.
[0052] Step S102: Process the layout image using the layout generation model to generate a target layout, wherein the target layout includes a background image and position information of one or more target regions on the background image, and each target region corresponds to a target object.
[0053] In this exemplary embodiment, the target region includes, but is not limited to, a rectangular region. As an example, the location information of the target region may include the coordinates of the center point, the length and width of the target bounding box.
[0054] In one implementation, the object and target object may include characters. Characters may include at least one or any combination of Chinese characters, letters, numbers, operators, punctuation marks, and other symbols. In one implementation, the object and target object may also include images. In one implementation, the object and target object may be a combination of images and characters.
[0055] As an example, the layout generation method is used for newspaper typesetting. In step S101 above, the layout image is an image of a published newspaper, and the objects on the layout image may include news articles, advertisements, etc. For example, each news article is considered as an object, which may include an image and one or more text paragraphs; another example is a set of news articles (usually a set of short news flashes, etc.) as an object. Correspondingly, the target objects may include news articles, advertisements, etc. Based on the published newspaper, a target layout (the newspaper to be published) is generated, with the target objects being the news articles, advertisements, etc. to be published.
[0056] As an example, the layout generation method is used for magazine typesetting. In step S101 above, the layout image is a page of a published magazine, and the objects on the layout image may include pictures, text paragraphs, a set of related text, etc. Based on the page of the published magazine, a target layout (the page of the magazine to be published) is generated, and the target objects are the pictures, text paragraphs, etc. to be published.
[0057] It should be understood that the layout generation method of this embodiment can also be used for the layout of posters, flyers, web pages, etc., and this embodiment does not limit it in this way.
[0058] As one embodiment, in the target layout, each target area can correspond to one target object. Target objects are arranged in target areas on the background image of the target layout, with each target area corresponding to one target object.
[0059] In some embodiments, such as Figure 2As shown, the layout generation model includes: a first sub-network, which includes a first encoder and a first decoder, and is configured to process a layout image to generate a background image; and a second sub-network, which includes a second encoder and a second decoder, and is configured to process the layout image to generate position information of one or more target regions on the background image.
[0060] As one implementation method, such as Figure 3 As shown, the first decoder includes at least two deconvolutional layers. Optionally, as... Figure 3 As shown, the first encoder includes a residual network, such as ResNet 18.
[0061] As one implementation method, such as Figure 3 As shown, the second decoder includes at least one convolutional layer and at least one fully connected layer. Optionally, as... Figure 3 As shown, the second encoder includes a residual network, such as ResNet 18.
[0062] In some embodiments, a first sub-network and a second sub-network of the layout generation model are trained respectively. Exemplary embodiments of the methods for training the first sub-network and the second sub-network are described below.
[0063] Figure 4 A flowchart illustrating a method for training a first sub-network according to an exemplary embodiment of this disclosure is shown, as follows: Figure 4 As shown, the method includes steps S401 to S402.
[0064] Step S401: Obtain first training data, wherein the first training data includes a first image and its first annotation information. The first image records one or more objects arranged in a specific layout. The first annotation information includes an image in which pixels outside the object area on the first image and pixels in the area where the first element in the object is located are set to preset values.
[0065] Step S402: Train the first sub-network based on the first training data.
[0066] Optionally, in step S402 above, the L1 loss function can be used to train the first sub-network.
[0067] In some possible implementations, before training the first sub-network based on the first training data, the method further includes: training a first generative model including a first encoder based on the first training data to pre-train the first encoder and obtain the network parameters of the first encoder. Optionally, an L1 loss function is used when training the first generative model. In step S402 above, the first sub-network is trained based on the first training data, and the pre-trained first encoder is fixed during the training process to train the first decoder and obtain the network parameters of the first decoder. Thus, the network parameters of the first sub-network are obtained.
[0068] The VAE model encodes the input using an encoder, then feeds the encoded data into a decoder. The decoder reconstructs the input, and in most cases, the reconstructed image closely resembles the original. Compared to other generative models, the VAE model is more stable to train and faster. As an example, the first generative model is the VAE model.
[0069] For example, such as Figure 5 As shown, the first generative model includes a first encoder and a third decoder, and the third decoder includes multiple convolutional layers and multiple deconvolutional layers. Figure 5 The diagram shows 3 convolutional layers and 4 deconvolutional layers. Further reference... Figure 5 As shown, the output of each stage of the first encoder and the four deconvolution layers of the decoder use a U-net-like structure. Specifically, the result of the deconvolution layer is connected to the output of the sub-module (stage) with the same resolution in the encoder and used as the input of the next sub-module (deconvolution layer) in the decoder.
[0070] In some possible implementations, the method for training the first sub-network further includes: processing a first image using a detection model to generate first location information of one or more objects on the first image, and second location information of first elements within the objects; processing the first image based on the first and second location information to obtain first annotation information. This generates the aforementioned first training data, reducing training costs.
[0071] Figure 6 A flowchart illustrating a method for training a second sub-network according to an exemplary embodiment of this disclosure is shown, such as... Figure 6 As shown, the method includes steps S601 to S602.
[0072] Step S601: Obtain second training data, wherein the second training data includes a second image and its second annotation information, the second image records the bounding boxes of one or more objects arranged in a layout, and the second annotation information includes the position information of the bounding boxes of one or more objects.
[0073] In one implementation, the second image consists of a solid color background (e.g., white) and a region bounding box (e.g., black, green, etc.). The region bounding box may include, but is not limited to, a rectangle, and its positional information includes, but is not limited to, the center point, width, and height of the region bounding box.
[0074] Step S602: Train the second sub-network based on the second training data.
[0075] Optionally, in step S602 above, the L1 loss function can be used to train the second sub-network.
[0076] In some possible implementations, before training the second sub-network based on the second training data, the method further includes: training a second generative model including a second encoder based on the second training data to pre-train the second encoder and obtain the network parameters of the second encoder. Optionally, an L1 loss function is used when training the second generative model. In step S602 above, the second sub-network is trained based on the second training data, and the network parameters of the pre-trained second encoder are fixed during the training process. The second decoder is then trained to obtain the network parameters of the second decoder, thereby training the network parameters of the second sub-network.
[0077] As an example, the second generative model is the VAE model.
[0078] For example, such as Figure 7 As shown, the second generative model includes a second encoder and a fourth decoder, the fourth decoder comprising multiple convolutional layers. Figure 7 (As shown in the diagram, there are 3 convolutional layers).
[0079] In some possible implementations, the method for training the second sub-network further includes: processing the third image using a detection model to generate third location information of one or more objects recorded on the third image; generating region boxes corresponding to the third location information on the solid color image to obtain the second image; and using the third location information of the one or more objects as second annotation information of the second image.
[0080] This exemplary embodiment also provides a method for generating workbooks, which can be applied to electronic devices such as servers and / or clients. This exemplary embodiment's workbook generation method can format questions to generate workbooks, reducing typesetting costs compared to manual typesetting.
[0081] Figure 8 A flowchart of a workbook generation method according to an exemplary embodiment of the present disclosure is shown, such as... Figure 8 As shown, the workbook generation method includes steps S801 to S803.
[0082] Step S801: Obtain the question set, which includes multiple questions.
[0083] Step S802: Generate multiple target layouts using a layout generation method, wherein each target layout includes a background image and position information of one or more target regions on the background image.
[0084] In this exemplary embodiment, the layout generation method is described in the foregoing description of this disclosure and will not be repeated here.
[0085] In step S802, the multiple target layouts mentioned above are generated using multiple layout images. The layout images can be pages from an already published workbook, and each page can contain one or more questions.
[0086] As one implementation, the multiple page images for generating multiple target pages in step S802 can be selected from a pre-collected set of pages from published workbooks. For example, they can be selected randomly.
[0087] In one implementation, each page image may correspond to a target page.
[0088] In one implementation, each target layout may correspond to a page in the workbook.
[0089] In this exemplary embodiment, the target region includes, but is not limited to, a rectangular region. As an example, the location information of the target region may include the coordinates of the center point, the length and width of the target bounding box.
[0090] Step S803: Arrange the above-mentioned multiple questions in the target areas of multiple target pages to obtain a workbook, wherein each target area corresponds to one question.
[0091] As one implementation method, parameters such as the font size of the question content are determined based on the size of the target area, and the questions are arranged in the target area.
[0092] The background image of each target page may include multiple target areas, which may differ in size and shape. As one implementation, the appropriate target area can be selected based on the content characteristics (e.g., number of words, length, etc.) of the questions arranged on the target page.
[0093] In this exemplary embodiment, the question may be one question or a group of questions (also referred to as a large question). For example, addition practice questions may include a group of addition expressions; for example, word problems may include a single question (e.g., there are 7 sheep on the grassland, and then 6 more come. How many sheep are there in total?). This embodiment does not limit the content of the questions.
[0094] In some embodiments, the question set is determined based on the association between questions and knowledge point categories. As one implementation, obtaining the question set in step S801 includes: receiving a workbook generation request, obtaining questions from a question bank based on the workbook generation request, and obtaining the question set. The information carried in the workbook generation request may include: one or more specified knowledge point categories, and obtaining questions corresponding to the knowledge point categories based on the association between questions and knowledge point categories. Optionally, the information carried in the workbook generation request may also include: the number of questions for each knowledge point category. A corresponding number of questions are obtained for each knowledge point category.
[0095] As one implementation method, the method for establishing the association between questions and knowledge point categories includes: obtaining question content; processing the question content using a question classification model to output the knowledge point category corresponding to the question content, wherein the knowledge point category belongs to multiple preset knowledge point categories.
[0096] As an example, the question classification model is a Transformer model, which classifies strings into predefined knowledge point categories. Specifically, the question classification model includes: an encoder, configured to encode the string corresponding to the question content to generate a feature vector; and a decoder, configured to decode the feature vector to output the knowledge point category corresponding to the question content. Compared to image-based classification, classification based on the string corresponding to the question content has higher accuracy.
[0097] In a typical Transformer model, the encoder, given a sequence (x1…xn) as input, maps it to a hidden layer H (h1…hn); the decoder, taking H as input, generates the decoded result (y1…ym) character by character. In this example, m is set to 1, meaning the entire string corresponding to the question content is decoded once to obtain the knowledge category corresponding to the output question content.
[0098] In this exemplary embodiment, the question can be identified from the question image, but is not limited thereto.
[0099] An exemplary embodiment of this disclosure provides a layout generation apparatus. In some embodiments, the layout generation apparatus may be implemented as computer program instructions.
[0100] Figure 9 A schematic block diagram of a layout generation apparatus according to exemplary embodiments of the present disclosure is shown, such as Figure 9As shown, the layout generation device 900 includes: an acquisition module 910 for acquiring a layout image, the layout image recording one or more typeset objects; and a generation module 920 for processing the layout image using a layout generation model to generate a target layout, wherein the target layout includes a background image and one or more target regions on the background image, each target region corresponding to a target object.
[0101] In some embodiments, the layout generation model includes: a first sub-network, which includes a first encoder and a first decoder and is configured to process the layout image to generate a background image; and a second sub-network, which includes a second encoder and a second decoder and is configured to process the layout image to generate position information of one or more target regions on the background image.
[0102] In one implementation, the first decoder includes at least two deconvolutional layers.
[0103] In one implementation, the second decoder includes at least one convolutional layer and at least one fully connected layer.
[0104] In one implementation, the first sub-network is trained by the following method: acquiring first training data, wherein the first training data includes a first image and its first annotation information, the first image recording one or more objects arranged in a specific layout, and the first annotation information including an image in which pixels outside the object area on the first image and pixels in the area where the first element in the object is located are set to preset values; and training the first sub-network based on the first training data.
[0105] As one implementation, before training the first sub-network based on the first training data, the method further includes: training a first generative model including a first encoder based on the first training data to pre-train the first encoder; wherein training the first sub-network based on the first training data includes: training the first sub-network based on the first training data, wherein the pre-trained first encoder is fixed during the training process to train the first decoder.
[0106] As one implementation, training the first sub-network further includes: processing the first image using a detection model to generate first location information of one or more objects on the first image, and second location information of first elements in the objects; processing the first image based on the first location information and the second location information to obtain first annotation information.
[0107] In one implementation, the second sub-network is trained by: acquiring second training data, wherein the second training data includes a second image and its second annotation information, the second image recording the bounding boxes of one or more objects arranged in a layout, and the second annotation information including the position information of the bounding boxes of one or more objects; and training the second sub-network based on the second training data.
[0108] As one implementation, before training the second sub-network based on the second training data, the method further includes: training a second generative model including a second encoder based on the second training data to pre-train the second encoder. Training the second sub-network based on the second training data includes: training the second sub-network based on the second training data, wherein, during training, the pre-trained second encoder is fixed, and the second decoder is trained.
[0109] As one implementation, training the second sub-network further includes: processing the third image using a detection model to generate third location information of one or more objects recorded on the third image; generating region boxes corresponding to the third location information on the solid color image to obtain the second image; and using the third location information of one or more objects as second annotation information of the second image.
[0110] An exemplary embodiment of this disclosure provides a workbook generation apparatus. In some embodiments, the workbook generation apparatus may be implemented as computer program instructions.
[0111] Figure 10 A schematic block diagram of a workbook generation apparatus according to exemplary embodiments of the present disclosure is shown, such as Figure 10 As shown, the workbook generating device 1000 includes:
[0112] Module 1010 is used to retrieve a set of questions, which includes multiple questions.
[0113] The layout generation module 1020 is used to generate multiple target layouts using the layout generation method of this disclosure, wherein each target layout includes a background image and one or more target regions on the background image;
[0114] The layout module 1030 is used to arrange the above-mentioned multiple questions in the target areas of multiple target pages to obtain a workbook, wherein each target area corresponds to one question.
[0115] In some embodiments, the set of questions is determined based on the association between questions and knowledge point categories.
[0116] As one implementation method, the association between questions and knowledge point categories is established as follows: obtain question content; process the question content using a question classification model to output the knowledge point category corresponding to the question content, wherein the knowledge point category belongs to multiple preset knowledge point categories.
[0117] As an example, the question classification model is a Transformer model, which includes: an encoder, configured to encode the string corresponding to the question content to generate a feature vector corresponding to the question content; and a decoder, configured to decode the feature vector to output the knowledge point category corresponding to the question content.
[0118] This disclosure also provides an exemplary embodiment of a system. This system is an example of a system implemented as a computer program on one or more computers at one or more locations.
[0119] Figure 11 A schematic block diagram of a system according to exemplary embodiments of the present disclosure is shown, such as Figure 11 As shown, the system 1100 includes: a question detection device 110, a question recognition device 120, a question classification device 130, a workbook generation device 140; an image database 210, a question image database 220, a question database 230; a detection model 310, a text recognition model 320, a question classification model 330, and a layout generation model 340.
[0120] Artificial intelligence technology has numerous applications in education, such as image-based question search and grading, as well as other atypical applications like automated problem-solving and various student support systems. Through text detection and recognition, offline books or text content are digitized and converted into online computer-processable strings. Therefore, hundreds of millions of digitized, editable questions or images of questions or workbooks have been accumulated; accumulating images is much easier than digitizing and structuring questions.
[0121] Image database 210 can store images such as test papers and workbooks. Each image can include one or more questions, and may also include content other than the questions. Question image database 220 can store question images extracted from workbooks, test papers, etc. Question database 230 can store digitized questions, as well as the association between questions and knowledge point categories.
[0122] A large number of question images were collected (these images can be directly obtained from the accumulated question images, including full-page images and single-question images, and in terms of image quality, they include normal images, as well as blurry, photocopied text images, and other text images). Then, manual annotation was performed. First, an arbitrary portion of the images (about 10%) was detected and annotated with two types of bounding boxes: the first type is the large bounding box of the entire question, and the second type is the printed single-line text box. Then, the content within the single-line text box was transcribed, i.e., the entire sequence was annotated, and then a dictionary was built based on the annotated strings. For each question, a category was assigned, classifying it into a knowledge point. Then, the data with detection annotations was used as dataset one. Then, based on the single-line detection boxes, smaller images were cropped from the large images and mapped to strings, which were used as dataset two. For the category of each question, dataset three was created.
[0123] Detection Model 310
[0124] Based on CenterNet, a detection model 310 is constructed, which can be divided into two parts. The first part uses ResNet18 to extract features. The second part includes three branches, each of which adopts three convolutional layers and two deconvolutional layers to output two-channel feature maps. The feature maps output by the first branch represent the center point score maps of the title box and the text line box, respectively. The feature maps output by the second branch represent the xy offset of the center point, respectively. The feature maps output by the third branch represent the width and height information of the predicted box, respectively. The three branches use Focal Loss, smooth-L1, and smooth-L1 loss functions, respectively, and are trained using dataset one. After training, the detection model 310 is obtained.
[0125] Text recognition model 320
[0126] A text recognition model 320 is constructed, and its specific structure can include three parts. The first part uses a ResNet18 network to extract features from the input text image. The second part uses a two-layer bidirectional LSTM network to further select and focus on the extracted features, while initially modeling context information. The third part includes a self-attention layer to calculate attention scores, obtain context vectors, and a GRU unit for recurrent decoding as the recognizer part. The loss function uses the multi-class cross-entropy loss function. The model is trained using the dataset 2 obtained above. After training, the text recognition model 320 is obtained.
[0127] Question Classification Model 330
[0128] Construct a question classification model 330. It uses the Transformer model, which includes an encoder and a decoder. Unlike the regular Transformer model, its encoder takes the entire question string as input, and the decoder decodes only once. The decoder's input is fixed, i.e., a specific vector, and the decoder's output is the knowledge point category to which the question belongs.
[0129] The detection device 110 processes the collected unlabeled data using the detection model 310 to obtain the question box and single-line text box on each image. Then, a smaller image is cropped based on the single-line text box, and the question recognition device 120 obtains the recognition result through the text recognition model 320. The question classification device 330 processes the information of each question using a question classification model to obtain the knowledge point category to which each question belongs.
[0130] Train the question classification model 330 so that it can process the question content and output the knowledge point category to which the question belongs.
[0131] The detection device 110 uses the detection model 310 to process the images containing the questions recorded in the image database 210 to obtain the question images, which are then stored in the question image database 220. The question recognition device 120 uses the text recognition model 320 to process the question images in the question image database 220 to obtain the question content corresponding to the question images, which is then stored in the question database 230.
[0132] Page Layout Generation Module 340
[0133] Based on dataset one, using detection model 310, the location of the question can be output. Then, other locations may contain handwritten answers. First, based on the location information of the question, change all pixel values of other locations to 0, and then mark the pixels at the text box location as 0 as image labels to obtain dataset four.
[0134] Construct a VAE model where the encoder part uses ResNet18 and the decoder part includes 3 convolutional layers and 4 deconvolutional layers. The output of each stage of the encoder and the 4 deconvolutional layers of the decoder use a U-Net-like structure. Train the model using the L1 loss function and dataset 4. After training, retain the encoder, which is called encoder 1.
[0135] Based on dataset one, using detection model 310, we can determine the bounding box for each question. Then, we draw the bounding boxes of the questions on a white background image and use this as input. The label is the position of each bounding box, resulting in dataset five.
[0136] Construct a VAE model with ResNet18 for the encoder and three convolutional layers for the decoder. Train it using the L1 loss function and dataset five. After training, retain the encoder, which is called encoder two.
[0137] A layout generation model 340 is constructed. Its encoder part comprises two branches: Encoder 1 and Encoder 2. Its decoder also comprises two branches: the first branch includes two deconvolutional layers (for obtaining the background image), and the second branch includes one convolutional layer and one fully connected layer (for obtaining the coordinate information of the target region). Encoder 1 and the first branch form the first sub-network, and Encoder 2 and the second branch form the second sub-network. The first and second sub-networks are trained separately. Both use the L1 loss function, with the encoder fixed, and are trained using the corresponding datasets (datasets four and five) to obtain the first and second sub-networks, thus resulting in the surface generation module 340.
[0138] Generate workbook
[0139] The knowledge point categories are determined (e.g., by teaching and research), and the number of questions designed for each knowledge point category is determined. The workbook generation device 140 then selects questions from the question database 230 based on the knowledge point category and the number of questions, determines the number of workbook pages, selects images from the image database 210 (e.g., randomly selects), uses the layout generation model 340 to generate the layout of each workbook page, then determines the font and size, renders the question strings to the corresponding positions, and obtains the workbook 400.
[0140] Optionally, the workbook can be proofread by curriculum and research staff. Once proofread, it can be published. Compared to the original technology, this can greatly increase the publishing speed while reducing labor costs.
[0141] When manually generating workbooks, the curriculum research department first determines the knowledge points and focus areas covered in the workbook. Then, questions are compiled based on these knowledge points. During this process, the graphic designer needs to design the background and overall layout of the workbook. Once the questions are compiled, colors and fonts are selected according to the designed layout, and the questions are pasted on. Then, the workbooks are manually proofread before finally being published and distributed. This disclosed embodiment of the workbook generation method utilizes existing question resources (a large number of questions and workbook images) to quickly generate a large number of workbooks with different needs and focuses (e.g., elementary school learning workbooks), achieving automatic workbook generation and reducing manual compilation and typesetting costs.
[0142] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.
[0143] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.
[0144] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.
[0145] refer to Figure 12 The present invention describes a structural block diagram of an electronic device 1200 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0146] like Figure 12 As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of the device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0147] Multiple components in electronic device 1200 are connected to I / O interface 1205, including: input unit 1206, output unit 1207, storage unit 1208, and communication unit 1209. Input unit 1206 can be any type of device capable of inputting information to electronic device 1200. Input unit 1206 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 1207 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1208 may include, but is not limited to, disk and optical disk. Communication unit 1209 allows electronic device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0148] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above. For example, in some embodiments, the layout generation method and the workbook generation method can be implemented as computer software programs, which are tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1200 via ROM 1202 and / or communication unit 1209. In some embodiments, the computing unit 1201 can be configured to perform the layout generation method and the workbook generation method by any other suitable means (e.g., by means of firmware).
[0149] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0150] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0151] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0153] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0154] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. A method for generating page layouts, characterized in that, include: Acquire a layout image, wherein the layout image records one or more objects arranged in a specific order; The layout image is processed using a layout generation model to generate a target layout, wherein the target layout includes a background image and position information of one or more target regions on the background image, and each target region corresponds to a target object; The layout generation model includes: A first sub-network, comprising a first encoder and a first decoder, is configured to process the layout image to generate the background image; The second sub-network, which includes a second encoder and a second decoder, is configured to process the layout image to generate position information of the one or more target regions on the background image. The method for training the first sub-network includes: acquiring first training data, wherein the first training data includes a first image and its first annotation information, the first image recording one or more objects arranged in a layout, the first annotation information including an image in which pixels outside the area where the objects are located on the first image and pixels in the area where a first element in the objects is located are set to preset values, the first element being a text box; and training the first sub-network based on the first training data.
2. The method as described in claim 1, wherein, The first decoder includes at least two deconvolutional layers; and / or the second decoder includes at least one convolutional layer and at least one fully connected layer.
3. The method as described in claim 1, wherein, Before training the first sub-network based on the first training data, the method further includes: A first generative model including the first encoder is trained based on the first training data to pre-train the first encoder; The process of training the first sub-network based on the first training data includes: fixing the pre-trained first encoder during the training process to train the first decoder.
4. The method as described in claim 1 or 3, wherein, The method for training the first sub-network also includes: The first image is processed using a detection model to generate first location information of the one or more objects on the first image, and second location information of a first element in the object; The first image is processed based on the first location information and the second location information to obtain the first annotation information.
5. The method of claim 1, wherein, The method for training the second sub-network includes: Acquire second training data, wherein the second training data includes a second image and its second annotation information, the second image records the bounding boxes of one or more objects arranged in a layout, and the second annotation information includes the position information of the bounding boxes of the one or more objects; The second sub-network is trained based on the second training data.
6. The method of claim 5, wherein, Before training the second sub-network based on the second training data, the following steps are also included: A second generative model including the second encoder is trained based on the second training data to pre-train the second encoder; The second sub-network is trained based on the second training data, which includes fixing the pre-trained second encoder and training the second decoder during the training process.
7. The method of claim 5 or 6, wherein, The method for training the second sub-network also includes: The third image is processed using a detection model to generate third location information of one or more objects recorded on the third image; A region bounding box corresponding to the third location information is generated on the solid color image to obtain the second image; The third location information of the one or more objects is used as the second annotation information of the second image.
8. The method according to any one of claims 1 to 3, wherein, The object and the target object include characters.
9. A method for generating workbooks, characterized in that, include: Obtain a set of questions, which includes multiple questions; A plurality of target layouts are generated using the layout generation method as described in any one of claims 1 to 8, wherein each of the plurality of target layouts includes a background image and position information of one or more target regions on the background image; The multiple questions are arranged in the target areas of the multiple target pages to obtain the workbook, wherein each target area corresponds to one question.
10. The method as described in claim 9, wherein the question set is determined based on the association between questions and knowledge point categories, wherein, The methods for establishing the aforementioned association include: Get the question content; The question content is processed using a question classification model to output the knowledge point category corresponding to the question content, wherein the knowledge point category belongs to multiple preset knowledge point categories.
11. The method of claim 10, wherein, The question classification model is the Transformer model, which includes: The encoder is configured to encode the string corresponding to the question content in order to generate a feature vector corresponding to the question content; The decoder is configured to decode the feature vector to output the knowledge point category corresponding to the question content.
12. A page layout generation device, characterized in that, include: The acquisition module is used to acquire a layout image, wherein the layout image records one or more objects arranged in a specific way; A generation module is used to process the layout image using a layout generation model to generate a target layout, wherein the target layout includes a background image and position information of one or more target regions on the background image, and each target region corresponds to a target object; The layout generation model includes: A first sub-network, comprising a first encoder and a first decoder, is configured to process the layout image to generate the background image; The second sub-network, which includes a second encoder and a second decoder, is configured to process the layout image to generate position information of the one or more target regions on the background image. The method for training the first sub-network includes: acquiring first training data, wherein the first training data includes a first image and its first annotation information, the first image recording one or more objects arranged in a layout, the first annotation information including an image in which pixels outside the area where the objects are located on the first image and pixels in the area where a first element in the objects is located are set to preset values, the first element being a text box; and training the first sub-network based on the first training data.
13. A workbook generating device, characterized in that, include: The acquisition module is used to acquire a set of questions, which includes multiple questions; A layout generation module is used to generate a plurality of target layouts using the layout generation method as described in any one of claims 1 to 8, wherein each of the plurality of target layouts includes a background image and position information of one or more target regions on the background image; A layout module is used to arrange the multiple questions in the target areas of the multiple target pages to obtain the workbook, wherein each target area corresponds to one question.
14. An electronic device comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.