Methods, apparatus, devices, and storage media for aggregating cross-page test questions

CN122290156BActive Publication Date: 2026-09-01HANGZHOU PLANET INTELLIGENT COMPUTING TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610729621.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-26
Publication Date
2026-09-01
Estimated Expiration
2046-05-26

AI Technical Summary

Technical Problem

在试题录入过程中,往往需要先对同一试题的不同要素进行聚类,但一些技术无法进行试题要素的自动聚类或聚类准确度低

Benefits of technology

[0009]在本申请一些实施例的技术方案中,在各个试题图像中,将属于同一试题要素的文本块划分为一个文本簇,并在基准文本簇和基准文本簇关联的候选文本簇中,基于文本簇所属的试题要素、文本簇中的文本块和文本块在试题图像中的位置,分别为各个文本簇生成标签编码。由于标签编码体现了文本簇的类别、位置和文本语义等特征,因此,基于标签编码,可以较为准确的评估各个候选文本簇与基准文本簇之间的关联度大小,进而可以基于关联度大小,对属于同一试题的文本簇进行自动聚类,且聚类准确度较高。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290156B_ABST
    Figure CN122290156B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technology and discloses a method, apparatus, device, and storage medium for cross-page test question aggregation. The method includes acquiring multiple test question images, each image including test question elements; extracting text blocks from the test question images and grouping text blocks belonging to the same test question and the same test question element into a text cluster, obtaining a text cluster set; searching for a baseline text cluster and multiple candidate text clusters in the text cluster set, where candidate text clusters are those related to the baseline text cluster; generating label codes for each text cluster in the baseline and candidate text clusters based on the test question element to which the text cluster belongs, the text blocks within the text cluster, and the position of the text blocks in the test question images; and searching for text clusters belonging to the same test question as the baseline text cluster in the candidate text clusters based on the label codes, obtaining a test question text cluster set. The method of this application can achieve automatic clustering of test question elements with high clustering accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for aggregating cross-page test questions. Background Technology

[0002] Test question digitization refers to the process of inputting test questions from teaching materials (such as workbooks and test papers) into a database. After digitization, the test questions in the database can be reorganized according to actual needs to obtain various different test question sets. For example, test questions can be recombined according to students or teaching plans to form test booklets or test papers that are more suitable for the current teaching scenario.

[0003] Each test question can include multiple elements. These elements refer to the components of a test question, including but not limited to the stem, options, explanations, and answer. Currently, in some digitized test scenarios, the elements of the same test question may be distributed in different locations within educational materials. For example, the stem and options might be located on pages 1-30 of a workbook, the answer on pages 31-40, and the explanation on pages 41-70. During the test question entry process, it is often necessary to first cluster the different elements of the same test question; however, some technologies cannot automatically cluster test elements or have low clustering accuracy. Summary of the Invention

[0004] This application provides a method, apparatus, computer device, and readable storage medium for aggregating cross-page test questions, which can achieve automatic clustering of test question elements with high clustering accuracy.

[0005] To achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, embodiments of this application provide a method for aggregating cross-page test questions, the method comprising: Multiple test question images are acquired. Each test question image includes test question elements, which are used to construct test questions. Each test question includes multiple test question elements, and at least some test question elements of the same test question are distributed in different test question images. Text blocks are extracted from the test question image, and text blocks belonging to the same test question and the same test question element are divided into a text cluster to obtain a set of text clusters; In the set of text clusters, a baseline text cluster and multiple candidate text clusters are searched, wherein the candidate text clusters refer to text clusters that are associated with the baseline text cluster; Based on the question elements to which the text cluster belongs, the text blocks in the text cluster, and the position of the text blocks in the question image, a label code is generated for each text cluster in the baseline text cluster and the candidate text cluster; Based on the tag encoding, a text cluster belonging to the same question as the baseline text cluster is searched in the candidate text cluster to obtain a set of question text clusters.

[0006] Secondly, embodiments of this application provide a cross-page test question aggregation device, the device comprising: An image acquisition module is used to acquire multiple test question images. The test question images include test question elements, which are used to construct test questions. Each test question includes multiple test question elements, and at least some test question elements of the same test question are distributed in different test question images. The text cluster segmentation module is used to extract text blocks from the test question image and divide text blocks belonging to the same test question and the same test question element into a text cluster to obtain a text cluster set; The text cluster search module is used to search for a baseline text cluster and multiple candidate text clusters in the text cluster set, wherein the candidate text clusters refer to text clusters that are associated with the baseline text cluster; The encoding generation module is used to generate label codes for the baseline text cluster and each text cluster in the candidate text cluster based on the test question elements to which the text cluster belongs, the text blocks in the text cluster, and the position of the text blocks in the test question image; The text clustering module is used to search for text clusters that belong to the same test question as the baseline text cluster in the candidate text clusters based on the label encoding, so as to obtain a set of test question text clusters.

[0007] Thirdly, embodiments of this application provide a computer device, including: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the cross-page question aggregation method as described above.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer instructions that cause a computer to execute the cross-page question aggregation method as described in any of the preceding claims.

[0009] In some embodiments of this application, text blocks belonging to the same question element are divided into text clusters within each question image. Then, based on the question element to which the text cluster belongs, the text blocks within the text cluster, and the position of the text blocks in the question image, a label encoding is generated for each text cluster within the baseline text cluster and its associated candidate text clusters. Since the label encoding reflects the category, position, and semantic features of the text cluster, the correlation between each candidate text cluster and the baseline text cluster can be evaluated relatively accurately. Therefore, based on the correlation, text clusters belonging to the same question can be automatically clustered with high accuracy. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a method for aggregating cross-page test questions according to some embodiments of this application; Figure 2 A schematic diagram illustrating a first typesetting method provided for some embodiments of this application; Figure 3 A schematic diagram illustrating a second typesetting method provided for some embodiments of this application; Figure 4 A schematic diagram illustrating a third typesetting method provided for some embodiments of this application; Figure 5 A schematic diagram illustrating a fourth typesetting method provided for some embodiments of this application; Figure 6 for Figure 5 A diagram illustrating the text block cache queue in the image; Figure 7 A schematic diagram illustrating the association between different test question images provided for some embodiments of this application; Figure 8 for Figure 7 A diagram illustrating the text block of the next question image in the test, after it has been loaded into the cache queue; Figure 9 A schematic diagram of the test element individualization segmentation process of the test element individualization segmentation system provided for some embodiments of this application; Figure 10 Schematic diagram of a cross-page question aggregation device provided for some embodiments of this application; Figure 11This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0012] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] Currently, in some technologies, the methods for clustering test item elements mainly include the following two: 1) Manual Clustering. In these methods, staff manually review teaching materials page by page, identifying the stem, options, explanations, and answers for the same question. Then, through manual input, copying, pasting, and categorization, different question elements for the same question are clustered, and the clustered elements are saved to a database. Manual clustering is extremely inefficient and prone to human errors such as omissions, mismatches, and duplicate classifications, making it unsuitable for large-scale question digitization needs.

[0014] 2) Conventional Text Matching Clustering. This method uses text retrieval methods based on keyword matching, string similarity, and simple rule matching to perform coarse clustering of test item elements. This clustering method relies solely on character similarity and lacks semantic understanding and logical reasoning capabilities. Without a unified question number association, it cannot identify deep logical relationships between test item elements, resulting in low clustering accuracy.

[0015] 3) Using a general text classification model to categorize test question elements, and then simply piecing them together based on question number or location information of the test question elements. These techniques do not optimize the general text classification model for the context of test question digitization, resulting in low accuracy in classifying test question elements such as question stems, explanations, and answers, and low clustering accuracy.

[0016] Therefore, this application provides a method for aggregating cross-page test questions, which can automatically cluster test question elements with high accuracy. This method can be applied to electronic devices, including but not limited to tablets, desktop computers, laptops, and servers. (See also...) Figure 1 This is a flowchart illustrating a method for aggregating cross-page test questions according to some embodiments of this application. Figure 1 In China, the method for aggregating cross-page test questions includes the following steps: Step S101: Obtain multiple test question images. Each test question image includes test question elements. The test question elements are used to construct test questions. Each test question includes multiple test question elements, and at least some test question elements of the same test question are distributed in different test question images.

[0017] Specifically, in practical applications, if the teaching materials are in paper or other non-electronic formats, multiple test question images can be obtained by taking photos or scanning them. If the teaching materials are in electronic format, multiple test question images can be obtained by taking screenshots, exporting pages, etc. For example, taking photos of each page of a workbook can yield multiple test question images. Similarly, taking photos of the front and back of a test paper can also yield multiple test question images.

[0018] The elements of a test question may include a stem, options, explanation, and answer. Each test question includes a stem, and each test question also includes one or more of the following: options, explanation, and answer. For example, the elements of test question A1 include a stem, options, and answer. The elements of test question A2 include a stem, explanation, and answer. The elements of test question A3 include a stem, options, explanation, and answer. If the elements of a test question are distributed across different test image pages, it means that the elements of the test question are distributed across different pages of the teaching materials, i.e., the test question is a page-spanning question. For example, suppose the stem of test question A1 is located in the first test image, the explanation is located in the 30th test image, and the answer is located in the 38th test image, then test question A1 is a page-spanning question.

[0019] A single test question image may include multiple test question elements of the same category, or multiple test question elements of different categories. For example, suppose test question image P1 includes the stems of questions 1-5, and test question image P2 includes the stem of question 6 and the explanations of questions 1-4. This means that the test question elements in test question image P1 are of the same category, while the test question elements in test question image P2 are of multiple different categories. Each test question element may contain text, tables, question images, etc. It should be noted that in addition to test question elements, test question images may also include non-test question elements. Non-test question elements may include, but are not limited to, answer instructions, answer requirements, and explanations of the question's score.

[0020] Step S102: Extract text blocks from the test question image, and divide text blocks belonging to the same test question and the same test question element into a text cluster to obtain a set of text clusters.

[0021] Specifically, for the text area in the test question image, the text area can be divided into multiple sub-regions, and the text content in each sub-region can be used as a text block. For example, assuming the text content in sub-region A1 is "Please solve the following equation", then the text block extracted from sub-region A1 is "Please solve the following equation".

[0022] For non-text areas in a test question image, these areas can be divided into multiple sub-regions. Text blocks are generated based on the content of each sub-region. The non-text areas include at least one of image areas and table areas, and the text blocks describe the content of each sub-region. Based on these text blocks, the non-text content within the sub-regions can be recovered. For example, assuming sub-region A2 contains two circles of different sizes, the text block extracted from sub-region A2 could include: the radii of the two circles and the coordinates of their centers in the test question image. Thus, based on the text blocks of sub-region A2, the circles within sub-region A2 can be recovered.

[0023] For any sub-region, after extracting a text block from that sub-region, the bounding box coordinates of that sub-region can be used as the position of the text block in the test question image. Based on the positional distribution characteristics of the text blocks in the test question image, text blocks belonging to the same test question and the same test question element are grouped into a text cluster. For example, if the distance between two text blocks is less than a preset value, the two text blocks can be considered to belong to the same test question and the same test question element. A better text clustering scheme is given in the subsequent description of this application, which will not be elaborated here.

[0024] After obtaining the text clusters, they can be input into the trained classification model to output the category of each text cluster. The text cluster category represents the question element to which the text cluster belongs. For example, text cluster C1 belongs to the question stem, and text cluster C2 belongs to the answer key. It's understandable that since elements belonging to the same question and the same question element are grouped into one cluster, each question element can correspond to multiple text clusters. For example, assuming there are 100 questions, and each question has a question stem and an answer, then 100 question stem text clusters and 100 answer text clusters can be obtained, forming a text cluster set.

[0025] Step S103: Search for a baseline text cluster and multiple candidate text clusters in the text cluster set. A candidate text cluster is a text cluster that has a relationship with the baseline text cluster.

[0026] Among them, the test element belonging to the benchmark text cluster is the test element that has a large degree of difference between test items. Based on the benchmark text cluster, different test items can be distinguished.

[0027] Specifically, since the question stems of different test questions vary significantly, the text cluster corresponding to the question stem can be used as the baseline text cluster. For each baseline text cluster, candidate text clusters that are related to it can be found. The candidate text clusters can be different for different baseline text clusters. For example, the candidate text clusters for baseline text cluster C1 are text clusters C11, C12, C13, and C14. The candidate text clusters for baseline text cluster C2 are text clusters C21, C22, C23, C24, and C25.

[0028] Step S104: Based on the test question elements to which the text cluster belongs, the text blocks in the text cluster, and the position of the text blocks in the test question image, generate label codes for each text cluster in the baseline text cluster and the candidate text cluster.

[0029] For example, assuming candidate text clusters C11, C12, C13, and C14 correspond to the baseline text cluster C1, then a label code can be generated for the baseline text cluster C1 based on the question elements to which it belongs, the text blocks within the baseline text cluster C1, and the positions of the text blocks in the question image. Simultaneously, a label code can be generated for the baseline text cluster C11 based on the question elements to which it belongs, the text blocks within the baseline text cluster C11, and the positions of the text blocks in the question image. This process can be repeated to obtain the label codes for the baseline text clusters C1, C12, C13, and C14.

[0030] Step S105: Based on tag encoding, search for text clusters in the candidate text clusters that belong to the same question as the baseline text cluster, and obtain the question text cluster set.

[0031] Specifically, label encoding can reflect the category, location, and semantic features of text clusters. Based on label encoding, the correlation between the baseline text cluster and its corresponding candidate text clusters can be determined, and then text clusters belonging to the same test question as the baseline text cluster can be found among the candidate text clusters. For example, assuming that among candidate text clusters C11, C12, C13, and C14, candidate text clusters C11, C12, and C13 have a high correlation with the baseline text cluster C1, while candidate text cluster C14 has a low correlation with the baseline text cluster C1, then it can be considered that the baseline text cluster C1 and candidate text clusters C11, C12, and C13 belong to the same test question.

[0032] In this embodiment, the baseline text cluster, candidate text clusters, and the label encodings of each text cluster can be input into the trained aggregation model. The aggregation model then outputs the labels of the text clusters that belong to the same test question as the baseline text cluster. In this way, text clusters are clustered on a test question basis, resulting in multiple text clusters corresponding to each test question.

[0033] Specifically, in this embodiment, a baseline model can be trained using sample data from the test question domain, and the parameters of the baseline model can be fine-tuned to obtain an aggregated model. The baseline model is a pre-trained large language model with conventional semantic processing capabilities. Because the aggregated model is trained based on sample data from the test question domain, the clustering results have high accuracy.

[0034] Based on the above description, in some embodiments of the technical solutions of this application, in each test question image, text blocks belonging to the same test question element are divided into a text cluster. Within the baseline text cluster and its associated candidate text clusters, label codes are generated for each text cluster based on the test question element to which the text cluster belongs, the text blocks within the text cluster, and the position of the text blocks in the test question image. Since the label codes reflect the category, position, and semantic features of the text clusters, the correlation between each candidate text cluster and the baseline text cluster can be evaluated relatively accurately. Therefore, based on the correlation, text clusters belonging to the same test question can be automatically clustered with high accuracy.

[0035] In some embodiments, a dedicated registration model with text reasoning capabilities designed for digital scenario matching of test questions can be used to perform a preliminary coarse screening of text clusters in the text cluster set, eliminating text clusters that are not related to other text clusters. Specifically, for any text cluster in the text cluster set, if the text cluster is not related to other text clusters in the text cluster set, it can be removed from the text cluster set. In this way, on the one hand, the amount of data in the text cluster set can be reduced, and on the other hand, when searching for baseline text clusters and candidate text clusters with relationships in the text cluster set, data interference can be reduced, and search accuracy can be improved.

[0036] In some embodiments, step S103, which involves searching for a baseline text cluster and multiple candidate text clusters in the text cluster set, includes: In the set of text clusters, find one of the text clusters that corresponds to the question stem, and use the found text cluster as the base text cluster; The baseline text cluster is compared with each other text cluster in the text cluster set, and candidate text clusters are obtained based on the similarity comparison results.

[0037] Specifically, since the question stems of different test questions vary greatly, using the text cluster corresponding to the question stem as the baseline text cluster can make the reference cluster (i.e., the baseline text cluster) of the clustering process globally unique, thereby improving the clustering accuracy.

[0038] In practical applications, a preliminary screening of text clusters can be performed first based on a dedicated registration model. Then, candidate text clusters matching the baseline text cluster are searched within the text cluster set using methods such as similarity comparison. Finally, based on the label encodings of the baseline and candidate text clusters, a high-precision text clustering is performed using an aggregation model. Thus, through hierarchical registration and a global retrieval mechanism, mismatch and omission rates can be reduced, improving the accuracy of the aggregation results. Furthermore, this method of preliminary screening followed by high-precision clustering reduces redundant computation and improves clustering efficiency.

[0039] In some embodiments, the above-mentioned comparison of similarity between the baseline text cluster and each other text cluster in the text cluster set, and the obtaining of candidate text clusters based on the similarity comparison results, include: For any target text cluster other than the baseline text cluster in the text cluster set, the similarity between the target text cluster and the baseline text cluster is used as the first similarity. If the first similarity score is greater than the first similarity threshold, then the target text cluster is taken as a candidate text cluster. If the first similarity is less than the first similarity threshold but greater than the second similarity threshold, then the target text cluster and the reference text cluster are compared to obtain the second similarity. Here, the reference text cluster refers to any text cluster other than the baseline text cluster corresponding to the question stem. The first similarity threshold is greater than the second similarity threshold. If the first similarity is less than the second similarity, then the target text cluster is used as a candidate text cluster for the reference text cluster.

[0040] For example, suppose the target text cluster is one of the answer text clusters. If the answer text cluster has a high similarity to the benchmark text cluster, it indicates a strong correlation between the two, and this answer text cluster can be considered a candidate text cluster for the benchmark. If the answer text cluster has a relatively low similarity to the benchmark, other question stem text clusters besides the benchmark can be used as reference text clusters, and the similarity between the answer text cluster and the reference text clusters can be compared. In this way, the answer text cluster can be used as a benchmark to search in reverse for other question stem text clusters with higher similarity to it. If they exist, the answer text cluster is considered a candidate text cluster for the reference text clusters; if they do not exist, the answer text cluster is considered a candidate text cluster for the benchmark.

[0041] In the above embodiments, reverse lookup can improve the accuracy of candidate text clusters and avoid the problem of reduced clustering accuracy caused by inaccurate candidate text clusters.

[0042] In some embodiments, step S102, which involves extracting text blocks from the test question image and grouping text blocks belonging to the same test question and the same test question element into a text cluster, includes the following steps: 1) Obtain the target test question image and extract the text block and the coordinates of the detection box of the text block from the target test question image.

[0043] Specifically, in this embodiment, for any sub-region in the target test question image, after extracting a text block from that sub-region, the bounding box coordinates of that sub-region can be used as the detection box coordinates of the text block. Here, the bounding box coordinates refer to the coordinates of the bounding box of the sub-region in the target test question image. For example, after extracting text block D1 from sub-region A1, the bounding box coordinates of sub-region A1 can be used as the detection box coordinates of text block D1. Thus, based on the detection box coordinates, the position of each text block in the target test question image can be located.

[0044] 2) Sort the text blocks according to the coordinates of the detection boxes to obtain the first sorting queue.

[0045] In this embodiment, the layout of the target test question image can be determined based on the distribution characteristics of the detection box coordinates. Based on the layout, the text blocks can be sorted to obtain a first sorting queue. The layout can include, but is not limited to, single-column layout, two-column layout, multi-column layout, mixed text and image layout, horizontal layout, and vertical layout.

[0046] For ease of understanding, please refer to the following: Figure 2 This is a schematic diagram illustrating a first typesetting method provided for some embodiments of this application. Figure 2 The layout is a two-column format. The detection box coordinates in a two-column layout have the following distribution characteristics: the detection box coordinates are distributed in the left and right regions of the question image, and the middle region of the question image may not have detection box coordinates. Meanwhile, in each region, the detection box coordinates on the left side (i.e., the region where the question number is located in each column) are relatively aligned, and the spacing between questions can be greater than the text line spacing within the same question.

[0047] If the coordinates of the detection box in the target test question image exist in a similar way... Figure 2 The distribution characteristics shown indicate that the target test question image is a two-column layout. In a two-column layout, the test questions are typically arranged from the left column to the right column, and within each column, the questions are arranged from top to bottom. Based on the test question layout characteristics of a two-column layout, Figure 2 The initial sorting of the text blocks can be: text block D1>text block D3>text block D5>text block D7>text block D2>text block D4>text block D6>text block D8.

[0048] See also Figure 3 This is a schematic diagram illustrating a second typesetting method provided for some embodiments of this application. Figure 3In the image, the layout is a single-column layout. The detection box coordinates of the single-column layout have the following distribution characteristics: the left side of the image can have more detection box coordinates, the right side can have fewer detection box coordinates, and the detection box coordinates in the left side can be relatively aligned.

[0049] If the coordinates of the detection box in the target test question image exist in a similar way... Figure 3 The distribution characteristics shown indicate that the target test question image is a single-column layout. In a single-column layout, the test questions are arranged sequentially from top to bottom. Based on the characteristics of single-column test question layout, Figure 3 The initial sorting of the text blocks can be: text block D1>text block D2>text block D3>text block D4>text block D5>text block D6>text block D7>text block D8.

[0050] 3) Based on the semantic relationship between text blocks, the order of text blocks in the first sorting queue is adjusted to obtain the second sorting queue.

[0051] Specifically, when the layout of the target test question image is complex, relying solely on the coordinates of the detection boxes may lead to sorting errors. For example, referring to... Figure 4 This is a schematic diagram illustrating a third typesetting method provided in some embodiments of this application. From Figure 4 It can be seen that the top area of ​​the target test question image is actually a single-column layout, but the area below the top area is actually a two-column layout. Under this layout, based on the detection box coordinate distribution characteristics, the layout may be determined to be a two-column layout, resulting in the initial text block order as: text block D1>text block D3>text block D5>text block D7>text block D9>text block D11>text block D2>text block D4>text block D6>text block D8>text block D10>text block D12. In this initial order, text blocks D1, D3, and D5 are arranged together, but their semantics are not coherent. Similarly, text blocks D2, D4, and D6 are arranged together, but their semantics are also not coherent. At the same time, text blocks D1 to D4 are semantically coherent sentences, but in this order, text blocks D1 to D4 are not arranged in adjacent positions. Therefore, based on the semantic relationships between text blocks, the order of text blocks in the first sorting queue can be fine-tuned to ensure semantic coherence among text blocks in the resulting second sorting queue, thereby preventing text block sorting errors. For example, it can be... Figure 4 The text blocks are sorted as follows: text block D1>text block D2>text block D3>text block D4>text block D5>text block D7>text block D9>text block D11>text block D6>text block D8>text block D10>text block D12.

[0052] 4) Load the text blocks in the second sorting queue into the cache queue of the grouping model. The grouping model is used to divide the text blocks in the cache queue that belong to the same test element into a group and output the first grouping result.

[0053] When loading text blocks into the cache queue, they need to be sorted according to the text blocks in the second sorting queue, and then loaded into the cache queue one by one. In other words, the sorting of text blocks in the cache queue needs to be the same as the sorting of text blocks in the second sorting queue. This maintains the semantic coherence between text blocks. It can be understood that after sorting, text blocks of the same question should be in adjacent positions in the cache queue.

[0054] The grouping model is a pre-trained large language model that can filter out target text blocks belonging to test item elements from the cache queue, and then group text blocks belonging to the same test item element into a group based on the semantic correlation between adjacent text blocks within the target text blocks.

[0055] For ease of understanding, please refer to the following: Figure 5 and Figure 6 . Figure 5 A schematic diagram illustrating a fourth typesetting method provided for some embodiments of this application. Figure 6 for Figure 5 A diagram illustrating the text block cache queue. The grouping model is based on... Figure 6 The cache queue shown can filter text blocks D5 to D12 as target text blocks. Among the target text blocks D5 to D12, since text blocks D5 and D6 have a relatively high semantic correlation, while text blocks D6 and D7 have a relatively low semantic correlation, text blocks D5 and D6 can be grouped together, indicating that text blocks D5 and D6 belong to the same question element. Following a similar principle, the grouping model can output a total of four groups: (D5, D6), (D7, D8), (D9, D10), and (D11, D12). Each group of text blocks can constitute a question element. Thus, the individualization of question elements is achieved. Each group of text blocks can be considered a text cluster.

[0056] 5) Based on the semantic relationship between adjacent groups, the result of the first group is verified, and if the verification passes, the text block in the last group is retained in the cache queue, and other text blocks outside the last group are deleted.

[0057] In this embodiment, text blocks from each group can be input into the trained verification model. The verification model performs semantic analysis and outputs a verification result indicating whether the verification passed or failed. Specifically, if any two adjacent groups have semantically discontinuous text blocks, the verification fails. Otherwise, the verification passes.

[0058] If the validation of the first grouping result fails, the first grouping result can be deleted, while all text blocks in the cache queue are retained. Then, the text blocks in the cache queue are regrouped. In this way, the first grouping result can be corrected, further improving the grouping accuracy.

[0059] If the validation of the first group result passes, the results of all other groups except the last group can be saved. The text blocks from the last group are retained in the cache queue, and all other text blocks except the last group are deleted. The last group refers to the last group at the end of the cache queue. For example, Figure 6 D11 and D12 can be considered as end groups. The text blocks in the end groups may have semantic relevance to the text blocks in the next question image. The next question image is the question image located after the target question image. For example, the target question image could be the question image on page 1 of the workbook, and the next question image could be the question image on page 2 of the workbook.

[0060] For ease of understanding, please refer to the following: Figure 7 This is a schematic diagram illustrating the relationship between different test question images provided in some embodiments of this application. Figure 7 As can be seen, after loading the text blocks from the target question image into the cache queue, the grouping model can output a total of 5 groups: (D5, D6), (D7, D8), (D9, D10), (D11, D12), and (D13). The final group (D13) contains only one text block, D13, and this text block D13, together with the text block D14 from the next question image, forms a single question; that is, text blocks D13 and D14 are semantically related. Logically, text blocks D13 and D14 should be grouped into the same group. However, text block D14 is located in the next question image, so grouping (D13, D14) cannot be obtained solely based on the text blocks from the target question image. Therefore, text block D13 from the final group can be retained in the cache queue so that after the text blocks from the next question image are loaded into the cache queue, text block D13 can be grouped with the text blocks from the next question image, thus obtaining the complete grouping (D13, D14). Meanwhile, the text blocks in other groups (D5,D6), (D7,D8), (D9,D10), and (D11,D12) are no longer missing, so these text blocks can be deleted from the cache queue. Additionally, non-question text blocks (such as D1–D4) can also be deleted from the cache queue. Thus, after grouping the target question image, the cache queue only contains the text blocks of the last group.

[0061] 6) Obtain the next question image after the target question image, and load the sorted text blocks extracted from the next question image into the cache queue, so that the grouping model can output the second grouping result based on the text blocks in the next question image and the retained text blocks in the cache queue.

[0062] The text blocks of the next test question image can be extracted and sorted in a similar way to the target test question image, and the sorted text blocks can be loaded into the cache queue.

[0063] See also Figure 8 ,for Figure 7 This is a diagram illustrating the text block of the next question image after it has been loaded into the cache queue. From Figure 8 As can be seen, since the target question image's group (D13) remains in the cache queue, text block D13 in the target question image and text blocks D14~D20 in the next question image form a new text block cache queue. Based on this new text block cache queue, the grouping model can group text blocks D13 and D14 together. This can solve the problem of question element breakage caused by question elements spanning multiple pages in some technologies.

[0064] 7) Based on the grouping results, one or more text clusters are obtained.

[0065] Specifically, for any given test question image, after removing the later-arriving groups from the text block groups of the test question image, the remaining groups can be used as the individual segmentation results of the test question elements extracted from the test question image. For example, using... Figure 7 For example, in the target test question image, group (D5,D6), (D7,D8), (D9,D10), (D11,D12), and (D13), group (D13) can be removed, and group (D5,D6), (D7,D8), (D9,D10), and (D11,D12) can be used as the individual segmentation results of the test question elements extracted from the target test question image.

[0066] A test element can be constructed based on the text blocks in each group. When a test element has an image or table, the group of the test element can include text blocks to describe the image or table. Based on these text blocks, the image or table can be recovered, and thus the test element can be constructed.

[0067] In the above technical solution, firstly, after sorting the text blocks according to the coordinates of the detection boxes, the order of the text blocks is adjusted based on the semantic correlation between them. This improves the accuracy of text block sorting, prevents semantic jumps between text blocks, and ensures that the sorting of text blocks conforms to human reading habits. Secondly, text blocks in the last group are retained in the cache queue, and grouping is performed based on the text blocks of the next question image and the retained text blocks in the cache queue. This prevents the problem of question element breakage due to question elements crossing pages, eliminating the need for manual correction. Finally, by grouping text blocks through a grouping model, the grouping model can learn the features of text blocks belonging to the same question element, dynamically adjust the text block grouping logic and segmentation window, etc., without the need for manual parameter setting. In this way, while improving the efficiency of question element individual segmentation, it can also overcome the problem of inaccurate segmentation caused by rigid question element individual segmentation logic and fixed segmentation windows in some other technologies. Based on the above aspects, the method of this application can improve the accuracy and efficiency of question element individual segmentation.

[0068] In some embodiments, loading text blocks from the second sorting queue into the cache queue of the grouping model includes: The total number of tokens is determined based on the total length of the text blocks in the second sorting queue. If the total number of tokens is greater than the maximum number of tokens allowed by the grouping model, then based on the maximum number of tokens, the first number of text blocks are extracted from the second sorting queue and loaded into the cache queue. Remove text blocks that have been loaded into the cache queue from the second sorting queue, and determine the number of remaining tokens based on the remaining text blocks in the second sorting queue; If the remaining number of tokens is greater than the maximum number of tokens allowed by the grouping model, then continue to extract the second number of text blocks from the second sorting queue based on the maximum number of tokens, and load the second number of text blocks into the cache queue; If the number of remaining tokens is not greater than the maximum number of tokens allowed by the grouping model, then all remaining text blocks in the second sorting queue will be loaded into the cache queue.

[0069] The total text length refers to the number of texts included in the second sorting queue. The total text length can be determined based on the total number of text blocks in the second sorting queue and the number of texts in each text block. There can be a conversion relationship between the total text length and the number of tokens. For example, 3 texts equal 1 token. Based on this conversion relationship, the total number of tokens in the second sorting queue can be obtained.

[0070] If the total number of tokens exceeds the maximum number of tokens allowed by the grouping model, the text blocks in the second sorting queue can be divided into multiple batches, and the text blocks can be loaded into the cache queue in batches. Conversely, if the total number of tokens does not exceed the maximum number of tokens, all text blocks in the second sorting queue can be loaded into the cache queue at once.

[0071] Among them, loading text blocks into the cache queue in batches means loading the text blocks of the previous batch into the cache queue, and after the grouping model completes the grouping of the text blocks of the previous batch, loading the text blocks of the next batch into the cache queue.

[0072] Specifically, when loading text blocks from the second sorting queue into the cache queue in batches, the number of text blocks that the cache queue can hold can be determined based on the maximum number of tokens allowed by the grouping model and the number of text blocks included in each text block in the second sorting queue. Then, according to the determined number of text blocks, the corresponding number of text blocks are extracted from the second sorting queue into the cache queue. After each batch of text blocks is extracted, the remaining text length in the second sorting queue can be calculated based on the remaining text blocks in the second sorting queue, and the remaining text length can be converted into the remaining number of tokens. If the remaining number of tokens is greater than the maximum number of tokens allowed by the grouping model, the remaining text blocks are continued to be loaded into the cache queue in batches. If the remaining number of tokens is not greater than the maximum number of tokens allowed by the grouping model, all remaining text blocks in the second sorting queue can be loaded into the cache queue at once.

[0073] It should be noted that since different text blocks may contain different amounts of text, the number of text blocks in different batches may also be different. That is, the first and second quantities mentioned above may be different.

[0074] In the above embodiments, on the one hand, when the number of tokens corresponding to the second sorting queue is greater than the maximum number of tokens allowed by the grouping model, controlling the number of text blocks that can be cached in the cache queue based on the maximum number of tokens can ensure the inference accuracy of the grouping model, thereby ensuring the grouping accuracy of text blocks. On the other hand, when the number of tokens corresponding to the second sorting queue is not greater than the maximum number of tokens allowed by the grouping model, loading all text blocks in the second sorting queue into the cache queue can improve the grouping efficiency of text blocks. That is, it is possible to ensure both the grouping accuracy and grouping efficiency of text blocks.

[0075] In some embodiments, when loading text blocks from the second sorting queue into the cache queue in batches, the above-mentioned retention of text blocks in the last group and deletion of other text blocks outside the last group in the cache queue includes: After the grouping model completes the grouping of the first number of text blocks, it retains the text blocks in the last group in the cache queue and deletes other text blocks outside the last group. After the remaining text blocks in the second sorting queue are loaded into the cache queue, the grouping model continues to group based on the remaining text in the second sorting queue and the retained text blocks in the cache queue.

[0076] Specifically, after the text blocks in the second sorting queue are divided into multiple batches, retaining the last grouped text blocks of each batch in the cache queue can prevent breakage of question elements within the same question image. For example, in conjunction with reference... Figure 7 Suppose that in the target question image, text blocks D1–D7 are divided into the first batch, and text blocks D8–D13 are divided into the second batch. If the last group of text blocks (D7) from the first batch is not retained in the cache queue, the stem of question 2 will be divided into two groups, resulting in a break in the question elements. Conversely, if the last group of text blocks (D7) from the first batch is retained in the cache queue, then after loading the text blocks from the second batch into the cache queue, the stem of question 2 can be divided into the same group, thus improving the accuracy of individual segmentation of question elements.

[0077] It should be noted that when retaining the last group of text blocks from the previous batch in the cache queue, the number of text blocks in the next batch can be determined as follows: Determine the number of tokens corresponding to the last group of text blocks in the previous batch; The target number of tokens is obtained by subtracting the number of tokens corresponding to the last group of text blocks in the previous batch from the maximum number of tokens allowed by the grouping model. The number of text blocks in the next batch is determined based on the target token count.

[0078] This is to prevent the number of tokens in the cache queue from exceeding the maximum number of tokens allowed by the grouping model.

[0079] In some embodiments, before the grouping model performs text block grouping, the method of this application further includes: Obtain the attribute features of test question elements; Based on the attribute features of the test question elements, prompt words are constructed and input into the grouping model so that the grouping model can group the text blocks in the cache queue based on the attribute features of the test question elements.

[0080] Specifically, the attribute features of test element can include the starting features and structure of the test element. For example, assuming that each test element starts with an Arabic numeral, then the prompt phrase "test element starts with an Arabic numeral" can be constructed. After inputting this prompt phrase into the grouping model, the grouping model can use this prompt phrase as a reference, thereby improving the accuracy of text block grouping.

[0081] In some embodiments, each text block has a text block identifier; the method further includes: For any text block, concatenate the text block, the coordinates of the detection box of the text block, and the text block identifier to obtain the first concatenated content of the text block; The first concatenated content of multiple text blocks is input into the sorting model, so that the sorting model outputs the first sorting result of the text block identifiers based on the detection box coordinates, which serves as the first sorting queue; For any text block, concatenate the text block and its identifier to obtain the second concatenated content of the text block; According to the text block order of the first sorting queue, the second concatenation content of each text block is input into the adjustment model so that the adjustment model outputs the second sorting result of the text block identifier, which is used as the second sorting queue; Following the text block order of the second sorting queue, the second concatenation content of each text block is input into the grouping model so that the grouping model outputs the grouping result of the text block identifier as the grouping result of the text block.

[0082] Specifically, the ranking and adjustment models can be the trained large language model. After inputting the text block identifiers and their concatenated contents into the model, the model can be prompted to rank them based on the text block identifiers. For example, assuming the text block identifiers of text blocks D1 to D7 are 1, 2, 3, 4, 5, 6, and 7 respectively, the ranking result of text blocks D1 to D7 can be represented as 1>3>2>7>4>5>6. In this way, the model can be prevented from modifying the text blocks during the ranking process.

[0083] See also Figure 9 This is a schematic diagram of the test element individualization process of a test element individualization segmentation system provided in some embodiments of this application. Figure 9The test item element individual segmentation system includes an image preprocessing module, a ranking model, an adjustment model, a buffer queue, and a grouping model. The image preprocessing module extracts text blocks and their bounding box coordinates from the test item images. The ranking model sorts the text blocks based on their bounding box coordinates, outputting a first ranked queue to the adjustment model. The adjustment model adjusts the order of text blocks in the first ranked queue based on semantic relationships between them, obtaining a second ranked queue, and then loads the text blocks from the second ranked queue into the buffer queue. The grouping model groups the text blocks in the buffer queue and outputs the grouping results for each test item image.

[0084] In practical applications, multiple test question images P1 to Pn can be obtained from teaching materials, and each test question image can be input into the individual question element segmentation system in the order of their sequence. The individual question element segmentation system can automatically complete the individual segmentation of each test question image without manual intervention, and can avoid problems such as broken test elements and segmentation errors.

[0085] Corresponding to the method, this application also provides a device for aggregating cross-page test questions. (See also...) Figure 10 This is a schematic diagram of a multi-page question aggregation device provided in some embodiments of this application. The multi-page question aggregation device includes: The image acquisition module 1001 is used to acquire multiple test question images. The test question images include test question elements, which are used to construct test questions. Each test question includes multiple test question elements, and at least some test question elements of the same test question are distributed in different test question images. The text cluster segmentation module 1002 is used to extract text blocks from the test question image and divide the text blocks belonging to the same test question and the same test question element into a text cluster to obtain a text cluster set; The text cluster search module 1003 is used to search for a baseline text cluster and multiple candidate text clusters in a text cluster set. A candidate text cluster is a text cluster that has a relationship with the baseline text cluster. The encoding generation module 1004 is used to generate label codes for each text cluster in the baseline text cluster and candidate text cluster based on the test question elements to which the text cluster belongs, the text blocks in the text cluster, and the position of the text blocks in the test question image; The text clustering module 1005 is used to find text clusters that belong to the same test question as the baseline text cluster in the candidate text clusters based on label encoding, so as to obtain a set of test question text clusters.

[0086] In some embodiments, the elements of a test question include a stem, options, explanations, and an answer. Each test question includes a stem, and each test question also includes one or more of the options, explanations, and answers. The text cluster search module 1003 is specifically used for: In the set of text clusters, find one of the text clusters that corresponds to the question stem, and use the found text cluster as the base text cluster; The baseline text cluster is compared with each other text cluster in the text cluster set, and candidate text clusters are obtained based on the similarity comparison results.

[0087] In some embodiments, the text cluster lookup module 1003 is specifically used for: For any target text cluster other than the baseline text cluster in the text cluster set, the similarity between the target text cluster and the baseline text cluster is used as the first similarity. If the first similarity score is greater than the first similarity threshold, then the target text cluster is taken as a candidate text cluster. If the first similarity is less than the first similarity threshold but greater than the second similarity threshold, then the target text cluster and the reference text cluster are compared to obtain the second similarity. Here, the reference text cluster refers to any text cluster other than the baseline text cluster corresponding to the question stem. The first similarity threshold is greater than the second similarity threshold. If the first similarity is less than the second similarity, then the target text cluster is used as a candidate text cluster for the reference text cluster.

[0088] In some embodiments, the text cluster segmentation module 1002 is specifically used for: Acquire the target test question image, and extract the text block and the coordinates of the detection box of the text block from the target test question image; Based on the coordinates of the detection boxes, the text blocks are sorted to obtain the first sorting queue; Based on the semantic relationships between text blocks, the order of text blocks in the first sorting queue is adjusted to obtain the second sorting queue; The text blocks in the second sorting queue are loaded into the cache queue of the grouping model. The grouping model is used to divide the text blocks in the cache queue that belong to the same test item element into a group and output the first grouping result. Based on the semantic relationship between adjacent groups, the result of the first group is verified, and if the verification passes, the text block in the last group is retained in the cache queue, and other text blocks outside the last group are deleted. The next question image after the target question image is obtained, and the sorted text blocks extracted from the next question image are loaded into the cache queue so that the grouping model can output the second grouping result based on the text blocks in the next question image and the retained text blocks in the cache queue. Based on the grouping results, one or more text clusters are obtained.

[0089] In some embodiments, the text cluster segmentation module 1002 is specifically used for: The total number of tokens is determined based on the total length of the text blocks in the second sorting queue. If the total number of tokens is greater than the maximum number of tokens allowed by the grouping model, then based on the maximum number of tokens, the first number of text blocks are extracted from the second sorting queue and loaded into the cache queue. Remove text blocks that have been loaded into the cache queue from the second sorting queue, and determine the number of remaining tokens based on the remaining text blocks in the second sorting queue; If the remaining number of tokens is greater than the maximum number of tokens allowed by the grouping model, then continue to extract the second number of text blocks from the second sorting queue based on the maximum number of tokens, and load the second number of text blocks into the cache queue; If the number of remaining tokens is not greater than the maximum number of tokens allowed by the grouping model, then all remaining text blocks in the second sorting queue will be loaded into the cache queue.

[0090] In some embodiments, the text cluster segmentation module 1002 is specifically used for: After the grouping model completes the grouping of the first number of text blocks, it retains the text blocks in the last group in the cache queue and deletes other text blocks outside the last group. After the remaining text blocks in the second sorting queue are loaded into the cache queue, the grouping model continues to group based on the remaining text in the second sorting queue and the retained text blocks in the cache queue.

[0091] In some embodiments, each text block has a text block identifier; the text cluster segmentation module 1002 is further configured to: For any text block, concatenate the text block, the coordinates of the detection box of the text block, and the text block identifier to obtain the first concatenated content of the text block; The first concatenated content of multiple text blocks is input into the sorting model, so that the sorting model outputs the first sorting result of the text block identifiers based on the detection box coordinates, which serves as the first sorting queue; For any text block, concatenate the text block and its identifier to obtain the second concatenated content of the text block; According to the text block order of the first sorting queue, the second concatenation content of each text block is input into the adjustment model so that the adjustment model outputs the second sorting result of the text block identifier, which is used as the second sorting queue; Following the text block order of the second sorting queue, the second concatenation content of each text block is input into the grouping model so that the grouping model outputs the grouping result of the text block identifier as the grouping result of the text block.

[0092] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0093] In this embodiment, the cross-page question aggregation device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0094] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application, such as... Figure 11 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 11 Take a processor 10 as an example.

[0095] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0096] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0097] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0098] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0099] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any embodiment of this application.

[0100] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for aggregating cross-page test questions, characterized in that, The method includes: Multiple test question images are acquired. Each test question image includes test question elements, which are used to construct test questions. Each test question includes multiple test question elements, and at least some test question elements of the same test question are distributed in different test question images. Text blocks are extracted from the test question image, and text blocks belonging to the same test question and the same test question element are divided into a text cluster to obtain a set of text clusters; In the set of text clusters, a baseline text cluster and multiple candidate text clusters are searched, wherein the candidate text clusters refer to text clusters that are associated with the baseline text cluster; Based on the question elements to which the text cluster belongs, the text blocks in the text cluster, and the position of the text blocks in the question image, a label code is generated for each text cluster in the baseline text cluster and the candidate text cluster; Based on the tag encoding, search for text clusters in the candidate text clusters that belong to the same question as the baseline text cluster to obtain a set of question text clusters; The step of extracting text blocks from the test question image and dividing text blocks belonging to the same test question and the same test question element into a text cluster includes: Acquire a target test question image, and extract text blocks and the coordinates of the detection boxes of the text blocks from the target test question image; Based on the coordinates of the detection box, the text blocks are sorted to obtain a first sorting queue; Based on the semantic relationships between text blocks, the order of text blocks in the first sorting queue is adjusted to obtain the second sorting queue; The text blocks in the second sorting queue are loaded into the cache queue of the grouping model. The grouping model is used to divide the text blocks in the cache queue that belong to the same test element into a group and output the first grouping result. Based on the semantic relationship between adjacent groups, the first group result is verified, and if the verification passes, the text block in the last group is retained in the cache queue, and other text blocks outside the last group are deleted; The next test question image is obtained after the target test question image, and the sorted text blocks extracted from the next test question image are loaded into the cache queue, so that the grouping model outputs a second grouping result based on the text blocks in the next test question image and the retained text blocks in the cache queue; Based on the grouping results, one or more text clusters are obtained.

2. The method according to claim 1, characterized in that, The test question elements include a stem, options, explanation, and answer. Each test question includes a stem and also includes one or more of the options, explanation, and answer. The step of searching for a baseline text cluster and multiple candidate text clusters in the text cluster set includes: In the set of text clusters, find one of the text clusters corresponding to the question stem, and use the found text cluster as the base text cluster; The baseline text cluster is compared with each of the other text clusters in the text cluster set, and the candidate text cluster is obtained based on the similarity comparison results.

3. The method according to claim 2, characterized in that, The step of comparing the baseline text cluster with each other text cluster in the text cluster set, and obtaining the candidate text cluster based on the similarity comparison results, includes: For any target text cluster in the text cluster set other than the baseline text cluster, the similarity between the target text cluster and the baseline text cluster is taken as the first similarity. If the first similarity is greater than the first similarity threshold, then the target text cluster is taken as the candidate text cluster; If the first similarity is less than the first similarity threshold and greater than the second similarity threshold, then the target text cluster and the reference text cluster are compared to obtain the second similarity. The reference text cluster refers to any text cluster other than the baseline text cluster corresponding to the question stem. The first similarity threshold is greater than the second similarity threshold. If the first similarity is less than the second similarity, then the target text cluster is used as a candidate text cluster for the reference text cluster.

4. The method according to claim 1, characterized in that, The step of loading text blocks from the second sorting queue into the cache queue of the grouping model includes: The total number of tokens is determined based on the total length of the text blocks in the second sorting queue. If the total number of tokens is greater than the maximum number of tokens allowed by the grouping model, then based on the maximum number of tokens, a first number of text blocks are extracted from the second sorting queue, and the first number of text blocks are loaded into the cache queue; Remove the text blocks that have been loaded into the cache queue from the second sorting queue, and determine the number of remaining tokens based on the remaining text blocks in the second sorting queue; If the remaining number of tokens is greater than the maximum number of tokens allowed by the grouping model, then based on the maximum number of tokens, a second number of text blocks are extracted from the second sorting queue, and the second number of text blocks are loaded into the cache queue; If the number of remaining tokens is not greater than the maximum number of tokens allowed by the grouping model, then all remaining text blocks in the second sorting queue are loaded into the cache queue.

5. The method according to claim 4, characterized in that, The step of retaining text blocks in the end group of the cache queue and deleting other text blocks outside the end group includes: After the grouping model completes the grouping of the first number of text blocks, it retains the text blocks in the last group in the cache queue and deletes other text blocks outside the last group. After the remaining text blocks in the second sorting queue are loaded into the cache queue, the grouping model continues to group based on the remaining text in the second sorting queue and the retained text blocks in the cache queue.

6. The method according to claim 5, characterized in that, Each of the text blocks has a text block identifier; the method further includes: For any of the text blocks, the text block, the coordinates of the detection box of the text block, and the text block identifier are concatenated to obtain the first concatenated content of the text block; The first concatenated content of multiple text blocks is input into the sorting model, so that the sorting model outputs the first sorting result of the text block identifier based on the detection box coordinates, which serves as the first sorting queue; For any of the text blocks, the text block and its text block identifier are concatenated to obtain the second concatenated content of the text block; According to the text block order of the first sorting queue, the second concatenation content of each text block is input into the adjustment model so that the adjustment model outputs the second sorting result of the text block identifier as the second sorting queue; According to the text block order of the second sorting queue, the second concatenation content of each text block is input into the grouping model, so that the grouping model outputs the grouping result of the text block identifier as the grouping result of the text block.

7. A device for aggregating cross-page test questions, characterized in that, The device includes: An image acquisition module is used to acquire multiple test question images. The test question images include test question elements, which are used to construct test questions. Each test question includes multiple test question elements, and at least some test question elements of the same test question are distributed in different test question images. The text clustering module is used to extract text blocks from the test question image and divide text blocks belonging to the same test question and the same test question element into a text cluster, obtaining a text cluster set. Specifically, it acquires the target test question image and extracts text blocks and their detection box coordinates from the target test question image; sorts the text blocks according to the detection box coordinates to obtain a first sorting queue; adjusts the order of text blocks in the first sorting queue according to the semantic relationship between text blocks to obtain a second sorting queue; and loads the text blocks in the second sorting queue into the cache queue of the grouping model, wherein the grouping model is used to sort the text blocks belonging to the same test question and the same test question element in the cache queue. Text blocks of the same test question element are grouped together, and the first grouping result is output. Based on the semantic relationship between adjacent groups, the first grouping result is validated. If the validation passes, text blocks from the last group are retained in the cache queue, and other text blocks outside the last group are deleted. The next test question image after the target test question image is obtained, and the sorted text blocks extracted from the next test question image are loaded into the cache queue. This allows the grouping model to output a second grouping result based on the text blocks in the next test question image and the retained text blocks in the cache queue. Based on the grouping result, one or more text clusters are obtained. The text cluster search module is used to search for a baseline text cluster and multiple candidate text clusters in the text cluster set, wherein the candidate text clusters refer to text clusters that are associated with the baseline text cluster; The encoding generation module is used to generate label codes for each text cluster in the baseline text cluster and the candidate text cluster based on the test question elements to which the text cluster belongs, the text blocks in the text cluster, and the position of the text blocks in the test question image; The text clustering module is used to search for text clusters that belong to the same test question as the baseline text cluster in the candidate text clusters based on the label encoding, so as to obtain a set of test question text clusters.

8. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the cross-page question aggregation method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute the cross-page question aggregation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Test question disassembling method and system based on test paper image, storage medium and equipment

    CN113610068A

  • Topic alignment method and device, computer equipment and storage medium

    CN119888770A