Question result generation method and device, electronic equipment and computer readable medium
Patent Information
- Application Number
- CN202310713263.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-15
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2043-06-15
AI Technical Summary
[0004]第一,做题结果的统计效率较低、且准确性同样也不能得到有效地保障;
[0013]The above-described embodiments of this disclosure have the following beneficial effects: the question-solving result generation method of some embodiments of this disclosure can accurately and efficiently generate question-solving results for question-solving template images. Specifically, the reason why the related question-solving results are not accurate and efficient is that the statistical efficiency of the question-solving results is low, and the accuracy cannot be effectively guaranteed. Based on this, the question-solving result generation method of some embodiments of this disclosure first scans the question-solving template image placed on the question set board at the target location to determine the question type, so as to facilitate the initiation of the question-solving operation. Then, the question type for the above-mentioned question-solving template image is determined to facilitate the subsequent determination of question-solving rule information. Next, the configured audio device is instructed to play the question-solving rule information for the above-mentioned question type to notify the question-solving sample image of the question-solving rule. Furthermore, in response to receiving the operation information for the operation push button in the above-mentioned question set board, the playback of the above-mentioned question-solving rule information is stopped, and the question-solving timing operation is executed. Here, by stopping the playback of the question-solving rule information in a timely manner, the subsequent impact on answering questions is avoided. And through the question-solving timing operation, the answering efficiency and subsequent calculation efficiency can be guaranteed. Finally, in response to receiving a completion message for the above-mentioned question template image within a predetermined time, at least one operation push button area and the question template image area corresponding to the question set board are rescanned to accurately generate the question-solving result for the above-mentioned question template image. Thus, by determining the question type and effectively playing and terminating the question-solving rule information, not only can effective answering be ensured, but also accurate question-solving results can be generated in a timely manner.
Smart Images

Figure CN118471034B_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to methods, apparatus, electronic devices, and computer-readable media for generating test results. Background Technology
[0002] Currently, with the rapid development of artificial intelligence, early childhood education products based on AI algorithms are increasingly appearing in people's lives. Regarding the determination of the corresponding test results for a test set, the sampling method is usually as follows: relevant personnel manually count the number of incorrect questions to generate the test results.
[0003] However, the inventors discovered that the following technical problems often arise when using the above method:
[0004] First, the statistical efficiency of the test results is low, and the accuracy cannot be effectively guaranteed either.
[0005] Second, the low dimensionality and precision of feature extraction lead to low accuracy in target recognition models.
[0006] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0007] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0008] Some embodiments of this disclosure provide methods, apparatus, electronic devices, and computer-readable media for generating test results to address one or more of the technical problems mentioned in the background section above.
[0009] In a first aspect, some embodiments of this disclosure provide a method for generating test results, including: scanning a test template image on a test set board placed at a target location; determining the test type for the test template image; instructing a configured audio device to play test rule information for the test type; in response to receiving operation information for operation buttons in the test set board, stopping the playback of the test rule information and performing a test timing operation; and in response to receiving response completion information for the test template image within a predetermined time, rescanning at least one operation button area and a test template image area corresponding to the test set board to generate a test result for the test template image.
[0010] Secondly, some embodiments of this disclosure provide a test result generation apparatus, comprising: a scanning unit configured to scan a test template image placed on a test set board at a target location; a determining unit configured to determine a test type for the test template image; an indicating unit configured to instruct a configured audio device to play test rule information for the test type; a stopping unit configured to stop playing the test rule information and perform a test timing operation in response to receiving operation information for an operation button in the test set board; and a generating unit configured to rescan at least one operation button area and a test template image area corresponding to the test set board in response to receiving response completion information for the test template image within a predetermined time period, so as to generate a test result for the test template image.
[0011] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.
[0012] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.
[0013] The above-described embodiments of this disclosure have the following beneficial effects: the question-solving result generation method of some embodiments of this disclosure can accurately and efficiently generate question-solving results for question-solving template images. Specifically, the reason why the related question-solving results are not accurate and efficient is that the statistical efficiency of the question-solving results is low, and the accuracy cannot be effectively guaranteed. Based on this, the question-solving result generation method of some embodiments of this disclosure first scans the question-solving template image placed on the question set board at the target location to determine the question type, so as to facilitate the initiation of the question-solving operation. Then, the question type for the above-mentioned question-solving template image is determined to facilitate the subsequent determination of question-solving rule information. Next, the configured audio device is instructed to play the question-solving rule information for the above-mentioned question type to notify the question-solving sample image of the question-solving rule. Furthermore, in response to receiving the operation information for the operation push button in the above-mentioned question set board, the playback of the above-mentioned question-solving rule information is stopped, and the question-solving timing operation is executed. Here, by stopping the playback of the question-solving rule information in a timely manner, the subsequent impact on answering questions is avoided. And through the question-solving timing operation, the answering efficiency and subsequent calculation efficiency can be guaranteed. Finally, in response to receiving a completion message for the above-mentioned question template image within a predetermined time, at least one operation push button area and the question template image area corresponding to the question set board are rescanned to accurately generate the question-solving result for the above-mentioned question template image. Thus, by determining the question type and effectively playing and terminating the question-solving rule information, not only can effective answering be ensured, but also accurate question-solving results can be generated in a timely manner. Attached Figure Description
[0014] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0015] Figure 1 This is a flowchart of some embodiments of the method for generating test results according to this disclosure;
[0016] Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the test result generation apparatus according to this disclosure;
[0017] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0023] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] refer to Figure 1 The flowchart 100 illustrates some embodiments of a method for generating test results according to the present disclosure. This method for generating test results includes the following steps:
[0025] Step 101: Scan the image of the problem template on the problem set board placed at the target location.
[0026] In some embodiments, the execution subject of the above-described method for generating test results (e.g., an electronic device) can scan the test template image placed on a test set board at a target location via a wired or wireless connection. The target location can be directly below the electronic device. The test set board can be a device for placing the test template image. Multiple sliding buttons are also provided around the perimeter of the test set board. Each button has a different color and / or shape.
[0027] Step 102: Determine the question type for the above-mentioned question template image.
[0028] In some embodiments, the executing entity may determine the question type for the aforementioned question template image. The question type may be the question type of each question to be answered associated with the question template image. In practice, the question type may include, but is not limited to, at least one of the following: multiple choice questions, fill-in-the-blank questions, and matching questions.
[0029] In some optional implementations of certain embodiments, determining the question type for the aforementioned question template image may include the following steps:
[0030] The first step is to perform text title recognition on the text in the above-mentioned sample image to obtain recognition information. This recognition information can be the title recognition information of the titles appearing in the sample image. Title recognition information can include: title location information, title semantic information, and title level information. Title location information represents the position of the title in the sample image. Title semantic information represents the semantic content of the title.
[0031] As an example, firstly, the aforementioned execution entity can input the sample image of the question into a pre-trained text recognition model to generate recognition information. The text recognition model can be a Transformer model.
[0032] The second step is to generate the above question types based on the identification information.
[0033] As an example, the aforementioned executing entity can determine the question type based on the identification information, including title location information, title semantic information, and title level information.
[0034] Step 103: Instruct the configured audio device to play the question-answering rules information for the above-mentioned question-answering type.
[0035] In some embodiments, the executing entity may instruct a configured audio device to play question-answering rule information for the aforementioned question type. The audio device may be a device pre-configured by the executing entity for playing the question-answering rule information. The question-answering rule information may be rule information characterizing how to answer each question corresponding to the question sample image. For example, for answering questions corresponding to the question sample image that are animal matching questions, the question-answering rule information may be information characterizing how to push the corresponding color or shape push button to the specific position on the question set board.
[0036] Step 104: In response to receiving operation information for the operation push button on the above-mentioned question set board, stop playing the above-mentioned question-solving rule information and execute the question-solving timer operation.
[0037] In some embodiments, in response to receiving operation information for the operation push button on the question set board, the execution entity may stop playing the question-solving rule information and perform a question-solving timing operation. The question-solving timing operation may be an operation to time the question-solving process.
[0038] Step 105: In response to receiving a response completion message for the above-mentioned question template image within a predetermined time period, rescan at least one operation push button area and question template image area corresponding to the above-mentioned question set board to generate a question-solving result for the above-mentioned question template image.
[0039] In some embodiments, in response to receiving completion information for the above-mentioned question template image within a predetermined time period, the executing entity may rescan at least one operation push button area and the question template image area corresponding to the question set board to generate a question-answering result for the above-mentioned question template image. The area corresponding to the question set board may include at least one operation push button area and a question template image area. The at least one operation push button area may be the push button area of various push buttons arranged around the perimeter. The question template image area may be the area where the question template image is placed. The completion information may be information indicating that each question corresponding to the question template image has been answered.
[0040] As an example, firstly, in response to receiving a completion message for the aforementioned question template image within a predetermined time period, the executing entity can obtain the answer information for the aforementioned question template image. Then, by rescanning at least one operation push button area and the question template image area corresponding to the question set board, a response result is generated. Finally, based on the answer information and the response result, a question-solving result is generated.
[0041] In some optional implementations of certain embodiments, after step 105, the steps may further include:
[0042] First, in response to the failure to receive the completion response information within the predetermined time period, the executing entity may instruct the audio device to play a continuation response confirmation message for the question template image. In practice, the predetermined time period may be a duration set specifically for the question template image. The continuation response confirmation message continues to provide information on each question associated with the question template image.
[0043] The second step involves receiving a confirmation message indicating no further response in response to the aforementioned confirmation message. The executing entity can then rescan at least one operation button area and a sample image area corresponding to the question set board to generate a solution for the sample image. The specific implementation of rescanning at least one operation button area and a sample image area corresponding to the question set board to generate a solution for the sample image will not be elaborated here; please refer to the implementation method described in step 105.
[0044] In some optional implementations of certain embodiments, the rescanning of at least one operation push button area and the question template image area corresponding to the question set board to generate a question-solving result for the question template image may include the following steps:
[0045] The first step is for the execution entity to rescan at least one operation push button area and the question template image area corresponding to the question set board to generate at least one operation push button area image and candidate question template image.
[0046] The second step involves the aforementioned executing entity performing a consistency check on the aforementioned candidate question template images and the aforementioned question template images to obtain the check results.
[0047] As an example, the aforementioned executing entity can compare the image pixel differences between the aforementioned candidate question template image and the aforementioned question template image to perform consistency verification and obtain the verification result.
[0048] Thirdly, in response to determining that the above verification result representation image is consistent and that the above question-solving type is one-way color matching, the above executing entity can determine the question-solving direction information corresponding to the above question-solving template image. Here, one-way color matching can represent the type of operation push button placed in one direction on the question set board, corresponding to the semantic matching of the question-solving template image at the corresponding position. The question-solving direction information can be the direction information of the operation push button that can be pushed in a certain direction on the question set board. For example, the question-solving direction information corresponding to the question-solving direction end can be one of the following: the left end of the question set board, the right end of the question set board, the top end of the question set board, or the bottom end of the question set board.
[0049] Fourth, the executing entity can determine the operation push button area image that corresponds to the question-answering direction information from at least one operation push button area image, and designate it as the target operation push button area image. The orientation of the target operation push button area image is the same as the orientation corresponding to the question-answering direction information. For example, the target operation push button area image is the image located at the left end of at least one operation push button area image. The question-answering direction information corresponds to the left end of the question set board.
[0050] Fifth, the aforementioned executing entity can input the image of the target operation push button area into a pre-trained target recognition model to generate a first push button position information set and a first push button color category information set. The target recognition model can be a model that recognizes push button information. In practice, the target recognition model can be a multi-layered, serially connected convolutional neural network. The first push button position information can be the position information of the push button on the sample image. The first push button color category information can be the color information of the push button. There is a one-to-one correspondence between the first push button position information in the first push button position information set and the first push button color category information in the first push button color category information set.
[0051] The sixth step is to input the above candidate question template images into the above target recognition model to generate the first subject location information set and the first subject color category information set.
[0052] Step 7: Based on the above-mentioned first push button position information set, the above-mentioned first push button color category information set, the above-mentioned first main body position information set, and the above-mentioned first main body color category information set, generate the above-mentioned question-solving results.
[0053] As an example, the aforementioned execution entity can compare each operation push button with the entity by comparing the first push button position information set with the first entity position information set, and by comparing the first push button color category information set with the first entity color category information set, in order to generate comparison results as the answer to the question.
[0054] Optionally, after determining the question-solving direction information corresponding to the question-solving template image in response to determining that the verification result representation image is consistent and that the question-solving type is unidirectional color matching, the method further includes the following steps:
[0055] The first step, in response to determining that the above verification result representation image is consistent and that the above question-solving type is bidirectional color matching, is to determine the question-solving direction end information group corresponding to the above question-solving template image. Here, bidirectional color matching can represent the type of operation push buttons of corresponding colors placed in both directions on the question set board, semantically matching them with the question-solving template image at the corresponding position. The question-solving direction end information group can be the direction end information group where operation push buttons can be pushed in two directions on the question set board. In practice, the question-solving direction end group corresponding to the question-solving direction end information group can include: the left direction end and the right direction end of the question set board.
[0056] The second step is to identify the group of operation push button area images that corresponds to the group of information about the direction of answering questions in at least one of the above operation push button area images.
[0057] The third step is to input the operation push button area image from the above operation push button area image group into the above target recognition model to generate the second push button position information set and the second push button color category information set, thus obtaining the second push button position information set group and the second push button color category information set group.
[0058] The fourth step is to generate the above-mentioned answer result based on the above-mentioned second push button position information set, the above-mentioned second push button color category information set, the above-mentioned first main body position information set, and the above-mentioned first main body color category information set.
[0059] As an example, the aforementioned execution entity can perform a one-to-one comparison between the corresponding two operation push buttons and the entity by comparing the second push button position information set group and the first entity position information set group, and by comparing the second push button color category information set group and the first entity color category information set group, in order to generate a comparison result as the result of answering the question.
[0060] Optionally, the consistency verification of the candidate question-answering template image and the question-answering template image obtained above may include the following steps:
[0061] The first step involves inputting the aforementioned sample problem image and the aforementioned candidate sample problem images into an image color region labeling model to generate first color labeling information for the sample problem image and second color labeling information for the candidate sample problem image. The image color region labeling model can be a model that colors each region of the candidate sample problem image. The first color labeling information represents the color information of each region corresponding to the sample problem image. The second color labeling information represents the color information of each region corresponding to the candidate sample problem image. In practice, the aforementioned image color region labeling model can be a residual neural network model based on an attention mechanism.
[0062] The second step involves, in response to the determination that the first color marking information and the second color marking information are the same, inputting the sample problem-solving image and the candidate sample problem-solving image into the image layout segmentation model to generate at least one first image layout information for the sample problem-solving image and at least one second image layout information for the candidate sample problem-solving image. The image layout segmentation model can be a model that performs layout segmentation based on the image content. In practice, the image layout segmentation model can be a semantic segmentation model. Specifically, the image layout segmentation model can be a U-Net neural network based on an attention mechanism. The semantic content corresponding to each first image layout information in at least one first image layout information differs significantly. Similarly, the semantic content corresponding to each second image layout information in at least one second image layout information differs significantly.
[0063] The third step is to input each of the first image layout information from the above-mentioned at least one first image layout information into the target recognition model to output a second subject position information set and a second subject color category information set, thereby obtaining at least one second subject position information set and at least one second subject color category information set.
[0064] The fourth step is to generate at least one set of first subject position information and at least one set of first subject color category information for the above-mentioned at least one second image layout information.
[0065] As an example, the aforementioned execution entity can divide the information of each first subject position information and each first subject color category information by using the first subject position information and the first subject color category information corresponding to each subject and the second image layout information corresponding to each subject, so as to obtain at least one set of first subject position information and at least one set of first subject color category information for the aforementioned at least one second image layout information.
[0066] Fifth step: Generate the above verification result based on the above at least one second subject location information set, the above at least one second subject color category information set, the above at least one first subject location information set, and the above at least one first subject color category information set.
[0067] As an example, the aforementioned executing entity can generate a comparison result as a verification result by comparing at least one second subject location information set with at least one first subject location information set and comparing at least one second subject color category information set with at least one first subject color category information set.
[0068] In some optional implementations of certain embodiments, the target recognition model includes: a feature extraction model, a first downsampled feature pyramid model, a first upsampled feature pyramid model, a second downsampled feature pyramid model, a second upsampled feature pyramid model, and a push button information output layer. The feature extraction model can be a model for extracting image feature information. In practice, the feature extraction model can be a multi-layered, serially connected convolutional neural network. The first downsampled feature pyramid model can include multiple serially connected first residual networks. The vector dimension of the network output vectors of the multiple serially connected first residual networks decreases sequentially. The first upsampled feature pyramid model can include multiple serially connected second residual networks. The vector dimension of the network output vectors of the multiple serially connected second residual networks increases sequentially. The second downsampled feature pyramid model can include multiple serially connected third residual networks. The vector dimension of the network output vectors of the multiple serially connected third residual networks decreases sequentially. The second upsampled feature pyramid model can include multiple serially connected fourth residual networks. The vector dimension of the network output vectors of the multiple serially connected fourth residual networks increases sequentially. The push button information output layer can be a multi-layered, serially connected fully connected layer.
[0069] Optionally, inputting the target operation button area image into a pre-trained target recognition model to generate a first button position information set and a first button color category information set may include the following steps:
[0070] The first step involves inputting the target operation push button area image into the feature extraction model to output first feature information. This first feature information characterizes the image features of the target operation push button area image.
[0071] The second step involves inputting the aforementioned first feature information into the aforementioned first downsampled feature pyramid model to output a sequence of first-level feature output information for each level of the aforementioned first downsampled feature pyramid model. This sequence of first-level feature output information can be the output information of each network corresponding to multiple serially connected first residual networks.
[0072] The third step is to perform feature fusion on the above-mentioned first-level feature output information sequence to generate the first fused feature information.
[0073] As an example, firstly, the aforementioned execution entity can unify the vector dimension of the first-level feature output information sequence to generate a unified hierarchical feature output information sequence. Then, the information of each unified hierarchical feature output information in the unified hierarchical feature output information sequence is added together to obtain the added information, which serves as the first fused feature information.
[0074] The fourth step is to generate a second-level feature output information sequence for each level of the first-level feature pyramid model based on the first-level feature output information sequence described above, using the first upsampled feature pyramid model described above.
[0075] As an example, firstly, the first target-level feature output information from the first-level feature output information sequence is input into the first-level second residual network of the multiple serially connected second residual networks included in the first upsampled feature pyramid model, to output residual results. The first target-level feature can be the last first-level feature output information in the first-level feature output information sequence. Then, the first residual result and the first-level feature output information with the same vector dimension in the first-level feature output information sequence are added together to obtain the sum result. Similarly, by inputting into the next-level residual network and adding the residual results, the residual results of each level of residual network are generated. Finally, the residual results of each level of residual network are determined as the second-level feature output information sequence.
[0076] The fifth step involves fusing the second-level feature output information sequence to generate the second fused feature information. The specific implementation method will not be detailed here; please refer to the section on fusing the first-level feature output information sequence for details.
[0077] Step 6: Based on the second-level feature output information sequence described above, and using the second downsampled feature pyramid model described above, generate third-level feature output information sequences for each level corresponding to the second downsampled feature pyramid model. For details on the implementation, please refer to Step 4, "Generation of the Second-Level Feature Output Information Sequence".
[0078] The seventh step involves fusing the aforementioned third-level feature output information sequence to generate third-level fused feature information. The specific implementation method will not be detailed here; please refer to the section on fusing the first-level feature output information sequence for details.
[0079] Step 8: Based on the third-level feature output information sequence described above, and using the second upsampled feature pyramid model described above, generate fourth-level feature output information sequences for each level corresponding to the second upsampled feature pyramid model. For details on the implementation, please refer to Step 4, "Generation of the Second-Level Feature Output Information Sequence".
[0080] The ninth step involves fusing the aforementioned fourth-level feature output information sequence to generate fourth-level fused feature information. The specific implementation method will not be detailed here; please refer to the section on fusing the first-level feature output information sequence for details.
[0081] Step 10: Perform a weighted summation of the first, second, third, and fourth fusion feature information to generate weighted summation feature information. For details on this implementation, please refer to Step 4, "Generation of the Second-Level Feature Output Information Sequence".
[0082] Step 11: Input the weighted summation feature information into the push button information output layer to output the first push button position information set and the first push button color category information set.
[0083] Optionally, the aforementioned target recognition model includes: a multi-dimensional feature extraction model, a third downsampled feature pyramid model, a third upsampled feature pyramid model, a deep feature extraction network, a feature fusion layer, and a push button information output layer. The multi-dimensional feature extraction model may include three feature extraction models. The feature extraction dimensions of the three feature extraction models are different, that is, the vector output dimensions of the corresponding feature extraction models are different. The third downsampled feature pyramid model may include a multi-layered, serially connected fifth residual network. The third upsampled feature pyramid model may include a multi-layered, serially connected sixth residual network. The deep feature extraction network includes a multi-layered, serially connected convolutional neural network. The feature fusion layer may be a network layer that fuses feature information.
[0084] Optionally, inputting the target operation button area image into a pre-trained target recognition model to generate a first button position information set and a first button color category information set may include the following steps:
[0085] The first step involves inputting the target operation push button area image into a multi-dimensional feature extraction model to output first image feature information, second image feature information, and third image feature information. The vector dimension corresponding to the first image feature information is greater than the vector dimension corresponding to the second image feature information. The vector dimension corresponding to the second image feature information is greater than the vector dimension corresponding to the third image feature information.
[0086] The second step is to input the first image feature information into the third downsampled feature pyramid model to output the fifth level feature output information sequence.
[0087] The third step is to input the third image feature information into the third upsampled feature pyramid model to output the sixth level feature output information sequence.
[0088] The fourth step is to input the second image feature information into the deep feature extraction network to output deep feature extraction information.
[0089] The fifth step involves using a feature fusion layer to fuse the various fifth-level feature outputs in the fifth-level feature output sequence to obtain the fifth fused feature information.
[0090] The sixth step involves using a feature fusion layer to fuse the various sixth-level feature outputs in the sixth-level feature output sequence to obtain the sixth-level fused feature information.
[0091] Step 7: Multiply the fifth fused feature information with the first feature coefficient to obtain the first multiplied feature information, and multiply the sixth fused feature information with the second feature coefficient to obtain the second multiplied feature information. The first feature coefficient characterizes the feature importance of the fifth fused feature information. The second feature coefficient characterizes the feature importance of the sixth fused feature information. The first and second feature coefficients can continuously change as the target recognition model is trained.
[0092] The eighth step involves using a feature fusion layer to fuse the first multiplied feature information, the second multiplied feature information, and the deep feature extraction information to generate the seventh fused feature information.
[0093] The ninth step is to input the seventh fusion feature information into the push button information output layer to output the first push button position information set and the first push button color category information set.
[0094] As one of the invention's key points, this invention addresses the second technical problem mentioned in the background art: "Due to the low dimensionality and precision of feature extraction, the recognition accuracy of the target recognition model is low." Based on this, the method for generating the test results disclosed herein firstly extracts a multi-depth, multi-dimensional set of feature information (i.e., first image feature information, second image feature information, and third image feature information) from the target operation push button area image using a multi-dimensional feature extraction model. Then, it uses a third downsampling feature pyramid model, a third upsampling feature pyramid model, and a deep feature extraction network to perform corresponding feature depth extraction on the first, second, and third image feature information, and to normalize the feature dimensions of the depth-extracted feature information into a unified dimension to facilitate subsequent feature information fusion. It should be noted that by setting the first and second feature coefficients, the importance of the first and third image feature information can be clearly reflected. Therefore, the feature fusion of the first, second, and third image feature information can be accurately achieved, allowing the resulting seventh fused feature information to more completely reflect the image features of the target operation push button area image. Finally, through the push button information output layer, the first push button position information set and the first push button color category information set can be accurately generated.
[0095] The above-described embodiments of this disclosure have the following beneficial effects: the question-solving result generation method of some embodiments of this disclosure can accurately and efficiently generate question-solving results for question-solving template images. Specifically, the reason why the related question-solving results are not accurate and efficient is that the statistical efficiency of the question-solving results is low, and the accuracy cannot be effectively guaranteed. Based on this, the question-solving result generation method of some embodiments of this disclosure first scans the question-solving template image placed on the question set board at the target location to determine the question type, so as to facilitate starting the question-solving operation. Then, the question type for the above-mentioned question-solving template image is determined to facilitate the subsequent determination of question-solving rule information. Next, the configured audio device is instructed to play the question-solving rule information for the above-mentioned question type to notify the question-solving sample image of the question-solving rule. Furthermore, in response to receiving the operation information for the operation push button in the above-mentioned question set board, the playback of the above-mentioned question-solving rule information is stopped, and the question-solving timing operation is executed. Here, by stopping the playback of the question-solving rule information in a timely manner, the answering is avoided. Through the question-solving timing operation, the answering efficiency and subsequent calculation efficiency can be guaranteed. Finally, in response to receiving a completion message for the above-mentioned question template image within a predetermined time, at least one operation push button area and the question template image area corresponding to the question set board are rescanned to accurately generate the question-solving result for the above-mentioned question template image. Thus, by determining the question type and effectively playing and terminating the question-solving rule information, not only can effective answering be guaranteed, but also accurate question-solving results can be generated in a timely manner.
[0096] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a problem-solving result generation device, which are similar to... Figure 1 Corresponding to the method embodiments shown, the problem-solving result generation device can be specifically applied to various electronic devices.
[0097] like Figure 2As shown, a problem-solving result generation device 200 includes: a scanning unit 201, a determining unit 202, an indicating unit 203, a stopping unit 204, and a generation unit 205. The scanning unit 201 is configured to scan a problem-solving template image placed on a problem set board at a target location; the determining unit 202 is configured to determine the problem type corresponding to the problem-solving template image; the indicating unit 203 is configured to instruct a configured audio device to play problem-solving rule information for the problem type; the stopping unit 204 is configured to stop playing the problem-solving rule information and perform a problem-solving timing operation in response to receiving operation information for operation buttons on the problem set board; the generation unit 205 is configured to rescan at least one operation button area and a problem-solving template image area corresponding to the problem set board in response to receiving response completion information for the problem-solving template image within a predetermined time period, to generate a problem-solving result for the problem-solving template image.
[0098] It is understandable that the units recorded in the problem-solving result generation device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the problem-solving result generation device 200 and the units contained therein, and will not be repeated here.
[0099] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0100] like Figure 3 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0101] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0102] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0103] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0104] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0105] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently without being assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: scan a sample problem image placed on a problem set board at a target location; determine the problem type corresponding to the sample problem image; instruct a configured audio device to play problem-solving rule information for the problem type; in response to receiving operation information for an operation button on the problem set board, stop playing the problem-solving rule information and perform a problem-solving timing operation; and in response to receiving a response completion information for the sample problem image within a predetermined time, rescan at least one operation button area and a sample problem image area corresponding to the problem set board to generate a problem-solving result for the sample problem image.
[0106] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0108] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including a scanning unit, a determining unit, an indicating unit, a stopping unit, and a generating unit. The names of these units do not necessarily limit the specific unit; for example, a scanning unit may also be described as "a unit that scans the image of a sample problem set placed on a problem set board at a target location."
[0109] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0110] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for generating problem-solving results, comprising: Scan the image of the problem template on the problem set board placed at the target location; Determine the question type for the provided question template image; The configured audio device is instructed to play the question-answering rule information for the specified question type; In response to receiving operation information for the operation push button in the question set board, stop playing the question-solving rule information and execute the question-solving timer operation; In response to receiving a completion message for the question template image within a predetermined time period, the system rescans at least one operation push button area and the question template image area corresponding to the question set board to generate a question-solving result for the question template image. The step of rescanning at least one operation push button area and a question template image area corresponding to the question set board to generate a question-solving result based on the question template image includes: Rescan at least one operation push button area and the question template image area corresponding to the question set board to generate at least one operation push button area image and candidate question template image; The candidate question-answering template image and the question-answering template image are subjected to consistency verification to obtain the verification result; In response to determining that the verification result representation image is consistent and that the question type is unidirectional color matching, the question-solving direction information corresponding to the question-solving template image is determined; Determine the operation push button area image that corresponds to the question-answering direction information from the at least one operation push button area image, and use it as the target operation push button area image; The image of the target operation push button area is input into a pre-trained target recognition model to generate a first push button position information set and a first push button color category information set; The candidate question template image is input into the target recognition model to generate a first subject location information set and a first subject color category information set; The answer result is generated based on the first push button position information set, the first push button color category information set, the first main body position information set, and the first main body color category information set. The step of performing a consistency check on the candidate question-answering template image and the question-answering template image to obtain the check result includes: The sample image and the candidate sample image are respectively input into the image color region marking model to generate first color marking information for the sample image and second color marking information for the candidate sample image; In response to determining that the first color marking information and the second color marking information are the same, the sample image for answering questions and the candidate sample image for answering questions are respectively input into the image layout segmentation model to generate at least one first image layout information for the sample image for answering questions and at least one second image layout information for the candidate sample image for answering questions. Each of the first image layout information in the at least one first image layout information is input into the target recognition model to output a second subject position information set and a second subject color category information set, thereby obtaining at least one second subject position information set and at least one second subject color category information set; Generate at least one first subject position information set and at least one first subject color category information set for the at least one second image layout information; The verification result is generated based on the at least one second subject location information set, the at least one second subject color category information set, the at least one first subject location information set, and the at least one first subject color category information set.
2. The method according to claim 1, wherein, The method further includes: If the response completion information is not received within the predetermined time period, the audio device is instructed to play a continuation response confirmation message for the question template image; In response to receiving a confirmation message indicating no further response, the system rescans at least one operation push button area and a question template image area corresponding to the question set board to generate a question-solving result based on the question template image.
3. The method according to claim 2, wherein, After determining the question-solving direction information corresponding to the question-solving template image in response to determining that the verification result representation image is consistent and that the question-solving type is unidirectional color matching, the method further includes: In response to determining that the verification result representation image is consistent and determining that the question type is bidirectional color matching, the question direction end information group corresponding to the question template image is determined; Determine the group of operation push button area images that corresponds to the question-answering direction end information group in the at least one operation push button area image; The operation push button area image in the operation push button area image group is input into the target recognition model to generate a second push button position information set and a second push button color category information set, thus obtaining a second push button position information set group and a second push button color category information set group. The question-answering result is generated based on the second push button position information set, the second push button color category information set, the first main body position information set, and the first main body color category information set.
4. The method according to claim 3, wherein, Determining the question type for the question template image includes: Text title recognition is performed on the text in the sample image to obtain recognition information; The question type is generated based on the identification information.
5. The method according to claim 4, wherein, The target recognition model includes: a feature extraction model, a first downsampled feature pyramid model, a first upsampled feature pyramid model, a second downsampled feature pyramid model, a second upsampled feature pyramid model, and a push button information output layer; and The step of inputting the target operation button area image into a pre-trained target recognition model to generate a first button position information set and a first button color category information set includes: The target operation push button area image is input into the feature extraction model to output the first feature information; The first feature information is input into the first downsampled feature pyramid model to output the first level feature output information sequence for each level corresponding to the first downsampled feature pyramid model; The first-level feature output information set is fused to generate first fused feature information; Based on the first-level feature output information sequence, and using the first upsampled feature pyramid model, a second-level feature output information sequence is generated for each level corresponding to the first upsampled feature pyramid model. The second-level feature output information sequence is fused to generate second fused feature information; Based on the second-level feature output information sequence, the third-level feature output information sequence for each level corresponding to the second downsampled feature pyramid model is generated using the second downsampled feature pyramid model. The third-level feature output information sequence is fused to generate third fused feature information; Based on the third-level feature output information sequence, the fourth-level feature output information sequence for each level corresponding to the second upsampled feature pyramid model is generated using the second upsampled feature pyramid model. The fourth-level feature output information sequence is fused to generate fourth fused feature information; The first fused feature information, the second fused feature information, the third fused feature information, and the fourth fused feature information are subjected to a weighted summation process to generate weighted summation feature information; The weighted summation feature information is input to the push button information output layer to output the first push button position information set and the first push button color category information set.
6. A device for generating test results, comprising: The scanning unit is configured to scan the problem template image placed on the problem set board at the target location; The determining unit is configured to determine the question type for the question template image; The instruction unit is configured to instruct the configured audio device to play question-answering rule information for the question type; The stop unit is configured to stop playing the question-solving rule information and perform a question-solving timer operation in response to receiving operation information for the operation push button in the question set board; The generation unit is configured to, in response to receiving a response completion information for the question template image within a predetermined time period, rescan at least one operation push button area and a question template image area corresponding to the question set board to generate a question-solving result for the question template image. The rescanning of at least one operation push button area and a question template image area corresponding to the question set board to generate a question-solving result for the question template image includes: rescanning at least one operation push button area and a question template image area corresponding to the question set board to generate at least one operation push button area image and a candidate question template image; performing a consistency check on the candidate question template image and the question template image to obtain a check result; and responding to determining the check result... The test results characterize the images as consistent, and the question type is unidirectional color matching. The question-solving direction information corresponding to the question-solving template image is determined. The operation push button area image corresponding to the question-solving direction information in the at least one operation push button area image is determined as the target operation push button area image. The target operation push button area image is input into a pre-trained target recognition model to generate a first push button position information set and a first push button color category information set. The candidate question-solving template image is input into the target recognition model to generate a first subject position information set and a first subject color category information set. Based on the first push button position information set, the first push button color category information set, the first subject position information set, and the... A first primary color category information set is used to generate the question-solving result. The step of performing a consistency check on the candidate question-solving template image and the question-solving template image to obtain a check result includes: inputting the question-solving template image and the candidate question-solving template image into an image color region marking model to generate first color marking information for the question-solving template image and second color marking information for the candidate question-solving template image; in response to determining that the first color marking information and the second color marking information are the same, inputting the question-solving template image and the candidate question-solving template image into an image layout segmentation model to generate at least one first image layout information for the question-solving template image and at least one first image layout information for the candidate question-solving template image. At least one second image layout information of the sample image for answering questions; input each of the at least one first image layout information into the target recognition model to output a second subject position information set and a second subject color category information set, thereby obtaining at least one second subject position information set and at least one second subject color category information set; generate at least one first subject position information set and at least one first subject color category information set for the at least one second image layout information; generate the verification result based on the at least one second subject position information set, the at least one second subject color category information set, the at least one first subject position information set, and the at least one first subject color category information set.
7. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Answering system, and client and client system used for answering
CN104933921A
Training board recognition method and device and robot
CN112200230A
Question-doing information display method and device, electronic equipment and computer readable medium
CN118471030A