Topic generation method and device based on large model, electronic equipment, readable storage medium and computer program product

By separating the problem-solving condition information and utilizing drawing scripts and large model generation methods, the problem of ignoring graphical information in multimodal large models during problem generation is solved, improving the diversity and adaptability of generated problems, and enhancing visual understanding and cross-modal alignment capabilities.

CN121809669APending Publication Date: 2026-04-07BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing multimodal large models tend to ignore graphical information during question generation, resulting in insufficient visual understanding and cross-modal alignment capabilities during training, and a lack of diversity and adaptability in the generated questions.

Method used

By separating the complete problem-solving conditions of the target knowledge point into numerically related secondary information and non-numerical tertiary information, using a drawing script to determine the graphical information, and combining predefined prompts with a large model to generate questions, the diversity and adaptability of the generated questions are ensured.

Benefits of technology

It improves the visual understanding and cross-modal alignment capabilities of large multimodal models, reduces the difficulty of synthesizing high-quality multimodal geometry problem data, and makes the generated problems more logically consistent and applicable at both the text and graphics levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809669A_ABST
    Figure CN121809669A_ABST
Patent Text Reader

Abstract

The invention provides a topic generation method and device based on a large model, electronic equipment, a readable storage medium and a computer program product, and relates to the field of data processing, in particular to the field of intelligent generation. According to the implementation scheme, in response to determining that first information of a target knowledge point comprises second information related to a numerical value, information, except the second information, in the first information serves as third information of the target knowledge point, the first information at least comprises description information about a complete problem solving condition of the target knowledge point; taking the numerical value in the second information as a drawing parameter, and determining graphic information corresponding to the target knowledge point by utilizing a drawing script; and based on the first information, the graphic information, the third information and a predefined first cue word, utilizing a large model to determine at least one type of question corresponding to the target knowledge point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing, and more particularly to the field of intelligent generation, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating questions based on a large model. Background Technology

[0002] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0003] Existing multimodal large-scale models can be improved in many ways in terms of question generation, such as model structure, data allocation and training scheme, and training data in the same domain. Based on current industry experience, collecting training data in the same domain is often the most direct and effective way to improve the performance of large-scale models in that domain.

[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating questions based on a large model.

[0006] According to one aspect of this disclosure, a method for generating questions based on a large model is provided, comprising: responding to the determination that first information of a target knowledge point includes second information related to numerical values, using information in the first information other than the second information as third information of the target knowledge point, wherein the first information includes at least descriptive information about complete problem-solving conditions for the target knowledge point; using the numerical values ​​in the second information as plotting parameters, and using a plotting script to determine graphical information corresponding to the target knowledge point; and based on the first information, the graphical information, the third information, and a predefined first prompt word, using a large model to determine at least one type of questions corresponding to the target knowledge point, wherein the first prompt word includes at least one of the following: question output format specification information and reference case information.

[0007] According to a second aspect of this disclosure, a problem generation apparatus based on a large model is provided, comprising: a description information determination module, configured to, in response to determining that first information of a target knowledge point includes second information related to numerical values, use information in the first information other than the second information as third information of the target knowledge point, wherein the first information includes at least description information of complete problem-solving conditions for the target knowledge point; a graphic information determination module, configured to use the numerical values ​​in the second information as drawing parameters and, using a drawing script, determine graphic information corresponding to the target knowledge point; and a problem generation module, configured to, based on the first information, the graphic information, the third information, and a predefined first prompt word, use a large model to determine at least one type of problem corresponding to the target knowledge point, wherein the first prompt word includes at least one of the following: problem output format specification information and reference case information.

[0008] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the above-described large-model-based question generation method.

[0009] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the problem generation method based on a large model as described above.

[0010] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described large-model-based question generation method.

[0011] According to one or more embodiments of this disclosure, by separating the first information containing complete problem-solving conditions of the target knowledge point into numerically related second information and non-numerical third information, and using the numerical value in the second information as a drawing parameter, the drawing script is used to determine the graphic information corresponding to the target knowledge point. Then, combined with a predefined first prompt word including at least question output format specification information and reference case information, a large model is used to determine at least one type of question corresponding to the target knowledge point, ensuring the diversity of question generation, thereby improving the adaptability of the generated questions to the application scenario.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown; Figure 2 A flowchart of a large-model-based question generation method according to an embodiment of the present disclosure is shown; Figure 3 A flowchart of a large-model-based question generation method according to an embodiment of the present disclosure is shown; Figure 4 A block diagram of a large-model-based question generation apparatus according to an embodiment of the present disclosure is shown; Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0016] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0017] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0018] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0019] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0020] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of content recommendation methods.

[0021] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0022] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0023] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to recommend content. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to users through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0024] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0025] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0026] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0027] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0028] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.

[0029] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0030] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0031] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0032] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0033] Figure 2 This is a flowchart illustrating a large-model-based question generation method 200 according to an embodiment of the present disclosure, such as... Figure 2 As shown, the question generation method may include: step S202, in response to the determination that the first information of the target knowledge point includes second information related to numerical problem-solving conditions, taking the information in the first information other than the second information as the third information of the target knowledge point, wherein the first information includes at least descriptive information about the complete problem-solving conditions of the target knowledge point; step S204, taking the numerical value in the second information as a drawing parameter, and using a drawing script to determine the graphic information corresponding to the target knowledge point; and step S206, based on the first information, the drawing parameter, the third information, and a predefined first prompt word, using a large model to determine at least one type of question corresponding to the target knowledge point, wherein the first prompt word includes at least one of the following: question output format specification information and reference case information.

[0034] In this embodiment, the first information containing complete problem-solving conditions of the target knowledge point is separated into numerically related second information and non-numerical third information. The numerical value in the second information is used as a drawing parameter. Using a drawing script, the graphic information corresponding to the target knowledge point is determined. Then, combined with a predefined first prompt word that includes at least question output format specification information and reference case information, a large model is used to determine at least one type of question corresponding to the target knowledge point, ensuring the diversity of question generation and thus improving the adaptability of the generated questions to the application scenario.

[0035] The method described in this application can provide rich training samples for large multimodal problem-generating models, significantly reducing the difficulty of synthesizing high-quality multimodal geometry problem data. Furthermore, by creating information gaps, it forces the model to combine text and image information to solve problems, effectively solving the pain point of the model ignoring images and relying solely on text during training, and significantly improving the visual understanding and cross-modal alignment capabilities of large multimodal models.

[0036] In some embodiments, in step S202, deep semantic analysis is performed on the initial information (first information) containing complete problem-solving conditions. When specific numerical problem-solving conditions (second information) are identified, a text subtraction operation is automatically performed. Specifically, the numerical parameters (such as line segment length, angle size, etc.) in the first information are extracted and marked, while retaining its logical framework, geometric topological relationships, and solution objectives, thereby forming text (third information) that lacks key numerical values.

[0037] This process ensures that missing conditions at the text level are filled in by conditions at the graphical level in the generated questions, thereby forcing the acquisition of complete solution data through cross-modal alignment during subsequent problem-solving. This effectively solves the pain point of relying solely on text and ignoring accompanying images when solving problems.

[0038] In step S204, the extracted second information is then converted into input parameters for the plotting script, and the information gap is filled by visual annotations in the geometric plot. These numerical parameters are then filled into the corresponding plotting script template (e.g., Python's Matplotlib, LaTeX's TikZ, or Asymptote scripts) according to preset logical mapping rules.

[0039] For example, for a "circle with a radius of 5", the Asymptote script is: draw(circle((0,0),5)); label("5",(2.5,0),N). This coded graphical information ensures the physical association between numerical labels and geometric entities.

[0040] In some embodiments, a large model can be used to dynamically generate an executable code or instruction set based on the geometric logical relationships (such as "parallel," "perpendicular," and "tangent") identified by the first information, combined with precise measurement values ​​provided by the second information. This parameter-driven script determination method ensures the consistency of the generated graphic information and text conditions in mathematical logic. It solves the problems of proportional distortion or annotation errors that are prone to occur in traditional image generation.

[0041] In some embodiments, step S206 involves using a large model as the core generation engine to reintegrate the various information fragments (complete conditions, extracted text, and plotting parameters) obtained previously, ultimately transforming them into structured question data. In this stage, the large model inputs the complete problem description (first information), extracted numerical values ​​(plotting parameters), and extracted skeleton text (third information) into the large model based on a predefined logical framework defined by the first prompt word. The first prompt word specifies the structured requirements of the output (such as JSON or a specific Markdown format) and provides reference examples to ensure the standardization of the generated data, facilitating subsequent automatic parsing by the system. Then, according to the instructions, the large model uses the same set of knowledge point materials to differentiate and determine different types of questions.

[0042] In some embodiments, the step of using information other than the second information in the first information as the third information of the target knowledge point includes: S302, determining the third information based on a predefined second prompt word using the large model, wherein the second prompt word is used to instruct the large model to identify numerical problem-solving conditions in the first information; and S304, using information other than the numerical problem-solving conditions as the third information of the target knowledge point, such that the complete problem-solving conditions of the target knowledge point cannot be determined solely based on the third information.

[0043] In this embodiment, a predefined second prompt word is used to control the large model to accurately extract numerical conditions, ensuring that the complete solution conditions cannot be obtained based solely on the third information. This logically eliminates the generation of questions based solely on plain text information and forces the generated questions to depend on other modal information.

[0044] In some embodiments, step S302 details the technical aspects of semantic parsing and conditional deconstruction using a Large Model (LLM). Specifically, a predefined second cue word is input into the LLM, which configures the LLM as a professional mathematical semantic parsing expert. The core technical logic is as follows: Intelligent identification of numerical conditions: The large model receives initial information (first information) containing complete problem-solving conditions. Guided by the second prompt, it accurately locates and extracts the problem-solving parameters (i.e., numerical problem-solving conditions) that involve specific numerical values, such as the measurement value of a line segment, a specific angle, and the coefficients of a function.

[0045] Automated condition stripping: After identifying numerical values, these specific values ​​are removed from the original text, while ensuring that the non-numerical logical framework of the problem (such as the topological relationship of geometric shapes, the solution objective, and logical judgments such as parallelism or perpendicularity) is completely preserved.

[0046] Generating information gap text: After the above stripping process, the information output by the large model is the third information. Since key values ​​have been removed, the final answer cannot be derived solely from this third information, thus artificially creating an information gap in the text modality.

[0047] Through this precise control based on the second cue word, the system can automatically transform traditional full-text questions into text drafts of questions with strong visual dependence, laying the foundation for the subsequent transfer of the extracted values ​​to explicit annotations in the accompanying figures.

[0048] In some embodiments, in step S304, by completely removing the identified second information (numerical conditions) from the original text, the logical framework, geometric background, and solution objective of the problem are preserved, thereby generating third information. At this point, the third information only contains qualitative descriptions (such as knowing that triangle ABC is a right triangle), lacking quantitative constraints, such as the side length AC being 3cm. This breaks the redundant text-image pattern: in traditional teaching data, graphics are often just repetitive explanations of text; however, through the above processing, the necessary and sufficient conditions for solving the problem cannot be met solely based on the third information; due to the unconsistent logical vacuum at the text level, it is forced that when solving the problem, one must deeply analyze the rendered geometric illustrations to obtain the numerical information labeled in the graphics. In this way, the shortcomings of ignoring illustrations and relying solely on text are solved from the data source.

[0049] In some embodiments, determining at least one type of question corresponding to the target knowledge point using a large model based on the first information, the graphic information, the third information, and a predefined first prompt word includes: instructing the large model to determine a first type of question corresponding to the knowledge point based on the first information, wherein the first type of question is a plain text question that does not contain the graphic information.

[0050] In this embodiment, a predefined first prompt word guides the large model to generate plain text questions without graphical information based solely on the first information (including complete solution conditions), thereby improving the rigor of the question logic and the completeness of the solution conditions in scenarios without graphical dependencies.

[0051] In some embodiments, the large model, through a deep understanding of the initial information, transforms it into standardized plain text questions (Type 1 questions). Since these questions contain complete solution conditions and do not involve graphical information, they constitute the logical reasoning benchmark for the model in the plain text modality. The generation of Type 1 questions ensures the integrity of the dataset. During the generation process, the initial prompts provide reference examples, limiting the randomness of the large model's generation and ensuring that the generated text questions meet industrial-grade training data standards in terms of professional expression, logical rigor, and the structured nature of the solution information (such as JSON or Markdown format).

[0052] In some embodiments, determining at least one type of question corresponding to the target knowledge point using a large model based on the first information, the graphic information, the third information, and a predefined first prompt word includes: instructing the large model to determine a second type of question corresponding to the target knowledge point based on the predefined first prompt word, using the first information and the graphic information, wherein the second type of question includes a first text determined based on the first information and a visual graphic determined based on the graphic information, and wherein the first text includes complete solution conditions for the target knowledge point, and the visual graphic serves as auxiliary information for the first text.

[0053] In this embodiment, by guiding the large model to collaboratively construct text-image questions based on the first information (including complete solution conditions) and accurately generated graphic information, it is ensured that the text carries all the solution logic while the graphics only serve as auxiliary explanations. This effectively eliminates the risk of text-image disconnect and provides text-image aligned question data. In this type of question, the images serve as auxiliary information (redundant), which helps to learn the conventional mapping relationship between text description and geometric features, significantly improving the visual guidance of the questions and the consistency of the solution logic.

[0054] In some embodiments, the large model is instructed to generate a second type of question with auxiliary visual information. The core technical logic is as follows: Question text is generated based on first information containing complete conditions, and a corresponding visual graphic is generated simultaneously by combining graphic information. In this mode, the question text (first text) already provides all the parameters sufficient to solve the problem, while the graphic serves as a visual repetition or supplementary explanation of the textual information. This type of question simulates a conventional illustrated question, where the main purpose of the graphic is to aid in understanding the question's background or geometric structure, rather than being the sole source of solution conditions. This type of data has significant benchmark value when training multimodal large models. It enables the model to learn how to align textual descriptions (such as right-angled triangles) with visual representations (such as graphics with right-angle symbols) given complete textual logic, achieving initial verification and understanding of cross-modal information. This step not only enriches the question types but also, through this information overlap design, provides the model with a smooth transition from simple cognition to complex image recognition, ensuring that the model can advance to higher-order questions with strong visual dependencies after acquiring basic alignment capabilities. In some embodiments, determining at least one type of question corresponding to the target knowledge point using a large model based on the first information, the graphic information, the third information, and a predefined first prompt word includes: instructing the large model to determine a third type of question corresponding to the knowledge point based on the predefined first prompt word, using the third information and the graphic information, wherein the third type of question includes a second text determined based on the third information and a visual graphic determined based on the graphic information, and wherein the complete solution conditions for the third type of question need to be determined by combining the numerical conditions provided by the second text and the visual graphic.

[0055] In this embodiment, the question type with complete solution conditions can only be determined by integrating third-party information and graphical information, effectively simulating the complex scenario of questions combining text and graphics, while improving the accurate communication of numerical conditions.

[0056] In some embodiments, the large model generates second text based on third information lacking key numerical values, and simultaneously renders a visual graph containing numerical annotations by combining graphical information. In this mode, the second text only retains the logical framework of the problem (such as geometric relationships and the solution objective), while the specific numerical conditions necessary for solving the problem (the second information) are uniquely contained in the visual graph. Because the second text is logically incomplete, the model or the solver cannot obtain the answer simply by reading the text. This forces the model to perform visual analysis, extracting numerical values ​​from the visual markers of the graph and associating and aligning them with the text logic to obtain complete solution conditions. This type of problem effectively solves the problem in existing technologies where models tend to be lazy and only look at the text while ignoring the graph. By creating this hard constraint that the problem cannot be solved without looking at the graph, the multimodal large model's ability to understand fine-grained features of geometric graphs is substantially trained, improving the accuracy of cross-modal reasoning.

[0057] In some embodiments, the visual graphics are determined using graphics rendering techniques based on the graphic information.

[0058] In this embodiment, the drawing script is efficiently converted into a visual graphic using graphics rendering technology.

[0059] In some embodiments, previously determined graphical information (essentially drawing instructions or script code) is input into a specific rendering engine, and the transformation from data to visual entities is achieved through graphics rendering technology. The rendering technology can accurately draw geometric figures on a white background or other preset surface according to the drawing parameters in the script, and re-mark the originally extracted numerical problem-solving conditions (secondary information) in the form of visual markers at the corresponding positions in the figure. For example, using drawing tools such as Asymptote and configuring specific rendering level parameters, high-quality, mathematically compliant visual geometric illustrations can be generated. Ensuring the physical realism and consistency of the problem: Through the rendering process, the system can verify the validity of the drawing script; if the rendering is successful, it indicates that the figure is logically and numerically self-consistent, thereby generating an image file containing real measurement information that can be read by the model.

[0060] In some embodiments, the method further includes using a PIL package to align the text image and the geometric image in a set manner, and setting the background color to pure white, so as to improve the visual recognition accuracy of multimodal large models.

[0061] Understandably, the generated visualizations are no longer mere embellishments, but rather a core component of the problem. In the third type of problem, the visualizations serve as the sole medium for carrying numerical information. Solvers must extract data from the rendered annotations through visual analysis to fill in the information gaps in the text, ultimately completing a cross-modal loop between logical and visual information.

[0062] In some embodiments, the first information includes at least one of the following: text information, image information, and video information.

[0063] In this embodiment, by supporting multimodal inputs such as text, images, and videos as the first information, the scope of scene coverage is significantly enhanced, effectively meeting the question-generating needs in different environments.

[0064] In some embodiments, the first information may be, for example, a textual description of “Given a right triangle ABC, angle C equals 90 degrees, side AC = 3 cm, side BC = 4 cm, find the length of the oblique AB.”

[0065] In some embodiments, the first information may be, for example, a JPEG image of a handwritten text title.

[0066] In some embodiments, the first information may be, for example, a 30-second video clip explaining a knowledge point.

[0067] In some embodiments, the method further includes: in response to determining that the first information is image information or video information, performing multimodal parsing on the first information to obtain text description information corresponding to the first information.

[0068] In this embodiment, by automatically generating text descriptions through multimodal parsing of the first information, such as images or videos, the semantic gap between non-text input and text information logic is effectively bridged, improving the coherence of subsequent large-scale model processing.

[0069] For example, for the JPEG image mentioned above, a multimodal large model can be used for OCR recognition and semantic analysis to convert it into text description information corresponding to the image.

[0070] For example, for the above-mentioned video clip explaining geometry problems, keyframe images containing clear problem presentations can be extracted; the keyframes can be parsed to obtain text descriptions, or numerical conditions in the teacher's explanation can be obtained using Automatic Speech Recognition (ASR); and the extracted complete problem-solving conditions (first information) can be standardized to finally generate a third type of training problem that conforms to strong visual dependence features.

[0071] In some embodiments, when processing the first information of the video, keyframes are extracted by calculating the pixel change rate between video frames, and a multimodal model is invoked to vectorize the geometric annotations in the keyframes and convert them into descriptive statements that can be understood by the text model.

[0072] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0073] Figure 4 This is a block diagram illustrating a large-model-based question generation apparatus 400 according to an embodiment of the present disclosure. In some embodiments, such as... Figure 4 As shown, the question generation device includes: a description information determination module 402, configured to, in response to determining that the first information of the target knowledge point includes second information related to numerical values, use the information in the first information other than the second information as the third information of the target knowledge point, wherein the first information includes at least description information about the complete problem-solving conditions of the target knowledge point; a graphic information determination module 404, configured to use the numerical values ​​in the second information as drawing parameters and, using a drawing script, determine the graphic information corresponding to the target knowledge point; and a question generation module 406, configured to, based on the first information, the graphic information, the third information, and predefined prompt words, use a large model to determine at least one type of question corresponding to the target knowledge point.

[0074] It should be noted that, Figure 4 The various modules of the device 400 shown can be connected to the reference. Figure 2The steps in method 200 described correspond to each other. Therefore, the operations, features, and advantages described above for method 200 also apply to apparatus 400 and its included modules and units. For the sake of brevity, some operations, features, and advantages will not be repeated here.

[0075] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0076] refer to Figure 5 The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0077] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0078] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.

[0079] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).

[0080] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0081] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0082] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0083] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0084] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0085] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0086] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0087] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. A question generation method based on a large model, comprising: In response to the fact that the first information for determining the target knowledge point includes second information related to numerical values, the information in the first information other than the second information is taken as the third information of the target knowledge point, wherein the first information includes at least descriptive information about the complete problem-solving conditions of the target knowledge point; Using the values ​​in the second information as drawing parameters, and employing a drawing script, determine the graphical information corresponding to the target knowledge point; and Based on the first information, the graphic information, the third information, and the predefined first prompt word, a large model is used to determine at least one type of question corresponding to the target knowledge point, wherein the first prompt word includes at least one of the following: question output format specification information and reference case information.

2. The method according to claim 1, wherein, The step of using information other than the second information in the first information as the third information of the target knowledge point includes: Based on a predefined second cue word, the large model is used to determine the third information, wherein the second cue word is used to instruct the large model to identify numerical problem-solving condition information in the first information; and The numerical problem-solving condition information is used as the second information, so that the complete problem-solving conditions of the target knowledge point cannot be determined based solely on the third information.

3. The method according to claim 1 or 2, wherein, Based on the first information, the graphic information, the third information, and the predefined first prompt word, the large model is used to determine at least one type of question corresponding to the target knowledge point, including: Based on the predefined first prompt word, the large model is instructed to determine a first type of question corresponding to the knowledge point based on the first information, wherein the first type of question is a plain text question that does not contain the graphic information.

4. The method according to any one of claims 1-3, wherein, Based on the first information, the graphic information, the third information, and the predefined first prompt word, the large model is used to determine at least one type of question corresponding to the target knowledge point, including: Based on the predefined first prompt word, the large model is instructed to determine a second type of question corresponding to the target knowledge point based on the first information and the graphic information. The second type of question includes a first text determined based on the first information and a visual graphic determined based on the graphic information. The first text includes complete solution conditions for the target knowledge point, and the numerical conditions provided by the visual graphic serve as auxiliary information for the first text.

5. The method according to any one of claims 1-4, wherein, Based on the first information, the graphic information, the third information, and the predefined first prompt word, the large model is used to determine at least one type of question corresponding to the target knowledge point, including: Based on the predefined first prompt word, the large model is instructed to determine a third type of question corresponding to the knowledge point based on the third information and the graphic information. The third type of question includes a second text determined based on the third information and a visual graphic determined based on the graphic information. The complete solution conditions for the third type of question need to be determined by combining the numerical conditions provided by the second text and the visual graphic.

6. The method according to claim 4 or 5, wherein, The visualized graphic is determined using graphic rendering technology based on the graphic information.

7. The method according to any one of claims 1-6, wherein, The first information includes at least one of the following: text information, image information, and video information.

8. The method according to claim 7, wherein, The method further includes: In response to determining that the first information is image information or video information, multimodal parsing is performed on the first information to obtain text description information corresponding to the first information.

9. A problem generation device based on a large model, comprising: The description information determination module is used to, in response to the determination that the first information of the target knowledge point includes second information related to numerical values, take the information in the first information other than the second information as the third information of the target knowledge point, wherein the first information includes at least description information of the complete problem-solving conditions of the target knowledge point; The graphic information determination module is used to use the values ​​in the second information as drawing parameters and, using a drawing script, determine the graphic information corresponding to the target knowledge point; and The question generation module is used to determine at least one type of question corresponding to the target knowledge point based on the first information, the graphic information, the third information, and the predefined first prompt word, using a large model. The first prompt word includes at least one of the following: question output format specification information and reference case information.

10. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

12. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-8.