Image-text matching method, electronic equipment and storage medium

By performing knowledge point entity recognition and deep semantic analysis on the test text, combined with knowledge graphs and deep learning, we have achieved efficient and accurate illustration selection for text-image matching, solving the problems of low efficiency and poor accuracy in existing technologies, and adapting to the diversified needs of educational content.

CN121579719AActive Publication Date: 2026-02-27ANHUI FEISHU INFORMATION TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511540515.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-02-27
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing technologies for image-text matching in educational informatization suffer from low efficiency and poor accuracy, especially in terms of insufficient deep semantic understanding and cross-modal association, resulting in a disconnect between the matched images and the content, and failing to meet the needs of different subjects and learning stages.

Method used

By performing knowledge point entity recognition, question type classification, and difficulty classification on the test question text, and by using knowledge graphs and deep learning models, combined with illustration requirement information and candidate illustration description information, the target illustration can be accurately matched.

Benefits of technology

It improves the efficiency and accuracy of image-text matching, adapts to the needs of different disciplines and learning stages, and supports the intelligent and large-scale production of educational content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579719A_ABST
    Figure CN121579719A_ABST
Patent Text Reader

Abstract

The invention provides an image-text matching method, electronic equipment and a storage medium. The method comprises the following steps: acquiring a test question text; performing knowledge point entity identification on the test question text to obtain knowledge point entities in the test question text and an entity relationship of the knowledge point entities, and performing question type classification and difficulty classification on the test question text to obtain a test question type and test question difficulty of the test question text; determining illustration demand information of the test question text based on knowledge point entities in the test question text, an entity relationship of the knowledge point entities, and a test question type and test question difficulty of the test question text; and based on the illustration demand information and the illustration description information of each candidate illustration, determining a target illustration matched with the test question text from each candidate illustration. According to the method, the electronic equipment and the storage medium provided by the invention, the image-text matching efficiency is improved, and the reliability and accuracy of image-text matching are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of educational informatization, and in particular to a picture-text matching method, an electronic device and a storage medium. BACKGROUND

[0002] With the development of digitalization and intelligentization of educational resources, picture-text combination has become the mainstream form of educational content production.

[0003] Currently, to match appropriate illustrations for test questions, professional art personnel usually manually design illustrations, which has a long production cycle and high cost. Although there are also schemes for automatically matching illustrations, most of them are based on matching keywords of test questions. Since there is a lack of deep semantic understanding of test questions, there are often problems such as disconnection between the matched illustrations and the actual content of the test questions, and even conceptual errors.

[0004] Therefore, how to ensure the efficiency of picture-text matching while optimizing the accuracy and reliability of picture-text matching is still a problem to be solved in the field of educational informatization. SUMMARY

[0005] The present application provides a picture-text matching method, an electronic device and a storage medium to solve the defect that the efficiency and accuracy of picture-text matching are difficult to be considered in related technologies.

[0006] The present application provides a picture-text matching method, comprising: obtaining a test question text; performing knowledge point entity recognition on the test question text to obtain knowledge point entities in the test question text and entity relationships of the knowledge point entities, performing question type classification and difficulty classification on the test question text to obtain a test question type and a test question difficulty of the test question text; determining illustration requirement information of the test question text based on the knowledge point entities in the test question text and the entity relationships of the knowledge point entities, the test question type and the test question difficulty of the test question text; determining a target illustration matched with the test question text from each candidate illustration based on the illustration requirement information and illustration description information of each candidate illustration.

[0007] According to the picture-text matching method provided by the present application, the knowledge point entity recognition on the test question text is performed to obtain the knowledge point entities in the test question text and the entity relationships of the knowledge point entities, comprising: performing entity recognition on the test question text to obtain entities in the test question text; mapping the entities to a knowledge graph to obtain the knowledge point entities and graph nodes of the knowledge point entities in the knowledge graph; The graph node of the knowledge point entity in the knowledge graph is subjected to relation path reasoning to obtain an entity relation of the knowledge point entity.

[0008] According to the method for matching pictures and texts provided by the application, the entity recognition on the test question text comprises: The word feature of each word unit in the test question text is coded, and the word feature comprises a word vector, a position vector and a segmentation vector of the word unit. The context feature of each word unit in the test question text is determined based on the dependency relation between the word features of each word unit in the test question text. The entity recognition is performed based on the context feature of each word unit in the test question text.

[0009] According to the method for matching pictures and texts provided by the application, the test question type classification on the test question text comprises: The macroscopic test question type classification, the subject special classification and the investigation target classification are performed on the test question text.

[0010] According to the method for matching pictures and texts provided by the application, the illustration description information comprises at least one of a subject knowledge point, a visual feature and a teaching attribute of the candidate illustration, and the illustration description information is generated based on a large language model.

[0011] According to the method for matching pictures and texts provided by the application, the target illustration matched with the test question text is determined from the candidate illustrations based on the illustration requirement information and the illustration description information of each candidate illustration, which comprises: The target illustration matched with the test question text is determined from the candidate illustrations based on the similarity between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration.

[0012] According to the method for matching pictures and texts provided by the application, the target illustration matched with the test question text is determined from the candidate illustrations based on the similarity between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration, which comprises: The target illustration matched with the test question text is determined from the candidate illustrations based on the similarity between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration, and the similarity between the text semantics of the test question text and the image features of each candidate illustration.

[0013] According to the method for matching pictures and texts provided by the application, the method further comprises: In a case where the target illustration matching the test question text does not exist in the candidate illustrations, a visual prompt is determined based on the illustration requirement information, and a target illustration matching the test question text is generated based on the visual prompt.

[0014] According to the method for matching illustrations and texts provided in the application, the method further comprises: A layout strategy is determined based on a test question type of the test question text and a device type of a display device. The test question text and the target illustration are displayed based on the layout strategy.

[0015] The application further provides a device for matching illustrations and texts, comprising the following modules: The application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for matching illustrations and texts according to any one of the above when executing the program.

[0016] The application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the method for matching illustrations and texts according to any one of the above.

[0017] The application further provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the method for matching illustrations and texts according to any one of the above.

[0018] The method for matching illustrations and texts, the electronic device, and the storage medium provided in the application perform knowledge point entity recognition, prompting classification, and difficulty classification on the test question text, implement deep semantic analysis on the test question text, and thus determine the illustration requirement information of the test question text. Based on the illustration requirement information and the illustration description information of each candidate illustration, the target illustration matching the test question text is determined from each candidate illustration, so as to improve the efficiency of matching illustrations and texts while ensuring the reliability and accuracy of matching illustrations and texts. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the application or the related art, the following will briefly introduce the drawings needed to be used in the embodiments or the related art description. Obviously, the drawings in the following description are some embodiments of the application, and those skilled in the art can also obtain other drawings according to these drawings without any creative effort.

[0020] Figure 1 is one of the flowcharts of the method for matching illustrations and texts provided in the application.

[0021] Figure 2 is the flowchart of obtaining the test question text provided in the application.

[0022] Figure 3 is a flowchart of the process of determining the illustration requirement information provided by the present application.

[0023] Figure 4 is a flowchart of the process of matching the test question text and the candidate illustrations provided by the present application.

[0024] Figure 5 is a flowchart of the process of generating the target illustration provided by the present application.

[0025] Figure 6 is a flowchart of the process of typesetting provided by the present application.

[0026] Figure 7 is a flowchart of the process of the image-text matching method provided by the present application.

[0027] Figure 8 is a structural diagram of the image-text matching system provided by the present application.

[0028] Figure 9 is a structural diagram of the image-text matching device provided by the present application.

[0029] Figure 10 is a structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0031] Currently, in order to match the test questions with appropriate illustrations, professional art personnel are usually required to manually design the illustrations. This process not only consumes time and effort, and the production cycle of a single illustration can be as long as several hours, but also the labor cost increases exponentially with the expansion of the test question bank. More seriously, manual illustration matching is difficult to meet the real-time illustration matching requirements of the massive test questions of online education platforms, and becomes a bottleneck restricting the large-scale production of educational content.

[0032] In order to improve the efficiency of image-text matching, the related art proposes the following three matching schemes: One is a keyword-based matching scheme. In this scheme, TF-IDF (term frequency-inverse document frequency) or Word2Vec can be used to extract keywords from the test questions, match the keywords with pre-set labels in the illustration library resources, and then use a rule engine to combine text and images. This scheme has three main defects: in terms of semantic understanding, it lacks deep semantic understanding of test questions and cannot distinguish between discipline-specific meanings of words. The matched illustrations may not match the actual content of the test questions, or even have conceptual errors. For example, keyword matching may ignore the difference between "force" in physics and its everyday usage. For example, matching the keyword "triangle" may not be able to distinguish between different types of triangle applications, resulting in inappropriate visual examples. In terms of context understanding, it is difficult to distinguish the specific reference of the same keyword in different test questions. For example, "circle" in a math test may represent a geometric shape or a function image. In terms of adaptability, keywords cannot reflect the learning stage of the test itself, so they cannot adjust the complexity of the illustrations according to the learning stage.

[0033] This mechanical matching cannot understand the knowledge system and teaching intentions behind the questions, resulting in a mismatch between the illustrations and the content. For example, "speed" in a math application question may refer to average speed or instantaneous speed, but the system cannot distinguish between them. "Revolution" in a history question may correspond to different historical events, but the illustrations are often the same. This shallow matching seriously affects teaching effectiveness and learning experience.

[0034] Another is a template rule-based scheme: by establishing a subject classification template library (such as mathematics, physics, chemistry, etc.), pre-setting fixed illustration templates for each subject's test questions, and using regular expressions to extract test question parameters for filling. This scheme has obvious shortcomings: first, the maintenance cost of the template library is high, and each new question type requires manual creation of corresponding templates; second, flexibility is severely limited, and it cannot automatically adjust the illustrations according to the test question parameters; finally, it lacks innovation and cannot meet the illustration needs of cross-disciplinary comprehensive test questions.

[0035] There is also a machine learning-based solution: using traditional deep learning models such as LSTM (Long Short-Term Memory) or CNN (Convolutional Neural Network) to train labeled question-image pair data and draw on collaborative filtering recommendation ideas to achieve image-text matching. This type of solution also has significant drawbacks: first, it relies heavily on labeled data and requires a large number of manually labeled training samples; second, it faces the cold start problem and performs poorly when encountering new knowledge point questions; third, the model decision-making process lacks explainability, making it difficult to ensure the accuracy of educational content; and most importantly, the alignment between text and image feature spaces is not precise enough.

[0036] In summary, current image-text matching solutions mostly use rule engines or simple machine learning models, which have the following limitations: First, they rely on manually defined hard rules and are difficult to cover complex and varied educational scenarios. In educational scenarios, questions of different subjects and difficulty levels need to match images of different complexity and style. Current matching solutions lack flexible adaptation mechanisms and do not consider cognitive rules and subject characteristics, making it difficult to adjust image complexity or distinguish between visual needs for science and liberal arts based on learning stages.

[0037] For example, science questions require precise formula charts, while liberal arts questions require situational illustrations. Simple diagrams are needed for primary school students, while more professional analysis charts are needed for high school students. However, current image-text matching solutions either use uniform templates or require manual rule setting, making it difficult to automatically adjust based on subject characteristics and cognitive rules, resulting in overly simple or overly complex images that lose their teaching value.

[0038] Second, they lack understanding of the deep semantics and teaching intent of questions, limiting the teaching value of generated image-text combinations. Additionally, there is a natural modality gap between text-based questions and visual-based images, making it difficult for traditional keyword or template-based methods to establish accurate cross-modal associations, which leads to a semantic gap between questions and images. The lack of multi-modal alignment precision directly affects the effectiveness of image-text matching.

[0039] Third, the system lacks scalability and is difficult to adapt to emerging educational forms such as AR (Augmented Reality) / VR (Virtual Reality) teaching, adaptive learning, etc. Each new question type requires the creation of a corresponding template, or each new knowledge point question faces the cold start problem of the model, which severely restricts the expansion and application of image-text matching solutions, making it difficult to scale and iterate, and affecting the quality and intelligent development of educational products.

[0040] To solve the above problems, the embodiment of the present application provides a text-image matching method. Figure 1 is one of the flowcharts of the text-image matching method provided by the present application, as shown in the figure, the method comprises: Figure 1 Step 110, obtaining a test question text.

[0041] Here, the test question text is the text form of the test question which needs to be matched with the image, and the test question text can be the text of the question in the test question.

[0042] The test question text can be directly extracted from the test question library, or directly received from the user input test question text, or extracted from the original test question data in the test question library, or received from the user input original test question data, and then standardized and cleaned and structured for the original test question data, so as to convert the original test question data into structured test question text, which is not limited in the embodiment of the present application. It should be noted that in the embodiment of the present application, the original test question data can be input in the form of text, voice, handwriting, image and other input channels.

[0043] Among them, the standardized cleaning for the original test question data can include removing redundant spaces, random codes, irrelevant symbols, etc., correcting spelling errors, and unifying term expressions. Among them, the symbol filtering can be realized based on regular expressions, the unified term replacement can be realized through an education special term library, and the spelling error correction can be realized based on context awareness. For example, the line feed character “\n” can be removed, the unit symbol can be unified (for example, “Ω” can be converted to “ohm”), etc. In addition, if there are mathematical formulas and chemical equations and other special content, they can be marked separately and converted to Markdown format.

[0044] Through standardized cleaning, the non-standardization of the original test question data input by the user can be compatible, and the problem of information loss or misconversion can be avoided, thereby ensuring the robustness of the text-image matching.

[0045] In addition, for the case that the original test question data is multi-modal data such as voice, handwriting and image, the standardized cleaning for the original test question data can also include converting the multi-modal test question data to text mode through a modal conversion tool, for example, a domain optimized ASR (Automatic Speech Recognition) system based on Conformer model can be used to realize the conversion of voice modal data; for example, a handwriting recognition system integrating StrokeNet neural network can be used to realize the conversion of handwriting modal data; for example, MinerU open source tool can be used to realize the conversion of image modal data.

[0046] ​Subsequently, the standardized and cleaned test question data is structured, which can be structured text. The structured text can include original text, cleaned text, and modal type, etc. The structured text can be regarded as test question text for subsequent application.

[0047] In some embodiments, Figure 2 is a flowchart of obtaining test question text provided by the present application, as Figure 2 shown, the flow of obtaining test question text can be to obtain original test question data, determine the modal of the original test question data as a supportable data modal, perform modal conversion on the original test question data, perform standardized cleaning on the test question data converted to text modal, and finally generate structured output of the test question data, i.e. obtain structured test question text.

[0048] For example, the process of converting the original test question data into structured test question text can be formalized as: wherein, is the original test question data, is the structured test question text, is a multi-level standardized cleaning and structuring process from the original test question data to the structured test question text. Based on the process, high-quality structured output can be obtained, thereby ensuring that the image-text matching method is compatible with different input modes and remains robust, providing a reliable foundation for subsequent semantic analysis of test question text.

[0049] wherein the original test question data wherein is the first test question data, and N is the number of original test question data. Through multi-level processing, the standardization and structuring of test questions can be achieved, i.e. , is the first structured test question text.

[0050] For example, the original test question data can be: The resistance in a circuit is 10Ω, when a 20V voltage is applied across it (1) Calculate the current size in the circuit according to Ohm's law; (2) To generate a 4A current through the circuit, what voltage needs to be adjusted? (Need to list equations to solve).

[0051] The test question text obtained after standardized cleaning and structuring can be: {"original text": "A circuit with a resistance of 10 Ω, when a voltage of 20 V is applied across it\n(1) Calculate the current in the circuit according to Ohm's law;\n(2) To generate a current of 4 A through the circuit, what voltage needs to be adjusted to? (Need to list equations to solve)", "cleaned text": "A circuit with a resistance of 10 ohms, when a voltage of 20 volts is applied across it (1) Calculate the current in the circuit according to Ohm's law; (2) To generate a current of 4 amperes through the circuit, what voltage needs to be adjusted to? (Need to list equations to solve)", "modal type": "text"} {"original text": "A circuit with a resistance of 10 Ω, when a voltage of 20 V is applied across it\n(1) Calculate the current in the circuit according to Ohm's law;\n(2) To generate a current of 4 A through the circuit, what voltage needs to be adjusted to? (Need to list equations to solve)", "cleaned text": "A circuit with a resistance of 10 ohms, when a voltage of 20 volts is applied across it (1) Calculate the current in the circuit according to Ohm's law; (2) To generate a current of 4 amperes through the circuit, what voltage needs to be adjusted to? (Need to list equations to solve)", "modal type": "text"}

[0052] Step 120, knowledge point entity recognition is performed on the test text to obtain knowledge point entities in the test text and entity relationships of the knowledge point entities, and the test text is classified by question type and difficulty to obtain the test type and difficulty of the test text.

[0053] Specifically, after obtaining the test text, semantic analysis can be performed on the test text. Here, the semantic analysis of the test text can be divided into knowledge point entity recognition, question type classification, and difficulty classification.

[0054] Among them, knowledge point entity recognition is used to identify knowledge point entities in the test text, and to extract relationships of knowledge point entities in the test text. Knowledge point entity recognition can be realized by pre-training NLP (Natural Language Processing) model. The knowledge point entity here is the knowledge point reflected in the test text, such as mathematical formula, physical law, historical event, etc. These knowledge point entities can be associated with the definition in the knowledge graph. And for cross-disciplinary questions, multiple domain knowledge points can be marked at the same time to ensure the accuracy of subsequent illustrations. The knowledge point entities and entity relationships obtained in this way can be stored in a structured form, such as including knowledge point name, subject, and context relationship in the test text in the structured output.

[0055] Question type recognition is used to classify the question type of the test text in detail. In this classification process, rule engine and machine learning can be combined to distinguish the question type and examination target of the test text. By obtaining the test type of the test text, the theme matching of subsequent figure-text matching can be ensured, such as matching the schematic diagram for geometry questions and matching the operation flowchart for experiment questions.

[0056] The difficulty classification is used to evaluate the difficulty of the test question text. The evaluation of the difficulty of the test question text can be multi-dimensional, for example, the difficulty of the test question text can be evaluated based on the text features, logical complexity, reference education standards, etc. of the test question text. For example, for the test question text belonging to the basic question, the difficulty of the test question text can be marked as L1 (memory / understanding), and for the test question text belonging to the comprehensive question, the difficulty of the test question text can be marked as L3 (application / analysis), and the level of the difficulty of the test question text can directly affect the complexity of the illustration.

[0057] In step 130, based on the knowledge point entity in the test question text and the entity relationship of the knowledge point entity, the test question type and the test question difficulty of the test question text, the illustration requirement information of the test question text is determined.

[0058] Specifically, after obtaining the results of the knowledge point entity recognition, the test type classification and the difficulty classification of the test question text, the results of the knowledge point entity recognition, the test type classification and the difficulty classification are integrated to generate the illustration requirement information of the test question text.

[0059] Here, the illustration requirement information of the test question text can include parameters of the illustration needed to be matched for the test question text, for example, can include the content type, complexity and interactivity, etc. of the required illustration. The illustration requirement information obtained in this way can provide clear constraints for subsequent image-text matching, thereby ensuring the teaching adaptability of image-text.

[0060] In addition, in some embodiments, the illustration requirement information can not only include the parameters of the illustration needed to be matched for the test question text, but also include the results of the knowledge point entity recognition, the test type classification and the difficulty classification, that is, the illustration requirement information not only reflects the requirement of the test question text for the matched illustration, but also reflects the semantics of the test question text itself, so as to more accurately perform image-text matching. For example, the illustration requirement information can be structured JSON data including four-dimensional labels of knowledge points, test question types, test question difficulties and illustration requirements.

[0061] For example, after obtaining the results of the knowledge point entity recognition, the test type classification and the difficulty classification of the test question text, the illustration complexity parameter can be determined based on the difficulty of the test question in the above results, combined with the Bloom cognitive level and the requirement of the school stage. The illustration requirement information generated in this way is stored in the form of structured JSON data as a set of metadata, and the illustration requirement information can include: Test question ID: used to uniquely identify the test question, which can adopt the coding rule of "subject-number", and can realize test question tracking and internal indexing.

[0062] Macro test type: generally takes values such as multiple-choice questions, fill-in-the-blank questions, true-or-false questions, experiment questions or answer questions, etc.

[0063] Parameters: Records entity information extracted from the test question text, along with corresponding descriptions.

[0064] Core Knowledge Points: This section contains object information for multiple knowledge point entities. Each entity's object information may include fields such as concept, subject, semantic role, and related concepts. Specifically, the concept identifies the name of the knowledge point entity involved in the test text, ensuring that subsequent illustrations accurately match the teaching content; the subject precisely locates the subject and branch to which the knowledge point entity belongs, thus influencing the professionalism and presentation style of the illustrations; the semantic role defines the function of the knowledge point entity in the test text, thereby determining the priority and presentation method of the illustrations, with common values ​​including key points of examination, calculation tools, and background knowledge; and related concepts list the relevant concepts upon which the problem-solving depends, ensuring the completeness of the illustrations and the coherence of the teaching.

[0065] Interdisciplinary connections: The main subject areas involved in this question are used to trigger interdisciplinary visualization strategies.

[0066] Difficulty Levels: Four levels of difficulty can be used, specifically divided into L1-L4. The difficulty level determines the complexity and detail of the illustrations.

[0067] Cognitive levels: Based on Bloom's taxonomy of educational objectives, common values ​​include memorization, comprehension, application, analysis, evaluation, and creation.

[0068] For example, the structured output of the illustration requirement information for a test question text can be represented in the following form: { "Problem ID": "PHYS-002", "Macroeconomic Question Types": "Short Answer Questions" "Parameter":[{"Entity":"Resistance","Value":"10", "Unit": "Ohms","Role":"Known Conditions"},{"Entity":"Voltage","Value":"20","Unit": "Volts", "Role":"Known Conditions"}, {"Entity":"Current","Value":"To be determined","Unit": "Ampere", "Role":"Solution objective (1)"},{"Entity":"Current","Value":"4","Unit": "Ampere","Role":"Given objective (2)"}, {"Entity":"Voltage","Value":"To be determined","Unit": "Volts", "Role":"Solution objective (2)"},], "Core Knowledge Points": [ { Concept: Ohm's Law Subject: Physics - Electricity "semantic role": "focus of examination", "associated concepts": ["current", "voltage", "resistance"] }, { "concept": "algebraic equation", "discipline": "mathematics-algebra", "semantic role": "computational tool" } ], "interdisciplinary associations": ["physics", "mathematics"], "difficulty level": "L2", "cognitive level": "application" } In the embodiments of the present application, step 120 deeply analyzes the semantic content of the test question text through NLP technology, and step 130 determines the illustration requirement information of the test question text based on the semantic content of the test question text. The processes of step 120 and step 130 can be formalized as: wherein, is the structured test question text, is the structured illustration requirement information, is the execution process, and in the execution process of , the accurate extraction of knowledge point entities and the association with the knowledge graph are realized, the question type and the difficulty are identified, and the parameterized illustration requirement information is dynamically generated. The application of steps 120 and 130 breaks through the limitations of traditional keyword matching, and the illustration requirement information obtained from the depth semantics of the test question text improves the relevance of the text and graphics and supports continuous learning and updating, providing an intelligent analysis basis for education text matching.

[0069] Figure 3 is the process diagram of the determination of the illustration requirement information provided by the present application, as shown in Figure 3 , semantic analysis can be performed in parallel for the test question text. Specifically, entity recognition, knowledge graph query and relationship extraction can be performed for the test question text, thereby determining the knowledge point entities and their entity relationships of the test question text. In addition, the test question type and the test question difficulty of the test question text are determined by classifying the test question type and the test question difficulty of the test question text, and the above various determination of the illustration requirement information is combined.

[0070] Step 140, based on the illustration requirement information and the illustration description information of each candidate illustration, determines a target illustration matched with the test question text from the candidate illustrations.

[0071] Specifically, the candidate illustration, i.e., a pre-collected educational illustration, can be obtained by acquiring resources such as educational test questions and educational illustrations. The acquisition of the candidate illustration supports the form of textbook scanning digitization and open source resource grabbing. Moreover, based on the acquisition of the candidate illustration, the candidate illustration can be standardized.

[0072] In addition, the candidate illustration can be pre-labeled with illustration description information. The illustration description information can be understood as a semantic label of the candidate illustration. The illustration description information can describe multidimensional information such as subject knowledge points, visual features, and teaching attributes of the candidate illustration, so as to realize intelligent classification and accurate retrieval of the candidate illustration.

[0073] After obtaining the illustration requirement information of the test question text, the illustration requirement information can be matched with the illustration description information of each candidate illustration, so as to select a candidate illustration matched with the test question text from each candidate illustration, which is recorded as a target illustration. Further, the selection of the target illustration can be realized by calculating the similarity between the illustration requirement information and the illustration description information of each candidate illustration, that is, the candidate illustration with the highest similarity can be selected as the target illustration, so as to realize accurate image-text matching in the educational scenario, taking into account the retrieval efficiency and matching accuracy.

[0074] The process of selecting the target illustration of the test question text based on the illustration requirement information can be formalized as: Wherein, is the illustration requirement information of the test question text, is the target illustration of the test question text, is a process of realizing the matching between the test question text and the candidate illustration. Selecting the target illustration of the test question text based on the illustration requirement information can improve the efficiency and accuracy of image-text matching and reduce the material search time.

[0075] Moreover, the quality of the illustration library where the candidate illustration is located can be continuously optimized through a dynamic updating mechanism, so as to ensure that the target illustration obtained by matching meets the scientific requirements of teaching and adapts to the educational needs of different school stages and regions, providing strong visual support for personalized teaching.

[0076] In the method provided in the embodiment of the application, knowledge point entity recognition, prompt classification and difficulty classification are performed on the test question text to realize deep semantic analysis of the test question text, so as to determine the illustration requirement information of the test question text. Based on the illustration requirement information and the illustration description information of each candidate illustration, the target illustration matched with the test question text is determined from each candidate illustration, so as to improve the efficiency of image-text matching while ensuring the reliability and accuracy of image-text matching.

[0077] Based on the above embodiment, in step 120, the knowledge point entity recognition is performed on the test question text to obtain the knowledge point entity in the test question text and the entity relationship of the knowledge point entity, including: performing entity recognition on the test question text to obtain the entity in the test question text; mapping the entity to a knowledge graph to obtain the knowledge point entity and a graph node of the knowledge point entity in the knowledge graph; performing relationship path reasoning on the graph node of the knowledge point entity in the knowledge graph to obtain the entity relationship of the knowledge point entity.

[0078] Specifically, for a test question text, an entity in the test question text can be recognized by a BERT (Bidirectional Encoder Representations from Transformers) model or other models that can be used for command entity recognition, and the type of the entity and other information are labeled. The type of the entity and other information labeled here can include the type of the entity in the test question, such as a parameter quantity, or a law, or a method, and can also include the value, unit, and role of the entity in the test question, such as a known condition or a solution target.

[0079] For example, for a test question text, the following entities and entity information can be extracted: { "parameter quantity": [ {"entity": "resistance", "value": "10", "unit": "ohm", "role": "known condition"}, {"entity": "voltage", "value": "20", "unit": "volt", "role": "known condition"}, {"entity": "current", "value": "to be solved", "unit": "ampere", "role": "solution target (1)"}, {"entity": "current", "value": "4", "unit": "ampere", "role": "given target (2)"}, {"entity": "voltage", "value": "to be solved", "unit": "volt", "role": "solution target (2)"}, ], "law / method": ["Ohm's law", "algebraic equation"] }。

[0080] After extracting entities from the test question text, these entities can be matched against a knowledge graph. Here, the knowledge graph used for matching can be a knowledge graph from various disciplines. By matching entities with the knowledge graph, entities can be mapped into the knowledge graph; that is, the corresponding graph node for each entity can be found within the knowledge graph. It can be understood that for an entity with a matching graph node in the knowledge graph, that is, based on its existence and semantic correctness as a knowledge point, it can be considered a knowledge point entity. The knowledge point entities identified in this way can eliminate potential ambiguities in cross-disciplinary terminology.

[0081] For example, the text representation of the entity can first be converted into a 768-dimensional vector. Through trainable matrices Reduce the dimensionality of this vector to the knowledge graph embedding space: in, It is a text representation of entities reduced to the knowledge graph embedding space, with an output dimension of 256, consistent with the embedding dimension of graph nodes in the knowledge graph; ReLU is an activation function used to enhance sparsity and filter noise. It is a trainable matrix.

[0082] Based on this, we can calculate Cosine similarity between the embedding representations of the nodes in the knowledge graph and the embedded representations of the nodes: In the formula, for Embedded representation of the i-th graph node in the knowledge graph The cosine similarity between them. N is the total number of graph nodes in the knowledge graph.

[0083] The similarity between each entity and a node in the knowledge graph can be obtained, and the three nodes with the highest similarity are selected as knowledge point nodes. For example, in the physics-electricity knowledge graph, the entity Ohm's law has a similarity of sim=0.95; in the physics-fundamentals knowledge graph, the entity current has a similarity of sim=0.82; and in the mathematics-algebra knowledge graph, the entity linear equation has a similarity of sim=0.76.

[0084] After mapping entities to a knowledge graph, relational path reasoning can be performed on the graph nodes based on the knowledge point entities' positions within the knowledge graph and the connections between these nodes. This quantifies the logical relationships between knowledge point entities, allowing the extraction of entity relationships from the knowledge graph, thus obtaining the entity relationships of the knowledge point entities. Specifically, the entity relationships of knowledge point entities can be represented as entities in the test text that are associated with the knowledge point entities.

[0085] Specifically, relationship path reasoning can be performed on the graph nodes based on a pre-set relationship matrix, thereby obtaining a path probability between entities , which can be expressed as: wherein, is a sigmoid function, is a matrix set of pre-defined education-specific relationship types, for example, subject relationship can include: belonging to (physics-electricity), cross-disciplinary association, etc.; teaching logic relationship can include: prerequisite knowledge (current → resistance), deepening concept, etc.; cognitive relationship can include: L1→L2 difficulty progression, memory→application level promotion, etc. Each type of relationship corresponds to a trainable relationship matrix .

[0086] The relationship path between the graph nodes in which the entities are mapped in the knowledge graph can be inferred through multi-hop path search, and in this process, some educational constraints can be set to filter invalid paths, for example, “historical event→physical formula” can be filtered as an invalid path; for each path meeting the requirements, the best path can be confirmed by calculating the confidence score: wherein, is an education weight.

[0087] By fusing the relationship path search system of pedagogy theory, the teaching logic can be calculated, and the entity association path meeting the cognitive development law can be automatically identified, for example, the teaching chain of “Ohm's law→series circuit→resistance calculation” is accurately connected in the physics question, avoiding the inference result of out-of-class or logical break.

[0088] After obtaining the relationship path, the relationship between entities can be extracted therefrom. Specifically, the relationship extraction can be performed based on the start entity and the end entity in the relationship path.

[0089] Firstly, the logical association between knowledge points can be quantified by modeling the entity-relation joint probability, to ensure that the relationship extracted by the entity meets the subject logic. Specifically, the conditional probability distribution of the relationship can be calculated by the vector of the start entity extracted by BERT, in combination with the context representation , through a weight matrix : ​​Then, by leveraging the global structural constraints of the knowledge graph, ambiguities arising from local text extraction can be avoided, thus enabling the establishment of entity relationships. Specifically, this can be achieved by utilizing predefined entity embeddings within the knowledge graph. And relation matrix Calculate triples The reasonableness score is calculated, with lower scores indicating a higher likelihood of a valid relationship. This leads to the relationship between the starting and ending entities in the relationship path.

[0090] In addition, classification weight matrix can also be used. This allows for hierarchical classification and semantic role analysis of entities, thereby clarifying the functional positioning of knowledge point entities within the test text and guiding subsequent visualization strategies. Specifically, entities can be... Embedded vector Mapped to semantic role category c: For example, the semantic role of "Ohm's Law" is classified as "a key point of examination", while the semantic role of "algebraic equations" is classified as "a computational tool".

[0091] Therefore, after completing the knowledge point entity recognition for the test text, the following structured information can be obtained to reflect the knowledge point entities, entity relationships, and semantic roles of knowledge point recognition in the test text. Entity relationships can be represented as "associated concepts," reflecting the logical dependencies between entities in the test text: { "Core Knowledge Points": [ { Concept: Ohm's Law Subject: Physics "Semantic Role": "Key Exam Focus", Related concepts: ["Current", "Voltage", "Resistance"] }, { Concept: Algebraic equations Subject: Mathematics Semantic Role: "Computational Tool" } ] }

[0092] Based on any of the above embodiments, in step 120, the entity recognition of the test question text includes: Encode the word features of each word in the test text, wherein the word features include the word vector, position vector and segmentation vector of the word; determine a context feature of each token in the test question text based on a dependency relationship between token features of each token in the test question text; perform entity recognition based on the context feature of each token in the test question text.

[0093] Specifically, when performing entity extraction on the test question text, a segmented coding based on the test question structure and a hierarchical expression fusion strategy are adopted, aiming to more accurately capture the semantic information of knowledge points contained in each component of the test question text.

[0094] First, when performing vector representation on the test question text, each token in the test question text can be encoded respectively, thereby obtaining a token feature of each token. Here, the token feature of each token in the test question text can be represented as follows: wherein the token sequence of the test question text is composed of n tokens, wherein the token feature of the jth token is the sum of the word vector , the position vector and the segment vector . Wherein represents encoding the key words in the test question text, such as “function”, “equation”; reflects the serial number of the jth token in the sequence, thereby retaining the word order information, such as “derivative” needs to rely on the context before and after. reflects the serial number of the jth token in the paragraph in the test question text. By superimposing the above three types of vectors, the test question text is converted into a numerical vector that can be processed. Wherein, can be obtained by BERT encoding.

[0095] It can be understood that the token feature combining the word vector, the position vector and the segment vector can not only reflect the meaning of the token itself and the position in the sequence, but also reflect the paragraph level of the token in the test question text. Therefore, in the entity recognition based on the token feature, the semantic information of knowledge points contained in each component of the test question can be better captured from the perspective of the paragraph level. Since the word vector is obtained by BERT encoding, the hierarchical token feature of the combined segment vector can be denoted as Hie-BERT.

[0096] After obtaining the token feature of each token, the serialized test question text can be semantically encoded based thereon, i.e. the context feature of each token in the test question text is extracted. Here, the extraction of the context feature can be realized by a multi-layer Transformer encoder, and the process can be represented as: wherein, is the context feature of the th word in the test text. In the process of extracting the context features of each word in the test text, the dependency between words can be calculated using the multi-head attention mechanism: wherein, Q, V, K are obtained by linear transformation of the word sequence of the test text, is a scaling factor. For example, in the mathematical test question "find the derivative of the function ", the word "derivative" needs to pay attention to the previous "function" and " ", and the attention mechanism can automatically learn this long-distance dependency, ensuring that "derivative" is correctly classified as a mathematical term.

[0097] After completing the context feature extraction of each word, the entities in the test text can be decoded based on this, thereby structuring the extraction of knowledge points. The probability distribution of entity recognition based on the context feature of each word can be represented as: wherein, and are classification layer parameters, denotes the predicted entity recognition result.

[0098] In the embodiments of the present application, based on the context feature of each word, not only entity recognition can be realized, but also entity aggregation and discipline adaptation can be used, thereby converting educational test questions into computable knowledge point data.

[0099] Based on any of the above embodiments, in step 120, the test text is classified by type, including: macro-type classification, discipline-specific classification, and examination target classification.

[0100] Specifically, the test text is classified by type, which can be achieved from three dimensions of macro-type classification, discipline-specific classification, and examination target classification.

[0101] Among them, the macro-type classification refers to determining the macro-type category to which the test text belongs, such as multiple-choice questions, fill-in-the-blank questions, true-or-false questions, experimental questions, or answer questions, etc. For example, the BERT model can be combined with a rule engine to achieve accurate macro-type classification of test text through semantic understanding and structured features.

[0102] For example, the test text can be encoded by a question type classification BERT model, thereby obtaining the context vector of the test text re-calculate the probability distribution of macro type : wherein, and are parameters of the type classification BERT model, denotes the probability that the test text belongs to the macro type.

[0103] Discipline-specific classification refers to determining the discipline to which the test text belongs and the type of discipline-specific category. For example, the test text "find the area of a circle" corresponds to the discipline-specific classification result: mathematics-geometry, the test text "calculate the electric field strength" corresponds to the discipline-specific classification result: physics-electromagnetism, and the test text "balance chemical equation" corresponds to the discipline-specific classification result: chemistry-chemical reaction.

[0104] For example, the discipline-specific classification BERT model can be used to classify the test text by discipline-specific category, and output the discipline-specific probability distribution : wherein, and are parameters of the discipline-specific classification BERT model; denotes the rule feature vector; denotes the probability that the test text belongs to the discipline-specific category. Examination target classification refers to determining the examination target of the test text, which can be memory, calculation, comprehensive application, etc.

[0105] For example, the multi-layer perceptron can be used to classify the examination target, and output the probability distribution of the examination target

[0106] : wherein, is a multi-layer perceptron prediction model, and the semantic vector of the test question extracted by the comprehensive BERT model , mathematical expression complexity , test text length and other information are used as input. The output , which represents the probability that the test text belongs to the examination target.

[0107] Based on any of the above embodiments, in step 120, the difficulty classification of the test text is performed, including: ​​​Based on the statistical features and semantic features of the test question text, the test question text is classified in difficulty.

[0108] The statistical features of the test question text are obtained by TF-IDF extraction of the test question text, and can be specifically represented as In addition, the semantic features of the test question text can be obtained based on the BERT model, and the semantic features can be represented as .

[0109] The statistical features and semantic features of the test question text can be fused as the feature vector of the test question text , for example, which can be represented as: The feature vector of the test question text obtained by fusion, the application of TF-IDF improves the weight on complex terms, and the application of BERT output activation value makes the distribution of the feature vector more close to the high-order cognitive dimension.

[0110] Subsequently, the feature vector of the test question text can be applied to test question difficulty classification: wherein, and are parameters of the difficulty classification model, is the level of test question difficulty. Here, the test question difficulty can reflect the logical complexity of the test question.

[0111] The difficulty level of the test question difficulty classification here can correspond to the Bloom cognitive level classification. The difficulty level can be divided into L1 to L4, and the Bloom cognitive level corresponding to the difficulty level can be obtained based on the following table.

[0112] Based on any of the above embodiments, the method further comprises: Collecting multi-source data to construct an illustration library, wherein the illustration library is used to store candidate illustrations.

[0113] Specifically, the illustration library can be a multi-modal educational illustration library, which can include static images, dynamic charts and other forms. The data source of the illustration library can be multi-source, for example, it can include open source educational resources, professional textbook illustration digitization processing, and teacher user uploaded content. All illustrations are standardized processed to ensure the uniformity of resolution, color mode and file format, and hierarchical storage index is established. At the same time, copyright compliance review is implemented, and commercial restricted resources are replaced or authorized to obtain, forming a safe and usable educational illustration library.

[0114] Further, the process of collecting data to construct the illustration library can be represented as: wherein, represents the combination of the test questions and images, represents the data source space, that is, the candidate illustrations can be collected from the teaching material scans, open source resources and user uploaded content respectively. is the original data space.

[0115] In the education scene, the teaching material scans Adopt the scanning resolution constraint to ensure the conversion accuracy of the illustrations; the user uploaded content ensures the availability of the images through the quality filter; the open source resources based on the Scrapy framework, a distributed crawler is built to implement the targeted crawling strategy for the education sites.

[0116] And for the collected images, data standardization processing can be performed: first, define the resampling function of the image to convert it into a standardized resolution output that meets the needs of the education scene, and its mathematical expression can be defined as: wherein, is the target resolution calculated dynamically, which is calculated in combination with the absolute threshold 300dpi and the relative size constraint to ensure the clarity requirement in the teaching scene. , is the height and width of the image.

[0117] Then, through a nonlinear mapping, the color data of the image is converted to the standard sRGB space to ensure the accurate reproduction of key visual elements in the teaching scene, and its mathematical expression can be defined as: wherein, is the result of the color conversion for the image .

[0118] Finally, define the format conversion function to realize the unified conversion of the illustration format, wherein is the education special format strategy: Thus, after image collection from multiple data sources, data standardization processing can be performed on the collected images, and the images after data standardization processing are stored in the illustration library as candidate illustrations.

[0119] Based on any of the above embodiments, the illustration description information includes at least one of the subject knowledge points, visual features, and teaching attributes of the candidate illustration, and the illustration description information is generated based on a large language model.

[0120] Specifically, for the collected candidate illustrations, prompt words can be constructed to guide a large language model to generate illustration description information for the candidate illustrations. Furthermore, prompt words can be constructed to guide the large language model to generate illustration description information for the candidate illustrations from at least one aspect of the candidate illustrations' subject knowledge points, visual features, and pedagogical attributes.

[0121] Among them, subject knowledge points refer to the subject to which the candidate illustration is adapted, as well as the knowledge points under that subject. Visual features can describe the core objects in the candidate illustration, the layout of the candidate illustration, color and other information. Teaching attributes can include information such as the difficulty level of the test questions and the grade level to which the candidate illustration is adapted.

[0122] For example, you can set the following prompts to retrieve illustration description information for candidate illustrations: You are an intelligent annotation system for educational resources. Please perform multimodal analysis on the given illustrations and complete the following tasks: [User Issues]: 'Educational Illustrations' [Your task]: 1. Visual feature extraction: Describe the core objects, layout, colors, etc. in the image.

[0123] 2. Text information recognition: Extract text, formulas, or annotations from the image.

[0124] 3. Teaching tag generation: Subject categories: Mathematics / Physics / Chemistry / Biology / History, etc. Key concepts: Specific concepts (such as 'Pythagorean theorem' and 'redox reaction') Difficulty Levels: L1 (Basic Memory) - L3 (Advanced Application) Educational stages (primary / middle / high school / university) 4. Semantic retrieval optimization: Generate 3-5 natural language query examples (e.g., 'Experiment diagram of Newton's first law in high school physics').

[0125] 5. Confidence Assessment and Recommendations: Fields with low confidence levels (<80%) should be manually reviewed. [Output format example]: { "visual_analysis": ["...", "..."], / / Visual feature description "text_analysis": ["...", "..."], / / Text / formula in the image "educational_tags": { "subject": "physics", / / subject "topic": "Newton's Second Law", / / Key Points "difficulty": "L2", / / Difficulty (L1-L3) "grade_level": "high school", / / grade level "interactivity": "interactive" / / Interactivity }, "search_queries": [ / / Search query examples] "Dynamic Demonstration of F=ma Experiment in High School Physics" Interactive simulation diagram of Newton's Second Law ], "confidence_scores": { / / Label confidence score (0-1) "subject": 0.95, "topic": 0.88, "difficulty": 0.75 }, "expert_review_suggestions": [ / / Fields requiring expert review] "difficulty" ] }

[0126] Based on similar cue words, large language models can infer from the input candidate illustrations and cue words, and then output illustration description information for the candidate illustrations. This process can be formalized as follows: in, Construct a constructor for the prompt words. For candidate illustrations, This refers to the illustration description information output by a large-scale language model, which can include subject knowledge points, visual features, teaching attributes, etc. Furthermore, the large-scale language model can also output confidence assessments and suggestions for the above illustration description information, to help determine whether to subsequently submit the illustration description information to experts for verification, thereby ensuring the accuracy of the illustration description information.

[0127] Alternatively, the process of annotating the illustration description information of the candidate illustration can be semi-automatic, for example, the associated description text of the candidate illustration can be obtained first, and the associated description text is parsed by an NLP model, thereby obtaining the basic label of the candidate illustration, which can include disciplines, knowledge points, etc.; in addition, the visual label of the candidate illustration can be extracted by a CV (Computer Vision) model, which can include color composition, composition features, etc. Subsequently, the teaching attribute label can also be manually annotated by an education expert, which can include applicable school stage, interaction demand, etc. The illustration description information of the candidate illustration is a kind of label, and the label system adopts a hierarchical ontology design, supporting flexible retrieval from coarse granularity to fine granularity.

[0128] Based on any of the above embodiments, in step 140, the target illustration matching the test question text is determined from the candidate illustrations based on the illustration requirement information and the illustration description information of each candidate illustration. The target illustration matching the test question text is determined from the candidate illustrations based on the similarity between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration.

[0129] Specifically, to realize the matching between the test question text and the candidate illustration, semantic feature extraction can be performed on the illustration requirement information and the illustration description information respectively. In the embodiment of the application, the semantic feature of the illustration requirement information is denoted as requirement semantics, and the semantic feature of the illustration description information is denoted as illustration semantics.

[0130] The requirement semantics and the illustration semantics can be obtained by a pre-trained semantic extraction model, for example, the BGE-small model can be used to extract the illustration description information of each candidate illustration in the illustration library and encode it to map it to a vector space of dimension The vectorization process of extracting the illustration semantics based on the illustration description information can be formalized as: wherein, is the dense semantic vector of the illustration description information of the candidate illustration, that is, the illustration semantics of the candidate illustration.

[0131] The embedding model similar to BGE-small is obtained by training on a large amount of corpus, and such a model can effectively capture the deep semantic features of the text, so that the illustration description information with similar semantics is closer to each other in the generated vector space of the illustration semantics. For example, the semantic similarity score between the illustration description information and of two candidate illustrations can be calculated by their illustration semantics and cosine similarity To measure: in, The L2 norm (Euclidean length) of a vector. The value of is between -1 and 1, and for non-negative vectors it is usually between 0 and 1. The closer the value is to 1, the more semantically similar it is. Due to the superiority of the BGE-small model, even and The length differences between them are significant, and this similarity score can reliably reflect the true semantic closeness.

[0132] Furthermore, demand semantics can be obtained based on an NLP encoder or a model of this type.

[0133] Based on this, by quantitatively analyzing the similarity between the semantic requirements of illustrations in the test question text and the semantics of illustration descriptions in each candidate illustration, high-precision automatic matching of test question text and illustrations can be achieved. For example, the candidate illustration with the highest similarity can be used as the target illustration to match the test question text.

[0134] Furthermore, in the process of retrieving target illustrations that match the test question text from the illustration library based on similarity, a hybrid retrieval strategy can be introduced, combining sparse retrieval for rapid initial screening and dense retrieval for fine ranking, thus balancing efficiency and accuracy.

[0135] Based on any of the above embodiments, step 140, which involves determining the target illustration matching the test question text from the candidate illustrations based on the similarity between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration, includes: Based on the similarity between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration, and the similarity between the text semantics of the test question text and the image features of each candidate illustration, a target illustration matching the test question text is determined from the candidate illustrations.

[0136] Specifically, when matching test text with candidate illustrations, not only can semantic matching be performed using illustration requirement information from the test text and illustration description information from the candidate illustrations, but cross-modal matching can also be performed directly using the test text and candidate illustrations.

[0137] Here, the cross-modal matching between the test text and the candidate illustrations can be implemented based on a cross-modal attention mechanism. First, semantic feature extraction can be performed on the test text and the candidate illustrations respectively. Here, the semantic feature of the test text is denoted as text semantics, and the semantic feature of the candidate illustrations is denoted as image feature. For example, the text semantics can be obtained by encoding the test text using a BERT model , and the image feature of the candidate illustrations can be extracted using a CNN model , which can be represented by the following formula: wherein, ( ) is a function for extracting text semantics, CNN( ) is a function for extracting image features, and the BERT encoding of the test text and the CNN encoding of the candidate illustrations can be mapped to a shared semantic space through a CLIP (Contrastive Language-Image Pre-training) model.

[0138] On this basis, similarity calculation can be performed on the text semantics and the image features in the shared semantic space.

[0139] Alternatively, before performing the similarity calculation, a knowledge graph revised attention rectification module (KARM) can be introduced. The KARM can query an educational knowledge graph, identify and strengthen cross-modal interactions related to core educational concepts. Specifically, key educational entities / concepts can be identified from the text semantics , and a subgraph of related concepts, attributes, and relationships can be obtained by querying the educational knowledge graph . The information of the subgraph G can be encoded into a knowledge context vector. The knowledge context vector can be used as a guide signal to participate in the calculation of the cross-modal attention between the text semantics and the image features. This process can be briefly formulated as: wherein, and are the text semantics and the image features that have been revised by the knowledge-guided attention mechanism and have more prominent educational semantics.

[0140] Subsequently, the similarity between the text semantics of the test text and the image features of the candidate illustrations can be calculated to match the test text with appropriate candidate illustrations. The similarity calculation can be represented by the following formula: wherein, is an education-enhanced term weight. is a preset parameter.

[0141] On this basis, the similarity between the demand semantics and the illustration semantics of the illustration description information of each candidate illustration, and the similarity between the text semantics and the image features of each candidate illustration, can be combined respectively, and the target illustration is determined from each candidate illustration. For example, the two types of similarity can be weighted and summed, and the target illustration is determined from each candidate illustration based on the similarity after the weighted sum.

[0142] In the embodiments of the present application, matching is performed based on the attention mechanism, which can ensure that the theme and details of the target illustration are highly consistent with the test question text. Specifically, in this process, a cross-modal attention mechanism of the Transformer architecture can be used to achieve fine-grained alignment of the test question-illustration, wherein higher attention weights can be assigned to key terms on the text side, and the semantic core area of the illustration is focused on the visual side. For composite questions, different modules such as experimental device diagrams, data recording tables, and result curve diagrams are matched respectively by multi-head attention, and the final matching scores are weighted and summarized.

[0143] Figure 4 is a matching process diagram of the test question text and the candidate illustration provided by the present application, as shown in Figure 4 The multi-modal data can be pre-collected, and data standardization processing is performed on the multi-modal data, thereby realizing collection of the candidate illustration and construction of the illustration library. In addition, the illustration description information of the candidate illustration is labeled. After obtaining the test question text, the illustration demand information of the test question text and the illustration description information of the candidate illustration are calculated for similarity, and the test question text and the candidate illustration are matched based on the cross-module attention mechanism. The candidate illustration matched with the test question text from the illustration library is determined as the target illustration.

[0144] Based on any of the above embodiments, after step 140, the method further comprises: In the case that there is no target illustration matched with the test question text in the candidate illustrations, a visualization prompt is determined based on the illustration demand information, and a target illustration matched with the test question text is generated based on the visualization prompt.

[0145] Specifically, for the case that the target illustration cannot be matched from the candidate illustrations, a generative artificial intelligence model can be used to generate a target illustration matched with the test question text. This process can be formalized as: wherein, is the illustration demand information of the test question text, A target illustration matching the test question text. Generate generative AI teaching illustrations for the execution process. The process can be based on a generative artificial intelligence model, and through intelligent generation and dynamic optimization mechanism, solve the core pain points of incomplete coverage of illustration resources and lagging update in traditional education content production.

[0146] Figure 5 The generation process of the target illustration provided by the present application is shown in Figure 5 The process can be divided into demand analysis and triggering, semantic-visual parameter conversion, generation model adaptation and optimization, and parameter dynamic adjustment stages.

[0147] Among them, demand analysis and triggering aim to trigger the automatic generation process of the target illustration. It is determined that there is no target illustration matching the test question text in the candidate illustration, which can be used as a condition for triggering the generation of the target illustration. It is determined that there is no target illustration matching the test question text in the candidate illustration, for example, the similarity between the test question text and each candidate illustration can be lower than a preset threshold.

[0148] For example, for the test question text to be matched, the similarity score between the test question text and each candidate illustration in the illustration library can be calculated. If the similarity score is lower than the preset threshold , it is determined that there is no matching candidate illustration in the illustration library, and the automatic generation process is triggered. The similarity threshold triggering condition can be expressed as: Semantic-visual parameter conversion aims to convert the structured test question text illustration requirement information into visual instructions executable by the artificial intelligence model, thereby realizing lossless conversion from abstract knowledge points to accurate teaching illustrations. Specifically, the JSON format test question text illustration requirement information can be converted into the parameterized prompt words of the generative artificial intelligence model, that is, the visualized prompt is obtained.

[0149] In specific operation, the JSON format test question text illustration requirement information can be parsed to extract key teaching elements such as physical formulas, mathematical parameters, and interactive components, and then map these key teaching elements to visual generation control parameters through multi-modal alignment technology and combine them in the visualized prompt.

[0150] Generation model adaptation and optimization aim to adapt and optimize the generative artificial intelligence model, so that the model can output illustrations that match the test question text according to the input visualized prompt.

[0151] For example, the generative artificial intelligence model can be an SD (Stable Diffusion) model that can generate precise and standardized illustrations according to input visual prompts. Specifically, a pre-trained SD model can be LoRA fine-tuned based on educational data (test text-illustration pairs) to adapt to academic illustration styles while retaining original generation capabilities. Additionally, the output style can be controlled by adding special trigger words. The LoRA fine-tuning parameter settings are as follows: train_config: pretrained_model: "stabilityai / stable-diffusion-2-1-base" dataset: "edu_dataset_v1" lora_rank: 64 batch_size: 8 learning_rate: 1e-4 text_encoder_lr: 5e-5 steps: 5000 trigger_word: "edu_diagram" By fine-tuning the SD model, the SD model can dynamically associate numerical variables in the test text with image elements in the illustration. For the SD model, ControlNet constraints can be combined to generate structures such as edge detection to ensure correct graph proportions. In addition, a text analysis module can be used to extract variable values and embed them into the visual prompt. Variable-visual binding can be expressed as: where, is the generated image; is the SD model generation function; is the visual prompt; is the Canny edge map extracted by ControlNet; is random noise; is the edge detection operator; M is the expected edge template; τ is the tolerance threshold.

[0152] On this basis, the CLIP model can be used to evaluate the semantic consistency of the generated target illustration and the test text, and based on the evaluation results of semantic consistency, the template of the visual prompt and the fine-tuning strategy of the SD model are iteratively optimized, thereby ensuring that the generated target illustration not only conforms to educational standards but also responds flexibly to variable changes. The index semantic alignment loss of iterative optimization can be expressed as: wherein, is an image encoder for CLIP; is a text encoder for CLIP; is a cosine similarity.

[0153] The end-to-end inference process based on the generative artificial intelligence model realizes the automatic generation from the test question text to the adapted target illustration, which can first extract key variable parameters from the illustration requirement information of the test question text, and inject these variable parameters into the structured prompt template, while combining the style description words of the educational illustration to form a complete generation instruction, that is, to obtain the visual prompt. In the diffusion generation stage, the generative artificial intelligence model not only generates content according to the visual prompt, but also is subject to the geometric constraints of ControlNet to ensure the accuracy of the graph structure of the generated target illustration. The whole process is supervised by the semantic alignment of the CLIP model, ensuring that the output image not only conforms to the mathematical relationship of the test question expression, but also maintains the normativity of the educational illustration, and finally generates a variable-labeled target illustration with accurate proportions as a teaching diagram. The inference process can be expressed as: wherein, is an SD original diffusion loss; is a controlNet constraint loss; is a final output image.

[0154] Parameter dynamic adjustment refers to the precise association of numerical variables in the test question text with visual elements through intelligent algorithms, thereby ensuring that the generated target illustration achieves the optimal scientific rigor and teaching applicability.

[0155] The parameter dynamic adjustment can be realized based on a multi-modal conditional control system. First, the key parameters (such as geometric dimensions, physical quantities, etc.) in the test question text can be parsed and dynamically mapped to multiple dimensions of the generative artificial intelligence model: the scale ratio in the target illustration is accurately controlled through the ControlNet (such as strictly matching the radius of a circle with the labeled number), and the generated style is fine-tuned using the LoRA weight adjuster (such as using different line widths and labeling densities for circles with radii of 5 cm and 10 cm), while the CLIP semantic guidance is used to ensure reasonable layout of variable labels (such as avoiding text from blocking key structures). In the multi-modal conditional system, a teaching knowledge rule library can be built in, thereby automatically optimizing the presentation form of the illustration (such as preferentially using the unit circle specification for trigonometric function illustrations), and performing multi-level checking on the generated target illustration (including numerical accuracy checking, scale verification, and teaching specification evaluation), thereby ensuring that the final output is a target illustration that meets both scientific principles and teaching demonstration requirements. The parameter dynamic adjustment can adapt to the variable generalization needs in the education scene, so that the same type of question can generate visual results that meet the discipline standards under different parameters.

[0156] In addition, when new knowledge points or special question parameter combinations are detected in the test library, professional illustrations that meet teaching specifications can also be automatically generated. These illustrations can be mathematical function graphs with accurate variable labels or chemical experiment flowcharts with step-by-step demonstrations, and the element layout and interaction design can be dynamically adjusted based on the test question requirements.

[0157] In the embodiments of the present application, the illustrations generated by the artificial intelligence model can be intelligently optimized and adjusted according to the specific parameters of the test questions: numerical precision correspondence is achieved through variable replacement, the time sequence relationship of experimental steps is presented through process decomposition, and the teaching effect is improved by adding interactive elements. The adjustment process uses three mechanisms: parameterized template modification, local condition regeneration, and multi-scheme optimization, for example, for the inclined plane slider problem, different inclination comparison illustrations can be automatically generated and associated with physical formulas. All optimization operations are recorded and analyzed to form a closed-loop learning system for continuous improvement, ensuring that the illustrations achieve the optimal state in terms of scientificity and teaching applicability.

[0158] Based on any of the above embodiments, after step 140, the method further comprises: determining a layout strategy based on the type of the test question text and the type of the display device; displaying the test question text and the target illustration based on the layout strategy.

[0159] Specifically, after the text and image matching is completed, the text and image of the matched test question can be laid out. In the embodiment of the present application, the text and image presentation mode can be dynamically optimized based on the test question type of the test question text, the image meaning requirement and the terminal device characteristics, thereby ensuring the logical coherence of the teaching content and improving the visual communication efficiency. The process can be formalized as: wherein is the structured test question text, is the target image of the test question text, is the test question data mixed with text and image, is the execution flow for implementing the mixed text and image layout. In the mixed text and image layout, the professional requirements of teachers for the beautiful layout can be met, the reading line of students can be optimized, the usability and dissemination effect of the digital teaching resources can be significantly improved, the visual presentation of complex knowledge points can be more in line with the cognitive rules, and finally the maximization of the teaching effect can be realized.

[0160] Specifically, Figure 6 is the flowchart of the text and image layout provided by the present application, as shown in Figure 6 The flowchart can be divided into intelligent layout matching, dynamic content optimization and multi-format output and compatibility.

[0161] The intelligent layout matching is used for intelligently adapting the text and image layout based on the test question type and the device type of the display device, i.e. obtaining the mixed text and image layout strategy.

[0162] The text and image presentation mode can be automatically optimized according to the test question type and the device type of the display device. For example, the optimal layout strategy can be automatically selected by a two-dimensional matrix of question type-device The layout mapping function can be defined as follows: wherein T is the test question type set, D is the device type, represents the layout strategy.

[0163] In addition, a multi-objective optimization algorithm can be used to balance the content density and the teaching effectiveness, for example, a mathematical proof question is automatically allocated 50% of the page width to the derivative image (such as f'(x) curve), and the formula font size is dynamically adjusted to ensure that the core content is not affected when the tablet is switched between horizontal and vertical screens.

[0164] Further, the complex content flow can be processed by a conflict detection engine. When a long formula in the test question text is detected to overlap with a chart, horizontal scrolling or page display can be automatically triggered. The content flow rearrangement strategy can be defined as follows: wherein, is a content block, which can be text, image, formula; ( ) is an overlap detection function.

[0165] Dynamic content optimization refers to intelligent optimization of display for special content such as long formulas and complex charts, so as to improve the readability of teaching content and user experience.

[0166] wherein, for a long formula, an intelligent line folding algorithm can be used to dynamically analyze the formula structure, and the long formula can be automatically split under the premise of ensuring the integrity of the mathematical semantics. Specifically, when it is detected that the length of the formula exceeds a threshold, the line can be broken at the relational operator or high-order operator, and an alignment symbol can be automatically added, so as to ensure the readability of the split formula.

[0167] For a complex chart, a dynamic rendering technology based on visual saliency can be used to identify the core teaching elements in the chart through a convolutional neural network, and a multi-level zoom view can be automatically generated, so as to realize adaptive rendering for complex charts.

[0168] Multi-format output and compatibility refers to supporting multiple formats such as PDF, HTML5, EPUB, etc., to ensure cross-platform and multi-browser display compatibility and consistency. That is, based on the typesetting strategy, the optimized test question text and target drawing can be displayed. And through the intelligent document conversion engine, the mixed image education content including the test question text and the target drawing can be output adaptively on multiple platforms.

[0169] For example, format conversion can be performed based on the conversion framework constructed by Apache FOP and Pandoc, and automatic generation of PDF, HTML5 and EPUB3 can be supported. Among them, the PDF output uses CMYK color gamut and 300 dpi resolution; HTML5 integrates MathJax formula rendering and SVG vector drawing, and supports touch interaction; EPUB realizes semantic label classification. Support responsive design to adapt to the display needs of multiple terminals (PC, tablet, mobile phone).

[0170] In the embodiments of the present application, the layout of text and images can be automatically optimized according to the type of the question and the display device. For example, the drawing of a multiple-choice question is usually inline with the options, and the drawing of a solution question is placed at the key step of solving the problem, etc. The typesetting engine uses responsive line folding or horizontal scrolling for special content. The final output supports multiple formats and meets the accessibility standards.

[0171] Based on any of the above embodiments, Figure 7 is a flowchart of the image-text matching method provided by the present application, as shown in Figure 7 the image-text matching method can be expressed as: in, , , These consist of the original test data (without illustrations), the intermediate processing steps, and the final output of the text and image data. Based on... Figure 2 The illustrated process, the image-text matching method, may include the following steps: First, the raw test question data without illustrations is input into the test question input construction module. Within this module, the raw test question data undergoes standardization preprocessing and structuring, resulting in structured test question text that provides high-quality input for subsequent semantic analysis.

[0172] Secondly, the test question text is input into the test question semantic analysis module. In the test question semantic analysis module, the content of the test question text can be deeply analyzed through natural language processing technology. Then, based on the knowledge point entities and their entity relationships, test question type and test question difficulty in the test question text, the illustration requirements of the test question text are determined.

[0173] Next, the illustration requirement information of the test question text is input into the candidate illustration matching module. In the candidate illustration matching module, the illustration requirement information of the test question text can be matched with the illustration description information of each candidate illustration, thereby obtaining the target illustration that matches the test question text.

[0174] For cases where the target illustration is not among the candidate illustrations, the illustration requirement information from the question text can be input into the dynamic illustration generation module. In this module, the illustration requirement information is converted into visual prompts, and a generative artificial intelligence model is invoked based on these prompts to generate the target illustration.

[0175] After obtaining the target illustration, the test question text and the target illustration can be combined. Figure 1 It also includes an adaptive text and image layout module. This module automatically optimizes the text and image layout based on the question type and the display device type, and supports multi-format output, thus balancing aesthetics and functionality to ensure readability and teaching effectiveness across different display devices.

[0176] Figure 8 This is a schematic diagram of the image-text matching system provided by the present invention. Figure 8 Each module shown is applied to Figure 7The shown figure-text matching method. Among them, the test question input construction module supports receiving manually inputted multi-modal test question data, and can standardize and clean and structure the inputted test question data. The test question semantic analysis module can realize semantic analysis on the test question text through knowledge point extraction, question type classification and difficulty classification, and determine the illustration demand information based on the results of the semantic analysis. In the process of knowledge point extraction, knowledge point entity extraction can be carried out based on an entity extraction model, and the entity relationship of the knowledge point entity can be obtained in combination with a knowledge graph. The candidate illustration matching module can construct an illustration library, and label the corresponding illustration description information for each candidate illustration in the illustration library, and perform similarity calculation on the illustration demand information of the received test question text and the labeled illustration description information of each candidate illustration, and can perform matching based on an attention mechanism on the text semantics of the test question text and the image features of the candidate illustration. The dynamic illustration generation module can call a generative artificial intelligence (AI) model to dynamically generate a target illustration of the test question text, and the content in the target illustration can be dynamically adjusted based on the parameter settings in the test question text. The adaptive figure-text layout module can automatically adjust the layout of the target illustration according to the test question type of the test question text, including the size and position of the target illustration, and can adapt to the display of multiple display devices.

[0177] The figure-text matching method provided by the embodiment of the present application realizes efficient automation of educational figure-text synthesis, can greatly reduce the time of manual picture matching, and improves the resource production efficiency. In terms of efficiency, the traditional picture matching process can be speeded up by more than 10 times, the single question picture matching time is controlled within 30 seconds, the human cost is reduced by more than 80%, and the large-scale application demand of millions of question banks is met. The method performs semantic analysis based on deep learning, breaks through the limitation of traditional keyword matching, and ensures that the illustrations are highly related to the content of the questions. The method is innovatively optimized for teaching scenarios, adopts a personalized adaptation mechanism, can dynamically adjust the illustration style and complexity according to the subject characteristics and question difficulty, has strong expansibility, can accurately retrieve and match illustrations from a preset illustration library, and can also dynamically create new illustrations that meet the requirements through generative AI, perfectly adapts to diversified education scenarios such as textbook writing, online question banks, AR / VR teaching, and realizes intelligent, large-scale and personalized production of education content. In addition, through the innovative incremental learning mechanism, the system can continuously absorb new teaching content and illustration styles, always maintain a knowledge point coverage rate of more than 95%, and provide an efficient, accurate and economical figure-text synthesis solution for scenarios such as textbook writing, online education and intelligent tutoring.

[0178] The figure-text matching device provided by the present application is described below. The figure-text matching device described below can be correspondingly referred to the figure-text matching method described above.

[0179] Figure 9 is a structural schematic diagram of the figure-text matching device provided by the present application. AsFigure 9 The apparatus includes: The acquisition unit 910 is configured to acquire a test question text. The semantic analysis unit 920 is configured to perform knowledge point entity recognition on the test question text to obtain knowledge point entities in the test question text and entity relationships of the knowledge point entities, and perform question type classification and difficulty classification on the test question text to obtain a test question type and a test question difficulty of the test question text. The requirement determination unit 930 is configured to determine an illustration requirement information of the test question text based on the knowledge point entities in the test question text and the entity relationships of the knowledge point entities, the test question type and the test question difficulty of the test question text. The requirement matching unit 940 is configured to determine a target illustration that matches the test question text from the candidate illustrations based on the illustration requirement information and illustration description information of each candidate illustration.

[0180] Based on the above embodiments, the semantic analysis unit is specifically configured to: Perform entity recognition on the test question text to obtain entities in the test question text. Map the entities to a knowledge graph to obtain knowledge point entities and graph nodes of the knowledge point entities in the knowledge graph. Perform relationship path reasoning on the graph nodes of the knowledge point entities in the knowledge graph to obtain entity relationships of the knowledge point entities.

[0181] Based on any of the above embodiments, the semantic analysis unit is specifically configured to: Encode token features of each token in the test question text, the token features including word vectors, position vectors and segmentation vectors of the tokens. Determine context features of each token in the test question text based on dependency relationships between the token features of each token in the test question text. Perform entity recognition based on the context features of each token in the test question text.

[0182] Based on any of the above embodiments, the semantic analysis unit is specifically configured to: Perform macro question type classification, subject special classification and examination target classification on the test question text.

[0183] Based on any of the above embodiments, the illustration description information includes at least one of subject knowledge points, visual features and teaching attributes of the candidate illustrations, and the illustration description information is generated based on a large language model.

[0184] Based on any of the above embodiments, the requirement matching unit is specifically configured to: determine a target illustration matching the test question text from the candidate illustrations based on similarity between a requirement semantic of the illustration requirement information and an illustration semantic of illustration description information of each candidate illustration.

[0185] Based on any of the above embodiments, the requirement matching unit is specifically configured to: determine a target illustration matching the test question text from the candidate illustrations based on similarity between a requirement semantic of the illustration requirement information and an illustration semantic of illustration description information of each candidate illustration, and similarity between a text semantic of the test question text and an image feature of each candidate illustration.

[0186] Based on any of the above embodiments, the apparatus further includes an illustration generation unit configured to: in a case where there is no target illustration matching the test question text in the candidate illustrations, determine a visualization prompt based on the illustration requirement information, and generate a target illustration matching the test question text based on the visualization prompt.

[0187] Based on any of the above embodiments, the apparatus further includes a layout unit configured to: determine a layout strategy based on a test question type of the test question text and a device type of a display device; display the test question text and the target illustration based on the layout strategy.

[0188] Figure 10 An example of a schematic diagram of a physical structure of an electronic device is shown in Figure 10 The electronic device can include a processor 1010, a communications interface 1020, a memory 1030, and a communications bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other through the communications bus 1040. The processor 1010 can invoke a logical instruction in the memory 1030 to execute a text-illustration matching method, which includes: obtaining a test question text; performing knowledge point entity recognition on the test question text to obtain knowledge point entities in the test question text and entity relationships of the knowledge point entities, and performing test question type classification and difficulty classification on the test question text to obtain a test question type and a test question difficulty of the test question text; determining illustration requirement information of the test question text based on the knowledge point entities in the test question text and the entity relationships of the knowledge point entities, the test question type, and the test question difficulty of the test question text; determine a target illustration from the candidate illustrations based on the illustration requirement information and illustration description information of the candidate illustrations.

[0189] In addition, the logic instructions in the memory 1030 described above can be realized in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application or the part that contributes to the related art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0190] In another aspect, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the graphic-text matching method provided by the above-mentioned methods, and the method includes: obtaining a test question text; performing knowledge point entity recognition on the test question text to obtain knowledge point entities in the test question text and entity relationships of the knowledge point entities, performing type classification and difficulty classification on the test question text to obtain a test question type and a test question difficulty of the test question text; determining illustration requirement information of the test question text based on the knowledge point entities in the test question text and the entity relationships of the knowledge point entities, the test question type and the test question difficulty of the test question text; determining a target illustration from the candidate illustrations based on the illustration requirement information and illustration description information of the candidate illustrations.

[0191] In still another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the graphic-text matching method provided by the above-mentioned methods, and the method includes: obtaining a test question text; performing knowledge point entity recognition on the test question text to obtain knowledge point entities in the test question text and entity relationships of the knowledge point entities, performing type classification and difficulty classification on the test question text to obtain a test question type and a test question difficulty of the test question text; determine the illustration requirement information of the test question text based on the knowledge point entity in the test question text and the entity relationship of the knowledge point entity, the test question type and the test question difficulty of the test question text; determine the target illustration from the candidate illustrations based on the illustration requirement information and the illustration description information of each candidate illustration.

[0192] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0193] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course, it can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0194] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of image-text matching, characterized by, The method comprises the following steps: acquiring a test question text; performing knowledge point entity recognition on the test question text to obtain knowledge point entities in the test question text and entity relationships of the knowledge point entities, performing test question type classification and difficulty classification on the test question text to obtain a test question type and a test question difficulty of the test question text; based on the knowledge point entities in the test question text and the entity relationships of the knowledge point entities, the test question type and the test question difficulty of the test question text, determining an illustration requirement information of the test question text; based on the illustration requirement information and illustration description information of each candidate illustration, determining a target illustration matching the test question text from the candidate illustrations.

2. The graph matching method of claim 1, wherein, The knowledge point entity recognition on the test question text comprises the following steps: performing entity recognition on the test question text to obtain entities in the test question text; mapping the entities to a knowledge graph to obtain the knowledge point entities and graph nodes of the knowledge point entities in the knowledge graph; performing relationship path reasoning on the graph nodes of the knowledge point entities in the knowledge graph to obtain entity relationships of the knowledge point entities.

3. The graph matching method of claim 2, wherein, The entity recognition on the test question text comprises the following steps: encoding word piece features of each word piece in the test question text, the word piece features comprising word vectors, position vectors and segmentation vectors of the word pieces; determining context features of each word piece in the test question text based on dependency relationships between the word piece features of each word piece in the test question text; performing entity recognition based on the context features of each word piece in the test question text.

4. The graph matching method of claim 1, wherein, The test question type classification on the test question text comprises the following steps: performing macro test question type classification, subject special classification and examination target classification on the test question text.

5. The graph matching method according to any one of claims 1 to 4, characterized in that, The illustration description information comprises at least one of subject knowledge points, visual features and teaching attributes of the candidate illustrations, and the illustration description information is generated based on a large language model.

6. The graph matching method according to any one of claims 1 to 4, wherein, The determination of the target illustration matching the test question text from the candidate illustrations based on the illustration requirement information and the illustration description information of each candidate illustration comprises the following steps: determining the target illustration matching the test question text from the candidate illustrations based on similarities between requirement semantics of the illustration requirement information and illustration semantics of the illustration description information of each candidate illustration.

7. The graph matching method of claim 6, wherein, The determination of the target illustration matching the test question text from the candidate illustrations based on the similarities between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration comprises the following steps: determining the target illustration matching the test question text from the candidate illustrations based on the similarities between the requirement semantics of the illustration requirement information and the illustration semantics of the illustration description information of each candidate illustration and similarities between text semantics of the test question text and image features of each candidate illustration.

8. The graph matching method of any one of claims 1 to 4, wherein, The method further comprises the following steps: In the case that there is no target illustration matching the test question text in the candidate illustrations, determining a visualization prompt based on the illustration requirement information, generating a target illustration matching the test question text based on the visualization prompt.

9. The graph matching method according to any one of claims 1 to 4, wherein, Also comprising: determining a layout strategy based on a test question type of the test question text and a device type of a display device; displaying the test question text and the target illustration based on the layout strategy.

10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor implements the computer program to realize the image-text matching method of any one of claims 1 to 9. 11.A non-transitory computer-readable storage medium having stored thereon a computer program. The computer program is executed by the processor to realize the image-text matching method of any one of claims 1 to 9.

12. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the image-text matching method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Text illustration method and device

    CN108733779A

  • Image searching method and device, electronic equipment and computer readable storage medium

    CN114741550A

  • Image illustration method and device, electronic equipment and storage medium

    CN118229809A

  • Scientific and technological text picture and text matching algorithm

    CN119782502A

  • Multi-modal ancient poetry knowledge graph construction method for mutual conversion of ancient poetry and image

    CN120179828A