Original title determination method and device, terminal equipment and storage medium

CN122734831APending Publication Date: 2026-09-11GUANGDONG XIAOTIANCAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610750279.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-27
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0003]然而,在实际应用过程中,题目文本可能因识别偏差出现字符错误、数字偏差、内容乱序等情况,仅通过文本相似程度计算难以准确捕捉题目核心含义,计算的相似度容易偏离真实情况,导致原题判断的准确率下降

Benefits of technology

[0017]本申请实施例中,同步获取待判定题目与候选题目的文本信息和图像信息,拓宽题目信息采集维度,避免仅依靠单一文本信息造成题目原始特征遗漏,将文本信息分别送入文本相似度模型与经指令微调的纯文本大语言模型,将图像信息输入经指令微调的多模态大语言模型,三者并行处理,由三条链路独立输出各自的判断结果,其中,文本相似度模型可以精准抓取文本表层字符特征,从表层匹配角度提供相似度参考,生成量化的文本相似度得分;纯文本大语言模型可从深层语义角度对文本信息进行理解,挖掘文本深层语义逻辑并输出第一分类结果,能够在文本信息存在字符错误、数字偏差或乱序等识别偏差时,仍准确捕捉题目核心含义,有效规避识别偏差对判断结果的影响,文本信息双重处理的方式可实现文本层面特征的互补校验;多模态大语言模型可解析图像内部视觉特征,从图文联合语义角度对图像信息进行理解并输出第二分类结果,能够充分提取和利用题目中实际包含的图形、公式等视觉要素,弥补单纯文本处理方式无法利用图形特征的缺陷。三条链路并行运行、优势互补,再根据题目类型、文本相似度得分、第一分类结果以及第二分类结果进行综合,使得最终判定结果融合了表层匹配、深层语义和图文联合理解的多维信息,从而有效提升原题判断的准确性和可靠性,有效规避识别偏差和特征缺失带来的影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122734831A_ABST
    Figure CN122734831A_ABST
Patent Text Reader

Abstract

This application relates to the field of smart terminal technology and provides a method, apparatus, terminal device, and storage medium for determining the original question. The method includes: acquiring text information and image information corresponding to the question to be determined and candidate questions, respectively; inputting the text information into a trained text similarity model to obtain a text similarity score between the question to be determined and the candidate questions; inputting the text information into a first language model and outputting a first classification result indicating whether it is the original question; inputting the image information into a second language model and outputting a second classification result indicating whether it is the original question; and determining the original question determination result based on the question type, text similarity score, first classification result, and second classification result of the question to be determined and the candidate questions. This application can improve the accuracy and reliability of original question determination and effectively avoid the impact of recognition bias and feature loss.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart terminal technology, and in particular to a method, apparatus, terminal device and storage medium for determining the original problem. Background Technology

[0002] In scenarios such as question bank management, test question deduplication, and online education, it is often necessary to determine whether a question is the same as the original question in order to achieve goals such as question bank deduplication and test question quality control. Currently, question determination mainly revolves around text information. By processing the question text, calculating the similarity between texts, and then determining whether it is the same question as the original based on the similarity.

[0003] However, in practical applications, the question text may contain character errors, numerical discrepancies, or disordered content due to recognition biases. Relying solely on text similarity calculations is insufficient to accurately capture the core meaning of the question, and the calculated similarity score can easily deviate from the actual situation, leading to a decrease in the accuracy of the original question's judgment. Furthermore, this method typically only processes textual information, failing to effectively extract and utilize visual elements such as graphics and formulas actually contained within the question, rendering graphic information ineffective in the original question's judgment. This results in biased judgments, consequently affecting the overall accuracy and reliability of the original question's judgment.

[0004] Therefore, how to improve the accuracy and reliability of original question judgment and effectively avoid the impact of identification bias and feature loss is a problem that needs to be solved. Summary of the Invention

[0005] This application provides a method, apparatus, terminal device, and storage medium for judging original questions, which can improve the accuracy and reliability of judging original questions and effectively avoid the impact of recognition bias and feature loss.

[0006] Firstly, embodiments of this application provide a method for determining the original problem, including: Obtain the text and image information corresponding to the questions to be judged and the candidate questions, respectively; The text information is input into a trained text similarity model to obtain the text similarity score between the judgment question and the candidate question; The text information is input into the first large language model, and the first large language model outputs a first classification result representing whether it is the original question. The first large language model is a pure text large language model that has been fine-tuned by instructions. The image information is input into the second large language model, and the second large language model outputs a second classification result representing whether it is the original question. The second large language model is a multimodal large language model that has been fine-tuned by instructions. The original question judgment result is determined based on the question type of the question to be judged and the candidate questions, the text similarity score, the first classification result and the second classification result.

[0007] In one possible implementation of the first aspect, determining the original question judgment result based on the question types of the question to be judged and the candidate questions, the text similarity score, the first classification result, and the second classification result includes: The fusion weight is determined based on the question types of the question to be judged and the candidate questions; The text similarity score, the first classification result, and the second classification result are weighted and fused using the fusion weights to obtain a comprehensive score. The original question judgment result is determined based on the comprehensive score and the preset judgment threshold.

[0008] In one possible implementation of the first aspect, determining the fusion weight based on the question types of the question to be judged and the candidate questions includes: When both the question to be judged and the candidate question are plain text questions, the first weight configuration is called to determine the fusion weight. In the first weight configuration, the weight of the first classification result is the highest, the weight of the text similarity score is the second highest, and the weight of the second classification result is the lowest. When the question type of the question to be judged is a question containing a picture, the second weight configuration is called to determine the fusion weight. In the second weight configuration, the weight of the second classification result is the highest.

[0009] In one possible implementation of the first aspect, when the question type of the question to be judged is a question containing images, before the weighted fusion of the text similarity score, the first classification result, and the second classification result using the fusion weights to obtain a comprehensive score, the method further includes: Based on the text information, calculate the text difference rate between the question to be judged and the candidate questions; When the text difference rate exceeds the preset difference threshold, it is directly determined to be a non-original question, and the weighted fusion is no longer performed; When the text difference rate is less than or equal to the preset difference threshold, the step of using the fusion weight to perform weighted fusion of the text similarity score, the first classification result and the second classification result to obtain a comprehensive score is executed.

[0010] In one possible implementation of the first aspect, the fine-tuning process of the first large language model includes: Construct a first original dataset, which contains question text samples labeled with the original question judgment results and the corresponding original question texts in the question bank. The question text samples include at least one of the following: text error, content disorder, or synonym rewriting. The samples in the first original dataset are formatted using a preset first instruction template to obtain training data in the first instruction format. The first instruction template contains a prompt message that requires the model to determine whether the input question text is the same as the original question text in the question bank. Using the training data in the first instruction format, the parameters of the pre-trained plain text large language model are adjusted to obtain the first large language model after instruction fine-tuning.

[0011] In one possible implementation of the first aspect, the fine-tuning process of the second large language model includes: Construct a second original dataset, which contains question image samples labeled with the original question judgment results and the corresponding original questions from the question bank. The question image samples include at least one of the following: text error, missing or deformed image. The samples in the second original dataset are formatted using a preset second instruction template to obtain training data in the second instruction format. The second instruction template contains prompts that require the model to determine whether the input question image is the same as the original question in the question bank and to output the basis for the determination. Using the training data in the second instruction format, the parameters of the pre-trained multimodal large language model are adjusted to obtain the second large language model after instruction fine-tuning.

[0012] In one possible implementation of the first aspect, obtaining the image information corresponding to the question to be judged and the candidate questions includes: When the question to be judged and the candidate questions contain images, the original image information of the questions to be judged and the candidate questions containing images is directly obtained; When neither the question to be judged nor the candidate question contains an image, the text information corresponding to the question to be judged and the candidate question are respectively converted into corresponding image information in the form of screenshots.

[0013] Secondly, embodiments of this application provide an original problem determination device, including: The information acquisition unit is used to acquire the text information and image information corresponding to the question to be judged and the candidate questions, respectively. The similarity assessment unit is used to input the text information into a trained text similarity model to obtain the text similarity score between the judgment question and the candidate question; The first classification unit is used to input the text information into the first large language model and output a first classification result representing whether it is the original question through the first large language model. The first large language model is a pure text large language model that has been fine-tuned by instructions. The second classification unit is used to input the image information into the second large language model, and output a second classification result representing whether it is the original question through the second large language model. The second large language model is a multimodal large language model that has been fine-tuned by instructions. The original question judgment unit is used to determine the original question judgment result based on the question type of the question to be judged and the candidate questions, the text similarity score, the first classification result and the second classification result.

[0014] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the original question judgment method as described in the first aspect above.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the original problem judgment method as described in the first aspect above.

[0016] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the original problem judgment method as described in the first aspect above.

[0017] In this embodiment, text and image information of the question to be judged and candidate questions are acquired simultaneously, broadening the dimensions of question information collection and avoiding the omission of original features of the questions due to relying solely on text information. The text information is fed into a text similarity model and a finely tuned pure text large language model, while the image information is input into a finely tuned multimodal large language model. The three are processed in parallel, and each of the three links independently outputs its own judgment result. Among them, the text similarity model can accurately capture the surface character features of the text, providing similarity reference from the surface matching perspective and generating a quantified text similarity score; the pure text large language model can analyze the text from a deep semantic perspective. The system performs a comprehensive analysis of textual information, uncovering its deep semantic logic and outputting a primary classification result. Even with errors such as character mistakes, numerical deviations, or disordered order in the text, it accurately captures the core meaning of the question, effectively mitigating the impact of these errors on the judgment. The dual processing of textual information allows for complementary verification of textual features. A multimodal large language model analyzes the internal visual features of images, understanding image information from a combined text-image semantic perspective and outputting a secondary classification result. This fully extracts and utilizes the visual elements such as graphics and formulas actually contained in the question, compensating for the inability of purely textual processing to utilize graphic features. These three parallel processes complement each other, and the results are then integrated based on question type, text similarity score, primary classification result, and secondary classification result. This results in a final judgment that combines multidimensional information from surface matching, deep semantics, and combined text-image understanding, effectively improving the accuracy and reliability of the original question judgment and mitigating the impact of recognition errors and feature omissions. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the implementation of the original problem judgment method provided in the embodiments of this application; Figure 2 This is a flowchart illustrating a specific implementation of the first major language model fine-tuning in the original question judgment method provided in this application embodiment; Figure 3 This is a flowchart illustrating a specific implementation of the second major language model fine-tuning in the original question judgment method provided in this application embodiment; Figure 4 This is a flowchart illustrating a specific implementation of step S105 in the problem-solving method provided in this application embodiment; Figure 5This is a flowchart illustrating a specific implementation of the original problem judgment method for determining fusion weights provided in the embodiments of this application. Figure 6 This is a flowchart illustrating a specific implementation of text difference rate verification in the original question judgment method provided in this application embodiment; Figure 7 This is a structural block diagram of the original problem judgment device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the terminal device provided in the embodiments of this application. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0021] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0022] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0023] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."

[0024] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0026] By way of example and not limitation, the original question determination method provided in this application is applicable to various types of terminal devices that need to perform original question determination. Specific terminal devices may include mobile phones, tablets, wearable devices, laptops, Ultra-Mobile Personal Computers (UMPCs), desktop computers, and servers, etc. This application does not impose any restrictions on the specific type of terminal device.

[0027] Figure 1 The implementation flow of the original problem determination method provided in this application embodiment is shown. The method flow includes steps S101 to S105. The specific implementation principle of each step is as follows: Step S101: Obtain the text information and image information corresponding to the question to be judged and the candidate questions respectively.

[0028] The question to be judged refers to the target question that needs to be determined whether it is the same as the original question. The candidate question refers to a question in the question bank that may be the same as the question to be judged, and is used to compare with the question to be judged to determine whether the two are the same original question. In one possible implementation, the questions to be judged come from scenarios such as user uploads and new additions to the question bank.

[0029] The text information includes all text content contained in the questions to be judged and the candidate questions, including the question stem, known conditions, problem description, options, formula textual descriptions, etc.; the image information includes all graphics and image content contained in the questions to be judged and the candidate questions, specifically including geometric figures, function graphs, schematic diagrams, table images, and page screenshots corresponding to the question text, etc.

[0030] In one possible implementation, when the question to be judged and the candidate questions contain graphics, the original image information of the graphics in the question to be judged and the candidate questions is directly obtained. Here, the original image information refers to the graphic image data carried by the question itself, without any editing or processing, and can completely preserve the original features of the graphics, such as shape, size, structural relationships, and symbol annotations. When the question to be judged and the candidate questions do not contain graphics, the text information corresponding to the question to be judged and the candidate questions is respectively generated as screenshots to form corresponding image information. The screenshots must completely preserve the visual features of the text, such as layout, font size, and symbol position, to ensure that the generated image information accurately reflects the original presentation state of the text.

[0031] For example, if the question to be judged is a mathematical geometry problem, with a text description and a triangle figure in the question stem, and the candidate question is a corresponding standard geometry problem in the question bank, then the text stems of the question to be judged and the candidate question are directly obtained as text information, and the original images of the triangle figures in the question to be judged and the candidate question are directly obtained as image information; if the question to be judged is a Chinese reading comprehension question, without any figures, and the candidate question is a corresponding standard reading comprehension question in the question bank, then the text content in the question to be judged and the candidate question is obtained as their respective text information, and at the same time, full-screen screenshots of the text content of the question to be judged and the candidate question are taken to generate their respective corresponding image information.

[0032] In this embodiment, by simultaneously collecting data from both text and image information corresponding to the question to be judged and the candidate questions, it is ensured that the collected information fully covers all the content of the question. Compared with the method of only acquiring text information, the acquisition of image information can preserve the original form of visual elements such as graphics and formulas in the question, avoid the omission of graphic features and the deviation caused by text recognition errors due to only extracting text, and help improve the accuracy and reliability of the original question judgment.

[0033] In one possible implementation, before inputting the text information into the trained text similarity model and the first language model, the text information corresponding to the question to be judged and the candidate questions is normalized. One possible implementation includes at least one of the following methods: converting full-width characters to half-width characters, converting Chinese punctuation marks to English punctuation marks or standardizing them, removing redundant spaces and line breaks, standardizing the case of letters, replacing or deleting special or invisible characters, standardizing the representation of symbols in mathematical formulas, and standardizing the representation of numbers in the text.

[0034] By normalizing the text information of the questions to be judged and the candidate questions, the text differences caused by non-semantic factors such as differences in character encoding, inconsistent punctuation, spaces and line breaks are eliminated. This enables the subsequent text similarity model and the first language model to perform similarity calculation and semantic judgment based on the content substance, avoiding interference with the accuracy of the original question judgment due to differences in text format.

[0035] Step S102: Input the text information into the trained text similarity model to obtain the text similarity score between the judgment question and the candidate question.

[0036] The trained text similarity model in this application refers to a model that has undergone parameter optimization using labeled data, possesses stable similarity calculation capabilities, and can calculate the degree of similarity between two texts. For example, the text similarity model of this application embodiment can be trained based on the traditional BERT model. However, this application embodiment does not impose any limitations on the specific type of text similarity model.

[0037] This text similarity model extracts surface character and lexical features from text information, maps the text into vector representations, calculates the similarity between the text information of the question to be judged and the candidate questions, and finally outputs a quantified text similarity score. The text similarity score is a quantified numerical value output by the text similarity model to characterize the surface similarity between the question to be judged and the candidate questions. The score ranges from 0 to 1; a higher score indicates a higher surface similarity, and a lower score indicates a lower surface similarity.

[0038] In this embodiment, a text similarity model is used to perform surface matching of text information, capturing the degree of similarity between two question texts at the character and vocabulary levels, and generating a quantified text similarity score. This text similarity model can effectively identify subtle differences at the character level and ensure the stability of text surface feature comparison.

[0039] Step S103: Input the text information into the first large language model, and output the first classification result representing whether it is the original question through the first large language model.

[0040] The first major language model is a finely tuned, pure text-based language model. Its architecture is based on a large-scale pre-trained neural network, possessing deep semantic understanding and logical reasoning capabilities. It can directly output a primary classification result, indicating whether the question to be judged and candidate questions are the same question, based on the input text information. For example, the first major language model can be obtained by fine-tuning Qwen3. The input and output of this pure text-based language model are both in text form, without involving the processing of other modalities such as images. The primary classification result refers to the binary label output by the first major language model, representing whether it is the original question or a different question.

[0041] In the embodiments of the present application, by utilizing the deep semantic understanding capability of the pure-text large language model, the core meaning and logical relationship of the to-be-determined question and the candidate question text are analyzed from the semantic level. When the text has recognition deviations such as character errors, numerical deviations or content out-of-order, the core test points and known conditions of the question can still be accurately captured, and the first classification result is output. Compared with the text similarity model that only performs matching based on surface character features, the pure-text large language model can understand the deep semantics of the text, and has stronger generalization ability for situations such as synonymous rewriting and differences in expression methods.

[0042] Instruction fine-tuning is a specific method of fine-tuning, which guides the model to learn to complete tasks in accordance with the requirements of instructions by constructing formatted training data including task instructions and input-output examples.

[0043] As a possible implementation of the present application, Figure 2 a specific implementation flow of fine-tuning the first large language model in the original question determination method provided by the embodiments of the present application is shown, which is described in detail as follows: A1: Construct a first original data set, where the first original data set includes question text samples labeled with original question determination results and corresponding original question texts in the question bank. The question text sample includes at least one of text errors, content out-of-order or synonymous rewriting.

[0044] The first original data set refers to the basic training data set used for instruction fine-tuning of the first large language model. Its core function is to provide targeted training samples for the original question determination task for the pure-text large language model, so as to ensure that the model can adapt to the original question determination scenario. The question text sample refers to the question text data collected from actual application scenarios, which may contain various text quality problems. The original question text in the question bank refers to the standard original question text corresponding to the question text sample in the question bank, which is used as a comparison benchmark, and its text information is complete and free of recognition errors. The original question determination result refers to the labeled label, which clearly marks whether the question text sample and the original question text in the question bank are the same question.

[0045] Text errors include character errors, numerical deviations and punctuation omissions that occur during text collection or recognition. For example, "triangle" is recognized as "triangie", and the number "30" is recognized as "3O". Content out-of-order refers to the situation where the arrangement order of characters, words or sentences in the question text sample is inconsistent with the standard expression. For example, "angle A equals 30 degrees" is presented as "30 degrees angle A equals". Synonymous rewriting refers to the rewriting of the question text sample that is semantically equivalent but has different expression forms. Its core meaning, test points and answering requirements are completely consistent with the original question text in the question bank. For example, "find the length of side BC" is rewritten as "calculate the length of BC side".

[0046] A2: Using a preset first instruction template, the samples in the first original dataset are formatted to obtain training data in the first instruction format. The first instruction template includes a prompt message requiring the model to determine whether the input question text is the same as the original question text in the question bank.

[0047] The first instruction template refers to a pre-defined text template used to convert raw data into an instruction format acceptable to the model. Formatting the samples in the first raw dataset involves organizing and filling the question text samples, original question texts from the question bank, and labeled original question judgment results in the first raw dataset according to the fixed format of the first instruction template, generating training samples that meet the model's input format requirements. Training data in the first instruction format refers to instruction format data that, after formatting, can be directly used for model training. In this embodiment, the prompt information refers to the explanatory text included in the first instruction template, used to guide the model to perform a specific task, clearly informing the model of the task it needs to complete: determining whether the input question text is the same question as the original question text in the question bank. For example, "Please determine whether the following question text is the same question as the original question text in the question bank, and output yes or no."

[0048] A3: Using the training data in the first instruction format, adjust the parameters of the pre-trained plain text large language model to obtain the first large language model after instruction fine-tuning.

[0049] A pre-trained plain text large language model refers to a plain text large language model that has been pre-trained on a large-scale general text corpus and possesses general language understanding and generation capabilities. It can only process plain text information and has general text semantic mining capabilities. In this embodiment, parameter adjustment refers to updating some or all of the model's parameters based on the pre-trained model parameters, using training data in the first instruction format, through optimization algorithms such as backpropagation and gradient descent. This allows the model to achieve better performance on the original question-judgment task and gradually adapt to it.

[0050] In this embodiment, a first original dataset is constructed to provide training data for fine-tuning the instructions of the first large language model. This training data includes various text anomalies and fits real-world application scenarios. The model is guided to clarify the training task through a preset first instruction template, ensuring that the model can accurately understand the core requirements of original question judgment. Then, through parameter adjustment, the pre-trained pure text large language model acquires a discrimination ability specifically for the original question judgment task. This strengthens the model's ability to mine deep semantics of text and resist text anomaly interference, enabling it to accurately determine whether two texts are the same question even when there are recognition errors, disordered content, or paraphrasing. This ensures that the first large language model can accurately output the first classification result representing whether it is the original question, improving the accuracy and robustness of the model in judging the original questions of pure text questions.

[0051] Step S104: Input the image information into the second language model, and output a second classification result representing whether it is the original question through the second language model.

[0052] The second major language model is a finely tuned multimodal language model capable of simultaneously processing multiple modalities, including images and text. Based on the input multimodal information, it directly outputs a secondary classification result indicating whether the question to be judged and candidate questions are the same question. This multimodal language model incorporates a visual encoder and a text encoder, capable of extracting visual features from images and semantic features from text, respectively. Internally, it fuses information from different modalities to achieve joint semantic understanding of images and text. For example, the second major language model can be obtained by fine-tuning Qwen2.5-VL. The secondary classification result refers to the binary label output by the second major language model, representing whether it is the original question or a different question.

[0053] In this embodiment, the image-text joint understanding capability of the multimodal large language model is utilized to extract visual features of the question graphics from image information, including shape, size, structural relationships, and symbol annotations. Within the model, the visual features are fused and jointly reasoned with the text semantic features in the image to output a second classification result. Compared with processing methods that rely solely on text information, the multimodal large language model can fully extract and utilize the visual elements such as graphics and formulas actually contained in the question, making up for the problem that graphic features cannot be utilized at the algorithm level. It has a stronger adaptability to situations where the question contains graphic elements but text processing methods cannot effectively utilize graphic features.

[0054] As one possible implementation of this application Figure 3 This paper illustrates a specific implementation process for fine-tuning the second language model in the original question judgment method provided in this application, detailed below: B1: Construct a second original dataset, which contains question image samples labeled with the original question judgment results and the corresponding original questions from the question bank. The question image samples include at least one of the following: text errors, missing images, or deformations.

[0055] The second original dataset refers to the basic training data set used for fine-tuning the second largest language model. Its core function is to provide targeted training samples for the original question judgment task in multimodal scenarios. Question image samples refer to question image data collected from real-world application scenarios, capable of fully representing the text content and graphic elements in the question, including complete images of the question stem, geometric figures, function graphs, etc. Original questions in the question bank refer to the standard original questions in the question bank corresponding to the question image sample; these can be standard images or standard text of the original questions.

[0056] Graphical omissions refer to situations where graphic elements in the sample image of the question are partially missing or obscured, resulting in the incomplete presentation of part of the graphic structure; graphic distortions refer to situations where graphic elements in the sample image of the question have distorted lines, size deviations, or incorrect symbol annotations, resulting in differences between the graphic and the standard graphic.

[0057] B2: Using a preset second instruction template, the samples in the second original dataset are formatted to obtain training data in the second instruction format. The second instruction template includes a prompt message requiring the model to determine whether the input question image is the same as the original question in the question bank and to output the basis for that determination.

[0058] The second instruction template refers to a pre-defined template used to convert raw image data into an instruction format acceptable to the model. Formatting the samples in the second raw dataset involves integrating the question image samples, original questions from the question bank, and the judgment results of those questions according to the fixed format of the second instruction template, forming standardized multimodal training samples. The training data in the second instruction format refers to multimodal instruction format data that has been formatted and can be directly used for model training. In this embodiment, the prompt information refers to the text content in the second instruction template used to inform the model of the task objective, explicitly requiring the model to compare the input question image with the original questions in the question bank, determine whether they are the same original question, and output the judgment criteria, i.e., explaining the core reasons for determining it as "original question" or "not original question." For example, pointing out specific reasons such as consistent graphic shapes, identical coordinate labels, and matching text content.

[0059] B3: Using the training data in the second instruction format, adjust the parameters of the pre-trained multimodal large language model to obtain the second large language model after instruction fine-tuning.

[0060] A pre-trained multimodal large language model refers to a multimodal large language model that has been pre-trained on large-scale text and image data and possesses general text and image understanding capabilities. It can process both textual and image information simultaneously and has dual capabilities in text semantic mining and image feature parsing. In this embodiment, parameter adjustment refers to optimizing and updating some or all parameters of the pre-trained multimodal large language model using training data in the second instruction format, enabling the model to gradually adapt to the original question judgment task in a multimodal scenario.

[0061] In this embodiment, by constructing a training dataset of question images containing various abnormalities such as text errors, missing or deformed graphics, the model can learn during fine-tuning to accurately determine whether a question is the same as the original question even when the image quality is poor or the graphic information is incomplete. The image samples are converted into a unified multimodal instruction format through a preset second instruction template, and the model is guided to learn to perform image-text joint reasoning by requiring the model to output prompts on the basis of judgment. This makes the model's judgment process interpretable and strengthens the model's ability to jointly focus on graphic visual features and text semantic features. Then, through parameter adjustment, the pre-trained general multimodal model acquires image-text joint discrimination ability specifically for the original question judgment task. This enables it to accurately determine whether two questions are the same question even when the question image contains text errors, missing or deformed graphics, and output a second classification result representing whether it is the original question and the basis for judgment.

[0062] Step S105: Determine the original question judgment result based on the question type of the question to be judged and the candidate questions, the text similarity score, the first classification result and the second classification result.

[0063] Question type refers to the category categorized based on whether the question contains graphic elements. Specifically, it is divided into plain text questions and questions containing graphics. Plain text questions are those that contain no graphics or image elements, consisting only of text information. Questions containing graphics include graphic elements such as graphs, function graphs, and diagrams. The original question judgment result refers to the final output conclusion indicating whether the question to be judged and the candidate questions are the same question.

[0064] In this embodiment, the terminal device integrates the processing results of the three links—text similarity model, first language model, and second language model—and makes a comprehensive decision based on the different question types to obtain the final original question judgment result. The text similarity score provides a quantitative reference for surface-level text matching, the first classification result provides a judgment at the deep semantic level, and the second classification result provides a judgment at the level of image-text joint understanding. These three factors independently judge whether a question is original from different dimensions. The comprehensive analysis integrates multi-dimensional information from surface matching, deep semantics, and image-text joint understanding, achieving collaborative judgment at the text and visual levels. This approach considers the adaptability to different question scenarios, compensates for the shortcomings of single judgment methods in specific scenarios, and thus ensures the accuracy and reliability of the final original question judgment result.

[0065] As one possible implementation of this application Figure 4 A specific implementation flow of step S105 in the original problem judgment method provided in the embodiment of this application is shown below: C1: Determine the fusion weight based on the question types of the question to be judged and the candidate questions.

[0066] The fusion weight is a weighted coefficient used to weight and fuse text similarity scores, primary classification results, and secondary classification results. Its core function is to balance the influence weights of different judgment results, making the comprehensive judgment result more consistent with the characteristics of the question type. Appropriate fusion weights are assigned to each link result according to the question type; that is, the value of the fusion weight is adapted and adjusted according to different question types, so that the comprehensive score can reflect the relative importance of the judgment ability of each link in different scenarios.

[0067] As one possible implementation of this application Figure 5 The following is a detailed description of a specific implementation process for determining the fusion weight in the original problem judgment method provided in this application: C11: When both the question to be judged and the candidate question are plain text questions, the first weight configuration is called to determine the fusion weight. In the first weight configuration, the weight of the first classification result is the highest, the weight of the text similarity score is the second highest, and the weight of the second classification result is the lowest.

[0068] Plain text questions refer to questions whose content consists solely of text and does not contain any graphic elements. The first weight configuration refers to the weight allocation scheme applicable to plain text question scenarios. Highest weight means the result of that link accounts for the largest proportion in the weighted calculation; second highest weight means the proportion is in the middle; and lowest weight means the proportion is the smallest.

[0069] In scenarios where both the question to be judged and the candidate questions are plain text questions, a preset first weight configuration is invoked to clarify the weight priority of the first classification result, text similarity score, and second classification result. The first classification result output by the first large language model is derived from deep semantic understanding and has stronger generalization ability for synonym rewriting and text errors, therefore it is given the highest weight. The text similarity score is derived from surface character matching and has stable reference value in plain text scenarios, therefore it is given the second highest weight. The second classification result output by the second large language model is derived from image information; in plain text scenarios, images are text screenshots with limited graphic features, therefore they are given the lowest weight. In one possible implementation, in the first weight configuration, the weight of the first classification result is set to 0.5, the weight of the text similarity score is set to 0.3, and the weight of the second classification result is set to 0.2.

[0070] In the pure text question scenario, the fusion weight is determined based on the first weight configuration mentioned above, so that the final fusion result relies more on deep semantic judgment and text surface matching, while retaining the auxiliary verification capability of the image dimension.

[0071] C12: When the question type of the question to be judged is a question containing a picture, the second weight configuration is called to determine the fusion weight. In the second weight configuration, the weight of the second classification result is the highest.

[0072] Questions containing graphics refer to questions whose content includes graphic elements such as geometric figures, function graphs, and diagrams. The second weighting configuration refers to the weighting allocation scheme applicable to scenarios involving questions containing graphics.

[0073] In scenarios where the question type to be judged is a question containing images, a preset second weight configuration is invoked, clearly indicating that the second classification result has the highest priority, while the first classification result and text similarity score have lower weights than the second classification result. This ensures that the judgment of questions containing images primarily relies on the image-text joint judgment result of the multimodal model. Specifically, the second classification result output by the second language model is derived based on image-text joint semantic understanding, which can fully extract and utilize the features of visual elements such as images and formulas in the question, thus assigning the second classification result the highest weight. The text similarity score and the first classification result provide references for surface-level text matching and deep semantics, respectively, as auxiliary judgment criteria. In one possible implementation, in the second weight configuration, the weight of the second classification result is set to 0.8, the weight of the text similarity score is set to 0.1, and the weight of the first classification result is set to 0.1.

[0074] In scenarios involving graphic questions, the fusion weights are determined based on the second weight configuration mentioned above, making the fusion results more dependent on the graphic-text joint judgment capability of the multimodal large language model, and giving full play to its advantages in extracting and utilizing graphic visual features.

[0075] In this embodiment, by combining the information features of different question types (pure text questions focus on text information, while graphic questions focus on image information), the weights of each judgment result are reasonably allocated to achieve differentiated configuration of fusion weights, ensuring that the subsequent weighted fusion results can accurately adapt to different question scenarios, thereby improving the accuracy of the original question judgment.

[0076] C2: The text similarity score, the first classification result, and the second classification result are weighted and fused using the fusion weights to obtain a comprehensive score.

[0077] Weighted fusion involves multiplying the output of each link by its corresponding fusion weight and then summing the results. The overall score is a comprehensive rating obtained after weighted fusion, used to comprehensively reflect the overall judgment result of the three links on whether the question is the original one.

[0078] In one possible implementation, the overall score is obtained according to the following formula (1): (1) in, This represents the overall score. This represents the text similarity score (with a value range of [0,1]). This indicates the first classification result (value is 0 or 1). This indicates the second classification result (value is 0 or 1). The weights representing the text similarity scores The weights representing the classification results. This represents the weight of the second classification result.

[0079] By integrating weights to balance the impact of data from various dimensions, the scattered judgment results are transformed into a unified quantitative comprehensive score.

[0080] C3: Determine the original question's judgment result based on the comprehensive score and the preset judgment threshold.

[0081] The preset judgment threshold refers to a pre-set critical value used to distinguish whether a question to be judged and a candidate question are the same original question. In one possible implementation, this value is determined and fixed based on a large amount of training data. By comparing the comprehensive score with the preset judgment threshold, the final judgment of whether a question is the same original question is made based on the comparison result. This realizes the transformation from quantitative scoring to classification judgment, providing a clear and executable decision standard for original question judgment, and avoiding the subjectivity and uncertainty of human judgment.

[0082] In one possible implementation, when the overall score is greater than or equal to the preset judgment threshold, the output represents the judgment result of the original question; when the overall score is less than the preset judgment threshold, the output represents the judgment result of the non-original question.

[0083] For example, the preset judgment threshold is set to 0.8. When the comprehensive score is ≥0.8, it means that the question to be judged and the candidate question are likely to be the same original question, and the output is "is the original question"; when the comprehensive score is <0.8, it means that the two are likely to be the same original question, and the output is "not the original question".

[0084] In this embodiment, by identifying the question types of the question to be judged and the candidate questions, the appropriate fusion weight is determined to achieve differentiated weight configuration. The fusion weight is used to perform weighted fusion of text similarity score, first classification result, and second classification result, integrating judgment data from three dimensions: text surface, text depth, and image-text combination. This breaks the limitations of single-dimensional judgment and achieves synergistic complementarity of multi-dimensional data, so that the comprehensive score can fully reflect the matching degree between the question to be judged and the candidate questions. Then, the comprehensive score is compared with the preset judgment threshold to transform the quantitative data into a clear binary classification judgment result, clarifying the judgment criteria, avoiding subjective errors, and ensuring the consistency and reliability of the judgment results.

[0085] As one possible implementation of this application, when the question type to be judged is a question containing a picture, such as Figure 6 As shown, before using the fusion weights to weight and fuse the text similarity score, the first classification result, and the second classification result to obtain a comprehensive score, a text difference rate verification is performed. This verification process includes: D1: Calculate the text difference rate between the question to be judged and the candidate questions based on the text information.

[0086] The text difference rate is a ratio used to quantify the degree of difference between the text information of the question to be judged and the candidate questions. It is calculated as the ratio of the text edit distance to the total number of characters in the longer text. The text edit distance refers to the minimum number of editing operations required to convert the text information of the question to be judged into the text information of the candidate questions; editing operations include character insertion, deletion, and replacement. In one possible implementation, the text edit distance is calculated based on normalized text information to ensure the accuracy of the calculation results.

[0087] The text difference rate ranges from 0 to 100%. A value of 0 indicates that the two pieces of text information are completely identical, while a value of 100% indicates that the two pieces of text information are completely different. The higher the text difference rate, the higher the degree of difference between the two pieces of text information; the lower the text difference rate, the lower the degree of difference between the two pieces of text information.

[0088] For example, the text of the question to be judged is "In triangle ABC, angle A is 30 degrees, find the length of side BC", and the text of the candidate question is "In triangle ABC, angle A is 35 degrees, find the length of side BC". The difference between the two texts is that the angle value changes from "30" to "35". "30" needs to be replaced with "35", which requires 1 replacement operation. Therefore, the text edit distance is 1, and the character difference rate corresponding to this text edit distance is 10%. If the candidate question text information is completely consistent with the question text information to be judged, the text edit distance is 0.

[0089] D2: When the text difference rate exceeds the preset difference threshold, it is directly determined to be a non-original question, and the weighted fusion is no longer performed.

[0090] The preset difference threshold refers to a pre-set critical value for the text difference rate, used to determine whether the difference between two question texts exceeds an acceptable range.

[0091] When the text difference rate exceeds a preset difference threshold, it indicates that there is a substantial difference between the question to be judged and the candidate question at the text level. Even if the multimodal large language model determines it as the original question based on image features, it should still be rejected at the text level to prevent misjudgment caused by substantial differences in text due to similar graphics. In this case, the judgment result of not being the original question is directly output, and subsequent weighted fusion and related operations are terminated.

[0092] D3: When the text difference rate is less than or equal to the preset difference threshold, perform the step of weighted fusion of the text similarity score, the first classification result and the second classification result using the fusion weight to obtain a comprehensive score.

[0093] When the text edit distance does not exceed the preset difference threshold, it indicates that the difference between the question to be judged and the candidate question at the text level is within an acceptable range. At this time, the normal weighted fusion process is entered. The second weight configuration corresponding to the question scenario with graphics is used to perform weighted fusion of the text similarity score, the first classification result and the second classification result to obtain a comprehensive score, which provides a basis for the final original question judgment.

[0094] In this embodiment, the relative difference between the two text segments is accurately quantified by calculating the text difference rate between the question to be judged and the candidate questions. This eliminates the interference of text length on the difference judgment, ensuring the objectivity and accuracy of the difference judgment. By comparing the text difference rate with a preset difference threshold, questions with excessively large text differences are directly judged as non-original questions, avoiding subsequent complex weighted fusion operations. When the text difference rate does not exceed the preset difference threshold, the weighted fusion process is entered. While making full use of the advantages of multimodal large language model image-text joint judgment, the judgment results are constrained by text-level difference measurement, effectively filtering questions with excessive text modifications and improving the rigor and accuracy of original question judgment in scenarios with image questions.

[0095] As can be seen from the above, in this embodiment, the text and image information of the question to be judged and the candidate questions are acquired simultaneously, broadening the dimensions of question information collection and avoiding the omission of original features of the questions due to relying solely on single text information. The text information is fed into the text similarity model and the finely tuned pure text large language model, respectively, while the image information is input into the finely tuned multimodal large language model. The three are processed in parallel, and each of the three links independently outputs its own judgment result. Among them, the text similarity model can accurately capture the surface character features of the text, providing similarity reference from the surface matching perspective and generating a quantified text similarity score; the pure text large language model can extract the deep character features from the surface matching perspective. From a semantic perspective, this approach understands textual information, uncovers its deep semantic logic, and outputs a primary classification result. Even with character errors, numerical deviations, or disordered order in the text, it accurately captures the core meaning of the question, effectively mitigating the impact of recognition bias on the judgment result. The dual processing of textual information enables complementary verification of textual features. A multimodal large language model analyzes the internal visual features of images, understanding image information from a combined text-image semantic perspective and outputting a secondary classification result. This fully extracts and utilizes the visual elements such as graphics and formulas actually contained in the question, compensating for the inability of pure text processing to utilize graphic features. These three processes operate in parallel, complementing each other. The results are then integrated based on question type, text similarity score, primary classification result, and secondary classification result, resulting in a final judgment that combines multidimensional information from surface matching, deep semantics, and combined text-image understanding. This effectively improves the accuracy and reliability of the original question judgment, effectively avoiding the impact of recognition bias and feature loss.

[0096] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0097] Corresponding to the original question judgment method described in the above embodiments, Figure 7 The diagram shows a structural block diagram of the original problem judgment device provided in the embodiments of this application. For ease of explanation, only the parts related to the embodiments of this application are shown.

[0098] Reference Figure 7 The original question judgment device includes: an information acquisition unit 71, a similarity assessment unit 72, a first classification unit 73, a second classification unit 74, and an original question judgment unit 75, wherein: The information acquisition unit 71 is used to acquire the text information and image information corresponding to the question to be judged and the candidate questions, respectively. The similarity assessment unit 72 is used to input the text information into the trained text similarity model to obtain the text similarity score between the judgment question and the candidate question; The first classification unit 73 is used to input the text information into the first large language model and output a first classification result representing whether it is the original question through the first large language model. The first large language model is a pure text large language model that has been fine-tuned by instructions. The second classification unit 74 is used to input the image information into the second large language model, and output a second classification result representing whether it is the original question through the second large language model. The second large language model is a multimodal large language model that has been fine-tuned by instructions. The original question judgment unit 75 is used to determine the original question judgment result based on the question type of the question to be judged and the candidate questions, the text similarity score, the first classification result and the second classification result.

[0099] As one possible implementation of this application, the original problem determination unit 75 includes: The weight determination module is used to determine the fusion weight based on the question types of the question to be judged and the candidate questions; The comprehensive scoring module is used to perform weighted fusion of the text similarity score, the first classification result and the second classification result using the fusion weights to obtain a comprehensive score. The original question determination module is used to determine the original question determination result based on the comprehensive score and the preset determination threshold.

[0100] As one possible implementation of this application, the weight determination module is specifically used for: When both the question to be judged and the candidate question are plain text questions, the first weight configuration is called to determine the fusion weight. In the first weight configuration, the weight of the first classification result is the highest, the weight of the text similarity score is the second highest, and the weight of the second classification result is the lowest. When the question type of the question to be judged is a question containing a picture, the second weight configuration is called to determine the fusion weight. In the second weight configuration, the weight of the second classification result is the highest.

[0101] As one possible implementation of this application, when the question type of the question to be judged is a question containing a picture, the original question judgment unit 75 further includes: The distance calculation module is used to calculate the text difference rate between the question to be judged and the candidate questions based on the text information. The direct judgment module is used to directly determine that the text difference rate exceeds a preset difference threshold and no longer perform the weighted fusion. The comprehensive scoring module is further configured to perform the step of weighted fusion of the text similarity score, the first classification result, and the second classification result using the fusion weights to obtain a comprehensive score when the text difference rate is less than or equal to the preset difference threshold.

[0102] As one possible implementation of this application, the original problem judgment device further includes a first fine-tuning unit, used for: Construct a first original dataset, which contains question text samples labeled with the original question judgment results and the corresponding original question texts in the question bank. The question text samples include at least one of the following: text error, content disorder, or synonym rewriting. The samples in the first original dataset are formatted using a preset first instruction template to obtain training data in the first instruction format. The first instruction template contains a prompt message that requires the model to determine whether the input question text is the same as the original question text in the question bank. Using the training data in the first instruction format, the parameters of the pre-trained plain text large language model are adjusted to obtain the first large language model after instruction fine-tuning.

[0103] As one possible implementation of this application, the original problem judgment device further includes a second fine-tuning unit, used for: Construct a second original dataset, which contains question image samples labeled with the original question judgment results and the corresponding original questions from the question bank. The question image samples include at least one of the following: text error, missing or deformed image. The samples in the second original dataset are formatted using a preset second instruction template to obtain training data in the second instruction format. The second instruction template contains prompts that require the model to determine whether the input question image is the same as the original question in the question bank and to output the basis for the determination. Using the training data in the second instruction format, the parameters of the pre-trained multimodal large language model are adjusted to obtain the second large language model after instruction fine-tuning.

[0104] As one possible implementation of this application, the information acquisition unit 71 includes an image information acquisition module, used for: When the question to be judged and the candidate questions contain images, the original image information of the questions to be judged and the candidate questions containing images is directly obtained; When neither the question to be judged nor the candidate question contains an image, the text information corresponding to the question to be judged and the candidate question are respectively converted into corresponding image information in the form of screenshots.

[0105] As can be seen from the above, in this embodiment, the text and image information of the question to be judged and the candidate questions are acquired simultaneously, broadening the dimensions of question information collection and avoiding the omission of original features of the questions due to relying solely on single text information. The text information is fed into the text similarity model and the finely tuned pure text large language model, respectively, while the image information is input into the finely tuned multimodal large language model. The three are processed in parallel, and each of the three links independently outputs its own judgment result. Among them, the text similarity model can accurately capture the surface character features of the text, providing similarity reference from the surface matching perspective and generating a quantified text similarity score; the pure text large language model can extract the deep character features from the surface matching perspective. From a semantic perspective, this approach understands textual information, uncovers its deep semantic logic, and outputs a primary classification result. Even with character errors, numerical deviations, or disordered order in the text, it accurately captures the core meaning of the question, effectively mitigating the impact of recognition bias on the judgment result. The dual processing of textual information enables complementary verification of textual features. A multimodal large language model analyzes the internal visual features of images, understanding image information from a combined text-image semantic perspective and outputting a secondary classification result. This fully extracts and utilizes the visual elements such as graphics and formulas actually contained in the question, compensating for the inability of pure text processing to utilize graphic features. These three processes operate in parallel, complementing each other. The results are then integrated based on question type, text similarity score, primary classification result, and secondary classification result, resulting in a final judgment that combines multidimensional information from surface matching, deep semantics, and combined text-image understanding. This effectively improves the accuracy and reliability of the original question judgment, effectively avoiding the impact of recognition bias and feature loss.

[0106] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0107] This application embodiment also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements... Figures 1 to 6 This represents the steps of any original problem-solving method.

[0108] This application embodiment also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements... Figures 1 to 6 This represents the steps of any original problem-solving method.

[0109] This application also provides a computer program product that, when run on a terminal device, causes the terminal device to execute the implementation of... Figures 1 to 6This represents the steps of any original problem-solving method.

[0110] Figure 8 This is a schematic diagram of a terminal device provided in an embodiment of this application. For example... Figure 8 As shown, the terminal device 8 in this embodiment includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80. When the processor 80 executes the computer program 82, it implements the steps in the various original problem judgment method embodiments described above, for example... Figure 1 Steps S101 to S105 are shown. Alternatively, when the processor 80 executes the computer program 82, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 7 The functions of units 71 to 75 shown.

[0111] For example, the computer program 82 may be divided into one or more modules / units, which are stored in the memory 81 and executed by the processor 80 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program 82 in the terminal device 8.

[0112] The terminal device 8 may include, but is not limited to, a processor 80 and a memory 81. Those skilled in the art will understand that... Figure 8 This is merely an example of terminal device 8 and does not constitute a limitation on terminal device 8. It may include more or fewer components than shown, or combine certain components, or different components. For example, terminal device 8 may also include input / output devices, network access devices, buses, etc.

[0113] The processor 80 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0114] The memory 81 can be an internal storage unit of the terminal device 8, such as a hard disk or memory of the terminal device 8. The memory 81 can also be an external storage device of the terminal device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal device 8. Furthermore, the memory 81 can include both internal and external storage units of the terminal device 8. The memory 81 is used to store the computer program and other programs and data required by the terminal device. The memory 81 can also be used to temporarily store data that has been output or will be output.

[0115] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0117] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0118] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0119] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for identifying original questions, characterized in that, include: Obtain the text and image information corresponding to the questions to be judged and the candidate questions, respectively; The text information is input into a trained text similarity model to obtain the text similarity score between the judgment question and the candidate question; The text information is input into the first large language model, and the first large language model outputs a first classification result representing whether it is the original question. The first large language model is a pure text large language model that has been fine-tuned by instructions. The image information is input into the second large language model, and the second large language model outputs a second classification result representing whether it is the original question. The second large language model is a multimodal large language model that has been fine-tuned by instructions. The original question judgment result is determined based on the question type of the question to be judged and the candidate questions, the text similarity score, the first classification result and the second classification result.

2. The method according to claim 1, characterized in that, The step of determining the original question judgment result based on the question types of the question to be judged and the candidate questions, the text similarity score, the first classification result, and the second classification result includes: The fusion weight is determined based on the question types of the question to be judged and the candidate questions; The text similarity score, the first classification result, and the second classification result are weighted and fused using the fusion weights to obtain a comprehensive score. The original question judgment result is determined based on the comprehensive score and the preset judgment threshold.

3. The method according to claim 2, characterized in that, The step of determining the fusion weight based on the question types of the question to be judged and the candidate questions includes: When both the question to be judged and the candidate question are plain text questions, the first weight configuration is called to determine the fusion weight. In the first weight configuration, the weight of the first classification result is the highest, the weight of the text similarity score is the second highest, and the weight of the second classification result is the lowest. When the question type of the question to be judged is a question containing a picture, the second weight configuration is called to determine the fusion weight. In the second weight configuration, the weight of the second classification result is the highest.

4. The method according to claim 2, characterized in that, When the question type to be judged is a question containing images, before the weighted fusion of the text similarity score, the first classification result, and the second classification result using the fusion weight to obtain the comprehensive score, the following steps are included: Based on the text information, calculate the text difference rate between the question to be judged and the candidate questions; When the text difference rate exceeds the preset difference threshold, it is directly determined to be a non-original question, and the weighted fusion is no longer performed; When the text difference rate is less than or equal to the preset difference threshold, the step of using the fusion weight to perform weighted fusion of the text similarity score, the first classification result and the second classification result to obtain a comprehensive score is executed.

5. The method according to claim 1, characterized in that, The fine-tuning process of the first large language model includes: Construct a first original dataset, which contains question text samples labeled with the original question judgment results and the corresponding original question texts in the question bank. The question text samples include at least one of the following: text error, content disorder, or synonym rewriting. The samples in the first original dataset are formatted using a preset first instruction template to obtain training data in the first instruction format. The first instruction template contains a prompt message that requires the model to determine whether the input question text is the same as the original question text in the question bank. Using the training data in the first instruction format, the parameters of the pre-trained plain text large language model are adjusted to obtain the first large language model after instruction fine-tuning.

6. The method according to claim 1, characterized in that, The fine-tuning process of the second major language model includes: Construct a second original dataset, which contains question image samples labeled with the original question judgment results and the corresponding original questions from the question bank. The question image samples include at least one of the following: text error, missing or deformed image. The samples in the second original dataset are formatted using a preset second instruction template to obtain training data in the second instruction format. The second instruction template contains prompts that require the model to determine whether the input question image is the same as the original question in the question bank and to output the basis for the determination. Using the training data in the second instruction format, the parameters of the pre-trained multimodal large language model are adjusted to obtain the second large language model after instruction fine-tuning.

7. The method according to any one of claims 1 to 6, characterized in that, The step of obtaining the image information corresponding to the question to be judged and the candidate questions includes: When the question to be judged and the candidate questions contain images, the original image information of the questions to be judged and the candidate questions containing images is directly obtained; When neither the question to be judged nor the candidate question contains an image, the text information corresponding to the question to be judged and the candidate question are respectively converted into corresponding image information in the form of screenshots.

8. A device for judging original questions, characterized in that, include: The information acquisition unit is used to acquire the text information and image information corresponding to the question to be judged and the candidate questions, respectively. The similarity assessment unit is used to input the text information into a trained text similarity model to obtain the text similarity score between the judgment question and the candidate question; The first classification unit is used to input the text information into the first large language model and output a first classification result representing whether it is the original question through the first large language model. The first large language model is a pure text large language model that has been fine-tuned by instructions. The second classification unit is used to input the image information into the second large language model, and output a second classification result representing whether it is the original question through the second large language model. The second large language model is a multimodal large language model that has been fine-tuned by instructions. The original question judgment unit is used to determine the original question judgment result based on the question type of the question to be judged and the candidate questions, the text similarity score, the first classification result and the second classification result.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the original question judgment method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the original question judgment method as described in any one of claims 1 to 7.