An anti-interference generation alignment method and device for education large model hallucination suppression

CN122838546APending Publication Date: 2026-09-29UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610999051.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0008]为了解决现有技术中存在生成链路复杂、推理成本较高、模型自身抗干扰能力不足、对权威证据辨识不充分、证据缺失情况下安全拒答能力较弱的问题,本发明提供一种面向教育大模型幻觉抑制的抗干扰生成对齐方法及装置

Benefits of technology

[0017]本发明的方法可有效降低教育大模型在教材资料、政策文件、理论文献等含噪上下文中生成事实错误、理论表述不准确或依据不充分回答的概率,提高问答结果的准确性、规范性、权威性和可靠性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838546A_ABST
    Figure CN122838546A_ABST
Patent Text Reader

Abstract

This invention discloses an anti-interference generation alignment method and apparatus for hallucination suppression in educational large-scale models, comprising the following steps: constructing a noise-free document set, a fully noisy document set, and a mixed document set; constructing the input of the educational large-scale model based on the query question and documents in the noise-free document set, and selecting ideal answers from the answers of the educational large-scale model; constructing the input of the educational large-scale model based on the query question and documents in the noise-free document set, the mixed document set, and the fully noisy document set, and training the educational large-scale model by combining ideal answer and rejection answer templates; constructing the input of the educational large-scale model based on the query question and documents in the noise-free document set, the mixed document set, and the fully noisy document set, calculating the reward for each candidate answer in the candidate answer set corresponding to each document, and updating the parameters of the educational large-scale model based on the principle of controlling the learning degree of the educational large-scale model on the candidate answers according to the reward size.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of anti-interference generation alignment technology, and particularly relates to an anti-interference generation alignment method and apparatus for suppressing illusions in educational large-scale models. Background Technology

[0002] The Educational Big Data Model is a product of the deep integration of generative artificial intelligence technology and education. It is a professional, domain-specific large language model serving mainstream ideological education and teaching, aiming to empower high-quality educational development with digital technology. As the application of large language models in education, theoretical learning guidance, course Q&A, and knowledge services deepens, the Educational Big Data Model is increasingly being used to assist students in learning theoretical courses, support teachers in conducting teaching Q&A, provide interpretations of policy documents, and generate teaching content. In these applications, the model typically needs to combine external knowledge such as textbooks, course outlines, policy documents, literature, authoritative statements, and current affairs materials to answer questions, thereby improving the accuracy, standardization, and timeliness of the responses.

[0003] Existing methods typically involve introducing relevance assessment, document filtering, retrieval reordering, or rule review modules before the model generates an answer. These modules screen educational materials, policy documents, textbook excerpts, or case studies that assist the model in answering the questions, minimizing irrelevant, erroneous, outdated, or value-biased content from entering the model context. Alternatively, after the model generates an answer, the methods involve fact-checking, policy consistency checks, sensitive expression detection, and error correction of the model output to reduce the risk of illusory content and inappropriate expressions in the final answer.

[0004] The existing technology has the following problems:

[0005] 1. Context-based filtering methods, such as relevance assessment, document filtering, or reordering, require the introduction of additional judgment or review modules into the educational large-scale model generation chain. While these methods can reduce the input of irrelevant, erroneous, and distracting content into the model to some extent, they also significantly increase system structural complexity, inference latency, and deployment and maintenance costs, which is not conducive to building lightweight, stable, and scalable question-answering and teaching assistance systems.

[0006] 2. Knowledge in the field of education has strong requirements for standardization, authority, and value orientation. Relying solely on document screening before generation is insufficient to completely avoid materials with similar semantics but inconsistent actual stances, timeliness, basis, or scope of application from entering the context. For example, the model may encounter policy fragments that are superficially related to the question but are actually inapplicable, theoretical statements taken out of context, outdated political materials, or content inconsistent with mainstream textbooks. In such cases, if the model itself lacks the ability to identify valid evidence and interfering information, it may still generate inaccurate, non-standard, or inconsistent answers based on incorrect context.

[0007] 3. Output-based verification and correction methods primarily address issues after model generation, mainly performing fact-checking, phrasing correction, or risk interception on already generated answers. They do not improve the educational model's ability to identify evidence, utilize authoritative sources, or safely refuse answers in noisy contexts from the training and alignment perspectives. Therefore, these methods struggle to fundamentally reduce the generation of illusory content and cannot guarantee that the model will proactively avoid overconfident answers when reliable evidence is lacking. Summary of the Invention

[0008] To address the problems of complex generation links, high inference costs, insufficient anti-interference capabilities of the model itself, inadequate identification of authoritative evidence, and weak ability to refuse answers safely in the absence of evidence in existing technologies, this invention provides an anti-interference generation alignment method and device for hallucination suppression in educational large-scale models.

[0009] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0010] The first aspect of this invention provides an interference-resistant generation alignment method for hallucination suppression in large-scale educational models, comprising the following steps:

[0011] Construct a noise-free document set, a full-noise document set, and a hybrid document set. In the noise-free document set, the context of all documents is more relevant to the educational question than a preset threshold. In the full-noise document set, the context of all documents is semantically similar to the educational question but not actually related to it. In the hybrid document set, the context of some documents is more relevant to the educational question than a preset threshold, and the context of the remaining documents is semantically similar to the educational question but not actually related to it.

[0012] The input to the education big data model is constructed based on the query question and the documents in the noise-free document set, and the ideal answer is selected from the candidate answer set of the education big data model according to the quality score.

[0013] Based on the query question and documents in the noise-free document set, mixed document set, and full-noise document set, an educational big data model is constructed as input. The model is then trained by combining ideal answer and rejection answer templates. This allows the educational big data model to strengthen the ideal answer when there is only valid evidence, eliminate interference answers that are closer to the ideal answer when there is invalid evidence, and reject the answer when there is only invalid evidence.

[0014] Based on the query question and documents in the noise-free document set, mixed document set, and full-noise document set, an educational big data model is constructed. The reward of each candidate answer in the candidate answer set corresponding to each document is calculated by combining the reward function corresponding to each document set. The average reward of each candidate answer set is calculated. The parameters of the educational big data model are updated based on the principle of improving the learning degree of the educational big data model for candidate answers with rewards greater than the average reward and suppressing the learning degree of the educational big data model for candidate answers with rewards less than the average reward.

[0015] A second aspect of the present invention provides an interference-resistant generation and alignment device for suppressing large-scale educational hallucinations, comprising a memory and a controller connected in sequence, wherein the memory stores a computer program, and the controller is used to read the computer program and execute the interference-resistant generation and alignment method for suppressing large-scale educational hallucinations as described in the first aspect.

[0016] The beneficial effects of this invention are:

[0017] The method of this invention can effectively reduce the probability of educational big data models generating factual errors, inaccurate theoretical statements, or insufficiently based answers in noisy contexts such as teaching materials, policy documents, and theoretical literature, thereby improving the accuracy, standardization, authority, and reliability of question and answer results.

[0018] The method of this invention does not introduce a judgment module or an audit module, and the link structure is simple and the reasoning cost is reduced. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of the anti-interference generation alignment method for hallucination suppression in educational large-scale models proposed in this invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0023] It should be noted that, unless otherwise specified, the embodiments and features described in this invention can be combined with each other.

[0024] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0025] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. They are only used for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0026] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0027] The first aspect of this invention discloses an anti-interference generation alignment method for suppressing hallucinations in educational large models, the method comprising steps S1 to S5.

[0028] Step S1: Construct a noise-free document set, a full-noise document set, and at least one hybrid document set, wherein the context of all documents in the noise-free document set is more relevant to the educational question than a preset threshold, the context of all documents in the full-noise document set is semantically similar to the educational question but actually unrelated, and the context of some documents in the hybrid document set is more relevant to the educational question than a preset threshold, and the context of the remaining documents is semantically similar to the educational question but actually unrelated.

[0029] Both the mixed document set and the fully noisy document set contain varying degrees of noise. Random noise introduced by randomly selecting irrelevant documents is often too simplistic and fails to effectively simulate the complex interference present in real-world retrieval systems; furthermore, it is easily identified and ignored by large-scale educational models. Therefore, the preferred method for constructing the mixed document set and the fully noisy document set is as follows:

[0030] Documents that are semantically similar to the query q but are actually unrelated are selected from the knowledge base D in the education field as interference noise to ensure that the introduced noise has high confusion and challenge, and is closer to the error distribution characteristics of the real retrieval system.

[0031] Specifically: For each query q, a search is performed, and all documents are sorted in descending order based on their semantic relevance to q, resulting in an ordered candidate list.

[0032]

[0033] Where simcos(·,·) is the cosine similarity calculation function. For the i-th document in the knowledge base D in the field of education, Set a lower bound N for the truncation of the document set consisting of query q. min With upper bound N max Filter the top-ranked candidate documents and construct a pool of noisy document candidates accordingly. :

[0034]

[0035] Then, m sets of noisy documents are randomly sampled from the noisy document candidate pool. :

[0036] .

[0037] This results in scene types with progressively increasing noise levels. The no-interference scene corresponds to a noise-free document set, the full-interference scene corresponds to a full-noise document set, and the mixed-interference scene corresponds to a mixed document set. Multiple mixed document sets can be set, and each set has a different noise level, meaning the proportion of documents whose contextual relevance to the educational question exceeds a preset threshold varies.

[0038] When setting only one mixed document set, it is preferable to set the proportion of documents whose contextual relevance to the educational issue is greater than a preset threshold to 50%. (=0.5)

[0039] At this point, the document composition, interference intensity, and expected behavior for the three scenarios are shown in Table 1 below.

[0040] Table 1

[0041]

[0042] Specifically, the input formats for the three scenarios are as follows:

[0043] Input in interference-free scenarios : = F (q, D+(q));

[0044] Input of mixed interference scene : = F (q, Shuffle (D+(q) ∪ D-(q)));

[0045] Input of a full interference scenario : = F (q, D-(q));

[0046] D+(q) represents a document whose context and query relevance are greater than a preset threshold, indicating that the document is highly relevant to the query.

[0047] Step S2: Construct the input of the education big data model based on the query question and the documents in the noise-free document set, and select the ideal answer from the answers of the education big data model according to the quality score.

[0048] Specifically, for each query q, the query q is combined with each document in the noise-free document set to form the input content, which is then input into the education big data model to generate the corresponding answer.

[0049] Because there are significant differences between the standard answer and the generation style of the educational model itself, it will affect the effectiveness of fine-tuning. Therefore, the educational model will be queried multiple times to obtain a batch of answers and construct a candidate answer set. ,in,

[0050] ,

[0051] In the formula, This represents the i-th answer to the query, where n is the number of samples. This represents the output probability distribution of the large-scale education model under given input conditions and in an interference-free scenario.

[0052] The OmniEval model evaluator is used to automatically score each candidate answer in the candidate answer set. The scoring is based on four dimensions: semantic correctness (accuracy), completeness (whether the answer covers all key aspects that a true answer should include), utilization rate (effectiveness of the model in utilizing retrieved documents), and standardization (consistency with educational theory, policy basis, or value orientation). Accuracy, completeness, utilization rate, and standardization are weighted equally, and the selected answer receives a comprehensive quality score. :

[0053]

[0054] In the formula, y represents the candidate answer. For the accuracy of candidate answers, For the completeness of candidate answers, The utilization rate of candidate answers, To ensure the standardization of candidate answers.

[0055] The answer with the highest quality score is selected as the ideal answer to the query in a non-interference scenario, and is denoted as the ideal answer. .

[0056] Step S3: Based on the query question and the documents in the noise-free document set, mixed document set, and full-noise document set, construct an educational big model input and train the educational big model by combining ideal answer and rejection answer templates, so that the educational big model can answer more closely to the ideal answer when there is valid evidence and reject the answer when there is no valid evidence.

[0057] This step involves multi-task joint training of the large-scale education model across three scenarios. Specifically, it includes steps S31 to S33.

[0058] Step S31: Based on the query question and the documents in the noise-free document set, construct an educational big model input and train the educational big model in combination with the ideal answer, so that the educational big model can directly generate the ideal answer when there is only valid evidence.

[0059] This step is baseline reinforcement, corresponding to a non-interference scenario, and aims to solidify the high-quality generation capability of the large-scale educational model under ideal conditions. Input = F(q, D+(q)), the education big model is based on The generated answers are compared with the ideal answers, making the educational big data model's answers closer to the ideal answers.

[0060] Step S32: Based on the query question and the documents in the mixed document set, construct an educational big model input and train the educational big model by combining it with the ideal answer, so that the educational big model can answer more closely to the ideal answer when there is interference.

[0061] This step involves noise-resistant generation, corresponding to a semi-interference scenario. When irrelevant or erroneous documents are mixed into the context input of the educational big data model, this step improves the model's accuracy in identifying errors and its ability to utilize valid information. Input = F (q, Shuffle (D+(q) ∪ D-(q))), the large-scale education model is based on The generated answers are compared with the ideal answers, so that the large-scale educational model can produce outputs that are close to the ideal answers even when faced with noise interference.

[0062] Step S33: Based on the query question and the documents in the full noisy document set, construct an educational big model input and train the educational big model in combination with the rejection response template, so that the educational big model learns to reject the response when there is a lack of valid evidence.

[0063] This step is for secure rejection, corresponding to a full-interference scenario. When the educational model's context is entirely erroneous, it establishes the model's ability to refuse answers, preventing it from providing incorrect responses to the user based on incorrect information. Input = F(q, D-(q)), the education big model is based on The generated responses are compared with standard rejection templates, allowing the educational big data model to learn to refuse to answer when there is a lack of valid evidence, and not to be overconfident and blindly generate incorrect responses.

[0064] The training data from the three scenarios above are combined to form a multi-task training dataset, and an autoregressive language modeling loss function is used. Optimization is being carried out, including

[0065]

[0066] x represents the input of the sample, and y represents the standard answer of the sample. Given input x and the previously generated correct prefix y, the educational big data model predicts that the t-th token is exactly the correct answer. The probability is calculated by averaging the loss across all samples in the training data. Through joint training in three scenarios, the large education model can simultaneously possess three capabilities: high-quality generation, noise resistance, and safe rejection.

[0067] S4. Based on the query question and the documents in the noise-free document set, mixed document set, and full-noise document set, construct an educational big model input and calculate the reward of each candidate answer in the candidate answer set corresponding to each document by combining the reward function corresponding to each document set. Calculate the average reward of each candidate answer set. Update the parameters of the educational big model based on the principle of improving the learning degree of the educational big model for candidate answers with rewards greater than the average reward and suppressing the learning degree of the educational big model for candidate answers with rewards less than the average reward.

[0068] Document sets with different levels of interference correspond to different reward functions. This step further adjusts the parameters of the large-scale education model based on the rewards. Specifically, this step includes steps S41 to S44.

[0069] Step S41: Based on the query question and the documents in the noise-free document set, the mixed document set, and the full-noise document set, construct an educational big data model input and calculate the reward for each candidate answer in each document candidate answer set by combining the reward function corresponding to each document set. Specifically, the reward for the candidate answer corresponding to the input of the query question and the documents in the noise-free document set is determined based on a general reward, calculated based on accuracy, completeness, utilization, and standardization. The reward for the candidate answer corresponding to the input of the query question and the documents in the mixed document set is determined based on a general reward and a consistency reward with the ideal answer, calculated based on the overlap at the lexical level and the closeness at the semantic level. The reward for the candidate answer corresponding to the input of the query question and the documents in the full-noise document set is determined based on a consistency reward with the rejection template and a consistency reward with the ideal answer.

[0070] Specifically, the reward is based on the input candidate answers constructed from the query question and documents in a noise-free document set. for:

[0071] ,

[0072] ,

[0073] In the formula, y represents the candidate answer. To find the standard answer to question q, , , , The weights are respectively for accuracy, completeness, utilization, and standardization.

[0074] Rewards for candidate answers corresponding to the input of the query question and the documents in the mixed document set for:

[0075] = +

[0076] ,

[0077] ,

[0078] In the formula, The weighting coefficient for consistency rewards. Candidate answers output for the large-scale education model With the ideal answer The consistency reward is calculated between the two texts. ROUGE-L(.,.) compares the overlap of the longest common subsequence of two texts, mainly comparing the degree of overlap at the lexical level. α is the weight of the lexical overlap, and E(.) is the text embedding function, which converts the model's output into an embedding vector. Cosine similarity is used to calculate semantic similarity. For text embedding functions, where, To convert the text y into an embedding vector, The text y_best is converted into an embedding vector for subsequent operations.

[0079] The above steps, by weighting the semantic proximity and the lexical overlap, can measure the overall similarity between texts from both lexical and semantic perspectives.

[0080] Rewards for constructing candidate answers from inputs based on the query question and documents in a set of noisy documents. for:

[0081] ,

[0082] ,

[0083] ,

[0084] In the formula, and These are the weighting coefficients for the refusal reward and the hallucination penalty, respectively. as candidate answers With Refusal Template Consistency is rewarded. In a fully perturbed scenario, consistency with the ideal answer can also be called a penalty.

[0085] Step S42: Normalize the rewards of the candidate answer set and calculate the average reward.

[0086] Specifically, the reward for the candidate answer set is normalized as follows:

[0087]

[0088] Let represent the i-th candidate answer in the candidate answer set, mean() takes the mean of the candidate answer set, and std represents the standard deviation of the candidate answer set. To prevent the numerically stable term from being divided by zero, The number of samples is the number of candidate answers obtained for this query. This represents the i-th reward after normalization.

[0089] Step S43: Based on the average reward, classify the corresponding candidate answer sets to obtain the first candidate answer set with a reward less than the average reward and the second candidate answer set with a reward greater than or equal to the average reward.

[0090] Step S44: Update the parameters of the educational big model based on the principle of improving the learning degree of the educational big model on the candidate answers in the second candidate answer set and suppressing the learning degree of the educational big model on the candidate answers in the first candidate answer set.

[0091] Specifically:

[0092]

[0093] By maximizing the above objective function The model can gradually learn the optimal behavioral strategy under different interference scenarios, making its output closer to the standard answer, thus achieving a simultaneous improvement in answer quality and rejection ability.

[0094] To demonstrate the effectiveness of this solution, for example, six educational models (Qwen3, LIama-3.2, Gemma-3, Hunyuan, Phi-4, and Ours) were trained based on three scenarios. The training results are as follows:

[0095]

[0096] In the table, ROUGE-L is the longest common subsequence between the generated text and the standard answer, F1 is the harmonic mean of precision (whether the generated content is accurate) and recall (whether content is omitted), and REF represents the rejection rate.

[0097] In summary, the proposed method achieved the highest semantic quality scores in both non-interference and semi-interference answerable scenarios, and obtained a near-perfect rejection detection score in the fully interference scenario. This demonstrates that the proposed progressive interference construction and two-stage alignment framework can form a more stable behavioral boundary between answer quality and safe rejection.

[0098] In the absence of interference, our proposed method significantly outperforms the baseline models in terms of accuracy (ACC=69.76%), completeness (COM=87.41%), and context utilization (UTL=73.51%). It also achieves the highest score in prescriptiveness (NAC=52.81%), indicating that the alignment of self-distillation-driven supervision signals with subsequent reinforcement learning can effectively improve the model's understanding of evidence and the reliability of its representation.

[0099] In the semi-interference scenario, when noisy documents with semantic similarity to the question but no substantial relevance are introduced into the retrieval context, the performance of all baseline models degrades to varying degrees, particularly in accuracy and prescriptiveness (e.g., Qwen3's NAC drops from 51.39% to 43.71%). In contrast, our proposed method not only maintains its leading position in accuracy (ACC=65.86%), completeness (COM=79.05%), and context utilization (UTL=72.59%) in this scenario, but also achieves a prescriptiveness NAC of 56.13%, higher than other baseline models and showing a significant improvement over the base model. Experimental results demonstrate that the noise-resistant generation task, by aligning the mixed noisy input with the ideal response in the interference-free scenario, enables the model to learn stable information filtering and evidence focusing capabilities.

[0100] In a fully interfering scenario, when the context consists entirely of irrelevant noise, experimental results show that the proposed method achieves a rejection detection score of 98.82%, which is significantly better than baseline models such as Llama-3.2 (50.70%) and Gemma-3 (65.20%). It also achieves a further improvement compared to the Qwen3 base model. The model can trigger the rejection strategy more stably when evidence is missing, thereby reducing the potential risks of unfounded generation.

[0101] The present invention discloses a non-interference generation and alignment device for suppressing hallucinations in educational large models, comprising a memory and a controller connected in sequence. The memory stores a computer program, and the controller is used to read the computer program and execute the non-interference generation and alignment method for suppressing hallucinations in educational large models as described in the first aspect.

[0102] Specifically, the operating principle of the device in the second aspect is detailed in the first aspect and will not be repeated here.

[0103] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A robust generative alignment method for suppressing hallucinations in a large-scale educational model, characterized in that, Includes the following steps: Construct a noise-free document set, a full-noise document set, and a hybrid document set. In the noise-free document set, the context of all documents is more relevant to the educational question than a preset threshold. In the full-noise document set, the context of all documents is semantically similar to the educational question but not actually related to it. In the hybrid document set, the context of some documents is more relevant to the educational question than a preset threshold, and the context of the remaining documents is semantically similar to the educational question but not actually related to it. The input to the education big data model is constructed based on the query question and the documents in the noise-free document set, and the ideal answer is selected from the answers of the education big data model according to the quality score. Based on the query question and documents in the noise-free document set, mixed document set, and full-noise document set, an educational big model is constructed as input. The educational big model is trained by combining ideal answer and rejection answer templates so that the educational big model can accurately identify the correct document content and answer more closely to the ideal answer when there is valid evidence, and reject the answer when there is only invalid evidence. Based on the query question and documents in the noise-free document set, mixed document set, and full-noise document set, an educational big data model is constructed. The reward of each candidate answer in the candidate answer set corresponding to each document is calculated by combining the reward function corresponding to each document set. The average reward of each candidate answer set is calculated. The parameters of the educational big data model are updated based on the principle of improving the learning degree of the educational big data model for candidate answers with rewards greater than the average reward and suppressing the learning degree of the educational big data model for candidate answers with rewards less than the average reward.

2. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 1, characterized in that... The hybrid document set is only one, and the proportion of documents in the hybrid document set whose context is more relevant to the educational issue than a preset threshold is 50%.

3. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 1, characterized in that... The process of constructing an educational big data model based on the query question and documents in the noise-free document set, and selecting ideal answers from the answers of the educational big data model based on quality scores, includes: The query q is combined with each document in the noise-free document set to form the input content, which is then input into the education big data model to generate the corresponding answer. The OmniEval model evaluator scores each candidate answer in the candidate answer set based on four equally weighted dimensions: accuracy, completeness, utilization, and standardization, to obtain a quality score for each answer. The answer with the highest quality score is selected as the ideal answer.

4. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 1, characterized in that... The educational big data model is constructed based on the query question and documents from noise-free, mixed, and fully noisy document sets. This model is then trained using ideal and rejection templates to ensure accurate identification of valid document content when valid evidence exists. Answers that are closer to the ideal answer are rejected when only invalid evidence exists, including: The educational big data model is constructed based on the query question and documents in a noise-free document set as input, and trained by combining ideal answers. This allows the educational big data model to strengthen model generation and directly generate ideal answers when only valid evidence exists. Furthermore, the educational big data model is constructed based on the query question and documents in a noise-free document set, a mixed document set, and a fully noisy document set as input, and trained by combining ideal answer and rejection answer templates. This allows the educational big data model to strengthen the model's ideal answer when only valid evidence exists, eliminate interfering answers that are closer to the ideal answer when invalid evidence exists, and reject answers when only invalid evidence exists. Based on the query question and documents in the noisy document set, an educational big data model is constructed as input and trained using a rejection template, so that the educational big data model learns to reject an answer when all evidence is invalid.

5. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 1, characterized in that... The process involves constructing an educational big data model based on the query question and documents from noise-free, mixed, and fully noisy document sets. This model is then used as input, and the reward for each candidate answer in the corresponding candidate answer set is calculated using the reward function for each document set. The average reward for each candidate answer set is calculated. The parameters of the educational big data model are updated based on the principle of increasing the learning degree of the educational big data model for candidate answers with rewards greater than the average reward and suppressing the learning degree of the educational big data model for candidate answers with rewards less than the average reward. This includes: Based on the query question and documents from noise-free document sets, mixed document sets, and fully noisy document sets, an educational big data model is constructed as input. The reward for each candidate answer in each document candidate answer set is calculated using the reward function corresponding to each document. Specifically, the reward for the candidate answer corresponding to the query question and documents constructed from the noise-free document set is determined based on a general reward, calculated based on accuracy, completeness, utilization, and standardization. The reward for the candidate answer corresponding to the query question and documents constructed from the mixed document set is determined based on a general reward and a consistency reward with the ideal answer, calculated based on the overlap at the lexical level and the closeness at the semantic level. The reward for the candidate answer corresponding to the query question and documents constructed from the fully noisy document set is determined based on a consistency reward with the rejection template and a consistency reward with the ideal answer. Normalize the rewards for the candidate response set and calculate the average reward; Based on the average reward, the corresponding candidate answer sets are classified to obtain the first candidate answer set with a reward less than the average reward and the second candidate answer set with a reward greater than or equal to the average reward; The parameters of the educational big model are updated based on the principle of improving the learning degree of the educational big model on the candidate answers in the second candidate answer set and suppressing the learning degree of the educational big model on the candidate answers in the first candidate answer set.

6. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 5, characterized in that... Rewards are given for constructing candidate answers to the input based on the query question and documents in a noise-free document set. for: , , In the formula, y represents the candidate answer. The standard answer to the query question, This is a collection of documents in a noise-free document set. , , , The weights are respectively for accuracy, completeness, utilization, and standardization.

7. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 5, characterized in that... Rewards are given for candidate answers corresponding to the input of the query question and the documents constructed from the mixed document set. for: = + , , In the formula, As candidate answers, The standard answer to the query question, For the ideal answer to the query question, For documents in a mixed document set, The weighting coefficient for consistency rewards. for This is a reward for consistency between the candidate answers and the ideal answer output by the large-scale education model. To compare the overlap of the longest common subsequence of two texts, The weighting of the degree of overlap at the lexical level. This is a text embedding function that converts the output of the large-scale education model into an embedding vector. This is a text embedding function.

8. The anti-interference generation alignment method for hallucination suppression in educational large-scale models according to claim 5, characterized in that... Rewards are given for constructing candidate answers to the input based on the query question and documents in a set of noisy documents. for: , , , In the formula, y represents the candidate answer. and Here are the weighting coefficients for the refusal reward and the hallucination penalty, respectively. as candidate answers With Refusal Template The consistency reward is calculated as follows: α is the weight of the overlap at the lexical level, and ROUGE-L(.,.) is used to compare the overlap of the longest common subsequence of two texts. For cosine similarity, For text embedding functions, as candidate answers With the ideal answer Consistency rewards.

9. An anti-interference generation alignment device for suppressing hallucinations in educational large-scale models, comprising a memory and a controller connected in sequence, wherein the memory stores a computer program, characterized in that: The controller is used to read the computer program and execute the anti-interference generation alignment method for hallucination suppression in educational large-scale models as described in any one of claims 1-6.