Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

39 results about "Factoid" patented technology

A factoid is either a false statement presented as a fact, or a true but brief or trivial item of news or information. The term was coined in 1973 by American writer Norman Mailer to mean a piece of information that becomes accepted as a fact even though it is not actually true, or an invented fact believed to be true because it appears in print. Since its creation in 1973, the term has evolved, now often being used to describe a brief or trivial item of news or information.

Generated text quality processing method based on large language model

The invention relates to the technical field of natural language processing, and discloses a generated text quality processing method and system based on a large language model, and the method comprises the steps: constructing a logic topological graph and a Laplacian matrix of an original generated text, and extracting a feature value sequence; recognizing a logic bearing wall based on the attention gradient and generating an anti-fact contrast text; calculating a logic collapse index by combining the difference between the map characteristic values of the original text and the contrast text and the logic polarity overturning condition; and judging the quality of the generated text according to the logic collapse index, blocking the text which does not pass the judgment, and shaping and resampling the output probability value of the error position by utilizing the spectrum difference information to generate a new text. The method does not need to depend on an external knowledge base, quantifies the stability of text logic through anti-fact interference and spectrum analysis, and identifies a high-risk illusion text; and a wrong logic path is automatically corrected through Logits shaping, so that the logic self-consistency of the generated content is improved on the premise of ensuring the semantic smoothness of the text language.
Owner:SHENZHEN HAIYUNAN NETWORK SECURITY TECH CO LTD

Multi-modal rumor detection method based on anti-factual reasoning and causal intervention

The invention discloses a multi-modal rumor detection method fusing texts, images and social propagation structures, and belongs to the technical field of natural language processing, computer vision and causal reasoning. Specifically, the invention provides a unified causal inference framework, and hybrid deviation in multi-modal data is effectively stripped by integrating text anti-fact causal inference and an image dot product causal intervention mechanism. Under the framework, social propagation structure features are further fused, and a multi-head collaborative attention mechanism is adopted, so that deep alignment and semantic enhancement in cross-modal features are realized. Adversarial samples are generated through projection gradient descent for adversarial training, and model parameters are optimized in combination with anti-fact loss, so that the classification accuracy and generalization ability of the model are improved. According to the rumor detection method, a causal reasoning normal form is introduced into a rumor detection task, the effectiveness of an anti-fact and intervention mechanism in a complex information scene is verified, and a new theoretical support and method path are provided for constructing a credible multi-modal information system.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Retrieval evidence enhancement-based interpretable false news detection method and system

The invention discloses an interpretable false news detection method and system based on retrieval evidence enhancement, the interpretable false news detection method combines a big language model with an external knowledge base retrieval mechanism, that is, factual supplementary evidence is provided for the big language model to generate final reasoning through a retrieval module. Screening candidate text evidences by utilizing a clustering algorithm, and performing credibility scoring and filtering on the candidate evidences in combination with a large language model; a unified reasoning Prompt is constructed to input a plurality of large language models, true and false judgment is output respectively, and detailed explanatory texts are generated with the assistance of natural language explanation, so that the false news detection method with factual support, language expression and user understandability is realized. And the judgment results of the models are fused through a majority voting mechanism, and explanation is generated from multiple perspectives, so that the risk caused by reasoning deviation of a single model is reduced, and the accuracy and stability of system output are effectively improved.
Owner:HANGZHOU NORMAL UNIVERSITY

Fine-tuning language models for reasoning with counterfactual feedback

Example solutions for fine-tuning a language model include: generating a dataset that includes a plurality of paired samples, each paired sample of the plurality of paired samples includes (i) a factual question and a true outcome for that factual question and (ii) a counterfactual question and a true outcome for that counterfactual question; submitting a factual query to an answer model, the factual query including the factual question and the true outcome of the factual question, the answer model generating a factual answer in response to the factual query; submitting a counterfactual query to the answer model, the counterfactual query including the counterfactual question and the true outcome of the counterfactual question, the answer model generating a counterfactual answer in response to the counterfactual query; and performing fine-tuning on a target model using at least the factual question paired with factual answer and the counterfactual question paired with counterfactual answer.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intelligent document question answering and knowledge base retrieval method and system based on large model

The invention relates to the technical field of intelligent retrieval, in particular to an intelligent document question answering and knowledge base retrieval method and system based on a large model. The method comprises the following steps: acquiring an initial query and performing semantic analysis, generating pseudo-fact fragments containing positive, negative and neutral contexts by using a first large language model in combination with a domain entity word list, and splicing the pseudo-fact fragments into a fused imaginary document; respectively vectorizing the query document and the imaginary document, carrying out weighted fusion according to a query-centrality coefficient to generate a retrieval vector, and retrieving in a knowledge base to obtain a candidate document set; generating a draft answer by using a second large language model, performing traceability comparison on each statement, and marking non-traceability or fact conflict items; and taking the draft and the mark as a correction instruction to be input into the model for correction, and generating a final answer. According to the scheme, the retrieval precision can be improved through multi-view document fusion, illusion is restrained through a closed-loop verification correction mechanism, and evidence reliability is ensured.
Owner:XIAN MINGFU CLOUD COMPUTING CO LTD

Atomic fact decomposition-based post-attribution question and answer method and system

The invention relates to the technical field of natural language processing and information retrieval, and particularly discloses a postmortem attribution question answering method and system based on atomic fact decomposition. Firstly, a user question is input into a large language model to generate an initial long answer; decomposing the long answer into molecular clauses and atomic facts through an instruction fine tuning model; performing dual-stage evidence retrieval and screening based on atomic facts; a verifier is used for conducting ternary judgment on the evidence and the atomic facts, and supporting, editing or re-retrieval operation is executed according to the judgment result; and the revised atomic facts are traced back to an original answer structure, a credible long answer is obtained through recombination, and a complete attribution report is generated. According to the method, through atomic-scale fine-grained decomposition and editing, high-precision attribution and minimization of intention disturbance of the answers are realized, and the credibility and continuity of the answers are improved. Meanwhile, evaluation indexes such as the attribute recall rate and the attribute accuracy rate are provided, and the accuracy and integrity of evidence support are effectively measured.
Owner:SHANXI UNIV

Large language model-based criminal name prediction method and device for easily-confused criminal affairs, and terminal

The invention relates to the field of artificial intelligence, and provides an easily-confused criminal name prediction method and device based on a big language model, and a terminal. The method comprises the following steps: obtaining a fact description of a target case, and determining a plurality of legal provisions related to the target case; generating a crime name prediction task instruction based on the fact description and the plurality of legal provisions, and inputting the crime name prediction task instruction into the large language model to obtain a predicted crime name of the target case output by the large language model; inputting a target element question set corresponding to the predicted crime name into a large language model under the condition that the predicted crime name is determined to be the easily confused crime name, and obtaining a target element answer set of the target element question set output by the large language model; and adjusting the predicted crime name based on the target element answer set, and determining a final predicted crime name of the target case. According to the big language model-based easily-confused criminal name prediction method provided by the invention, easily-confused criminal names can be distinguished, and the accuracy, reliability and interpretability of criminal name prediction are improved.
Owner:CHINA MOBILE COMM LTD RES INST +1

Regulation and regulation question and answer illusion suppression fine tuning method based on anti-fact negative sample

The invention belongs to the technical field of artificial intelligence and natural language processing, and particularly relates to a rule and regulation question and answer illusion suppression fine tuning method based on an anti-fact negative sample. According to the method, three types of anti-fact negative samples of error units, error thresholds and error versions are constructed, and cooperative training is carried out in combination with boundary maximization loss of evidence perception, numerical value consistency regularization, head rejection and temperature calibration and a controlled generation mechanism, so that the discrimination capability of the model on high-similarity error information is effectively improved; the accuracy of numerical answers is ensured, and the answers can be reliably rejected when evidences are insufficient or unreliable, so that the illusion phenomenon of a large language model in professional questions and answers is remarkably inhibited.
Owner:GUANGZHOU CITY UNIV OF TECH

Minimum evidence span alignment and accurate reference generation method and system

The invention relates to the technical field of information retrieval and natural language processing, and particularly discloses a minimum evidence span alignment and accurate reference generation method and system. According to the method, candidate terms are obtained through direct numbering and semantic recall, Span-level evidence alignment is achieved through fusion of sequence labeling and interpretable attribution, the boundary robustness is optimized in combination with anti-fact boundary learning, closed-loop mending is triggered based on the minimum span coverage rate, and finally checkable reference with a standardized RefTag is output. According to the method, the problems of coarse evidence granularity, boundary drift and unreviewable reference of a traditional method are effectively solved, the generation illusion is remarkably reduced, and the auditing traceability is improved.
Owner:GUANGZHOU CITY UNIV OF TECH

Intelligent question answering and action execution method, system and equipment based on intention recognition, medium and product

The invention provides an intelligent question answering and action execution method, system and device based on intention recognition, a medium and a product, an automatic fact checking step is executed on initial reply content, then a fact completeness score is generated, and based on the fact completeness score, a part with a low fact completeness score is intercepted and converted into manual intervention, so that the success rate of question answering and action execution is improved. The part with the high fact integrity score is returned to the user, so that the accuracy of the reply fact is improved, and the problem that the semantic consistency of the generated reply and the original fact source cannot be ensured due to the lack of a mandatory real-time after-event verification mechanism in the prior art is solved; by executing the corresponding steps of the action execution intention and the business irrelevant intention, the defect that the system function is limited to static information query is overcome, and the operation from an information interface to business handling can be realized.
Owner:BEIJING WISDOM TOOTH TECH CONSULTING CO LTD +1

A fact checking method based on multi-modal evidence fusion and discourse relation enhancement

ActiveCN119089381BMedicineFactoid
The present application belongs to the field of computer science and technology, and specifically relates to a fact verification method based on multi-modal evidence fusion and discourse relationship enhancement, comprising the following steps: step 1, evidence retrieval: for the statement to be detected, relevant evidence is retrieved from the evidence database; step 2, construction of evidence relationship graph: using a natural language processing tool, discourse relationships are extracted from the relevant evidence, a discourse relationship heterogeneous graph is constructed based on the extracted discourse relationships, and the features of the discourse relationship heterogeneous graph are fused and extracted through a graph neural network to form the evidence features of the discourse relationship heterogeneous graph; step 3, authenticity judgment: using the complete relevant evidence as a unit, the statement features and the evidence features are fused, and it is analyzed whether the relevant evidence supports the statement to be judged. The present application improves the accuracy and reliability of the multi-modal fact verification model in judging the authenticity of the statement, and can effectively judge the statement to be detected.
Owner:BEIHANG UNIV

A legal document quality adversarial inspection method based on logical game

The application discloses a kind of legal document quality confrontation test method based on logic game, it is related to legal artificial intelligence technical field, including, collection legal document full-text data, and extract fact triple, evidence citation fragment, law article number and conclusion claim, generate structured argumentation set;Based on the confidence attenuation sequence of multiple rounds of game, in combination with the vulnerability propagation coefficient corresponding to different types of edges in directed heterogenous graph, carry out path vulnerability contribution value calculation, and generate legal document global vulnerability index by stratified accumulation;Based on logic breakpoint list, contradiction conflict list and legal document global vulnerability index, comprehensive evaluation legal document's logic integrity, argumentation consistency and conclusion robustness, generate structured check report.The application provides a complete and feasible implementation path for the automatic check of legal documents combined with natural language processing.
Owner:SUYUAN TECHNOLOGY (HUNAN) CO LTD

Method and system for generating longform technical question and answer dataset

Conventional Question and Answer (QA) datasets are created for generating factoid questions only and the present disclosure generates longform technical QA dataset from textbooks. Initially, the system receives a technical textbook document and extracts a plurality of contexts. Further, a first plurality of questions are generated based on the plurality of contexts. A plurality of answerable questions are generated further based on the plurality of contexts using an unsupervised template-based matching technique. Further, a combined plurality of questions are generated by combining the first plurality of questions and the plurality of answerable questions. Further, an answer for the combined plurality of questions are generated using an autoregressive language model and a mapping score is computed. Further, a plurality of optimal answers are selected based on the corresponding mapping score. Finally, a longform technical question and answer dataset is generated based on the combined plurality of questions and optimal answers.
Owner:TATA CONSULTANCY SERVICES LTD

Automatic summarization method and system for legal documents based on semantic segmentation

The application discloses a kind of legal document automatic abstract method and system based on semantic segmentation, belong to natural language processing field.The application obtains civil first instance adjudication document as input, using the method of continuous sentence classification, the semantic segmentation is carried out to adjudication document, and the text paragraph of five parts, such as dispute category, plaintiff claim, defendant statement, fact and reason, judgment basis, judgment main text and tail portion is divided into;Using the method of generative text abstract is obtained abstract for each text paragraph after segmentation;For the abstract generated by each segmented paragraph of the same adjudication document, the final result is obtained by splicing in order.The application carries out automatic abstract to legal document, using the method of semantic segmentation, shortens the text length of single input generation abstract model, and can retain complete original text semantic structure features.
Owner:ZHEJIANG UNIV

A method for movie and television role playing based on sentiment retrieval and role consistency control

The application discloses a film and television role playing method based on emotional retrieval and role consistency control, and relates to the technical field of artificial intelligence and natural language processing, which aims at the problems of role behavior deviating from the setting, lack of emotional response resonance and inconsistent retrieval content style in the prior art, constructs a role memory knowledge graph and calculates semantic and emotional vectors, carries out emotional retrieval according to user query, modifies the preliminary response after atomic fact analysis and consistency verification, and finally outputs a natural language reply conforming to the role setting and having emotional resonance, and is mainly used for improving the reality and immersion of film and television role interaction.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Fine-tuning language models for reasoning with counterfactual feedback

Example solutions for fine-tuning a language model include: generating a dataset that includes a plurality of paired samples, each paired sample of the plurality of paired samples includes (i) a factual question and a true outcome for that factual question and (ii) a counterfactual question and a true outcome for that counterfactual question; submitting a factual query to an answer model, the factual query including the factual question and the true outcome of the factual question, the answer model generating a factual answer in response to the factual query; submitting a counterfactual query to the answer model, the counterfactual query including the counterfactual question and the true outcome of the counterfactual question, the answer model generating a counterfactual answer in response to the counterfactual query; and performing fine-tuning on a target model using at least the factual question paired with factual answer and the counterfactual question paired with counterfactual answer.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method and system for automatically generating Wiki-style article

The invention belongs to the technical field of natural language processing and artificial intelligence, and particularly relates to a method and system for automatically generating a Wiki-style article. The invention aims to solve the problems of logic fragmentation, insufficient fact accuracy, limited content coverage, poor format normalization and the like in the existing automatic text generation. An outline with a strict structure is constructed by combining internal knowledge and external references obtained through multi-view retrieval; generating a preliminary article according to the outline, and performing segment-by-segment refinement and fact calibration on the preliminary article by using references; and finally, carrying out comprehensive evaluation and iterative optimization on the article by adopting a multi-dimensional review mechanism to ensure that the article is logically coherent and accurate in content and completely meets the specification requirements of Wikipedia. According to the method, through multi-source knowledge fusion and multi-stage quality control, irrelevant information interference is effectively reduced, the efficiency, reliability and accuracy of automatically generating the Wiki-style article are remarkably improved, and meanwhile the highest publishing standard of the Wiki-style article can be met.
Owner:BEIJING INST OF TECH

Event authenticity determination method and system, computer equipment and storage medium

The invention discloses an event authenticity determination method and system, computer equipment and a storage medium, and relates to the technical field of natural language processing, and the method comprises the following steps: generating a derived event document of any event; and obtaining a keyword in the derived event document of the any event, and judging whether the any event is real or not according to the keyword. According to the method, by identifying the keywords, facts and misleading information can be better distinguished, and the capability of identifying false news is improved.
Owner:CHINA NAT PETROLEUM CORP +1

Data processing method and device

The embodiment of the invention provides a data processing method and device.The method comprises the steps that a to-be-verified text is obtained, and the to-be-verified text is generated through a data processing model; a question text corresponding to the to-be-verified text is generated, the to-be-verified text and the question text are input into the data processing model, a target reply generated by the data processing model is obtained, and the question text is a text which puts forward a question according to the fact basis of the to-be-verified text; determining target credibility corresponding to the to-be-verified text according to the target reply, and determining a correctness result of the to-be-verified text according to the target credibility; the data processing model carries out introspection on the output to-be-verified text through the questioning text, the target credibility of the to-be-verified text is determined according to the internal uncertainty signal of the data processing model, the determination of the target credibility is completed in the data processing model, the cost is extremely low, an external knowledge base does not need to be used for verification, and the verification efficiency is greatly improved. And insufficient coverage of an external knowledge base does not need to be worried.
Owner:ALI HEALTH TECH CO LTD

Ffact verification task-oriented data processing method, system and equipment and medium

The invention provides a fact verification task-oriented data processing method, system and device and a medium, and relates to the field of fact verification, the data processing method comprises the following steps: obtaining first language data information based on news corpus, and carrying out data preprocessing to obtain target language data information; designing a zero sample prompt template for carrying out a single round of dialogue based on target language data information, constructing a false declaration generator and a false declaration quality verifier, generating a plurality of false declarations through the false declaration generator and the false declaration quality verifier based on a preset generation target, and screening to obtain a target false declaration, a data source is provided for a fact verification task. By comparing the real news headline with the generated target false declaration, model training can be carried out to enable the model to learn differences of the real news headline and the generated target false declaration in expression styles and fact logic, so that AI generated false content is more accurately identified in practical application.
Owner:MINZU UNIVERSITY OF CHINA

Method of extracting data from written articles and associated system

A method and system for extracting data from a written article are disclosed. An initial question related to the article is provided to a first large language model (LLM) to generate one or more secondary questions. Numerical representations of the initial and secondary questions are compared against numerical representations of text chunks from the article to select a relevant subset of chunks. A context is generated from this subset, and a second LLM generates an answer to the initial question based on the context. A user interface displays the question, the generated answer, and an indication of where in the article supporting evidence for the answer is located. This allows a user to rapidly verify the factual basis of the AI-generated answer, improving the efficiency and trustworthiness of the data extraction process.
Owner:DISTILLERSR INC

Multi-modal false news detection method based on retrieval enhancement and large language model

The invention belongs to the technical field of deep learning and big data, and particularly relates to a multi-modal false news detection method based on retrieval enhancement and a big language model, which comprises the following steps: constructing an external knowledge base; obtaining to-be-detected multi-modal news data, and constructing an external corpus of the to-be-detected multi-modal news data according to the external knowledge base; extracting a core fact assertion triple set of the multi-modal news data to be detected by utilizing a pre-trained large language model; performing fact consistency comparison on the core fact assertion triple set and an external corpus to obtain a fact consistency score; performing cross-modal consistency comparison on the to-be-detected multi-modal news data to obtain a cross-modal consistency score; generating a news authenticity probability according to the fact consistency score and the cross-modal consistency score; according to the method, the latest authoritative information related to the news can be obtained in time by adopting a layered retrieval enhancement mechanism, and the capability of distinguishing emergencies and real-time news is remarkably improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A fake news oriented detection method and related device

The application provides a fake news detection method and related equipment, and belongs to the technical field of natural language processing and information authenticity detection. The method comprises the following steps: preprocessing and segmenting an original input text to obtain a sentence set and score a candidate claim set obtained; converting the sentence into a single sentence declarative expression to obtain a normalized claim, and structurally constructing the normalized claim; simultaneously performing information retrieval according to the structural construction mode to obtain a candidate evidence, performing multidimensional consistency verification on the structured claim and the candidate evidence, and scoring to obtain a total verification score; and correcting the total verification score by using the matching degree of a news site and an IP address, and outputting a verification result. The application prepositions noise processing and fact positioning, adopts a structured claim for fine-grained verification, and corrects through cross-information source matching of an IP address and a news site, so that the accuracy, interpretability and identification ability for numerical / time tampering of fake news detection are improved.
Owner:SOUTH CHINA UNIV OF TECH

Method and device for identifying fake news based on traceability reasoning, equipment and medium

The application discloses a false news identification method and device based on traceability reasoning, equipment and medium. Based on text classification, text similarity and traceability natural language generation technology, a new false news identification process is designed, which solves the problem that new news cannot be identified due to the lag of fact database update in the fact-based method to a certain extent.
Owner:10TH RES INST OF CETC

An agent-based knowledge question answering system

The application provides an agent-based knowledge question answering system, and relates to the technical field of knowledge question answering, which determines a target intent label from a plurality of preset intent labels based on received question text; takes the semantic feature vector in the API vector sub-library corresponding to the preset intent label consistent with the target intent label as an API comparison vector, and determines a plurality of API description texts strongly related to the question text based on the semantic feature vector of the question text and the API comparison vector; sends the API prompt text obtained by splicing the plurality of API description texts to the agent together with the question text to obtain an answer text corresponding to the question text; effectively controls the length of the API prompt text, excludes the interference of irrelevant or low-relevance API description texts, significantly reduces the risk that the content generated by the large language model does not match the facts or is logically self-contradictory, and improves the accuracy of the generated answer text.
Owner:MOBILE TECH COMPANY CHINA TRAVELSKY HLDG

Rumor detection method based on user cognition deviation mining

The invention relates to the technical field of natural language processing and deep learning, in particular to a rumor detection method based on user cognitive deviation mining. The method comprises the following steps: performing cognitive information extraction on user comments based on a big language model, and obtaining a standing judgment for news contents and a corresponding explanatory basis; respectively extracting embedded features of the news text, the news image and the user cognitive information by means of a CLIP multi-mode model, and carrying out feature fusion; the initial standing site label is optimized and calibrated through a knowledge distillation mechanism, so that the reliability of standing site identification is improved; a triple loss function is constructed, and the deviation degree between user cognition and news authenticity is mined and measured; and realizing end-to-end rumor classification and discrimination by fusing multi-modal features and introducing cognitive deviation. According to the method, the sensitivity of the model to cognitive deviation can be enhanced, the rumor detection accuracy and robustness are improved, good interpretability is achieved, and the method is suitable for a network public opinion analysis and fact verification system.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

False news detection method and system based on fact-sentiment dual uncertainty

This invention belongs to the field of data detection and provides a method and system for detecting fake news based on dual uncertainty of fact and sentiment. The method involves acquiring the text and image of each news item from a social network; extracting features from the text and image to obtain text embeddings and image embeddings; calculating Gaussian uncertainty representations by Gaussian reweighting of the text and image embeddings respectively; filtering the Gaussian uncertainty representations based on a variational information bottleneck strategy to obtain multimodal uncertainty representations; processing the multimodal uncertainty representations to generate fact inconsistency representations; processing the multimodal uncertainty representations based on a sentiment manipulation graph convolutional network to generate sentiment inconsistency representations; fusing the fact inconsistency representations and sentiment inconsistency representations to generate fused features; and classifying based on the fused features to obtain the fake news detection classification result. This invention alleviates the limitations of cross-modal complementarity caused by multimodal data uncertainty.
Owner:SHANDONG JIAOTONG UNIV

A news synthesis method and system based on traditional algorithms and large models

The application provides a news synthesis method and system based on traditional algorithms and large models, and relates to the technical field of natural language processing and text generation. The method comprises: obtaining news texts of different data sources and preprocessing; using an improved text segmentation and sorting algorithm, performing semantic synthesis on the preprocessed news texts, generating a comprehensive news summary text and performing semantic role labeling to obtain a structured fact anchor point; constructing a structured prompt word template and configuring a low randomness parameter, and calling a large-scale language model to generate a preliminary news draft that meets the fact boundary according to the fact anchor point; performing quality detection on the news preliminary draft to obtain a comprehensive score report and a quality defect label of the news preliminary draft; using a large-scale language model to generate structured feedback of the news preliminary draft, and iteratively optimizing the news preliminary draft according to the feedback result to generate a news manuscript. The application realizes automatic generation of news with high accuracy, high consistency and clear structure.
Owner:NORTHEASTERN UNIV CHINA

Two-stage document filtering and robust fine tuning method based on graph attention network

The invention discloses a two-stage document filtering and robust fine tuning method based on a graph attention network, and relates to the technical field of natural language processing and deep learning. The problems that an existing retrieval enhancement generation system is difficult to accurately identify useful documents in a mixed document environment and is easily interfered by anti-fact information are solved. The method comprises the following steps of: constructing a multi-class document pool comprising correct documents, anti-fact documents, noise documents and irrelevant documents, and constructing a semantic graph on a paragraph level; the method comprises the following steps: sequentially filtering irrelevant documents and noise documents by adopting a two-stage graph attention network to obtain a reference document set, constructing a document discrimination training sample and a question and answer training sample based on the reference document set, and combining the document discrimination training sample and the question and answer training sample into joint fine tuning data to perform instruction fine tuning on a large language model. And the model has a document reliability discrimination capability and keeps stable output in a mixed document situation, so that the factuality and credibility of a retrieval enhancement generation system in an anti-fact attack and noise environment are improved.
Owner:CHANGCHUN UNIV OF SCI & TECH