Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Factoid" patented technology

A factoid is either a false statement presented as a fact, or a true but brief or trivial item of news or information. The term was coined in 1973 by American writer Norman Mailer to mean a piece of information that becomes accepted as a fact even though it is not actually true, or an invented fact believed to be true because it appears in print. Since its creation in 1973, the term has evolved, now often being used to describe a brief or trivial item of news or information.

Generated text quality processing method based on large language model

The invention relates to the technical field of natural language processing, and discloses a generated text quality processing method and system based on a large language model, and the method comprises the steps: constructing a logic topological graph and a Laplacian matrix of an original generated text, and extracting a feature value sequence; recognizing a logic bearing wall based on the attention gradient and generating an anti-fact contrast text; calculating a logic collapse index by combining the difference between the map characteristic values of the original text and the contrast text and the logic polarity overturning condition; and judging the quality of the generated text according to the logic collapse index, blocking the text which does not pass the judgment, and shaping and resampling the output probability value of the error position by utilizing the spectrum difference information to generate a new text. The method does not need to depend on an external knowledge base, quantifies the stability of text logic through anti-fact interference and spectrum analysis, and identifies a high-risk illusion text; and a wrong logic path is automatically corrected through Logits shaping, so that the logic self-consistency of the generated content is improved on the premise of ensuring the semantic smoothness of the text language.
Owner:SHENZHEN HAIYUNAN NETWORK SECURITY TECH CO LTD

Multi-modal rumor detection method based on anti-factual reasoning and causal intervention

The invention discloses a multi-modal rumor detection method fusing texts, images and social propagation structures, and belongs to the technical field of natural language processing, computer vision and causal reasoning. Specifically, the invention provides a unified causal inference framework, and hybrid deviation in multi-modal data is effectively stripped by integrating text anti-fact causal inference and an image dot product causal intervention mechanism. Under the framework, social propagation structure features are further fused, and a multi-head collaborative attention mechanism is adopted, so that deep alignment and semantic enhancement in cross-modal features are realized. Adversarial samples are generated through projection gradient descent for adversarial training, and model parameters are optimized in combination with anti-fact loss, so that the classification accuracy and generalization ability of the model are improved. According to the rumor detection method, a causal reasoning normal form is introduced into a rumor detection task, the effectiveness of an anti-fact and intervention mechanism in a complex information scene is verified, and a new theoretical support and method path are provided for constructing a credible multi-modal information system.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Fine-tuning language models for reasoning with counterfactual feedback

PCT designated stageWO2026072135A1Semantic analysisComputer security arrangementsPaired samplesData set
Example solutions for fine-tuning a language model include: generating a dataset that includes a plurality of paired samples, each paired sample of the plurality of paired samples includes (i) a factual question and a true outcome for that factual question and (ii) a counterfactual question and a true outcome for that counterfactual question; submitting a factual query to an answer model, the factual query including the factual question and the true outcome of the factual question, the answer model generating a factual answer in response to the factual query; submitting a counterfactual query to the answer model, the counterfactual query including the counterfactual question and the true outcome of the counterfactual question, the answer model generating a counterfactual answer in response to the counterfactual query; and performing fine-tuning on a target model using at least the factual question paired with factual answer and the counterfactual question paired with counterfactual answer.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Intelligent document question answering and knowledge base retrieval method and system based on large model

The invention relates to the technical field of intelligent retrieval, in particular to an intelligent document question answering and knowledge base retrieval method and system based on a large model. The method comprises the following steps: acquiring an initial query and performing semantic analysis, generating pseudo-fact fragments containing positive, negative and neutral contexts by using a first large language model in combination with a domain entity word list, and splicing the pseudo-fact fragments into a fused imaginary document; respectively vectorizing the query document and the imaginary document, carrying out weighted fusion according to a query-centrality coefficient to generate a retrieval vector, and retrieving in a knowledge base to obtain a candidate document set; generating a draft answer by using a second large language model, performing traceability comparison on each statement, and marking non-traceability or fact conflict items; and taking the draft and the mark as a correction instruction to be input into the model for correction, and generating a final answer. According to the scheme, the retrieval precision can be improved through multi-view document fusion, illusion is restrained through a closed-loop verification correction mechanism, and evidence reliability is ensured.
Owner:XIAN MINGFU CLOUD COMPUTING CO LTD

Atomic fact decomposition-based post-attribution question and answer method and system

The invention relates to the technical field of natural language processing and information retrieval, and particularly discloses a postmortem attribution question answering method and system based on atomic fact decomposition. Firstly, a user question is input into a large language model to generate an initial long answer; decomposing the long answer into molecular clauses and atomic facts through an instruction fine tuning model; performing dual-stage evidence retrieval and screening based on atomic facts; a verifier is used for conducting ternary judgment on the evidence and the atomic facts, and supporting, editing or re-retrieval operation is executed according to the judgment result; and the revised atomic facts are traced back to an original answer structure, a credible long answer is obtained through recombination, and a complete attribution report is generated. According to the method, through atomic-scale fine-grained decomposition and editing, high-precision attribution and minimization of intention disturbance of the answers are realized, and the credibility and continuity of the answers are improved. Meanwhile, evaluation indexes such as the attribute recall rate and the attribute accuracy rate are provided, and the accuracy and integrity of evidence support are effectively measured.
Owner:SHANXI UNIV

Regulation and regulation question and answer illusion suppression fine tuning method based on anti-fact negative sample

The invention belongs to the technical field of artificial intelligence and natural language processing, and particularly relates to a rule and regulation question and answer illusion suppression fine tuning method based on an anti-fact negative sample. According to the method, three types of anti-fact negative samples of error units, error thresholds and error versions are constructed, and cooperative training is carried out in combination with boundary maximization loss of evidence perception, numerical value consistency regularization, head rejection and temperature calibration and a controlled generation mechanism, so that the discrimination capability of the model on high-similarity error information is effectively improved; the accuracy of numerical answers is ensured, and the answers can be reliably rejected when evidences are insufficient or unreliable, so that the illusion phenomenon of a large language model in professional questions and answers is remarkably inhibited.
Owner:GUANGZHOU CITY UNIV OF TECH

Minimum evidence span alignment and accurate reference generation method and system

The invention relates to the technical field of information retrieval and natural language processing, and particularly discloses a minimum evidence span alignment and accurate reference generation method and system. According to the method, candidate terms are obtained through direct numbering and semantic recall, Span-level evidence alignment is achieved through fusion of sequence labeling and interpretable attribution, the boundary robustness is optimized in combination with anti-fact boundary learning, closed-loop mending is triggered based on the minimum span coverage rate, and finally checkable reference with a standardized RefTag is output. According to the method, the problems of coarse evidence granularity, boundary drift and unreviewable reference of a traditional method are effectively solved, the generation illusion is remarkably reduced, and the auditing traceability is improved.
Owner:GUANGZHOU CITY UNIV OF TECH

Intelligent question answering and action execution method, system and equipment based on intention recognition, medium and product

The invention provides an intelligent question answering and action execution method, system and device based on intention recognition, a medium and a product, an automatic fact checking step is executed on initial reply content, then a fact completeness score is generated, and based on the fact completeness score, a part with a low fact completeness score is intercepted and converted into manual intervention, so that the success rate of question answering and action execution is improved. The part with the high fact integrity score is returned to the user, so that the accuracy of the reply fact is improved, and the problem that the semantic consistency of the generated reply and the original fact source cannot be ensured due to the lack of a mandatory real-time after-event verification mechanism in the prior art is solved; by executing the corresponding steps of the action execution intention and the business irrelevant intention, the defect that the system function is limited to static information query is overcome, and the operation from an information interface to business handling can be realized.
Owner:BEIJING WISDOM TOOTH TECH CONSULTING CO LTD +1

A fact checking method based on multi-modal evidence fusion and discourse relation enhancement

ActiveCN119089381BMedicineFactoid
The present application belongs to the field of computer science and technology, and specifically relates to a fact verification method based on multi-modal evidence fusion and discourse relationship enhancement, comprising the following steps: step 1, evidence retrieval: for the statement to be detected, relevant evidence is retrieved from the evidence database; step 2, construction of evidence relationship graph: using a natural language processing tool, discourse relationships are extracted from the relevant evidence, a discourse relationship heterogeneous graph is constructed based on the extracted discourse relationships, and the features of the discourse relationship heterogeneous graph are fused and extracted through a graph neural network to form the evidence features of the discourse relationship heterogeneous graph; step 3, authenticity judgment: using the complete relevant evidence as a unit, the statement features and the evidence features are fused, and it is analyzed whether the relevant evidence supports the statement to be judged. The present application improves the accuracy and reliability of the multi-modal fact verification model in judging the authenticity of the statement, and can effectively judge the statement to be detected.
Owner:BEIHANG UNIV

A legal document quality adversarial inspection method based on logical game

The application discloses a kind of legal document quality confrontation test method based on logic game, it is related to legal artificial intelligence technical field, including, collection legal document full-text data, and extract fact triple, evidence citation fragment, law article number and conclusion claim, generate structured argumentation set;Based on the confidence attenuation sequence of multiple rounds of game, in combination with the vulnerability propagation coefficient corresponding to different types of edges in directed heterogenous graph, carry out path vulnerability contribution value calculation, and generate legal document global vulnerability index by stratified accumulation;Based on logic breakpoint list, contradiction conflict list and legal document global vulnerability index, comprehensive evaluation legal document's logic integrity, argumentation consistency and conclusion robustness, generate structured check report.The application provides a complete and feasible implementation path for the automatic check of legal documents combined with natural language processing.
Owner:SUYUAN TECHNOLOGY (HUNAN) CO LTD

Method and system for generating longform technical question and answer dataset

Conventional Question and Answer (QA) datasets are created for generating factoid questions only and the present disclosure generates longform technical QA dataset from textbooks. Initially, the system receives a technical textbook document and extracts a plurality of contexts. Further, a first plurality of questions are generated based on the plurality of contexts. A plurality of answerable questions are generated further based on the plurality of contexts using an unsupervised template-based matching technique. Further, a combined plurality of questions are generated by combining the first plurality of questions and the plurality of answerable questions. Further, an answer for the combined plurality of questions are generated using an autoregressive language model and a mapping score is computed. Further, a plurality of optimal answers are selected based on the corresponding mapping score. Finally, a longform technical question and answer dataset is generated based on the combined plurality of questions and optimal answers.
Owner:TATA CONSULTANCY SERVICES LTD

A method for movie and television role playing based on sentiment retrieval and role consistency control

The application discloses a film and television role playing method based on emotional retrieval and role consistency control, and relates to the technical field of artificial intelligence and natural language processing, which aims at the problems of role behavior deviating from the setting, lack of emotional response resonance and inconsistent retrieval content style in the prior art, constructs a role memory knowledge graph and calculates semantic and emotional vectors, carries out emotional retrieval according to user query, modifies the preliminary response after atomic fact analysis and consistency verification, and finally outputs a natural language reply conforming to the role setting and having emotional resonance, and is mainly used for improving the reality and immersion of film and television role interaction.
Owner:ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU

Fine-tuning language models for reasoning with counterfactual feedback

PendingUS20260087368A1Natural language translationBiological modelsPaired samplesData set
Example solutions for fine-tuning a language model include: generating a dataset that includes a plurality of paired samples, each paired sample of the plurality of paired samples includes (i) a factual question and a true outcome for that factual question and (ii) a counterfactual question and a true outcome for that counterfactual question; submitting a factual query to an answer model, the factual query including the factual question and the true outcome of the factual question, the answer model generating a factual answer in response to the factual query; submitting a counterfactual query to the answer model, the counterfactual query including the counterfactual question and the true outcome of the counterfactual question, the answer model generating a counterfactual answer in response to the counterfactual query; and performing fine-tuning on a target model using at least the factual question paired with factual answer and the counterfactual question paired with counterfactual answer.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method and system for automatically generating Wiki-style article

The invention belongs to the technical field of natural language processing and artificial intelligence, and particularly relates to a method and system for automatically generating a Wiki-style article. The invention aims to solve the problems of logic fragmentation, insufficient fact accuracy, limited content coverage, poor format normalization and the like in the existing automatic text generation. An outline with a strict structure is constructed by combining internal knowledge and external references obtained through multi-view retrieval; generating a preliminary article according to the outline, and performing segment-by-segment refinement and fact calibration on the preliminary article by using references; and finally, carrying out comprehensive evaluation and iterative optimization on the article by adopting a multi-dimensional review mechanism to ensure that the article is logically coherent and accurate in content and completely meets the specification requirements of Wikipedia. According to the method, through multi-source knowledge fusion and multi-stage quality control, irrelevant information interference is effectively reduced, the efficiency, reliability and accuracy of automatically generating the Wiki-style article are remarkably improved, and meanwhile the highest publishing standard of the Wiki-style article can be met.
Owner:BEIJING INST OF TECH

Data processing method and device

The embodiment of the invention provides a data processing method and device.The method comprises the steps that a to-be-verified text is obtained, and the to-be-verified text is generated through a data processing model; a question text corresponding to the to-be-verified text is generated, the to-be-verified text and the question text are input into the data processing model, a target reply generated by the data processing model is obtained, and the question text is a text which puts forward a question according to the fact basis of the to-be-verified text; determining target credibility corresponding to the to-be-verified text according to the target reply, and determining a correctness result of the to-be-verified text according to the target credibility; the data processing model carries out introspection on the output to-be-verified text through the questioning text, the target credibility of the to-be-verified text is determined according to the internal uncertainty signal of the data processing model, the determination of the target credibility is completed in the data processing model, the cost is extremely low, an external knowledge base does not need to be used for verification, and the verification efficiency is greatly improved. And insufficient coverage of an external knowledge base does not need to be worried.
Owner:ALI HEALTH TECH CO LTD

Method of extracting data from written articles and associated system

A method and system for extracting data from a written article are disclosed. An initial question related to the article is provided to a first large language model (LLM) to generate one or more secondary questions. Numerical representations of the initial and secondary questions are compared against numerical representations of text chunks from the article to select a relevant subset of chunks. A context is generated from this subset, and a second LLM generates an answer to the initial question based on the context. A user interface displays the question, the generated answer, and an indication of where in the article supporting evidence for the answer is located. This allows a user to rapidly verify the factual basis of the AI-generated answer, improving the efficiency and trustworthiness of the data extraction process.
Owner:DISTILLERSR INC

Multi-modal false news detection method based on retrieval enhancement and large language model

The invention belongs to the technical field of deep learning and big data, and particularly relates to a multi-modal false news detection method based on retrieval enhancement and a big language model, which comprises the following steps: constructing an external knowledge base; obtaining to-be-detected multi-modal news data, and constructing an external corpus of the to-be-detected multi-modal news data according to the external knowledge base; extracting a core fact assertion triple set of the multi-modal news data to be detected by utilizing a pre-trained large language model; performing fact consistency comparison on the core fact assertion triple set and an external corpus to obtain a fact consistency score; performing cross-modal consistency comparison on the to-be-detected multi-modal news data to obtain a cross-modal consistency score; generating a news authenticity probability according to the fact consistency score and the cross-modal consistency score; according to the method, the latest authoritative information related to the news can be obtained in time by adopting a layered retrieval enhancement mechanism, and the capability of distinguishing emergencies and real-time news is remarkably improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A fake news oriented detection method and related device

The application provides a fake news detection method and related equipment, and belongs to the technical field of natural language processing and information authenticity detection. The method comprises the following steps: preprocessing and segmenting an original input text to obtain a sentence set and score a candidate claim set obtained; converting the sentence into a single sentence declarative expression to obtain a normalized claim, and structurally constructing the normalized claim; simultaneously performing information retrieval according to the structural construction mode to obtain a candidate evidence, performing multidimensional consistency verification on the structured claim and the candidate evidence, and scoring to obtain a total verification score; and correcting the total verification score by using the matching degree of a news site and an IP address, and outputting a verification result. The application prepositions noise processing and fact positioning, adopts a structured claim for fine-grained verification, and corrects through cross-information source matching of an IP address and a news site, so that the accuracy, interpretability and identification ability for numerical / time tampering of fake news detection are improved.
Owner:SOUTH CHINA UNIV OF TECH

An agent-based knowledge question answering system

The application provides an agent-based knowledge question answering system, and relates to the technical field of knowledge question answering, which determines a target intent label from a plurality of preset intent labels based on received question text; takes the semantic feature vector in the API vector sub-library corresponding to the preset intent label consistent with the target intent label as an API comparison vector, and determines a plurality of API description texts strongly related to the question text based on the semantic feature vector of the question text and the API comparison vector; sends the API prompt text obtained by splicing the plurality of API description texts to the agent together with the question text to obtain an answer text corresponding to the question text; effectively controls the length of the API prompt text, excludes the interference of irrelevant or low-relevance API description texts, significantly reduces the risk that the content generated by the large language model does not match the facts or is logically self-contradictory, and improves the accuracy of the generated answer text.
Owner:MOBILE TECH COMPANY CHINA TRAVELSKY HLDG

Rumor detection method based on user cognition deviation mining

The invention relates to the technical field of natural language processing and deep learning, in particular to a rumor detection method based on user cognitive deviation mining. The method comprises the following steps: performing cognitive information extraction on user comments based on a big language model, and obtaining a standing judgment for news contents and a corresponding explanatory basis; respectively extracting embedded features of the news text, the news image and the user cognitive information by means of a CLIP multi-mode model, and carrying out feature fusion; the initial standing site label is optimized and calibrated through a knowledge distillation mechanism, so that the reliability of standing site identification is improved; a triple loss function is constructed, and the deviation degree between user cognition and news authenticity is mined and measured; and realizing end-to-end rumor classification and discrimination by fusing multi-modal features and introducing cognitive deviation. According to the method, the sensitivity of the model to cognitive deviation can be enhanced, the rumor detection accuracy and robustness are improved, good interpretability is achieved, and the method is suitable for a network public opinion analysis and fact verification system.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A news synthesis method and system based on traditional algorithms and large models

The application provides a news synthesis method and system based on traditional algorithms and large models, and relates to the technical field of natural language processing and text generation. The method comprises: obtaining news texts of different data sources and preprocessing; using an improved text segmentation and sorting algorithm, performing semantic synthesis on the preprocessed news texts, generating a comprehensive news summary text and performing semantic role labeling to obtain a structured fact anchor point; constructing a structured prompt word template and configuring a low randomness parameter, and calling a large-scale language model to generate a preliminary news draft that meets the fact boundary according to the fact anchor point; performing quality detection on the news preliminary draft to obtain a comprehensive score report and a quality defect label of the news preliminary draft; using a large-scale language model to generate structured feedback of the news preliminary draft, and iteratively optimizing the news preliminary draft according to the feedback result to generate a news manuscript. The application realizes automatic generation of news with high accuracy, high consistency and clear structure.
Owner:NORTHEASTERN UNIV CHINA

Two-stage document filtering and robust fine tuning method based on graph attention network

The invention discloses a two-stage document filtering and robust fine tuning method based on a graph attention network, and relates to the technical field of natural language processing and deep learning. The problems that an existing retrieval enhancement generation system is difficult to accurately identify useful documents in a mixed document environment and is easily interfered by anti-fact information are solved. The method comprises the following steps of: constructing a multi-class document pool comprising correct documents, anti-fact documents, noise documents and irrelevant documents, and constructing a semantic graph on a paragraph level; the method comprises the following steps: sequentially filtering irrelevant documents and noise documents by adopting a two-stage graph attention network to obtain a reference document set, constructing a document discrimination training sample and a question and answer training sample based on the reference document set, and combining the document discrimination training sample and the question and answer training sample into joint fine tuning data to perform instruction fine tuning on a large language model. And the model has a document reliability discrimination capability and keeps stable output in a mixed document situation, so that the factuality and credibility of a retrieval enhancement generation system in an anti-fact attack and noise environment are improved.
Owner:CHANGCHUN UNIV OF SCI & TECH

A Two-Stage Document Filtering and Robust Fine-Tuning Method Based on Graph Attention Networks

ActiveCN121997918BLinguistic modelFactoid
A two-stage document filtering and robust fine-tuning method based on graph attention networks is proposed, relating to the fields of natural language processing and deep learning. It addresses the problems of existing retrieval augmentation generation systems struggling to accurately identify useful documents in mixed document environments and being susceptible to counterfactual information interference. This method constructs a multi-class document pool including correct documents, counterfactual documents, noisy documents, and irrelevant documents, and builds a semantic graph at the paragraph level. A two-stage graph attention network is used to sequentially filter irrelevant and noisy documents to obtain a reference document set. Based on this reference document set, document discrimination training samples and question-answering training samples are constructed. These two sets are combined into joint fine-tuning data to fine-tune a large language model, enabling the model to possess document reliability discrimination capabilities and maintain robust output in mixed document scenarios. This improves the factual accuracy and credibility of the retrieval augmentation generation system under counterfactual attacks and noisy environments.
Owner:CHANGCHUN UNIV OF SCI & TECH

Agent performance verification method, device and product based on structured information embedding

The application relates to the technical field of artificial intelligence, and proposes an intelligent agent performance verification method based on structured information embedding, an electronic device and a computer program product. The method comprises the following steps: converting each truth record in a structured truth record set into a corresponding unstructured sentence, and the truth record set is generated according to a predefined data structure; inserting each unstructured sentence into a background text stream to form a test text stream; inputting the test text stream and an aggregation query instruction generated based on the data structure into a to-be-tested intelligent agent, and obtaining a natural language answer output by the to-be-tested intelligent agent; converting the natural language answer into a structured answer record set; and determining a performance verification result of the to-be-tested intelligent agent by comparing the truth record set with the answer record set. The method can verify whether the intelligent agent can restore discrete distributed fact information from the test text stream without omission, so as to verify the aggregation completeness of the intelligent agent for key information.
Owner:HANGZHOU HIGH-TECH ZONE (BINJIANG) INSTITUTE OF BLOCKCHAIN & DATA SECURITY

Intelligent document question answering and knowledge base retrieval method and system based on large model

The present application relates to the technical field of intelligent retrieval, in particular to an intelligent document question answering and knowledge base retrieval method and system based on a large model. The method comprises the following steps: obtaining an initial query and performing semantic analysis, combining a domain entity word table to generate pseudo-fact segments containing positive, negative and neutral contexts by using a first large language model, and splicing the pseudo-fact segments into a fusion fictional document; vectorizing the query and the fictional document respectively, generating a retrieval vector by weighted fusion according to a query-center degree coefficient, and retrieving a candidate document set in a knowledge base; generating a draft answer by using a second large language model, performing trace comparison and marking items that cannot be traced or have fact conflicts for each statement, inputting the draft together with the marks as a correction instruction into a model to perform correction, and generating a final answer. That is, the scheme of the present application can improve retrieval accuracy by multi-perspective document fusion, and can inhibit hallucinations by using a closed-loop verification and correction mechanism to ensure reliable evidence.
Owner:XIAN MINGFU CLOUD COMPUTING CO LTD

Multi-type document question and answer pair generation method based on PaddleOCR and large model cue word engineering

The invention discloses a multi-type document question-answer pair generation method based on PaddleOCR and large model cue word engineering, and the method comprises the following steps: employing an OCR technology to recognize multi-type document contents, and obtaining a unified structured text, the multi-type document contents comprising cross-type documents in the field of power operation management; processing the unified structured text by adopting a semantic partitioning method to obtain standardized text blocks; based on the standardized text block, using a large language model to extract atomic facts to obtain an atomic fact set; based on the atomic fact set, performing subject classification by adopting a large language model to obtain a hierarchical knowledge structure; based on the atomic fact set and the hierarchical knowledge structure, adopting a large language model to generate question and answer pairs to obtain structured question and answer data; and based on the structured question and answer data, processing by adopting a quality verification method to obtain a high-quality question and answer pair set.
Owner:INFORMATION & COMM CO OF STATE GRID SHAANXI ELECTRIC POWER CO LTD

News synthesis method and system based on traditional algorithm and large model

The invention provides a news synthesis method and system based on a traditional algorithm and a large model, and relates to the technical field of natural language processing and text generation. The method comprises the steps that news texts of different data sources are acquired and preprocessed; carrying out semantic synthesis on the preprocessed news text by adopting an improved text segmentation and sorting algorithm, generating a comprehensive news abstract text, and carrying out semantic role labeling to obtain a structured fact anchor point; constructing a structured cue word template and configuring low-randomness parameters, and generating first draft news meeting a fact boundary according to the fact anchor point by calling a large-scale language model; performing quality detection on the first news draft to obtain a comprehensive score report and a quality defect label of the first news draft; and utilizing the large-scale language model to generate structured feedback of the first news manuscript, and performing iterative optimization on the first news manuscript according to a feedback result to generate the news manuscript. According to the method, automatic news generation with high accuracy, high consistency and clear structure is realized.
Owner:NORTHEASTERN UNIV CHINA

Large language model watermark embedding and detecting method based on attention head perception

The invention discloses a large language model watermark embedding and detecting method based on attention head perception, and relates to the technical field of natural language processing and digital watermarking. The method comprises the following four core steps of: firstly, screening an attention head set which plays a key role in factual information coding through a'pure-pollution-recovery 'three-stage causal intervention experiment; secondly, a binary linear fact detector is trained based on the set, and accurate distinguishing of token factuality and non-factuality is achieved; then, in a model reasoning stage, a dynamic green list is only generated for non-factual tokens, watermark hidden embedding is completed through logits bias intervention, and bias intensity is adaptively adjusted along with text fact density; and finally, in a detection stage, reproducing an embedding process to screen non-factual tokens, and quantifying an abnormal degree by counting a matching ratio and a z score to realize accurate detection of the watermark. The text factual accuracy is prevented from being damaged through the attention head perception mechanism, the watermark robustness and concealment are improved through the dynamic green list design, the method is suitable for scenes such as copyright tracing and authenticity verification of the text generated by the large language model, and the engineering feasibility is high.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Method and system for fact determination

A fact determination method is provided, the method comprising receiving target text, generating one or more question prompts using a prompt generation model to induce extraction of information associated with the target text, obtaining answers to the respective question prompts by inputting the question prompts into a language model and outputting a result of determining whether the target text is factual using the language model.
Owner:SAMSUNG SDS CO LTD