Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

5 results about "Document filtering" patented technology

Document filters are responsible for: reading an original document in from a file in a specific format (e.g. different filters exist for handling plain text and OpenDocument/OpenOffice); extracting translatable content from a document;

A code security scanning method

PendingCN122333483AData setAlgorithm
This invention discloses a code security scanning method, comprising: project file filtering: filtering project files in a source code repository according to a preset filtering strategy to generate a set of files to be scanned; scan scheduling: scheduling a first scanning tool and / or a second scanning tool to perform code security scanning on the set of files to be scanned according to the scan mode selected by the user, and obtaining a set of scan alert results, wherein the first scanning tool is a static scanning tool based on regular expressions or abstract syntax trees, and the second scanning tool is a deep scanning tool based on graph analysis; result standardization: deduplicating the scan alert result set to obtain a standardized scan alert dataset; report generation: generating a standardized scan report based on the standardized scan alert dataset. By scheduling the corresponding scanning tool to perform code security scanning on the files to be scanned according to the scan mode selected by the user, multiple scanning modes are provided to improve scanning accuracy.
Owner:深圳市和讯华谷信息技术有限公司

Method, system and medium for multi-modal document re-ranking based on answer quality driving

The application discloses a multi-modal document reordering method and system based on answer quality driving and a medium, relates to the technical field of artificial intelligence, and inputs a query question into a rearrangement large model to output a document rearrangement result; the training process of the rearrangement large model is as follows: an initial candidate document set corresponding to the query question is acquired through a multi-modal retriever and is taken as the input of a rearranger, and through screening and rearrangement, a rearranged document subset and a thinking process are generated; the document subset and the thinking process are input into a sandbox, the sandbox is internally deployed with a frozen parameter answer model, is used for parallel processing of multiple different rearrangement results of the same question, and generates corresponding multiple answers; a format consistency reward is set by using the output of the rearranger, a document filtering reward is set by using the rearranged document subset, and an answer quality reward is set by using the similarity between the answers and standard answers, so as to optimize the parameters of the rearranger; and the reordering method improves the document sorting quality and the question and answer accuracy.
Owner:ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)

A Two-Stage Document Filtering and Robust Fine-Tuning Method Based on Graph Attention Networks

ActiveCN121997918BLinguistic modelFactoid
A two-stage document filtering and robust fine-tuning method based on graph attention networks is proposed, relating to the fields of natural language processing and deep learning. It addresses the problems of existing retrieval augmentation generation systems struggling to accurately identify useful documents in mixed document environments and being susceptible to counterfactual information interference. This method constructs a multi-class document pool including correct documents, counterfactual documents, noisy documents, and irrelevant documents, and builds a semantic graph at the paragraph level. A two-stage graph attention network is used to sequentially filter irrelevant and noisy documents to obtain a reference document set. Based on this reference document set, document discrimination training samples and question-answering training samples are constructed. These two sets are combined into joint fine-tuning data to fine-tune a large language model, enabling the model to possess document reliability discrimination capabilities and maintain robust output in mixed document scenarios. This improves the factual accuracy and credibility of the retrieval augmentation generation system under counterfactual attacks and noisy environments.
Owner:CHANGCHUN UNIV OF SCI & TECH

Method, device and electronic equipment for acquiring motion information

PendingCN122285875Aimprove accuracySolve technical problems with low accuracyThe InternetEngineering
This application discloses a method, apparatus, and electronic device for acquiring motion information. Relating to the field of artificial intelligence, the method includes: querying M source documents in a database and generating N query statements based on the requirements for acquiring target motion information; performing query operations on the M source documents using each query statement to obtain a set of N source documents; performing tag matching operations on each source document in the N source document sets to obtain tag matching results, and performing document filtering operations on the N source document sets based on the tag matching results to obtain P target documents; splitting the P target documents according to the document structure of each target document to obtain multiple text blocks, and identifying the multiple text blocks as motion information of the target motion. This application solves the problem of low accuracy in determining motion information based on knowledge documents, literature, and other information collected from the internet in related technologies.
Owner:BEIJING CALORIE INFORMATION TECH CO LTD