Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

251 results about "Source document" patented technology

A source document is a document in which data collected for a clinical trial is first recorded. This data is usually later entered in the case report form. The International Conference on Harmonisation of Technical Requirements for Registration of Pharmaceuticals for Human Use (ICH-GCP) guidelines define source documents as "original documents, data, and records." Source documents contain source data, which is defined as "all information in original records and certified copies of original records of clinical findings, observations, or other activities in a clinical trial necessary for the reconstruction and evaluation of the trial."

Large model-based standard document automatic generation and multi-dimensional auditing method and system

The invention provides a standard document automatic generation and multi-dimensional auditing method and system based on a large model, and relates to the technical field of artificial intelligence, and the method comprises the steps: 1, building a distributed database of a multi-source document, and analyzing a heterogeneous text through natural language processing to obtain a standardized knowledge network; step 2, extracting index elements based on the standardized knowledge network, and forming a structured parameter library through verification and verification; and step 3, based on the structured parameter library, constructing a template library, analyzing user demands in combination with semantic matching, and automatically generating a standard document outline. The document generation efficiency and quality are improved, the manual auditing cost is reduced, and the auditing comprehensiveness and accuracy are enhanced.
Owner:浙江金汇数字技术有限公司

Methods and systems for retrieval-augmented generation using synthetic question embeddings

Methods and systems for retrieval-augmented generation are described. Responsive to a user input, an input embedding associated with the user input is obtained. A synthetic question embedding is retrieved from an embeddings database, based on a similarity to the input embedding. The synthetic question embedding is used to obtain a relevant source text based on a stored mapping between the synthetic question embedding and the source text. A prompt is provided to a large language model (LLM) to generate and display a textual response to the user input, based on the user input and the source text. The disclosed methods and systems effectively narrow the pool of source documents based on similarity measures between the user input embedding and the synthetic question embedding, to enable the retrieval of more relevant sources for use in response generation.
Owner:SHOPIFY INC

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Retrieval enhancement generated document screening system and method fusing verification mechanism

The invention discloses a retrieval enhancement generated document screening system and method fusing a verification mechanism, and relates to the technical field of document screening, the system comprises a user input and query analysis module for extracting key information through natural language processing, and converting the key information into a high-dimensional semantic vector, a meta-tag and a keyword set; the multi-source document retrieval module is used for obtaining documents from multiple data sources through mixed retrieval and generating a candidate set through preliminary screening and sorting; the credibility evaluation and security verification module is used for generating scores and labels after multi-dimensional evaluation and screening qualified documents; the document consistency detection module is used for detecting document conflicts, processing and sequencing, and ensuring logic consistency; the document acquisition and generation module is used for inputting qualified documents into a generation model and generating answers with references; and the result output and tracing module is used for outputting answers and recording whole-process data to ensure traceability. The invention aims to ensure the accuracy and credibility of the generated content through a multi-dimensional verification mechanism.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Knowledge base technical method and system based on RAG retrieval enhancement

The invention discloses a knowledge base technical method and system based on RAG retrieval enhancement, and relates to the technical field of knowledge bases. The method comprises the steps that a source document is analyzed into structured text fragments, and semantic vectors are generated; splitting the text segments into minimum knowledge units, extracting concepts and behavior trigger words to construct a concept association graph, and storing session abstracts in a fixed-length annular structure; the method comprises the following steps: receiving user query, generating a query vector, retrieving a most relevant text fragment, constructing and de-duplicating candidate knowledge units by combining hierarchical diffusion of a concept association graph and an annular abstract matching result, deeply splicing original text paragraphs according to a graph path, generating a dynamic prompt box, calling a generative model to generate a preliminary answer, and executing verification. And the verification state is marked on the final answer. By constructing a concept association map, the activeness weight and the time sequence fingerprint of a map edge are updated in real time, and accurate capture of deep semantics and logical relationships of query intentions is realized.
Owner:深圳市华磊迅拓科技有限公司

Generating structured documents with traceable source lineage

Systems and methods disclosed herein are enabled to dynamically generate structured documents using one or more artificial intelligence models. A computing device receives an output generation request and uses a first AI model to retrieve data chunks from source documents and applicable templates. A second AI model ranks the retrieved chunks based on one or more metrics, such as vector similarity, keyword density, and temporal relevance. A third AI model subsequently generates a response using the ranked chunks, templates, and predefined operational boundaries for each chunk. The generated response is tagged with source identifiers to enable the traceability of the response back to corresponding chunks. The system transmits, via the computing device, the response, the retrieved chunks, and / or the source identifiers.
Owner:CITIBANK N A

Computer-Implemented Methods and Systems for Generative Text Painting

A system and method for transforming text within documents using, such as by using large language models (LLMs). Users can select source text from a source document, in response to which a painting configuration is identified or generated based on the source text, such as by providing the source text and a source prompt to a large language model to produce source output, and selecting or generating the painting configuration based on the source output. The user can select destination text, in response to which the painting configuration is applied to the destination text, such as by selecting or generating a destination action definition based on the painting configuration and the destination text, and providing the destination action definition to a large language model to produce destination output. The destination text may be replaced with the destination output, or output derived therefrom. In this way, the system can extract a variety of sophisticated properties, such as style or tone, from user-selected text source text, and apply those properties to user-selected destination text, with minimal user input.
Owner:QUABBIN PATENT HOLDINGS INC

Multi-agent traceable analysis method, device, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-agent traceable analysis method, device, equipment and medium, and the method comprises the steps: receiving a target theme and a data source list, collecting a multi-source document, and carrying out the preprocessing of the multi-source document to generate a preprocessing document set; configuring an analysis agent based on a semantic clustering result, and setting an analysis direction to form an analysis agent set; generating a structured note and index data table, and executing cross-document comparison to form an analysis output set; and receiving a feedback instruction to adjust the analysis agent set, triggering incremental processing to update the analysis output set, generating a theme research and judgment report, and keeping mapping consistency. According to the method, the multi-source document is fused through semantic clustering and a multi-agent cooperation mechanism, semantic association and traceable analysis are achieved, agent configuration is optimized in combination with interactive feedback, and the accuracy and the intelligent level of report generation are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Vertical field data construction method based on large model

The invention provides a vertical field data construction method based on a large model, which belongs to the technical field of data processing and artificial intelligence, and comprises the following steps: converting a vertical field source document into an intermediate format text, and segmenting the intermediate format text into a plurality of text blocks; inputting the text blocks into a pre-trained generative language model, guiding the pre-trained generative language model according to pre-designed cue words to generate a plurality of candidate questions according to the content of the text blocks, and performing preliminary screening and fine screening on each candidate question to obtain a question set; pre-defining a mode of a knowledge graph according to field characteristics of the vertical field, processing all text blocks based on an information extraction model, and constructing a field knowledge graph; and performing local context retrieval on each final question in the question set based on the text block of the question source, performing global knowledge retrieval based on the domain knowledge graph, and generating a final answer and a final thinking chain. The method is suitable for different vertical fields, the data quality can be effectively improved, and the problem generation accuracy is guaranteed.
Owner:PEKING UNIV

Ranking-augmented generation for long documents

A computer-implemented method comprising: receiving, as input, a query and a source document intended for a content-grounded question-answering or multi-turn conversation task by a specified large language model (LLM) which has a context window size limit, wherein the source document has a size which exceeds the context window size limit; dividing the source document into a plurality of segments; applying a language model to each of the segments, to assign to each of the segments a relevance score; selecting the k-top segments having the highest the relevance scores; combining the selected k-top segments into a virtual document having a size which complies with the context window size limit; and feeding the virtual document as input to the specified LLM, to generate a response that is grounded in the content of the virtual document.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Generation method and generation device of batch record table and storage medium

The invention relates to a batch record table generation method and device and a storage medium. The method comprises the following steps: slicing historical table data of a historical batch record file according to rows to obtain multiple rows of sample fragment data, extracting a sample semantic vector of each row of sample fragment data and storing the sample semantic vector into a vector database, and slicing a template table file according to rows to obtain multiple rows of template fragment data; and respectively extracting the format attribute of each row of template fragment data. And then constructing guide information according to the vector database, format attributes of the multiple rows of template fragment data and source document data, and generating an editing behavior description corresponding to each cell in the template table file based on the guide information, so as to edit each cell in the template table file and generate a target batch record table file. In this way, the generated target batch record table file can meet the requirements for format fidelity and content accuracy at the same time.
Owner:CHENGDU HONGRUI TECH +1

Synthetic data set construction method and electronic equipment

The invention discloses a synthetic data set construction method and electronic equipment, and relates to the technical field of artificial intelligence, and the synthetic data set construction method comprises the following steps: dividing an original multi-source document of a target field into a plurality of word segmentation units by using a word segmentation device; obtaining representative scores of the plurality of word segmentation units on the original multi-source document; based on the representative scores, determining the word segmentation units with the representative scores higher than a first score threshold as candidate keywords; determining importance degree scores of the candidate keywords based on the representative scores of the candidate keywords; based on the importance score, determining the candidate keyword of which the importance score is higher than a second score threshold as a target keyword; and calling a pre-training language model, and based on the target keyword, generating a question and answer pair corresponding to the target keyword to obtain a synthetic data set of the target field. The technical problem that the data coverage rate and the field correlation of the generated synthetic data set are low in the prior art is solved, and the technical effect of improving the data coverage rate and the field correlation of the generated synthetic data set is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Providing generative answers including citations to source documents

Systems and methods include pre-processing documents in cloud storage using query embeddings, providing personalized prompts to users based on documents in cloud storage, real-time anticipation of user interest in information contained in documents in cloud storage, and providing generative answers including citation to source documents in cloud storage. The system and methods generate generative machine learning model (MLM) prompts based on document portions of documents in a cloud-based content management platform. The systems and methods use the generative MLM to generate responses to prompts, and the responses include citations to the document portions used to generate the responses in order for users to verify the responses.
Owner:GOOGLE LLC

Document data general conversion method and device

The invention provides a general document data conversion method and device, and the method comprises the steps: carrying out the business scene analysis and rule field extraction of an input text of a user through a large model business flow processor and an intelligent agent, generating a recommendation rule template, and carrying out the processing of the recommendation rule template; updating a configuration field of the recommendation rule template according to an adjustment requirement text input by a user to generate a target document conversion rule template; and according to the target document conversion rule template, performing general data mapping conversion on the source document data to generate a document data model of the target document, and realizing dynamic code-free automatic configuration by configuring the document conversion rule template, so that the reusability is improved, and the user experience is improved. Automatic conversion of isomorphic data and heterogeneous data to a general document model in a complex scene can be met, the delivery period is shortened, and the delivery cost is reduced; multiple cancel-after-verification relationships can be recorded according to the cancel-after-verification relationship table, an arbitrary cancel mechanism is supported, the bill redundancy is reduced, and the flexibility of cancel-after-verification relationship processing in bill conversion is improved.
Owner:BEIJING JOIN CHEER SOFTWARE

Context-aware information retrieval

Certain aspects of the disclosure provide for information retrieval that exploits context derived from document structure. Source documents can be preprocessed to identify fields and determine context attributes related to each field based on the structural layout of a source document. Resource documents can also be preprocessed to segment a resource document into passages and determine context related to the passages based on structural layout. Queries pertaining to a field can be enhanced by adding context metadata associated with the field. A query embedding can be generated and compared with previously generated passage embeddings to locate candidate matches based on similarity. A machine learning model can be provided with the top-ranked passages and tasked with re-ranking the passages based on relevancy to the original query. The highest re-ranked passage or set of passages can be output in response to the query.
Owner:INTUIT INC

Electronic document standardization modeling processing system

The invention belongs to the technical field of electronic document processing, and particularly relates to an electronic document standardization modeling processing system which comprises a document input preprocessing module, a format standardization conversion module, an analysis reconstruction module, a metadata extraction management module, a quality inspection correction module, an experience optimization module and a security privacy protection module. According to the application, the acceptability of documents of different sources is improved through the document input preprocessing module, it is ensured that various types of electronic documents can be processed, the problem of typesetting disorder caused by format differences is reduced through the format standardization conversion module, and the document appearance consistency is kept through accurate style mapping; the analysis precision is improved through the analysis and reconstruction module, so that the original layout and style are better reserved; through the metadata extraction management module, the document management and retrieval capability is enhanced, and meanwhile, an adjustment space is provided to adapt to requirements under special conditions.
Owner:CHINA AEROSPACE STANDARDIZATION INST

Multi-language version automatic generation and synchronization system of international trade document

The invention relates to the technical field of natural language processing, in particular to a multi-language version automatic generation and synchronization system for international trade documents. Comprising a document processing unit, an intelligent translation unit, a version synchronization unit and a collaborative review unit. A term library is updated in real time, the multiplexing value of a translation memory library is optimized, an advanced deep learning mechanism and strict quality control are applied, high-quality translation results closely following industry dynamics are output, meanwhile, a version synchronization unit saves resources by means of monitoring source document changes in real time and incremental translation, and updating records are safely stored by means of the block chain technology; automatic generation and real-time synchronization of the multi-language document are achieved, the document quality, safety and compliance are guaranteed, and the accuracy, efficiency and collaboration of international trade document processing are improved.
Owner:EAST CHINA JIAOTONG UNIVERSITY

File anti-desensitization self-learning recognition system and method based on information entropy

The invention discloses a file anti-desensitization self-learning recognition system and method based on information entropy, belongs to the technical field of intersection of natural language processing and content security recognition, and is applied to document screening and risk recognition in a multi-task scene. The implementation method comprises the following steps of: 1, performing character recognition and noise reduction processing on an original file to form a data set; 2, training labeled sample data through small samples, respectively adopting probability distribution of a data sliding window and information entropy to carry out maximum and minimum normalization screening, and further utilizing a fitted linear regression model to form an anti-desensitization word list; 3, screening the anti-desensitization degrees of the sentence segments of the data set by adopting a dictionary tree Trie structure to form an anti-desensitization sentence segment table; 4, marking the chapter-level anti-desensitization degree data set text fragments by using the large model; 5, generating an anti-desensitization report according to the anti-desensitization word and the anti-desensitization degree of the marked anti-desensitization file; compared with the prior art, the anti-desensitization file screening method and device have the advantage that the anti-desensitization file screening accuracy is improved.
Owner:BEIJING INST OF TECH

Word document format conversion method based on Java

The invention particularly relates to a Word document format conversion method based on Java. The Word document format conversion method based on Java comprises the following steps: analyzing the content of an original. Doc file, and extracting text paragraphs, tables, pictures, style information and document metadata; the extracted content is divided into different categories, and the XML node type corresponding to each category of elements in the target. Docx document is established; the method comprises the following steps of: constructing a pattern mapping rule base, constructing a new. Docx document structure by using an XWPF Document object model according to an Office Open XML (Extensible Markup Language) specification, sequentially inserting paragraphs, tables and pictures, and applying corresponding pattern configuration; and outputting and generating a. Docx file, detecting and processing abnormal conditions, and recording a conversion log at the same time. The Word document format conversion method based on Java is efficient and accurate, has good compatibility, expandability and cross-platform capability, is suitable for enterprise-level document management systems, cloud services and batch document processing scenes, and remarkably improves document compatibility and processing efficiency of office automation systems.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Document-based presentation generation

A method, apparatus, non-transitory computer readable medium, and system for natural language processing include obtaining a source document and a user characteristic that indicates a complexity preference of a user. A topic description is generated, using a language generation model, based on the source document and the user characteristic. The language generation model is trained based on an objective function that measures a complexity of the topic description.
Owner:ADOBE INC

Automatically extracting tabular data included within a source document

Systems and methods are disclosed for automatically extracting relevant information from various source document that include tabular data. The tabular data in various forms can be received as mixed with other dissimilar data. Tabular data can appear in different orientations, document types, and can be fragmented horizontally or vertically. The proposed technique automatically detects table header data in certain regions of the received source document and associates values to the extracted headers. The proposed system is capable of combining different snippets of smaller tables into a single cohesive and monolithic table with headers designated by a set of keywords and all the values in the various columns (and / or rows) included under or along proper headers.
Owner:USHUR INC

Generalized validation framework for retrieval augmented generation (RAG)

The method involves a process to validate text generated by a RAG system. The method receives text that the RAG system has rephrased in response to a query. The method finds and extracts relevant sections from a source document that match the rephrased text. Both the rephrased text and the source sections are transformed into semantic vectors using NLP techniques. A sliding window technique is applied to the source document vectors, moving sentence by sentence to calculate a semantic similarity score with the rephrased text at each step. Sentences are ranked by similarity, and the ones with the closest match are identified. If the similarity score is above a set threshold, the rephrased text is deemed semantically congruent and validated.
Owner:INTUIT INC

Method and system for generating secure verification documents

In methods and systems for generating secure verification documents are disclosed, a processor that is associated with a multifunction print device will receive one or more source electronic document files, each of which includes content of one or more source documents, each associated with a unique content creator. The processor will cause a print engine of a multifunction device to print a plurality of verification document sheets, each of which comprises data from at least one of the source documents and includes a unique identifier (ID). After printing each verification document sheet, the processor will cause a scanner of the multifunction device to scan the verification document sheet to capture a digital image of the verification document sheet. The processor will then save the digital images of the verification document sheets to a data store.
Owner:XEROX CORP

Server-free document processing and real-time report method and system

The invention provides a server-free document processing and real-time reporting method and system, relates to the technical field of data processing, and starts from responding to the operation that a user uploads a source document to an object storage server, and routing an event to a pre-subscribed document processing server-free function. And after the function is triggered, obtaining a source document, performing analysis and data extraction on the source document to generate structured data, and storing the structured data in a database. The event bus then routes the event to a pre-subscribed report-generating server-less function. And automatically generating a notice containing an access link and pushing the notice to a specified user or terminal. According to the method, automation, real-time performance and flexibility of the whole process from document uploading to report generation are achieved, and through tight combination of event driving and a server-free architecture, the processing efficiency, the resource utilization rate and the system response speed are improved.
Owner:ANHUI FUXING SOFTWARE CO LTD

Automatic generation of handouts from multi-modal documents

Embodiments of the present disclosure include generating a summary of a source document. Some embodiments generate a set of topics based on the summary and a predetermined number of topics. An expanded text is generated for each of the plurality of topics. An image is selected from the source document for each of the set of topics by computing a similarity score between the image and the expanded text. Then, a summary document is generated based on the plurality of topics and the expanded text.
Owner:ADOBE INC

Generating content items based on source document metadata using a generative neural network

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating content items based on source document metadata using a generative neural network. One of the methods include: receiving, from a user, a request to generate a content item using a generative neural network conditioned on a context input, wherein the context input comprises content derived from a source electronic document; obtaining metadata associated with the source electronic document; generating a prompt for the generative neural network based on the context input and the metadata associated with the source electronic document; processing the prompt using the generative neural network to generate the content item; and providing the content item for presentation to the user.
Owner:GOOGLE LLC

Automatic generation of presentations from documents

Embodiments of the present disclosure include extracting structured text from a source document. The structured text comprises a plurality of source sections. Some embodiments generate a semantic outline based on the structured text. In some examples, the semantic outline comprises a plurality of output headings. Some embodiments generate text content corresponding to each of the plurality of output headings. An image is selected from the source document for each of the plurality of output headings by computing a similarity score between the image and the text content. Then, an output document is generated based on the semantic outline, where the output document comprises a plurality of output sections corresponding to the plurality of output headings, respectively.
Owner:ADOBE INC

Title level identification large model training method, title identification method, system and program product

According to the title level identification large model training method, the title identification method and system and the program product provided by the invention, the first title information covering all texts of the original file is constructed through the multi-modal semantic model, and the limitation of plain text identification is made up in combination with the element matching page picture; second title information with semantic and visual features is generated through multi-modal fusion, so that the title judgment accuracy is improved; in the training process, effective title objects are screened in combination with original title objects to optimize pre-training data, and a large model which is high in precision and adapts to complex scenes is cultivated; in the identification process, the title information to be identified is constructed based on the valid title object. And performing dynamic branch processing according to a calling condition, if not, directly outputting an answer, if yes, generating accurate final title information by means of a trained large model, and finally constructing and outputting the answer, thereby realizing training and identification full-link coordination, considering complex scene adaptability and efficient and accurate identification, and comprehensively improving the structuralization and practicability of document title identification.
Owner:SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD +1

Multi-source document management method and device based on knowledge construction and fusion storage

The invention discloses a multi-source document management method and device based on knowledge construction and fusion storage, and relates to the technical field of document management, and the method comprises the steps: receiving a multi-source document to an object storage system, and recognizing the document type; performing document analysis and structure extraction according to the document type; extracting pictures in the image-text mixed content, uploading the pictures to an object storage system, generating a mapping dictionary, and inserting picture marks in a text part of the mapping dictionary; carrying out structured processing on contents of the table key value pairs to extract summaries and abstracts; performing standardization processing and semantic slicing on the processed image-text mixed content, table key value pair content and / or plain text content to generate knowledge fragments; constructing a knowledge extracting questions from knowledge fragments; the knowledge fragments are stored in a first index, and the questions and IDs of the corresponding knowledge fragments are stored in a second index; generating vectors for each knowledge fragment and question, embedding and writing the vectors into a vector database; and performing document management based on the knowledge base. The document management efficiency is improved.
Owner:DIGITAL CHINA SYST INTEGRATION SERVICE

Graphical user interface for syntax and policy compliance review

The present disclosure provides a document review and approval system comprising a user interface with a source selector, an objective definer, and a corpus selector for receiving source documents, user-defined objectives, and policy documents, respectively. An artificial intelligence engine processes the received inputs. Output interfaces include a syntax auditor for identifying and facilitating correction of syntax errors, a policy auditor for identifying non-compliance issues with policies, and a reporter for generating reports based on findings. The system enables efficient review and approval of documents for both syntax and policy compliance, utilizing artificial intelligence to process natural language inputs and generate outputs, including explanations for identified issues and natural language summaries in generated reports.
Owner:CONDUCTORAI CORP