Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

204 results about "Source document" patented technology

A source document is a document in which data collected for a clinical trial is first recorded. This data is usually later entered in the case report form. The International Conference on Harmonisation of Technical Requirements for Registration of Pharmaceuticals for Human Use (ICH-GCP) guidelines define source documents as "original documents, data, and records." Source documents contain source data, which is defined as "all information in original records and certified copies of original records of clinical findings, observations, or other activities in a clinical trial necessary for the reconstruction and evaluation of the trial."

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Retrieval enhancement generated document screening system and method fusing verification mechanism

The invention discloses a retrieval enhancement generated document screening system and method fusing a verification mechanism, and relates to the technical field of document screening, the system comprises a user input and query analysis module for extracting key information through natural language processing, and converting the key information into a high-dimensional semantic vector, a meta-tag and a keyword set; the multi-source document retrieval module is used for obtaining documents from multiple data sources through mixed retrieval and generating a candidate set through preliminary screening and sorting; the credibility evaluation and security verification module is used for generating scores and labels after multi-dimensional evaluation and screening qualified documents; the document consistency detection module is used for detecting document conflicts, processing and sequencing, and ensuring logic consistency; the document acquisition and generation module is used for inputting qualified documents into a generation model and generating answers with references; and the result output and tracing module is used for outputting answers and recording whole-process data to ensure traceability. The invention aims to ensure the accuracy and credibility of the generated content through a multi-dimensional verification mechanism.
Owner:INSPUR TIANYUAN COMM INFORMATION SYST CO LTD

Generating structured documents with traceable source lineage

Systems and methods disclosed herein are enabled to dynamically generate structured documents using one or more artificial intelligence models. A computing device receives an output generation request and uses a first AI model to retrieve data chunks from source documents and applicable templates. A second AI model ranks the retrieved chunks based on one or more metrics, such as vector similarity, keyword density, and temporal relevance. A third AI model subsequently generates a response using the ranked chunks, templates, and predefined operational boundaries for each chunk. The generated response is tagged with source identifiers to enable the traceability of the response back to corresponding chunks. The system transmits, via the computing device, the response, the retrieved chunks, and / or the source identifiers.
Owner:CITIBANK N A

Multi-agent traceable analysis method, device, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-agent traceable analysis method, device, equipment and medium, and the method comprises the steps: receiving a target theme and a data source list, collecting a multi-source document, and carrying out the preprocessing of the multi-source document to generate a preprocessing document set; configuring an analysis agent based on a semantic clustering result, and setting an analysis direction to form an analysis agent set; generating a structured note and index data table, and executing cross-document comparison to form an analysis output set; and receiving a feedback instruction to adjust the analysis agent set, triggering incremental processing to update the analysis output set, generating a theme research and judgment report, and keeping mapping consistency. According to the method, the multi-source document is fused through semantic clustering and a multi-agent cooperation mechanism, semantic association and traceable analysis are achieved, agent configuration is optimized in combination with interactive feedback, and the accuracy and the intelligent level of report generation are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Vertical field data construction method based on large model

The invention provides a vertical field data construction method based on a large model, which belongs to the technical field of data processing and artificial intelligence, and comprises the following steps: converting a vertical field source document into an intermediate format text, and segmenting the intermediate format text into a plurality of text blocks; inputting the text blocks into a pre-trained generative language model, guiding the pre-trained generative language model according to pre-designed cue words to generate a plurality of candidate questions according to the content of the text blocks, and performing preliminary screening and fine screening on each candidate question to obtain a question set; pre-defining a mode of a knowledge graph according to field characteristics of the vertical field, processing all text blocks based on an information extraction model, and constructing a field knowledge graph; and performing local context retrieval on each final question in the question set based on the text block of the question source, performing global knowledge retrieval based on the domain knowledge graph, and generating a final answer and a final thinking chain. The method is suitable for different vertical fields, the data quality can be effectively improved, and the problem generation accuracy is guaranteed.
Owner:PEKING UNIV

Ranking-augmented generation for long documents

A computer-implemented method comprising: receiving, as input, a query and a source document intended for a content-grounded question-answering or multi-turn conversation task by a specified large language model (LLM) which has a context window size limit, wherein the source document has a size which exceeds the context window size limit; dividing the source document into a plurality of segments; applying a language model to each of the segments, to assign to each of the segments a relevance score; selecting the k-top segments having the highest the relevance scores; combining the selected k-top segments into a virtual document having a size which complies with the context window size limit; and feeding the virtual document as input to the specified LLM, to generate a response that is grounded in the content of the virtual document.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Generation method and generation device of batch record table and storage medium

The invention relates to a batch record table generation method and device and a storage medium. The method comprises the following steps: slicing historical table data of a historical batch record file according to rows to obtain multiple rows of sample fragment data, extracting a sample semantic vector of each row of sample fragment data and storing the sample semantic vector into a vector database, and slicing a template table file according to rows to obtain multiple rows of template fragment data; and respectively extracting the format attribute of each row of template fragment data. And then constructing guide information according to the vector database, format attributes of the multiple rows of template fragment data and source document data, and generating an editing behavior description corresponding to each cell in the template table file based on the guide information, so as to edit each cell in the template table file and generate a target batch record table file. In this way, the generated target batch record table file can meet the requirements for format fidelity and content accuracy at the same time.
Owner:CHENGDU HONGRUI TECH +1

Synthetic data set construction method and electronic equipment

The invention discloses a synthetic data set construction method and electronic equipment, and relates to the technical field of artificial intelligence, and the synthetic data set construction method comprises the following steps: dividing an original multi-source document of a target field into a plurality of word segmentation units by using a word segmentation device; obtaining representative scores of the plurality of word segmentation units on the original multi-source document; based on the representative scores, determining the word segmentation units with the representative scores higher than a first score threshold as candidate keywords; determining importance degree scores of the candidate keywords based on the representative scores of the candidate keywords; based on the importance score, determining the candidate keyword of which the importance score is higher than a second score threshold as a target keyword; and calling a pre-training language model, and based on the target keyword, generating a question and answer pair corresponding to the target keyword to obtain a synthetic data set of the target field. The technical problem that the data coverage rate and the field correlation of the generated synthetic data set are low in the prior art is solved, and the technical effect of improving the data coverage rate and the field correlation of the generated synthetic data set is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Context-aware information retrieval

Certain aspects of the disclosure provide for information retrieval that exploits context derived from document structure. Source documents can be preprocessed to identify fields and determine context attributes related to each field based on the structural layout of a source document. Resource documents can also be preprocessed to segment a resource document into passages and determine context related to the passages based on structural layout. Queries pertaining to a field can be enhanced by adding context metadata associated with the field. A query embedding can be generated and compared with previously generated passage embeddings to locate candidate matches based on similarity. A machine learning model can be provided with the top-ranked passages and tasked with re-ranking the passages based on relevancy to the original query. The highest re-ranked passage or set of passages can be output in response to the query.
Owner:INTUIT INC

Electronic document standardization modeling processing system

The invention belongs to the technical field of electronic document processing, and particularly relates to an electronic document standardization modeling processing system which comprises a document input preprocessing module, a format standardization conversion module, an analysis reconstruction module, a metadata extraction management module, a quality inspection correction module, an experience optimization module and a security privacy protection module. According to the application, the acceptability of documents of different sources is improved through the document input preprocessing module, it is ensured that various types of electronic documents can be processed, the problem of typesetting disorder caused by format differences is reduced through the format standardization conversion module, and the document appearance consistency is kept through accurate style mapping; the analysis precision is improved through the analysis and reconstruction module, so that the original layout and style are better reserved; through the metadata extraction management module, the document management and retrieval capability is enhanced, and meanwhile, an adjustment space is provided to adapt to requirements under special conditions.
Owner:CHINA AEROSPACE STANDARDIZATION INST

Word document format conversion method based on Java

The invention particularly relates to a Word document format conversion method based on Java. The Word document format conversion method based on Java comprises the following steps: analyzing the content of an original. Doc file, and extracting text paragraphs, tables, pictures, style information and document metadata; the extracted content is divided into different categories, and the XML node type corresponding to each category of elements in the target. Docx document is established; the method comprises the following steps of: constructing a pattern mapping rule base, constructing a new. Docx document structure by using an XWPF Document object model according to an Office Open XML (Extensible Markup Language) specification, sequentially inserting paragraphs, tables and pictures, and applying corresponding pattern configuration; and outputting and generating a. Docx file, detecting and processing abnormal conditions, and recording a conversion log at the same time. The Word document format conversion method based on Java is efficient and accurate, has good compatibility, expandability and cross-platform capability, is suitable for enterprise-level document management systems, cloud services and batch document processing scenes, and remarkably improves document compatibility and processing efficiency of office automation systems.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Document-based presentation generation

A method, apparatus, non-transitory computer readable medium, and system for natural language processing include obtaining a source document and a user characteristic that indicates a complexity preference of a user. A topic description is generated, using a language generation model, based on the source document and the user characteristic. The language generation model is trained based on an objective function that measures a complexity of the topic description.
Owner:ADOBE INC

Automatically extracting tabular data included within a source document

Systems and methods are disclosed for automatically extracting relevant information from various source document that include tabular data. The tabular data in various forms can be received as mixed with other dissimilar data. Tabular data can appear in different orientations, document types, and can be fragmented horizontally or vertically. The proposed technique automatically detects table header data in certain regions of the received source document and associates values to the extracted headers. The proposed system is capable of combining different snippets of smaller tables into a single cohesive and monolithic table with headers designated by a set of keywords and all the values in the various columns (and / or rows) included under or along proper headers.
Owner:USHUR INC

Generalized validation framework for retrieval augmented generation (RAG)

The method involves a process to validate text generated by a RAG system. The method receives text that the RAG system has rephrased in response to a query. The method finds and extracts relevant sections from a source document that match the rephrased text. Both the rephrased text and the source sections are transformed into semantic vectors using NLP techniques. A sliding window technique is applied to the source document vectors, moving sentence by sentence to calculate a semantic similarity score with the rephrased text at each step. Sentences are ranked by similarity, and the ones with the closest match are identified. If the similarity score is above a set threshold, the rephrased text is deemed semantically congruent and validated.
Owner:INTUIT INC

Method and system for generating secure verification documents

In methods and systems for generating secure verification documents are disclosed, a processor that is associated with a multifunction print device will receive one or more source electronic document files, each of which includes content of one or more source documents, each associated with a unique content creator. The processor will cause a print engine of a multifunction device to print a plurality of verification document sheets, each of which comprises data from at least one of the source documents and includes a unique identifier (ID). After printing each verification document sheet, the processor will cause a scanner of the multifunction device to scan the verification document sheet to capture a digital image of the verification document sheet. The processor will then save the digital images of the verification document sheets to a data store.
Owner:XEROX CORP

Server-free document processing and real-time report method and system

The invention provides a server-free document processing and real-time reporting method and system, relates to the technical field of data processing, and starts from responding to the operation that a user uploads a source document to an object storage server, and routing an event to a pre-subscribed document processing server-free function. And after the function is triggered, obtaining a source document, performing analysis and data extraction on the source document to generate structured data, and storing the structured data in a database. The event bus then routes the event to a pre-subscribed report-generating server-less function. And automatically generating a notice containing an access link and pushing the notice to a specified user or terminal. According to the method, automation, real-time performance and flexibility of the whole process from document uploading to report generation are achieved, and through tight combination of event driving and a server-free architecture, the processing efficiency, the resource utilization rate and the system response speed are improved.
Owner:ANHUI FUXING SOFTWARE CO LTD

Generating content items based on source document metadata using a generative neural network

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating content items based on source document metadata using a generative neural network. One of the methods include: receiving, from a user, a request to generate a content item using a generative neural network conditioned on a context input, wherein the context input comprises content derived from a source electronic document; obtaining metadata associated with the source electronic document; generating a prompt for the generative neural network based on the context input and the metadata associated with the source electronic document; processing the prompt using the generative neural network to generate the content item; and providing the content item for presentation to the user.
Owner:GOOGLE LLC

Title level identification large model training method, title identification method, system and program product

According to the title level identification large model training method, the title identification method and system and the program product provided by the invention, the first title information covering all texts of the original file is constructed through the multi-modal semantic model, and the limitation of plain text identification is made up in combination with the element matching page picture; second title information with semantic and visual features is generated through multi-modal fusion, so that the title judgment accuracy is improved; in the training process, effective title objects are screened in combination with original title objects to optimize pre-training data, and a large model which is high in precision and adapts to complex scenes is cultivated; in the identification process, the title information to be identified is constructed based on the valid title object. And performing dynamic branch processing according to a calling condition, if not, directly outputting an answer, if yes, generating accurate final title information by means of a trained large model, and finally constructing and outputting the answer, thereby realizing training and identification full-link coordination, considering complex scene adaptability and efficient and accurate identification, and comprehensively improving the structuralization and practicability of document title identification.
Owner:SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD +1

Multi-source document management method and device based on knowledge construction and fusion storage

The invention discloses a multi-source document management method and device based on knowledge construction and fusion storage, and relates to the technical field of document management, and the method comprises the steps: receiving a multi-source document to an object storage system, and recognizing the document type; performing document analysis and structure extraction according to the document type; extracting pictures in the image-text mixed content, uploading the pictures to an object storage system, generating a mapping dictionary, and inserting picture marks in a text part of the mapping dictionary; carrying out structured processing on contents of the table key value pairs to extract summaries and abstracts; performing standardization processing and semantic slicing on the processed image-text mixed content, table key value pair content and / or plain text content to generate knowledge fragments; constructing a knowledge extracting questions from knowledge fragments; the knowledge fragments are stored in a first index, and the questions and IDs of the corresponding knowledge fragments are stored in a second index; generating vectors for each knowledge fragment and question, embedding and writing the vectors into a vector database; and performing document management based on the knowledge base. The document management efficiency is improved.
Owner:DIGITAL CHINA SYST INTEGRATION SERVICE

Graphical user interface for syntax and policy compliance review

The present disclosure provides a document review and approval system comprising a user interface with a source selector, an objective definer, and a corpus selector for receiving source documents, user-defined objectives, and policy documents, respectively. An artificial intelligence engine processes the received inputs. Output interfaces include a syntax auditor for identifying and facilitating correction of syntax errors, a policy auditor for identifying non-compliance issues with policies, and a reporter for generating reports based on findings. The system enables efficient review and approval of documents for both syntax and policy compliance, utilizing artificial intelligence to process natural language inputs and generate outputs, including explanations for identified issues and natural language summaries in generated reports.
Owner:CONDUCTORAI CORP

Document translation feasibility analysis systems and methods

Automatic post-editing of machine translated content is disclosed herein. An example method includes receiving a source document for translation, calculating semantic signatures of text chunks within the source document using an artificial intelligence (AI) model, matching the semantic signatures against a repository of previous translations, and displaying a document-based visualization that includes both semantic and exact matches found in the repository of previous translations.
Owner:SDL LTD

Intelligent security traceability method and system based on steganography watermark and hierarchical signature

The invention relates to the technical field of data encryption hiding, in particular to an intelligent security traceability method and system based on steganography watermarking and hierarchical signatures, and the method comprises the steps: carrying out the first-level watermarking embedding of an original file, and obtaining a first-level watermarking file; then storing signature event data and device fingerprints in the watermark embedding process of the first level in a block chain; transmitting the watermark file of the first level to a second level to obtain a first transmission watermark file; performing hierarchical watermark signature chain verification according to the first transmission watermark file; when the verification is not passed, terminating file transmission and freezing file access permission, and when the verification is passed, calculating a real-time risk index at the current moment, triggering a hierarchical response according to the real-time risk index, and performing a subsequent hierarchical watermark embedding operation according to a hierarchical response result; and when the hierarchical response is high in risk, querying an abnormal point through the block chain, and positioning a leakage source. According to the invention, the accuracy of intelligent safety traceability is improved.
Owner:SHAANXI ZHIYUAN INTERNET SOFTWARE CO LTD

Data set generation method and device, storage medium and electronic equipment

The invention belongs to the field of artificial intelligence, and provides a data set generation method and device and electronic device.The method comprises the steps that through a document parser, the type of each element in each page of content in a target document is recognized, and each element is obtained; and combining the elements in sequence to generate a text in a first preset format, and processing the text in the first preset format through a first generator to generate a corresponding fine tuning data set, and processing the text in the first preset format through a second generator to generate a corresponding evaluation data set. Through the generation method, the fine tuning data set and the evaluation data set are automatically generated based on the multi-source document.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Variable edge document encoding for secure document storage and retrieval

PendingUS20260017472A1Digital marking by printing code marksPictoral communicationElectronic documentComputer graphics (images)
This document discloses a method and system for assessing integrity of a stack of ballots or other documents. A processor receives a first set of electronic source document files, each of which includes content of a corresponding source document. The processor generates a verification image, segments the verification image into slices, and causes a print device to print verification document sheets. Each verification document sheet includes data from one of the source electronic document files and, on an edge of the sheet, a unique one of the slices of the verification image. The verification document sheets are stacked so that the slices appear on a side of the stack and, when all of the verification document sheets have been placed onto the stack, the slices collectively form the verification image.
Owner:XEROX CORP

Method and system for securely generating document verification records

Methods for providing security in a document printing process are disclosed. In various embodiments, when a system comprising a print device receives source document files, then before printing it may ensure that the system satisfies various security conditions before it will print verification document sheets based on the source document files. For example, the system may require that data from all prior print jobs be cleared from the memory. The system also may require that the print device be restricted from external communications, browser application usage, or both. After printing all verification document sheets for all of the source electronic document files, the system may restrict the print device from processing any future print jobs until all data associated with the source electronic document files and the verification document sheets have been removed from the print device.
Owner:XEROX CORP

Document analysis method, device and equipment and computer readable storage medium

The invention discloses a document analysis method, device and equipment and a computer readable storage medium. The method comprises the following steps: when the type of an original file is a spreadsheet file type and a table in the original file has merged cells, determining the content of each cell in the table, the row number i and the column number j of each non-merged cell in the table and the area coordinate of each merged cell in the table; for each non-merged cell, filling the content of the non-merged cell into the cells in the ith row and the jth column in the correction table; for each merged cell, filling the content of the merged cell into a target cell corresponding to the region coordinate in the correction table; and after all the cells are traversed, embedding the obtained correction table into the plain text document. By means of the method and device, the situation that the format of the content embedded into the plain text document is disordered or information is lost when compared with that of a table in an original file is avoided to a great extent.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

Method for resisting printing and scanning digital watermarking for text image

The invention discloses an anti-printing and anti-scanning digital watermark method for a text image, which comprises the following steps of: embedding and extracting a watermark on the basis of a stroke-like thought, segmenting characters and determining the average skeleton quality ASM of each character during embedding, and forming a group by every two characters, changing the ASM of each group of characters according to the to-be-embedded information to embed the watermark and replace the corresponding characters of the source document; during extraction, two characters are also taken as a group, ASM is extracted, and watermark information is extracted by comparing ASM among the characters, so that 1-bit information can be embedded into the two characters, the language type and font of the characters are not limited, the watermark information amount of the document is ensured, and meanwhile, the watermark information is extracted. The problems that algorithms among different languages are limited, and the robustness is poor and the capacity is small when digital watermarks in a space domain and a transform domain resist printing and scanning attacks are effectively solved.
Owner:浣江实验室

Date offset in document

A method of anonymizing a digital document includes: receiving a digital source document containing text including at least one date; referring to a user-defined rule set, wherein the rule set includes at least a date offset function; and applying modifications to the digital source document according to the user-defined rule set, so as to anonymize the digital source document, the modifications including changing the least one date to a different date in accordance with the date offset function; and saving the modified digital source document as an output document.
Owner:TRIALASSURE INC

Synthetic document generation using machine learning based language model

A system generates synthetic documents from source documents. The system receives a source document and generates prompts for generating sections of a synthetic document from the source document. The system receives sections of the synthetic document based on execution of the machine learning based language model. For each of one or more portions of the synthetic document, the system determines a snippet of the source document that provides support for the portion of the synthetic document. The system sends the synthetic document for display via a user interface. The system receives a request via the user interface to inspect a portion of the synthetic document. The system identifies a snippet of the source document corresponding to the portion of the synthetic document and sends the snippet of the source document for display via the user interface.
Owner:BENCH IQ INC

Updating support documentation for developer platforms with topic clustering of feedback

A method includes obtaining at least one feedback response from a developer regarding an answer and corresponding source documents of the answer in a discussion thread initiated by the developer. The method further includes converting the discussion thread into a topic model clustering input, responsive to the feedback response specifying an unsatisfactory category of feedback responses, to obtain a multitude of topic clustering model inputs. The method further includes periodically processing the multitude of topic clustering model inputs by a thread-topic clustering model to obtain a multitude of candidate topics. The method further includes processing, by an answer generation model, a first candidate topic of the multitude of candidate topics to obtain a multitude of corresponding documentation recommendations for the candidate topic. The method further includes presenting the first candidate topic and the multitude of corresponding documentation recommendations.
Owner:INTUIT INC