Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

125 results about "Source document" patented technology

A source document is a document in which data collected for a clinical trial is first recorded. This data is usually later entered in the case report form. The International Conference on Harmonisation of Technical Requirements for Registration of Pharmaceuticals for Human Use (ICH-GCP) guidelines define source documents as "original documents, data, and records." Source documents contain source data, which is defined as "all information in original records and certified copies of original records of clinical findings, observations, or other activities in a clinical trial necessary for the reconstruction and evaluation of the trial."

Multi-agent traceable analysis method, device, equipment and medium

The invention relates to the technical field of data analysis, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-agent traceable analysis method, device, equipment and medium, and the method comprises the steps: receiving a target theme and a data source list, collecting a multi-source document, and carrying out the preprocessing of the multi-source document to generate a preprocessing document set; configuring an analysis agent based on a semantic clustering result, and setting an analysis direction to form an analysis agent set; generating a structured note and index data table, and executing cross-document comparison to form an analysis output set; and receiving a feedback instruction to adjust the analysis agent set, triggering incremental processing to update the analysis output set, generating a theme research and judgment report, and keeping mapping consistency. According to the method, the multi-source document is fused through semantic clustering and a multi-agent cooperation mechanism, semantic association and traceable analysis are achieved, agent configuration is optimized in combination with interactive feedback, and the accuracy and the intelligent level of report generation are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Vertical field data construction method based on large model

The invention provides a vertical field data construction method based on a large model, which belongs to the technical field of data processing and artificial intelligence, and comprises the following steps: converting a vertical field source document into an intermediate format text, and segmenting the intermediate format text into a plurality of text blocks; inputting the text blocks into a pre-trained generative language model, guiding the pre-trained generative language model according to pre-designed cue words to generate a plurality of candidate questions according to the content of the text blocks, and performing preliminary screening and fine screening on each candidate question to obtain a question set; pre-defining a mode of a knowledge graph according to field characteristics of the vertical field, processing all text blocks based on an information extraction model, and constructing a field knowledge graph; and performing local context retrieval on each final question in the question set based on the text block of the question source, performing global knowledge retrieval based on the domain knowledge graph, and generating a final answer and a final thinking chain. The method is suitable for different vertical fields, the data quality can be effectively improved, and the problem generation accuracy is guaranteed.
Owner:PEKING UNIV

Context-aware information retrieval

Certain aspects of the disclosure provide for information retrieval that exploits context derived from document structure. Source documents can be preprocessed to identify fields and determine context attributes related to each field based on the structural layout of a source document. Resource documents can also be preprocessed to segment a resource document into passages and determine context related to the passages based on structural layout. Queries pertaining to a field can be enhanced by adding context metadata associated with the field. A query embedding can be generated and compared with previously generated passage embeddings to locate candidate matches based on similarity. A machine learning model can be provided with the top-ranked passages and tasked with re-ranking the passages based on relevancy to the original query. The highest re-ranked passage or set of passages can be output in response to the query.
Owner:INTUIT INC

Server-free document processing and real-time report method and system

The invention provides a server-free document processing and real-time reporting method and system, relates to the technical field of data processing, and starts from responding to the operation that a user uploads a source document to an object storage server, and routing an event to a pre-subscribed document processing server-free function. And after the function is triggered, obtaining a source document, performing analysis and data extraction on the source document to generate structured data, and storing the structured data in a database. The event bus then routes the event to a pre-subscribed report-generating server-less function. And automatically generating a notice containing an access link and pushing the notice to a specified user or terminal. According to the method, automation, real-time performance and flexibility of the whole process from document uploading to report generation are achieved, and through tight combination of event driving and a server-free architecture, the processing efficiency, the resource utilization rate and the system response speed are improved.
Owner:ANHUI FUXING SOFTWARE CO LTD

Generating content items based on source document metadata using a generative neural network

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating content items based on source document metadata using a generative neural network. One of the methods include: receiving, from a user, a request to generate a content item using a generative neural network conditioned on a context input, wherein the context input comprises content derived from a source electronic document; obtaining metadata associated with the source electronic document; generating a prompt for the generative neural network based on the context input and the metadata associated with the source electronic document; processing the prompt using the generative neural network to generate the content item; and providing the content item for presentation to the user.
Owner:GOOGLE LLC

Title level identification large model training method, title identification method, system and program product

According to the title level identification large model training method, the title identification method and system and the program product provided by the invention, the first title information covering all texts of the original file is constructed through the multi-modal semantic model, and the limitation of plain text identification is made up in combination with the element matching page picture; second title information with semantic and visual features is generated through multi-modal fusion, so that the title judgment accuracy is improved; in the training process, effective title objects are screened in combination with original title objects to optimize pre-training data, and a large model which is high in precision and adapts to complex scenes is cultivated; in the identification process, the title information to be identified is constructed based on the valid title object. And performing dynamic branch processing according to a calling condition, if not, directly outputting an answer, if yes, generating accurate final title information by means of a trained large model, and finally constructing and outputting the answer, thereby realizing training and identification full-link coordination, considering complex scene adaptability and efficient and accurate identification, and comprehensively improving the structuralization and practicability of document title identification.
Owner:SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD +1

A content intelligent referencing method for document editing

This invention provides a method for intelligent content referencing in document editing, relating to the technical field of intelligent document referencing technology. The method includes: displaying a source document selection interface in response to a user's insertion operation in the main document; displaying at least one page of the source document in response to the user's selection of the source document; determining the selected target page in response to the user's selection of the page; parsing the document structure of the target page, identifying and visually displaying at least one content block within the page; recommending a corresponding insertion method based on the type of the content block; and inserting the content block into the main document according to the selected insertion method in response to the user's selection of the content block and confirmation of the insertion method. This invention solves the problem of low content referencing efficiency, thereby improving the efficiency and accuracy of content referencing.
Owner:HUNAN ZHUOZHI INFORMATION TECHNOLOGY CO LTD

Providing personalized prompts to users based on documents in cloud storage

Systems and methods include pre-processing documents in cloud storage using query embeddings, providing personalized prompts to users based on documents in cloud storage, real-time anticipation of user interest in information contained in documents in cloud storage, and providing generative answers including citation to source documents in cloud storage. The system and methods generate generative machine learning model (MLM) prompts based on document portions of documents in a cloud-based content management platform. The systems and methods use the generative MLM to generate responses to prompts, and the responses include citations to the document portions used to generate the responses in order for users to verify the responses.
Owner:GOOGLE LLC

Extracting definitions from documents utilizing definition-labeling-dependent machine learning background

This disclosure describes methods, non-transitory computer readable storage media, and systems that extract a definition for a term from a source document by utilizing a single machine-learning framework to classify a word sequence from the source document as including a term definition and to label words from the word sequence. To illustrate, the disclosed system can receive a source document including a word sequence arranged in one or more sentences. The disclosed systems can utilize a machine-learning model to classify the word sequence as comprising a definition for a term and generate labels for the words from the word sequence corresponding to the term and the definition. Based on classifying the word sequence and the generated labels, the disclosed system can extract the definition for the term from the source document.
Owner:ADOBE INC

Context-based retrieval enhancement generation optimization method and system

The invention discloses a context-based retrieval enhancement generation optimization method and system, and the method comprises the steps: screening and obtaining a source document from a terminal, and carrying out the structural processing and conversion processing of the source document, and obtaining a directory tree; according to the directory tree, performing content segmentation on the source document to generate a context with a hierarchical relationship; performing semantic association extension according to the context to generate a knowledge block set; preprocessing all knowledge blocks in the knowledge block set to generate a knowledge block list; wherein all the subordinate knowledge blocks of the knowledge block list are configured with attribute tags; performing metadata judgment on the knowledge block list, and constructing a prompt word bank; and matching the retrieval request with the prompt word bank, and searching a corresponding knowledge result in the knowledge base list. Semantic association is performed through a tree adjacent context strategy, so that a knowledge block range is expanded, a structured directory tree and a dynamic splicing strategy enable knowledge base segmentation and splicing to better conform to cognitive logic, semantic breakage caused by fixed length segmentation is avoided, and the understanding depth of long documents is improved.
Owner:HUIZHIAN INFORMATION TECH CO LTD

A document parsing and exporting method and device based on a multi-modal large model, equipment and a storage medium

PendingCN122472022ADocument analysisAlgorithm
The application discloses a document analysis and export method and device based on a multi-modal large model, equipment and a storage medium, and relates to the technical field of computers, which comprises the following steps: generating a task object based on the original file name of a to-be-processed document, determining a to-be-processed task; rendering the current page into a raster image, setting the text content of the current page as a page text clue, determining the to-be-processed analysis main text of the current page based on a prompt word template and in combination with the raster image and the page text clue, obtaining an update request for the analysis main text, calling a multi-modal large model based on the request, the to-be-processed analysis main text and the prompt word template to generate target analysis main text, marking the current page as confirmed based on a confirmation instruction, and storing the corresponding target analysis main text and raster image, generating a document file according to a merging and exporting instruction, and sequentially writing the original file name row, the page number row, the raster image and the target analysis main text into each confirmed page, so that the efficiency of document analysis and export is improved.
Owner:XINTONG EMPOWERMENT (CHANGSHA) ARTIFICIAL INTELLIGENCE IND APPLICATION SYSTEM CO LTD

Watermark embedding method, equipment, storage medium and device

The invention discloses a watermark embedding method, watermark embedding equipment, a storage medium and a watermark embedding device, and relates to the technical field of watermark embedding, and the watermark embedding method comprises the following steps: analyzing user identification information in an identity authentication token of a current user in a file viewing request; generating a target fan-shaped watermark based on the user identification information and a preset fan-shaped watermark generation rule; a target fan-shaped watermark is embedded into an original file based on a preset rendering engine, a file stream with the watermark is generated, and compared with a traditional watermark embedding method which cannot be effectively embedded into a document and is changed, the file leakage risk is high, and the watermark embedding efficiency is improved. According to the method, the to-be-checked file is subjected to watermark embedding through the user identity verification information and the preset fan-shaped watermark generation rule, so that the watermark is effectively prevented from being changed, and the risk of file leakage is reduced.
Owner:中邮消费金融有限公司

Computer-implemented methods and systems for generative text painting

A system and method for transforming text within documents using, such as by using large language models (LLMs). Users can select source text from a source document, in response to which a painting configuration is identified or generated based on the source text, such as by providing the source text and a source prompt to a large language model to produce source output, and selecting or generating the painting configuration based on the source output. The user can select destination text, in response to which the painting configuration is applied to the destination text, such as by selecting or generating a destination action definition based on the painting configuration and the destination text, and providing the destination action definition to a large language model to produce destination output. The destination text may be replaced with the destination output, or output derived therefrom. In this way, the system can extract a variety of sophisticated properties, such as style or tone, from user-selected text source text, and apply those properties to user-selected destination text, with minimal user input.
Owner:QUABBIN PATENT HOLDINGS INC

Data retrieval method, server, terminal, and storage medium

The application discloses a data retrieval method, a server, a terminal and a storage medium, which are used for providing a knowledge base question and answer retrieval function for VDI users through a server without occupying the storage space of the server, and providing a safer, more reliable and more intelligent question and answer retrieval scheme. The method comprises the following steps: acquiring first text information from a first remote terminal in one or more remote terminals, wherein the first text information corresponds to a first embedding vector; if the similarity of the first embedding vector and a second embedding vector in a vector database is greater than or equal to a threshold value, determining a retrieval result of the first text information according to the second text information corresponding to the second embedding vector and the first embedding vector, wherein the second text information is obtained by text segmentation according to source documents of the one or more remote terminals.
Owner:RUIJIE NETWORKS CO LTD

Self-evolution demonstration document generation method and system based on multi-modal and knowledge graph

PendingCN122262361AAvoid Distortion of Detailsavoid inconsistent styleSemantic analysisSpecial data processing applicationsEngineeringContinual learning
The application provides a self-evolution presentation document generation method and system based on multi-modal and knowledge graph, wherein the self-evolution presentation document generation method comprises the following steps: constructing an enterprise-level picture meta-database containing a vector index; parsing a source document to obtain structured content and generating a presentation document outline; searching for matching candidate pictures in the enterprise-level picture meta-database based on the target page core content in the outline; querying a design decision knowledge graph to generate layout and color matching planning; assembling a presentation document according to the planning and synchronously generating design decision metadata; driving knowledge graph updating by displaying the design decision metadata and obtaining user modification feedback on the design; and the self-evolution presentation document generation system comprises functional modules for realizing the above method. The application solves the technical problems of poor picture quality, inaccurate image-text matching, rigid design, high copyright risk and difficulty in continuous learning and evolution from use in the existing automatic presentation document generation technology.
Owner:CHINA RAILWAY TUNNEL GROUP CO LTD +1

Modification, sharing, and querying of authoritative source documents

The aspects described herein pertain at least to a method implemented by a computer for responding to a request addressed to a global marker of an electronic document. The method includes receiving a request addressed to a global marker of an electronic document, processing, by execution of the electronic document, the request, and responding to, by the execution of the electronic document, the request.
Owner:FACTIFY TECHNOLOGIES INC

Method and device for fusing, storing and managing massive small files and data records

The invention discloses a method and device for fusing, storing and managing massive small files and data records, and the method comprises the steps: building an external metadatabase which is used for managing a hierarchical relation of scientific observation data sources and a mapping relation between an aggregation file and an original file of scientific observation data; according to a preset aggregation rule based on data product categories and timelines, the multiple original files are aggregated and packaged in a target file container in a column storage format, and the content of the original files and structured data records obtained by analyzing the original files are stored in the target file container at the same time; and based on the external metadatabase and the target file container, providing file-level access aiming at the original file content and record-level retrieval service aiming at the structured data record. According to the method, the aspects of system simplification, performance optimization, space saving and service capability are remarkably improved, and a new way with high expandability is provided for efficient storage and multi-level service of space weather scientific data.
Owner:NAT SPACE SCI CENT CAS

Method for automatic generation of frequently asked questions

Methods and systems for generating a frequently asked questions are provided, which include defining, by a computer program executed by a computer, a first large language model (LLM) with a user query and a feedback of the user query; refining, by the computer program, the user query to a question set based on the feedback, the question set comprising one or more sentences; defining, by the computer program, a second LLM to generate a first set of question and answer pairs from a source document; defining, by the computer program, a third LLM to generate a content set from the source document based on a rewriting of the source document; selecting, by the computer program, top questions from the content set to be provided to the second LLM; and generating, by the second LLM, a second set of question and answer pairs based on the top questions.
Owner:JPMORGAN CHASE BANK NA

Systems and methods for data masking

PendingUS20260037673A1Digital data protectionData miningData masking
The present disclosure is related to systems and methods for data masking. The method includes obtaining at least one original file and a hierarchical relationship that is associated with data in the at least one original file. The method includes obtaining a masking template for the data in the at least one original file. The method includes masking the data in the at least one original file based on the masking template, to generate at least one target file. The method includes storing the at least one target file based on the hierarchical relationship.
Owner:SHANGHAI UNITED IMAGING METAHEALTHCARE CO LTD

Intelligent agent creating and optimizing method for sintering field

The invention discloses an intelligent agent creating and optimizing method for the sintering field, and relates to the technical field of intelligent industries, and the method comprises the specific steps: collecting a sintering field multi-source document, and constructing a structured knowledge base after processing; splitting and preprocessing the document to generate question and answer pairs; building a retrieval framework, processing the text by using a vectorization model, and calculating semantic similarity matching; finely adjusting the base model according to the question-answer pair to obtain a special large language model; finally, the model is deployed in a localized mode, performance is verified through an evaluation algorithm, a knowledge base, the model or retrieval logic is optimized according to a result, and a closed loop is formed; according to the method, a structured knowledge base in the sintering field is constructed by integrating multi-source knowledge, accurate query is realized through hierarchical problem tree and scene perception retrieval, a large language model is optimized by adopting a field adaptation fine tuning algorithm, closed-loop optimization is formed by combining a dual evaluation algorithm, professional and reliable intelligent agent output is ensured, efficient intelligent support is provided for iron and steel enterprises, and the method is suitable for popularization and application. And intelligent development of sintering production is promoted.
Owner:HEBEI UNIVERSITY OF ECONOMICS AND BUSINESS

A visual-based document parsing method and system

The application discloses a kind of based on vision's document analysis method and system, applied to intelligent identification technical field, method includes: first artificial intelligence model is split into multiple fragment documents with target source document, and obtains the meta-information corresponding to fragment document;Second artificial intelligence model is parsed according to meta-information to all fragment documents to generate corresponding analysis fragment and join pagination anchor in analysis fragment, third artificial intelligence model is abnormal analysis and repair to analysis fragment;Pagination anchor is the page number identification corresponding to each fragment document;The final analysis fragment obtained is spliced to form the complete analysis document of target source document.This application is combined by intelligent splitting and parallel analysis, and is supplemented with comprehensive abnormal monitoring and recovery mechanism, effectively overcome the context length limit of existing large artificial intelligence model, guarantee the semantic and structural integrity of content unit after splitting, improve processing efficiency and system stability.
Owner:CIVIL AVIATION FLIGHT UNIV OF CHINA

Decentralized storage dynamic auditing method and system based on zero-knowledge proof

The invention belongs to the technical field of information security, and particularly relates to a decentralized storage dynamic auditing method and system based on zero-knowledge proof, and the method comprises the steps: partitioning an original file, constructing an index Merkle hash tree, generating a unique identifier, uploading a data block to a storage node, and recording file information in a version chain; when the smart contract initiates verification, the storage node uses the first zero-knowledge proof circuit to generate an audit proof and return the audit proof, and the contract verifies the proof to confirm that the file is correctly stored; after receiving an update request, the storage node performs aggregation update such as insertion, deletion or modification on the file according to an operation instruction, and generates an update proof by using a second zero-knowledge proof circuit; and the smart contract verifies the proof, and records the unique identifier of the new version to the version chain after the verification is passed. According to the invention, the credibility and security of decentralized storage are improved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Text clustering with heuristic and multi-metric control

Implementations generally relate to text clustering with heuristic and multi-metric control. In some implementations, a method includes receiving an electronic source document containing text. The method further includes dividing the text into text units, encoding the text units, and transforming the text units into numerical values. The method further includes generating a graph of the text units based on the numeric values, where the graph includes nodes corresponding to the text units and edges corresponding to pairs of the text units. The method further includes ordering the text units into text clusters based on the graph of the text units. The method further includes generating an electronic target document that presents the text clusters based on one or more preference heuristics.
Owner:JPMORGAN CHASE BANK NA

Information processing system, information processing method, and information processing program

This system more accurately detects duplicate evidence and prevents the inclusion of duplicate documents. [Solution] The duplicate detection unit P5 detects a duplicate between the source document information Ss and the referenced document information Sd when confirmed information F included in the source document information Ss is input, and the confirmed information F included in the source document information Ss is common with at least one of the confirmed information F or estimated information U included in the referenced document information Sd.
Owner:FREEE

A method and system for fully localized hybrid retrieval and knowledge graph construction in the research and production of standard gases.

This invention relates to a fully localized hybrid retrieval and knowledge graph construction method and system for standard gas R&D and production, belonging to the interdisciplinary field of artificial intelligence and industrial control. The method includes: acquiring multi-source heterogeneous documents and performing text parsing, mixed Chinese and English word segmentation, and terminology standardization; employing a hybrid retrieval strategy that weights and fuses BM25 keyword retrieval and vector semantic retrieval to obtain fused retrieval results; automatically extracting entities and relations from the fused retrieval results to generate triplet data; constructing a knowledge graph based on the triplet data, with nodes bearing source document identifiers and version tags; and deploying the embedded model, text units, semantic vectors, and knowledge graph on a local server. This invention achieves the fusion of precise matching and semantic recall, ensures core data security, transforms unstructured knowledge into a computable and traceable knowledge network, and significantly improves the efficiency and intelligence level of knowledge retrieval in standard gas R&D and production.
Owner:重庆朝阳气体有限公司

Method for automatic generation of frequently asked questions

Methods and systems for generating a frequently asked questions are provided, which include defining, by a computer program executed by a computer, a first large language model (LLM) with a user query and a feedback of the user query; refining, by the computer program, the user query to a question set based on the feedback, the question set comprising one or more sentences; defining, by the computer program, a second LLM to generate a first set of question and answer pairs from a source document; defining, by the computer program, a third LLM to generate a content set from the source document based on a rewriting of the source document; selecting, by the computer program, top questions from the content set to be provided to the second LLM; and generating, by the second LLM, a second set of question and answer pairs based on the top questions.
Owner:JPMORGAN CHASE BANK NA

Intelligent auxiliary system for bid inviting and purchasing business

The invention discloses an intelligent auxiliary system for bid inviting and purchasing business. Comprising the following steps: S1, responding to a bidding document making request of a user, and loading a preset data model associated with a specific purchasing project machine; s2, receiving structured bidding data input by a user through a data field, and receiving an unstructured original file uploaded by the user; s3, based on the structured bidding data and the unstructured original file, generating a unified bidding data packet; s4, performing encryption and fragmentation processing on the unified bidding data packet, and uploading the unified bidding data packet to a bidding server; s5, the received data fragments are recombined, decrypted and verified, and structured bidding data are separated from the bidding data packet passing verification; according to the method, the self-contained data object of deep mapping of the structured metadata and the unstructured load data is used as a core to drive subsequent secure transmission and back-end intelligent processing, so that linkage and collaborative optimization of the whole data processing flow are realized.
Owner:HENAN TENGLONG INFORMATION ENG

Multimodal retrieval based on relational modeling of atomic document elements

An atomic relational retrieval system can determine a type of modality for each document of a plurality of documents having unstructured data. The system can route each document to a parser based on the type of modality. The system can parse at least the unstructured data of each document according to an atomic unit type to extract a plurality of atomic units from the document and a plurality of attributes of each atomic unit. The system can update a table in a relational database to include a record for each atomic unit, the record including a unique identifier of the atomic unit, a document identifier linking the atomic unit to its source document, and the plurality of attributes. The system can output, in response to a request for a chunk of one or more atomic units, at least one record corresponding to the chunk, the chunk is dynamically defined.
Owner:COHERE HEALTH INC

An intelligent document understanding and automatic form filling system based on a multi-modal large model

The application discloses an intelligent document understanding and automatic form filling system based on a multimodal large model, comprising a multimodal document coding module, a document-form semantic alignment module, an information fusion module, a form constraint checking module and an intelligent form filling module. The application fuses multimodal information such as document images, text content, table structures, form fields and knowledge bases, constructs a multimodal document encoder, a document-form semantic alignment module, a cross-document information fusion module, a form constraint understanding module and an intelligent form filling module, realizes extraction of structured information from a source document, understanding of field meanings and filling rules of a target form, automatic matching and filling of information, supports cross-document information completion, format conversion and consistency verification.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD