Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

212 results about "Text document" patented technology

A text document may have a TXT, DAT, LOG or HTML file extension. You can create one in a text editor of your choice. Since text documents do not include special formatting, they appear as plain text within an application.

Large language model (LLM) message generation

A large language model (LLM) message generation system and method tor generating a response to a user message. The method includes: obtaining a text document specified by a user; chunking the text document into a plurality of chunks; generating text gloss data for each of the plurality of chunks based on the chunk and a predetermined prompt; storing the text gloss data for each chunk into a vector data store along with an identifier for the chunk and / or tine chunk itself; receiving a user message; querying the vector data store with a vector data store query, wherein the vector data store query is generated based on the user message; obtaining chunk(s) based on querying the vector data store with the vector data store query; generating a language model input based on the one or more identified chunks; and generating a response message based on inputting the language model input.
Owner:FORGEN AI LLC

Structured output of duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

Techniques described herein provide for generation of structured output for documents identified using automated near-duplicate detection. In one example, a system can receive a set of documents including at least one pair of similar documents determined to be similar to one another based on similarity scores generated using a predefined similarity scoring technique. The system can generate document groups by merging together pairs of documents that share at least one document. The system can, for each of the document groups, identify a representative document for the document group. The system can generate an output for display including a section for each document group, in which each section includes the representative document for the document group and, for each document in the document group, the similarity score relative to the representative document for the document group.
Owner:SAS INSTITUTE INC

Generative search engine text documents

This disclosure describes utilizing a generative document system to dynamically build and provide generative text documents using one or more generative artificial intelligence (AI) models. For example, the generative document system efficiently utilizes various systems and one or more generative AI models to determine intents and topics, curate topic sections, and generate a generative text document that includes a directed answer along with select curated topic sections for search queries. In various implementations, the generative document system performs additional actions that enhance the efficiency and accuracy of operations used to produce generative text documents. Additionally, in many cases, these generative text documents provide a foundation for providing an interactive, intuitive, wide-ranging, and flexible curation of answers to users that address the corresponding search queries.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Automated near-duplicate detection for text documents

Techniques described herein provide for automated detection of near-duplicate documents. In one example, a system can cluster documents into a set of clusters based on character frequencies associated with the documents. For a given cluster, the system can generate first similarity scores associated with every pair of documents in the cluster. The system can then select a filtered group of documents associated with first similarity scores that meet or exceed a first predefined similarity threshold. Next, the system can convert the filtered group of documents into matrix representations. The system can generate second similarity scores for every pair of matrix representations. The system can then identify documents, from among the filtered group of documents, associated with second similarity scores that meet or exceed a second predefined similarity threshold. The identified documents can be duplicate or near-duplicate text documents.
Owner:SAS INSTITUTE INC

Efficient and accurate regional explanation technique for NLP models

Herein are techniques for topic modeling and content perturbation that provide machine learning (ML) explainability (MLX) for natural language processing (NLP). A computer hosts an ML model that infers an original inference for each of many text documents that contain many distinct terms. To each text document (TD) is assigned, based on terms in the TD, a topic that contains a subset of the distinct terms. In a perturbed copy of each TD, a perturbed subset of the distinct terms is replaced. For the perturbed copy of each TD, the ML model infers a perturbed inference. For TDs of a topic, the computer detects that a difference between original inferences of the TDs of the topic and perturbed inferences of the TDs of the topic exceeds a threshold. Based on terms in the TDs of the topic, the topic is replaced with multiple, finer-grained new topics. After sufficient topic modeling, a regional explanation of the ML model is generated.
Owner:ORACLE INT CORP

Retrieval enhancement generation method based on multi-modal document

The invention discloses a retrieval enhancement generation method based on a multi-modal document. The method comprises the following steps: S1, data construction; s2, feature extraction of a multi-modal knowledge searcher is carried out; s3, carrying out feature mapping on the multi-modal knowledge retriever; s4, calculating the relevancy of the multi-modal knowledge retriever; and S5, multi-modal answer generation: the large language model generates a text reply according to multi-modal input. According to the method, a multi-modal document formed by combining pictures and texts is used as a knowledge carrier, and a multi-modal retrieval enhancement generation scheme is designed. Compared with an existing end-to-end model scheme, the scheme is based on a retrieval enhancement generation framework, and the accuracy and interpretability of answers are guaranteed; compared with a retrieval enhancement generation scheme using a text document as a knowledge carrier, the scheme has the advantages that visual information is added to the document to construct the multi-modal document, the knowledge retriever and the answer generator are improved to utilize the multi-modal document, and the accuracy of the knowledge-intensive visual question and answer task is improved.
Owner:RENMIN UNIVERSITY OF CHINA

Multi-field scientific knowledge base automatic construction method, system, equipment and medium

The invention belongs to the technical field of natural language processing and knowledge engineering, and discloses a method, a system, equipment and a medium for automatically constructing a multi-field scientific knowledge base, and the method comprises the following steps: based on an input identifier list or a field search word, retrieving a full text of a literature, analyzing and converting the full text into a structured text; combining a pre-configured large language model with a field cue word, extracting key information in the structured text, and outputting structured data; the automatic script filters irrelevant, repeated or incomplete structured data according to a pre-configured filtering rule; performing standardization processing on the filtered structured data so as to realize the consistency and comparability of the data; and inserting the standardized structured data into an interactive knowledge base or database, establishing data association, and verifying association logic. The method supports cross-field rapid adaptation, solves the problems of low efficiency and high error rate of a traditional method, and can be widely applied to the fields of biomedicine, material science, synthetic biology and the like.
Owner:TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI

Intelligent model selection system for style-specific digital content generation

Aspects of the present disclosure provide systems, methods, and computer-readable storage media that support intelligent model selection for style-specific digital content generation. For example, a system that provides a digital content generation service may include a trained style detection model may receive reference digital content items from a user and extract a user style embedding that represents a style preference of the user. In some implementations, the reference digital content items may include text documents or images provided or selected by the user. The system may compare the user style embedding to a plurality of model style embeddings that each correspond to a respective generative artificial intelligence (AI) model to generate a ranked list of generative AI models. The system may access one or more highest ranked generative AI models from the ranked list to generate novel digital content based on a prompt from the user.
Owner:ACCENTURE GLOBAL SOLUTIONS LTD

Information extraction for unstructured text documents

Training and using a machine learning model for data extraction is provided. The method comprises receiving keys of interest received from a user through an interface and receiving a batch of documents containing unstructured text. Unstructured text of a first document is processed to extract structured text. The model predicts text classifications of the structured text according to the keys of interest. The predicted text classifications are output to the user through the interface. Annotations to correct any incorrect predictions are received from the user, and the model is retrained according to the annotations. The above steps are repeated for less than ten additional documents from the batch until the model has been trained to predict text classifications with a specified level of accuracy. The trained model then classifies extracted structured text in the remaining documents in the batch.
Owner:S&P GLOBAL INC

Mutually generative artificial intelligence system based on multi-dimensional spatiotemporal information vector graphics

Disclosed is a mutually generative artificial intelligence system based on multi-dimensional spatiotemporal information vector graphics, relating to the field of artificial intelligence and engineering applications. Based on a geographic information system (GIS) or a computer-aided design (CAD) platform and a data source, a multi-dimensional vector spatiotemporal large model terminal, a multi-dimensional spatiotemporal information processing agent terminal, and an intelligent information system application terminal are constructed. For multi-dimensional spatiotemporal data such as two-dimensional and three-dimensional vectors and temporal states, a multimodal spatiotemporal large model having an understanding capacity for an engineering professional knowledge system, a data processing flow, multi-dimensional vector graphics, and thematic graphics-text documents is pre-trained to achieve the mutual expression and generation of engineering multi-dimensional vector graphics and thematic graphics-text documents and to form an intelligent engineering graphic data processing application.
Owner:BEIJING LONGRUAN TECHNOLOGIES INC +1

Document image-text integrated intelligent understanding and processing method and system

The invention provides a document image-text integrated intelligent understanding and processing method and system, and relates to the technical field of artificial intelligence. Comprising the following steps: separating image-text elements and establishing context association; calling a graph vectorization engine, and executing intelligent vector conversion processing on the rasterized illustration; performing hierarchical classification, attribute endowing and structured reconstruction on the vector primitives in combination with text semantic analysis to generate vector entity objects with complete attributes; the text elements and the vector entity objects are jointly input into a pre-trained multi-mode large language model, unified knowledge expression containing vectors, topology, attributes and document contexts is generated, and applications such as question answering, editing, abstracting and reporting of image-text integration are supported. According to the method, the problems of graph semantic loss, difficulty in fine analysis and the like caused by vector-to-grid illustration in papers, reports and other types of documents are solved, the overall understanding depth and interpretability of AI for complex image-text documents are remarkably improved, and high-performance semantic understanding capacity is provided for document-based training and intelligent processing.
Owner:BEIJING LONGRUAN TECHNOLOGIES INC +1

System and method for automated software development through ai based requirements gathering for application design, reverse engineering and redesign

A system and method for an AI-based Automated Software Application Development Platform which intelligently gathers requirements from a business analyst and generates a runnable software application from those requirements. The AI-based platform enables the business analyst with no technical knowledge to visually model their requirements rapidly without requiring any programming Using the abstract requirements architecture and the platform, the analyst creates requirements models for the application's business processes and business tasks for each of the processes, including defining the supporting information needs. The system intelligently assists the user in defining the information needs for each of the business processes and their tasks. Once the requirements models are created by the analyst with the help of the platform, the platform automatically designs and creates an appropriate persistent canonical data model for the application from the defined information needs. The platform then creates a runnable software application. The system also provides a template for the user to input their requirements in descriptive text, and accepts requirements document for applications as written free form text or text documents written in English as input, and in a non-English language that the system automatically translates them into English. The system provides a method for reverse engineering an existing application, including generating a requirements document for the existing application such as a legacy modernization or digital transformation an old applications for redesigning, updating or revising them to accommodate a new requirements set for such applications.
Owner:PRISM X INC

Prompt engineering and automated quality assessment for large language models

Various embodiments of the present disclosure provide prompt engineering and text quality assessment techniques for improving generative text outputs. The techniques may include identifying an initial document subset for a generative text request that includes a request to generate a generative text document based on one or more request text fields. The techniques may include generating a contextual classification for the one or more request text fields and identifying a refined document subset based on the contextual classification. The techniques may include generating one or more request field embeddings respectively corresponding to the one or more request text fields and identifying a prompt document subset based on the one or more request field embeddings. The techniques may include generating, using a large language model, one or more generative text fields using a generative model prompt based on the prompt document subset and the one or more request text fields.
Owner:UNITEDHEALTH GROUP INC

Text string comparison for duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

ActiveUS20250231993A1Pattern recognitionBoilerplate text
Techniques described herein provide for text string comparison for documents identified using automated near-duplicate detection. In one example, a system can receive a pair of documents. The system can extract text strings from the documents. The system can normalize the extracted text strings using a predefined normalization scheme. The system can identify boilerplate text segments in the normalized text strings. The system can remove the boilerplate text segments from the normalized text strings to generate filtered text strings. The system can divide the filtered text strings by identifying section indicators. The system can, for each section, generate groupings of text strings and determine a similarity score between each pair of corresponding groupings to identify matching groupings of text strings. The system can generate an output for display showing the visual indications of the matched groupings of text strings.
Owner:SAS INSTITUTE INC

User question and answer method and device, equipment, storage medium and product

The invention discloses a user question and answer method and device, equipment, a storage medium and a product, and relates to the technical field of artificial intelligence, the method comprises the following steps: analyzing a knowledge text document to obtain text knowledge; dividing the text knowledge into blocks according to the dynamically adjustable model parameter threshold and generating a knowledge base; performing similarity retrieval on the obtained user input question and knowledge blocks in a knowledge base to obtain associated knowledge blocks; and inputting the associated knowledge blocks and the user input question into a large language generation model to generate a question and answer result. According to the method, the text knowledge is partitioned according to the dynamically adjustable model parameter threshold value, so that the defect of low accuracy during fixed parameter partitioning in the existing large model retrieval enhancement technology is overcome, the partitioning method is more flexible, and dynamic adjustment can be performed according to different texts and contexts, so that the accuracy and effectiveness of partitioning are improved. And then a knowledge base is generated for retrieval based on knowledge blocks obtained by partitioning, and a question and answer result with higher accuracy is generated.
Owner:中国移动通信有限公司政企客户分公司 +1

Hallucination detection and remediation in text generation interface systems

Enumerated source text passages may be determined based on one or more source text documents. The enumerated source text passages may include source text passage identifiers uniquely identifying the passages. A novel text passage including novel text portions may be determined based on a query and the enumerated source text passages. One or more of the novel text portions may be verified by a large language model to produce text verification information. A novel text generation message including novel text generated by the large language model may be determined based on the text verification information and sent to a client machine.
Owner:CASETEXT INC

Large language model fine tuning and evaluation method based on retrieval enhancement generation

The invention relates to the technical field of artificial intelligence, and provides a retrieval enhancement generation-based large language model fine tuning and evaluation method, which comprises the following steps of: segmenting an input original text document content into a plurality of independent paragraphs according to a preset character level segmentation rule; extracting entities and relationships among the entities from the plurality of independent paragraphs, and constructing a knowledge graph; based on the knowledge graph, noise data, complex data, rejection data and multi-hop data are generated through a noise sampling module, a complex attribution module, a rejection analysis module and a thinking reasoning module; performing format conversion on the noise data, the complex data, the rejection data, the multi-hop data and the general data to obtain fine-tuning data; performing fine tuning on the large language model according to the fine tuning data to obtain a fine-tuned retrieval enhancement generation model; according to the answer correlation index and the answer similarity index, the fine-tuned retrieval enhancement generation model is evaluated, and the accuracy, reliability and reasoning ability of the retrieval enhancement generation model are improved.
Owner:NINGBO TELIAN INFORMATION TECH CO LTD

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Systems, methods, and graphical user interfaces for predicting and analyzing action likelihood

A computer-implemented system, computer-implemented method, and computer-program product includes obtaining a text document that includes text describing an action; extracting one or more action tokens from the text document; executing a plurality of linguistic pattern searches that search the text document for one or more likelihood tokens associated with the one or more action tokens; classifying the action to a likelihood category associated with a respective linguistic pattern search of the plurality of linguistic pattern searches that identified the one or more likelihood tokens; classifying the text document to a respective domain; computing a priority value of the action described in the text document based on an input of the likelihood category and the respective domain; and generating a priority summary artifact that visually prioritizes the text document over one or more other text documents when the priority value of the action satisfies a predefined maximum priority threshold value.
Owner:SAS INSTITUTE INC

Expediting automated near-duplicate detection for new text documents

Techniques described herein provide for automated near-duplicate detection for new text documents given text documents that were previously processed using automated near-duplicate detection for text documents. In one example, a system can receive new documents and documents that were previously processed using a predefined processing technique for automated near-duplicate detection. The system can process the new documents and cluster the new documents into multiple predefined clusters previously identified using the predefined processing technique. For each predefined cluster including at least one new document, the system can generate document groups by determining similarity scores using the predefined processing technique as applied to the documents in the predefined clusters. The system can identify a representative document for each document group and generate an output data structure including the document groups and the representative document for each group.
Owner:SAS INSTITUTE INC

Electric power project material inventory intelligent generation method and system based on multi-modal document analysis and knowledge graph semantic mapping

The invention discloses an electric power project material inventory intelligent generation method and system based on multi-modal document analysis and knowledge graph semantic mapping, and belongs to the field of artificial intelligence and electric power engineering. Comprising the steps that in the first stage, unified analysis and semantic coding are conducted on text documents, engineering drawings and table data, and a semi-structured original information set is generated; in the second stage, named entity recognition and type labeling are carried out on the original information set based on a two-way long-short term memory network and a conditional random field hybrid model, a structured entity set is generated, an electric power material field knowledge graph is constructed on the basis, and a standard key point rule base is established; semantic mapping and information completion are carried out through a multi-level matching mechanism and knowledge reasoning, and a standardized material inventory is generated after quality control. According to the invention, intelligent and automatic generation of the electric power project material inventory is realized, and the accuracy and efficiency of material inventory generation are improved.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD

Multi-language intelligent analysis system for medical documents

The invention provides a medical document multi-language intelligent analysis system, relates to the field of language translation, and improves the accuracy and efficiency of professional term translation. The method comprises the following steps of: firstly, performing language recognition on an original text document by a recognition module through text unitization and context vector generation, and matching a corresponding corpus; then, a translation module carries out lexical element alignment on the source lexical elements through term bank injection and an AI model, the translation process is automatically optimized, and accurate translation of the terminologies is ensured; and finally, the reconstruction module accurately replaces corresponding contents in the original text document with translation output through the mapping file, so as to ensure that the document format and typesetting are consistent. Through the automatic and optimized translation process, the quality and efficiency of professional term translation are remarkably improved, manual intervention is reduced, and the translation requirement of a high professional standard is met.
Owner:LUNAN PHARMA GROUP CORPORATION +2

Multi-modal retrieval enhancement generation method and system based on multi-expert model

The invention discloses a multi-modal retrieval enhancement generation method and system based on a multi-expert model. The multi-modal retrieval enhancement generation method comprises the following steps: S1, processing input multi-modal data by using a visual language model and an audio understanding model as professional understanding models, and converting image, video and audio information into uniform text representation; s2, generating a structured text description through a cross-modal feature alignment mechanism, and forming a standardized text unit of multi-modal information; s3, based on a BERT generation embedded type, performing vectorization processing on a text query input by a user and a converted text; s4, performing similarity search in the vector database, and retrieving a text document most relevant to query; and S5, splicing the retrieved text and the user query, and inputting the spliced text and user query into the large language model to generate a final answer. According to the multi-modal retrieval enhancement generation method, a professional understanding model is innovatively introduced to serve as a different-modal encoder, the multi-modal information processing capacity of a large model is improved, and real cross-modal understanding is achieved.
Owner:GUANGZHOU BINGO SOFTWARE

Business demand analysis method and device and computer readable storage medium

The invention provides a service demand analysis method and device and a computer readable storage medium, and the method comprises the steps: obtaining all dialogue voice data discussed by a service demand, and converting all dialogue voice data into corresponding dialogue text documents; generating a preliminary demand document based on all dialogue text documents; performing information extraction on the preliminary demand document by using a big and small language model according to a set demand template to obtain specific information, and filling the specific information to a position corresponding to the set demand template to obtain a demand document; and performing function point analysis on the demand document by using the big and small language models to generate a function demand specification. The problem that in the prior art, work such as collection and summarization, key point extraction, induction analysis and content checking of massive multi-source heterogeneous original documents consumes a large amount of time cost is solved.
Owner:中国邮政储蓄银行股份有限公司

Text classification via term mapping and machine-learning classification model

A method, apparatus, and computer-readable medium are described that identify subject matter of text using identified groups and a machine-learning model. Using the combination of the identified subject matter and the machine-learning model, classifications may be adjusted over time. Based on the adjusted classifications, the machine-learning model may be retrained to better classify previously unclassified text. One or more benefits may include better summarization of text documents and / or better training of machine-learning models that are then used to assist in the summarization of the text documents. The resulting classifications may be used to improve resource allocations for future tasks.
Owner:CAPITAL ONE SERVICES LLC

College student competition intelligent auxiliary system based on AI large model

The invention belongs to the technical field of education informatization, and discloses an AI large model-based college student competition intelligent auxiliary system, which provides a one-stop intelligent tutoring service for college student competition through a material uploading module, a multi-modal analysis module, a case library retrieval module, an AI project diagnosis module, a material generation module and the like. The system can automatically identify text documents, images and code contents and generate structured data of projects; and based on a preset case library, project similarity matching is realized, and reference experience information is extracted. The AI large model is used for performing multi-dimensional scoring on the project, a score deduction basis is provided, operable improvement suggestions are generated, and meanwhile, competition documents such as road performance materials, project abstracts or commercial schedules can be automatically generated. According to the method, the project preparation efficiency can be remarkably improved, the material quality is optimized, accurate and personalized competition support is provided for college students, and the method has good practical value and popularization significance.
Owner:BEIJING MORNINGSTAR VENTURE CAPITAL TECHNOLOGY CO LTD

System and method for software development on mobile devices

An interactive coding environment is provided via an integrated mobile application. The interactive coding environment includes multiple user interface input elements including, but not limited to, a text editor, a virtual keyboard, a virtual autocomplete toolbar, a quick actions toolbar, and a virtual joystick. A user interacts with the coding environment via their user device, such as a mobile device. While coding within the text editor, the user is provided with auto-complete code suggestions and / or code snippets. Suggested inputs are obtained from a predictive model based on the user's input. A virtual joystick is provided to enable easy navigation within a text document shown within the text editor.
Owner:REPLIT INC

Method and system for extracting information from documents with varying formats

Certain aspects of the disclosure provide a method for extracting attributes from documents with varying formats, layouts and complexities. The method displays a user interface (UI) that enables a user to obtain an unstructured document from a knowledge base. The method converts the unstructured document into a text document using a text recognition. The method obtains, as output from a large language model (LLM), an extracted page attribute from the text document. The extracted page attribute contains a first type of information recorded in text on a single page of the text document. The extracted document attribute contains a second type of information recorded in text on more than one page of the text document. The method obtains, as output from the LLM, an extracted document attribute from the text document. The extracted page attribute and the extracted document attribute are displayed in the UI.
Owner:SCHLUMBERGER TECH CORP

Document analysis method, device and equipment and computer readable storage medium

The invention discloses a document analysis method, device and equipment and a computer readable storage medium. The method comprises the following steps: when the type of an original file is a spreadsheet file type and a table in the original file has merged cells, determining the content of each cell in the table, the row number i and the column number j of each non-merged cell in the table and the area coordinate of each merged cell in the table; for each non-merged cell, filling the content of the non-merged cell into the cells in the ith row and the jth column in the correction table; for each merged cell, filling the content of the merged cell into a target cell corresponding to the region coordinate in the correction table; and after all the cells are traversed, embedding the obtained correction table into the plain text document. By means of the method and device, the situation that the format of the content embedded into the plain text document is disordered or information is lost when compared with that of a table in an original file is avoided to a great extent.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD