Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

159 results about "Text document" patented technology

A text document may have a TXT, DAT, LOG or HTML file extension. You can create one in a text editor of your choice. Since text documents do not include special formatting, they appear as plain text within an application.

Large language model (LLM) message generation

A large language model (LLM) message generation system and method tor generating a response to a user message. The method includes: obtaining a text document specified by a user; chunking the text document into a plurality of chunks; generating text gloss data for each of the plurality of chunks based on the chunk and a predetermined prompt; storing the text gloss data for each chunk into a vector data store along with an identifier for the chunk and / or tine chunk itself; receiving a user message; querying the vector data store with a vector data store query, wherein the vector data store query is generated based on the user message; obtaining chunk(s) based on querying the vector data store with the vector data store query; generating a language model input based on the one or more identified chunks; and generating a response message based on inputting the language model input.
Owner:FORGEN AI LLC

Generative search engine text documents

This disclosure describes utilizing a generative document system to dynamically build and provide generative text documents using one or more generative artificial intelligence (AI) models. For example, the generative document system efficiently utilizes various systems and one or more generative AI models to determine intents and topics, curate topic sections, and generate a generative text document that includes a directed answer along with select curated topic sections for search queries. In various implementations, the generative document system performs additional actions that enhance the efficiency and accuracy of operations used to produce generative text documents. Additionally, in many cases, these generative text documents provide a foundation for providing an interactive, intuitive, wide-ranging, and flexible curation of answers to users that address the corresponding search queries.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Efficient and accurate regional explanation technique for NLP models

Herein are techniques for topic modeling and content perturbation that provide machine learning (ML) explainability (MLX) for natural language processing (NLP). A computer hosts an ML model that infers an original inference for each of many text documents that contain many distinct terms. To each text document (TD) is assigned, based on terms in the TD, a topic that contains a subset of the distinct terms. In a perturbed copy of each TD, a perturbed subset of the distinct terms is replaced. For the perturbed copy of each TD, the ML model infers a perturbed inference. For TDs of a topic, the computer detects that a difference between original inferences of the TDs of the topic and perturbed inferences of the TDs of the topic exceeds a threshold. Based on terms in the TDs of the topic, the topic is replaced with multiple, finer-grained new topics. After sufficient topic modeling, a regional explanation of the ML model is generated.
Owner:ORACLE INT CORP

Multi-field scientific knowledge base automatic construction method, system, equipment and medium

The invention belongs to the technical field of natural language processing and knowledge engineering, and discloses a method, a system, equipment and a medium for automatically constructing a multi-field scientific knowledge base, and the method comprises the following steps: based on an input identifier list or a field search word, retrieving a full text of a literature, analyzing and converting the full text into a structured text; combining a pre-configured large language model with a field cue word, extracting key information in the structured text, and outputting structured data; the automatic script filters irrelevant, repeated or incomplete structured data according to a pre-configured filtering rule; performing standardization processing on the filtered structured data so as to realize the consistency and comparability of the data; and inserting the standardized structured data into an interactive knowledge base or database, establishing data association, and verifying association logic. The method supports cross-field rapid adaptation, solves the problems of low efficiency and high error rate of a traditional method, and can be widely applied to the fields of biomedicine, material science, synthetic biology and the like.
Owner:TIANJIN INST OF IND BIOTECH CHINESE ACADEMY OF SCI

Intelligent model selection system for style-specific digital content generation

Aspects of the present disclosure provide systems, methods, and computer-readable storage media that support intelligent model selection for style-specific digital content generation. For example, a system that provides a digital content generation service may include a trained style detection model may receive reference digital content items from a user and extract a user style embedding that represents a style preference of the user. In some implementations, the reference digital content items may include text documents or images provided or selected by the user. The system may compare the user style embedding to a plurality of model style embeddings that each correspond to a respective generative artificial intelligence (AI) model to generate a ranked list of generative AI models. The system may access one or more highest ranked generative AI models from the ranked list to generate novel digital content based on a prompt from the user.
Owner:ACCENTURE GLOBAL SOLUTIONS LTD

Mutually generative artificial intelligence system based on multi-dimensional spatiotemporal information vector graphics

Disclosed is a mutually generative artificial intelligence system based on multi-dimensional spatiotemporal information vector graphics, relating to the field of artificial intelligence and engineering applications. Based on a geographic information system (GIS) or a computer-aided design (CAD) platform and a data source, a multi-dimensional vector spatiotemporal large model terminal, a multi-dimensional spatiotemporal information processing agent terminal, and an intelligent information system application terminal are constructed. For multi-dimensional spatiotemporal data such as two-dimensional and three-dimensional vectors and temporal states, a multimodal spatiotemporal large model having an understanding capacity for an engineering professional knowledge system, a data processing flow, multi-dimensional vector graphics, and thematic graphics-text documents is pre-trained to achieve the mutual expression and generation of engineering multi-dimensional vector graphics and thematic graphics-text documents and to form an intelligent engineering graphic data processing application.
Owner:BEIJING LONGRUAN TECHNOLOGIES INC +1

Document image-text integrated intelligent understanding and processing method and system

The invention provides a document image-text integrated intelligent understanding and processing method and system, and relates to the technical field of artificial intelligence. Comprising the following steps: separating image-text elements and establishing context association; calling a graph vectorization engine, and executing intelligent vector conversion processing on the rasterized illustration; performing hierarchical classification, attribute endowing and structured reconstruction on the vector primitives in combination with text semantic analysis to generate vector entity objects with complete attributes; the text elements and the vector entity objects are jointly input into a pre-trained multi-mode large language model, unified knowledge expression containing vectors, topology, attributes and document contexts is generated, and applications such as question answering, editing, abstracting and reporting of image-text integration are supported. According to the method, the problems of graph semantic loss, difficulty in fine analysis and the like caused by vector-to-grid illustration in papers, reports and other types of documents are solved, the overall understanding depth and interpretability of AI for complex image-text documents are remarkably improved, and high-performance semantic understanding capacity is provided for document-based training and intelligent processing.
Owner:BEIJING LONGRUAN TECHNOLOGIES INC +1

System and method for automated software development through ai based requirements gathering for application design, reverse engineering and redesign

A system and method for an AI-based Automated Software Application Development Platform which intelligently gathers requirements from a business analyst and generates a runnable software application from those requirements. The AI-based platform enables the business analyst with no technical knowledge to visually model their requirements rapidly without requiring any programming Using the abstract requirements architecture and the platform, the analyst creates requirements models for the application's business processes and business tasks for each of the processes, including defining the supporting information needs. The system intelligently assists the user in defining the information needs for each of the business processes and their tasks. Once the requirements models are created by the analyst with the help of the platform, the platform automatically designs and creates an appropriate persistent canonical data model for the application from the defined information needs. The platform then creates a runnable software application. The system also provides a template for the user to input their requirements in descriptive text, and accepts requirements document for applications as written free form text or text documents written in English as input, and in a non-English language that the system automatically translates them into English. The system provides a method for reverse engineering an existing application, including generating a requirements document for the existing application such as a legacy modernization or digital transformation an old applications for redesigning, updating or revising them to accommodate a new requirements set for such applications.
Owner:PRISM X INC

Prompt engineering and automated quality assessment for large language models

Various embodiments of the present disclosure provide prompt engineering and text quality assessment techniques for improving generative text outputs. The techniques may include identifying an initial document subset for a generative text request that includes a request to generate a generative text document based on one or more request text fields. The techniques may include generating a contextual classification for the one or more request text fields and identifying a refined document subset based on the contextual classification. The techniques may include generating one or more request field embeddings respectively corresponding to the one or more request text fields and identifying a prompt document subset based on the one or more request field embeddings. The techniques may include generating, using a large language model, one or more generative text fields using a generative model prompt based on the prompt document subset and the one or more request text fields.
Owner:UNITEDHEALTH GROUP INC

Hallucination detection and remediation in text generation interface systems

Enumerated source text passages may be determined based on one or more source text documents. The enumerated source text passages may include source text passage identifiers uniquely identifying the passages. A novel text passage including novel text portions may be determined based on a query and the enumerated source text passages. One or more of the novel text portions may be verified by a large language model to produce text verification information. A novel text generation message including novel text generated by the large language model may be determined based on the text verification information and sent to a client machine.
Owner:CASETEXT INC

Large language model fine tuning and evaluation method based on retrieval enhancement generation

The invention relates to the technical field of artificial intelligence, and provides a retrieval enhancement generation-based large language model fine tuning and evaluation method, which comprises the following steps of: segmenting an input original text document content into a plurality of independent paragraphs according to a preset character level segmentation rule; extracting entities and relationships among the entities from the plurality of independent paragraphs, and constructing a knowledge graph; based on the knowledge graph, noise data, complex data, rejection data and multi-hop data are generated through a noise sampling module, a complex attribution module, a rejection analysis module and a thinking reasoning module; performing format conversion on the noise data, the complex data, the rejection data, the multi-hop data and the general data to obtain fine-tuning data; performing fine tuning on the large language model according to the fine tuning data to obtain a fine-tuned retrieval enhancement generation model; according to the answer correlation index and the answer similarity index, the fine-tuned retrieval enhancement generation model is evaluated, and the accuracy, reliability and reasoning ability of the retrieval enhancement generation model are improved.
Owner:NINGBO TELIAN INFORMATION TECH CO LTD

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Electric power project material inventory intelligent generation method and system based on multi-modal document analysis and knowledge graph semantic mapping

The invention discloses an electric power project material inventory intelligent generation method and system based on multi-modal document analysis and knowledge graph semantic mapping, and belongs to the field of artificial intelligence and electric power engineering. Comprising the steps that in the first stage, unified analysis and semantic coding are conducted on text documents, engineering drawings and table data, and a semi-structured original information set is generated; in the second stage, named entity recognition and type labeling are carried out on the original information set based on a two-way long-short term memory network and a conditional random field hybrid model, a structured entity set is generated, an electric power material field knowledge graph is constructed on the basis, and a standard key point rule base is established; semantic mapping and information completion are carried out through a multi-level matching mechanism and knowledge reasoning, and a standardized material inventory is generated after quality control. According to the invention, intelligent and automatic generation of the electric power project material inventory is realized, and the accuracy and efficiency of material inventory generation are improved.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD

Multi-language intelligent analysis system for medical documents

The invention provides a medical document multi-language intelligent analysis system, relates to the field of language translation, and improves the accuracy and efficiency of professional term translation. The method comprises the following steps of: firstly, performing language recognition on an original text document by a recognition module through text unitization and context vector generation, and matching a corresponding corpus; then, a translation module carries out lexical element alignment on the source lexical elements through term bank injection and an AI model, the translation process is automatically optimized, and accurate translation of the terminologies is ensured; and finally, the reconstruction module accurately replaces corresponding contents in the original text document with translation output through the mapping file, so as to ensure that the document format and typesetting are consistent. Through the automatic and optimized translation process, the quality and efficiency of professional term translation are remarkably improved, manual intervention is reduced, and the translation requirement of a high professional standard is met.
Owner:LUNAN PHARMA GROUP CORPORATION +2

Multi-modal retrieval enhancement generation method and system based on multi-expert model

The invention discloses a multi-modal retrieval enhancement generation method and system based on a multi-expert model. The multi-modal retrieval enhancement generation method comprises the following steps: S1, processing input multi-modal data by using a visual language model and an audio understanding model as professional understanding models, and converting image, video and audio information into uniform text representation; s2, generating a structured text description through a cross-modal feature alignment mechanism, and forming a standardized text unit of multi-modal information; s3, based on a BERT generation embedded type, performing vectorization processing on a text query input by a user and a converted text; s4, performing similarity search in the vector database, and retrieving a text document most relevant to query; and S5, splicing the retrieved text and the user query, and inputting the spliced text and user query into the large language model to generate a final answer. According to the multi-modal retrieval enhancement generation method, a professional understanding model is innovatively introduced to serve as a different-modal encoder, the multi-modal information processing capacity of a large model is improved, and real cross-modal understanding is achieved.
Owner:GUANGZHOU BINGO SOFTWARE

College student competition intelligent auxiliary system based on AI large model

The invention belongs to the technical field of education informatization, and discloses an AI large model-based college student competition intelligent auxiliary system, which provides a one-stop intelligent tutoring service for college student competition through a material uploading module, a multi-modal analysis module, a case library retrieval module, an AI project diagnosis module, a material generation module and the like. The system can automatically identify text documents, images and code contents and generate structured data of projects; and based on a preset case library, project similarity matching is realized, and reference experience information is extracted. The AI large model is used for performing multi-dimensional scoring on the project, a score deduction basis is provided, operable improvement suggestions are generated, and meanwhile, competition documents such as road performance materials, project abstracts or commercial schedules can be automatically generated. According to the method, the project preparation efficiency can be remarkably improved, the material quality is optimized, accurate and personalized competition support is provided for college students, and the method has good practical value and popularization significance.
Owner:BEIJING MORNINGSTAR VENTURE CAPITAL TECHNOLOGY CO LTD

System and method for software development on mobile devices

An interactive coding environment is provided via an integrated mobile application. The interactive coding environment includes multiple user interface input elements including, but not limited to, a text editor, a virtual keyboard, a virtual autocomplete toolbar, a quick actions toolbar, and a virtual joystick. A user interacts with the coding environment via their user device, such as a mobile device. While coding within the text editor, the user is provided with auto-complete code suggestions and / or code snippets. Suggested inputs are obtained from a predictive model based on the user's input. A virtual joystick is provided to enable easy navigation within a text document shown within the text editor.
Owner:REPLIT INC

Method and system for extracting information from documents with varying formats

Certain aspects of the disclosure provide a method for extracting attributes from documents with varying formats, layouts and complexities. The method displays a user interface (UI) that enables a user to obtain an unstructured document from a knowledge base. The method converts the unstructured document into a text document using a text recognition. The method obtains, as output from a large language model (LLM), an extracted page attribute from the text document. The extracted page attribute contains a first type of information recorded in text on a single page of the text document. The extracted document attribute contains a second type of information recorded in text on more than one page of the text document. The method obtains, as output from the LLM, an extracted document attribute from the text document. The extracted page attribute and the extracted document attribute are displayed in the UI.
Owner:SCHLUMBERGER TECH CORP

Document analysis method, device and equipment and computer readable storage medium

The invention discloses a document analysis method, device and equipment and a computer readable storage medium. The method comprises the following steps: when the type of an original file is a spreadsheet file type and a table in the original file has merged cells, determining the content of each cell in the table, the row number i and the column number j of each non-merged cell in the table and the area coordinate of each merged cell in the table; for each non-merged cell, filling the content of the non-merged cell into the cells in the ith row and the jth column in the correction table; for each merged cell, filling the content of the merged cell into a target cell corresponding to the region coordinate in the correction table; and after all the cells are traversed, embedding the obtained correction table into the plain text document. By means of the method and device, the situation that the format of the content embedded into the plain text document is disordered or information is lost when compared with that of a table in an original file is avoided to a great extent.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

Automated topic modelling and visualization based upon service phase

PendingUS20260211935A1Data visualizationDocumentation
Disclosed in some examples are methods, systems, devices, and machine-readable mediums which create various data visualizations from a corpus of raw data collected from one or more sources. For example, natural language text documents of a corpus may be classified based upon the topic of the data. The documents may be labeled with the phase during which the documents were collected or observed. A visualization may then be generated which shows a correlation between the phase and the topics observed. For example, a number of times a particular topic appeared in a particular phase. The visualization may be two-dimensional, three-dimensional, or the like.
Owner:WELLS FARGO BANK NA

Automatic stamping method and system based on invisible watermarking

The application provides a kind of automatic stamping method and system based on invisible watermark, comprising: the electronic text of automatic office unit uploading stamping file and the position information needing stamping;Automatic office unit replaces the electronic text without invisible electronic watermark in process with electronic text;Server informs automatic stamping robot unit to automatically print the electronic text in automatic office unit electronic stamping process;Automatic stamping robot extracts token information in hidden electronic watermark from paper text page by page;If each page token authentication passes, pass through OCR and image matching;Automatic stamping robot stamps information in server through the token of each page text, confirms whether each page text is stamped and where to stamp;Automatic stamping is carried out in the corresponding position of text document, and after stamping is completed, paper document is output.The application can efficiently realize the comparison of paper document and electronic document, and avoid the inconsistency between stamping file and electronic document in approval process.
Owner:COWA TECHNOLOGY CO LTD +1

A large model-based knowledge base index construction optimization method and device

This application discloses a method and apparatus for optimizing knowledge base index construction based on a large model. The method divides text documents into multiple configuration blocks, performs text block vectorization and embedding to obtain a vector database, and extracts graph elements and merges entities to obtain a merged result. Entity parsing is performed on the merged result to obtain a structured entity and relational data graph. Based on the closeness of relationships between the structured entity and relational data graphs, similar characteristic community structures are obtained by grouping them. A community summary report is generated using a large model and integrated into the graph database. Finally, based on the naive RAG method and knowledge graph technology, the vector database and graph database are combined to obtain a multidimensional knowledge base index. The method uses knowledge graphs to construct a knowledge base index from complex connections and implicit relationships in knowledge. Simultaneously, large model technology intelligently assists the knowledge graph generation process, constructing knowledge communities to fill in the explicit and implicit relationships between knowledge points, thereby improving the efficiency and accuracy of knowledge retrieval and generation.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Automated topic modelling and visualization based upon service phase

Disclosed in some examples are methods, systems, devices, and machine-readable mediums which create various data visualizations from a corpus of raw data collected from one or more sources. For example, natural language text documents of a corpus may be classified based upon the topic of the data. The documents may be labeled with the phase during which the documents were collected or observed. A visualization may then be generated which shows a correlation between the phase and the topics observed. For example, a number of times a particular topic appeared in a particular phase. The visualization may be two-dimensional, three-dimensional, or the like.
Owner:WELLS FARGO BANK NA

Method and device for intelligently generating report

The invention provides a method and device for intelligently generating a report, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an industry database, constructing an outline vector library and an outline word segmentation library based on the classified industry database, segmenting a text document in the industry database, generating text blocks, and storing the text blocks in the industry database; constructing a text block vector library and a text block word segmentation library based on the text blocks; generating a related report outline based on the report question, the report type and an outline generation model; and analyzing the related report outline through a regular matching technology, generating a multi-level title based on an analysis result, generating related text blocks only for leaf node titles based on the similarity between the leaf node titles and a text block vector library and a text block word segmentation library, and generating related report contents by combining the related text blocks through a report generation model. And the data volume of the industry database is used as support, so that the field adaptability of the generated report content can be effectively improved.
Owner:UNIV OF SCI & TECH BEIJING +1

Self-adaptive wavelet denoising method and system based on mass spectrum signal processing

PendingCN121636903AData setWavelet thresholding
The invention provides a self-adaptive wavelet noise reduction method and system based on mass spectrum signal processing, and the method comprises the following steps: S1, collecting a data set outputted by a mass spectrometer, and inputting the data set into a noise reduction process in the form of a text document; s2, reading related data in the text document, analyzing spectral peak characteristic parameters, and initializing related parameters; s3, performing multilayer wavelet decomposition on the document data, and outputting an approximate component and a detail component of each layer; s4, calculating the noise intensity of different layers according to the output approximate component and the detail component, and adaptively calculating a wavelet threshold value based on the signal length and the noise intensity; s5, threshold processing is applied to the detail components, and wavelet signals are reconstructed; and S6, calculating signal-to-noise ratios under different decomposition layer numbers, and selecting the optimal decomposition layer number to output the mass spectrum data after noise reduction. According to the method, the related spectrogram information output by the mass spectrometer is denoised, and the signal-to-noise ratio of the mass spectrum data is improved while the spectrum peak information is reserved to a great extent.
Owner:NATIONAL INSTITUTE OF METROLOGY CHINA

Unsupervised record extraction system and method under limited data and resource scene

The invention relates to an unsupervised record extraction system and method in a limited data and resource scene, and the system comprises a data conversion module which converts a paper court trial record file into a readable text document through a visual large model ocr mode; the data extraction module is used for extracting structured fact data according to the court trial record document by utilizing a teacher model based on the set extraction subject and extraction cue word; the data enhancement module is used for carrying out corresponding cue word and theme disassembly on the structured fact data on the basis of the structured fact data, and carrying out data reconstruction on the basis of the disassembled fact data and the corresponding theme to form new enhanced data; and the model distillation module takes the structured fact data extracted by the data extraction module and the enhanced data generated by the data enhancement module as distillation data, and performs instruction fine tuning on the student model to improve the extraction and instruction following ability of the student model. And efficient and deployable court trial record structured extraction is realized.
Owner:SUZHOU UNIV OF SCI & TECH +1

Structured output of duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

Techniques described herein provide for generation of structured output for documents identified using automated near-duplicate detection. In one example, a system can receive a set of documents including at least one pair of similar documents determined to be similar to one another based on similarity scores generated using a predefined similarity scoring technique. The system can generate document groups by merging together pairs of documents that share at least one document. The system can, for each of the document groups, identify a representative document for the document group. The system can generate an output for display including a section for each document group, in which each section includes the representative document for the document group and, for each document in the document group, the similarity score relative to the representative document for the document group.
Owner:SAS INSTITUTE INC

Detection enhancement method, device, equipment, medium and program product

The invention provides a retrieval enhancement method which can be applied to the technical field of big data and artificial intelligence. The method is applied to a distributed cluster, the distributed cluster comprises a document analysis cluster, a vector database cluster and a retrieval type generation cluster, and the method comprises the following steps: querying a preset knowledge base in the vector database cluster based on a received retrieval request from a user to obtain a query result, the preset knowledge base is established by cutting an image-text document on a document analysis cluster based on a preset text cutting strategy; and generating a retrieval formula based on the retrieval request and the query result in a retrieval formula generation cluster so as to use the retrieval formula as the input of a large language model service when the large language model service is called. The invention further provides a retrieval enhancement device and equipment, a storage medium and a program product.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Industrial control programming language implementation method based on text coding

The invention discloses an industrial control programming language implementation method based on text coding, relates to an information processing technology, and is used for solving the problems of low manual code writing efficiency, high error rate and high threshold in traditional control logic development. Comprising the steps of pre-constructing an explanation text and a language prompt text; loading the description text and the language prompt into an AI dialogue model; inputting a demand and a purpose to the AI dialogue model; the AI dialogue model analyzes the demand and the purpose, carries out processing based on the description text and the language prompt text, generates a corresponding final code, and constructs the final code into a text document; and performing corresponding equipment control based on the text document. According to the method, automatic analysis and code generation are performed through the AI dialogue model, the manual coding time is shortened, and the response is faster.
Owner:JINGJIANG XINGHUO MICROCOMPUTER APPLICATION RESEARCH INSTITUTE