Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Text document" patented technology

A text document may have a TXT, DAT, LOG or HTML file extension. You can create one in a text editor of your choice. Since text documents do not include special formatting, they appear as plain text within an application.

Intelligent model selection system for style-specific digital content generation

Aspects of the present disclosure provide systems, methods, and computer-readable storage media that support intelligent model selection for style-specific digital content generation. For example, a system that provides a digital content generation service may include a trained style detection model may receive reference digital content items from a user and extract a user style embedding that represents a style preference of the user. In some implementations, the reference digital content items may include text documents or images provided or selected by the user. The system may compare the user style embedding to a plurality of model style embeddings that each correspond to a respective generative artificial intelligence (AI) model to generate a ranked list of generative AI models. The system may access one or more highest ranked generative AI models from the ranked list to generate novel digital content based on a prompt from the user.
Owner:ACCENTURE GLOBAL SOLUTIONS LTD

Mutually generative artificial intelligence system based on multi-dimensional spatiotemporal information vector graphics

Disclosed is a mutually generative artificial intelligence system based on multi-dimensional spatiotemporal information vector graphics, relating to the field of artificial intelligence and engineering applications. Based on a geographic information system (GIS) or a computer-aided design (CAD) platform and a data source, a multi-dimensional vector spatiotemporal large model terminal, a multi-dimensional spatiotemporal information processing agent terminal, and an intelligent information system application terminal are constructed. For multi-dimensional spatiotemporal data such as two-dimensional and three-dimensional vectors and temporal states, a multimodal spatiotemporal large model having an understanding capacity for an engineering professional knowledge system, a data processing flow, multi-dimensional vector graphics, and thematic graphics-text documents is pre-trained to achieve the mutual expression and generation of engineering multi-dimensional vector graphics and thematic graphics-text documents and to form an intelligent engineering graphic data processing application.
Owner:BEIJING LONGRUAN TECHNOLOGIES INC +1

Document image-text integrated intelligent understanding and processing method and system

The invention provides a document image-text integrated intelligent understanding and processing method and system, and relates to the technical field of artificial intelligence. Comprising the following steps: separating image-text elements and establishing context association; calling a graph vectorization engine, and executing intelligent vector conversion processing on the rasterized illustration; performing hierarchical classification, attribute endowing and structured reconstruction on the vector primitives in combination with text semantic analysis to generate vector entity objects with complete attributes; the text elements and the vector entity objects are jointly input into a pre-trained multi-mode large language model, unified knowledge expression containing vectors, topology, attributes and document contexts is generated, and applications such as question answering, editing, abstracting and reporting of image-text integration are supported. According to the method, the problems of graph semantic loss, difficulty in fine analysis and the like caused by vector-to-grid illustration in papers, reports and other types of documents are solved, the overall understanding depth and interpretability of AI for complex image-text documents are remarkably improved, and high-performance semantic understanding capacity is provided for document-based training and intelligent processing.
Owner:BEIJING LONGRUAN TECHNOLOGIES INC +1

Prompt engineering and automated quality assessment for large language models

Various embodiments of the present disclosure provide prompt engineering and text quality assessment techniques for improving generative text outputs. The techniques may include identifying an initial document subset for a generative text request that includes a request to generate a generative text document based on one or more request text fields. The techniques may include generating a contextual classification for the one or more request text fields and identifying a refined document subset based on the contextual classification. The techniques may include generating one or more request field embeddings respectively corresponding to the one or more request text fields and identifying a prompt document subset based on the one or more request field embeddings. The techniques may include generating, using a large language model, one or more generative text fields using a generative model prompt based on the prompt document subset and the one or more request text fields.
Owner:UNITEDHEALTH GROUP INC

Hallucination detection and remediation in text generation interface systems

Enumerated source text passages may be determined based on one or more source text documents. The enumerated source text passages may include source text passage identifiers uniquely identifying the passages. A novel text passage including novel text portions may be determined based on a query and the enumerated source text passages. One or more of the novel text portions may be verified by a large language model to produce text verification information. A novel text generation message including novel text generated by the large language model may be determined based on the text verification information and sent to a client machine.
Owner:CASETEXT INC

Large model information extraction and structure restoration system for long text document

The invention belongs to the technical field of intelligent document processing, and particularly relates to a large model information extraction and structure restoration system for a long text document. The system comprises an extraction point defining and text preprocessing module which is used for carrying out deep preprocessing and analysis on an input unstructured document, extracting text and coordinate information and identifying and protecting a special structure; meanwhile, an intelligent dynamic blocking strategy based on semantic boundaries is adopted, a dynamic overlapping mechanism is combined, and an input document is divided into text blocks keeping semantic integrity; the high-concurrency processing and scheduling control module is used for processing the text blocks in batches, calling a large language model to carry out distributed reasoning and outputting a dispersion result; and the extraction reasoning and result fusion module is used for performing semantic deduplication, entity alignment and confidence fusion on the dispersion result returned by the large language model to generate globally consistent structured output. The method has the advantages of being high in precision, high in speed, controllable in cost and high in robustness.
Owner:浙江实在智能科技有限公司

Electric power project material inventory intelligent generation method and system based on multi-modal document analysis and knowledge graph semantic mapping

The invention discloses an electric power project material inventory intelligent generation method and system based on multi-modal document analysis and knowledge graph semantic mapping, and belongs to the field of artificial intelligence and electric power engineering. Comprising the steps that in the first stage, unified analysis and semantic coding are conducted on text documents, engineering drawings and table data, and a semi-structured original information set is generated; in the second stage, named entity recognition and type labeling are carried out on the original information set based on a two-way long-short term memory network and a conditional random field hybrid model, a structured entity set is generated, an electric power material field knowledge graph is constructed on the basis, and a standard key point rule base is established; semantic mapping and information completion are carried out through a multi-level matching mechanism and knowledge reasoning, and a standardized material inventory is generated after quality control. According to the invention, intelligent and automatic generation of the electric power project material inventory is realized, and the accuracy and efficiency of material inventory generation are improved.
Owner:STATE GRID LIAONING ELECTRIC POWER CO LTD

Multi-language intelligent analysis system for medical documents

The invention provides a medical document multi-language intelligent analysis system, relates to the field of language translation, and improves the accuracy and efficiency of professional term translation. The method comprises the following steps of: firstly, performing language recognition on an original text document by a recognition module through text unitization and context vector generation, and matching a corresponding corpus; then, a translation module carries out lexical element alignment on the source lexical elements through term bank injection and an AI model, the translation process is automatically optimized, and accurate translation of the terminologies is ensured; and finally, the reconstruction module accurately replaces corresponding contents in the original text document with translation output through the mapping file, so as to ensure that the document format and typesetting are consistent. Through the automatic and optimized translation process, the quality and efficiency of professional term translation are remarkably improved, manual intervention is reduced, and the translation requirement of a high professional standard is met.
Owner:LUNAN PHARMA GROUP CORPORATION +2

College student competition intelligent auxiliary system based on AI large model

The invention belongs to the technical field of education informatization, and discloses an AI large model-based college student competition intelligent auxiliary system, which provides a one-stop intelligent tutoring service for college student competition through a material uploading module, a multi-modal analysis module, a case library retrieval module, an AI project diagnosis module, a material generation module and the like. The system can automatically identify text documents, images and code contents and generate structured data of projects; and based on a preset case library, project similarity matching is realized, and reference experience information is extracted. The AI large model is used for performing multi-dimensional scoring on the project, a score deduction basis is provided, operable improvement suggestions are generated, and meanwhile, competition documents such as road performance materials, project abstracts or commercial schedules can be automatically generated. According to the method, the project preparation efficiency can be remarkably improved, the material quality is optimized, accurate and personalized competition support is provided for college students, and the method has good practical value and popularization significance.
Owner:BEIJING MORNINGSTAR VENTURE CAPITAL TECHNOLOGY CO LTD

Method and system for extracting information from documents with varying formats

Certain aspects of the disclosure provide a method for extracting attributes from documents with varying formats, layouts and complexities. The method displays a user interface (UI) that enables a user to obtain an unstructured document from a knowledge base. The method converts the unstructured document into a text document using a text recognition. The method obtains, as output from a large language model (LLM), an extracted page attribute from the text document. The extracted page attribute contains a first type of information recorded in text on a single page of the text document. The extracted document attribute contains a second type of information recorded in text on more than one page of the text document. The method obtains, as output from the LLM, an extracted document attribute from the text document. The extracted page attribute and the extracted document attribute are displayed in the UI.
Owner:SCHLUMBERGER TECH CORP

Automated topic modelling and visualization based upon service phase

PendingUS20260211935A1Data visualizationDocumentation
Disclosed in some examples are methods, systems, devices, and machine-readable mediums which create various data visualizations from a corpus of raw data collected from one or more sources. For example, natural language text documents of a corpus may be classified based upon the topic of the data. The documents may be labeled with the phase during which the documents were collected or observed. A visualization may then be generated which shows a correlation between the phase and the topics observed. For example, a number of times a particular topic appeared in a particular phase. The visualization may be two-dimensional, three-dimensional, or the like.
Owner:WELLS FARGO BANK NA

Automatic stamping method and system based on invisible watermarking

The application provides a kind of automatic stamping method and system based on invisible watermark, comprising: the electronic text of automatic office unit uploading stamping file and the position information needing stamping;Automatic office unit replaces the electronic text without invisible electronic watermark in process with electronic text;Server informs automatic stamping robot unit to automatically print the electronic text in automatic office unit electronic stamping process;Automatic stamping robot extracts token information in hidden electronic watermark from paper text page by page;If each page token authentication passes, pass through OCR and image matching;Automatic stamping robot stamps information in server through the token of each page text, confirms whether each page text is stamped and where to stamp;Automatic stamping is carried out in the corresponding position of text document, and after stamping is completed, paper document is output.The application can efficiently realize the comparison of paper document and electronic document, and avoid the inconsistency between stamping file and electronic document in approval process.
Owner:COWA TECHNOLOGY CO LTD +1

Automated topic modelling and visualization based upon service phase

Disclosed in some examples are methods, systems, devices, and machine-readable mediums which create various data visualizations from a corpus of raw data collected from one or more sources. For example, natural language text documents of a corpus may be classified based upon the topic of the data. The documents may be labeled with the phase during which the documents were collected or observed. A visualization may then be generated which shows a correlation between the phase and the topics observed. For example, a number of times a particular topic appeared in a particular phase. The visualization may be two-dimensional, three-dimensional, or the like.
Owner:WELLS FARGO BANK NA

Method and device for intelligently generating report

The invention provides a method and device for intelligently generating a report, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an industry database, constructing an outline vector library and an outline word segmentation library based on the classified industry database, segmenting a text document in the industry database, generating text blocks, and storing the text blocks in the industry database; constructing a text block vector library and a text block word segmentation library based on the text blocks; generating a related report outline based on the report question, the report type and an outline generation model; and analyzing the related report outline through a regular matching technology, generating a multi-level title based on an analysis result, generating related text blocks only for leaf node titles based on the similarity between the leaf node titles and a text block vector library and a text block word segmentation library, and generating related report contents by combining the related text blocks through a report generation model. And the data volume of the industry database is used as support, so that the field adaptability of the generated report content can be effectively improved.
Owner:UNIV OF SCI & TECH BEIJING +1

Self-adaptive wavelet denoising method and system based on mass spectrum signal processing

PendingCN121636903AData setWavelet thresholding
The invention provides a self-adaptive wavelet noise reduction method and system based on mass spectrum signal processing, and the method comprises the following steps: S1, collecting a data set outputted by a mass spectrometer, and inputting the data set into a noise reduction process in the form of a text document; s2, reading related data in the text document, analyzing spectral peak characteristic parameters, and initializing related parameters; s3, performing multilayer wavelet decomposition on the document data, and outputting an approximate component and a detail component of each layer; s4, calculating the noise intensity of different layers according to the output approximate component and the detail component, and adaptively calculating a wavelet threshold value based on the signal length and the noise intensity; s5, threshold processing is applied to the detail components, and wavelet signals are reconstructed; and S6, calculating signal-to-noise ratios under different decomposition layer numbers, and selecting the optimal decomposition layer number to output the mass spectrum data after noise reduction. According to the method, the related spectrogram information output by the mass spectrometer is denoised, and the signal-to-noise ratio of the mass spectrum data is improved while the spectrum peak information is reserved to a great extent.
Owner:NATIONAL INSTITUTE OF METROLOGY CHINA

Structured output of duplicate or near-duplicate text documents identified using automated near-duplicate detection for text documents

Techniques described herein provide for generation of structured output for documents identified using automated near-duplicate detection. In one example, a system can receive a set of documents including at least one pair of similar documents determined to be similar to one another based on similarity scores generated using a predefined similarity scoring technique. The system can generate document groups by merging together pairs of documents that share at least one document. The system can, for each of the document groups, identify a representative document for the document group. The system can generate an output for display including a section for each document group, in which each section includes the representative document for the document group and, for each document in the document group, the similarity score relative to the representative document for the document group.
Owner:SAS INSTITUTE INC

Combined small model-based electric power rich text document identification method and system

The invention discloses an electric power rich text document identification method and system based on a combined small model. The method comprises the steps of obtaining metadata information of a power rich text document, converting the metadata information into a picture format, and performing preprocessing to obtain an image data set; constructing layout analysis model test indexes, and selecting the model with the highest index from the open source layout recognition models as a layout analysis model; constructing a formula recognition model test index, and selecting a formula recognition model with the highest index from the open source layout recognition models as the formula recognition model; constructing table recognition model test indexes, and selecting the table recognition model with the highest index from the open source table recognition models as a table recognition model; and recognizing data in the image data set by using a layout analysis model, a formula recognition model, a table recognition model and an open source character recognition model, and converting an obtained result into a Markdown format to finish the recognition of the power rich text document. According to the method, the multi-stage model evaluation and the document analysis algorithm are combined together to form a complete processing flow, and recognition of the electric power rich text document is achieved.
Owner:STATE GRID HUNAN ELECTRIC POWER CO +2

Detection of Sensitive Information in a Text Document

An apparatus (300) for detecting sensitive information in a first text document representative of a first topic is provided. The apparatus (300) is configured to generate a first updated text document by tagging a segment of text in the first text document using a list of one or more types of sensitive information for a second topic: train a language model on text representative of the first topic and on a list of one or more types of sensitive information for a third topic, wherein the language model is a transformer-based machine learning model; and generate a second updated text document by classifying as sensitive a segment of text in the first updated text document using the trained language model representative of relationships between the tagged segment, one or more types of sensitive information for the third topic, and the text representative of the first topic.
Owner:TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)

Tree-based hierarchical rich text slice construction method and system

The invention discloses a tree-based hierarchical rich text slice construction method and system, belongs to the technical field of data processing, and aims to solve the technical problem of how to improve slice quality and retrieval recall accuracy for a rich text document with a good target hierarchy. Comprising the following steps: mapping a rich text document into a model-friendly mark format according to an original layout structure; analyzing the text subjected to format conversion into an element tree through an extractor, and representing a hierarchical structure of the document through the element tree; fragmenting contents in the element tree through a fragmenter to generate document slices suitable for model processing; text optimization is carried out on the document slices through index construction, a word embedding model is called to convert a text into vector representation, and indexes of the document slices are constructed through an index technology.
Owner:INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD

Report text generation method and device, equipment, storage medium and program product

The invention relates to a report text generation method and device, equipment, a storage medium and a program product. The method comprises the following steps: constructing a reusable filling task library, newly creating a report, and selecting and associating tasks needing to be filled from the filling task library; defining a report text framework by adopting a tree data structure, transmitting the attribute of each node in the form of a JSON character string, and analyzing and storing the attribute in a database; adjusting a node sequence through a visual dragging engine, driving a task node to issue a bound filling task to a handling user, setting different approval processes and multi-stage approval routes, and marking the task node as a completed state after approval is passed; if the approval is not passed, returning to the handled user for modification; and when all the task nodes are completed, triggering a text generation instruction, and according to a newest node sequence, integrating contents of the title node and the completed task nodes to generate a report text document meeting a format requirement. By adopting the method, the differentiated structure requirements of different types of periodic reports can be met.
Owner:SHANGHAI PUDONG DEVELOPMENT BANK

System-guided collaborative editing of standardized documents

According to techniques described herein, a system / platform may implement a method for collaboratively editing a standardized text document. The platform may receive a user selection of a standardized text document, receive a first party set of values, receive a second party set of values, execute an algorithm for generating a final set of values for the standardized text document, and transmit the final set of values to the first party and the second party.
Owner:BONTERMS INC

Ai assisted ADA content compliance workflow

PCT designated stageWO2026136345A1Natural language data processingWebsite content managementWeb Content Accessibility GuidelinesEngineering
A system for converting digital documents into American Disabilities Act (ADA) Web Content Accessibility Guidelines (WCAG) compliant content includes a content upload system capable of receiving a digital document. An optical character recognition program converts the digital document to a plain text document. A language module structurally organizes the plain text document while maintaining the content of the digital document. A hypertext markup language (HTML) module builds an HTML based document having a structure that is WCAG compliant.
Owner:PEACHJAR

A text chunking method and device based on HTML node path resolution

Embodiments of the present specification relate to the technical field of text processing, and provide a text blocking method and device based on HTML node path analysis, comprising: performing initial blocking on the HTML document to obtain an initial HTML document blocking result, and recording an XPath path expression corresponding to each piece of text in the target HTML document and an initial HTML document block to which the text belongs; performing preprocessing on each initial HTML document block to obtain a plurality of initial pure text document blocks; performing merging or segmentation operations on the initial pure text document blocks to obtain a plurality of final pure text document blocks; and blocking the target HTML document according to the XPath path expression corresponding to each piece of text in the final pure text document blocks and the initial HTML document block to which the text belongs, to obtain a final HTML document blocking result. Through the embodiments of the present specification, the accuracy of HTML text blocking can be improved.
Owner:CHINA EVERBRIGHT BANK

Clustering and dynamic re-clustering of similar textual documents

A computer-implemented method includes obtaining a plurality of textual records divided into clusters and a residual set of the textual records, where a machine learning (ML) clustering model has divided the plurality of textual records into the clusters based on a similarity metric. The method also includes receiving, from a client device, a particular textual record representing a query and determining, by way of the ML clustering model and based on the similarity metric, that the particular textual record does not fit into any of the clusters. The method additionally includes, in response to determining that the particular textual record does not fit into any of the clusters, adding the particular textual record to the residual set of the textual records. The method can additionally include identifying, by way of the ML clustering model, that the residual set of the textual records contains a further cluster.
Owner:SERVICENOW INC

Training a machine learned model based on a training vector associated with a text document embedding and an attribute set

PendingUS20260187520A1Ground truthFeature set
Various embodiments of the present disclosure provide for training and / or deploying a machine learned model based on a training vector associated with a text document embedding and an attribute set. The techniques may include receiving a user attribute feature set associated with a user identifier for a user-text pair, generating an input text document embedding for the user identifier, generating an input vector by concatenating the user attribute feature set with the input text document embedding, generating, using a machine learned model that is trained based on a ground truth label and a training vector associated with a text document embedding generated by a text encoder model, a classification score for the input text document embedding based on the input vector, and generating a transcript corresponding to the input text document embedding upon determining that the classification score satisfies a predetermined threshold.
Owner:OPTUM INC

Target processing node evaluation method and answer generation method of language generation model

The embodiment of the invention relates to the technical field of artificial intelligence, in particular to a target processing node evaluation method of a language generation model and an answer generation method.The target processing node evaluation method of the language generation model comprises the steps that a text document in a document data set is analyzed, and multiple pieces of query data are generated, a first keyword and a second keyword group corresponding to each piece of query data in the plurality of pieces of query data are generated, and each piece of query data comprises a reference question and a reference answer corresponding to the reference question; generating prediction answers corresponding to the reference questions in the query data by using a target processing node of a language generation model; and according to the first keyword and the second keyword group corresponding to each piece of query data, the reference question and the reference answer in each piece of query data, and the predicted answer corresponding to each reference question, evaluating the target processing node to obtain an evaluation result of the target processing node.
Owner:ALIBABA (CHINA) CO LTD

Method for supporting large file questions and answers in large language model

The invention discloses a method for supporting large file questions and answers in a large language model. The method comprises the steps of document reading and character extraction, text segmentation and feature vector conversion, a text feature vector storage system, the large language model and feature vector conversion of content questions. According to the method for supporting large file questions and answers in the large language model, the software is adopted, rapid indexing and searching characteristics of a vector database can be effectively combined, the software is matched with the large language model, and when a user needs to enable the large language model to understand a long text document, the large file questions and answers can be obtained by utilizing the technology and scheme adopted by the patent. According to the method, accurate question answering for the document can be efficiently and rapidly supported, and in the question answering process, the ability of the large language model is utilized, and the ability of the large language model such as multi-language translation can also be achieved. The efficiency of analysis, summarization, translation and question answering of professional document content can be greatly improved.
Owner:SHENZHEN YUANSHIJIE SOFTWARE TECH CO LTD

Real-time visualization method and system for functional function modules

ActiveCN114896918BReal-time visual displayQuick and efficient replacementCAD circuit designSpecial data processing applicationsComputer graphics (images)Engineering
The application provides a real-time visualization method, system, device and computer readable storage medium for a function function module. One of the real-time visualization methods for the function function module can specifically include: converting a manual file in basicProcedure into a text document; classifying the function function module contained in the text document through category; selecting the function function module which needs to perform a parameter modification operation; visualizing and presenting the function function module according to the visualized analysis of the function function module; and updating the visualized presentation in real time during the parameter modification operation of the function function module. Compared with manual input to build a template, the manual combined with graphical function viewing is more intuitive and more convenient.
Owner:PRIMARIUS TECH CO LTD

Intelligent education knowledge graph construction method based on large language model

The invention discloses an intelligent education knowledge graph construction method based on a large language model, and the method comprises the steps: S1, obtaining curriculum text data from an intelligent education video platform, and cleaning and integrating the curriculum text data into a long text document; s2, performing semantic slicing on the long text document by using a natural language processing model to obtain semantic fragments; s3, constructing a course category included in the course knowledge graph, and obtaining extended noun explanation of the course category to obtain an explanation text; s4, finding a corresponding course category for each semantic segment by calculating the similarity between the embedded vectors; s5, decomposing the semantic fragment into a plurality of knowledge points by utilizing a large language model and CoT cue words, and marking; and S6, storing the knowledge points in the neo4j graph database in a defined format to form a multi-course knowledge graph. According to the method, illusion problems and errors caused by completely depending on a large model can be avoided as much as possible, and construction of the course atlas is realized by efficiently utilizing partial open source technical means.
Owner:SOUTH CHINA UNIV OF TECH