Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

140 results about "Plain text" patented technology

In computing, plain text is a loose term for data (e.g. file contents) that represent only characters of readable material but not its graphical representation nor other objects (floating-point numbers, images, etc.). It may also include a limited number of characters that control simple arrangement of text, such as spaces, line breaks, or tabulation characters (although tab characters can "mean" many different things, so are hardly "plain"). Plain text is different from formatted text, where style information is included; from structured text, where structural parts of the document such as paragraphs, sections, and the like are identified); and from binary files in which some portions must be interpreted as binary objects (encoded integers, real numbers, images, etc.).

Multi-modal retrieval system based on similarity between image and text feature vectors

A multi-modal retrieval system based on the similarity between image and text feature vectors, which system belongs to the technical field of information retrieval. In the present invention, advanced deep learning technology is used to extract high-dimensional feature vectors of image data and text data, and the similarity between these feature vectors is computed by means of a specific similarity algorithm. The system can process various types of query inputs, which take the form of plain text, plain images and a combination thereof, thereby providing a flexible search mode.
Owner:INSPUR CLOUD INFORMATION TECH CO LTD

An electronic binder system (ebinder) for processing source data to EDC systems

The present invention provides a method and system for automatically and seamlessly processing clinical trial source data into electronic data capture (EDC) systems. In one embodiment, a file structure is defined for an electronic binder system (eBinder); source data is uploaded to the eBinder; the source data is encrypted, Patient Identifiable Information in the source data is masked; the source data is converted into machine readable plain text in the JavaScript Object Notation (JSON) format using Natural Language Processing (NPL) technologies; the JSON data is converted into tabulated machine readable data in the HyperText Markup Language (HTML) format using NPL technologies; the HTML data is converted into machine understandable data using NPL technologies; the machine understandable data is populated into EDC datasets using NPL technologies; the source data and converted data are displayed side-by-side for source data verification; and a platform is provided for regulatory data verification or auditing.
Owner:XIE TAI +1

Title level identification large model training method, title identification method, system and program product

According to the title level identification large model training method, the title identification method and system and the program product provided by the invention, the first title information covering all texts of the original file is constructed through the multi-modal semantic model, and the limitation of plain text identification is made up in combination with the element matching page picture; second title information with semantic and visual features is generated through multi-modal fusion, so that the title judgment accuracy is improved; in the training process, effective title objects are screened in combination with original title objects to optimize pre-training data, and a large model which is high in precision and adapts to complex scenes is cultivated; in the identification process, the title information to be identified is constructed based on the valid title object. And performing dynamic branch processing according to a calling condition, if not, directly outputting an answer, if yes, generating accurate final title information by means of a trained large model, and finally constructing and outputting the answer, thereby realizing training and identification full-link coordination, considering complex scene adaptability and efficient and accurate identification, and comprehensively improving the structuralization and practicability of document title identification.
Owner:SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD +1

Industrial automation design environment prompt engineering for generative AI

An integrated development environment (IDE) for designing, programming, and configuring aspects of an industrial automation system uses a generative artificial intelligence (AI) model and associated neural networks to generate portions of an industrial automation project in accordance with functional requirements provided to the industrial IDE system in intuitive formats, such as spoken or written plain language text. The system uses generative AI to translate plain language requests or functional specifications into industrial control code, human-machine interface (HMI) applications, device configuration settings, or other aspects of an industrial control project.
Owner:ROCKWELL AUTOMATION TECH INC

Description generation method based on continuous zero sample

The invention provides a description generation method based on continuous zero samples. The description generation method comprises the following steps: sequencing plain text training corpora according to the difficulty of continuous learning; setting that each training step t only can access the t-th corpus, and for the t-th learning step, generating a synthetic image for an input text of the t-th corpus by using a diffusion model; extracting image embedding of the synthesized image by using a CLIP image encoder; retrieving a group of text description embedding with similar semantics in a text corpus by using image embedding; image embedding and text description embedding are input into a fusion module network, and the output of a fusion module is defined as a soft prompt of a large language model; for an input text, key entities are extracted from the text through word frequency statistics so as to construct a soft prompt and a hard prompt of the large language model; hard prompts and soft prompts are input into a large language model network, and key parts in input are concerned through an attention mechanism.
Owner:JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE

Project establishment repeatability detection method and device based on ocean engineering scientific research project

The invention relates to the technical field of natural language processing, in particular to a project establishment repeatability detection method and device based on ocean engineering scientific research projects. According to the method, word segmentation processing is carried out on historical scientific research projects by constructing an ocean engineering terminology dictionary, and meanwhile, independent processing is carried out on numerical parameters, so that the problems of low terminology recognition accuracy, misrecognition and the like are solved; when the text similarity is detected, semantic similarity compensation is introduced, invisible correlation is effectively recognized, and the matching limitation is broken through; in addition, the similarity of item attributes is considered, repeatability detection is comprehensively carried out based on text similarity and attribute similarity, and the defects of plain text detection are effectively overcome.
Owner:SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD

Secure network identification for active scanning device

An example operation may include one or more of storing a hash of a service set identifier (SSID) of a wireless network via an apparatus, receiving a probe request message transmitted from a network device, wherein the probe request message comprises a hash value, determining that the hash value within the probe request is a valid SSID based on the hash of the SSID of the wireless network stored in the storage device, and controlling the network interface to transmit a probe response with a plain text name of the SSID to the network device in response to the determination.
Owner:KYNDRYL INC

Prompt input tuning with any unstructured data

Systems and methods are directed to tuning prompt inputs for a large language model (LLM). A prompt tuning system accesses multiple types of unstructured data for use in generating a fusion prompt input embedding. The prompt tuning system then generates a respective embedding for each of the multiple types of unstructured data. The respective embedding for each of the multiple types of unstructured data is provided to a fusion graph neural network (GNN). The fusion GNN generates the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding. The fusion prompt input embedding provides context for a pure text input that comprises instructions for an output from the LLM. The fusion prompt input embedding and a token embedding representing the pure text are inputted to the LLM. An output of the LLM is then displayed. (FIG. 6)
Owner:EBAY INC +7

Network module verification process generation method and system and storage medium

The invention relates to a network module verification process generation method and system and a storage medium, and relates to the technical field of embedded communication, and the method comprises the steps: S1, obtaining a to-be-analyzed AT instruction document, and carrying out the preprocessing of the AT instruction document, and obtaining a plurality of plain text segments; s2, inputting the plain text segment into a preset semantic analysis engine for semantic extraction to obtain a semantic triple; s3, constructing an AT instruction knowledge graph according to the semantic triple; and S4, generating a network module verification process corresponding to the AT instruction document according to the AT instruction knowledge graph. Compared with the prior art, the method has the advantages that zero manual reading can be realized, multi-language and multi-format documents are supported, the time sequence, dependency and conditional branches among instructions are understood, an executable graphical verification process is automatically generated, and incremental updating is supported.
Owner:SOUTH SURVEYING & MAPPING INSTR

Robot control method and device based on image-text interleaving instruction, equipment and medium

The invention relates to the technical field of artificial intelligence, provides a robot control method and device based on an image-text interleaving instruction, equipment and a medium, is applied to financial and medical health care service scenes, and can construct an initial model which takes an image-text mixed format as an input data format and takes a visual language model as a backbone model. The limitation that a traditional vision-language-action model can only process image observation and plain text instructions is broken through; performing language-action pre-training on the initial model based on a potential action distribution function and a flow matching mechanism, so that the model can learn potential representation of action intention from a language; instruction tuning training is carried out based on a hybrid expert mechanism to ensure that the model is flexibly switched between language reasoning and action prediction, so that the model can accurately reasone from a complex and natural image-text instruction and carry out action planning; and generating a predicted action sequence by using the robot action generation model, so as to accurately control the robot based on the image-text interlacing instruction.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-subject personalized image generation method, system and device and storage medium

The invention discloses a multi-subject personalized image generation method, system and device and a storage medium, which are corresponding schemes, and in the scheme, based on diffusion blueprint generation and initial noise record of text prior, semantic-level spatial layout prior in a model can be extracted from plain text prompt; a region constraint image cross attention mechanism is introduced, based on the space offset of a blueprint mask, visual features of each subject reference image are forced to take effect only in a mask region specified by a diffusion blueprint, and cross-border interaction of appearance features of each subject is prevented; furthermore, the noise adding-denoising reversibility of the diffusion model is combined with the recorded initial noise, so that the constructed mixed initial noise is obtained; in addition, background image information is injected in a delayed mode in the later stage, and finally natural fusion of the foreground and the background is achieved. The whole scheme is completely constructed based on the pre-training model, a specific subject or scene does not need to be finely adjusted, and good universality and expansibility are achieved.
Owner:UNIV OF SCI & TECH OF CHINA

Canonicalization of Unicode Prompt Injections

A prompt for a generative artificial intelligence (GenAI) model is received which includes unicode. Unicode fonts in the prompt are identified and then translated into a plaintext representation. Further, unicode characters in the prompt are identified which each have an associated unicode tag. It is determined, based on the associated unicode tags, whether at least a portion of the unicode characters are valid. When at least a portion of the unicode characters are determined to be valid, the unicode characters in the prompt are converted into a plaintext representation. The prompt with the translated fonts and the converted unicode fonts are passed into the GenAI model. When at least a portion of the unicode characters are not determined to be valid, the unicode characters are removed from the prompt. This prompt with the translated unicode fonts, after the unicode characters are removed, is input into the GenAI model.
Owner:HIDDENLAYER INC

Multi-modal large model reliable inference method and system for medical image assisted diagnosis

PendingCN122290955AEngineeringMedical diagnosis
This invention relates to a multimodal large-scale model reliable reasoning method and system for medical image-assisted diagnosis, comprising: designing a multimodal thought chain data generation mechanism; using plain text thought chain data and generated multimodal thought chain data to supervise and fine-tune the large model to achieve a cold start effect; further stimulating the reasoning ability of the large model through reinforcement learning, wherein the reward function of reinforcement learning includes multi-dimensional considerations; the calculated reward function values ​​of each dimension are connected to a dynamic reward scheduling mechanism, which adaptively adjusts the weight coefficients between the reward functions to achieve a balanced optimization of the objective function with multiple reward objectives, and outputs a reliable reasoning medical diagnosis result. This invention eliminates redundant information in the answer while accurately retaining the key information needed to solve the problem, thereby ensuring the accuracy and reliability of the answer. Compared with the prior art, this invention has the advantages of high accuracy, strong robustness, excellent reliability, and high reasoning efficiency.
Owner:TONGJI UNIV

Generative AI for industrial automation control design environment

An integrated development environment (IDE) for designing, programming, and configuring aspects of an industrial automation system uses a generative artificial intelligence (AI) model and associated neural networks to generate portions of an industrial automation project in accordance with functional requirements provided to the industrial IDE system in intuitive formats, such as spoken or written plain language text. The system uses generative AI to translate plain language requests or functional specifications into industrial control code, human-machine interface (HMI) applications, device configuration settings, or other aspects of an industrial control project.
Owner:ROCKWELL AUTOMATION TECH INC

Entity identification method based on semantic analysis and label visual feature extension

The invention belongs to the technical field of information extraction, and provides an entity recognition method based on semantic analysis and label visual feature extension. The problem that an entity boundary is difficult to determine is solved by calculating the similarity between candidate entities and text global semantics and extracting interaction features of the candidate entities and the text global semantics to implicitly construct embedded representation containing sentence structure information, visual features are systematically introduced into an entity recognition task, and a text-visual cross-modal entity recognition normal form is constructed. Context semantic information represented by embedding is enriched. The entity recognition method combines character-level, sentence-level and visual feature triple semantic representation, not only considers local entity semantics, but also integrates global context information, and introduces visual features to extend limitation of embedded representation based on plain texts.
Owner:DALIAN UNIV OF TECH +1

Method and system for standard-compliant, confidential receipt of data for encrypted data processing

A method for standard-compliant, confidential receipt of data is set forth. The method is executable via a system for data processing, in particular a system for standard-compliant, confidential receipt of data. The data is divided into secret shares already upon receipt, so that none of the servers involved can see the received data in plain text and the method thus enables encrypted data processing and secure data receipt.
Owner:BERGISCHE UNIV WUPPERTAL

Split-aggregate form storage method and system

PendingUS20260252728A1Data fileEngineering
A document management system uploads documents and stores them. The documents include customer-filled forms having a form template and fields within each form having variable data. The data within the fields is extracted and stored in a database. The data can be plain text, confidential text, a signature, or an image. The data for the confidential text and the signature is encrypted prior to storage. The extracted data is placed in a customer form data file that is stored with other data files for the plurality of filled forms along with the form template. The form template and data file is retrieved to regenerate the form with the data within the fields.
Owner:KYOCERA DOCUMENT SOLUTIONS INC

Network data packet analysis system

PendingCN121334014ATransmissionData packLogical combination
The invention relates to the technical field of networks, in particular to a network data packet analysis system, and provides a system for network data packet analysis, which comprises a protocol decoding engine, a graph optimization engine and a graph traversal engine, in a protocol decoding engine, a plain text concise language specification is adopted to describe basic logic and logic combination of formal languages, description of logic combination and complex nested structures is supported, high expression ability and good readability are achieved, the design is concise enough, the manufacturing difficulty of decoding scripts and the total number of codes can be greatly reduced, and the decoding efficiency is improved. In the graph optimization engine and the graph traversal engine, grammar logic of an original graph structure processed by the protocol decoding engine is converted into graph logic by adopting a graph algorithm to realize performance optimization during operation, and flexible analysis of different types of protocols is supported for a multi-graph engine architecture; and a graph structure designed in a classified manner comprises a character string graph, a regular graph and the like, and a standardized graph traversal mechanism is realized by unifying a multi-graph traversal engine.
Owner:科来网络技术股份有限公司

Explanation language generation method and device suitable for protocol decoding and medium

The invention relates to the technical field of networks, in particular to an interpretation language generation method and device suitable for protocol decoding and a medium, and provides an interpretation language suitable for protocol decoding. The method supports the description of logic combination and complex nested structures, has strong expression ability and good readability, is simple enough in design, greatly reduces the manufacturing difficulty and reading difficulty of decoding scripts and the total amount of codes, and greatly shortens the development period. Meanwhile, in protocol decoding, each network protocol is regarded as a formal language, abstraction is carried out to describe most protocols, an interpretation language which is concise in design is designed to standardize and describe basic logic and combination of the formal language, a plain text protocol decoding script is used for replacing a traditional hard coding decoder, and a complete three-level grammar system of words, sentences and segments is constructed.
Owner:科来网络技术股份有限公司

Knowledge graph construction method and system

The application discloses a knowledge graph construction method and system, and relates to the technical field of knowledge graph construction.The knowledge graph construction method comprises the following steps: performing a streaming decoupling operation on a received unstructured document to obtain a pure text data stream and a visual data stream; performing semantic direct reading and extraction on the pure text data stream and the visual data stream to obtain a triple pool; performing induction and standardization processing on the triple pool to obtain standardized triple data; mapping the standardized triple data into nodes and directed edges; and constructing a knowledge graph according to the nodes and the directed edges.The knowledge graph construction method can improve the completeness and numerical accuracy of data extraction from a complex visual chart.
Owner:启元实验室

Online community creator feedback prediction method and system based on dynamic reputation graph and text analysis

PendingCN122020314ASolve the problem of failing to reflect changes in user statusreduce mistakesSemantic analysisPagerank algorithmEngineering
The invention relates to an online community creator feedback prediction method and system based on a dynamic reputation graph and text analysis. Relates to the technical field of feedback prediction. The method comprises the following steps: S1, acquiring and preprocessing historical interaction data containing comment texts, user IDs, timestamps and like and treading records; s2, constructing a like and treading double-view interaction map; s3, segmenting the data according to the day, and calculating a daily positive and negative reputation value through a PageRank algorithm; s4, introducing a time decay function, and carrying out weighted summation on the daily granularity reputation value in the close time window to obtain a dynamic reputation; s5, text semantic features are extracted through RoBERTa, the text semantic features and the dynamic reputation features are spliced and fused, and the full-connection neural network is input to output the like or treading probability. The method gives consideration to positive and negative reputation and time dynamics, makes up for the defects of plain text prediction, remarkably improves the feedback prediction precision, especially optimizes the click behavior prediction effect, and can provide support for community atmosphere guidance and network violent early warning.
Owner:ZANAO (SUZHOU) TECHNOLOGY CO LTD

Auditable structural consistency receipt system and single-step repair method for edge AI reasoning pipeline

The invention discloses an auditable structural consistency receipt system and a single-step repair method for an edge AI reasoning pipeline, and belongs to the technical field of control and data auditing during operation of artificial intelligence equipment, and the auditable structural consistency receipt system comprises a sequence exchangeable rectangle and a loop conservative triangle (structure), and according to a predetermined level (geometry-framework-data field), a single correction adapter is selected to restore inconsistency. Each repair strictly ensures that the inconsistent measure mu is reduced by 1 (delta mu = 1), and single-step convergence is realized. The system serializes and records each step operation as an additional plain text receipt, and performs integrity protection on a receipt chain through a deterministic natural number fence value without depending on an encryption algorithm. And finally, result judgment is carried out through six Boolean condition'judgment gate 'outside the assembly line, wherein the six Boolean conditions comprise check in the aspects of parity check, phase alignment, sequence exchange (torsion), label consistency, scale benchmark, gluing closing and the like.
Owner:GUANGZHOU KINGPIN IND CO LTD

Webpage information extraction and classification method and device

The application provides a webpage information extraction and classification method and device, and belongs to the technical field of artificial intelligence. The webpage information extraction and classification method comprises the following steps: converting the source code of a target webpage into a dom tree; processing each node of the dom tree to obtain four feature matrices, i.e., a text feature matrix, an Xpath feature matrix, a layout feature matrix and a visual feature matrix; inputting the four feature matrices into an encoding network respectively to obtain four representation vectors; performing feature fusion on the four representation vectors to obtain fused features; inputting the fused features into a classification network to obtain and store the classification result of an information unit of the target webpage, wherein the classification result comprises at least one of the following: a table, a form to be filled, a text unit that needs to be linked to a next webpage for display, a navigation bar or a display bar, pure text, an advertisement and useless information. The technical scheme of the application can improve the universality of the webpage information extraction and classification scheme.
Owner:CHINA MOBILE COMM LTD RES INST +1

Chinese webpage interest point retrieval method and device, and electronic equipment

The present disclosure discloses a Chinese webpage interest point retrieval method and device and electronic equipment, and relates to the technical field of neural networks and the technical field of cloud computing. The technical problem that in the model training process, only the pure text content of the webpage is pre-trained, the position structure information in the webpage data is ignored, and then the model learning representation is too single, which affects the accuracy of information extraction of the downstream task, is solved. The specific implementation scheme is: in response to the instruction of the user selecting the interest point, obtaining the target webpage data of the target webpage containing the interest point; inputting the target webpage data into the Chinese webpage pre-training model to obtain the target attribute information corresponding to the interest point; and displaying the target attribute information.
Owner:BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD

Large visual language model hallucination mitigation method and apparatus

The present application belongs to the technical field of artificial intelligence and multi-modal large model, and particularly relates to a large visual language model hallucination reduction method and device. The method comprises the following steps: obtaining complete visual tokens of an original image and text tokens of a text prompt, and connecting the same to jointly input a large language model decoder; calculating an attention score matrix of the text tokens and all the visual tokens based on a cross-modal dynamic sampling strategy to sample key visual tokens; obtaining classification tokens of the original image, and screening significant visual tokens based on the attention scores of the visual tokens in the classification tokens and the complete visual tokens; performing adaptive attention enhancement on the significant visual tokens and the key visual tokens, and subtracting the logits distribution influence of pure text input from the logits distribution of visual information enhancement through a comparison decoding strategy to obtain a final target text output. The present application aims to reduce the hallucination problem in a large visual language model.
Owner:SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1

Systems and methods for providing explainability of natural language processing

Systems or techniques that facilitate systems and methods for providing explainability of natural language processing are provided. In various embodiments, a system can access a plain text clinical sentence. In various aspects, the system can generate, via execution of a first machine learning model, an assertion status classification label for a word of interest in the plain text clinical sentence. In various instances, the system can extract, from a hidden attention layer of the first machine learning model, word-wise attention scores corresponding to the plain text clinical sentence and render, on an electronic display, both the assertion status classification label and a graphical representation of the word-wise attention scores. In various cases, the system can determine, via execution of a second machine learning model, a reliability score for the assertion status classification label, based on the word-wise attention scores, and can render the reliability score on the electronic display.
Owner:GE PRECISION HEALTHCARE LLC

Ai assisted ADA content compliance workflow

PCT designated stageWO2026136345A1Natural language data processingWebsite content managementWeb Content Accessibility GuidelinesEngineering
A system for converting digital documents into American Disabilities Act (ADA) Web Content Accessibility Guidelines (WCAG) compliant content includes a content upload system capable of receiving a digital document. An optical character recognition program converts the digital document to a plain text document. A language module structurally organizes the plain text document while maintaining the content of the digital document. A hypertext markup language (HTML) module builds an HTML based document having a structure that is WCAG compliant.
Owner:PEACHJAR

Text sentence breaking method, system and related device based on combination of acoustics and semantics

The application provides a text punctuation method and system based on the combination of acoustics and semantics, and related equipment. The method comprises: obtaining audio data containing voice instructions; processing the audio data based on a preset voice recognition engine and a semantic model, identifying at least one candidate punctuation point in the text corresponding to the audio data, and obtaining a semantic punctuation probability of each candidate punctuation point; for each candidate punctuation point, extracting an acoustic feature set corresponding to the candidate punctuation point in the audio data; calculating a fusion punctuation score for the candidate punctuation point according to the semantic punctuation probability and the acoustic feature set; based on the fusion punctuation score, determining whether to perform a punctuation operation at the candidate punctuation point, and outputting the final text punctuation result of the audio data. The application fuses and calibrates pure text semantic analysis by introducing acoustic information, reduces the ambiguity of punctuation, and improves the accuracy of complex voice instruction punctuation.
Owner:SHENZHEN TONGXINGZHE TECH