Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

224 results about "Plain text" patented technology

In computing, plain text is a loose term for data (e.g. file contents) that represent only characters of readable material but not its graphical representation nor other objects (floating-point numbers, images, etc.). It may also include a limited number of characters that control simple arrangement of text, such as spaces, line breaks, or tabulation characters (although tab characters can "mean" many different things, so are hardly "plain"). Plain text is different from formatted text, where style information is included; from structured text, where structural parts of the document such as paragraphs, sections, and the like are identified); and from binary files in which some portions must be interpreted as binary objects (encoded integers, real numbers, images, etc.).

Retrieval augmented generative question and answer boosting

Systems or techniques are provided for facilitating retrieval augmented generative question and answer boosting. In various embodiments, a system can access a plain text question regarding a scientific instrument. In various aspects, the system can generate, via a large language model that references a document-graph repository, a structured or unstructured answer for the plain text question. In various instances, the document-graph repository can comprise a plurality of document-graphs that respectively correspond to a plurality of technical documents. In various cases, for a first document-graph that corresponds to a first technical document, leaf nodes of the first document-graph can represent respective text blocks written in the first technical document, and non-leaf nodes of the first document-graph can respectively represent a document title, one or more section headings, and one or more scientific instrument identifiers written in the first technical document and beneath which the respective text blocks are nested.
Owner:PPD DEVELOPMENT LP +2

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Multi-modal retrieval system based on similarity between image and text feature vectors

A multi-modal retrieval system based on the similarity between image and text feature vectors, which system belongs to the technical field of information retrieval. In the present invention, advanced deep learning technology is used to extract high-dimensional feature vectors of image data and text data, and the similarity between these feature vectors is computed by means of a specific similarity algorithm. The system can process various types of query inputs, which take the form of plain text, plain images and a combination thereof, thereby providing a flexible search mode.
Owner:INSPUR CLOUD INFORMATION TECH CO LTD

Data circulation trusted system and device based on block chain and privacy calculation

The invention discloses a data circulation trusted system and device based on a block chain and privacy calculation, and relates to the technical field of trusted data. Comprising a data supply end, a data demand end and a credible supervision end, the data supply end comprises a digital identity identification module which is used for generating a dynamic digital identity according to an enterprise unified social credit code and an ISO industry classification code and submitting the dynamic digital identity to a first block chain of the credible supervision end for authentication; the mixed desensitization module is used for carrying out first 3-bit plaintext reservation and hash replacement on the sensitive field by adopting an AES-256 algorithm, binding a UTC timestamp to generate a desensitization record and uploading the desensitization record to a first block chain; and the zero-knowledge proof generation module is used for converting the privacy calculation result into a zero-knowledge proof voucher and pushing verification parameters to a sandbox environment of the data demand end. According to the invention, the security and privacy of data in a circulation link and the credibility of transactions are guaranteed, and the problems of privacy leakage, uncredible transactions, difficult supervision and the like possibly existing in data circulation are solved.
Owner:GUIZHOU DIGITAL INNOVATION HLDG (GRP) CO LTD

Multi-dimensional storage system for cultural heritage archives

The invention discloses a cultural heritage file multi-dimensional storage system, which relates to the field of cultural heritage digital protection and intelligent storage and comprises a data acquisition and processing module, a double-library management module, a multi-dimensional index retrieval engine, a security and convergence module and a backup and disaster recovery module. According to the method, structured texts and unstructured files are uniformly accessed through various interfaces, a special tool chain is adopted for format standardization, audios and videos are transcoded into an H.264 / AAC format through FFmpeg, images are subjected to resolution adjustment and compression through ImageMagick, plain texts are extracted from texts through ApacheTika and are converted into UTF-8 codes, standardized data are uniformly converted into an intermediate format, and the format of the intermediate format is converted into the intermediate format. According to the technical scheme, the method comprises the following steps of: standardizing input data by adopting a three-layer duplicate removal strategy, facilitating other modules to receive data, performing refined redundancy processing on the standardized input data by adopting a three-layer duplicate removal strategy, remarkably reducing storage redundancy, partitioning a cultural heritage data main library according to three layers of ''project type-medium type-age'', and partitioning an inheritor information library according to two levels of ''province-inheritor number prefix''.
Owner:济宁市退役军人服务中心

Encryption Key Distribution System

Customers of a software platform, such as a unified communications as a service platform, are enabled to control their own encryption keys used to encrypt and decrypt data from various communication services in the software platform. A key connector service is used to coordinate communications with a key management server for generating plaintext keys for data encryption and encrypted keys based on the plain text keys, and with a key broker server for storing and retrieving copies of the encrypted keys. A context identifier may be associated with an encrypted key and used to track which data has been encrypted with an underlying plaintext key. Examples of data encrypted may include conference recordings, webinar recordings, phone call recordings, voicemails, emails, and calendar tokens.
Owner:ZOOM COMMUNICATIONS INC

Large visual language model illusion mitigation method and device

The invention belongs to the technical field of artificial intelligence and multi-modal large models, and particularly relates to a large visual language model illusion relieving method and device. The method comprises the following steps of: acquiring a complete visual token of an original image and a text token of a text prompt, connecting the tokens and jointly inputting the tokens into a large language model decoder; calculating an attention score matrix of the text token and all the visual tokens based on a cross-modal dynamic sampling strategy so as to sample key visual tokens; obtaining classification tokens of the original image, and screening significant visual tokens based on the classification tokens and attention scores of the visual tokens in the complete visual tokens; and carrying out adaptive attention enhancement on the significant visual token and the key visual token, and subtracting the logs distribution influence of plain text input from the logs distribution with enhanced visual information by comparing decoding strategies so as to obtain final target text output. The objective of the invention is to alleviate illusion problems in large-scale visual language models.
Owner:HANGZHOU PENGUIN TECH CO LTD

Intelligent extraction method and system for input text containing mathematical formula

The present invention belongs to the technical field of text processing, and relates to an intelligent extraction method and system for an input text containing a mathematical formula. The method comprises: 1) performing format determination, conversion and preprocessing on an input text; 2) performing angle correction on the preprocessed image-formatted text; 3) performing formula detection; 4) performing layout analysis; 5) for inline formulas, according to a formula detection box, determining whether a corrected OCR detection box contains an inline formula, and splitting the OCR detection box containing the inline formula, so as to obtain a text-only OCR detection box; 6) performing formula recognition to obtain a formula recognition result; 7) performing text recognition to obtain a text recognition result; and 8) on the basis of a layout analysis box and a layout category thereof, performing same-line detection box determination and merging on the formula recognition result and the text recognition result, so as to obtain an extraction result for the input text. The present invention can effectively improve the extraction efficiency and accuracy for input texts containing mathematical formulas.
Owner:BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD

An electronic binder system (ebinder) for processing source data to EDC systems

The present invention provides a method and system for automatically and seamlessly processing clinical trial source data into electronic data capture (EDC) systems. In one embodiment, a file structure is defined for an electronic binder system (eBinder); source data is uploaded to the eBinder; the source data is encrypted, Patient Identifiable Information in the source data is masked; the source data is converted into machine readable plain text in the JavaScript Object Notation (JSON) format using Natural Language Processing (NPL) technologies; the JSON data is converted into tabulated machine readable data in the HyperText Markup Language (HTML) format using NPL technologies; the HTML data is converted into machine understandable data using NPL technologies; the machine understandable data is populated into EDC datasets using NPL technologies; the source data and converted data are displayed side-by-side for source data verification; and a platform is provided for regulatory data verification or auditing.
Owner:XIE TAI +1

Open domain-oriented adaptive public opinion data classification method and system

The invention discloses a self-adaptive public opinion data classification method and system for an open domain, and the method comprises the steps: collecting original text data from a plurality of data sources, carrying out the preprocessing of the original text data, obtaining a plain text list, converting the plain text list into high-dimensional semantic vectors in batches, and enabling all high-dimensional semantic vectors to form an embedded matrix; calculating a minimum clustering number and a maximum clustering number according to the number of texts in the plain text list, generating various clustering schemes corresponding to the embedded matrix through all clustering thresholds in a clustering range, calculating a comprehensive score of each clustering scheme, and selecting the clustering scheme with the highest comprehensive score as an optimal scheme; generating subject terms of all clustering clusters based on a large model; and integrating the optimal parameters, the texts in each cluster and the subject terms in each cluster, and outputting a structured classification result. The open domain public opinion data can be efficiently, intelligently and interpretably classified.
Owner:CHENGDU SPACEON IND CO LTD

Title level identification large model training method, title identification method, system and program product

According to the title level identification large model training method, the title identification method and system and the program product provided by the invention, the first title information covering all texts of the original file is constructed through the multi-modal semantic model, and the limitation of plain text identification is made up in combination with the element matching page picture; second title information with semantic and visual features is generated through multi-modal fusion, so that the title judgment accuracy is improved; in the training process, effective title objects are screened in combination with original title objects to optimize pre-training data, and a large model which is high in precision and adapts to complex scenes is cultivated; in the identification process, the title information to be identified is constructed based on the valid title object. And performing dynamic branch processing according to a calling condition, if not, directly outputting an answer, if yes, generating accurate final title information by means of a trained large model, and finally constructing and outputting the answer, thereby realizing training and identification full-link coordination, considering complex scene adaptability and efficient and accurate identification, and comprehensively improving the structuralization and practicability of document title identification.
Owner:SHANGHAI HUNDSUN JUYUAN DATA SERVICE CO LTD +1

Rich text data transmission method, electronic device and program product

The invention provides a rich text data transmission method, electronic equipment and a program product. The rich text data transmission method comprises the steps of obtaining a plain text and rendering information of the plain text included in to-be-transmitted rich text data; obtaining rendering description data of the rendering information based on the plain text; generating structured rendering information based on the rendering description data; and transmitting the plain text and the structured rendering information to a receiving end to execute transmission of the rich text data to be transmitted.
Owner:KE COM (BEIJING) TECHNOLOGY CO LTD

Hybrid retrieval method, system and equipment based on multi-algorithm fusion and storage medium

The invention provides a mixed retrieval method, system and device based on multi-algorithm fusion and a storage medium, and the method comprises the steps: obtaining various types of document data, preprocessing the document data, and obtaining plain text data; segmenting the plain text data into a plurality of pieces of block node data, and generating a plurality of derivative problems for each piece of block node data by using a large language model; carrying out vectorization coding on the block node data and the derivation problem, and storing the block node data and the derivation problem in a vector database; receiving a user query question, and generating a plurality of related preset questions for the query question by using the large language model; converting the query question and the preset question into a vector form, executing keyword retrieval and vector retrieval, and obtaining corresponding candidate results in a vector database; and carrying out weighted fusion on the candidate results through a weighted reciprocal rearrangement algorithm to obtain most relevant retrieval data. According to the method, the efficiency of information retrieval and the correlation of results can be improved, and the retrieval requirement of a user in a complex semantic scene is met.
Owner:XIAMEN INTRETECH

Multi-source document management method and device based on knowledge construction and fusion storage

The invention discloses a multi-source document management method and device based on knowledge construction and fusion storage, and relates to the technical field of document management, and the method comprises the steps: receiving a multi-source document to an object storage system, and recognizing the document type; performing document analysis and structure extraction according to the document type; extracting pictures in the image-text mixed content, uploading the pictures to an object storage system, generating a mapping dictionary, and inserting picture marks in a text part of the mapping dictionary; carrying out structured processing on contents of the table key value pairs to extract summaries and abstracts; performing standardization processing and semantic slicing on the processed image-text mixed content, table key value pair content and / or plain text content to generate knowledge fragments; constructing a knowledge extracting questions from knowledge fragments; the knowledge fragments are stored in a first index, and the questions and IDs of the corresponding knowledge fragments are stored in a second index; generating vectors for each knowledge fragment and question, embedding and writing the vectors into a vector database; and performing document management based on the knowledge base. The document management efficiency is improved.
Owner:DIGITAL CHINA SYST INTEGRATION SERVICE

Industrial automation design environment prompt engineering for generative AI

An integrated development environment (IDE) for designing, programming, and configuring aspects of an industrial automation system uses a generative artificial intelligence (AI) model and associated neural networks to generate portions of an industrial automation project in accordance with functional requirements provided to the industrial IDE system in intuitive formats, such as spoken or written plain language text. The system uses generative AI to translate plain language requests or functional specifications into industrial control code, human-machine interface (HMI) applications, device configuration settings, or other aspects of an industrial control project.
Owner:ROCKWELL AUTOMATION TECH INC

PDF text extraction method and system capable of maintaining text reading sequence, and device

A PDF text extraction method and system capable of maintaining a text reading sequence, and a device. The method comprises: for a PDF text type, performing text extraction to obtain text block coordinate information; on the basis of the text block coordinate information, performing coordinate alignment and calculating coordinates, a row number row_no, a column number col_no, a row processing sequential number row_sequential_number and a column processing sequential number col_sequential_number of each text block; and on the basis of a text recovery type and calculation results, performing text recovery to obtain plain text content.
Owner:FUJIAN FOXIT SOFTWARE DEV LTD

Automatic standard file classification method based on artificial intelligence

The invention discloses a standard file automatic classification method based on artificial intelligence, and relates to the technical field of text classification, and the method comprises the steps: obtaining original data of a to-be-classified file, and carrying out the analysis and preprocessing of the original data of the to-be-classified file, and obtaining metadata, chapter structure information and plain text content; inputting the plain text content and the chapter structure information into a multi-granularity semantic pyramid model for analysis and fusion, and generating a document feature vector; constructing a standard classification system knowledge graph by using the standard classification system data to obtain a classification name mapping table, and performing reinforcement learning on the standard classification system knowledge graph by using a graph neural network technology to generate an enhanced feature vector; and calculating the semantic similarity between the document feature vector and the enhanced feature vector. According to the method, hierarchical reasoning is performed in the knowledge graph based on the similarity to obtain the target classification code, and the classification result is matched and output through the mapping table.
Owner:CHINA STANDARD TECH DEV CORP

Description generation method based on continuous zero sample

The invention provides a description generation method based on continuous zero samples. The description generation method comprises the following steps: sequencing plain text training corpora according to the difficulty of continuous learning; setting that each training step t only can access the t-th corpus, and for the t-th learning step, generating a synthetic image for an input text of the t-th corpus by using a diffusion model; extracting image embedding of the synthesized image by using a CLIP image encoder; retrieving a group of text description embedding with similar semantics in a text corpus by using image embedding; image embedding and text description embedding are input into a fusion module network, and the output of a fusion module is defined as a soft prompt of a large language model; for an input text, key entities are extracted from the text through word frequency statistics so as to construct a soft prompt and a hard prompt of the large language model; hard prompts and soft prompts are input into a large language model network, and key parts in input are concerned through an attention mechanism.
Owner:JIANGSU VOCATIONAL & TECHNICAL UNIVERSITY OF ARCHITECTURE

Project establishment repeatability detection method and device based on ocean engineering scientific research project

The invention relates to the technical field of natural language processing, in particular to a project establishment repeatability detection method and device based on ocean engineering scientific research projects. According to the method, word segmentation processing is carried out on historical scientific research projects by constructing an ocean engineering terminology dictionary, and meanwhile, independent processing is carried out on numerical parameters, so that the problems of low terminology recognition accuracy, misrecognition and the like are solved; when the text similarity is detected, semantic similarity compensation is introduced, invisible correlation is effectively recognized, and the matching limitation is broken through; in addition, the similarity of item attributes is considered, repeatability detection is comprehensively carried out based on text similarity and attribute similarity, and the defects of plain text detection are effectively overcome.
Owner:SHANGHAI INVESTIGATION DESIGN & RES INST CO LTD

RAG-oriented document analysis method and system and computer equipment

The invention relates to the field of artificial intelligence, and provides an RAG-oriented document analysis method and system and computer equipment. The RAG-oriented document analysis method comprises the following steps: uniformly normalizing obtained documents in different formats to obtain Markdown formats of all the documents; the method comprises the following steps: extracting plain text content in a document in a Markdown format, segmenting the plain text content according to a Markdown semantic structure to obtain a plurality of text segments, and vectorizing all the text segments; extracting non-text content in the document in the Markdown format, associating the extracted non-text content with the text fragment vector, and storing the non-text content and the text fragment vector in a relational database; and according to a query request input by a user, retrieving the text fragment vector in the relational database and the non-text content associated with the text fragment vector, and generating a context fragment. And the retrieval accuracy and the integrity of generated answers are improved.
Owner:INSPUR GENERSOFT CO LTD

Document content marking method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a document content marking method, device and equipment and a medium. Matching a target text in the plain text and positioning a matching position; converting the matching position into an original position index by utilizing a mapping relationship; carrying out continuity judgment on the original position index to generate continuous index segments; generating insertion point pairs according to the index segments; and inserting preset starting and ending marking labels at corresponding positions, and generating a target document after marking processing. According to the method, plain text extraction and position mapping are combined, label disorder caused by operation of structured documents is avoided, and the problems of inaccurate cross-label matching and high structure damage rate are effectively solved in cooperation with continuous index aggregation and inverted-order insertion.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Document analysis method, device and equipment and computer readable storage medium

The invention discloses a document analysis method, device and equipment and a computer readable storage medium. The method comprises the following steps: when the type of an original file is a spreadsheet file type and a table in the original file has merged cells, determining the content of each cell in the table, the row number i and the column number j of each non-merged cell in the table and the area coordinate of each merged cell in the table; for each non-merged cell, filling the content of the non-merged cell into the cells in the ith row and the jth column in the correction table; for each merged cell, filling the content of the merged cell into a target cell corresponding to the region coordinate in the correction table; and after all the cells are traversed, embedding the obtained correction table into the plain text document. By means of the method and device, the situation that the format of the content embedded into the plain text document is disordered or information is lost when compared with that of a table in an original file is avoided to a great extent.
Owner:CHONGQING CHANGAN AUTOMOBILE CO LTD

Block chain authentication algorithm model based on textile and clothing supply chain

The invention belongs to the field of textile and clothing supply chains, and particularly relates to a textile and clothing supply chain-based block chain authentication algorithm model, which comprises the following steps of: establishing a block chain network, and distributing a unique digital identity for each participant of the supply chain; classifying the original data into public data and sensitive data by each participant; each participant performs data uplink, public data plaintext uplink and sensitive data encryption uplink through the corresponding digital identity; agreed verification is performed on the public data and the sensitive data on the block chain, and display fragments and key fragments are distributed to participants if agreed verification is satisfied; each operation on the block chain is signed through a digital identity, the authenticity of data is automatically verified in combination with an intelligent contract, full-process tamper-proof traceability and classified chaining of public data and sensitive data are realized, the transparency degree of the block chain is high, the public data can be verified and invisible through a zero-knowledge proof pool, and the sensitive data can be authorized to backtrack; sensitive data leakage is prevented, and the advantages of the block chain are fully utilized.
Owner:无锡物联网创新促进中心

Dialog agents with two-sided modeling

A central learning model is deployed as a user model and as an assistant model. Sensitive information utterances from a corpus of previously stored conversation language corresponding to user queries and chat agent responses thereto are used to train the user model to become an updated user model and to train the assistant model to become an updated assistant model, respectively. The user model provides user contexts corresponding to user queries to the assistant model and the assistant model provides assistant contexts corresponding to chat agent responses to the user model. During training, the user model does not provide plain-text queries to the assistant model and the assistant model does not provide plain-text responses to the user model. The updated assistant model may facilitate a federated training process produce an updated central model. An updated central model may be used to provide real-time chat agent responses to live user queries.
Owner:THE HONG KONG UNIV OF SCI & TECH

Secure network identification for active scanning device

An example operation may include one or more of storing a hash of a service set identifier (SSID) of a wireless network via an apparatus, receiving a probe request message transmitted from a network device, wherein the probe request message comprises a hash value, determining that the hash value within the probe request is a valid SSID based on the hash of the SSID of the wireless network stored in the storage device, and controlling the network interface to transmit a probe response with a plain text name of the SSID to the network device in response to the determination.
Owner:KYNDRYL INC

Prompt input tuning with any unstructured data

Systems and methods are directed to tuning prompt inputs for a large language model (LLM). A prompt tuning system accesses multiple types of unstructured data for use in generating a fusion prompt input embedding. The prompt tuning system then generates a respective embedding for each of the multiple types of unstructured data. The respective embedding for each of the multiple types of unstructured data is provided to a fusion graph neural network (GNN). The fusion GNN generates the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding. The fusion prompt input embedding provides context for a pure text input that comprises instructions for an output from the LLM. The fusion prompt input embedding and a token embedding representing the pure text are inputted to the LLM. An output of the LLM is then displayed. (FIG. 6)
Owner:EBAY INC +7

Network module verification process generation method and system and storage medium

The invention relates to a network module verification process generation method and system and a storage medium, and relates to the technical field of embedded communication, and the method comprises the steps: S1, obtaining a to-be-analyzed AT instruction document, and carrying out the preprocessing of the AT instruction document, and obtaining a plurality of plain text segments; s2, inputting the plain text segment into a preset semantic analysis engine for semantic extraction to obtain a semantic triple; s3, constructing an AT instruction knowledge graph according to the semantic triple; and S4, generating a network module verification process corresponding to the AT instruction document according to the AT instruction knowledge graph. Compared with the prior art, the method has the advantages that zero manual reading can be realized, multi-language and multi-format documents are supported, the time sequence, dependency and conditional branches among instructions are understood, an executable graphical verification process is automatically generated, and incremental updating is supported.
Owner:SOUTH SURVEYING & MAPPING INSTR

A method and device for parsing emails with missing content

The present application discloses a method and device for parsing emails with missing content, relating to the field of email technology. The method comprises: extracting and classifying the text content of the email text to obtain the category of the text content; the category includes plain text format, web page format, and attachment; using the text content in the plain text format and the web page format as the text to be decoded; decoding the text to be decoded according to the corresponding encoding method to obtain a decoded string; performing segmented character set detection on the decoded string according to a set byte length, and adding the character set corresponding to the character set detection result that meets the set conditions to a character set list; performing character conversion on the decoded string by traversing the character sets in the character set list until the decoding result passes the semantic coherence judgment to obtain the original email text. The present application realizes the effective parsing of emails with missing content.
Owner:NAT UNIV OF DEFENSE TECH

Robot control method and device based on image-text interleaving instruction, equipment and medium

The invention relates to the technical field of artificial intelligence, provides a robot control method and device based on an image-text interleaving instruction, equipment and a medium, is applied to financial and medical health care service scenes, and can construct an initial model which takes an image-text mixed format as an input data format and takes a visual language model as a backbone model. The limitation that a traditional vision-language-action model can only process image observation and plain text instructions is broken through; performing language-action pre-training on the initial model based on a potential action distribution function and a flow matching mechanism, so that the model can learn potential representation of action intention from a language; instruction tuning training is carried out based on a hybrid expert mechanism to ensure that the model is flexibly switched between language reasoning and action prediction, so that the model can accurately reasone from a complex and natural image-text instruction and carry out action planning; and generating a predicted action sequence by using the robot action generation model, so as to accurately control the robot based on the image-text interlacing instruction.
Owner:PING AN TECH (SHENZHEN) CO LTD

Multi-subject personalized image generation method, system and device and storage medium

The invention discloses a multi-subject personalized image generation method, system and device and a storage medium, which are corresponding schemes, and in the scheme, based on diffusion blueprint generation and initial noise record of text prior, semantic-level spatial layout prior in a model can be extracted from plain text prompt; a region constraint image cross attention mechanism is introduced, based on the space offset of a blueprint mask, visual features of each subject reference image are forced to take effect only in a mask region specified by a diffusion blueprint, and cross-border interaction of appearance features of each subject is prevented; furthermore, the noise adding-denoising reversibility of the diffusion model is combined with the recorded initial noise, so that the constructed mixed initial noise is obtained; in addition, background image information is injected in a delayed mode in the later stage, and finally natural fusion of the foreground and the background is achieved. The whole scheme is completely constructed based on the pre-training model, a specific subject or scene does not need to be finely adjusted, and good universality and expansibility are achieved.
Owner:UNIV OF SCI & TECH OF CHINA