Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

268 results about "Sequence processing" patented technology

Sequence processing. This subsection of the Sequence section indicates if the canonical sequence displayed by default in the entry is in its mature form or if it represents the precursor.

Machine-Learned User Interface Command Generator Using Pretrained Image Processing Model

An example method can include providing a natural language instruction and user interface image data to a machine-learned sequence processing model that is configured to process image data and generate commands for controlling the target computing device, wherein the machine-learned sequence processing model has parameters learned using an interface recognition objective based on an evaluation of an interface recognition output generated based on processing a rendered training interface from a pre-training dataset and an interface navigation objective based on an evaluation of a user interface command generated based on processing a rendered training interface from a fine-tuning dataset; receiving, from the machine-learned sequence processing model, a command indicating an interaction with the user interface to implement the natural language instruction; and generating, based on the command, a control signal configured to initiate the interaction.
Owner:GOOGLE LLC

Direct posterior preference fine-tuning

Provided is a methodology for direct supervised preference fine-tuning of sequence processing models such as, for example, so-called large language models (LLMs) and large multimodal models (LMMs). The proposed approaches can fine-tune the model to directly predict the posterior token probabilities conditioned on a positive preference of the sequence for which the token is the last token on a sequence of tokens that are the prefix to the sequence. This method offers a simpler fine-tuning approach that directly generates the desired posteriors for use in decoding, without requiring additional inference per vocabulary token at decoding time.
Owner:GDM HOLDING LLC

Metering error correction method and system based on electric power data acquisition

The invention discloses a metering error correction method and system based on electric power data acquisition, and the method comprises the steps: detecting an operation state through a self-inspection and data acquisition module, and obtaining environment data, electrical parameters and error information; key feature data are selected and standardized through a data preprocessing module; an LSTM model and an SVR model are constructed through a model establishment and fusion module, and prediction results are fused to generate error compensation output; real-time error compensation is carried out on a combined prediction result by utilizing an ARIMA model and error time sequence processing through a dynamic error compensation module; the data smoothing and evaluation module is used for smoothing the voltage data and evaluating the error compensation effect; an error compensation effect is monitored through a monitoring and optimizing module, and feature selection and model parameters are optimized regularly; through the fault diagnosis and emergency module, an abnormal condition is detected, and a corresponding standby scheme is triggered. According to the invention, the accuracy and anti-interference capability of electric power metering are effectively improved, and the method is suitable for complex and changeable metering environments.
Owner:STATE GRID NINGXIA ELECTRIC POWER CO LTD MARKETING SERVICE CENT STATE GRID NINGXIA ELECTRIC POWER CO LTD METERING CENT

Image restoration and super-resolution reconstruction system and method based on deep learning

The invention provides an image restoration and super-resolution reconstruction system and method based on deep learning, and belongs to the technical field of digital image processing. The invention aims to solve the problems of high calculation complexity and resource consumption, limitation of long sequence processing, high training difficulty and texture scene deficiency when a multi-scale residual network based on a Transform architecture is used for image resolution conversion. The reconstruction system comprises: an image preprocessing module performing window division and video memory optimization on an input low-resolution image; the multi-layer fusion network dynamically adjusts the characteristics of the low-resolution image, captures channel information in different scenes, performs interactive fusion, performs comparison supervision, establishes an information communication channel, dynamically adjusts and optimizes parameters through negative feedback, and obtains a super-resolution image. And the loss function module maximizes the similarity of the super-resolution image and the high-resolution image in the segmentation feature space to obtain a final super-resolution image.
Owner:QIQIHAR UNIVERSITY

Posterior Preference Optimization

Provided is a framework for fine-tuning pre-trained sequence processing models to human preferences and / or other objective(s). Instead of using reinforcement learning to fine-tune the LLM parameters towards the human preferences, example systems take a Bayesian approach which can preserve the learned prediction distributions of the pre-trained model, but adds explicit sequential preference tuned predictions in a multi-objective model fine-tuning training setup. The model can be tuned to predict posterior token probabilities conditioned on the human preferences.
Owner:GOOGLE LLC

Robot time sequence imitation learning method and system based on Mama coding complete history

The invention belongs to the related technical field of artificial intelligence, and discloses a robot time sequence imitation learning method and system based on a Mama coding complete history, and the method comprises the steps: receiving a multi-modal observation sequence in a task execution process of a robot, the multi-modal observation sequence comprising observation data of at least one sensor; processing the multi-modal observation sequence by using a sequence processing module based on a state space model, and updating a time sequence output of complete historical information of one code up to the current time step at each time step; and predicting the next step or a series of future actions of the robot based on the time sequence output of the current time step so as to control the robot to simulate. The time sequence processing module based on the state space model is utilized to process and encode the complete observation history in the task execution process of the robot, so that a non-Markov decision-making imitation learning method is realized, and the learning efficiency and the execution success rate of the robot in a complex and state-dependent long time sequence operation task are improved.
Owner:HUAZHONG UNIV OF SCI & TECH

Visual task generation method based on Token

The invention discloses a Token-based visual task generation method, and belongs to the technical field of intelligent task automation, and the method comprises the steps: S1, cross-modal alignment; s2, performing visual Token processing; s3, constructing a task description Token sequence: constructing the task description Token sequence based on a predefined visual task template library according to requirements of a user or a specific application scene; s4, checking task feasibility; s5, task priority scheduling; s6, training a task generation model; s7, dynamic task allocation; and S8, optimizing the model. According to the method, the long sequence processing capability is optimized through a hierarchical merging strategy, the compactness of feature expression is realized while space position information is reserved, linear projection and enhanced position coding are combined to form a visual Token sequence with strong representation capability, local detail features are contained, a global context relationship is kept, and the method is suitable for the visual Token sequence with high representation capability. High-information-density feature input is provided for subsequent task processing, and the processing precision of various visual algorithms is effectively improved.
Owner:BEIJING DIGITAL FUTURE TECHNOLOGY CO LTD

Cross-Modal Adapters for Machine-Learned Sequence Processing Models

A machine-learned system for aligning textual and image representations prior to input to a sequence processing model is described. The system includes a machine-learned image embedding model configured to receive image data and generate one or more image embeddings and a machine-learned text embedding model configured to receive text data and the one or more image embeddings and generate one or more text embeddings. The system includes a machine-learned cross-modal adapter configured to generate one or more text tokens aligned with one or more image tokens based at least in part on aligning data associated with the one or more text embeddings and the one or more image tokens. The system includes a machine-learned sequence processing model configured to generate an output based at least in part on the one or more text tokens and the one more image tokens.
Owner:GOOGLE LLC

Aligning Sequence Processing Models with Recommendation Knowledge

The present disclosure provides systems and methods that align sequence processing models with recommendation knowledge. Example training systems can generate natural language prompts, which can be referred to as ‘auxiliary prompts’, that encode different types of recommendation-related knowledge, such as item attributes and user preferences. These auxiliary prompts encode into natural language format various operations and losses that can be used to impart recommendation knowledge to a sequence processing model, including item embedding, Bayesian personalized ranking (BPR), and masked item modeling.
Owner:GOOGLE LLC

Hierarchical Machine-Learned Agents For Performing Mixed Sequence Processing Tasks

A computing device can obtain a first machine-learned sequence processing model configured to use a plurality of first tools, wherein at least one first tool of the plurality of first tools is a second machine-learned sequence processing model configured to use one or more second tools. The computing device can obtain an input context. The computing device can select, using the first machine-learned sequence processing model based at least in part on the input context, a first tool of the plurality of first tools, wherein the first tool selected is the second machine-learned sequence processing model. The computing device can select, using the second machine-learned sequence processing model, at least one second tool of the one or more second tools. The computing device can generate, using the at least one second tool of the one or more second tools, a first output.
Owner:GOOGLE LLC

Double-branch electroencephalogram emotion recognition method and system based on brain region topology and space-time

The invention belongs to the field of artificial intelligence and electroencephalogram emotion recognition, and provides a double-branch electroencephalogram emotion recognition method and system based on brain region topology and time-space, and the method comprises the steps: preprocessing a to-be-recognized electroencephalogram signal to obtain a plurality of electroencephalogram fragments, and extracting a difference entropy sequence of each electroencephalogram fragment and a Spearman correlation coefficient matrix between channels; based on the Spearman correlation coefficient matrix, utilizing a bridging dynamic graph attention network module to extract topological features of a brain region; processing the differential entropy sequence by using a multi-scale space-time mixed attention module to obtain multi-scale space-time features; carrying out residual mutual cross attention fusion on the topological features of the brain region and the multi-scale spatial-temporal features to obtain fusion features; and performing classification based on the fusion features, and determining an emotion recognition result corresponding to the electroencephalogram signal. According to the method, the accuracy and robustness of emotion recognition are improved by utilizing the spatial topology characteristics and the multi-topology time dynamic characteristics of the electroencephalogram signals, and the defects of modeling spatial dependence and time dynamic are overcome.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Automata-based constraints for model decoding

PCT designated stageWO2025213022A1Natural language translationMachine learningPushdown automatonSequence processing
Provided are systems and methods for enhancing the output accuracy and consistency of sequence processing models, particularly when generating text that must conform to a specific formal language. Systems and methods can integrate finite-state or push-down automata with the decoding process to ensure that the generated sequences adhere to the desired syntax or grammar rules. A system can generate or otherwise leverage a finite-state or push-down automaton that encodes constraints associated with a particular formal language or syntax. At each decoding iteration of a sequence processing model, the decoding system can use the finite-state or push-down automaton to evaluate the token validity for each of a number of possible output tokens. The output of the sequence processing model can then be limited to or otherwise guided towards tokens that are indicated as valid by the finite-state or push-down automaton.
Owner:GOOGLE LLC

SoC interrupt processing method based on hardware Sequence

The invention relates to the technical field of system-on-chip service interrupt processing, in particular to a hardware Sequence-based SoC interrupt processing method, which comprises the following steps of: pre-programming an interrupt processing instruction sequence into an instruction random access memory (RAM), and configuring an instruction set, a sequence processor and an interrupt priority strategy of a task scheduler Sequence; receiving the interrupt signal by a task scheduler, and selecting a corresponding sequence processor according to an interrupt type; for the interrupt with low priority, the Sequencer directly executes the pre-stored instruction sequence to complete the interrupt response; and for high-priority interruption, the Sequencer caches the state of the register and sends a message packet to notify the CPU to perform cooperative processing. According to the method, interrupt autonomous response is achieved through hardware Sequence, zero CPU intervention is achieved in low-priority interrupt, a lightweight message mechanism is adopted in high-priority interrupt, and bandwidth occupation and power consumption are reduced while efficiency is improved.
Owner:SHANGHAI FANGYI WANQIANG MICROELECTRONICS CO LTD

Error-Resistant Insight Summarization Using Generative AI

Systems and methods for machine-learned generation of data insight summaries are provided. A computing system can obtain numerical time series data comprising a plurality of numerical values associated with a plurality of times. The computing system can identify, based on the numerical time series data, one or more first mathematical relationships in the numerical time series data. The computing system can generate, based at least in part on the mathematical relationships, a first input context comprising first natural language data indicative of the mathematical relationships. The computing system can provide the first input context to a first machine-learned sequence processing model. The first machine-learned sequence processing model can generate, based at least in part on the first input context, one or more outputs describing the one or more first mathematical relationships. The computing system can output the one or more outputs.
Owner:GOOGLE LLC

Tied Preference Optimization for Sequence Processing Models

Provided are systems and methods for fine-tuning sequence processing models to human preferences. The approaches can account for tied preferences between pairs of sequences and, therefore, can be referred to as Tied Preference Optimization (TPO). Example sequence processing models include so-called large language models (LLMs), large multimodal models (LMMs), and other models that are configured to process inputs and / or generate outputs that are structured as a series of data elements such as tokens.
Owner:GDM HOLDING LLC

Learner cognitive level fine-grained tracking method and system based on state space model

The invention relates to the technical field of education intelligent analysis, and particularly discloses a learner cognition level fine-grained tracking method and system based on a state space model. The method aims at solving the problems that an existing knowledge tracking method is low in cognitive level modeling granularity, weak in long sequence processing capacity, poor in educational interpretation and the like. According to the method, a Bloom cognitive classification system and a state space modeling technology are combined, and fine-grained and multi-level dynamic modeling and future learning performance prediction of the knowledge mastering state of the learner are achieved. The core steps of the method comprise: constructing a semantic mapping relationship between knowledge points and cognitive hierarchies (S101); collecting and encoding multi-source learning behavior features (S102); learning a cognitive state evolution trajectory based on the state space model (S103); and outputting the cognitive hierarchy classification and the answer performance prediction (S104). According to the method, by fusing the multi-dimensional learning behavior data and the state space modeling capability, the accuracy and personalized analysis depth of cognitive tracking are improved, and technical support is provided for precise teaching and intelligent decision making.
Owner:HUAZHONG NORMAL UNIV

Machine Learned Models For Generative User Interfaces

Aspects of the disclosed technology include machine-learning systems and methods for generating user interface elements that allow user control over generative content creation by machine-learned generative models. A generative user interface (UI) system is configured to generate, as output of one or more machine-learned sequence processing models, computer-executable functional code to process a user query in association with a content item. The system is configured to generate computer-executable interface code for a user interface that includes a user interface element associated with at least one parameter of the computer-executable functional code for modifying the content item. The system is configured to determine data for the at least one parameter of the computer-executable functional code based at least in part on a user input to the user interface element and generate a modified content item using the computer-executable functional code and the data for the at least one parameter.
Owner:GOOGLE LLC

Machine-Learning Systems and Methods for Conversational Recommendations

Aspects of the disclosed technology include computer-implemented systems and methods for conversational recommendation systems, such as conversational chatbots that are configured to process user queries and generate responses. A recommendation system includes a conversational user interface configured to receive a user query and provide a recommendation response and a machine-learned sequence processing model that has been trained on training data including a plurality of triplets. Each triplet includes an example query, an example model reasoning plan associated with the example query, and an example response associated with the example query and the example model reasoning plan. The sequence processing model can be trained to provide conversational-based recommendations using a multi-stage recommendation process that includes a planning stage, a conversation stage, and a retrieval stage.
Owner:GOOGLE LLC

Learner knowledge cognition level diagnosis method and system based on cross-scale learning performance dynamic modeling

The invention belongs to the technical field of education data mining and personalized learning, discloses a learner knowledge cognition level diagnosis method and system based on cross-scale learning performance dynamic modeling, and has higher accuracy in the aspects of learner cognition state prediction and knowledge point difficulty assessment. Through a selective state space modeling mechanism and cross-scale historical learning income feature engineering, cognitive change tracks of students in various learning scenes can be accurately captured; and the robustness, convergence efficiency and long sequence processing capability of the model in learner performance prediction are improved. The method can be widely applied to a personalized education platform, a self-adaptive learning system and an intelligent teaching auxiliary tool, provides accurate student learning state analysis for teachers, optimizes learning path design, and improves the teaching effect.
Owner:HUAZHONG NORMAL UNIV

Millimeter wave signal blind source separation and reconstruction system for complex electromagnetic environment

The invention relates to the technical field of signal reconstruction, in particular to a complex electromagnetic environment-oriented millimeter wave signal blind source separation and reconstruction system, which comprises a frequency spectrum trend division module, a path fading construction module, an initial cluster label generation module, a multi-solution path screening module and a fusion reconstruction execution module. According to the method, power spectral density sequence processing is carried out on millimeter wave frequency domain data, a multi-dimensional feature group is formed according to a path loss factor, an angle of arrival and the like, and an initial feature cluster is screened through an Euclidean distance, so that the accuracy of signal source classification is effectively improved; after the frequency domain response and the phase contour are continuously subjected to point comparison, path screening is completed by integrating a mean square error residual value, the path misjudgment probability is reduced, time domain resampling and phase frequency offset standardization are executed in fusion reconstruction, the consistency and fidelity of signal reconstruction are enhanced, and the reconstruction precision is improved. The whole process improves the separation accuracy and reconstruction precision of mixed signals in a complex electromagnetic environment.
Owner:DONGGUAN UNIV OF TECH +1

ADHD graph convolution model construction method based on fMRI spatial-temporal characteristics

The invention belongs to the technical field of deep learning and brain science, and particularly relates to an ADHD graph convolution model construction method based on fMRI spatial-temporal characteristics, and the method comprises the following steps: processing resting state functional magnetic resonance data; fMRI sequence processing: randomly cutting fMRI sequences from different sites to obtain fMRI sequences with consistent sequence lengths, and obtaining a functional connection matrix according to Pearson's correlation; model input data: the model input data is composed of an fMRI sequence, ADHD phenotype information and a functional connection matrix, and the input data is divided to obtain a training set, a verification set and a test set; constructing a model; and model training: model evaluation. According to the method, the gated feature fusion module is proposed to fuse the captured fMRI features to obtain the spatial-temporal features, the phenotypic information of the subject is considered, the ADHD graph convolution model based on the fMRI spatial-temporal features is constructed, and the classification accuracy of ADHD diseases is improved.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Efficient Estimation & Verification with Early Exits

One example aspect is directed to a computer-implemented method for performing model decoding with reduced latency. The method includes obtaining a pre-trained sequence processing model comprising a plurality of layers. The method includes modifying the sequence processing model to contain an adapter layer that is configured to receive and process an intermediate representation generated by a particular intermediate layer of the plurality of layers to predict an output token. The method includes training the adapter layer while holding the plurality of layers of the sequence processing model frozen. The method includes deploying the sequence processing model for speculative decoding in which the adapter layer, the particular intermediate layer, and the plurality of layers that precede the particular intermediate layer perform speculative token decoding and the plurality of layers that are subsequent to the particular intermediate layer perform token verification.
Owner:GOOGLE LLC

Labor dispatch outsourcing management and intelligent manpower dispatch outsourcing system

The invention relates to the technical field of human resource management, in particular to a labor dispatch outsourcing management and intelligent human dispatch outsourcing system, which comprises a process disassembly module, a condition check module, a capability construction module, a connection screening module and a dispatch adaptation module. According to the method, the post task is subjected to nodal sequence processing, digital expression of the post operation process is realized in combination with action labels and operation stage mapping, deviation nodes in the operation process are identified by means of device numbers and sorting verification of operation sections, and the task process precision and the quality control capability are improved; post boundary condition constraint verification is completed by fusing post equipment permission and an operation section period table, the probability of occurrence of task matching conflicts is remarkably reduced, and the coherence consistency of personnel and post actions is improved by comparing behavior tags with a post action sequence, so that accurate matching and efficient coherence deployment of tasks and personnel are realized, and the task matching efficiency is improved. And the intelligence, the matching degree and the performability of labor dispatching are enhanced.
Owner:SHANXI IDEAL FUTURE HUMAN RESOURCES CO LTD

Text generation method and device based on pre-trained language model, equipment and medium

The invention provides a text generation method and device based on a pre-trained language model, equipment and a medium, and can be applied to the technical field of text generation. The method comprises the following steps: inputting a target question text into a pre-trained language model to generate a target path; generating a confusion value based on the conditional probability of each step in the logic step sequence; in response to determining that the confusion value is greater than the preset threshold value, correcting at least one step in the logic step sequence by utilizing a pre-trained language model to obtain a corrected logic step sequence; and processing the target question text and generating the target reply text according to the corrected logic step sequence by using the pre-trained language model, so that the correctness of each step in the logic step sequence is ensured, the reasoning efficiency is improved, the consumption of computing resources is reduced, and meanwhile, the robustness and interpretability of the target reply text are improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Multi-modal emotion recognition method based on Mama state space model and cross-modal self-distillation

The invention belongs to the technical field of artificial intelligence and multi-modal emotion calculation, and discloses a multi-modal emotion recognition method based on a Mama state space model and cross-modal self-distillation. Through the organic combination of the efficient sequence modeling capability of the Mamba state space model and the knowledge sharing mechanism of cross-modal self-distillation, the advantages of the state space model in the aspects of time sequence modeling and calculation efficiency are fully played, and meanwhile, the limitation of a single model architecture is made up through a cross-modal attention mechanism; the technical bottlenecks of an existing multi-modal emotion recognition method in the aspects of long sequence processing efficiency, cross-modal information fusion and knowledge transfer sufficiency are effectively solved, and an efficient and reliable technical solution is provided for further development and practical application of the multi-modal emotion recognition technology.
Owner:NORTHEASTERN UNIV CHINA

Multi-domain feature fusion network and bearing fault diagnosis method and system

The invention provides a multi-domain feature fusion network and a bearing fault diagnosis method and system, and relates to the technical field of deep learning. The network comprises a feature extraction module, a feature fusion module and a classification module. The feature extraction module performs feature extraction on input data by three branch lines to obtain bearing vibration time domain features, frequency domain features and time-frequency domain features; and the feature fusion module performs cross-domain fusion and dynamic adjustment fusion on the three features through the multi-head attention mechanism module and the feature calibration mechanism module to obtain accurate multi-domain fusion features, so that deep mining of input data is realized, and information of a fault state is captured more comprehensively. And a self-adaptive multi-pooling fusion structure is designed in the classification module, so that the diagnosis robustness of the model is enhanced, and the accuracy and comprehensiveness of fault classification are improved. The multi-domain feature fusion network is remarkably superior to an existing model in the aspects of noise robustness, variable working condition adaptability and short sequence processing capacity.
Owner:HEFEI UNIV OF TECH

Sequence processing method, electronic equipment, storage medium and program product

The invention relates to the technical field of artificial intelligence, and provides a sequence processing method, electronic equipment, a storage medium and a program product.The method comprises the steps that a preset alternative block size set is obtained, and the alternative block size set comprises a plurality of different alternative block sizes; obtaining a sequence length of a target sequence to be processed, and determining a target block size corresponding to the target sequence from the alternative block size set based on the sequence length, the target block size is the alternative block size with the minimum invalid filling data volume generated when the target sequence is subjected to blocking processing in the alternative block size set; and performing attention calculation on the target sequence based on the target block size. According to the method, the optimal block size capable of minimizing invalid calculation is dynamically matched for the sequences with different lengths, and self-adaptive optimization of the calculation load and the sequence length is realized, so that the calculation efficiency and the throughput capacity can be remarkably improved, and meanwhile, the consumption of memory resources is reduced.
Owner:SHANGHAI BIREN TECH CO LTD

Machine-Learned Model Alignment With Synthetic Data

Aspects of the disclosed technology include computer-implemented systems and methods for adapting machine-learned models using high-quality synthetic data that is tailored to elicit improved instruction-following abilities for particular target instruction distributions and models. A model adaptation system can obtain instruction metadata indicative of at least one use case and at least one skill associated with a particular computing task to be performed by a target machine-learned model. The system can generate a metadata-conditioned synthetic instruction by prompting a generative model system including one or more machine-learned generative models with the instruction metadata as one or more constraints. The system can generate a model response by prompting the generative model system with the metadata-conditioned synthetic instruction. The system can modify a target sequence processing model based at least in part on a data pair including the metadata-conditioned synthetic instruction and the model response.
Owner:GOOGLE LLC

Transverse mixed attention mechanism model training method, medium, device and program product

The invention provides a model training method for a transverse mixed attention mechanism, a medium, equipment and a program product, and the method comprises the steps: obtaining a data set containing a plurality of sample sequences, each sample sequence in the data set being formed by arranging a plurality of Token sequences obtained through word segmentation; constructing a to-be-trained model based on the pre-trained full attention model, and adding newly added parameters for linear attention calculation; in the same transverse mixed attention layer, executing total attention calculation on a Token set in a preset total attention calculation range, executing linear attention calculation on all Tokens, and fusing results of the total attention calculation and the linear attention calculation to obtain transverse mixed attention output used for forward reasoning and loss calculation; and based on the output and prediction result, only updating the newly added parameters to optimize the to-be-trained model until the to-be-trained model converges. According to the method, the calculation complexity and video memory occupation of long text sequence processing are reduced, and the reasoning speed and the resource utilization rate are improved.
Owner:BEIJING JIBU QIANLI TECHNOLOGY CO LTD

Background coherent story picture book generation method based on diffusion model

PendingCN121392035A2D-image generationBiological modelsFrame (artificial intelligence)Linguistic model
The invention discloses a background coherent story picture book generation method based on a diffusion model, and belongs to the field of computer vision and generative artificial intelligence. The method comprises a training stage and a testing stage: in the training stage, bidirectional cross attention fusion and modal soft selection are carried out on text, background and role multi-modal conditions through a feature enhancement fusion module, model optimization is carried out by utilizing joint alignment loss, and efficient parameter fine tuning is carried out on a diffusion model by adopting an efficient parameter fine tuning method; in the test stage, a reference image input by a user and a text sequence are processed into a fusion condition, a large language model is driven to generate an image mark, the image mark is converted into a diffusion condition through a mapper, and each frame of image is generated step by step by combining an autoregression mode with a multi-condition injection diffusion network. According to the method, the problem of inconsistency of cross-frame backgrounds, styles and roles is effectively solved, and high-quality and coherent generation of the long-sequence story picture book is realized.
Owner:JIANGXI NORMAL UNIV