Method, apparatus, device and storage medium for information processing

CN122654470APending Publication Date: 2026-08-28BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238303.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0002]随着计算机水平的发展,扩散模型可以用于处理各种模态的输入内容,并且利用注意力机制提升输出结果的准确性,但是针对输入内容的内容过多或者内容复杂的情况,扩散模型的处理效率较低,因此需要更为优化的扩散模型,以提升模型整体的效率

Benefits of technology

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122654470A_ABST
    Figure CN122654470A_ABST
Patent Text Reader

Abstract

According to an embodiment of the present disclosure, a method, an apparatus, a device and a storage medium for information processing are provided. The method comprises: obtaining input content of a diffusion model; determining a first attention matrix by using an attention unit at a first denoising time step of the diffusion model; determining a plurality of target blocks from a plurality of candidate blocks based on weight information corresponding to the plurality of candidate blocks in the first attention matrix; determining a second attention matrix based on an attention mask by using the attention unit at a second denoising time step of the diffusion model; and generating target content corresponding to the input content based on at least the second attention matrix. In this way, the embodiments of the present disclosure can reuse the attention mask generated at the previous denoising time step in other denoising time steps of the diffusion model, so that the diffusion model can improve the processing efficiency of the overall model while maintaining the quality of the output result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The exemplary embodiments disclosed herein generally relate to the field of computers, and more particularly to methods, apparatus, devices and computer-readable storage media for information processing. Background Technology

[0002] With the development of computer technology, diffusion models can be used to process input content of various modalities and improve the accuracy of output results by using attention mechanisms. However, when the input content is too large or complex, the processing efficiency of diffusion models is low. Therefore, more optimized diffusion models are needed to improve the overall efficiency of the model. Summary of the Invention

[0003] In a first aspect of this disclosure, an information processing method is provided. The method includes: acquiring input content of a diffusion model, the diffusion model including a converter unit, the converter unit including an attention unit; at a first denoising time step of the diffusion model, using the attention unit to determine a first attention matrix of a first query feature and a first key feature, the first attention matrix corresponding to full attention; determining multiple target blocks from the multiple candidate blocks based on weight information corresponding to multiple candidate blocks in the first attention matrix; at a second denoising time step of the diffusion model, using the attention unit to determine a second attention matrix of a second query feature and a second key feature based on an attention mask, the attention mask being determined based on the multiple target blocks; and generating target content corresponding to the input content, at least based on the second attention matrix.

[0004] In a second aspect of this disclosure, an apparatus for information processing is provided. The apparatus includes: an acquisition module configured to acquire input content of a diffusion model, the diffusion model including a converter unit, the converter unit including an attention unit; a first determination module configured to, at a first denoising time step of the diffusion model, use the attention unit to determine a first attention matrix of a first query feature and a first key feature, the first attention matrix corresponding to full attention; a second determination module configured to determine a plurality of target blocks from the plurality of candidate blocks based on weight information corresponding to the plurality of candidate blocks in the first attention matrix; a third determination module configured to, at a second denoising time step of the diffusion model, use the attention unit to determine a second attention matrix of a second query feature and a second key feature based on an attention mask, the attention mask being determined based on the plurality of target blocks; and a generation module configured to generate target content corresponding to the input content, at least based on the second attention matrix.

[0005] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes computer-executable instructions that, when executed by a processor, implement the method according to a first aspect of this disclosure.

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram of an example environment in which embodiments of the present disclosure can be implemented is shown;

[0011] Figure 2 A flowchart illustrating an information processing procedure according to some embodiments of the present disclosure is shown;

[0012] Figure 3 A schematic diagram of a full attention mechanism according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic diagram of attention masks according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A flowchart illustrating the processing of information in a diffusion model according to some embodiments of the present disclosure is shown;

[0015] Figure 6 A schematic structural block diagram of an apparatus for information processing according to certain embodiments of the present disclosure is shown;

[0016] Figure 7 A block diagram of an electronic device capable of implementing several embodiments of the present disclosure is shown. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should be noted that the headings of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and embodiments of any type may be included under any section / subsection. Furthermore, embodiments described in any section / subsection may be combined in any way with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0019] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0020] The embodiments of this disclosure may involve user data, data acquisition, and / or use. These aspects all comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the user should be informed of the data or data type, scope of use, and usage scenarios that may be involved, and their authorization should be obtained, in accordance with relevant laws and regulations and through appropriate means. The specific methods of notification and / or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.

[0021] In this specification and the embodiments, any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal data beyond that required for basic functions will not affect the user's use of basic functions.

[0022] Traditionally, diffusion models are content-based diffusion models that generate new samples by simulating the process of gradually transforming data from noise into a target distribution. The core idea is to progressively add noise to transform the data distribution into a simpler one, and then recover the data from the noise through a reverse process. Because applying attention mechanisms to diffusion models can more accurately accomplish various generation tasks, they are widely used. However, if the input content for diffusion models is too large or too complex, such as a long video, the computational cost of the diffusion model increases with the increase in video resolution or duration, ultimately leading to lower processing efficiency.

[0023] The embodiments of this disclosure propose an information processing scheme. According to the scheme, input content of a diffusion model is obtained. The diffusion model includes a converter unit, and the converter unit includes an attention unit. At a first denoising time step of the diffusion model, the attention unit determines a first attention matrix for a first query feature and a first key feature, the first attention matrix corresponding to full attention. Based on the weight information corresponding to multiple candidate blocks in the first attention matrix, multiple target blocks are determined from the multiple candidate blocks. At a second denoising time step of the diffusion model, the attention unit determines a second attention matrix for a second query feature and a second key feature based on an attention mask, the attention mask being determined based on the multiple target blocks. And at least based on the second attention matrix, target content corresponding to the input content is generated.

[0024] Based on this approach, embodiments of this disclosure can reuse the attention mask generated in previous denoising time steps in other denoising time steps of the diffusion model, thereby enabling the diffusion model to improve the overall processing efficiency of the model while maintaining the quality of the output results.

[0025] Example Environment

[0026] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. For example... Figure 1 As shown, example environment 100 may include electronic device 110.

[0027] In some embodiments, the electronic device 110 can acquire input content from the diffusion model 120, which includes a converter unit 130 and an attention unit 140. In a first denoising time step of the diffusion model 120, the attention unit 140 determines a first attention matrix for a first query feature and a first key feature, the first attention matrix corresponding to full attention. Based on the weight information corresponding to multiple candidate blocks in the first attention matrix, multiple target blocks are determined from the multiple candidate blocks. In a second denoising time step of the diffusion model, the attention unit 140 determines a second attention matrix for a second query feature and a second key feature based on an attention mask, the attention mask being determined based on the multiple target blocks. And, at least based on the second attention matrix, target content corresponding to the input content is generated. The diffusion model 120 can be deployed on the electronic device 110, or on other devices, which will not be elaborated here.

[0028] In some embodiments, the electronic device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 may also support any type of interface for the target user (such as "wearable" circuitry).

[0029] Electronic device 110 can also be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Electronic device 110 may include, for example, computing systems / servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.

[0030] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.

[0031] The following description will continue with reference to the accompanying drawings, which will provide some exemplary embodiments of this disclosure.

[0032] Example process

[0033] Figure 2 A flowchart of an information processing procedure 200 according to some embodiments of the present disclosure is shown. Procedure 200 can be implemented at electronic device 110. Reference is made below. Figure 1 Describe the process 200.

[0034] In box 210, the electronic device can acquire the input content of the diffusion model.

[0035] In some embodiments, the input content can be any form of content that can be input into the diffusion model for information processing. The input content can include content of one modality or content of multiple modalities. For example, the input content can be text content in natural language processing or image content in computer vision, or the input content can also include a combination of multiple modalities such as text and video, images and videos, etc., which will not be elaborated here.

[0036] In some embodiments, a diffusion model is a network model that generates a target result by simulating the process of gradually transforming data from noise into a target distribution. Its core idea is to gradually add noise to transform the data distribution into a simple distribution, and then recover the data from the noise through a reverse process. For example, diffusion models can achieve tasks such as text-to-image synthesis, realistic video generation, and 3D content creation.

[0037] In some embodiments, to enhance the efficiency of the diffusion model in extracting important features, enabling it to automatically focus on key features or regions in the input content, an attention mechanism can be applied to the diffusion model in this application embodiment. The attention mechanism is a deep learning technique that simulates human attention allocation, allowing the model to dynamically focus on key parts while ignoring unimportant information when processing input data. The core idea of ​​this mechanism is to calculate the relevance of different parts of the input data to the current task, assigning different weights to each part, thereby extracting key information more effectively.

[0038] In some embodiments, the diffusion model may include a converter unit, which includes an attention unit. For example, the diffusion model may be a DiT (diffusion transformer) model. The DiT model can optimize predictions through the diffusion process and process multimodal data such as video and text with the help of an attention mechanism. The attention mechanism can capture spatial, temporal and cross-modal dependencies.

[0039] In box 220, electronic device 110 uses attention units to determine a first attention matrix for the first query feature and the first key feature at the first denoising time step of the diffusion model.

[0040] Since the diffusion model transforms the data distribution into a simple distribution by gradually adding noise, and then recovers the data from the noise through a reverse process, the content generation process using the diffusion model involves multiple denoising time steps. This embodiment takes the first denoising time step as an example. For instance, the electronic device 110 can input the input content into the diffusion model, and then perform data preprocessing on the input content. For example, it can perform normalization, feature extraction, and dimensionality transformation on the input content to convert the input data into a format suitable for subsequent attention calculation and obtain the first feature sequence.

[0041] In some embodiments, since the input content may include more than one modality, a three-dimensional full attention mechanism can be used to calculate the first attention matrix in order to improve the fusion performance between different modalities. The three-dimensional full attention mechanism is an attention mechanism for handling content generation tasks. It captures global dependencies in the spatial and temporal dimensions by converting the input content into a sequence and applying self-attention to the entire sequence.

[0042] for example, Figure 3 This illustrates the working principle of the 3D full attention mechanism in DiT. The input content includes video and text. After the input content is transformed into a first feature sequence, it includes a first video frame sequence 310, a second video frame sequence 320, and a text sequence 330. Then, query features 350 and key features 340 can be obtained based on the first feature sequence. Finally, a first attention matrix 360 is constructed through the full attention mechanism. The number of frames in the video can be defined as f, the spatial resolution of each frame is h×e, and t is the length of the text sequence, where f·h·w >> t. The total sequence length L corresponding to the input content is shown below:

[0043] L=f·h·w+t

[0044] In some embodiments, a multi-head attention mechanism can also be used in the diffusion model. This mechanism, through parallel and independent attention mechanisms, allows the model to simultaneously focus on different feature subspaces of the input content, thereby enabling the model to learn multi-dimensional features. For example, the first feature sequence can be linearly transformed to obtain query features Q, key features K, and value features V, where Q, K, V ∈ R. H×L×DThe query features represent the content that needs attention, the key features represent the features of each element in the first feature sequence, and the value features represent the actual content of each element. H is the number of attention heads, L is the length of the first feature sequence, and D is the dimension of each attention head. Then, the similarity between the query features and the key features is calculated to obtain a score matrix. To avoid the gradient vanishing or exploding problem caused by excessively large dot product results, this score matrix can be scaled. The scaled scores are then normalized using the softmax function to obtain the first attention matrix W. attn ∈R L×L The first attention matrix represents the attention weights, and the calculation formula for the first attention matrix can be as follows:

[0045]

[0046] In box 230, electronic device 110 determines multiple target blocks from multiple candidate blocks based on the weight information corresponding to multiple candidate blocks in the first attention matrix.

[0047] In some embodiments, the computational cost of applying the attention mechanism in a diffusion model is very high, especially for long videos. Therefore, generating high-fidelity long videos is often limited by significant latency. For example, the first attention matrix might be an L×L matrix, which would result in O(L...)... 2 The computational complexity is high, both in terms of time and memory, and the computational cost increases with the video resolution or duration, which is particularly unfriendly for input content with a large amount of information.

[0048] To address the issues of high computational complexity and redundancy in attention mechanisms when processing long sequences, a sparse attention mechanism can be adopted in the attention module. The sparse attention mechanism can significantly reduce computational costs and process long sequence data more efficiently by limiting the scope or pattern of attention computation, while maintaining or improving performance.

[0049] In some embodiments, attention masks are typically used to control parameter updates in a neural network or to filter specific elements in a tensor. It is implemented by introducing a binary mask matrix where most elements are zero and only a few positions are non-zero. In this application embodiment, the attention mask indicates that multiple target blocks among multiple candidate blocks are retained for attention operations. For example, in attention mechanisms, the attention mask can indicate which interactions between elements can be omitted and which interactions need to be retained to reduce computational load. For instance, interactions with smaller weights can be ignored to reduce computational complexity, thus not only ensuring model performance but also significantly improving model efficiency.

[0050] In some embodiments, there are multiple sparsity patterns for sparse attention mechanisms. However, since the sparsity of the DiT diffusion model varies considerably depending on the input content, sparse patterns for offline search of DiT lack good portability and accuracy. Furthermore, because the sparse indexes in DiT are complex, and the key regions are scattered rather than concentrated and continuous, sparse patterns for approximate search of DiT cannot accurately estimate the sparse indexes in DiT.

[0051] Furthermore, since the sparse attention mechanism of the DiT diffusion model exhibits significant hierarchical features within and between different modalities, for example, due to the significant differences in the interaction between video frames, the global attention weights show obvious block-based features based on frames, and the interaction between different frame blocks differs significantly, with a stronger aggregation tendency within specific frame blocks. Therefore, the diffusion model can effectively simulate block-based patterns.

[0052] Therefore, this application employs a block-based approach for sparse attention computation, which eliminates the resource consumption associated with selecting a suitable sparse mode. The block-based approach reduces computational complexity and improves efficiency by dividing the input sequence into multiple blocks and performing sparse computations across these blocks. The core idea of ​​this approach is to leverage the sparsity of the attention mechanism, performing computations only between specific blocks, rather than performing full attention computations across the entire sequence.

[0053] In some embodiments, the electronic device 110 can acquire a first parameter indicating the size and / or number of candidate blocks, which may be, for example, hyperparameters of the model. The electronic device 110 can then divide the first attention matrix into multiple candidate blocks based on the first parameter. This method avoids loading the entire first attention matrix into high-bandwidth memory at once, significantly reducing the number of read / write operations on high-bandwidth memory and thus improving memory access efficiency. It also avoids explicitly storing the complete first attention matrix, making the attention mechanism more efficient when processing long sequences.

[0054] In some embodiments, the electronic device 110 can sort multiple candidate blocks based on the weight information corresponding to the multiple candidate blocks, and then determine multiple target blocks from the multiple candidate blocks based on the sorting result of the multiple candidate blocks. For example, if the size of the first attention matrix is ​​L×L and the size of a candidate block is B×B, the attention mask can be defined as... Where M ij =1 indicates that element i is interested in element j, while M ij =0 indicates that the interaction between element i and element j is ignored, M ij The set of indices S where 1 = 1 can be called a sparse index set. For ease of description, we can use M...s Will be expanded to The masked attention matrix after applying an attention mask to the first attention matrix is ​​expressed by the following formula;

[0055]

[0056] It is a relatively large negative bias, and c can be set to a sufficiently large number. Understandably, M... s The value can be 0 or 1, and correspondingly, The value can also be 0 or 1. When When the value is 0, the value within softmax is negative infinity, and the final calculated value is... That is, 0; when When the value is 0, the values ​​in the softmax function are the weights calculated based on the attention mechanism. This representation allows the mask attention matrix to ignore unimportant parts.

[0057] In some embodiments, applying different attention masks to the first attention matrix will result in different masked attention matrices. Since the goal of sparse attention mechanisms is to reduce computational cost while maintaining accuracy, the smaller the difference between the first attention matrix and the masked attention matrix, the more accurate the application of the attention mask. Mathematically, the desired attention mask can also be obtained by constructing an attention loss mechanism. For example, a mask representing the first attention matrix W can be defined. attn The concept of W, the sum of the weights of all candidate blocks. sum-attn The formula is expressed as follows:

[0058]

[0059] Therefore, given the sparsity, the sparse index set S can be represented as follows:

[0060]

[0061] W can be calculated sum-attn And the top-k operation is used to obtain the optimal attention mask. This method reduces the computational complexity from O(L...) 2 d) Reduced to O((1-sparsity)L 2 d), thereby achieving a significant acceleration effect.

[0062] In some embodiments, there are various methods for determining the target block from multiple candidate blocks based on the weight information corresponding to multiple candidate blocks. For example, the target mask can be obtained by methods such as fixed threshold or adaptive threshold. However, since the diffusion model needs to go through multiple denoising time steps, and the attention mask can be reused in other denoising time steps, the computational load is greatly reduced. Therefore, the embodiments of this application can use the top-k method to filter out k target blocks, thereby obtaining more accurate data.

[0063] In one embodiment, log-sum-exponential (LSE) can be applied to process the attention score after masking, thereby ensuring numerical stability during the attention calculation process. Log-sum-exponential (LSE) is a technique for numerically stable computation; its core objective is to convert the exponential sum of a set of numbers into the logarithmic field, thus avoiding numerical overflow or underflow problems that may occur during direct calculation.

[0064] In some embodiments, the electronic device 110 can determine the logarithmic summation-exponential LSE based on the weight information of multiple target blocks in the first attention matrix, and generate reference weight information based on the logarithmic summation-exponential LSE. For example, when calculating the attention score, an attention mask can be applied to ensure that only the elements at positions 1 in the attention mask participate in the calculation. Then, the LSE can be stably calculated by subtracting the maximum value of each row, avoiding numerical problems in exponential operations. It can be defined... and use z j Let j represent the j-th component of the row vector Z. The formula for calculating LSE can be as follows:

[0065] LSE(z)=log∑ j exp(z j )

[0066] =max j z j +log∑ j exp(z j -max k z k )

[0067] This method ensures that no numerical overflow or underflow occurs when calculating softmax, which can be represented as follows:

[0068]

[0069] The numerically stable mask attention matrix can then be expressed as follows:

[0070]

[0071] After obtaining the reference weight information based on LSE, it can be cached so that it can be used directly in subsequent steps.

[0072] In some embodiments, during the process of filtering target blocks from candidate blocks, the number of target blocks selected is determined based on a second parameter, which indicates a preset sparsity. Sparsity refers to the proportion or percentage of non-zero elements in data or a model, and can be, for example, a hyperparameter of the model. For instance, electronic device 110 can obtain a preset sparsity, then determine the proportion of target blocks in the candidate blocks based on the preset sparsity, and then determine the target number of target blocks from the multiple candidate blocks based on the weight information corresponding to the multiple candidate blocks in the first attention matrix.

[0073] In some embodiments, for multi-head attention mechanisms, since not all attention heads have the same sparsity characteristics—some attention heads may perform well when retaining less content, while others may perform well when retaining more content—it is unreasonable to reuse the same attention mask for all attention heads. Therefore, in a multi-head attention mechanism, the attention unit in the diffusion model includes multiple attention heads, and each attention head corresponds to an independent attention mask. That is, in a multi-head attention mechanism, the attention mask for each attention head is calculated separately, rather than reusing or referencing other attention masks.

[0074] In some embodiments, since recall can measure the extent to which sparse patterns can retain attention, the effectiveness of sparse attention mechanisms can also be evaluated through recall. If the recall indicates that the effect obtained through the sparse attention mechanism is not good, the overall model performance can be improved by adjusting the sparsity. The electronic device 110 can determine a first set of attention masks for multiple attention heads based on a preset sparsity, then determine the recall information corresponding to the multiple attention heads based on the first set of attention masks, and then increase the first sparsity corresponding to the first attention head among the multiple attention heads and decrease the second sparsity of the second attention head among the multiple attention heads based on the recall information, wherein the first recall corresponding to the first attention head is higher than the second recall corresponding to the second attention head.

[0075] For example, in a multi-head attention mechanism, the attention mask for each attention head is calculated separately. These attention masks can form the first set of attention masks, and then the recall information for each attention head can be obtained separately. The recall formula can be as follows:

[0076]

[0077] A higher recall indicates better preservation of the original attention score. The sparsity of the n attention heads with the highest recall can then be increased to... And reduce the sparsity of the n attention heads with the lowest recall to This method can effectively reduce redundancy among attention heads with high recall while improving the accuracy of attention heads with low recall, thus reducing the waste of computational resources and improving the overall accuracy of the model.

[0078] In some embodiments, since the attention weight matrix obtained in the attention mechanism has a clear boundary between the text modality and the pure video modality, it exhibits different degrees of text sinking effect. Therefore, the electronic device 110 can determine multiple additional blocks associated with the text prompt words from multiple candidate blocks in the first attention matrix, and then construct an attention mask based on the multiple target blocks and the multiple additional blocks.

[0079] For example, if the input content is detected to include both text and video modalities, then it can be done as follows: Figure 4 As shown, Figure 4 The multiple rectangular regions 410 in the image represent the target blocks selected according to the segmentation mode. Then, the candidate blocks in the rectangular region 420 on the right and the rectangular region 430 at the bottom of the image are determined as additional blocks. This can enhance the perception of the video modality on the text modality, thereby achieving better results.

[0080] In some embodiments, ensuring that each query feature pays attention to approximately the same number of key features in the attention mechanism can improve the coherence of the generated video; otherwise, some regions deemed unimportant might never be addressed, resulting in artifacts. Therefore, uniform selection by row can be enforced in block sparse mode.

[0081] In box 204, the electronic device 110 can use the attention unit to determine the second attention matrix of the second query feature and the second key feature based on the attention mask at the second denoising time step of the diffusion model.

[0082] In one embodiment, since the same sparsity method is applied in the sparse attention mechanism, the resulting recall value will vary depending on the attention head and the different converter units in the diffusion model. However, it will not change much with the change of the denoising time step. In other words, the attention mask that has been obtained can be reused in other denoising time steps.

[0083] For example, when applying the same attention mask in a sparse attention mechanism, the recall values ​​obtained in the 10th denoising time step, layer 0, and the 6th attention head are similar to those obtained in the 20th denoising time step, layer 0, and the 6th attention head; the recall values ​​obtained in the 10th denoising time step, layer 0, and the 6th attention head are different from those obtained in the 10th denoising time step, layer 0, and the 18th attention head; and the recall values ​​obtained in the 10th denoising time step, layer 0, and the 6th attention head are different from those obtained in the 10th denoising time step, layer 15, and the 6th attention head.

[0084] In some embodiments, in the first denoising time step of the diffusion model, an attention mask can be obtained by determining the target block. Then, in the second denoising time step of the diffusion model, the attention mask obtained in the first denoising time step can be applied again, thus omitting the process of re-obtaining the attention mask and greatly improving the efficiency of the overall model. Furthermore, since the recall rate obtained by applying the same attention mask in different denoising time steps remains basically unchanged, it will not affect the overall model performance.

[0085] In some embodiments, since the same LSE is applied in the sparse attention mechanism, the resulting recall value may vary depending on the attention head or the network layer, but it will not change much with the denoising time step. That is, the LSE already obtained can be reused in other denoising time steps. For example, in the second denoising time step of the diffusion model, the electronic device 110 can use the attention unit to determine the second initial matrix of the second query feature and the second key feature, then reuse the reference weight information to perform data stabilization processing on the second initial matrix to obtain the optimized matrix, and then reuse the attention mask on the optimized matrix to obtain the second attention matrix.

[0086] For example, if the reference weight information has already been calculated in the first denoising time step of the diffusion model, it can be applied again in the second denoising time step, thus eliminating the need to re-acquire the reference weight information and significantly reducing additional search time. Furthermore, since the recall rate obtained by applying the same LSE in different denoising time steps remains essentially unchanged, accurate searching can also be ensured.

[0087] In box 250, electronic device 110 generates target content corresponding to the input content based at least on the second attention matrix.

[0088] In some embodiments, since the recall value obtained by applying the same attention mask in the sparse attention mechanism will still change due to the change in the denoising time step, the attention mask can be recalculated after several denoising time steps to ensure overall accuracy. For example, the electronic device 110 can determine the second attention mask in the third denoising time step of the diffusion model, and then use the attention unit to determine the third attention matrix based on the second attention mask in the fourth denoising time step of the diffusion model to generate the target content.

[0089] For example, it can be pre-set that the attention mask is re-acquired at the 10th denoising time step. This 10th denoising time step is also the third denoising time step. At the 10th denoising time step, the step of determining the attention matrix using attention units to identify query and key features is re-executed. Then, based on the weight information corresponding to multiple candidate blocks in the attention matrix, multiple target blocks are determined from the candidate blocks, and a new second attention mask is generated based on the target blocks. Then, in the fourth denoising time step following the 10th denoising time step, this second attention mask is reused to obtain the third attention matrix for the third query and third key features. This obtained third attention matrix is ​​then used in subsequent denoising time steps for information processing, ultimately yielding the target content.

[0090] In some embodiments, since the recall value obtained by applying the same reference weight information in the sparse attention mechanism changes very little due to the change in the denoising time step, the same reference weight information can be reused in all denoising time steps, which greatly improves efficiency while ensuring accuracy. The electronic device 110 can obtain the cached reference weight information and then determine the second attention mask based on the cached reference weight information in the third denoising time step of the diffusion model.

[0091] For example, the cached reference weight information can be obtained, which is determined based on the first attention matrix. Then, in the denoising time step after the first denoising time step, this reference weight information can be reused to obtain the attention mask.

[0092] In some embodiments, Figure 5 A flowchart illustrating the structure of information processing using a diffusion model is shown. The denoising time step T in the full attention stage can be defined. w ={1,2,…,t w}, and select k denoising time steps. Perform precise online searches, and From denoising time step 1 to denoising time step t w-1 employs a full attention mechanism 510, in the denoising time step t w At that time, the application performs full attention computation using the fusion online search 520, thereby generating a first attention mask, which can then be passed to subsequent denoising time steps. To perform head-adaptive hierarchical block sparse attention 530 computation. Subsequently, for each For i>1, the preorder traversal can be used. The found cached LSE performs an online LSE cache search 540 to obtain a second attention mask, which is then passed to subsequent denoising time steps. To complete the computation of head-adaptive hierarchical block sparse attention 550.

[0093] In some embodiments, since the technical solution of this application works better when applied to input content including video, the information processing method can be applied to input content including video.

[0094] In some embodiments, a plug-and-play plugin can also be provided that allows for seamless integration into DiT without fine-tuning or data profiling, and it is independent of other acceleration techniques such as parallelization, cache reuse, and token merging.

[0095] By combining head-adaptive hierarchical block sparse attention with online search techniques, latency is significantly reduced while maintaining high-quality results. Furthermore, since it requires no additional fine-tuning or dataset analysis, it can be provided as a plug-and-play plugin, seamlessly integrating into existing diffusion models.

[0096] Example devices and equipment

[0097] Embodiments of this disclosure also provide corresponding apparatus for implementing the above methods or processes. Figure 6 A schematic structural block diagram of an apparatus 600 for information processing according to certain embodiments of the present disclosure is shown. The apparatus 600 may be implemented as or included in the electronic device 110 discussed above. The various modules / components in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0098] like Figure 6As shown, the apparatus 600 includes an acquisition module 610 configured to acquire input content of a diffusion model, the diffusion model including a converter unit, the converter unit including an attention unit; a first determination module 620 configured to determine a first attention matrix of a first query feature and a first key feature using the attention unit at a first denoising time step of the diffusion model, the first attention matrix corresponding to full attention; a second determination module 630 configured to determine multiple target blocks from multiple candidate blocks based on weight information corresponding to multiple candidate blocks in the first attention matrix; a third determination module 640 configured to determine a second attention matrix of a second query feature and a second key feature using the attention unit based on an attention mask at a second denoising time step of the diffusion model, the attention mask being determined based on multiple target blocks; and a generation module 650 configured to generate target content corresponding to the input content based at least on the second attention matrix.

[0099] In some embodiments, the information processing apparatus 600 further includes a fourth determining module 640 configured to determine a second attention mask at a third denoising time step of the diffusion model; and a fifth determining module 650 configured to use an attention unit to determine a third attention matrix of a third query feature and a third key feature based on the second attention mask at a fourth denoising time step of the diffusion model, for use in generating target content.

[0100] In some embodiments, the fourth determining module 640 is further configured to obtain cached reference weight information, which is determined based on the first attention matrix; and to determine a second attention mask based on the cached reference weight information at the third denoising time step of the diffusion model.

[0101] In some embodiments, the reference weight information includes the log-sum-exponential (LSE) determined based on the weight information of multiple target blocks in the first attention matrix.

[0102] In some embodiments, an attention mask indicates that multiple target blocks among multiple candidate blocks are retained for attention operations.

[0103] In some embodiments, the second determining module 630 is further configured to sort multiple candidate blocks based on the weight information corresponding to the multiple candidate blocks; and to determine multiple target blocks from the multiple candidate blocks based on the sorting result of the multiple candidate blocks.

[0104] In some embodiments, the information processing device 600 is further configured to acquire a first parameter indicating the size and / or number of candidate blocks; and to divide the first attention matrix into a plurality of candidate blocks based on the first parameter.

[0105] In some embodiments, the number of multiple target blocks is determined based on a second parameter, which indicates a preset sparsity.

[0106] In some embodiments, an attention unit includes a plurality of attention heads, and the plurality of attention heads correspond to independent attention masks.

[0107] In some embodiments, the information processing apparatus 600 is further configured to determine a first set of attention masks for a plurality of attention heads based on a preset sparsity; determine recall information corresponding to the plurality of attention heads based on the first set of attention masks; and increase the first sparsity corresponding to the first attention head among the plurality of attention heads and decrease the second sparsity of the second attention head among the plurality of attention heads based on the recall information, wherein the first recall corresponding to the first attention head is higher than the second recall corresponding to the second attention head.

[0108] In some embodiments, the information processing device 600 is further configured to determine multiple additional blocks associated with text prompt words from multiple candidate blocks in a first attention matrix; and to construct an attention mask based on multiple target blocks and multiple candidate blocks.

[0109] In some embodiments, the target content includes video content generated by the diffusion model.

[0110] In some embodiments, the diffusion model includes multiple converter units, and the attention units in different converter units correspond to independent attention masks.

[0111] The units included in device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units may be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 600 may be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that may be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0112] Figure 7 A block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that... Figure 7 The electronic device 700 shown is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to achieve Figure 1The electronic device 110 shown.

[0113] like Figure 7 As shown, electronic device 700 is in the form of a general-purpose electronic device. Components of electronic device 700 may include, but are not limited to, one or more processors or processing units 710, memory 720, storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. Processing unit 710 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 720. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 700.

[0114] Electronic device 700 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 720 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 730 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store data and / or data (e.g., training data for training) and can be accessed within electronic device 700.

[0115] Electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 7 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 720 may include computer program product 725 having one or more program modules configured to perform various methods or actions of various embodiments of this disclosure.

[0116] The communication unit 740 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 700 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 700 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.

[0117] Input device 750 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 760 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 700 can also communicate with one or more external devices (not shown) via communication unit 740 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 700, or with any device that enables electronic device 700 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0118] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0119] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0120] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable information processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable information processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable information processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0121] Computer-readable program instructions can be loaded onto a computer, other programmable information processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable information processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable information processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0123] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. An information processing method, comprising: The input content of the diffusion model is obtained, wherein the diffusion model includes a converter unit, and the converter unit includes an attention unit; In the first denoising time step of the diffusion model, the attention unit is used to determine a first attention matrix for the first query feature and the first key feature, and the first attention matrix corresponds to full attention. Based on the weight information corresponding to multiple candidate blocks in the first attention matrix, multiple target blocks are determined from the multiple candidate blocks; In the second denoising time step of the diffusion model, the attention unit is used to determine a second attention matrix based on the second query feature and the second key feature, wherein the attention mask is determined based on the plurality of target blocks; as well as At least based on the second attention matrix, target content corresponding to the input content is generated.

2. The method according to claim 1, wherein the attention mask is a first attention mask, and before generating target content corresponding to the input content, the method further includes: In the third denoising time step of the diffusion model, the second attention mask is determined; as well as In the fourth denoising time step of the diffusion model, the attention unit determines a third attention matrix based on the second attention mask, which is a third query feature and a third key feature, to generate the target content.

3. The method according to claim 2, wherein, In the third denoising time step of the diffusion model, determining the second attention mask includes: Obtain cached reference weight information, the reference weight information being determined based on the first attention matrix; and In the third denoising time step of the diffusion model, the second attention mask is determined based on the cached reference weight information.

4. The method of claim 3, wherein the reference weight information includes a log-sum-exponential (LSE) determined based on the weight information of the plurality of target blocks in the first attention matrix.

5. The method of claim 1, wherein the attention mask indicates that the plurality of target blocks among the plurality of candidate blocks are retained for attention computation.

6. The method according to claim 1, wherein, The step of determining multiple target blocks from the multiple candidate blocks based on the weight information corresponding to the multiple candidate blocks in the first attention matrix includes: Based on the weight information corresponding to the multiple candidate blocks, the multiple candidate blocks are sorted; and Based on the sorting results of the multiple candidate blocks, the multiple target blocks are determined from the multiple candidate blocks.

7. The method according to claim 1, further comprising: Obtain a first parameter, which indicates the size and / or number of candidate blocks; as well as Based on the first parameter, the first attention matrix is ​​divided into the multiple candidate blocks.

8. The method of claim 1, wherein the number of the plurality of target blocks is determined based on a second parameter, the second parameter indicating a preset sparsity.

9. The method of claim 1, wherein the attention unit comprises a plurality of attention heads, and the plurality of attention heads correspond to independent attention masks.

10. The method of claim 9, further comprising: Based on a preset sparsity, a first set of attention masks is determined for the multiple attention heads; Based on the first set of attention masks, the recall information corresponding to the multiple attention heads is determined; as well as Based on the recall information, the first sparsity corresponding to the first attention head among the plurality of attention heads is increased, and the second sparsity of the second attention head among the plurality of attention heads is decreased, wherein the first recall corresponding to the first attention head is higher than the second recall corresponding to the second attention head.

11. The method according to claim 1, wherein the input content includes text prompts, and the method further includes: Multiple additional blocks associated with the text prompt words are determined from the multiple candidate blocks in the first attention matrix; as well as The attention mask is constructed based on the multiple target blocks and the multiple additional blocks.

12. The method of claim 1, wherein the target content includes video content generated by the diffusion model.

13. The method of claim 1, wherein the diffusion model comprises a plurality of the converter units, and the attention units in different converter units correspond to independent attention masks.

14. An apparatus for information processing, comprising: The acquisition module is configured to acquire the input content of the diffusion model, which includes a converter unit and an attention unit. The first determining module is configured to, at the first denoising time step of the diffusion model, use the attention unit to determine a first attention matrix of a first query feature and a first key feature, wherein the first attention matrix corresponds to full attention. The second determining module is configured to determine multiple target blocks from the multiple candidate blocks based on the weight information corresponding to the multiple candidate blocks in the first attention matrix; The third determining module is configured to, at the second denoising time step of the diffusion model, use the attention unit to determine a second attention matrix based on the attention mask for the second query feature and the second key feature, wherein the attention mask is determined based on the plurality of target blocks; as well as The generation module is configured to generate target content corresponding to the input content, at least based on the second attention matrix.

15. An electronic device comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 13.

16. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method according to any one of claims 1 to 13.

17. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 13.