Content generation method and device, equipment and storage medium

By generating the model to process the input content and calculating the correlation information between the token and the subnet, determining the routing information to activate the appropriate subnet, the problem of poor routing strategy in the diffusion model when processing visual information is solved, and the overall effect is improved.

CN120106130APending Publication Date: 2025-06-06DOUYIN VISION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192392.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When processing visual information, existing diffusion models require more optimized routing strategies to improve overall results due to spatial redundancy and complexity of denoising tasks.

Method used

By processing the input content by generating the model, determining the token corresponding to the input content, and calculating the correlation information between the token and the subnet, based on this, the routing information is determined to activate the appropriate subnet and generate the target content.

Benefits of technology

The routing strategy between the token and the subnet is optimized, improving the overall effect of the generative model, especially when processing complex visual information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106130A_ABST
    Figure CN120106130A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a content generation method and device, equipment and a storage medium. The method comprises the following steps: processing input content by utilizing a generative model to determine a first group of tokens corresponding to the input content; determining correlation information between the first group of tokens and a plurality of sub-networks in the generative model; determining routing information based on the correlation information; activating at least one sub-network indicated by the routing information to process corresponding tokens indicated by the routing information to generate a second set of tokens; and generating target content corresponding to the input content based on the second group of tokens. Based on the mode, according to the embodiment of the invention, the routing strategy of the generation model can be effectively optimized by obtaining the correlation information comprising the multiple correlation degrees corresponding to the multiple token-sub-network pairs and adopting the mode of determining the routing information through one-time screening, and the overall effect of the generation model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for content generation. Background Art

[0002] With the development of artificial intelligence, network models can be used to perform a variety of complex tasks. However, since the amount of tasks and the degree of difficulty of tasks at different stages of the network model are different, more optimized routing strategies are needed to improve the overall effect of the network model. Summary of the invention

[0003] In a first aspect of the present disclosure, a method for content generation is provided. The method includes: processing input content using a generation model to determine a first group of tokens corresponding to the input content; determining correlation information between the first group of tokens and multiple subnetworks in the generation model, the correlation information indicating multiple degrees of association corresponding to multiple token-subnetwork pairs; determining routing information based on the correlation information, the routing information indicating at least one token-subnetwork pair determined from multiple token-subnetwork pairs, wherein at least one token-subnetwork pair is determined based on the ranking of multiple degrees of association, or at least one token-subnetwork pair is determined based on a global threshold corresponding to the correlation information; activating at least one subnetwork indicated by the routing information to process the corresponding tokens indicated by the routing information to generate a second group of tokens; and generating target content corresponding to the input content based on the second group of tokens.

[0004] In a second aspect of the present disclosure, a device for content generation is provided. The device includes: a calling module configured to process input content using a generation model to determine a first group of tokens corresponding to the input content; a first determination module configured to determine correlation information between the first group of tokens and multiple sub-networks in the generation model, the correlation information indicating multiple degrees of association corresponding to multiple token-sub-network pairs; a second determination module configured to determine routing information based on the correlation information, the routing information indicating at least one token-sub-network pair determined from multiple token-sub-network pairs, wherein at least one token-sub-network pair is determined based on the ranking of multiple degrees of association, or at least one token-sub-network pair is determined based on a global threshold corresponding to the correlation information; a token processing module configured to activate at least one sub-network indicated by the routing information to process the corresponding token indicated by the routing information to generate a second group of tokens; and a content generation module configured to generate target content corresponding to the input content based on the second group of tokens.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory is coupled to the at least one processing unit and stores instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and the computer program can be executed by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer executable instructions, and when the instructions are executed by a processor, the method according to the first aspect of the present disclosure is implemented.

[0008] It should be understood that the contents described in this content section are not intended to limit the key features or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:

[0010] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0011] Figure 2 A flowchart showing a process of content generation according to some embodiments of the present disclosure;

[0012] Figure 3 A schematic diagram of a network structure of a generation module according to some embodiments of the present disclosure is shown;

[0013] Figure 4 A schematic diagram of the structure of a hybrid expert network module according to some embodiments of the present disclosure is shown;

[0014] Figure 5 A schematic diagram of pre-layer regularization according to some embodiments of the present disclosure is shown;

[0015] Figure 6 A schematic structural block diagram of an apparatus for content generation according to some embodiments of the present disclosure is shown;

[0016] Figure 7 A block diagram of an electronic device capable of implementing various embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0018] It should be noted that the titles of any sections / subsections provided herein are not restrictive. Various embodiments are described throughout this article, and any type of embodiment may be included under any section / subsection. In addition, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or different sections / subsections.

[0019] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.

[0020] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects are subject to the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user knows and confirms. Accordingly, when implementing each embodiment of the present disclosure, the type, scope of use, usage scenario, etc. of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method can vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.

[0021] In this specification and the embodiments, if personal information processing is involved, it will be processed on the premise of having a legal basis (such as obtaining the consent of the subject of personal information, or it is necessary to perform a contract, etc.), and will only be processed within the scope of regulations or agreements. If a user refuses to process personal information other than the necessary information for basic functions, it will not affect the user's use of basic functions.

[0022] Traditionally, the diffusion model is a content generation model that generates new samples by simulating the process of data gradually transforming from noise to target distribution. The core idea is to transform the data distribution into a simple distribution by gradually adding noise, and then recover the data from the noise through the inverse process. However, the visual information processed by the diffusion model usually has high spatial redundancy. For example, there is a significant difference in the information density between the background and foreground areas of visual information. The foreground area usually contains more key details. In addition, the complexity of the denoising task will show temporal changes at different time steps. It is relatively simple to predict the noise at the beginning of the denoising process, but it is much more complicated to predict the noise near the end because finer details need to be reconstructed in the later stages. These characteristics require the diffusion model to adopt a more optimized routing strategy.

[0023] The embodiment of the present disclosure proposes a solution for content generation. According to the solution, input content is processed using a generation model to determine a first group of tokens corresponding to the input content; correlation information between the first group of tokens and multiple subnetworks in the generation model is determined, and the correlation information indicates multiple degrees of association corresponding to multiple token-subnetwork pairs; based on the correlation information, routing information is determined, and the routing information indicates at least one token-subnetwork pair determined from multiple token-subnetwork pairs, wherein at least one token-subnetwork pair is determined based on the ranking of multiple degrees of association, or at least one token-subnetwork pair is determined based on a global threshold corresponding to the correlation information; at least one subnetwork indicated by the routing information is activated to process the corresponding tokens indicated by the routing information to generate a second group of tokens; and based on the second group of tokens, target content corresponding to the input content is generated.

[0024] Since the routing information can be determined from the correlation information by a one-time screening method, such as by sorting the correlation degree or by screening by a global threshold, the routing strategy between the tokens and the sub-networks can be optimized, thus improving the overall effect of the generated model.

[0025] Example Environment

[0026] Figure 1 1 is a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Figure 1 As shown, example environment 100 may include electronic device 110 .

[0027] In some embodiments, the electronic device 110 can process the input content using the generation model 120 to determine a first set of tokens corresponding to the input content, determine the correlation information between the first set of tokens and multiple sub-networks in the generation model, determine routing information, activate at least one sub-network indicated by the routing information to process the corresponding tokens indicated by the routing information to generate a second set of tokens, and generate target content corresponding to the input content based on the second set of tokens. The generation model 120 can be deployed on the electronic device 110, and can also be deployed on other devices, which will not be described in detail here.

[0028] In some embodiments, the electronic device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable game terminals, VR / AR devices, personal communication systems (PCS) devices, personal navigation devices, personal digital assistants (PDA), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the electronic device 110 can also support any type of interface for the target user (such as "wearable" circuits, etc.).

[0029] The electronic device 110 may also be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The electronic device 110 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.

[0030] It should be understood that the structure and function of the various elements in the environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure.

[0031] Some example embodiments of the present disclosure will be described below with continued reference to the accompanying drawings.

[0032] Example Process

[0033] Figure 2FIG. 2 is a flowchart of a process 200 for content generation according to some embodiments of the present disclosure. For example, the process 200 may be implemented in Figure 1 The electronic device 110 shown in the figure may be implemented by a combination of the electronic device 110 and other computing devices. Figure 1 The process 200 is described below.

[0034] In process 200 , the electronic device 110 performs content generation processing on input content.

[0035] In block 210 , the electronic device 110 processes the input content using the generative model to determine a first set of tokens corresponding to the input content.

[0036] In some embodiments, the input content may be any form of content that can be input into the network model for content processing. For example, the input content may be media content in the form of text, image, video, etc., which will not be elaborated here.

[0037] In some embodiments, the generation model is a network model that can be used to implement the content generation function. The specific structure and type of the generation model are not limited in the embodiments of the present application. For example, the generation model can be a diffusion model, such as the DiT model (Diffusion Transformer). The diffusion model is a content generation model that generates new samples by simulating the process of data gradually transforming from noise to target distribution. The core idea is to transform the data distribution into a simple distribution by gradually adding noise, and then recover the data from the noise through the inverse process.

[0038] In some embodiments, the electronic device 110 can input the input content into the generation model and call the generation model to convert the input content into a series of tokens, i.e., a first group of tokens, which are the basis for content understanding and content processing of the network model for subsequent operations. The embodiment of the present application does not limit the form of the input content.

[0039] There are many ways to convert the input content into tokens. For example, if the input content is in text form, the input text can be segmented into subwords or words; if the input content is in image form, the input image can be segmented into multiple blocks, so as to convert the input content into a format that can be processed by the network model.

[0040] In some embodiments, since the visual information processed by the diffusion model usually has a high spatial redundancy, for example, there is a significant difference in information density between the background area and the foreground area in the image, and the foreground area usually contains more detail information than the background area, and in the diffusion model, the complexity of the denoising task will also show considerable differences at different time steps. For example, it is relatively simple to predict the noise at the beginning of the denoising process, while it is much more complicated to predict the noise near the end of the denoising process because more refined detail reconstruction operations are required in the later stages.

[0041] It is precisely because of the above characteristics of the diffusion model itself that it is necessary to provide a new routing strategy for the diffusion model to improve the overall effect of the diffusion model in processing content. Figure 3 A network structure diagram of a generative model is shown, based on which process 300 can be performed, the network structure diagram includes a scaling layer 310, a hybrid expert network module 320, a scaling and translation layer 330, a layer normalization layer 340, and a multi-head self-attention mechanism 350. The generative model replaces the multi-layer perceptron (MLP) module in the traditional DiT model with Figure 3 The hybrid expert network module 320 shown, Figure 4 A schematic diagram of the structure of the hybrid expert network module 320 is shown, and the sub-network in the generated model is the expert network. Based on the hybrid expert network module structure diagram, process 400 can be executed. This is a method of introducing the idea of ​​the hybrid expert network model into the diffusion model, which can complete the allocation of sub-networks for tokens at one time. Each token can be allocated to any number of sub-networks, and each sub-network can also process any number of tokens. In this way, the routing module of the hybrid expert network model can flexibly allocate and combine tokens and expert networks according to the prediction difficulty, thereby achieving the purpose of allocating more expert networks for more challenging tasks, so as to make more reasonable task allocation in the diffusion model and ultimately improve the overall model performance.

[0042] Among them, the hybrid expert network model is a machine learning network model built based on the divide-and-conquer idea. The hybrid expert network model is a neural network layer that consists of multiple expert networks {E i} and a routing module R, with a total of N E The core idea is to decompose a complex task into multiple subtasks and process them separately by different expert networks, each of which specializes in processing a certain part of the input token.

[0043] Among them, the expert network is an independent network model, each expert network specializes in processing a certain part of the input content, different expert networks usually have different parameters and structures, and expert networks can be divided into many types, such as neural networks, linear models, or other types of models, etc. Since different expert networks are good at different functions, choosing a more suitable expert network for processing different tasks will achieve better results.

[0044] Among them, the routing module is responsible for deciding how to assign input tokens to each expert network for processing. The routing module can map the input content to the correlation information between the first group of tokens and multiple expert networks in the generative model, and then process it through a gating function. After processing, the input token will be assigned to the top K expert networks with the highest scores for processing, and the output is the weighted sum of the outputs from these expert networks. In order to increase the number of choices for the top K expert networks while keeping the number of activation parameters unchanged, the hybrid expert network model usually splits the hidden dimensions of the middle layer of the expert network according to the values ​​of the top K expert networks. "x-in-y" means that there are a total of y candidate expert networks, of which the top x expert networks are activated, and the hidden dimensions of the middle layer of the expert network will be divided by x.

[0045] In block 220 , the electronic device 110 determines correlation information between the first set of tokens and a plurality of sub-networks in the generative model.

[0046] In some embodiments, the electronic device 110 can call the generation model to determine the correlation information between the first group of tokens and multiple sub-networks in the generation model, and the correlation information indicates multiple association degrees corresponding to multiple token-sub-network pairs, and the association degree is the importance score of each sub-network when processing tokens. In subsequent operations, the correlation information can be used to calculate the output of the gating network, that is, to determine how to allocate sub-networks for the first group of tokens and determine the contribution of each sub-network to the final result.

[0047] Among them, the token-subnetwork pair refers to the mapping relationship between the token and the subnetwork. For example, token T1 and subnetwork E1 can constitute a token-subnetwork pair, token T3 and subnetwork E2 can constitute a token-subnetwork pair, and so on.

[0048] In some embodiments, the electronic device 110 may determine an association matrix between the first group of tokens and the plurality of sub-networks; and determine correlation information based on the association matrix. For example, the electronic device 110 may determine the number of networks of the sub-networks in the generation model and the number of tokens in the first group of tokens, and then call the generation model to generate first score information corresponding to each token-sub-network pair, and then construct an association matrix including two dimensions of the number of networks of the sub-networks and the number of tokens in the first group of tokens based on the first score information, and then determine the correlation information between the first group of tokens and the plurality of sub-networks in the generation model based on the association matrix.

[0049] In some embodiments, in order to ensure the independence between rows and columns and improve the flexibility of the routing module in sub-network allocation. The electronic device 110 can first determine the number of networks in the sub-network in the generation model and the number of tokens in the first group of tokens. Then call the generation model to generate the first score information corresponding to each token-sub-network pair, and the first score information is used to indicate the degree of correlation between the token and the sub-network. Then, the association matrix S is constructed based on the multiple first score information obtained, and the association matrix includes two dimensions: the number of networks in the sub-network and the number of tokens in the first group of tokens. Figure 4 A form of the correlation matrix 420 is shown in FIG. Then, the correlation information between the first group of tokens and the multiple sub-networks in the generation model is determined according to the correlation matrix. Figure 4 One form of correlation information 410 is shown in FIG. B represents the size of the candidate pool of the expert network, D A Indicates the number of parallel select operations.

[0050] In some embodiments, the electronic device 110 may also construct a single-dimensional correlation tensor based on multiple correlations in the correlation matrix as correlation information. Figure 4 A structural diagram of the hybrid expert network module is shown. The first group of tokens includes four tokens T1, T2, T3, and T4. The current generation model includes three sub-networks E1, E2, and E3. Then, the generation model can be called to generate the first score information corresponding to each token-sub-network pair, and then a correlation matrix 420 is constructed based on the multiple first score information obtained. Then, a one-dimensional correlation tensor is constructed based on the correlation matrix 420 as the correlation information 410 between the first group of tokens and the multiple sub-networks in the generation model.

[0051] Among them, each element in the association matrix 420 indicates the association between the token and the sub-network corresponding to the position of the element. For example, the element located at the position (1,1) of the association matrix 420 can indicate the association between token T1 and sub-network E1, and the element located at the position (2,3) of the association matrix 420 can indicate the association between token T2 and sub-network E3, and so on.

[0052] In some embodiments, the electronic device 110 may call the generation model to generate second score information corresponding to each token-subnetwork pair, and then perform a mapping operation on the second score information to obtain the first score information. The mapping operation may be implemented by applying an identity function.

[0053] The identity function formula refers to a function that has the same value as the parameter, and its function formula is G(x) = x. By using this method, we avoid using softmax for normalization, which not only avoids disrupting the score relationship between different tokens, but also greatly reduces the computational cost and avoids the risk of numerical underflow as the sequence length in the network model increases, thereby significantly improving the overall effect of the model generation results.

[0054] In some embodiments, a sigmoid function may be further applied to perform a mapping operation on the second score information. The embodiment of the present application does not limit the type of function applied in the mapping operation.

[0055] In block 230 , the electronic device 110 may determine routing information based on the correlation information.

[0056] The routing information indicates at least one token-subnetwork pair determined from a plurality of token-subnetwork pairs, and a routing strategy for the token allocation subnetwork can be determined according to the routing information.

[0057] In some embodiments, at least one token-subnetwork pair in the routing information can be determined based on the numerical order of the correlation degrees corresponding to multiple token-subnetwork pairs in the correlation information. For example, the electronic device 110 can sort the correlation degrees corresponding to multiple token-subnetwork pairs in the correlation information according to the numerical size, and then screen the token-subnetwork pairs according to the sorting result to determine the routing information.

[0058] In some embodiments, routing information can also be determined based on a pre-acquired global threshold. For example, the electronic device 110 can obtain a pre-set global threshold, and then compare the value of the correlation degree corresponding to each token-subnetwork pair in the correlation information with the value of the global threshold, and then screen the token-subnetwork pairs according to the comparison result to determine the routing information.

[0059] In some embodiments, for a hybrid expert network model or other similar network models, during the batch training phase, multiple samples are processed simultaneously, these samples will affect each other's routing selection, and the time steps are randomly sampled, while in the inference phase, the samples are processed independently, so there is no mutual influence between samples within the batch, and the time steps are fixed or processed in a consistent manner. This difference will lead to inconsistencies between batch training and inference. Since the time step directly controls the noise mixing level, this inconsistency will reduce the generation quality and may cause model failure. At the same time, the mutual influence between samples during the routing selection process will cause the inference process to be unstable.

[0060] Therefore, in order to mitigate these effects, the embodiment of the present application proposes a global threshold τ applied in the model reasoning stage, which is determined based on the data obtained in the model training stage, and then the global threshold can be directly used for data screening in the model reasoning stage to ensure sample independence and consistency. The electronic device 110 can determine the routing information according to the numerical sorting of the correlation degrees corresponding to multiple token-subnetwork pairs in the correlation information during the training stage of the generated model; and determine the routing information according to the comparison result of the numerical value of the correlation degree corresponding to each token-subnetwork pair in the correlation information and the numerical value of the global threshold during the reasoning stage of the generated model.

[0061] In some embodiments, in practical scenarios such as generative model training, the complexity of data generation varies due to two key factors: the denoising time step B and the spatial image area L. In order to cope with this computational heterogeneity, the routing module needs to dynamically allocate more sub-networks to tokens with greater generation requirements. In order to maximize the flexibility of the routing module to learn adaptive allocation patterns so that sub-networks can allocate tokens according to computational requirements, the existing routing strategy can be generalized by using a top-K selection mechanism in the correlation information.

[0062] In some embodiments, the correlation information between the first group of tokens and the multiple subnetworks in the generation model can be defined as an overall candidate pool. The electronic device 110 can screen K token-subnetwork pairs that meet the preset screening conditions from the overall candidate pool at one time according to the value of the correlation degree corresponding to each token-subnetwork pair in the overall candidate pool. The screened K token-subnetwork pairs can constitute an effective selection area of ​​the candidate pool. According to the effective selection area of ​​the candidate pool, the routing information can be determined. The size of the effective selection area of ​​the candidate pool can be defined as:

[0063]

[0064] Among them, K represents the effective selection area size of the candidate pool, k represents the number of sub-networks expected to be activated by each token, and k can control the sparsity of the hybrid expert network module, N E represents the number of sub-networks in the generative model, D B Represents the size of the overall candidate pool. Then the process of the routing module applying the top-K selection mechanism to screen the overall candidate pool can be achieved by the following maximization formula:

[0065]

[0066] Among them, T i represents the index set of the top K correlations of the values ​​in the i-th row of the correlation information S', D A represents the number of parallel selection operations. Applying the above maximization formula, we can set D A =1 provides an optimal solution for the above maximization formula, which ensures that the values ​​of the correlations corresponding to the K token-subnetwork pairs selected are globally maximized through one-time screening. For example, Figure 4 shows the structural diagram of the hybrid expert network module. Figure 4 It can be seen that according to the maximization formula, the top four token-subnetwork pairs can be screened out from the correlation information 410, which are represented by bold lines respectively, and then the routing information can be determined according to these four token-subnetwork pairs.

[0067] In some embodiments, the electronic device 110 may determine multiple training associations corresponding to multiple token-subnetwork pairs during the training phase, and then determine a reference association corresponding to a preset sorting position by sorting the multiple training associations, and then determine a global threshold to be applied to the reasoning phase based on the reference association.

[0068] For example, the electronic device 110 selects four token-subnetwork pairs from the correlation information, and the correlation degrees corresponding to these four token-subnetwork pairs are multiple training correlation degrees, and the routing information is determined based on these four token-subnetwork pairs. Then, the four token-subnetwork pairs can be sorted according to the numerical values ​​of the training correlation degrees corresponding to them to obtain the sorting results, and then the token-subnetwork pair with the fourth largest correlation value is determined from the four token-subnetwork pairs. The correlation degree corresponding to the token-subnetwork pair with the fourth largest correlation value is the reference correlation degree, and then the model weights of the generated model are updated according to the reference correlation degree to obtain the global threshold applied to the inference stage of the generated model.

[0069] In some embodiments, an exponential moving average can be applied to update the model weights, wherein the exponential moving average is to reduce the fluctuation of the network model during the training phase by maintaining the sliding average of the network model weights, making the training of the network model more stable, and because the smooth weight update can usually better capture the long-term trend of the network model, it also improves the generalization ability of the network model. For example, the calculation formula of the global threshold τ can be as follows:

[0070]

[0071] Among them, S′ i,K represents the correlation degree corresponding to the token-subnetwork pair with the Kth largest correlation value in the i-th row of S', τ represents the global threshold, and D A Represents the number of parallel selection operations, and m represents the smoothing coefficient. The average value of the network model weights is gradually updated during the training phase of the network model using exponential moving average to obtain a global threshold, which is then used during the inference phase of the network model to provide a more stable performance for the network model. This approach can minimize the impact of differences between the batch processing phase and the inference phase.

[0072] In some embodiments, the electronic device 110 may also adopt BL routing strategy, BE routing strategy, and LE routing strategy to obtain routing information. These three strategies are obtained by selecting paired combinations of dimensions, where B represents batch, L represents sequence length, and E represents the number of subnetworks. The above routing strategies can introduce different degrees of routing selection flexibility, thereby improving the performance of the network model.

[0073] In block 240 , the electronic device 110 may activate at least one sub-network indicated by the routing information to process corresponding tokens indicated by the routing information to generate a second set of tokens.

[0074] In some embodiments, the electronic device 110 can decide which subnetwork or subnetworks to apply to process which token or tokens based on the routing information, and then perform an activation operation on the subnetwork that needs to be activated, and use at least one activated subnetwork among multiple subnetworks to process corresponding tokens in the first group of tokens indicated by the routing information to generate a second group of tokens, which can be applied to subsequent content generation operations.

[0075] In block 250 , the electronic device 110 generates target content corresponding to the input content based on the second set of tokens.

[0076] In some embodiments, since the present application hopes to apply a generation model to perform content generation operations on the input content, after the electronic device 110 obtains the second set of tokens, it can apply the second set of tokens to subsequent content generation operations to generate target content corresponding to the input content.

[0077] As an example, the target content may include content of any appropriate modality. For example, in a scenario where media content is generated based on text, the input content may include text content, and the generated target content may include media content such as audio, video, image, etc.

[0078] In some embodiments, the diversity of routing strategies can be increased by constructing a first loss indicating the similarity of routing strategies of the tokens for the sub-networks, and the collapse of sub-networks due to following the same token allocation rule can be avoided as much as possible to improve the overall network performance. The electronic device 110 can determine the first loss based on the correlation information, the first loss indicating the similarity of routing strategies of the tokens for the sub-networks; and train the generation model based on the first loss.

[0079] In some embodiments, the electronic device 110 may obtain the correlation information generated by the generation model during the model training phase, and determine the correlation matrix S∈R according to the correlation information. (B×L)×E , and the routing information corresponding to the association matrix, and then apply the normalization function softmax along the sub-network dimension to obtain the normalized probability matrix P, and calculate the following two correlation matrices:

[0080] M′=M T M

[0081] P′=P T P

[0082] Where M is the indicator matrix, if subnetwork j selects the i-th token, then M i,j =1, otherwise 0, P is the normalized probability matrix. Since the off-diagonal elements in the probability matrix P represent the similarity between each pair of sub-networks based on the token selection pattern in the current batch, P i ' ,j Represents the geometric mean of the balanced loss, so the diversity between sub-networks can be improved by calculating the cross-correlation matrix and minimizing its off-diagonal elements to maximize the specialization of the sub-networks. The first loss function formula constructed based on the correlation matrix generated during the model training phase is as follows:

[0083]

[0084] Among them, T represents the number of the first group of tokens corresponding to the generative model in the training phase, E represents the number of sub-networks in the generative model, and P i ' ,j represents the joint probability that the token is routed to subnetworks i and j, and W(i,j) is a weighting function, which can be defined as follows:

[0085]

[0086] Applying this first loss to train the generative model can standardize the consistent common selection patterns between sub-networks in the generative model and promote the combination of different sub-networks to avoid different tokens being processed by exactly the same sub-network, thereby allocating more reasonable sub-networks to tokens and reducing the occurrence of collapse.

[0087] In some embodiments, a projection layer can be introduced into the generative model through a pre-layer regularization mechanism to enhance the gradient in a supervised manner without changing the core network structure, thereby alleviating the problem of the generative model weakening the shallow network component due to the rapid increase in the output norm of the deep network. The electronic device 110 can determine a second loss based on the target content and the reference content corresponding to the training sample, the second loss indicating the degree of difference between the target content and the reference content; and train the generative model based on the second loss.

[0088] for example, Figure 5 The process 500 of pre-layer regularization is shown. The electronic device 110 may give the hidden input 570 of the lth layer as h l The generation model includes a projection layer and a routing module, and the projection layer is located after the routing module. The routing module includes a two-layer MLP router 560, wherein the first layer applies a linear transformation and a GELU activation function while retaining the original hidden dimension; the second layer is divided into two branches, wherein the first branch is a gated head branch 530, which is used to generate the correlation between the first set of tokens and multiple sub-networks in the generation model; the second branch is a target head branch 550, including the projection layer H:R L′×d →R L×d , the reference content H(h l )540, and according to the target content y∈R corresponding to the generated model L×d 510 constructs a second loss function. This dual-head design allows routing decisions and target prediction to be performed simultaneously. The second loss function formula can be as follows:

[0089]

[0090] Where N is the total number of patches, n is the patch index, L and L′ represent the number of patches before and after the patch operation, and then the second loss function is applied to train the generated model. l )540 and the target content y∈R L× d The second loss between 510 can be calculated separately at each layer. By aligning the reference content output by the intermediate layer with the final target content, the contribution of the shallow network in the generative model to the final output result of the model can be effectively enhanced, and the optimization process of the network model can be significantly accelerated in the pre-normalized architecture, thereby improving the overall network performance.

[0091] In some embodiments, after executing the routing strategy according to the routing information, each token will be assigned a corresponding subnetwork combination, which may include one subnetwork or multiple subnetworks. For example, token T1 may be assigned subnetwork E1, and token T1 may be assigned subnetworks E2 and E3 at the same time, etc. The usage frequencies of these subnetwork combinations are different. Therefore, by analyzing the usage frequencies of the subnetwork combinations, it is possible to understand which subnetwork combinations are frequently used and which subnetwork combinations are not frequently used. Subnetwork combinations that are not frequently used may mean that the collaboration between these subnetworks is not effective enough. By adjusting the gating mechanism or the subnetwork structure to optimize the subnetwork combinations that are not frequently used, the purpose of optimizing the computing efficiency and resource allocation of the network model can be achieved.

[0092] In some embodiments, the electronic device 110 may determine a subnetwork combination allocated to at least one token based on routing information, the subnetwork combination including one or more subnetworks, and then determine a combined usage rate of each subnetwork combination; and optimize the generation model based on the combined usage rate.

[0093] In some embodiments, in order to estimate the number of sub-network combinations that are actually activated and participate in the calculation, the electronic device 110 can obtain the distribution between tokens and sub-networks in the first group of tokens through routing information, and count the number of times each sub-network combination corresponding to each token in the first group of tokens is selected, construct a corresponding histogram, and then sort the number of times these sub-network combinations are selected in descending order and normalize them to obtain a sorted normalized histogram, and then calculate the cumulative sum of the normalized histogram, and count the number of bins with a cumulative sum less than 95%. The ratio of the number of bins with a cumulative sum less than 95% to the total number of bins is the combination usage rate, that is, a sub-network combination with a selection number less than 5% of the total number is considered to be a sub-network combination whose combination usage rate does not meet the preset combination usage conditions, and then optimize the sub-network combination whose combination usage rate does not meet the preset combination usage conditions by adjusting the gating mechanism or the sub-network structure, thereby effectively improving the utilization rate of the sub-network combination.

[0094] In addition, the embodiment of the present application does not impose any specific restrictions on the preset combination use conditions, nor does it impose any restrictions on the numerical value of the cumulative sum. The numerical values ​​in the preset combination use conditions can be adjusted accordingly according to actual needs.

[0095] In some embodiments, the electronic device 110 can use the AdamW optimizer to set the batch size of the generation model in the training phase to 256, and use a constant learning rate of 1e-4, and do not apply weight decay. During the initialization process, all adaptive layers can be normalized using zero initialization, and all linear layers can be initialized using a uniformly distributed Xavier method. When training the generation model, all feedforward neural networks can be replaced with a hybrid expert network module, and the number of parameters activated can be ensured to be the same. In addition, a smaller initialization range can be set for each subnetwork, and the internal hidden dimension of each subnetwork is set to 1 / k of its dense corresponding model to ensure that its initialization range is the same as the initialization range of its dense corresponding model. It is also possible to use per-layer regularization, with its weight set to 1e-2, and use a first loss function with a weight of 1e-4. In addition, in the training phase of the network model, the exponential moving average can also be applied to update the network model weight to obtain a global threshold, and the global threshold is applied in the reasoning phase of the network model.

[0096] The embodiments of the present disclosure can provide a new routing strategy for the generative model by integrating the hybrid expert network module into the generative model. The routing strategy achieves higher routing flexibility by performing top-K selection in the entire routing space across batch, sequence and expert network dimensions, thereby providing the generative model with greater optimization freedom and significantly improving the performance of the network model. In order to meet the challenges brought about by the increased flexibility, the distribution changes caused by the time step in the training and reasoning stages of the network model can also be alleviated by introducing a global threshold based on the exponential moving average, thereby ensuring the consistency of network generation. In addition, a pre-layer regularization mechanism can be used to improve the stability of model training, apply the first loss to promote the combination of different sub-networks, and ensure better load balancing.

[0097] Example devices and equipment

[0098] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 6 Schematic structural block diagram of an apparatus 600 for content generation according to some embodiments of the present disclosure is shown. The apparatus 600 may be implemented as or included in the electronic device 110 as discussed above. Each module / component in the apparatus 600 may be implemented by hardware, software, firmware, or any combination thereof.

[0099] like Figure 6 As shown, the apparatus 600 includes a calling module 610, configured to process input content using a generative model to determine a first group of tokens corresponding to the input content; a first determining module 620, configured to determine correlation information between the first group of tokens and a plurality of subnetworks in the generative model, the correlation information indicating a plurality of degrees of association corresponding to a plurality of token-subnetwork pairs; a second determining module 630, configured to determine routing information based on the correlation information, the routing information indicating at least one token-subnetwork pair determined from a plurality of token-subnetwork pairs, wherein at least one token-subnetwork pair is determined based on a ranking of a plurality of degrees of association, or at least one token-subnetwork pair is determined based on a global threshold corresponding to the correlation information; a token processing module 640, configured to activate at least one subnetwork indicated by the routing information to process the corresponding tokens indicated by the routing information to generate a second group of tokens; and a content generating module 650, configured to generate target content corresponding to the input content based on the second group of tokens.

[0100] In some embodiments, the device 600 also includes a first token-subnetwork pair determination module 660, which is configured to determine at least one token-subnetwork pair based on the ranking of multiple association degrees during the training phase of the generated model; and / or a second token-subnetwork pair determination module 670, which is configured to determine at least one token-subnetwork pair based on a global threshold corresponding to the correlation information during the inference phase of the generated model, and the global threshold is determined based on the training phase.

[0101] In some embodiments, the second token-subnetwork pair determination module 670 is further configured to determine multiple training associations corresponding to multiple token-subnetwork pairs during the training phase; determine a reference association corresponding to a preset sorting position by sorting the multiple training associations; and determine a global threshold applied to the inference phase based on the reference association.

[0102] In some embodiments, the first determination module 620 is further configured to determine an association matrix between the first group of tokens and the plurality of sub-networks; and determine the correlation information based on the association matrix.

[0103] In some embodiments, the first determination module 620 is further configured to determine an association matrix between the first group of tokens and multiple sub-networks; and construct a single-dimensional association tensor based on multiple associations in the association matrix as correlation information.

[0104] In some embodiments, the apparatus 600 further includes a first loss module configured to determine a first loss based on the correlation information, the first loss indicating the similarity of the routing strategies of the tokens for the sub-networks; and train the generation model based on the first loss.

[0105] In some embodiments, the device 600 also includes a second loss module, which is configured to determine a second loss based on the target content and the reference content corresponding to the training sample, the second loss indicating the degree of difference between the target content and the reference content; and train the generation model based on the second loss.

[0106] In some embodiments, the device 600 also includes an optimization module, which is configured to determine a subnetwork combination to allocate at least one token based on the routing information, wherein the subnetwork combination includes one or more subnetworks; determine a combination usage rate of each subnetwork combination; and optimize the generation model based on the combination usage rate.

[0107] The units included in the device 600 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units in the device 800 can be implemented at least in part by one or more hardware logic components. As an example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.

[0108] Figure 7 1 shows a block diagram of an electronic device 700 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 7 The electronic device 700 shown is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. Figure 7 The electronic device 700 shown can be used to implement Figure 1 An electronic device 110 is shown.

[0109] like Figure 7 As shown, the electronic device 700 is in the form of a general electronic device. The components of the electronic device 700 may include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be an actual or virtual processor and is capable of performing various processes according to a program stored in the memory 720. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the electronic device 700.

[0110] The electronic device 700 typically includes a plurality of computer storage media. Such media may be any accessible media that is accessible to the electronic device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 may be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (e.g., a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 may be a removable or non-removable medium, and may include a machine-readable medium, such as a flash drive, a disk, or any other medium, which may be capable of being used to store information and / or data (e.g., training data for training) and may be accessed within the electronic device 700.

[0111] The electronic device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 7 As shown in , a disk drive for reading or writing from a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading or writing from a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules that are configured to perform various methods or actions of various embodiments of the present disclosure.

[0112] The communication unit 740 implements communication with other electronic devices through a communication medium. Additionally, the functions of the components of the electronic device 700 can be implemented with a single computing cluster or multiple computing machines that can communicate through a communication connection. Therefore, the electronic device 700 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.

[0113] The input device 750 may be one or more input devices, such as a mouse, a keyboard, a tracking ball, etc. The output device 760 may be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 700 may also communicate with one or more external devices (not shown) through the communication unit 740 as needed, such as a storage device, a display device, etc., communicate with one or more devices that allow a user to interact with the electronic device 700, or communicate with any device that allows the electronic device 700 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).

[0114] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0115] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices, equipment, and computer program products implemented according to the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0116] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0117] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0118] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple implementations of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction includes one or more executable instructions for realizing the logical function of the specification. In some implementations as replacements, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.

[0119] The above descriptions of various implementations of the present disclosure are exemplary, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The selection of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating content, comprising: Processing the input content using the generative model to determine a first set of tokens corresponding to the input content; Determining correlation information between the first group of tokens and a plurality of sub-networks in the generative model, the correlation information indicating a plurality of association degrees corresponding to a plurality of token-sub-network pairs; Determine routing information based on the correlation information, the routing information indicating at least one token-subnetwork pair determined from the plurality of token-subnetwork pairs, wherein the at least one token-subnetwork pair is determined based on the ranking of the plurality of association degrees, or the at least one token-subnetwork pair is determined based on a global threshold corresponding to the correlation information; activating at least one sub-network indicated by the routing information to process corresponding tokens indicated by the routing information to generate a second set of tokens; and Based on the second set of tokens, target content corresponding to the input content is generated.

2. The method according to claim 1, wherein: During the training phase of the generative model, the at least one token-subnetwork pair is determined based on the ranking of the plurality of association degrees; and / or In the inference phase of the generative model, the at least one token-subnetwork pair is determined based on the global threshold corresponding to the correlation information, and the global threshold is determined based on the training phase.

3. The method according to claim 2, wherein: The global threshold is determined based on the following process: In the training phase, determining a plurality of training association degrees corresponding to the plurality of token-subnetwork pairs; By sorting the multiple training association degrees, determining a reference association degree corresponding to a preset sorting position; as well as Based on the reference association degree, the global threshold applied to the reasoning stage is determined.

4. The method according to claim 1, wherein: The determining of correlation information between the first group of tokens and a plurality of sub-networks in the generative model comprises: determining an affinity matrix between the first set of tokens and the plurality of sub-networks; and Based on the association matrix, the correlation information is determined.

5. The method according to claim 4, wherein: The determining the correlation information based on the association matrix includes: Based on the multiple associations in the association matrix, a single-dimensional association tensor is constructed to serve as the correlation information.

6. The method according to claim 1, wherein: The input content is a training sample obtained in the training phase, and the method further includes: determining a first loss based on the correlation information, the first loss indicating how similar the routing strategies of the tokens for the sub-networks are; and The generative model is trained based on the first loss.

7. The method according to claim 1, wherein: The input content is a training sample obtained in the training phase, and the method further includes: determining a second loss based on the target content and reference content corresponding to the training sample, the second loss indicating a degree of difference between the target content and the reference content; and Based on the second loss, the generative model is trained.

8. The method according to claim 1, further comprising: Determine a subnetwork combination allocated to at least one token based on the routing information, wherein the subnetwork combination includes one or more of the subnetworks; determining a combined usage rate of each of the sub-network combinations; as well as Based on the combined usage, the generation model is optimized.

9. A device for content generation, comprising: A calling module configured to process the input content using the generative model to determine a first set of tokens corresponding to the input content; A first determination module is configured to determine correlation information between the first group of tokens and a plurality of sub-networks in the generative model, wherein the correlation information indicates a plurality of association degrees corresponding to a plurality of token-sub-network pairs; A second determination module is configured to determine routing information based on the correlation information, wherein the routing information indicates at least one token-subnetwork pair determined from the plurality of token-subnetwork pairs, wherein the at least one token-subnetwork pair is determined based on the ranking of the plurality of association degrees, or the at least one token-subnetwork pair is determined based on a global threshold corresponding to the correlation information; a token processing module configured to activate at least one sub-network indicated by the routing information to process corresponding tokens indicated by the routing information to generate a second set of tokens; and The content generation module is configured to generate target content corresponding to the input content based on the second group of tokens.

10. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processing unit.

11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.

12. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.