Image generation method and device and electronic equipment

By using multi-chip collaborative processing technology in electronic devices, the image generation process is divided into two parts: noise feature extraction and image reconstruction, which solves the problems of high resource occupation, high cost and low efficiency in the prior art, and realizes efficient and low-cost image generation services.

CN120014091APending Publication Date: 2025-05-16艾酷软件技术(上海)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510134529.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing image generation technology has huge amount of calculation model parameters, too much resource occupies, and takes a long time to calculate, resulting in high cost and low efficiency, especially when electronic devices cannot use cellular networks, which can't interact, affecting the image generation efficiency.

Method used

By using at least two chips in electronic devices, namely an embedded neural network processor (NPU) and a microcontroller unit (MCU) or arithmetic logic unit (ALU), the image generation process is divided into two parts: noise feature extraction and image reconstruction, which are performed on NPU and MCU/ALU respectively, to achieve full utilization of resources and reduce costs.

Benefits of technology

It realizes efficient image generation without the need for a cellular network, reduces the cost and resource occupancy of image generation, and improves the efficiency and quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014091A_ABST
    Figure CN120014091A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device and electronic equipment, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining user demand information through a first chip of the electronic equipment; determining a first noise feature in the first image based on the user demand information through a first chip of the electronic device, the first noise feature being a noise feature expected to appear in the first image; inputting the first noise feature into a second chip of the electronic equipment; generating, by a second chip, a second image based on the first noise feature and the first image; and determining a user demand image according to the second image through the first chip, the user demand image being an image including semantic features of the user demand information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an image generation method, device and electronic equipment. Background Art

[0002] Image generation technology has become one of the research fields of artificial intelligence (AI) and is widely used in image generation scenarios such as speech, image, and text. Since the computational models involved in image generation technology, such as diffusion models, have a huge number of related parameters, occupy too many resources, and take a long time to calculate, the computational models are usually deployed on servers to provide corresponding services to users.

[0003] In the related art, electronic devices can obtain user needs and send them to the server through the cellular network. The server generates service information related to the user needs based on the user needs and the calculation model, and feeds it back to the electronic device. However, users need to pay the corresponding fees to open the cellular network, which increases the cost of image generation. When the electronic device cannot use the cellular network, it cannot interact with the server, so that the server cannot obtain user needs, affecting the efficiency of image generation. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide an image generation method, device, electronic device and storage medium, which can reduce the image generation cost and improve the image generation efficiency.

[0005] In a first aspect, an embodiment of the present application provides an image generation method, which is applied to an electronic device, wherein the electronic device includes a first chip and a second chip, including:

[0006] Obtain user demand information;

[0007] Determining, by the first chip, a first noise feature in the first image based on user demand information, where the first noise feature is a noise feature that is expected to appear in the first image;

[0008] inputting the first noise characteristic into the second chip;

[0009] Generate a second image based on the first noise feature and the first image using a second chip;

[0010] A user demand image is determined according to the second image through the first chip, where the user demand image is an image including semantic features of user demand information.

[0011] In a second aspect, an embodiment of the present application provides an image generating device, which is applied to an electronic device, wherein the electronic device includes a first chip and a second chip, and the image generating device includes:

[0012] Acquisition module, used to obtain user demand information;

[0013] A determination module, configured to determine, through the first chip and based on user requirement information, a first noise feature in the first image, where the first noise feature is a noise feature that is expected to appear in the first image;

[0014] An input module, used for inputting the first noise characteristic into the second chip;

[0015] A generating module, configured to generate a second image based on the first noise feature and the first image through a second chip;

[0016] The determination module is also used to determine, through the first chip, a user demand image according to the second image, where the user demand image is an image including semantic features of user demand information.

[0017] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the image generation method shown in the first aspect.

[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the image generation method shown in the first aspect are implemented.

[0019] In a fifth aspect, an embodiment of the present application provides a chip, the chip including a processor and a display interface, the display interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the image generation method shown in the first aspect.

[0020] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the image generation method shown in the first aspect.

[0021] In the embodiment of the present application, user demand information can be obtained, and the first noise feature in the first image can be determined based on the user demand information through the first chip, and the first noise feature is the noise feature expected to appear in the first image; the first noise feature is input into the second chip; and the second image is generated based on the first noise feature and the first image through the second chip; the user demand image is determined according to the second image through the first chip, and the user demand image is an image including the semantic features of the user demand information. In this way, the data interaction between the chips of the electronic device can be fully utilized, and the relevant data of image generation can be deployed in the electronic device, so that the user can provide the user with the service of generating images without opening the cellular network, which reduces the cost of image generation, and when the electronic device cannot use the cellular network, the user demand information can be obtained at any time to generate the user demand image, which improves the image generation efficiency. And, by dividing the image generation process into two parts, namely the part of extracting noise features and the part of reconstructing the user demand image, two different chips are jointly called to execute separately, namely the first chip executes the part of extracting noise features and the second chip executes the part of reconstructing the user demand image, which can avoid the problem of high chip resource occupancy and poor image quality caused by the image generation process being executed on the same chip, so that the electronic device can provide users with high-quality, high-efficiency and low-cost image generation services. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 A structural diagram of an electronic device provided for some embodiments of the present application;

[0023] Figure 2 A flowchart of an image generation method provided for some embodiments of the present application;

[0024] Figure 3 A schematic diagram of an interface of an image generation method provided in some embodiments of the present application;

[0025] Figure 4 A schematic diagram of an algorithm for generating a second image by an MCU of an image generating method provided in some embodiments of the present application;

[0026] Figure 5 A schematic diagram of an algorithm for generating a second image by an MCU of an image generating method provided in some embodiments of the present application;

[0027] Figure 6 A schematic diagram of MCU storing data in an image generation method provided in some embodiments of the present application;

[0028] Figure 7 A schematic diagram of a sampling algorithm type of an image generation method provided for some embodiments of the present application;

[0029] Figure 8A schematic diagram of the structure of an image generating device provided in some embodiments of the present application;

[0030] Fig. 9 A schematic diagram of the structure of an electronic device provided for some embodiments of the present application;

[0031] Fig.10 A schematic diagram of the hardware structure of an electronic device provided for some embodiments of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0033] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0034] In order to solve the problems arising in the above-mentioned related technologies, the embodiments of the present application provide a method, an apparatus and an electronic device.

[0035] The following is combined with Figures 1 to 7 , the image generation method provided in the embodiment of the present application is described in detail through specific embodiments and their application scenarios.

[0036] Figure 1 A schematic diagram of the structure of an electronic device provided for some embodiments of the present application.

[0037] The electronic device provided in the embodiment of the present application may include at least two chips, at least two chips may include an embedded neural network processor (Neural Processing Unit, NPU), and the second chip is a microcontroller unit (Microcontroller Unit, MCU). Alternatively, at least two chips may include an NPU and an arithmetic logic unit (Arithmetic Logic Unit, ALU). Alternatively, at least two chips may include an NPU, an MCU, and an ALU.

[0038] Here, the image generation method provided in the embodiment of the present application can use NPU and MCU and / or ALU to collaboratively implement the algorithm path of the overall diffusion model, which can improve the computing efficiency, reuse the chip resources of the electronic device, and provide users with a good user experience. Among them, the scheduling capability of MCU is better than that of NPU and ALU, NPU and ALU have more computing resources, and NPU and ALU have better computing performance than MCU. The diffusion model and the sampler are connected in series in the overall algorithm path. If all are deployed in NPU, the NPU computing power is under high pressure and runs at high load for a long time, then additional chip area needs to be added to support the fast calculation of the sampler, which will increase the cost. If all MCU or ALU are deployed, the computing resources are weak and cannot reach the time-consuming performance index of the diffusion model alone. Therefore, the embodiment of the present application will deploy the diffusion model in NPU, and the sampler in MCU and ALU, and NPU, MCU and ALU communicate and coordinate scheduling with each other to achieve full utilization of resources and reduce additional cost investment.

[0039] Based on this, the following Figure 1 A detailed description of the NPU, MCU and / or ALU is provided.

[0040] Specifically, the diffusion model of the NPU in the embodiment of the present application is a forward diffusion process, that is, noise is gradually added to the first image according to the user demand information to make it a pure noise feature of a Gaussian distribution, and its noise feature is input into the MCU. Among them, the diffusion model in the embodiment of the present application can use methods such as model compression, mixed precision calculation, and hierarchical design, including but not limited to Stable Diffusion Lite, MobileDiffusion, SnapFusion and other models. Since the diffusion model calculates too many parameters, it needs to be deployed on the NPU unit to ensure operating efficiency.

[0041] The sampler running on the MCU and ALU is a reverse diffusion process, that is, a completely random first image is gradually denoised according to the noise characteristics output by the NPU until a high-quality image that meets the distribution of the training data, that is, a user-required image, is generated. Specifically, the MCU starts with pure noise features, gradually removes the noise, obtains the second image, and finally generates the user-required image through the NPU and the variational autoencoder (VAE). Among them, different sampling methods can affect the quality, speed and diversity of image generation. The MCU generates images through a process called denoising, which involves gradually extracting meaningful image features from random noise in the latent space. Among them, the denoising sampling algorithm of the sampler in the embodiment of the present application includes but is not limited to: Markov chain sampling algorithm (DDPM), non-Markov chain sampling algorithm (DDIM), stochastic differential equation sampling algorithm (Stochastic Differential Equations, SDE), ordinary differential equation sampling algorithm (Ordinary Differential Equation, ODE), sampling algorithm (flow matching), Euler, Ancestral Sampling, DPMSolver, unipc and other sampling algorithms. The detailed calculation details inside different samplers are different, but the ultimate goal is to iteratively generate the user's required image from the noise.

[0042] Based on this, the steps of the image generation method in the embodiment of the present application executed by the aforementioned at least two chips can be as follows.

[0043] First, the electronic device obtains user demand information, which may include first prompt information, wherein the first prompt information is information expressed in an image generated by a user instruction, such as a cat lying in a doghouse. The first prompt information may be input into a multimodal pre-training model (Contrastive Language-Image Pre-training, CLIP), which is trained through a large number of text-image pairs to understand images that semantically match the first prompt information, and then the prompt features output by the CLIP model are input into the NPU.

[0044] The NPU obtains the hint feature and generates a completely random noise image, namely the first image, in the latent space. Based on the hint feature, the NPU determines the first noise feature of the first image, where the first noise feature is the noise feature expected to appear in the first image. The output first noise feature is passed to the MCU as the input of the sampler.

[0045] The MCU may generate a second image based on the first noise feature and the first image, feed the second image back to the NPU, and perform an iterative process using the second image as the first image.

[0046] Among them, the iterative process is that the NPU determines the first noise feature of the first image based on the second image output by the MCU, that is, the first image of the current iteration and the prompt feature, and passes the output first noise feature to the MCU as the input of the sampler. The MCU can generate a second image based on the first noise feature and the second noise feature of the first image, and input the second image generated in the current iteration into the NPU as the first image of the next iteration. These second images gradually change from random noise to clearer and clearer images. When the number of iterations of the iterative process is greater than the reference number of iterations, the second image output by the MCU is input into the VAE. Through the VAE, the second image obtained by the iterative process, that is, the second image corresponding to the number of iterations of the iterative process greater than the reference number of iterations, is subjected to image dimensionality reduction processing to obtain a clean, denoised user demand image, which reflects the content described in the user demand information. It should be noted that the reference number of iterations in the embodiment of the present application refers to the upper limit of the number of iterations pre-set during the iterative process. In the embodiment of the present application, the reference number of iterations can be determined based on the second image output by the second chip or based on artificial experience. The reference number of iterations in the embodiment of the present application can be dynamically updated. For example, when the second image output by the second chip can be used to restore the semantic features including the user demand information, the iteration can be stopped in advance before the reference number of iterations. Thus, in the entire image generation process, the GPU runs a diffusion model to predict noise, and the MCU runs a sampler to restore the image content from the noise picture. The two run alternately to gradually restore high-quality image content from the original noise picture. In addition, the overall image generation process is implemented on the electronic device, and the user can enjoy the image generation application service without connecting to the Internet, reducing the user's cost of use. Using the MCU and NPU single-module collaborative processing method, the respective advantages of the two are fully utilized under the existing resources, the NPU resource occupancy rate of the electronic device chip is reduced, and the MCU and NPU jointly complete the image generation algorithm process path. In addition, the sampling algorithm is deployed on the MCU of the electronic device chip to reduce the NPU resource occupancy rate of the electronic device chip, accelerate the deployment of image generation technology on the electronic device, and thus improve the efficiency of image generation.

[0047] It should be noted that the ALU in the embodiment of the present application can replace the MCU to generate the second image based on the first noise feature and the first image. Also, the ALU in the embodiment of the present application can also be combined with the MCU to generate the second image.

[0048] Figure 2A flowchart of an image generation method provided for some embodiments of the present application.

[0049] like Figure 2 As shown, the image generation method provided in the embodiment of the present application can be applied to an electronic device, and the electronic device includes a first chip and a second chip. Based on this, the image generation method can include steps 210 to 250 as shown below.

[0050] Step 210, obtaining user demand information; Step 220, determining, through the first chip, a first noise feature in the first image based on the user demand information, the first noise feature being a noise feature expected to appear in the first image; Step 230, inputting the first noise feature into the second chip; Step 240, generating, through the second chip, a second image based on the first noise feature and the first image; Step 250, determining, through the first chip, a user demand image based on the second image, wherein the user demand image is an image including semantic features of the user demand information.

[0051] It should be noted that the first chip in the embodiment of the present application is an NPU, and the second chip includes at least one of the following: an MCU, an ALU.

[0052] Exemplarily, user demand information is obtained, and the user demand information may include first prompt information, wherein the first prompt information is information expressed in the image generated by the user's instruction, such as a cat lying in a doghouse. The prompt feature and a completely random noise image, i.e., the first image, generated in the latent space are input into the NPU, and the first noise feature in the first image is determined by the NPU based on the user demand information, wherein the first noise feature is a noise feature expected to appear in the first image, and the output first noise feature is input into the MCU. Then, the second image is generated by the MCU based on the first noise feature and the second noise feature of the first image, and the second noise feature is the noise feature in the first image. The second image is fed back to the NPU, and the second image is used as the first image to perform an iterative process. The iterative process includes the NPU determining the first noise feature of the first image based on the second image output by the MCU, i.e., the first image of the current iteration, and the prompt feature, and passing the output first noise feature to the MCU as the input of the sampler, and the MCU can generate the second image based on the first noise feature and the second noise feature of the first image, and the second image generated in the current iteration is input into the NPU as the first image of the next iteration. These second images gradually change from random noise to increasingly clear images. When the number of iterations of the iterative process is greater than the reference number of iterations 20, the second image output by the MCU is determined as a user requirement image, which reflects the content described in the user requirement information.

[0053] Thus, the data interaction between the chips of the electronic device can be fully utilized, and the relevant data of image generation can be deployed in the electronic device, so that the user can be provided with the service of generating images without opening a cellular network, thereby reducing the cost of image generation, and when the electronic device cannot use the cellular network, the user's demand information can be obtained at any time to generate the user's demand image, thereby improving the efficiency of image generation. In addition, by dividing the image generation process into two parts, namely the part of extracting noise features and the part of reconstructing the user's demand image, two different chips are jointly called to execute separately, namely the first chip executes the part of extracting noise features and the second chip executes the part of reconstructing the user's demand image, thereby avoiding the problem of high chip resource occupancy and poor image quality caused by the image generation process being executed on the same chip, and thus providing users with high-quality, high-efficiency, and low-cost image generation services through electronic devices.

[0054] It should be noted that the image generation method provided in the embodiment of the present application can be applied to related fields such as electronic device terminal photography, video recording, video calling, live broadcasting, etc., and can provide users with high-quality, efficient, and low-cost content generation services. The above steps are described in detail below, as shown below.

[0055] First, step 210 is involved. In some embodiments of the present application, the user requirement information includes first prompt information and one of the following information: second prompt information, a reference image, the number of iterations of the iterative process, a random seed, and an image size of the second image;

[0056] The first prompt information is information expressed in the image generated by the user's instruction; the second prompt word is information not expressed in the second image generated by the user's instruction.

[0057] For example, Figure 3 As shown, the first prompt information may be a cat lying in a doghouse. The second prompt information is that no dog can appear. The number of iterations of the iterative process is, for example, 20 times. Based on this, when the number of iterations of the iterative process is greater than the reference number of iterations 20, the second image output by the MCU at the 20th iteration is determined to be the user required image. The random seed is used to limit the style of the image subject in each second image output during the iterative process to remain consistent. For example, the cat breeds are all ragdolls, and the image style is all cartoon. The image size of the second image can be set to 5 inches.

[0058] Based on this, the user demand information in the embodiment of the present application can be displayed as an option on the user demand interface, for example, you can still refer to Figure 3 , Figure 3The user requirement interface shown in the figure may include a first display area and a second display area. The first display area is used to display the first prompt information and the second prompt information input by the user, and the second display area is used to display the options of the number of iterations of the iteration process, the options of the random seed, and the options of the image size of the second image, so that the user can configure the user requirement information. After the user configures the user requirement information, the user can click the "Image Generation" control to generate the user requirement image.

[0059] Next, step 220 is involved. In some embodiments of the present application, the first chip includes a deep learning model (U-Net) for image segmentation. Based on this, step 220 may specifically include:

[0060] Determining, by a deep learning model, a first noise feature in the first image based on user demand information;

[0061] Before executing the iteration process, the first image is a random noise image, and during the iteration process, the first image is the second image output by the second chip in the last iteration process.

[0062] Here, in the embodiment of the present application, when processing complex and high-resolution images, U-Net can better restore the detailed information of the image through its unique skip connection and upsampling design, thereby providing a more accurate segmentation result and improving the accuracy of extracting the first noise feature. Furthermore, in some embodiments of the present application, the second chip includes a denoising sampling algorithm, based on which, step 240 can specifically include:

[0063] Step 2401, generating a second image based on the first noise feature and the second noise feature of the first image through a denoising sampling algorithm.

[0064] Here, if Figure 4 As shown, the embodiment of the present application can be explained by taking DDIM as an example, because the DDIM sampling method is based on denoising iteration to eliminate noise and restore the image. DDIM relies on multiple state information for solving each state based on denoising. In the DDIM sampler description, the feature solution of each state can be solved by difference between two state information, and there is no need to rely on migration state information for solution. Therefore, the calculation time is greatly reduced in the solution process, and the calculation efficiency is improved. Furthermore, the above step 2401 can specifically include steps 24011 to 24014, as shown below.

[0065] Step 24011, through the denoising sampling algorithm, discretize N time nodes from the time axis in the reverse diffusion process.

[0066] Step 24012: Determine the image obtained by removing the first noise feature from the second noise feature of the first image as the noise image of the ith time node among the N time nodes.

[0067] Step 24013, sampling the noise image of the i-th time node and the estimated noise image of the j-th time node, to obtain the noise image of the i-1-th time node among the N time nodes; wherein the j-th time node is at least two time nodes before the i-th time point among the N time nodes. wherein the noise features of the noise image of the i-1-th time node are less than the noise features of the noise images of the first i time nodes.

[0068] Step 24014, determine the noise image of the time node before the i-th time node as the second image.

[0069] For example, DDIM is used as a sampler for description and analysis, where the DDIM sampler algorithm is to iteratively restore the original high-quality image from the original Gaussian noise. Specifically, the core content of the DDIM sampler reverse reasoning process is the Bayesian conditional probability as shown in formula (1). At the same time, in (x t ,x0) under the condition x t-1 The probability of occurrence follows the normal distribution as shown in formula (2). From this, we can deduce that about (x t ,x0) is shown in formula (3).

[0070]

[0071] in,

[0072]

[0073] The reverse reasoning process of DDIM sampler is as follows Figure 5 As shown, take x0, x3 as an example to explain. When calculating x1, DDIM is not calculated based on x2, but is calculated based on x0, x3 for interpolation. n ) state, solve any state x m process.

[0074] Here, it should be noted that the random number α and the random number β can be calculated and provided by the MCU. When the ALU combines with the MCU to generate the second image, the random number α and the random number β in the DDIM sampler involved in the above formula (3) can also be provided by the ALU.

[0075] In some embodiments of the present application, the second chip includes at least two storage spaces, and the at least two storage spaces include a first storage space and a second storage space. Based on this, before the above step 24013, the image generation method may further include step 24015:

[0076] The noise image of the i-th time node is cached in the first storage space, and the image noise feature is cached in the second storage space, where the image noise feature includes at least one of the following: a first noise feature and a second noise feature.

[0077] Based on this, step 24013 may specifically include step 240131 and step 240132.

[0078] Step 240131, call at least one noise image of a time node from the first storage space and call at least one image noise feature from the second storage space.

[0079] Step 240132, based on the noise image of at least one time node and at least one image noise feature, sample the noise image of the i-th time node and the estimated noise image of the j-th time node to obtain the noise image of the i-1-th time node among the N time nodes.

[0080] Exemplarily, according to the calculation principle description of the above formula (1) to formula (3), the algorithm flow of the diffusion sampling algorithm is explained by taking DDIM as an example in the embodiment of the present application.

[0081] Step 1: Discretize the inference time nodes from the time axis of the back-diffusion process , and assume that the input noise image x n (The feature after VAE encoding) is the time node t n Noise image (features after VAE encoding). Step 2: Initialize i=N, the first storage space and the second storage space are empty. Step 3: Set the time node at t i The noise image and the protected noise are cached in the first and second storage spaces respectively. Step 4: Based on the first storage space and the second storage space, the non-Markov chain sampling algorithm of formula (3) is used to perform the transfer time pair (t i ,t i-1 ) to obtain the image sampling at time node t i-1 The noise image and the noise contained in it. Step 5: When i≠0, let i=i-1 and return to step 3; when i=0, take the noise image at time node t0 as the restored image of the input noise image, that is, the second image.

[0082] Therefore, the characteristic solution of each state in the DDIM sampler description can be solved by the difference of two state information, without having to rely on the migration state information for solution. Therefore, the calculation time is greatly reduced in the solution process and the calculation efficiency is improved.

[0083] In some embodiments of the present application, at least two storage spaces in the embodiments of the present application may include static random access memory (SRAM), data lifecycle management (DLM) and information lifecycle management (ILM). Based on this, after the above step 24015, the image generation method may further include steps 24016 and 24017.

[0084] Step 24016, based on the frequency with which the noise image of at least one time node and at least one image noise feature are called from the storage space, store the variable data having a call frequency greater than a preset call threshold into the first storage space, wherein the variable data includes at least one of the following: a noise image of at least one time node and at least one image noise feature.

[0085] Step 24017, storing variable data whose call frequency is less than or equal to the preset call threshold into the second storage space.

[0086] Exemplarily, still taking the above example as an example, the 5 steps of the sampler are run on the MCU side. Since the feature size of the diffusion model output is 4*128*18, and the data type is floating point 32 type, one feature needs to occupy 256K. In the overall sampler algorithm, 2 input features are required, including image features (256K) and expected noise features (256K); 2 intermediate temporary storage variables (256K*2=512K). If all these 4 main features are placed in the DLM memory, the storage cost will increase. Therefore, the embodiment of the present application will count the frequency of use of the variables used in the sampler, and store the variables with high utilization in the DLM storage space according to the frequency of use, and the variables with lower frequency of use in the SRAM storage space. The statistical summary is as follows Figure 6As shown. Among them, latents is the image input feature, g_pred_epsilon and g_pred_original_sample are temporary variables, and the three variables are 768K in total, which exceeds the DLM256K limit. Therefore, the three variables are converted to float16 for storage while ensuring accuracy, occupying a total of 384K, and linebufer storage is used. The input features are divided into 2 segments, each occupying 192K. The DLM memory stores the segmented linebufer data, reducing the DLM storage space, and the remaining variables are stored in the SRAM memory.

[0087] Therefore, in the embodiments of the present application, the performance of each chip of the electronic device can be fully utilized, the sampler algorithm is deployed on the MCU, the diffusion model is deployed on the NPU side, and the data segments are stored in different memories by category, so as to make full use of the chip modules of the electronic device, jointly call, and locally deploy the model, provide efficient hardware support for image generation technology, and reduce user usage costs. In chip design, data is classified and stored in different memories (SRAM, DLM, ILM) according to the frequency of use, which reduces storage space and application costs while ensuring the real-time performance of the algorithm. In some embodiments of the present application, the second chip includes a first sampling algorithm and a second sampling algorithm. Based on this, step 24013 can specifically include step 240133 and step 240134.

[0088] Step 240133, through the first sampling algorithm, sample the single feature pixel in the noise image of the i-th time node and the expected noise image of the j-th time node respectively, and obtain the noise image of the i-1-th time node among the N time nodes.

[0089] Step 240134, through the second sampling algorithm, sample at least two characteristic pixels in the noise image of the i-th time node and the estimated noise image of the j-th time node respectively to obtain the noise image of the i-1-th time node among the N time nodes.

[0090] For example, in order to improve the efficiency of MCU in generating the second image, the visual processing unit (VPU) or digital signal processor (DSP) may be used in the embodiment of the present application. The VPU operation rate is M times that of the DSP operation rate, but the disadvantage is that the area is increased by N times. In practical applications, the VPU or DSP implementation method is determined based on the current terminal cost and chip area limitations. The algorithms running on the main modules in the MCU-NPU are as follows: Figure 7 shown.

[0091] Then, step 250 is involved. In some embodiments of the present application, the second image may be generated in an iterative manner to improve the accuracy of determining the image required by the user. Based on this, step 250 may specifically include:

[0092] An iterative process is performed through the first chip using the second image as the first image. When an iterative condition is met, a user required image is determined based on the second image obtained through the iterative process. The iterative process includes: inputting the first noise feature into the second chip; generating the second image based on the first noise feature and the first image through the second chip; and the iterative condition includes that the number of iterations of the iterative process is greater than the reference number of iterations.

[0093] Here, the process of determining the first noise feature in the first image based on the user demand information by the first chip can refer to the above step 220, which will not be repeated here. And the step of inputting the first noise feature into the second chip can refer to the above step 230, which will not be repeated here. And the process of generating the second image based on the first noise feature and the first image by the second chip can refer to the above step 240, which will not be repeated here.

[0094] It should be noted that the first noise feature involved in the first iteration involved in the iterative process can be determined by the following steps: determining the first noise feature in the first image based on the user requirement information through the first chip. The first noise feature involved in each iteration from the second iteration to the kth iteration corresponding to the number of iterations involved in the iterative process is the noise feature in the second image obtained based on the previous iteration.

[0095] In some embodiments of the present application, the first chip includes a variational autoencoder. Based on this, the step of determining the user required image according to the second image obtained by the iterative process involved in step 250 may specifically include:

[0096] Through the variational autoencoder, the second image obtained in the iterative process is subjected to image dimension reduction processing to obtain the user required image.

[0097] Therefore, the high-dimensional data can be compressed into a low-dimensional hidden representation through the variational autoencoder, and then restored to the original high-dimensional data through the decoder. This method not only reduces the dimension of the data, but also retains important feature information, so that the data after dimensionality reduction can still reconstruct the original data well, improving the accuracy of determining the image required by the user.

[0098] In some embodiments of the present application, before step 250, the image generation method may further include step 260, in which, when the number of iterations of executing the iterative process is greater than the reference number of iterations, it is determined that the iteration condition is satisfied.

[0099] Therefore, the embodiments of the present application can deploy a diffusion model in an electronic device as the ultimate goal, and improve the overall content generation efficiency through the connectivity of various chip modules under the premise of ensuring the sampling quality of the diffusion model, so that the chip modules can be calculated in a coordinated manner, reduce memory usage costs, improve reading efficiency, and reduce the cost of user image generation.

[0100] It should be noted that the image generation method provided in the embodiments of the present application can be executed by electronic devices such as mobile phones, tablet computers, laptop computers, PDAs, wearable devices, etc. In some embodiments of the present application, the image generation method provided in the embodiments of the present application is described by taking the electronic device as the execution subject to execute the image generation method as an example.

[0101] The image generation method provided in the embodiment of the present application can be executed by an image generation device. In the embodiment of the present application, an image generation device executing the image generation method is taken as an example to illustrate the device of the image generation method provided in the embodiment of the present application.

[0102] The present application also provides an image generation device. Figure 8 Provide detailed explanation.

[0103] Figure 8 A schematic diagram of the structure of an image generating device provided in some embodiments of the present application.

[0104] like Figure 8 As shown, the image generating device 80 can be applied to an electronic device, the electronic device includes a first chip and a second chip, and the image generating device 80 can specifically include:

[0105] Acquisition module 801, used to acquire user demand information;

[0106] A determination module 802 is configured to determine, through a first chip, a first noise feature in a first image based on user requirement information, where the first noise feature is a noise feature that is expected to appear in the first image;

[0107] An input module 803, used for inputting the first noise characteristic into the second chip;

[0108] A generating module 804 is used to generate a second image based on the first noise feature and the first image by using a second chip;

[0109] The determination module 802 is further configured to determine, through the first chip, a user requirement image according to the second image, where the user requirement image is an image including semantic features of user requirement information.

[0110] The image generating device 80 in the embodiment of the present application is described in detail below, as shown below.

[0111] In some embodiments of the present application, the first chip is an embedded neural network processor, and the second chip includes at least one of the following: a microcontroller unit and an arithmetic logic unit.

[0112] In some embodiments of the present application, the determination module 802 may be specifically configured to, when the first chip includes a deep learning model for image segmentation, determine a first noise feature in the first image based on user demand information by using the deep learning model;

[0113] Before executing the iteration process, the first image is a random noise image, and during the iteration process, the first image is the second image output by the second chip in the last iteration process.

[0114] In some embodiments of the present application, the determination module 802 may be specifically configured to, through the first chip, perform an iterative process using the second image as the first image, and when an iterative condition is met, determine the user required image according to the second image obtained through the iterative process;

[0115] The iterative process includes: determining a first noise feature in a first image based on user requirement information through a first chip; inputting the first noise feature into a second chip; generating a second image based on the first noise feature and the first image through the second chip;

[0116] The iteration condition includes that the number of iterations of executing the iteration process is greater than the reference number of iterations.

[0117] In some embodiments of the present application, the generation module 804 may be specifically configured to generate a second image based on the first noise feature and the second noise feature of the first image by using the denoising sampling algorithm when the second chip includes the denoising sampling algorithm.

[0118] In some embodiments of the present application, the generating module 804 may be specifically used to discretize N time nodes from the time axis in the reverse diffusion process through a denoising sampling algorithm;

[0119] Determine an image obtained by removing the first noise feature from the second noise feature of the first image as a noise image at the i-th time node among the N time nodes;

[0120] Sampling the noise image at the i-th time node and the estimated noise image at the j-th time node to obtain the noise image at the i-1-th time node among the N time nodes; wherein the j-th time node is at least two time nodes before the i-th time point among the N time nodes;

[0121] The noise image at a time node before the i-th time node is determined as the second image.

[0122] In some embodiments of the present application, the image generating device 80 may further include a storage module, a calling module and a sampling module; wherein,

[0123] A storage module, configured to cache the noise image of the i-th time node into the first storage space and cache the image noise feature into the second storage space when the second chip includes at least two storage spaces, and the at least two storage spaces include a first storage space and a second storage space, wherein the image noise feature includes at least one of the following: a first noise feature and a second noise feature;

[0124] A calling module, used to call at least one noise image of a time node from the first storage space and call at least one image noise feature from the second storage space;

[0125] The sampling module is used to sample the noise image of the i-th time node and the estimated noise image of the j-th time node based on the noise image of at least one time node and at least one image noise feature, so as to obtain the noise image of the i-1-th time node among the N time nodes.

[0126] In some embodiments of the present application, the image generating device 80 may further include a storage module, which is used to store variable data having a call frequency greater than a preset call threshold into the first storage space according to the call frequency of the noise image at at least one time node and at least one image noise feature from the storage space, wherein the variable data includes at least one of the following: a noise image at at least one time node, at least one image noise feature;

[0127] The variable data whose calling frequency is less than or equal to the preset calling threshold is stored in the second storage space.

[0128] In some embodiments of the present application, the image generating device 80 may further include a sampling module, which is used to sample the single characteristic pixel in the noise image at the i-th time node and the estimated noise image at the j-th time node respectively by using the first sampling algorithm when the second chip includes the first sampling algorithm and the second sampling algorithm, so as to obtain the noise image at the i-1-th time node among the N time nodes;

[0129] By using the second sampling algorithm, at least two characteristic pixels in the noise image at the i-th time node and the estimated noise image at the j-th time node are sampled respectively to obtain the noise image at the i-1-th time node among the N time nodes.

[0130] In some embodiments of the present application, the determination module 802 can be specifically used to, when the first chip includes a variational autoencoder, perform image dimensionality reduction processing on the second image obtained in the iterative process through the variational autoencoder to obtain the user required image.

[0131] In some embodiments of the present application, the user demand information includes the first prompt information and one of the following information: the second prompt information, the reference image, the number of iterations of the iterative process, the random seed, and the image size of the second image;

[0132] The first prompt information is information expressed in the image generated by the user's instruction; the second prompt word is information not expressed in the second image generated by the user's instruction.

[0133] The image generating device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0134] The image generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0135] The image generation device provided in the embodiment of the present application can achieve Figures 1 to 7 The various processes implemented in the illustrated embodiment of the image generation method achieve the same technical effect, and will not be described again here to avoid repetition.

[0136] Based on this, the image generation device provided in the embodiment of the present application can obtain user demand information, and determine the first noise feature in the first image based on the user demand information through the first chip, the first noise feature is the noise feature expected to appear in the first image; input the first noise feature into the second chip; and, through the second chip, generate the second image based on the first noise feature and the first image; through the first chip, determine the user demand image according to the second image, the user demand image is an image including the semantic features of the user demand information. In this way, the data interaction between the chips of the electronic device can be fully utilized, and the relevant data of image generation can be deployed in the electronic device, so that the user can provide the user with the service of generating images without opening a cellular network, reducing the cost of image generation, and when the electronic device cannot use the cellular network, the user demand information can be obtained at any time to generate the user demand image, thereby improving the efficiency of image generation. Furthermore, by dividing the image generation process into two parts, namely the part of extracting noise features and the part of reconstructing the image required by the user, two different chips are jointly called to execute them separately, namely the first chip executes the part of extracting noise features and the second chip executes the part of reconstructing the image required by the user. This can avoid the problem of high chip resource occupancy and poor image quality caused by the image generation process being executed on the same chip. In this way, high-quality, high-efficiency and low-cost image generation services can be provided to users through electronic devices.

[0137] Optional, such as Fig. 9 As shown, an embodiment of the present application further provides an electronic device 90, including a processor 901 and a memory 902, wherein the memory 902 stores programs or instructions that can be executed on the processor 901, and when the program or instructions are executed by the processor 901, the various steps of the above-mentioned image generation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.

[0138] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0139] Fig.10 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.

[0140] The electronic device 1000 includes but is not limited to: a radio frequency unit 1001, a network module 1002, an audio output unit 1003, an input unit 1004, a sensor 1005, a display unit 1006, a user input unit 1007, an interface unit 1008, a memory 1009, a processor 1010 and other components.

[0141] Those skilled in the art will appreciate that the electronic device 1000 may also include a power source (such as a battery) for supplying power to each component, and the power source may be logically connected to the processor 1010 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system.

[0142] Fig.10 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.

[0143] Among them, in the embodiment of the present application, the processor 1010 is used to obtain user demand information. The processor 1010 is also used to determine, through the first chip, a first noise feature in the first image based on the user demand information, the first noise feature being a noise feature expected to appear in the first image. The processor 1010 is also used to input the first noise feature into the second chip. The processor 1010 is also used to generate a second image based on the first noise feature and the first image through the second chip. The processor 1010 is also used to determine, through the first chip, a user demand image based on the second image, the user demand image being an image including semantic features of the user demand information.

[0144] The electronic device 1000 is described in detail below, as shown below.

[0145] In some embodiments of the present application, the first chip is an embedded neural network processor, and the second chip includes at least one of the following: a microcontroller unit and an arithmetic logic unit.

[0146] In some embodiments of the present application, the processor 1010 may be specifically configured to, when the first chip includes a deep learning model for image segmentation, determine a first noise feature in the first image based on user demand information through the deep learning model;

[0147] Before executing the iteration process, the first image is a random noise image, and during the iteration process, the first image is the second image output by the second chip in the last iteration process.

[0148] In some embodiments of the present application, the processor 1010 may be specifically configured to, through the first chip, perform an iterative process using the second image as the first image, and when an iterative condition is met, determine a user required image according to the second image obtained through the iterative process;

[0149] The iterative process includes: determining a first noise feature in a first image based on user requirement information through a first chip; inputting the first noise feature into a second chip; generating a second image based on the first noise feature and the first image through the second chip;

[0150] The iteration condition includes that the number of iterations of executing the iteration process is greater than the reference number of iterations.

[0151] In some embodiments of the present application, the processor 1010 may be specifically configured to generate a second image based on the first noise feature and the second noise feature of the first image by using the denoising sampling algorithm when the second chip includes the denoising sampling algorithm.

[0152] In some embodiments of the present application, the processor 1010 may be specifically configured to discretize N time nodes from a time axis in a back diffusion process by using a denoising sampling algorithm;

[0153] Determine an image obtained by removing the first noise feature from the second noise feature of the first image as a noise image at the i-th time node among the N time nodes;

[0154] Sampling the noise image at the i-th time node and the estimated noise image at the j-th time node to obtain the noise image at the i-1-th time node among the N time nodes; wherein the j-th time node is at least two time nodes before the i-th time point among the N time nodes;

[0155] The noise image at a time node before the i-th time node is determined as the second image.

[0156] In some embodiments of the present application, the processor 1010 is configured to cache the noise image of the i-th time node into the first storage space and cache the image noise feature into the second storage space when the second chip includes at least two storage spaces, and when the at least two storage spaces include a first storage space and a second storage space, wherein the image noise feature includes at least one of the following: a first noise feature and a second noise feature;

[0157] Recalling at least one noise image of a time node from the first storage space and recalling at least one image noise feature from the second storage space;

[0158] Based on the noise image of at least one time node and at least one image noise feature, the noise image of the i-th time node and the estimated noise image of the j-th time node are sampled to obtain the noise image of the i-1-th time node among the N time nodes.

[0159] In some embodiments of the present application, the processor 1010 is configured to store variable data having a call frequency greater than a preset call threshold in a first storage space according to a call frequency of the noise image at at least one time node and at least one image noise feature from the storage space, wherein the variable data includes at least one of the following: a noise image at at least one time node and at least one image noise feature;

[0160] The variable data whose calling frequency is less than or equal to the preset calling threshold is stored in the second storage space.

[0161] In some embodiments of the present application, the processor 1010 is configured to, when the second chip includes the first sampling algorithm and the second sampling algorithm, sample the single feature pixel in the noise image at the i-th time node and the estimated noise image at the j-th time node respectively by using the first sampling algorithm to obtain the noise image at the i-1-th time node among the N time nodes;

[0162] By using the second sampling algorithm, at least two characteristic pixels in the noise image at the i-th time node and the estimated noise image at the j-th time node are sampled respectively to obtain the noise image at the i-1-th time node among the N time nodes.

[0163] In some embodiments of the present application, the processor 1010 can be specifically used to, when the first chip includes a variational autoencoder, perform image dimensionality reduction processing on the second image obtained in the iterative process through the variational autoencoder to obtain the user required image.

[0164] In some embodiments of the present application, the user demand information includes the first prompt information and one of the following information: the second prompt information, the reference image, the number of iterations of the iterative process, the random seed, and the image size of the second image;

[0165] The first prompt information is information expressed in the image generated by the user's instruction; the second prompt word is information not expressed in the second image generated by the user's instruction.

[0166] It should be understood that the input unit 1004 may include a graphics processor (Graphics Processing Unit, GPU) 10041 and a microphone 10042, and the graphics processor 10041 processes the image data of the static image or video obtained by the image capture device (such as a camera) in the video capture mode or the image capture mode. The display unit 1006 may include a display panel, and the display panel may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch display. Other input devices 10072 may include, but are not limited to, a physical keyboard, function keys (such as a volume display button, a switch button, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.

[0167] The memory 1009 can be used to store software programs and various data. The memory 1009 can mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area can store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 1009 can include a volatile memory or a non-volatile memory, or the memory 1009 can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 1009 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.

[0168] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to the operating system, user interface, and application programs, and the modem processor mainly processes wireless display signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 1010.

[0169] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned image generation method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0170] The processor is the processor in the electronic device in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0171] In addition, an embodiment of the present application further provides a chip, which includes a processor and a display interface, the display interface and the processor are coupled, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned image generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0172] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0173] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned image generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0174] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0175] In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in a reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.

[0176] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of each embodiment of the present application.

[0177] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. An image generation method, characterized in that: Applied to an electronic device, the electronic device includes a first chip and a second chip, and the image generation method includes: Obtain information about user needs; determining, by the first chip, a first noise feature in the first image based on the user demand information, where the first noise feature is a noise feature that is expected to appear in the first image; inputting the first noise characteristic into the second chip; generating a second image based on the first noise feature and the first image by the second chip; A user demand image is determined according to the second image through the first chip, where the user demand image is an image including semantic features of the user demand information.

2. The method according to claim 1, characterized in that The first chip is an embedded neural network processor, and the second chip includes at least one of the following: a micro control unit and an arithmetic logic unit.

3. The method according to claim 1 or 2, characterized in that: The first chip includes a deep learning model for image segmentation; and determining, by the first chip, a first noise feature in the first image based on the user demand information, includes: Determining, by the deep learning model, a first noise feature in the first image based on the user demand information; Before executing the iterative process, the first image is a random noise image, and during the iterative process, the first image is a second image output by the second chip when the iterative process was last executed.

4. The method according to claim 1, characterized in that: The determining, by the first chip and according to the second image, the user required image includes: By using the first chip, performing an iterative process using the second image as the first image, and when an iterative condition is met, determining the user demand image according to the second image obtained by the iterative process; The iterative process includes: inputting a first noise feature in the first image into the second chip; generating the second image based on the first noise feature and the first image through the second chip; The iteration condition includes that the number of iterations of executing the iteration process is greater than a reference number of iterations.

5. The method according to claim 1 or 4, characterized in that: The second chip includes a denoising sampling algorithm; and generating a second image based on the first noise feature and the first image by the second chip includes: A second image is generated based on the first noise feature and the second noise feature of the first image through the denoising sampling algorithm.

6. The method according to claim 5, characterized in that The step of generating a second image based on the first noise feature and the second noise feature of the first image by using the denoising sampling algorithm includes: N time nodes are discretized from the time axis in the reverse diffusion process by the denoising sampling algorithm; Determine an image obtained by removing the first noise feature from the second noise feature of the first image as a noise image at an i-th time node among the N time nodes; Sampling the noise image of the i-th time node and the estimated noise image of the j-th time node to obtain the noise image of the i-1-th time node among the N time nodes; wherein the j-th time node is at least two time nodes before the i-th time point among the N time nodes; The noise image at a time node before the i-th time node is determined as the second image.

7. The method according to claim 6, characterized in that The second chip includes at least two storage spaces, and the at least two storage spaces include a first storage space and a second storage space; Before sampling the noise image at the i-th time node and the estimated noise image at the j-th time node to obtain the noise image at the i-1-th time node among the N time nodes, the method further includes: Cache the noise image of the i-th time node into the first storage space, and cache the image noise feature into the second storage space, wherein the image noise feature includes at least one of the following: the first noise feature and the second noise feature; The sampling of the noise image at the i-th time node and the estimated noise image at the j-th time node to obtain the noise image at the i-1-th time node among the N time nodes includes: Recalling at least one noise image of a time node from the first storage space and recalling at least one image noise feature from the second storage space; Based on the noise image of the at least one time node and the at least one image noise feature, the noise image of the i-th time node and the estimated noise image of the j-th time node are sampled to obtain the noise image of the i-1-th time node among the N time nodes.

8. The method according to claim 7, characterized in that After caching the noise image of the i-th time node into the first storage space and caching the image noise feature into the second storage space, the method further includes: According to the frequency at which the noise image at the at least one time node and the at least one image noise feature are called from the storage space, storing variable data having a calling frequency greater than a preset calling threshold into the first storage space, the variable data comprising at least one of the following: the noise image at the at least one time node and the at least one image noise feature; The variable data whose calling frequency is less than or equal to the preset calling threshold is stored in the second storage space.

9. The method according to claim 6, characterized in that The second chip includes a first sampling algorithm and a second sampling algorithm; the sampling of the noise image at the i-th time node and the estimated noise image at the j-th time node to obtain the noise image at the i-1-th time node among the N time nodes includes: By using the first sampling algorithm, sampling the single characteristic pixel in the noise image at the i-th time node and the estimated noise image at the j-th time node respectively, to obtain the noise image at the i-1-th time node among the N time nodes; By using the second sampling algorithm, at least two characteristic pixels in the noise image of the i-th time node and the estimated noise image of the j-th time node are sampled respectively to obtain the noise image of the i-1-th time node among the N time nodes.

10. The method according to claim 4, characterized in that The first chip includes a variational autoencoder; and determining the user demand image according to the second image obtained by the iterative process includes: The variational autoencoder is used to perform image dimension reduction processing on the second image obtained in the iterative process to obtain the user demand image.

11. The method according to claim 1, characterized in that: The user requirement information includes first prompt information and one of the following information: second prompt information, a reference image, the number of iterations of the iterative process, a random seed, and an image size of the second image; The first prompt information is information expressed in the image generated by the user's instruction; the second prompt word is information not expressed in the second image generated by the user's instruction.

12. An image generating device, characterized in that: Applied to an electronic device, the electronic device comprises a first chip and a second chip, and the image generating device comprises: Acquisition module, used to obtain user demand information; a determination module, configured to determine, through the first chip and based on the user requirement information, a first noise feature in the first image, where the first noise feature is a noise feature that is expected to appear in the first image; An input module, used for inputting the first noise characteristic into the second chip; A generating module, configured to generate a second image based on the first noise feature and the first image by using the second chip; The determination module is further used to determine, through the first chip and according to the second image, a user requirement image, where the user requirement image is an image including semantic features of the user requirement information.

13. An electronic device, characterized in that: include: A processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the image generation method according to any one of claims 1 to 11.

14. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the image generation method according to any one of claims 1 to 11 are implemented.

15. A computer program product, characterized in that The program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the image generating method according to any one of claims 1 to 11.