Image generation model training method and device, image generation method and device, equipment and medium
By encoding the position rotation matrix and adjusting the loss value of the panoramic image samples, the problem that existing algorithms cannot generate high-quality panoramic images is solved, and efficient generation of panoramic images is achieved.
Patent Information
- Application Number
- CN202510294423.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-04
AI Technical Summary
The existing image generation algorithm cannot be effectively applied to 360-degree panoramic images, and cannot represent the texture distribution of the panoramic images well, resulting in poor generation effect.
By obtaining the training panoramic image samples and description text, the input vector is positionally encoded using the preset position rotation matrix formula, the output vector is calculated, and the parameters of the image generation model are adjusted based on the loss value, and the position rotation matrix is obtained to determine the distance information of each pixel point in the panoramic image.
The generation effect of panoramic images is improved, the distance information of each pixel point in the panoramic image is accurately determined, and the quality of image generation is improved.
Smart Images

Figure CN120259461A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a training of an image generation model, an image generation method, an apparatus, a device, and a medium. Background Art
[0002] Currently, image generation algorithms have developed rapidly and have been widely applied in multiple fields.
[0003] All existing image generation algorithms can only be applied to the generation of perspective views and cannot be applied to 360-degree panoramic images. Specifically, the texture distribution of panoramic images is relatively different from that of perspective views. The position encoding method for panoramic images in existing image generation algorithms cannot well represent the texture distribution of panoramic images, resulting in relatively poor panoramic image generation effects. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a training of an image generation model, an image generation method, an apparatus, a device, and a medium.
[0005] An embodiment of the present disclosure provides a method for training an image generation model. The method includes:
[0006] Obtaining a training data pair; wherein the training data pair includes a training panoramic image sample and a description text corresponding to the training panoramic image sample;
[0007] Inputting the training panoramic image sample and the description text into a pre-constructed image generation model to be trained, and performing position encoding on an input vector corresponding to the training panoramic image sample through a pre-set position rotation matrix formula to obtain a position rotation matrix corresponding to the input vector, and calculating an output vector based on the position rotation matrix corresponding to the input vector;
[0008] Adjusting model parameters of the image generation model to be trained based on a loss value between a training vector of the training panoramic image sample and the output vector to obtain an image generation model.
[0009] An embodiment of the present disclosure further provides an image generation method. The method includes:
[0010] Obtaining an image generation description text;
[0011] Inputting the image generation description text into an image generation model, and performing position encoding on a target vector corresponding to a preset noise vector through a pre-set position rotation matrix formula to obtain a position rotation matrix corresponding to the target vector, and calculating a generated vector based on the position rotation matrix corresponding to the target vector, and decoding the generated vector to obtain a target panoramic image;
[0012] Among them, the image generation model is obtained according to the training method of the image generation model described in any one of the foregoing embodiments.
[0013] The embodiments of the present disclosure also provide a training device for an image generation model, and the device includes:
[0014] A first acquisition module, configured to acquire a training data pair; wherein, the training data pair includes a training panoramic image sample and a description text corresponding to the training panoramic image sample;
[0015] An input module, configured to input the training panoramic image sample and the description text into a pre-constructed image generation model to be trained, so as to perform position encoding on an input vector corresponding to the training panoramic image sample through a pre-set position rotation matrix formula, obtain a position rotation matrix corresponding to the input vector, and calculate an output vector based on the position rotation matrix corresponding to the input vector;
[0016] A training module, configured to adjust model parameters of the image generation model to be trained based on a loss value between a training vector of the training panoramic image sample and the output vector, and obtain an image generation model.
[0017] The embodiments of the present disclosure also provide an image generation device, and the device includes:
[0018] A second acquisition module, configured to acquire an image generation description text;
[0019] A generation module, configured to input the image generation description text into an image generation model, so as to perform position encoding on a target vector corresponding to a preset noise vector through a pre-set position rotation matrix formula, obtain a position rotation matrix corresponding to the target vector, calculate a generated vector based on the position rotation matrix corresponding to the target vector, and decode the generated vector to obtain a target panoramic image;
[0020] Among them, the image generation model is obtained according to the training method of the image generation model described in any one of the foregoing embodiments.
[0021] The embodiments of the present disclosure also provide an electronic device, and the electronic device includes: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the training of the image generation model and the image generation method provided in the embodiments of the present disclosure.
[0022] The embodiments of the present disclosure also provide a computer-readable storage medium, and the storage medium stores a computer program, and the computer program is used to execute the training of the image generation model and the image generation method provided in the embodiments of the present disclosure.
[0023] An embodiment of the present disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, is for training an image generation model and an image generation method provided by the embodiment of the present disclosure.
[0024] The technical solution provided by the embodiment of the present disclosure has the following advantages compared with the prior art: In the training solution of the image generation model provided by the embodiment of the present disclosure, a pair of training data is obtained; wherein, the pair of training data includes a training panoramic image sample and a description text corresponding to the training panoramic image sample; the training panoramic image sample and the description text are input into a pre-constructed image generation model to be trained, so as to perform position encoding on the input vector corresponding to the training panoramic image sample through a pre-set position rotation matrix formula to obtain a position rotation matrix corresponding to the input vector, and calculate an output vector based on the position rotation matrix corresponding to the input vector; the model parameters of the image generation model to be trained are adjusted based on the loss value between the training vector of the training panoramic image sample and the output vector to obtain an image generation model. By adopting the above technical solution, the pixel positions of the panoramic image are represented by the position rotation matrix to train the image generation model, so that in the process of generating a panoramic image, the position rotation matrix is obtained based on the image generation model to determine the distance information of each pixel in the generated panoramic image, thereby improving the generation effect of the panoramic image. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.
[0026] Figure 1 is a schematic flow chart of a method for training an image generation model provided by an embodiment of the present disclosure;
[0027] Figure 2A is a schematic diagram of a panoramic spherical coordinate system provided by an embodiment of the present disclosure;
[0028] Figure 2B is a panoramic unfolded coordinate system provided by an embodiment of the present disclosure;
[0029] Figure 3 is a schematic diagram of training an image generation model provided by an embodiment of the present disclosure;
[0030] Figure 4 is a schematic diagram of a Self-attention module provided by an embodiment of the present disclosure;
[0031] Figure 5 is a schematic flow chart of another method for training an image generation model provided by an embodiment of the present disclosure;
[0032] Figure 6 A schematic diagram of image generation provided by an embodiment of the present disclosure;
[0033] Figure 7 A flowchart of an image generation method provided by an embodiment of the present disclosure;
[0034] Figure 8 Another schematic diagram of image generation provided by an embodiment of the present disclosure;
[0035] Figure 9 A schematic structural diagram of a training device for an image generation model provided by an embodiment of the present disclosure;
[0036] Figure 10 A schematic structural diagram of an image generation device provided by an embodiment of the present disclosure;
[0037] Figure 11 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0038] The embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0039] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0040] The term "including" and its variations used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0041] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0042] It should be noted that the modifiers "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless clearly specified otherwise in the context, it should be understood as "one or more".
[0043] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0044] Generally, the texture distribution of panoramic images is relatively different from that of perspective views. Specifically, compared with perspective views, the texture structure of panoramic images has the following characteristics: the aspect ratio of the image is 2:1, the horizontal field of view is 360 degrees, and the vertical field of view is 180 degrees; due to the 360-degree continuous field of view in the horizontal direction, the left and right sides of the panoramic image are continuous; due to the 180-degree field of view in the horizontal direction, the upper and lower edges of the image are not continuous; the image represents that the imaging environment is projected on a unit sphere, so the horizontal structure will be distorted into an arc; however, the existing position encoding methods for panoramic images in image generation algorithms cannot well represent the texture distribution of panoramic images, resulting in poor panoramic image generation effects.
[0045] To address the above problems, this disclosure proposes a training scheme for an image generation model, which obtains training data pairs; where the training data pairs include training panoramic image samples and the description texts corresponding to the training panoramic image samples; inputs the training panoramic image samples and the description texts into a pre-constructed image generation model to be trained, and performs position encoding on the input vector corresponding to the training panoramic image sample through a pre-set position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector, and calculates the output vector based on the position rotation matrix corresponding to the input vector; adjusts the model parameters of the image generation model to be trained based on the loss value between the training vector and the output vector to obtain the image generation model. By adopting the above technical solution, the pixel positions of panoramic images are represented by the position rotation matrix to train the image generation model, so that during the panoramic image generation process, the position rotation matrix is obtained based on the image generation model to determine the distance information of each pixel in the generated panoramic image, improving the generation effect of panoramic images.
[0046] Figure 1 FIG. is a schematic flow chart of a method for training an image generation model provided by an embodiment of this disclosure. This method can be executed by a training device for an image generation model, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 1 shown, this method includes:
[0047] Step 101, obtain training data pairs; where the training data pairs include training panoramic image samples and the description texts corresponding to the training panoramic image samples.
[0048] Among them, the training panoramic image sample can be any panoramic image. The embodiments of the present disclosure do not limit the training panoramic image sample. For example, the training panoramic image sample can be a panoramic image stitched from multiple rooms of a show flat of a certain real estate project, or a panoramic image stitched from the rooms of a house for sale.
[0049] Among them, the description text refers to the text information used to introduce the training panoramic image sample. The image content recognition result of the training panoramic image sample can be obtained through image content recognition of the training panoramic image sample and used as the description text. For example, when recognizing a panoramic image sample stitched from a room, the image content recognition result is "a living room including a sofa and a TV", and "a living room including a sofa and a TV" is used as the description text of the panoramic image sample stitched from the room; or the text information input by the user for the training panoramic image sample can also be used as the description text, etc.
[0050] Specifically, during the training of the image generation model, multiple training data pairs can be obtained. Each training data pair includes a training panoramic image sample and the description text corresponding to the training panoramic image sample.
[0051] Step 102: Input the training panoramic image sample and the description text into a pre-constructed image generation model to be trained, so as to perform position encoding on the input vector corresponding to the training panoramic image sample through a pre-set position rotation matrix formula, obtain the position rotation matrix corresponding to the input vector, and calculate the output vector based on the position rotation matrix corresponding to the input vector.
[0052] Among them, the image generation model to be trained is pre-constructed, such as Diffusion Transformer (DiT, a deep learning architecture based on the diffusion model) or SDXL (Stable Diffusion XL, an open-source text-to-image framework), a network model architecture composed of multiple VAEs (Variational Autoencoders). It is specifically selected and set according to the actual application scenario.
[0053] Specifically, after inputting the training panoramic image sample and the description text into the pre-constructed image generation model to be trained, first, the training panoramic image sample is encoded based on VAE to obtain a training vector. The training vector refers to converting the training panoramic image sample into a numerical form that can be processed, that is, extracting features from the training panoramic image sample and encoding them into a vector representation, and then inputting the training vector and the description text into DiT or SDXL for processing. Among them, DiT or SDXL includes a Self-attention module. The vector input into the Self-attention module can be understood as a vector obtained by performing network encoding and other processing on the training vector, and the vector processed by the Self-attention module can better learn the distance relationship between each pixel.
[0054] It can be understood that before inputting into the Self-attention module, the training panoramic image sample undergoes a series of network encodings to obtain an input vector proportional to the scale of the training panoramic image sample (the original image). The input vector is position-encoded through a pre-set position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector, and the output vector is calculated based on the position rotation matrix corresponding to the input vector.
[0055] Among them, the input vector refers to the vector input into the Self-attention module for processing, which is an input vector proportional to the scale of the training panoramic image sample. For example, the size of the input vector X of the Self-attention module is (c, h, w), where c represents the number of channels, and h and w represent the vector dimensions. This vector dimension is in a proportional relationship with the size of the original image (the training panoramic image sample); the position rotation matrix formula refers to the calculation formula for calculating the position rotation matrix corresponding to the input vector; the position rotation matrix corresponding to the input vector refers to the rotation matrix of each sub-vector of the input vector. Continuing with the above example, the position rotation matrix corresponding to the input vector, for example, the rotation matrix of the sub-vector at the (h i , w i ) of the input vector (c, h, w).
[0056] Among them, the position rotation matrix formula is pre-set. Specifically, each position rotation matrix parameter in the position rotation matrix formula is determined, as well as the calculation formula for the parameter value corresponding to each position rotation matrix parameter.
[0057] In the embodiments of the present disclosure, after obtaining the input vector, position encoding is performed on the input vector through the position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector. As an example, the input vector corresponding to the training panoramic image sample is obtained, each sub-vector corresponding to the input vector is obtained, the pixel position information corresponding to each sub-vector is encoded according to the position rotation matrix formula, and multiple sub-position rotation matrices are obtained and combined to obtain the position rotation matrix corresponding to the input vector; as another example, for each sub-vector in the input vector, the parameter values corresponding to the respective position rotation matrix parameters in the position rotation matrix formula are determined based on the pixel position coordinates in the target coordinate system, and multiple sub-position rotation matrices are obtained and combined to obtain the position rotation matrix corresponding to the input vector. The above two methods are only examples of performing position encoding on the input vector through the position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector, and the present disclosure does not limit the specific implementation manner of performing position encoding on the input vector through the position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector.
[0058] Therefore, after performing position encoding on the input vector, it is input into the Self-attention module for processing. For example, the input vector passes through a linear layer to obtain a first vector, a second vector, and a third vector respectively. The first vector and the second vector perform relevant calculations, and the relevant calculation results are applied to the third vector and then a new vector is output. Thus, when the first vector and the second vector perform relevant calculations, the position rotation matrix corresponding to the input vector can be used to better learn the distance between the two vectors, thereby learning the pixel distances between the respective pixel points in the training panoramic image sample, thereby improving the training effect of the image generation model and the generation effect of the subsequent panoramic image generated based on the image generation model.
[0059] It can be understood that the vector output by the Self-attention module is further processed in the DiT or SDXL network architecture and then the vector is output.
[0060] In the embodiments of the present disclosure, after obtaining the training panoramic image sample and the description text corresponding to the training panoramic image sample, the training panoramic image sample and the description text can be input into the pre-constructed image generation model to be trained, so as to perform position encoding on the input vector corresponding to the training panoramic image sample through the pre-set position rotation matrix formula, obtain the position rotation matrix corresponding to the input vector, and calculate the output vector based on the position rotation matrix corresponding to the input vector.
[0061] Step 103: Adjust the model parameters of the image generation model to be trained based on the loss value between the training vector of the training panoramic image sample and the output vector, so as to obtain the image generation model.
[0062] Among them, the training vector refers to converting the training panoramic image sample into a numerical form that can be processed, that is, extracting features from the training panoramic image sample and encoding them into a vector representation.
[0063] In the embodiments of the present disclosure, to obtain the training vector corresponding to the training panoramic image sample, specifically, the training panoramic image sample is encoded to obtain the training vector; more specifically, the training panoramic image sample is encoded by a VAE (Variational Auto-Encoder) to obtain the training vector.
[0064] In the embodiments of the present disclosure, the similarity between the training vector and the output vector can be calculated, and the loss value is calculated for the similarity through a preset loss function (such as the cross-entropy loss function, etc.). When the loss value is greater than or equal to the preset loss threshold, the model parameters of the image generation model to be trained are adjusted. Then, the loss value is calculated again for the similarity between the training vector and the generated vector in the new training data pair and compared with the loss threshold until the loss value is less than the loss threshold, and the image generation model is obtained.
[0065] Specifically, after obtaining the output vector, the model parameters of the image generation model to be trained are adjusted based on the loss value between the training vector of the training panoramic image sample and the output vector, and the image generation model is obtained.
[0066] The training scheme of the image generation model provided by the embodiments of the present disclosure includes obtaining a training data pair; among them, the training data pair includes a training panoramic image sample and the description text corresponding to the training panoramic image sample; the training panoramic image sample and the description text are input into a pre-constructed image generation model to be trained, and the input vector corresponding to the training panoramic image sample is position-encoded through a preset position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector, and the output vector is calculated based on the position rotation matrix corresponding to the input vector; the model parameters of the image generation model to be trained are adjusted based on the loss value between the training vector of the training panoramic image sample and the output vector, and the image generation model is obtained. By adopting the above technical solution, the pixel positions of the panoramic image are represented by the position rotation matrix to train the image generation model, so that in the process of generating the panoramic image, the position rotation matrix is obtained based on the image generation model to determine the distance information of each pixel point in the generated panoramic image, and the generation effect of the panoramic image is improved.
[0067] In some embodiments, the input vector corresponding to the training panoramic image sample is position-encoded by a preset position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector, including: obtaining the input vector corresponding to the training panoramic image sample; obtaining each sub-vector corresponding to the input vector, encoding the pixel position information corresponding to each sub-vector according to the preset position rotation matrix formula, and obtaining a plurality of sub-position rotation matrices for combination to obtain the position rotation matrix corresponding to the input vector.
[0068] In the embodiments of the present disclosure, after encoding the training panoramic image sample, the training vector is input into DiT and SDXL. After a series of network encoding processes in DiT and SDXL, the input vector is obtained. After position-encoding the input vector, it is input into the Self-attention module to learn the pixel position relationship between each pixel point in the training panoramic image sample and then a series of network encoding outputs the vector.
[0069] Specifically, each sub-vector corresponding to the input vector is obtained, and the pixel position information corresponding to each sub-vector is encoded according to the preset position rotation matrix formula, and a plurality of sub-position rotation matrices are obtained for combination to obtain the position rotation matrix corresponding to the input vector.
[0070] In the embodiments of the present disclosure, there are many ways to encode the pixel position information corresponding to each sub-vector according to the position rotation matrix formula and obtain a plurality of sub-position rotation matrices for combination to obtain the position rotation matrix corresponding to the input vector. As an example, the parameters of each position rotation matrix are determined based on the position rotation matrix formula, and the parameter values corresponding to the parameters of each position rotation matrix are determined based on the pixel position information corresponding to each sub-vector. The parameter values corresponding to the parameters of each position rotation matrix are input into the position rotation matrix formula to obtain the sub-position rotation matrix corresponding to each sub-vector for combination to obtain the position rotation matrix corresponding to the input vector.
[0071] As another example, the target pixel position coordinates of each sub-vector in the target coordinate system are determined based on the pixel position information corresponding to each sub-vector; the first position rotation matrix parameter value is determined based on the preset attitude parameters, the target pixel position coordinates of each sub-vector, and the preset first matrix parameter calculation formula; the second position rotation matrix parameter value is determined based on the target pixel position coordinates of each sub-vector and the preset second matrix parameter calculation formula; the first position rotation matrix parameter value and the second position rotation matrix parameter value are input into the position rotation matrix formula to obtain the sub-position rotation matrix of each sub-vector, and a plurality of sub-position rotation matrices are combined to obtain the position rotation matrix corresponding to the input vector.
[0072] In the embodiments of the present disclosure, the target coordinate system refers to a polar coordinate system. It can be understood that the image pixel coordinates in the training panoramic image sample are usually in the Cartesian coordinate system, that is, the initial coordinate system is the Cartesian coordinate system. Therefore, coordinate system conversion processing is required.
[0073] Specifically, based on the pixel position information corresponding to each sub-vector, obtain the initial pixel position coordinates of each sub-vector in the initial coordinate system; convert the initial pixel position coordinates of each sub-vector according to the preset coordinate system conversion formula to obtain the target pixel position coordinates of each sub-vector in the target coordinate system.
[0074] Specifically, the coordinate system of the training panoramic image sample is as Figure 2A shown, where each pixel corresponds to a point on the unit sphere, and this point can be expressed as (x, y, z) in the Cartesian coordinate system, where x, y, z ∈ [-1, 1], and can also be expressed in the polar coordinate system as where θ ∈ [-π / 2, π / 2],
[0075] Specifically, the transformation relationship between the two coordinate systems is shown in formulas (1) and (2).
[0076]
[0077]
[0078] Specifically, the pose of an object in space can be represented by rotation (R) and translation (t), that is, where RR T = I, det(R) = 1,
[0079] At the same time, the rotation matrix can be calculated using the preset pose parameters (yaw, pitch, roll). Let pitch = θ, roll = 0 (the definition of other orders is the same), then the rotation matrix of the position of any pixel in the training panoramic image sample can be represented as shown in formula (3).
[0080]
[0081] In addition, since the horizontal direction of the panoramic image is continuous left and right, and the vertical direction is not continuous up and down, the absolute position information in the θ direction is introduced in the displacement as shown in formula (4).
[0082] t = [θ / (π / 2), θ / (π / 2), θ / (π / 2)] T (4)
[0083] Therefore, the preset rotation position encoding formula is as shown in Formula (5).
[0084]
[0085] Among them, the first position rotation matrix parameter R comes from Formula (3), and the second position rotation matrix parameter t comes from Formula (4).
[0086] Therefore, the first position rotation matrix parameter value and the second position rotation matrix parameter value can be determined based on the pixel position information corresponding to each sub-vector, so as to obtain the sub-position rotation matrix of each sub-vector. Finally, multiple sub-position rotation matrices are combined to obtain the position rotation matrix corresponding to the input vector.
[0087] It can be understood that when obtaining the position rotation matrix corresponding to the input vector and performing subsequent correlation calculations through the position rotation matrix corresponding to the input vector, the pixel distance between each pixel in the panoramic image sample can be learned and trained. Among them, since all pixel points of the panoramic image are distributed on a unit sphere, a complete representation of the panoramic image pixel space can be realized (that is, each point has and only has one rotation matrix corresponding to it). Therefore, based on the characteristics of the rotation matrix, the distance between two pixel points can be represented by the included angle between the two viewing directions, and the included angle calculation formula is as shown in Formula (6).
[0088]
[0089] In the above solution, the position rotation matrix of each sub-vector corresponding to the input vector can be obtained, so that the pixel distance between each pixel point in the training panoramic image sample can be obtained in subsequent processing, so that the correlation between each pixel point in the training panoramic image sample can be learned and trained quickly and effectively. Therefore, in the process of generating a panoramic image by the trained image generation model, based on the position rotation matrix, the distance information of each pixel point in the generated panoramic image can be determined, improving the generation effect of the panoramic image.
[0090] Based on the description of the above embodiments, the pre-constructed image generation model to be trained can be trained with training data, and during the training process, the input vector corresponding to the training panoramic image sample is position-encoded through the preset position rotation matrix formula to obtain an image generation model to improve the generation effect of the generated panoramic image.
[0091] Specifically, as Figure 3 shown, the pre-constructed image generation model to be trained is the Diffusion Transformer (DiT) framework. During the training process, first, the input image is encoded by the VAE to obtain the image hidden vector latent, and then the image hidden vector latent and the description text such as Figure 3The input "A living room with tv and sofa" is subjected to the noise addition and diffusion process by DiT, and finally the output hidden vector is obtained. By decoding through the decoder of the VAE, an image similar to the original image can be obtained. However, the above are all designed for perspective views and cannot learn the texture distortion characteristics of panoramic images, resulting in a relatively poor generation effect of panoramic images.
[0092] Specifically, the Self-attention module is widely used in image and language models and is also one of the key modules of the Diffusion network. It is applied in text-to-image frameworks (such as DiT, SDXL, etc.). The Self-attention module is as Figure 4 shown. The input vector X passes through a linear layer to obtain three vectors {Q, K, V} respectively. Q and K are subjected to a correlation calculation with each other, such as Figure 4 the matmul operation (a matrix multiplication operation), scale operation (amplification and reduction), and softmax operation of the normalization exponential function shown, and the relevant calculation results are applied to V, such as Figure 4 the matmul operation shown to obtain the output This Self-attention module realizes the self-correlation calculation between input sequences. The formula of the Self-attention module is as shown in formula (7).
[0093]
[0094] In the embodiments of the present disclosure, it is proposed to represent the pixel positions of panoramic images through a position rotation matrix to train an image generation model. Thus, during the generation process of panoramic images, based on the position rotation matrix, the distance information of each pixel in the generated panoramic image can be determined, improving the generation effect of panoramic images.
[0095] Specifically, Rotary Position Embedding (RoPE) can be applied to image generation models, such as for position encoding of one-dimensional sequences, with characteristics such as strong scalability to sequence length and faster convergence, that is, through the method of absolute position encoding, the effect of relative position encoding is realized.
[0096] As shown in formula (8), the correlation relationship between Q and K is obtained through an inner product calculation. In RoPE, a certain function f is designed to improve this inner product calculation, so that the final result equivalently realizes relative position encoding.
[0097] <f(Q m ,m),f(K n ,n)>=g(Q m ,Kn , m - n) (8)
[0098] Among them, Q and K correspond to Q and K in the self - attention module, m and n represent positions. Assuming the lengths of Q and K are L, then m, n ∈ [0, L - 1]. That is, by designing a certain function f and then through the same inner - product calculation, it equivalently realizes g(Q m , K n , m - n), and this function is related to the relative position (m - n).
[0099] In the related RoPE encoding method, f is designed as shown in formula (9).
[0100]
[0101] Among them, α represents a preset fixed parameter. By splitting Q and K into two channels and applying a 2x2 rotation matrix, the effect of f encoding is achieved; through derivation, the inner - product calculation process of formula (9) can be converted into that shown in formula (10).
[0102] (R m Q m ) T (R n K n ) = Q m T R m T R n K n = Q m T R m-n K n (10)
[0103] That is, g(Q m , K n , m - n) = Q m T R m-n K n .
[0104] When RoPE is used in image tasks, the horizontal and vertical dimensions of the image are decoupled respectively, and encoded in a one - dimensional manner in two directions. This method can be applied to general perspective images, but for the horizontal and vertical dimensions of panoramic images, they are not continuous in the real space, and it cannot be well expressed in this way.
[0105] During the training of the image generation model according to the embodiments of the present disclosure, three-dimensional rotation position encoding can effectively represent the texture structure of panoramic images. In the form of absolute position encoding, it realizes the effects of relative position encoding and absolute position encoding at the same time, and can be conveniently applied to generation models (such as DiT, SDXL, etc.) to realize the generation of panoramic images. Specifically, it is combined with Figure 5 for detailed description.
[0106] Figure 5 FIG. is a schematic flowchart of another method for training an image generation model provided by the embodiments of the present disclosure. On the basis of the above embodiments, the method for training the above image generation model is further optimized. As Figure 5 shown, the method includes:
[0107] Step 201, obtain a training data pair; wherein, the training data pair includes a training panoramic image sample and a description text corresponding to the training panoramic image sample.
[0108] It should be noted that step 201 is the same as step 101. For specific reference, see the detailed description of step 101, and details are not described here again.
[0109] Step 202, input the training panoramic image sample and the description text into a pre-constructed image generation model to be trained to obtain an input vector corresponding to the training panoramic image sample, and obtain each sub-vector corresponding to the input vector.
[0110] Step 203, encode the pixel position information corresponding to each sub-vector according to the position rotation matrix formula, obtain a plurality of sub-position rotation matrices, combine them to obtain the position rotation matrix corresponding to the input vector, and calculate the output vector based on the position rotation matrix corresponding to the input vector.
[0111] In the embodiments of the present disclosure, encoding the pixel position information corresponding to each sub-vector according to the position rotation matrix formula, obtaining a plurality of sub-position rotation matrices, and combining them to obtain the position rotation matrix corresponding to the input vector includes: determining the target pixel position coordinates of each sub-vector in the target coordinate system based on the pixel position information corresponding to each sub-vector, determining the first position rotation matrix parameter value based on the preset attitude parameter, the target pixel position coordinates of each sub-vector, and the preset first matrix parameter calculation formula, determining the second position rotation matrix parameter value based on the target pixel position coordinates of each sub-vector and the preset second matrix parameter calculation formula, inputting the first position rotation matrix parameter value and the second position rotation matrix parameter value into the position rotation matrix formula to obtain the sub-position rotation matrix of each sub-vector, and combining a plurality of sub-position rotation matrices to obtain the position rotation matrix corresponding to the input vector.
[0112] In the embodiments of the present disclosure, determining the target pixel position coordinates of each sub-vector in the target coordinate system based on the pixel position information corresponding to each sub-vector includes: obtaining the initial pixel position coordinates of each sub-vector corresponding to the initial coordinate system based on the pixel position information corresponding to each sub-vector; and converting the initial pixel position coordinates of each sub-vector according to a preset coordinate system conversion formula to obtain the target pixel position coordinates of each sub-vector in the target coordinate system.
[0113] Specifically, before inputting into the Self-attention module, the input image undergoes a series of network encodings to obtain an input vector X that is proportional to the scale of the original image; assuming the size of the input image (training panoramic image sample) is (c, H, W), and the size of the input vector X is (c, h, w), where (c, h, w) = (c, H / N, W / N); N is a rational number; since the input vector X is aligned with the original image in terms of width and height dimensions, the input vector X can be encoded in the same way as each pixel of the above panoramic image, that is, any position (h i , w i ) in the input vector X corresponds to a vector X i ∈ R^(c×1), and the position encoding Referring to Figure 4 the structure of the Self-attention module, the corresponding Recombining {Q i , K i} into When the number of channels C of the input vector X is a multiple of 4, R pano_i can be applied to {Q i , K i}, as shown in formula (11).
[0114] f(Q i , i) = R pano_i Q i (10)
[0115] Thus, by representing the pixel positions of the panoramic image in the form of a rotation matrix to generate a subsequent training image generation model, the generation effect of the panoramic image is achieved.
[0116] Step 204: Adjust the model parameters of the to-be-trained image generation model based on the loss value between the training vector and the output vector of the training panoramic image sample to obtain an image generation model.
[0117] It should be noted that step 204 is the same as step 103. For the specific details, please refer to the detailed description of step 103 and will not be elaborated here.
[0118] Specifically, an image generation model to be trained is constructed according to network parameter design, the input vector of each Self-attention module is calculated, and the rotation matrix R at each position of the input vector is calculated based on formula (5). pano ; During the training and inference processes, referring to the RoPE encoding method, apply R pano to {Q, K} of the Self-attention module; Train and infer the network according to the original model architecture to obtain the image generation model.
[0119] The training scheme of the image generation model provided by the embodiments of the present disclosure obtains training data pairs; wherein, the training data pairs include training panoramic image samples and description texts corresponding to the training panoramic image samples. Input the training panoramic image samples and description texts into a pre-constructed image generation model to be trained to obtain the input vector corresponding to the training panoramic image sample, obtain each sub-vector corresponding to the input vector, encode the pixel position information corresponding to each sub-vector according to the position rotation matrix formula, obtain multiple sub-position rotation matrices and combine them to obtain the position rotation matrix corresponding to the input vector, and calculate the output vector based on the position rotation matrix corresponding to the input vector. Adjust the model parameters of the image generation model to be trained based on the loss value between the training vector of the training panoramic image sample and the output vector to obtain the image generation model. By adopting the above technical solution, the pixel position of each pixel point in the panoramic image is represented by the position rotation matrix, so that the pixel distance between each pixel point in the panoramic image can be accurately determined to train the image generation model. Thus, during the panoramic image generation process, the position rotation matrix is obtained based on the image generation model to determine the distance information of each pixel point in the generated panoramic image, improving the generation effect of the panoramic image.
[0120] Based on the description of the foregoing embodiments, after obtaining the image generation model, image generation can be performed based on the image generation model. Exemplarily, corresponding to Figure 3 as, for example, Figure 6 shown, replace the image latent vector with a noise vector of the same size ( Figure 6 the noise shown), and through the decoder of the DiT network and VAE, an image conforming to the text description "A livingroom with tv and sofa" can be obtained. The following will be described in detail with reference to Figure 7 .
[0121] Figure 7 is a schematic flowchart of an image generation method provided by the embodiments of the present disclosure. This method can be executed by an image generation device, where the device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 7 shown, this method includes:
[0122] Step 301: Obtain the image generation description text.
[0123] Step 302: Input the image generation description text into the image generation model to perform position encoding on the target vector corresponding to the preset noise vector through a preset position rotation matrix formula, obtain the position rotation matrix corresponding to the target vector, calculate the generated vector based on the position rotation matrix corresponding to the target vector, and decode the generated vector to obtain the target panoramic image.
[0124] In the embodiments of the present disclosure, the image generation model is obtained through the training method of the image generation model described in the foregoing embodiments.
[0125] In the embodiments of the present disclosure, the image generation description text is, for example, "A living room with tv and sofa", which is specifically input according to the actual generation requirements, and relevant noise vectors are preset.
[0126] Specifically, performing position encoding on the target vector corresponding to the preset noise vector through a preset position rotation matrix formula to obtain the position rotation matrix corresponding to the target vector includes: obtaining the target vector corresponding to the noise vector, obtaining each sub-target vector corresponding to the target vector, encoding the pixel position information corresponding to each sub-target vector according to the position rotation matrix formula, and obtaining a plurality of sub-target position rotation matrices and combining them to obtain the position rotation matrix corresponding to the target vector.
[0127] More specifically, determining the target pixel position coordinates of each sub-target vector in the target coordinate system based on the pixel position information corresponding to each sub-target vector, determining the first position rotation matrix parameter value based on the preset attitude parameters, the target pixel position coordinates of each sub-target vector, and the preset first matrix parameter calculation formula, determining the second position rotation matrix parameter value based on the target pixel position coordinates of each sub-target vector and the preset second matrix parameter calculation formula, inputting the first position rotation matrix parameter value and the second position rotation matrix parameter value into the position rotation matrix formula to obtain the sub-target position rotation matrix of each sub-target vector, and combining a plurality of sub-target position rotation matrices to obtain the position rotation matrix corresponding to the target vector.
[0128] In some embodiments, the method further includes: calculating based on the sub-target position rotation matrices corresponding to any two sub-target vectors and a preset angle calculation formula to obtain the target angle, determining the pixel distance information between any two sub-target vectors based on the target angle, and determining the pixel position relationship of the target panoramic image based on the pixel distance information.
[0129] Among them, the angle calculation formula is as shown in formula (6). By inputting the sub-goal position rotation matrix corresponding to any two sub-goal vectors into formula (6), the target angle can be obtained. The target angle is used to represent the distance between two pixel points, so that the pixel distance information between any two sub-goal vectors can be determined, and the pixel position relationship of the target panoramic image can be determined based on the pixel distance information. It can be understood that the panoramic image can be rotated arbitrarily in the left-right direction, while having absolute positions at the South Pole and the North Pole. Therefore, the absolute position information of the second position rotation matrix parameter is added to the position rotation matrix, and finally the relative position effect in the horizontal direction and the absolute position effect in the vertical direction are achieved.
[0130] Specifically, during the image generation process, the pixel distance between each pixel point in the generated target panoramic image can be accurately determined through the encoding method of the position rotation matrix, so that the pixel position relationship between each pixel point in the target panoramic image can be determined, improving the panoramic image generation effect.
[0131] Exemplarily, as Figure 8 shown, input the image generation description text "A living room with tv and sofa", and process the noise vector based on DiT. Perform position encoding on the target vector corresponding to the noise vector through the position encoding method of the present disclosure to obtain the position rotation matrix corresponding to the target vector, calculate the generated vector based on the position rotation matrix corresponding to the target vector, and decode the generated vector to obtain the target panoramic image.
[0132] The image generation scheme provided by the embodiments of the present disclosure obtains an image generation description text, inputs the image generation description text into an image generation model, performs position encoding on the target vector corresponding to a preset noise vector through a preset position rotation matrix formula to obtain the position rotation matrix corresponding to the target vector, calculates the generated vector based on the position rotation matrix corresponding to the target vector, and decodes the generated vector to obtain the target panoramic image. Thus, during the panoramic image generation process, by encoding according to the encoding method of the position rotation matrix, the target panoramic image is generated, improving the generation effect of the panoramic image.
[0133] Figure 9 FIG. is a schematic structural diagram of a training device for an image generation model provided by an embodiment of the present disclosure. The device can be implemented by software and / or hardware and is generally integrated in an electronic device. As Figure 9 shown, the device includes:
[0134] A first acquisition module 401, configured to acquire a training data pair; wherein, the training data pair includes a training panoramic image sample and a description text corresponding to the training panoramic image sample;
[0135] An input module 402, configured to input the training panoramic image sample and the description text into a pre-constructed image generation model to be trained, perform position encoding on an input vector corresponding to the training panoramic image sample through a preset position rotation matrix formula to obtain a position rotation matrix corresponding to the input vector, and calculate an output vector based on the position rotation matrix corresponding to the input vector;
[0136] A training module 403, configured to adjust model parameters of the image generation model to be trained based on a loss value between a training vector of the training panoramic image sample and the output vector to obtain an image generation model.
[0137] Optionally, the input module 402 includes:
[0138] A first acquisition unit, configured to input the training panoramic image sample and the description text into a pre-constructed image generation model to be trained to obtain an input vector corresponding to the training panoramic image sample;
[0139] A second acquisition unit, configured to acquire each sub-vector corresponding to the input vector;
[0140] An encoding unit, configured to encode pixel position information corresponding to each sub-vector according to the position rotation matrix formula, and combine multiple sub-position rotation matrices to obtain a position rotation matrix corresponding to the input vector.
[0141] Optionally, the encoding unit is specifically configured to:
[0142] Determine target pixel position coordinates of each sub-vector in a target coordinate system based on the pixel position information corresponding to each sub-vector;
[0143] Determine a first position rotation matrix parameter value based on a preset attitude parameter, the target pixel position coordinates of each sub-vector, and a preset first matrix parameter calculation formula;
[0144] Determine a second position rotation matrix parameter value based on the target pixel position coordinates of each sub-vector and a preset second matrix parameter calculation formula;
[0145] Input the first position rotation matrix parameter value and the second position rotation matrix parameter value into the position rotation matrix formula to obtain a sub-position rotation matrix of each sub-vector, and combine the multiple sub-position rotation matrices to obtain a position rotation matrix corresponding to the input vector.
[0146] Optionally, the determining target pixel position coordinates of each sub-vector in a target coordinate system based on the pixel position information corresponding to each sub-vector includes:
[0147] Obtain the initial pixel position coordinates of each sub-vector corresponding to the initial coordinate system based on the pixel position information corresponding to each sub-vector;
[0148] Convert the initial pixel position coordinates of each sub-vector according to a preset coordinate system conversion formula to obtain the target pixel position coordinates of each sub-vector in the target coordinate system.
[0149] The training device of the image generation model provided by the embodiments of the present disclosure can execute the training method of the image generation model provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0150] Figure 10 It is a schematic structural diagram of an image generation device provided by an embodiment of the present disclosure. This device can be implemented by software and / or hardware, and is generally integrated in an electronic device. As Figure 10 shown, this device includes:
[0151] A second acquisition module 501, configured to acquire an image generation description text;
[0152] A generation module 502, configured to input the image generation description text into an image generation model, perform position encoding on a target vector corresponding to a preset noise vector through a preset position rotation matrix formula to obtain a position rotation matrix corresponding to the target vector, calculate a generated vector based on the position rotation matrix corresponding to the target vector, and decode the generated vector to obtain a target panoramic image;
[0153] Wherein, the image generation model is obtained according to the training method of the image generation model described in the foregoing embodiments.
[0154] Optionally, the generation module 502 is specifically configured to:
[0155] Input the image generation description text into an image generation model to obtain a target vector corresponding to the noise vector;
[0156] Obtain each sub-target vector corresponding to the target vector;
[0157] Encode the pixel position information corresponding to each sub-target vector according to the position rotation matrix formula, and obtain a position rotation matrix corresponding to the target vector by combining multiple sub-target position rotation matrices.
[0158] Optionally, the device further includes:
[0159] A calculation module, configured to calculate a target angle based on the sub-target position rotation matrices corresponding to any two of the sub-target vectors and a preset angle calculation formula;
[0160] A first determination module, configured to determine pixel distance information between any two of the sub-target vectors based on the target angle;
[0161] A second determination module, configured to determine a pixel position relationship of the target panoramic image based on the pixel distance information.
[0162] The image generation device provided by the embodiments of the present disclosure can execute the image generation method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.
[0163] The embodiments of the present disclosure further provide a computer program product, including a computer program / instructions, which when executed by a processor, implement the training method of the image generation model provided by any embodiment of the present disclosure.
[0164] Figure 11 It is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. Specifically refer to Figure 11 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The electronic device 600 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is only an example, and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0165] As Figure 11 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0166] Typically, the following devices can be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 11 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0167] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above functions defined in the training method of the image generation model of the embodiment of the present disclosure are executed.
[0168] It should be noted that the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0169] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0170] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately and not be assembled into the electronic device.
[0171] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a training panoramic image sample and a description text corresponding to the training panoramic image sample, input them into an image generation model to be trained, perform position encoding on the input vector corresponding to the training panoramic image sample to obtain a position rotation matrix corresponding to the input vector, calculate an output vector based on the position rotation matrix corresponding to the input vector, and adjust the model parameters of the image generation model to be trained based on the loss value between the training vector of the training panoramic image sample and the output vector to obtain an image generation model.
[0172] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0173] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0174] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself.
[0175] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0176] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disk read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0177] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, including:
[0178] A processor;
[0179] A memory for storing executable instructions of the processor;
[0180] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of any of the image generation models provided by the present disclosure.
[0181] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium storing a computer program for executing the training method of any of the image generation models provided by the present disclosure.
[0182] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.
[0183] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0184] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A training method for an image generation model, characterized in that, Including: Obtain training data pairs; wherein, the training data pairs include training panoramic image samples and description texts corresponding to the training panoramic image samples; Input the training panoramic image samples and the description texts into a pre-constructed image generation model to be trained, so as to perform position encoding on the input vector corresponding to the training panoramic image sample through a preset position rotation matrix formula, obtain the position rotation matrix corresponding to the input vector, and calculate an output vector based on the position rotation matrix corresponding to the input vector; Adjust the model parameters of the image generation model to be trained based on the loss value between the training vector of the training panoramic image sample and the output vector to obtain an image generation model.
2. The method according to claim 1, characterized in that, The performing position encoding on the input vector corresponding to the training panoramic image sample through a preset position rotation matrix formula to obtain the position rotation matrix corresponding to the input vector includes: Obtain the input vector corresponding to the training panoramic image sample; Obtain each sub-vector corresponding to the input vector; Encode the pixel position information corresponding to each sub-vector according to the position rotation matrix formula, and obtain a plurality of sub-position rotation matrices for combination to obtain the position rotation matrix corresponding to the input vector.
3. The method according to claim 2, wherein The encoding the pixel position information corresponding to each sub-vector according to the position rotation matrix formula, and obtaining a plurality of sub-position rotation matrices for combination to obtain the position rotation matrix corresponding to the input vector includes: Determine the target pixel position coordinates of each sub-vector in the target coordinate system based on the pixel position information corresponding to each sub-vector; Determine the first position rotation matrix parameter value based on a preset attitude parameter, the target pixel position coordinates of each sub-vector, and a preset first matrix parameter calculation formula; Determine the second position rotation matrix parameter value based on the target pixel position coordinates of each sub-vector and a preset second matrix parameter calculation formula; Input the first position rotation matrix parameter value and the second position rotation matrix parameter value into the position rotation matrix formula to obtain the sub-position rotation matrix of each sub-vector, and combine the plurality of sub-position rotation matrices to obtain the position rotation matrix corresponding to the input vector.
4. The method according to claim 3, characterized in that, The determining the target pixel position coordinates of each sub-vector in the target coordinate system based on the pixel position information corresponding to each sub-vector includes: Obtain the initial pixel position coordinates of each sub-vector in the initial coordinate system based on the pixel position information corresponding to each sub-vector; Convert the initial pixel position coordinates of each sub-vector according to a preset coordinate system conversion formula to obtain the target pixel position coordinates of each sub-vector in the target coordinate system.
5. An image generation method, characterized in that, Including: Obtain an image generation description text; Input the image generation description text into the image generation model, so as to perform position encoding on the target vector corresponding to a preset noise vector through a preset position rotation matrix formula, obtain the position rotation matrix corresponding to the target vector, calculate a generated vector based on the position rotation matrix corresponding to the target vector, and decode the generated vector to obtain a target panoramic image; Among them, the image generation model is obtained according to the training method of the image generation model described in any one of claims 1-4.
6. The method according to claim 5, wherein The position encoding of the target vector corresponding to the preset noise vector by the preset position rotation matrix formula to obtain the position rotation matrix corresponding to the target vector includes: Obtaining the target vector corresponding to the noise vector; Obtaining each sub-target vector corresponding to the target vector; Encoding the pixel position information corresponding to each sub-target vector according to the position rotation matrix formula, and obtaining a plurality of sub-target position rotation matrices for combination to obtain the position rotation matrix corresponding to the target vector.
7. The method according to claim 6, wherein The method further includes: Calculating based on the sub-target position rotation matrices corresponding to any two of the sub-target vectors and a preset angle calculation formula to obtain a target angle; Determining the pixel distance information between any two of the sub-target vectors based on the target angle; Determining the pixel position relationship of the target panoramic image based on the pixel distance information.
8. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the image generation model described in any one of claims 1-4 or the image generation method described in any one of claims 5-7.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the training method of the image generation model described in any one of claims 1-4 or the image generation method described in any one of claims 5-7.
10. A computer program product, characterized in that, Including a computer program, wherein the computer program, when executed by a processor, implements the training method of the image generation model described in any one of claims 1-4 or the image generation method described in any one of claims 5-7.