Image generation and model training method, computing device, storage medium and product
By acquiring mapping information and utilizing an image generation model in combination with a target mask to expand an image, the problems of high cost and unstable quality of extended image generation in the prior art are solved, and low-cost, high-quality extended image generation is achieved.
Patent Information
- Application Number
- CN202510824633.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-17
AI Technical Summary
In the existing technology, the generation cost of extended images is high, the scenes are limited, and the quality is unstable. It is impossible to generate extended images with low cost and high quality in multiple scenes.
By obtaining the mapping information between the target original image and the sample extended image, the target extended image is generated by combining the target mask extended image with the image generation model. The image generation model is trained based on multiple training sample data, and the diffusion model is used to simulate the diffusion process to generate the extended image.
The production cost of extended images is reduced, the production threshold is lowered, and the generated extended images have real spatial mapping effects and high-quality content information, which improves production efficiency.
Smart Images

Figure CN120807273A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular to an image generation and model training method, a computing device, a storage medium and a product. BACKGROUND
[0002] With the development of computer technology, extended images capable of capturing a wider field of view and providing richer visual information are increasingly popular in various industries. An extended image is an image obtained by extending the image of a regular field of view, for example, a panoramic image, a wide-angle image, etc.
[0003] Currently, the production of extended images usually relies on professional cameras, and the use scenarios are limited, for example, non-realistic scenes cannot be produced, and at the same time, professional cameras are expensive, resulting in high production costs. In addition, extended images can also be generated through image processing algorithms, but the quality of the extended images generated by this method is unstable due to the limitation of algorithm accuracy.
[0004] Therefore, how to generate extended images with low cost and high quality in multiple scenarios has become a problem to be solved. SUMMARY
[0005] Embodiments of the present application provide an image generation and model training method, a computing device, a storage medium and a product to solve the problems of high cost, limited use scenarios and unstable quality in the generation process of extended images in the prior art.
[0006] In a first aspect, an image generation method is provided in embodiments of the present application, comprising:
[0007] obtaining a target original image and first mapping information; the first mapping information is obtained based on a first sample extended image; the first mapping information includes a mapping relationship between a pixel point in a first dimensional space and a position point in a second dimensional space of the first sample extended image;
[0008] generating a target mask extended image corresponding to the target original image;
[0009] generating a target extended image based on the first mapping information, the target original image and the target mask extended image using an image generation model;
[0010] wherein the image generation model is trained based on a plurality of training sample data; each training sample data includes a sample original image and second mapping information, and a second sample extended image corresponding to the sample original image.
[0011] In a second aspect, an embodiment of the present application provides an image generation method, comprising:
[0012] Acquire a target original image and first mapping information; the first mapping information is constructed based on the first sample panoramic image; the first mapping information includes a mapping relationship between pixel points of the first sample panoramic image in a two-dimensional space and position points in a spherical coordinate system;
[0013] Generating a target mask panoramic image corresponding to the target sample image;
[0014] Generate a target panoramic image by using an image generation model based on the first mapping information and combining the target original image and the target mask panoramic image;
[0015] The image generation model is trained based on a plurality of training sample data; each training sample data includes a sample original image and second mapping information, as well as a second sample panoramic image corresponding to the sample original image.
[0016] In a third aspect, an embodiment of the present application provides a model training method, comprising:
[0017] Acquire second mapping information; the second mapping information is constructed based on the third sample extended image; the second mapping information includes a mapping relationship between pixel points of the third sample extended image in the first dimensional space and position points in the second dimensional space;
[0018] Acquire a second sample extended image, and a sample original image and a sample mask extended image corresponding to the second sample extended image;
[0019] Generate a predicted extended image by using an image generation model based on the second mapping information and combining the second sample extended image, the sample original image, and the sample mask extended image;
[0020] training the image generation model according to difference information between the predicted extended image and the second sample extended image;
[0021] The image generation model is used to generate a target extended image based on the first mapping information, in combination with the target original image and the target mask extended image corresponding to the target original image.
[0022] In a fourth aspect, an embodiment of the present application provides a model training method, comprising:
[0023] Obtaining training samples corresponding to the distorted extended image type, the non-distorted extended image type, and the original image type, respectively; the training samples include sample training images, or sample training images and description text;
[0024] extract sample features corresponding to the training samples of different image types respectively by using a feature extraction model; the sample features include sample image features, or sample image features and sample text features;
[0025] train the feature extraction model based on the sample features, similarity between sample features corresponding to the same image type, and similarity between sample features corresponding to different image types meeting training requirements as a training target;
[0026] The feature extraction model is used to extract target image features of a target extended image, and extract first image features and / or first text features corresponding to a non-distorted extended image type; the target image features, and the first image features and / or first text features are used to determine whether the target extended image meets distortion requirements; the target extended image is generated by using an image generation model based on first mapping information, combined with a target original image and a target mask extended image corresponding to the target original image; the first mapping information is obtained based on a first sample extended image; the first mapping information includes a mapping relationship between a pixel point of the first sample extended image in a first dimensional space and a position point in a second dimensional space.
[0027] In a fifth aspect, an embodiment of the present application provides a computing device, including a processing component, a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component, to implement the image generation method of the first aspect or the second aspect, or implement the model training method of the third aspect or the fourth aspect.
[0028] In a sixth aspect, an embodiment of the present application provides a computer storage medium, storing a computer program, when the computer program is executed by a computer, the image generation method of the first aspect or the second aspect is implemented, or the model training method of the third aspect or the fourth aspect is implemented.
[0029] In a seventh aspect, an embodiment of the present application provides a computer program product, including a computer program or instructions, when the computer program or instructions are executed by a computer, the image generation method of the first aspect or the second aspect is implemented, or the model training method of the third aspect or the fourth aspect is implemented.
[0030] In the embodiment of the present application, the target original image is obtained, and the first mapping information including the mapping relationship between the pixel points in the first dimension space and the position points in the second dimension space of the first sample extended image is obtained. The target mask extended image corresponding to the target original image is generated. Then, the image generation model is used to generate the target extended image based on the first mapping information, in combination with the target original image and the target mask extended image. The image generation model is trained based on a plurality of training sample data. Each training sample data includes a sample original image, second mapping information, and a second sample extended image corresponding to the sample original image. In the process of generating the target extended image by using the image generation model, the professional camera shooting is not needed, the production cost is reduced, and the production threshold of the extended image is reduced without being limited by the shooting scene. In addition, in the process of training or generating the target extended image, the image generation model of the embodiment is guided by the mapping information constructed based on the sample extended image, so that the image generation model perceives the mapping relationship between the pixel points in the first dimension space and the position points in the second dimension space, so that the target extended image generated has a real spatial mapping effect. In addition, the target original image and the target mask extended image corresponding thereto are used as the guide of the extended content, the accuracy of the content information in the target extended image is improved, the quality of the target extended image generated is greatly improved, and the production efficiency is improved.
[0031] These aspects or other aspects of the present application will be more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0032] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, illustrate the exemplary embodiments of the present application and the description thereof, and serve to explain the present application, and do not limit the present application in any way. In the drawings:
[0033] Figure 1 A flowchart of one embodiment of an image generation method provided by the present application is shown;
[0034] Figure 2 A structural schematic diagram of an image generation model provided by the present application is shown;
[0035] Figure 3 A flowchart of one embodiment of a model training method provided by the present application is shown;
[0036] Figure 4 A flowchart of one embodiment of another model training method provided by the present application is shown;
[0037] Figure 5 A structural schematic diagram of a feature extraction model provided by the present application is shown;
[0038] Figure 6 A structural diagram of an image generation model provided by the present application in actual application is shown;
[0039] Figure 7 A flow chart of one embodiment of an image generation method provided by the present application in actual application is shown;
[0040] Figure 8 A structural diagram of one embodiment of an image generation device provided by the present application is shown;
[0041] Figure 9 A structural diagram of one embodiment of a model training device provided by the present application is shown;
[0042] Figure 10 A structural diagram of one embodiment of another model training device provided by the present application is shown;
[0043] Figure 11 A structural diagram of one embodiment of a computing device provided by the present application is shown. DETAILED DESCRIPTION
[0044] In order to make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in conjunction with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative work fall within the scope of protection of the present application.
[0045] It should be noted that in the case where the embodiments of the present application involve user information, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation portal for user to choose authorization or refusal. In addition, the various models (including but not limited to language models or large models) involved in the present application are in compliance with relevant legal and standard regulations.
[0046] The technical solutions of the embodiments of the present application can be applied to the scene of extending the field of view of a common target original image to generate a target extended image, for example, can be applied to the scene of extending the field of view of a two-dimensional space image under a standard field of view to generate a panoramic image corresponding to the two-dimensional space image. As described in the background, the production method of the target extended image usually depends on professional camera shooting, and the use scene is limited, for example, it is impossible to produce a non-realistic scene, at the same time, the professional camera is expensive and needs to be operated by a professional, resulting in high shooting cost. In addition, the target extended image can also be generated through an image processing algorithm. For example, an artificial intelligence model can be used to generate six perspective images corresponding to the six projection faces of a cube based on a cube projection algorithm, and then an equirectangular projection algorithm is used to stitch the six generated perspective images to obtain the target extended image. However, due to the limitation of algorithm accuracy, the spatial mapping effect of the target extended image generated through the image processing algorithm can be inaccurate, thereby affecting the image quality and subsequent application. For example, if the target extended image needs to be applied to a perspective view under a certain specific field of view, due to the inaccurate spatial mapping effect of the target extended image, the perspective view obtained by projecting the target extended image will have distortion problems and cannot be used continuously.
[0047] With the development of artificial intelligence generated content (AIGC) technology, especially the continuous development of generative technology represented by diffusion models such as Stable Diffusion, generative technology has been widely applied in artistic expression and product design to personalized content creation. The inventors thought that on this basis, a Stable Diffusion can be used as the main network, and by simulating the diffusion process, the random noise is gradually converted into the target expansion image, so as to realize the automatic generation of the target expansion image with low cost, high quality and scalability, and reduce the cost of target expansion image production and improve the production efficiency. Therefore, the inventors have proposed the technical scheme of the present application after a series of researches: obtaining a target original image and first mapping information including a mapping relationship between a pixel point of a first sample expansion image in a first dimensional space and a position point in a second dimensional space; generating a target mask expansion image corresponding to the target original image; and then using an image generation model to generate a target expansion image based on the first mapping information, the target original image and the target mask expansion image. The image generation model is trained according to the second mapping information, the second sample expansion image, the sample original image corresponding to the second sample expansion image and the sample mask expansion image. In the process of using the image generation model to produce the target expansion image, the present application does not need to rely on professional camera shooting, which reduces the production cost and is not limited by the shooting scene, thereby reducing the production threshold of the expansion image. In addition, the image generation model of the present embodiment uses the mapping information constructed by the sample expansion image as a guide in the process of training and generating the target expansion image, so that the image generation model perceives the mapping relationship between the pixel point in the first dimensional space and the position point in the second dimensional space, so that the generated target expansion image has a real spatial mapping effect. In addition, the target original image and the target mask expansion image corresponding thereto are used as a guide for expansion content, which improves the accuracy of the content information in the target expansion image, greatly improves the quality of the generated target expansion image, and improves the production efficiency.
[0048] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0049] The implementation details of the technical solutions of the embodiments of the present application will be described in detail below.
[0050] Figure 1A flowchart of an embodiment of an image generation method provided in the present application, the technical solution of the embodiment can be applied to a server.
[0051] The server can include a server providing various services, such as a server providing instant messaging, a server providing audio and video communication, a proxy server providing gateway services, etc. It should be noted that the server can be implemented as a distributed server cluster composed of multiple servers, or as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content distribution networks (CDN), and big data and artificial intelligence platforms, etc. Basic cloud computing services, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0052] Figure 1 The image generation method shown can include the following steps:
[0053] S101, obtaining a target original image and first mapping information.
[0054] The target original image can be any two-dimensional image to be processed. That is, on the basis of the target original image, the field of view and the content in the horizontal and / or vertical directions can be expanded to obtain a target expanded image. For example, if the target expanded image to be generated is a panoramic image, the target original image at this time can be a perspective image in a two-dimensional space. By expanding the field of view in the horizontal direction to 360 degrees and the field of view in the vertical direction to 180 degrees and adding the real space mapping effect of the corresponding three-dimensional spherical space and two-dimensional space of the equirectangular projection in the expansion, a target panoramic image can be obtained.
[0055] The first mapping information can include the mapping relationship between the pixel points in the first dimensional space and the position points in the second dimensional space of the first sample expanded image. The first dimensional space can be the dimensional space corresponding to the image content of the expanded image in the image, and the second dimensional space can be the dimensional space representing the actual space corresponding to the image content of the expanded image. For example, if the expanded image is a panoramic image, especially a panoramic image with a field of view of 360 degrees in the horizontal direction and 180 degrees in the vertical direction, since the panoramic image is a two-dimensional image, the first dimensional space can be a two-dimensional space. Since the content in the panoramic image is essentially data in a three-dimensional spherical space, the second dimensional space can be a three-dimensional spherical space. Through projection, the three-dimensional spherical space content can be represented by two-dimensional space data for convenient storage and viewing, etc.
[0056] The first mapping information can be obtained based on a first sample extended image. The first sample extended image can be any one of the extended images with the same size as the target extended image to be generated. It should be noted that in the embodiment, the first mapping information corresponding to the extended images with the same size is the same, so the first mapping information needs to be constructed only once for the extended images with the same size. The size of the extended image in the embodiment is related to the dimension of the image. For example, if the extended image is a two-dimensional image, the size at this time can refer to the length and width of the extended image.
[0057] Optionally, the manner of obtaining the target original image in the embodiment can be to shoot a two-dimensional image in a real scene as the target original image, or to draw a two-dimensional perspective view of a virtual scene based on the perspective principle, which is used to represent the spatial relationship of a three-dimensional object in a two-dimensional space as the target original image.
[0058] The manner of obtaining the mapping relationship can be: based on the size of the target extended image to be generated, determining whether the first mapping information has been constructed for the size, if yes, the first mapping information constructed in advance for the size can be directly obtained. If not, an existing extended image with the same size as the size, i.e. the first sample extended image, can be selected, and then the first mapping information is constructed based on the first sample extended image.
[0059] Optionally, the manner of constructing the first mapping information in the embodiment can be: determining the mapping relationship between the pixel points of the first sample extended image in the first dimensional space and the position points in the second dimensional space; performing spatial edge alignment processing on the mapping relationship, and obtaining the first mapping information based on the mapping relationship after the spatial edge alignment processing.
[0060] Specifically, the mapping relationship between the pixel points of the first sample extended image in the first dimensional space and the position points in the second dimensional space can be determined based on the relationship definition between the pixel points of the first sample extended image in the first dimensional space and the position points in the second dimensional space. Since the first dimensional space and the second dimensional space are in different dimensions, the edge of the mapping relationship determined based on the relationship definition can be not aligned (i.e. the image edge should be continuous, but the mapping relationship is not continuous). At this time, the mapping relationship can be processed by position coding to perform spatial edge alignment, and the mapping relationship after the spatial edge alignment can be taken as the first mapping information.
[0061] For example, if the first sample extended image is a panoramic image I∈R H×W×3 At this time, the relationship between the pixel points (i,j) of the panoramic image in the two-dimensional space and the position points (θ,φ,r) in the three-dimensional spherical space S is as follows formula (1)-(4).
[0062] S(θ,φ,r)=I(i,j) (1)
[0063] θ=(2i / H-1)π (2)
[0064] φ=(2j / W-1)π / 2 (3)
[0065] r=1 (4)
[0066] Among them, S(θ, φ, r) is the position representation of the position point in the three-dimensional spherical space S in the spherical coordinate system. θ is the azimuth angle in the spherical coordinate system, which ranges from -π to π. φ is the elevation angle in the spherical coordinate system, which ranges from -π / 2 to π / 2. r is the radius of the sphere. The center of the spherical coordinate system is the camera position. Since the r value is fixed, the mapping relationship D∈R between the pixel points in the panoramic image and the position points in the spherical coordinate system can be obtained. H×W×2D The following formula (5):
[0067] D(i,j)=(θ,φ) (5)
[0068] Since the left and right edges of the three-dimensional sphere space are seamlessly aligned in the two-dimensional representation, the current mapping relationship D does not exhibit this characteristic (i.e., the -π and π of the leftmost and rightmost edges are discontinuous). In this case, the first-order Taylor expansion position encoding γ(·) can be introduced to perform spatial edge alignment on the mapping relationship D to make it continuous. The final first mapping information is obtained as follows: Formulas (6)-(8).
[0069] D(i,j)=(γ(θ),γ(φ)) (6)
[0070] γ(θ)=[sin(2 0 πθ),cos(2 0 πθ)] (7)
[0071] γ(φ)=[sin(2 0 πφ),cos(2 0 πφ)] (8)
[0072] Wherein, γ(θ) and γ(φ) are the results of first-order Taylor expansion position encoding of the corresponding formulas of azimuth angle θ and elevation angle φ in the spherical coordinate system (i.e., formulas (2) and (3)).
[0073] By the above method, the first mapping information D∈R is constructed H×W×4D, the left and right edges are aligned (i.e., sin(-π) and sin(π), cos(-π) and cos(π) are equal), the accuracy of the first mapping information construction is improved, and the quality of the target extended image generated based on the first mapping information is improved.
[0074] S102, a target mask extended image corresponding to the target original image is generated.
[0075] The target mask extended image guides the image generation model to perform content filling in the process of generating the target extended image. For example, the image generation model can be guided to perform content filling in the mask region of the target mask extended image. The target mask extended image includes a partial extended image and an outer drawing region mask image. The partial extended image is an extended image containing only part of the image content, i.e., only the central region contains image content, and the surrounding region (i.e., the outer drawing region) to be expanded is masked. The outer drawing region mask image can be an image obtained by masking the central region and the outer drawing region in different ways.
[0076] Optionally, the manner in which the embodiment generates the target mask extended image corresponding to the target original image can be: performing projection processing (such as inverse equirectangular projection) on the target original image, taking the processed image as the central region, expanding the surrounding region in a mask processing manner based on the size of the target extended image to be generated (for example, masking the expanded region to black), and obtaining a partial extended image. Then, the central region containing image content in the partial extended image is masked in a different manner from the outer drawing region (for example, masking the central region to white), and an outer drawing region mask image is obtained.
[0077] S103, using an image generation model, generating a target extended image based on the first mapping information, combining the target original image and the target mask extended image.
[0078] The image generation model can be a neural network model for generating a target extended image. For example, it can be a diffusion model. The image generation model can be trained based on a plurality of training sample data; each training sample data can include a sample original image and second mapping information, and a second sample extended image corresponding to the sample original image. Alternatively, when training the image generation model based on each training sample data, a sample mask extended image corresponding to the sample original image can be generated based on the sample original image in each training sample data. A sample mask extended image corresponding to the sample original image can also be generated based on the second sample extended image in each training sample data. Alternatively, the sample original image can be pre-configured to correspond to the second sample extended image, and of course it can also be generated based on the second sample extended image when training the image generation model. Then, the image generation model can be trained based on the second mapping information, the second sample extended image, the sample original image, and the sample mask extended image. The second mapping information is obtained based on a third sample extended image; the second mapping information includes the mapping relationship between the pixel points in the first dimensional space and the position points in the second dimensional space of the third sample extended image. The second mapping information is obtained in a manner similar to the first mapping information described in the above embodiments. The generation of the sample original image and the sample mask extended image, and the specific training process of the model will be described in detail in subsequent embodiments.
[0079] Alternatively, the first mapping information, the target original image, and the target mask extended image are input as control conditions to guide the image generation model to simulate a diffusion process, gradually convert random noise into a target extended image, so that the generated target extended image meets the mapping position condition constrained by the first mapping information, and meets the extended content condition constrained by the target original image and the target mask extended image. Specifically, the image generation model can be used to generate extended content corresponding to the extended position in the second dimensional space in combination with the target original image and the target mask extended image, and the extended content can be mapped to the corresponding pixel point position in the first dimensional space according to the first mapping information, thereby generating a target panoramic image. In the case of a panoramic image as the target extended image, the target mask extended image is a target mask panoramic image. The image generation model can determine panoramic content in a three-dimensional spherical space in combination with the target original image and the target mask panoramic image, and project the panoramic content to the two-dimensional space according to the mapping relationship to generate a target panoramic image.
[0080] It should be noted that various models involved in this paper, including the image generation model of the present embodiment and the feature extraction model that may appear in the corresponding embodiments below, can be a language model (LM) or a multimodal model (MM) based on artificial intelligence, and the present embodiment does not limit the number of model parameters supported by the model to meet the actual needs.
[0081] The scheme of the present embodiment obtains a target original image, and first mapping information including a mapping relationship between a pixel point of a first sample extended image in a first dimensional space and a position point in a second dimensional space; generates a target mask extended image corresponding to the target original image; and then generates a target extended image based on the first mapping information, in combination with the target original image and the target mask extended image, by using an image generation model. The image generation model is trained based on a plurality of training sample data. In the process of using the image generation model to make the target extended image, the present embodiment does not need to rely on professional cameras for shooting, thereby reducing the production cost and lowering the production threshold of the extended image, which is not limited by the shooting scene. In addition, in the process of training or generating the target extended image, the image generation model of the present embodiment is guided by the mapping information constructed by the sample extended image, so that the image generation model perceives the mapping relationship between the pixel point in the first dimensional space and the position point in the second dimensional space, thereby making the generated target extended image have a real spatial mapping effect. In addition, the target original image and the target mask extended image corresponding thereto are used as a guide for the extension content, thereby improving the accuracy of the content information in the target extended image and greatly improving the quality of the generated target extended image and the production efficiency.
[0082] In some embodiments, as shown in Figure 2 The image generation model of the present embodiment includes three network branches, i.e., a position generation network 21, a content generation network 22, and an image generation network 23. Among them, the position generation network 21 and the content generation network 22 are two control networks of the image generation model, which are used to guide the image generation network 23 to generate a target extended image that meets the constraint condition. For example, the position generation network 21, the content generation network 22, and the image generation network 23 of the present embodiment can be a U-Net (a convolutional neural network used for image segmentation) built based on a Block (a basic unit for building a network) of a Transformer (a deep learning model based on a self-attention mechanism) structure.
[0083] Based on the image generation model shown in Figure 2 The generation of the target extended image based on the first mapping information, in combination with the target original image and the target mask extended image, can include the following four steps:
[0084] Step one, determining the position guide information based on the first mapping information by using the position generation network 21.
[0085] The position guide information can be information used to guide the mapping position of the extended content in the target extended image during the generation of the target extended image, so that the generated target extended image has a real spatial mapping effect, that is, there is no distortion problem after the target extended image is converted into a perspective view under a certain perspective.
[0086] In this embodiment, the first mapping information can be input into the position generation network 21, and the position generation network 21 determines the position guide information corresponding to the projection relationship between the pixel points in the first dimensional space and the position points in the second dimensional space based on the first mapping information.
[0087] In some embodiments, as shown in Figure 2 The position generation network 21 of this embodiment at least includes a position encoding sub-network 211, which includes a plurality of position encoding blocks, that is Figure 2 The solid rectangles in the position encoding sub-network 211 correspond to the position encoding blocks, and the dashed rectangles correspond to the zero convolution blocks. The position encoding sub-network 211 is used to analyze the position guide information required when generating the target extended image. The position encoding sub-network 211 can be composed of a plurality of U-Net structure Transformer Blocks. Each Transformer Block (i.e., the Block of the Transformer structure) can correspond to one of the plurality of position encoding blocks. For example, the first Transformer Block in the position encoding sub-network 211 corresponds to the first position encoding block in the plurality of position encoding blocks. Each position encoding block can be composed of multiple layers of neural networks, specifically, can be composed of convolution layers, attention layers, and residual layers, etc., for processing the received data and passing the processing results to the next position encoding block.
[0088] In this embodiment, the noise information can be input into the first position encoding block of the position encoding sub-network, and the first mapping information can be input into the plurality of position encoding blocks respectively, to obtain the position guide information output by the position encoding sub-network. The noise information can be random noise. As shown in Figure 2As shown, the noise information can be input as the backbone network of the position encoding subnetwork 211, and at the same time, the first mapping information is injected into each Transformer Block (i.e., position encoding block) of the position encoding subnetwork 211 in a position encoding manner as a control condition, that is, the first Transformer Block processes the received noise information and the first mapping information (such as determining the dependency between data, extracting local features and global features through attention mechanism, etc.), and the processing result is transmitted as the output of the first Transformer Block. The transfer feature is input into the second Transformer Block, and at the same time, the first mapping information is also injected into the second Transformer Block in a position encoding manner as a control condition. At this time, the second Transformer Block also further processes the received transfer feature and the first mapping information (such as further determining the dependency between data, extracting local features, global features, etc. through attention mechanism), and the processing result is transmitted as the output of the second Transformer Block. The transfer feature is input into the third Transformer Block, and at the same time, the first mapping information is injected into the third Transformer Block in a position encoding manner as a control condition. In this way, the outputs of these Transformer Blocks are finally processed by a zero-roller and respectively input into the corresponding Transformer Blocks in the U-Net model 232. In this embodiment, the first mapping information is used as position encoding, and the first mapping information is injected into each Transformer Block in the position encoding subnetwork 211, that is, each Transformer Block is injected with the control condition related to the mapping position, so that the position generation network 21 can more accurately determine the position guide information, and further ensure that the target expansion image has a real spatial mapping effect in the subsequent process of generating the target expansion image.
[0089] In some embodiments, as Figure 2As shown, the position generation network 21 of the embodiment can also include a first encoding sub-network 212. The first encoding sub-network 212 can be used to convert the first mapping information into a format that can be processed by the position encoding sub-network 211 while preserving the spatial perception of the first mapping information. Accordingly, when the first mapping information is input into the position encoding sub-network 211, the mapping features of the first mapping information can be extracted by the first encoding sub-network 212, and the mapping features can be input into the multiple position encoding blocks of the position encoding sub-network 211, respectively. Specifically, the first mapping information can be first input into the first encoding sub-network 212, the first encoding sub-network 212 can perform feature encoding on the received first mapping information, complete the conversion of the first mapping information into a sequence, and inject spatial information to obtain mapping features with aligned dimensions, and then the mapping features with aligned dimensions can be further input into each Transformer Block (i.e., each position encoding block) of the position encoding sub-network 211. The advantage of this setting is that the first mapping information can be input in a format suitable for the position encoding sub-network 211, and the spatial perception of the first mapping information can be preserved, which facilitates the multiple Transformer Blocks of the position encoding sub-network 211 to more accurately extract position guidance information.
[0090] Optionally, the first encoding sub-network 212 of the embodiment can further include a first encoding layer (DistortEncoder) and a first embedding layer (Distort Embedding). The first encoding layer is used to convert the first mapping information into a feature sequence with spatial information through block processing, linear projection, and position encoding. The first embedding layer is used to enable the multiple position encoding blocks to understand the first mapping information through high-dimensional vector representation. Specifically, the first mapping information can be first input into the first encoding layer, and the output result of the first encoding layer can be further input into the first embedding layer to obtain mapping features output by the first embedding layer.
[0091] Step two, using the content generation network 22 to determine content guidance information based on the target original image and the target mask extended image.
[0092] The content guidance information can be information used to guide the image content under the extended field of view during the generation of the target extended image, so that the target extended image has real image content.
[0093] The embodiment can input both the target original image and the target mask extended image into the content generation network 22, determine the field of view range to be extended based on the target mask extended image by the content generation network 22, and determine the content guidance information corresponding to the field of view range to be extended according to the image content in the target original image.
[0094] In some embodiments, as Figure 2 As shown, the content generation network 22 of this embodiment includes at least a second encoding subnetwork 221, a third encoding subnetwork 222, and a content encoding subnetwork 223; the content encoding subnetwork 223 includes multiple content encoding blocks. The structure of the multiple content encoding blocks in this embodiment is similar to the structure of the multiple position encoding blocks, and is used to parse the content guidance information required to generate the target extended image. The content encoding subnetwork 223 can be composed of multiple Transformer Blocks with a U-Net structure. Each Transformer Block (i.e., a Block of the Transformer structure) can correspond to one of the multiple content encoding blocks. For example, the first Transformer Block in the content encoding subnetwork 223 corresponds to the first content encoding block in the multiple content encoding blocks. The second encoding subnetwork 221 can be used to perform content encoding on the target masked extended image, and the first encoding features obtained by encoding are adapted to the multiple content encoding blocks. The third encoding subnetwork 222 can be used to perform content encoding on the image content of the target original image and align the dimensions of the encoding result with the text input dimensions of the multiple content encoding blocks to facilitate subsequent conditional injection of the text input of the content encoding subnetwork 223.
[0095] This embodiment can be to use the second encoding sub-network 221 to extract the first encoding feature of the target mask expansion image; use the third encoding sub-network 222 to extract the second encoding feature of the target original image; align the dimension of the second encoding feature with the text input dimension of multiple content encoding blocks; input the first encoding feature and noise information into the first content encoding block (i.e., the first Transformer Block) in the content encoding sub-network 223, and use the second encoding feature as the text input of the corresponding content encoding block in the content encoding sub-network 223 to obtain content guidance information output by multiple content encoding blocks. Figure 2As shown, the embodiment can be that the target mask extended image is input to the second encoding sub-network 221, the content of the target mask extended image is encoded by the second encoding sub-network 221, the conversion from the image content to the sequence is completed, and the spatial information is injected to obtain the first encoding feature suitable for the input of multiple content encoding blocks. The target original image is input to the third encoding sub-network 222, the image content of the target original image is encoded by the third encoding sub-network 222, and the dimension of the encoding result is aligned with the text input dimension of the multiple content encoding blocks to obtain the second encoding feature. Then, after the first encoding feature is spliced with the noise information (the noise information of this step is the same as the noise information input in the above step one), it is input to the first content encoding block (i.e. the first Transformer Block) of the content encoding sub-network 223, and the second encoding feature aligned with the text input dimension of the first content encoding block is input to the first content encoding block instead of the text input of this dimension. At this time, the first content encoding block will process the received data (such as determining the dependency between the data, extracting local features and global features through the attention mechanism, etc.), and the processing result is input to the second content encoding block (i.e. the second Transformer Block) as the transfer feature output by the first content encoding block. At the same time, the second encoding feature aligned with the text input dimension of the second content encoding block is input to the second content encoding block instead of the text input of this dimension. At this time, the second content encoding block will also further process the received data (such as further determining the dependency between the data, extracting local features and global features through the attention mechanism, etc.), and the processing result is input to the third content encoding block (i.e. the third Transformer Block) as the transfer feature output by the second content encoding block. The second encoding feature aligned with the text input dimension of the third content encoding block is input to the third content encoding block instead of the text input of this dimension. In this way, finally, the outputs of these Transformer Blocks are processed by a zero volume machine and respectively input to the corresponding Transformer Blocks of the U-Net model 232. The target mask extended image and the target original image are encoded and processed by the encoding sub-network in the present application to ensure the input requirements of the multiple content encoding blocks. In addition, after the encoding features of the target original image are aligned with the text input dimension of the multiple content encoding blocks, the text input is replaced and injected into the corresponding multiple content encoding blocks, which improves the accuracy of the content guide information determination.
[0096] Optionally, the third encoding sub-network 222 of the embodiment can further include a second encoding layer, a projection layer and a second embedding layer. The second encoding layer is configured to convert the image content of the target original image into a feature sequence with spatial information through block processing, linear projection and position encoding, etc. For example, the second encoding layer can be a variational auto-encoder (VAE). The projection layer is configured to align the dimension of the encoding result of the target original image with the text input dimension of the plurality of content encoding blocks. For example, the projection layer can be a multilayer perceptron (MLP). The second embedding layer is configured to enable the plurality of content encoding blocks to understand the semantic and positional relationship of the image content through high-dimensional vector representation. Specifically, the target original image can be input to the second encoding layer first, and the output image features of the second encoding layer are input to the projection layer to align the image features with the text input dimension of the plurality of content encoding blocks, and then the image features are input to the second embedding layer to obtain the second encoding features output by the second embedding layer.
[0097] Step three, using the image generation network 23 to generate the target extended image by combining the position guide information and the content guide information.
[0098] In this embodiment, the noise information (the same as the noise information in steps one and two) is input to the image generation network 23, and the position guide information and the content guide information are also input to the image generation network 23 as two control conditions. The image generation network 23 simulates the diffusion process by combining the control of the position guide information and the content guide information, and gradually converts the noise information into the target extended image. For example, Figure 3As shown, the image generation network 23 can include a fourth encoding sub-network 231, a U-Net model 232, and a decoding sub-network 233 corresponding to the fourth encoding sub-network 231. In generating the target extended image, the fourth encoding sub-network 231 adds noise to an incomplete second-dimensional image, i.e., the central region of the incomplete second-dimensional image is mapped from the given first-dimensional image (i.e., the target original image) and the other regions are pure noise, and the thus-incomplete second-dimensional image after overall noise addition is taken as the input of the U-Net model 232. Meanwhile, the position guidance information and the content guidance information are both taken as two control conditions and input to the corresponding network layers of the U-Net model 232. Since the position encoding sub-network 211 and the content encoding sub-network 223 outputting the position guidance information and the content guidance information, and the U-Net model 232 are all based on multiple Transformer Blocks, and the Transformer Blocks in the position encoding sub-network 211 and the content encoding sub-network 223 have corresponding Transformer Blocks in the U-Net model 232, the position guidance information and the content guidance information can be input to the corresponding Transformer Blocks of the U-Net model 232 as two control conditions, so that each Transformer Block in the U-Net model 232 processes based on the input information to obtain the final output result of the U-Net model 232. The output result of the U-Net model 232 is input to the decoding sub-network 233, and the result after decoding by the decoding sub-network 233 is the finally generated target extended image.
[0099] In this embodiment, the position generation network 21 and the content generation network 22 are introduced by controlling the network, i.e., the position guidance information and the content guidance information are introduced simultaneously in the generation process of the target extended image, so as to realize precise control of the generation process of the extended image.
[0100] In some embodiments, the target extended image generated by the above-mentioned embodiments of the present application can be used in subsequent various scenarios. For example, the target extended image can be projected and converted at any field of view to obtain a perspective image under the field of view, and the perspective image can be studied or displayed. At this time, the quality of the target extended image directly affects whether the converted perspective image is distorted or can be normally used. Therefore, after the image generation model is used to generate the target extended image, a pre-trained feature extraction model can be further used to evaluate the quality of the generated target extended image.
[0101] Optionally, the feature extraction model of the embodiment at least includes an image encoder, and can further include a text encoder. The image encoder can extract image features of images of different image types (i.e., distorted extended image type, non-distorted extended image type, and original image type), and the text encoder can extract text features of description texts of different image types. The feature extraction model can minimize the similarity between any two image features corresponding to the same image type, and maximize the similarity between any two image features corresponding to different image types. In addition, the feature extraction model can minimize the similarity between any image feature and any text feature corresponding to the same image type, and maximize the similarity between any image feature and any text feature corresponding to different image types.
[0102] Optionally, the feature extraction model can be a trained multi-modal contrastive learning model (Contrastive Language-Image Pre-training, CLIP). The training process of the feature extraction model can be that the feature extraction model extracts sample features corresponding to the distorted extended image type, the non-distorted extended image type, and the original image type, respectively, and trains the feature extraction model based on the sample features, so that the similarity between sample features corresponding to the same image type and the similarity between sample features corresponding to different image types meet the training requirements. The sample features include sample image features or sample image features and sample text features. The specific training process will be described in detail in subsequent embodiments.
[0103] It should be noted that the non-distorted extended image type can be an image type to which an extended image having a real space mapping effect belongs, i.e., the perspective image converted from the extended image of this type does not have distortion. The distorted extended image type can be a type to which an extended image obtained by randomly distorting part of the non-distorted extended image belongs, i.e., the perspective image converted from the extended image of this type has distortion. The original image type can be an image type to which the target original image belongs. The purpose of the above embodiment is to generate an extended image of the non-distorted image type using the image generation model, so that the process of evaluating the quality of the target extended image can be to evaluate whether the generated target extended image belongs to the non-distorted extended image type.
[0104] The specific evaluation process can include the following three steps:
[0105] Step A, using the feature extraction model to extract the target image feature of the target extended image.
[0106] The embodiment can input the target extended image generated in the above embodiment into the trained feature extraction model, and the feature extraction model can perform image feature extraction on the target extended image to obtain target image features of the target extended image. Specifically, the target image features of the target extended image can be extracted by inputting the target extended image into an image encoder of the feature extraction model.
[0107] Step B, determining the first image features and / or the first text features corresponding to the non-distorted extended image type.
[0108] The first image features are obtained by extracting any image corresponding to the non-distorted extended image type by using the feature extraction model. For example, the embodiment can input an image of the non-distorted extended image type into the trained feature extraction model (such as an image encoder of the feature extraction model), and extract the image features of the input image by using the feature extraction model (such as the image encoder of the feature extraction model) as the first image features corresponding to the non-distorted extended image type. The first text features are obtained by extracting the description text corresponding to the non-distorted extended image type by using the feature extraction model; for example, the embodiment can input the description text corresponding to the non-distorted extended image type into the trained feature extraction model (such as a text encoder of the feature extraction model), and extract the text features of the input description text by using the feature extraction model (such as the text encoder of the feature extraction model) as the first text features corresponding to the non-distorted extended image type. The description text of the non-distorted extended image type can be “this is a non-distorted extended image”.
[0109] This step can only determine the first image features, only determine the first text features, or simultaneously determine the first image features and the first text features. Specifically, an image of the non-distorted extended image type can be optionally selected, and the first image features and / or the first text features of the image are extracted by using the trained feature extraction model. Alternatively, the first image features and / or the first text features corresponding to the non-distorted extended image type can be obtained by using the feature extraction model.
[0110] Step C, determining whether the target extended image meets the distortion requirement according to the first similarity between the target image features and the first image features and / or the second similarity between the target image features and the first text features.
[0111] Since the feature extraction model of the embodiment can minimize the similarity between any two image features corresponding to the same image type, if the target extended image belongs to the non-distorted extended image type, the similarity between the first image feature and the target image feature should be higher at this time. Therefore, one implementation manner of the step can be: determining whether the target extended image meets the distortion requirement according to the first similarity between the target image feature and the first image feature. Specifically, the similarity (such as cosine similarity) between the first image feature and the target image feature can be calculated as the first similarity, and it is determined whether the first similarity is greater than a first preset threshold. If yes, it is determined that the target extended image meets the distortion requirement, that is, the target extended image belongs to the non-distorted extended image type.
[0112] Since the feature extraction model of the embodiment can also minimize the similarity between any image feature and any text feature corresponding to the same image type, if the target extended image belongs to the non-distorted extended image type, the similarity between the first text feature and the target image feature should be higher at this time. Therefore, another implementation manner of the step can be: determining whether the target extended image meets the distortion requirement according to the second similarity between the target image feature and the first text feature. Specifically, the similarity (such as cosine similarity) between the target image feature and the first text feature can be calculated as the second similarity, and it is determined whether the second similarity is greater than a second preset threshold. If yes, it is determined that the target extended image meets the distortion requirement, that is, the target extended image belongs to the non-distorted extended image type. It should be noted that the first preset threshold and the second preset threshold can be the same or different.
[0113] On the basis of the above-mentioned embodiments, since the feature extraction model of the present embodiment minimizes the similarity between any two image features corresponding to the same image type while maximizing the similarity between any two image features corresponding to two different image types. Therefore, another implementable manner of the present step can also be: determining a second image feature corresponding to a reference image type; at this time, when determining whether the target expanded image meets the distortion requirement according to the first similarity, it can be that the target expanded image meets the distortion requirement is determined according to the first similarity between the target image feature and the first image feature, and the third similarity between the target image feature and the second image feature. Wherein, the reference image type includes: the distortion expanded image type and / or the original image type; the second image feature is extracted from any image corresponding to the reference image type by using the feature extraction model (such as the image encoder of the feature extraction model). Specifically, at least one of the distortion expanded image type and the original image type can be selected as the reference image type, and the second image feature of any image of the reference image type is extracted by using the trained feature extraction model (such as the image encoder of the feature extraction model), or the second image feature corresponding to the reference image type is obtained by the feature extraction model (such as the image encoder of the feature extraction model) generated in advance. And calculate the similarity (such as the cosine similarity) between the target image feature and the second image feature corresponding to the reference image type as the third similarity, judge whether the first similarity is greater than the third similarity, if yes, determine that the target expanded image meets the distortion requirement, that is, the target expanded image belongs to the non-distortion expanded image type. It should be noted that if the reference image type simultaneously selects the distortion expanded image type and the original image type, the present embodiment will determine one third similarity for the distortion expanded image type and the original image type respectively, at this time, the target expanded image can be determined to meet the distortion requirement when the first similarity is greater than the two third similarities. The present embodiment compares the target expanded image with the image features of different image types based on the pre-trained feature extraction model to determine the image type to which the target expanded image belongs, which can more accurately and flexibly realize the quality evaluation of the target expanded image.
[0114] Based on the above embodiment, since the feature extraction model can maximize the similarity between any image feature and any text feature corresponding to two different image types while minimizing the similarity between any image feature and any text feature corresponding to the same image type, another possible implementation of this step is to determine the second text feature corresponding to the reference image type; in this case, when determining whether the target extended image meets the distortion requirement based on the second similarity, it can be determined based on the second similarity between the target image feature and the first text feature, and the fourth similarity between the target image feature and the second text feature. Wherein, the reference image type includes: a distorted extended image type and / or an original image type; the second text feature is obtained by extracting the descriptive text corresponding to the reference image type using the feature extraction model (such as the text encoder of the feature extraction model); wherein, the descriptive text corresponding to the distorted extended image type can be "This is a distorted extended image"; the descriptive text corresponding to the original image type can be "This is an original image".
[0115] Specifically, at least one of the distorted extended image type and the original image type may be selected as the reference image type, and a trained feature extraction model (such as a text encoder of the feature extraction model) may be used to extract the second text features of the descriptive text of the reference image type, or the second text features corresponding to the reference image type generated in advance by the feature extraction model (such as a text encoder of the feature extraction model) may be obtained. A similarity (such as cosine similarity) between the target image features and the second text features corresponding to the reference image type is calculated as a fourth similarity, and a determination is made as to whether the second similarity is greater than the fourth similarity. If so, the target extended image is determined to meet the distortion requirement, i.e., the target extended image belongs to the non-distorted extended image type. It should be noted that if both the distorted extended image type and the original image type are selected as the reference image type, this embodiment will determine a fourth similarity for each of the distorted extended image type and the original image type. In this case, the target extended image is determined to meet the distortion requirement when the second similarity is greater than both fourth similarities. This embodiment is based on a pre-trained feature extraction model, and compares the image features of the target extended image with the text features of different image types to determine the image type to which the target extended image belongs. This can more accurately and flexibly achieve quality assessment of the target extended image.
[0116] It should be noted that if the target extended image is determined whether to meet the distortion requirement according to the first similarity between the target image feature and the first image feature and the second similarity between the target image feature and the first text feature at the same time, the way of determining whether the target extended image meets the distortion requirement according to the first similarity and the way of determining whether the target extended image meets the distortion requirement according to the second similarity can be combined to further improve the accuracy of the evaluation result. For example, it can be judged whether the first similarity is greater than the third similarity and whether the second similarity is greater than the fourth similarity at the same time. If the above two conditions are met at the same time, it is determined that the target extended image meets the distortion requirement, that is, the target extended image belongs to the non-distortion extended image type.
[0117] It should be noted that since the embodiment only needs to distinguish whether the target extended image meets the distortion requirement, for the description text of different image types, only the corresponding image type needs to be provided, and the detailed description of the image does not need to be provided.
[0118] In addition, the embodiment of determining whether the target extended image meets the distortion requirement means evaluating whether the target extended image has a real space mapping effect, that is, whether the perspective image obtained after converting the target extended image into a perspective image under any field of view exists distortion phenomenon.
[0119] The embodiment of the application based on the pre-trained feature extraction model to evaluate the quality of the target extended image generated by the image generation model, judge whether it meets the distortion requirement, if it meets, it means that the target extended image generated by the image generation model has a real space mapping effect, and there is no distortion problem when it is mapped into a perspective image, which belongs to the non-distortion extended image type. If it does not meet, it means that the target extended image generated by the image generation model has not very good space mapping effect, if it is mapped into a perspective image, there may be distortion problem. The image generation model needs to be further trained. So as to ensure the quality of the generated target extended image.
[0120] Figure 3 The flowchart of one embodiment of the model training method provided in the application, the technical solution of the embodiment can be applied to the server. The image generation model introduced in the above embodiment is trained by the server. Figure 2 The model training method shown can include the following steps:
[0121] S301, acquiring second mapping information.
[0122] The second mapping information is obtained based on a third sample extended image; and the second mapping information comprises a mapping relationship between a pixel point of the third sample extended image in a first dimension space and a position point in a second dimension space. The third sample extended image has the same size as the second sample extended image and the first sample extended image. The third sample extended image and the second sample extended image can have the same content or different content. The third sample image can be a second sample extended image in a plurality of training sample data or can not be.
[0123] Optionally, the embodiment can determine whether the corresponding second mapping information has been constructed in advance according to the size of the target extended image generated by the image generation model to be trained. If the second mapping information has been constructed, the second mapping information corresponding to the size can be directly obtained. If the second mapping information has not been constructed, any third sample extended image of the size can be obtained, and the second mapping information is constructed based on the third sample extended image. The construction method has been introduced in the above embodiment, and will not be described here.
[0124] S302, obtaining a second sample extended image, and a sample original image and a sample mask extended image corresponding to the second sample extended image.
[0125] The first sample extended image and the second sample extended image are both non-distorted extended image type images, and have the same image size and can have the same or different image content.
[0126] The embodiment can obtain a plurality of non-distorted extended image type images as the second sample extended image. Then, for any one of the obtained second sample extended images, the corresponding sample original image and sample mask extended image are further determined. Specifically, for any one of the second sample extended images, the corresponding sample original image can be obtained by converting the image of the central region of the second sample extended image into a two-dimensional image of a normal field angle based on a projection processing algorithm (such as an equirectangular projection algorithm), that is, obtaining the sample original image corresponding to the second sample extended image.
[0127] The sample mask extended image corresponding to any one of the second sample extended images of the embodiment can include two parts of a partial extended image and an outer drawing area mask image. The partial extended image can be obtained by retaining the image content of the center area of the second sample extended image, using the first mask algorithm to perform mask processing on other areas except the center area, and obtaining the partial extended image. The outer drawing area mask image can be obtained by using the second mask algorithm to perform mask processing on the center area of the partial extended image. Another obtaining method of the sample mask extended image can also be based on the sample original image. The specific implementation process is similar to the method of generating the target mask extended image corresponding to the target original image in the above-mentioned image generation method embodiment, and will not be described here.
[0128] In S303, the image generation model is used to generate a predicted extended image based on the second mapping information, in combination with the second sample extended image, the sample original image, and the sample mask extended image.
[0129] Optionally, the embodiment can be used to input the second mapping information, the sample original image, and the sample mask extended image as control condition, guide the image generation model to simulate the diffusion process based on the second sample extended image, learn how to generate accurate position guide information and content guide information, and how to convert noise into a target extended image based on the position guide information and the content guide information. It should be noted that the process of generating a predicted extended image based on the second mapping information, in combination with the second sample extended image, the sample original image, and the sample mask extended image by using the image generation model is similar to the process of generating a target extended image based on the first mapping information, in combination with the target original image and the target mask extended image by using the image generation model as described in the above embodiment, and the difference lies in that Figure 2 As shown in the figure, the fourth encoding sub-network 231 in the training stage is to add noise information (such as Gaussian noise) on the second sample extended image, and the fourth encoding sub-network 231 in the inference stage is to add noise on the incomplete second-dimensional image, that is, the center area of the incomplete second-dimensional image is mapped from the given first-dimensional image (i.e., the target original image) and the other areas are pure noise, and the overall noise is added on such an incomplete second-dimensional image. That is, for the fourth encoding sub-network 231 in the inference stage, the center area of the second-dimensional image is given by the first-dimensional image (i.e., the target original image), and the other areas are pure noise. Figure 4The image generation model shown, when training the model, similar to the process of generating the target extended image, the input of the first encoding sub-network 212 is the second mapping information, the input of the second encoding sub-network 221 is the sample mask extended image, and the input of the third encoding sub-network 222 is the sample original image. The difference is that the fourth encoding sub-network 231 is an incomplete second-dimensional image in the target extended image generation process, and adds noise to the incomplete second-dimensional image and inputs it to the position encoding sub-network 211, the content encoding sub-network 223 and the U-Net model 232 composed of multiple Transformer Blocks. In the model training phase, the input of the fourth encoding sub-network 231 is the second sample extended image, and the second sample extended image with added noise information is input to the position encoding sub-network 211, the content encoding sub-network 223 and the U-Net model 232. The process of how each network in the image generation model cooperates to generate the predicted extended image is similar to the process of generating the target extended image, and will not be described here.
[0130] S304, training the image generation model according to the difference information between the predicted extended image and the second sample extended image.
[0131] Optionally, the embodiment can determine the reconstruction loss based on the difference information between the predicted extended image and the second sample extended image, and train the image generation model based on the reconstruction loss.
[0132] It should be noted that the image generation model trained in the embodiment is used to implement the above-mentioned embodiment of generating a target extended image based on first mapping information, in combination with a target original image and a target mask extended image corresponding to the target original image.
[0133] The scheme of the embodiment, when training the image generation model, obtains the second mapping information, the second sample extended image, and the sample original image and the sample mask extended image corresponding to the second sample extended image, and uses the obtained data information to guide the image generation model to learn the mapping relationship between the pixel points in the first-dimensional space and the position points in the second-dimensional space with the second mapping information as the control condition, so as to guide the image generation model to learn how to generate a target extended image with a real spatial mapping effect, avoid image distortion during mapping conversion, and further guide the image generation model to learn the target original image and the target mask extended image corresponding thereto as the control condition of the extended content, so as to guide the image generation model to learn how to accurately extend the image content, greatly improve the model training precision, and provide a guarantee for subsequent low-cost, high-quality and efficient generation of target extended images.
[0134] In some embodiments, in order to further improve the accuracy of model training, the embodiment can introduce the feature extraction model introduced in the above embodiment to optimize the training loss in the process of training the image generation model. That is, the above S304 step can specifically include: extracting the predicted image features of the predicted extended image by using the feature extraction model; determining the fifth similarity between the predicted image features and the sample text features corresponding to the distorted extended image type, determining the sixth similarity between the predicted image features and the sample text features corresponding to the non-distorted extended image type, and determining the seventh similarity between the predicted image features and the sample text features corresponding to the original image type; determining the distortion loss according to the fifth similarity, the sixth similarity and the seventh similarity; determining the reconstruction loss according to the difference information between the predicted extended image and the second sample extended image; and training the image generation model according to the distortion loss and the reconstruction loss.
[0135] It should be noted that the sample text features corresponding to the distorted extended image type, the non-distorted extended image type and the original image type are respectively extracted by using the feature extraction model on the description text corresponding to the distorted extended image type, the non-distorted extended image type and the original image type. The specific extraction method has been described in the above embodiment, and will not be described here. In addition, the way of extracting the predicted image features of the predicted extended image by using the feature extraction model is similar to the way of extracting the target image features of the target extended image by using the feature extraction model as described in the above embodiment, and the way of determining the fifth similarity, the sixth similarity and the seventh similarity is similar to the way of determining the third similarity and the fourth similarity as described in the above embodiment, which will not be described here.
[0136] After the fifth similarity, the sixth similarity and the seventh similarity are determined, the distortion loss can be determined based on the following formula (9). The reconstruction loss is determined based on the difference information between the predicted extended image and the second sample extended image. Then, the final loss is constructed according to the following formula (10) to train the image generation model.
[0137]
[0138] wherein, the final loss is L, the reconstruction loss is Lr, the distortion loss is Ld; x is the predicted image feature; the sample text feature corresponding to the non-distorted extended image type is x2; the sample text feature corresponding to the original image type is x3; the sample text feature corresponding to the non-distorted extended image type is x2, the sixth similarity is s6; the seventh similarity is s7, wherein, the fifth similarity is a similarity between the first image feature and the second image feature; and the parameter variable is a parameter variable of the fifth similarity.
[0139] In the training of the image generation model, the reconstruction loss and the distortion loss are introduced, which greatly improves the accuracy of model training and ensures that the trained image generation model can generate target expanded images with more realistic spatial mapping effects.
[0140] Figure 4 The technical scheme of the embodiment of the model training method provided in the application can be applied to a server. The feature extraction model introduced in the above embodiment is trained by the server. Figure 5 The model training method shown can include the following steps:
[0141] S401, acquiring training samples corresponding to a distortion expanded image type, a non-distortion expanded image type, and an original image type respectively.
[0142] The training samples include sample training images or sample training images and description texts.
[0143] This step can be acquiring training samples corresponding to the distortion expanded image type, the non-distortion expanded image type, and the original image type respectively. It should be noted that if the training target of training the feature extraction model in this embodiment is to minimize the similarity between any two image features corresponding to the same image type and to maximize the similarity between any two image features corresponding to two different image types (i.e., training target one), then the image encoder of the feature extraction model can be trained at this time, and therefore, the training samples acquired can be multiple sample training images corresponding to the distortion expanded image type, the non-distortion expanded image type, and the original image type respectively.
[0144] If the training target of training the feature extraction model in this embodiment is to minimize the similarity between any image feature and any text feature corresponding to the same image type and to maximize the similarity between any image feature and any text feature corresponding to two different image types (i.e., training target two), or the above training target one and training target two are the final training targets, then the image encoder and the text encoder of the feature extraction model need to be trained at this time, and therefore, the training samples acquired need to include not only multiple sample training images corresponding to the distortion expanded image type, the non-distortion expanded image type, and the original image type respectively, but also description samples corresponding to the distortion expanded image type, the non-distortion expanded image type, and the original image type respectively.
[0145] Optionally, when the sample training image in the training sample is obtained in this embodiment, the sample training image corresponding to the non-distorted extended image type can be obtained first, and then the sample training image corresponding to the non-distorted extended image type is randomly distorted to obtain the sample training image corresponding to the distorted extended image type. For the sample training image corresponding to the original image type, it can be directly drawn or photographed, or it can be obtained by projecting the sample training image corresponding to the non-distorted extended image type.
[0146] S402, the sample features corresponding to the training samples of different image types are extracted by using the feature extraction model.
[0147] It should be noted that in this embodiment, the sample image features and the sample text features corresponding to the same image type are aligned with each other, so that the similarity calculation can be performed.
[0148] Optionally, if the training sample obtained in S401 includes only multiple sample training images corresponding to the three image types (i.e., the distorted extended image type, the non-distorted extended image type, and the original image type) respectively, then at this time, the sample image features corresponding to each sample training image need to be extracted by using the feature extraction model (such as the image encoder of the feature extraction model) for the multiple sample training images corresponding to the three image types in the training sample, as the sample features corresponding to the training samples of different image types respectively.
[0149] If the training sample obtained in S401 includes sample training images and description texts, then at this time, not only the sample image features corresponding to each sample training image need to be extracted by using the feature extraction model (such as the image encoder of the feature extraction model) for the multiple sample training images corresponding to the three image types in the training sample, but also the sample text features corresponding to the three image types need to be extracted by using the feature extraction model (such as the text encoder of the feature extraction model) for the description texts corresponding to the three image types in the training sample. That is, the sample features corresponding to each image type not only include the sample image features, but also include the sample text features.
[0150] It should be noted that for each of the three image types, the image content of the second sample extended image corresponding to each image type is different, so the sample image features corresponding to each image type are different. However, the description texts corresponding to each image type are the same, so the sample text features corresponding to each image type are the same.
[0151] S403, training the feature extraction model based on the sample features, the similarity between the sample features corresponding to the same image type, and the similarity between the sample features corresponding to different image types satisfying the training requirements as the training target.
[0152] Optionally, if the sample features are sample image features, then at this time, sample training images of the same image type may be used as positive sample pairs, and sample training images of different image types may be used as negative sample pairs. By using a contrastive learning method, the feature similarity between sample features corresponding to the same image type (i.e., sample image features) is optimized, while the feature similarity between sample features corresponding to different image types (i.e., sample image features) is expanded to train the feature extraction model (i.e., train the image encoder of the feature extraction model). At this time, the feature extraction model (i.e., train the image encoder of the feature extraction model) may be trained by minimizing the feature similarity between sample features corresponding to the same image type (i.e., sample image features) and maximizing the feature similarity between sample features corresponding to different image types (i.e., sample image features); or the feature extraction model (i.e., train the image encoder of the feature extraction model) may be trained by taking the feature similarity between sample features corresponding to the same image type (i.e., sample image features) as the training goal to be greater than the feature similarity between sample features corresponding to different image types (i.e., sample image features).
[0153] If the sample features are sample image features and sample text features, one situation is: taking sample training images and description texts of the same image type as positive sample pairs, and sample training images and description texts of different image types as negative sample pairs, using contrastive learning to optimize the feature similarity between sample features corresponding to the same image type (i.e., sample image features and sample text features), and at the same time expand the feature similarity between sample features corresponding to different image types (i.e., sample image features and sample text features), to train the feature extraction model (i.e., train the text encoder and image encoder of the feature extraction model). At this time, the feature extraction model (i.e., the text encoder and image encoder of the feature extraction model) can be trained by minimizing the feature similarity between sample features corresponding to the same image type (i.e., sample image features and sample text features) and maximizing the feature similarity between sample features corresponding to different image types (i.e., sample image features and sample text features) as the training objectives; or the feature extraction model (i.e., the text encoder and image encoder of the feature extraction model) can be trained with the feature similarity between sample features corresponding to the same image type (i.e., sample image features and sample text features) being greater than the feature similarity between sample features corresponding to different image types (i.e., sample image features and sample text features) as the training objectives.
[0154] In another case, if the sample features include sample image features and sample text features, the feature extraction model is trained based on the sample image features and sample text features, similarity between sample image features corresponding to the same image type, and similarity between sample image features corresponding to different image types satisfying a first training requirement as a first training target, and similarity between sample text features and sample image features corresponding to the same image type, and similarity between sample text features and sample image features corresponding to different image types satisfying a second training requirement as a second training target. This case can be a combination of the above two cases, i.e., training the text encoder and the image encoder of the feature extraction model in a contrast learning manner by taking two sample training images of the same image type and the sample training image and the description text of the same image type as positive sample pairs, and taking two sample training images of different image types and the sample training image and the description text of different image types as negative sample pairs. The advantage of this setting is to further improve the accuracy of model training.
[0155] As shown in FIG. 5, Figure 6 The feature extraction model 50 includes an image encoder 501 and a text encoder 502. P1 and P2 in the figure are sample image features of the non-distorted extended image type extracted by the image encoder 501; TP1 and TP2 are sample text features of the non-distorted extended image type extracted by the text encoder 502; N1 and N2 are sample image features of the original image type extracted by the image encoder 501; TN1 and TN2 are sample text features of the original image type extracted by the text encoder 502; R1 and R2 are sample image features of the distorted extended image type extracted by the image encoder 501; and TR1 and TR2 are sample text features of the distorted extended image type extracted by the text encoder 502. This embodiment can calculate the similarity for any two different sample image features, such as P1P2. The similarity between any sample image feature and sample text feature is also calculated, such as P1TP1.
[0156] Then, the loss function of the image encoder 501 is designed according to the following formula (11)
[0157]
[0158] wherein, is the training loss of the image encoder 501; x and y are two different sample image features; N is the total number of sample image features; x i is the i-th sample image feature; y j is the j-th sample image feature; l ijan index label representing whether the ith sample image feature and the jth sample image feature belong to the same image type, if yes, l ij = 1, otherwise l ij = 0.
[0159] The loss function of the text encoder 502 is designed according to the following formula (12)
[0160]
[0161] wherein, is the training loss of the text encoder 502; x and z represent sample image features and sample text features respectively; N is the total number of sample image features and sample text features; it should be noted that for each sample training image, there is a corresponding sample text feature, so the number of sample image features and sample text features is the same, that is, both are N. i is the ith sample image feature; z j is the jth sample text feature; l ij an index label representing whether the ith sample image feature and the jth sample image feature belong to the same image type, if yes, l ij = 1, otherwise l ij = 0.
[0162] The embodiment can combine formulas (11) and (12) to determine the overall loss function of the feature extraction model 50 that is, the following formula (13).
[0163]
[0164] Because in the embodiment, the distortion types of the sample image features and the sample text features are mutually aligned. In addition, in order to identify the different roles of the image encoder and the text encoder, the gradient flowing to the image encoder can be cut off when adjusting the text encoder. That is, when adjusting the parameters of the image encoder, the parameters of the text encoder are not adjusted, and when adjusting the parameters of the text encoder, the parameters of the image encoder are not adjusted. Thus, each encoder can fully exert its potential.
[0165] It should be noted that the feature extraction model trained in this embodiment can be used to evaluate the quality of the target extended image generated by the image generation model, that is, to extract the target image features of the target extended image, and to extract the first image features and / or first text features corresponding to the non-distorted extended image type; the target image features and the first image features and / or first text features are used to determine whether the target extended image meets the distortion requirement; the target extended image is generated by using the image generation model, based on the first mapping information, in combination with the target original image and the target mask image corresponding to the target original image; the first mapping information is obtained based on the first sample extended image; the first mapping information includes the mapping relationship between the pixel points of the first sample extended image in the first dimensional space and the position points in the second dimensional space. The specific implementation manner has been described in the above embodiment, and will not be described here.
[0166] In addition, the feature extraction model trained in this embodiment can also be used to determine the distortion loss of the image generation model when training the image generation model. The specific application manner has been described in the above embodiment and will not be described here.
[0167] In this embodiment, a feature extraction model capable of optimizing the similarity between sample features with the same image type and expanding the difference between sample features of different image types is trained in a contrast learning manner, which provides a guarantee for subsequent quality evaluation of the target extended image generated by the image generation model and improvement of the training precision of the image generation model.
[0168] In some embodiments, the extended image in the embodiments of the present application can be a panoramic image. Next, this embodiment is directed to the actual application scenario of generating a panoramic image, and the panoramic image generation method shown in Figure 7 The image generation model in the actual application is introduced. Figure 6 The panoramic image generation method shown in the method can include the following steps:
[0169] S701, obtaining a target original image and first mapping information.
[0170] The first mapping information is obtained based on a first sample panoramic image; and the first mapping information includes a mapping relationship between pixel points of the first sample panoramic image in a two-dimensional space and position points in a spherical coordinate system.
[0171] The target original image of this embodiment can be a two-dimensional image taken in a real scene, or a target perspective image. The perspective image is a drawing method based on the perspective principle, which is used to represent the three-dimensional spherical space relationship of a three-dimensional object on a two-dimensional plane. The acquisition method of the target original image and the first mapping information can be referred to the description in the corresponding embodiments described above, and will not be described here.
[0172] S702, generate a target mask panorama image corresponding to the target original image.
[0173] Optionally, the embodiment can be inverse processing of equirectangular projection on the target original image (i.e. Figure 6 , to obtain the target mask panorama image.
[0174] S703, using an image generation model, generating a target panorama image based on the first mapping information, combining the target original image and the target mask panorama image.
[0175] Wherein, the image generation model can be trained based on a plurality of training sample data; wherein each training sample data can include a sample original image and second mapping information, and a second sample panorama image corresponding to the sample original image. Optionally, when training the image generation model based on each training sample data, a corresponding sample mask panorama image can be generated based on the sample original image in each training sample data. A sample mask panorama image corresponding to the sample original image can also be generated based on the second sample panorama image in each training sample data. Optionally, the sample original image can be pre-set corresponding to the second sample panorama image, of course, it can also be generated based on the second sample panorama image when training the image generation model. Then, the image generation model can be trained according to the second mapping information, the second sample panorama image, the sample original image, and the sample mask panorama image corresponding to the sample original image. The specific training process has been introduced in the above embodiment, which will not be repeated here.
[0176] At this time, the overall idea of this step is to obtain the first mapping information, and to guide the model to perceive the relationship between the corresponding points in the two-dimensional space and the three-dimensional spherical space based on the first mapping information, so as to guide the image generation model to generate a target panorama image with real space mapping effect in the form of equirectangular projection.
[0177] Specifically, as Figure 8 shown, for the first mapping information, the embodiment can perform position condition injection in each position encoding block in the plurality of position encoding blocks. That is, the formula for position condition injection of the backbone network U-NET of the image generation network is as follows
[0178]
[0179] Wherein is the position condition injection; is random noise; c t is the description text "this is a distorted image"; cd is the first mapping information; DE() is the first encoding layer (Distort Encoder) corresponding function. de is the first encoding layer (Distort Encoder) processing result; Proj b () is the first embedding layer (Distort Embedding) corresponding function; Z is the parameter variable; DN() is the multiple position encoding block (DistortNet) corresponding function; dn b is the bth position encoding block processing result.
[0180] For the target original image, it is converted into a feature vector through the second encoding layer, i.e., the variational autoencoder (VAE Encoder), and then the dimension of the feature vector is aligned with the text input dimension of the multiple content encoding blocks through the mapping layer (Projection) to obtain the second encoding feature cn', i.e., the following formula (15)
[0181] cn' = Proj(VAE(c n )) (15)
[0182] Wherein, cn' is the second encoding feature; Proj() is the mapping layer (Projection) corresponding function; VAE() is the variational autoencoder (VAE Encoder) corresponding function; c n is the target original image.
[0183] On this basis, the embodiment can inject the content condition into the backbone network U-NET of the image generation network by using the following formula (16).
[0184]
[0185] Wherein, is the content condition injection; is random noise; cn' is the second encoding feature; c p and are the partial panorama and outer drawing area mask maps in the target mask panorama image respectively; cn 0 is the output result of the first content encoding block; cn b is the output result of the bth content encoding block; CN() is the multiple content encoding block (Content Net) corresponding function; CE is the content encoding block (Content Encoder) corresponding function corresponding to the first encoding subnetwork of the above embodiment; B is the number of content encoding blocks.
[0186] When the backbone network U-Net is added with the position condition injection and the content condition injection, the total output flow of the U-Net is defined as formula (17), and out is the target panoramic image finally output by the image generation model.
[0187]
[0188] wherein out is the target panoramic image output by the variational auto-encoder (VAE Encoder) corresponding to the variational auto-decoder (VAE Decoder); is the output result of the backbone network U-Net; is the content condition injection; is the position condition injection; and Z is a parameter variable.
[0189] Optionally, the application scenarios of the target panoramic image generated by the embodiment are various. For example, the target panoramic image can be sent to a user end for display. Since the target panoramic image can show a complete scene of 360 degrees in the horizontal direction and 180 degrees in the vertical direction, a user can control the direction of browsing the target panoramic image on the user end through a mouse or other interactive methods, as if watching the object or scene in person. When the user has a perspective image viewing requirement for a certain field of view, the user end can also send a perspective image viewing request for the field of view to the server end. At this time, the server end can generate a perspective image for the field of view based on the equirectangular projection algorithm, and send the perspective image to the user end for display. Since the image generation model in the embodiment combines the first mapping information when generating the target panoramic image, the target panoramic image has a real spatial mapping effect, and distortion can be avoided when converting to a perspective image under any field of view, thereby ensuring a better image browsing experience for the user.
[0190] It should be noted that after the target panoramic image is generated, the feature extraction model trained in the above embodiment can be further used to evaluate whether the target panoramic image meets the distortion requirement. This process is similar to the process of evaluating whether the target panoramic image meets the distortion requirement described in the above embodiment, and will not be described in detail here.
[0191] The detailed implementation and beneficial effects of each step in the embodiment have been described in detail in the foregoing embodiments, and will not be described in detail here.
[0192] It should be noted that in some of the processes described in the above embodiments and accompanying drawings, a plurality of operations are included in a specific order, but it should be clearly understood that these operations can be executed or in parallel, not in the order in which they appear in this article. The serial numbers of the operations such as 301, 302, etc. are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and the operations can be executed in sequence or in parallel. It should be noted that the "first", "second" and the like described herein are used to distinguish different messages, devices, modules, etc., and do not represent the order or limit the types of "first" and "second".
[0193] Figure 8 A structural diagram of an image generation device provided for an exemplary embodiment of the present application is shown. The device includes:
[0194] The first acquisition module 801 is configured to acquire a target original image and first mapping information. The first mapping information is obtained based on a first sample extended image. The first mapping information includes a mapping relationship between a pixel point of the first sample extended image in a first dimensional space and a position point in a second dimensional space.
[0195] The mask processing module 802 is configured to generate a target mask extended image corresponding to the target original image.
[0196] The image generation module 803 is configured to generate a target extended image based on the first mapping information, in combination with the target original image and the target mask extended image, by using an image generation model.
[0197] The image generation model is trained based on a plurality of training sample data. Each training sample data includes a sample original image and second mapping information, and a second sample extended image corresponding to the sample original image.
[0198] In an optional embodiment, the apparatus further includes an image evaluation module configured to extract a target image feature of the target extended image by using a feature extraction model; determine a first image feature and / or a first text feature corresponding to a non-distorted extended image type; and determine whether the target extended image meets a distortion requirement according to a first similarity between the target image feature and the first image feature and / or a second similarity between the target image feature and the first text feature; wherein the first image feature is extracted from any image corresponding to the non-distorted extended image type by using the feature extraction model; the first text feature is extracted from a description text corresponding to the non-distorted extended image type by using the feature extraction model; the feature extraction model extracts sample features corresponding to the distorted extended image type, the non-distorted extended image type, and the original image type respectively, and trains the feature extraction model based on the sample features, similarities between sample features corresponding to the same image type, and similarities between sample features corresponding to different image types meeting a training requirement as a training target; and the sample features include sample image features or sample image features and sample text features.
[0199] In some embodiments, the apparatus further includes a feature determination module configured to determine a second image feature corresponding to a reference image type; the reference image type includes the distorted extended image type and / or the original image type; and the second image feature is extracted from any image corresponding to the reference image type by using the feature extraction model.
[0200] The image evaluation module is specifically configured to determine whether the target extended image meets the distortion requirement according to the first similarity between the target image feature and the first image feature and a third similarity between the target image feature and the second image feature.
[0201] In some embodiments, the feature determination module is further configured to determine a second text feature corresponding to a reference image type; the reference image type includes the distorted extended image type and / or the original image type; and the second text feature is extracted from a description text corresponding to the reference image type by using the feature extraction model.
[0202] The image evaluation module is specifically configured to determine whether the target extended image meets the distortion requirement according to the second similarity between the target image feature and the first text feature and a fourth similarity between the target image feature and the second text feature.
[0203] In some embodiments, the image generation model includes a position generation network, a content generation network, and an image generation network.
[0204] The image generation module 803 is specifically configured to determine position guidance information based on the first mapping information by using the position generation network; determine content guidance information based on the target original image and the target mask extended image by using the content generation network; and generate a target extended image by combining the position guidance information and the content guidance information by using the image generation network.
[0205] In some embodiments, the position generation network comprises a position encoding subnetwork, and the position encoding subnetwork comprises a plurality of position encoding blocks.
[0206] The image generation module 803 is specifically configured to input noise information to a first position encoding block of the position encoding subnetwork, and input the first mapping information to the plurality of position encoding blocks respectively, to obtain position guidance information output by the position encoding subnetwork.
[0207] In some embodiments, the position generation network further comprises a first encoding subnetwork.
[0208] The image generation module 803 is specifically configured to extract mapping features of the first mapping information by using the first encoding subnetwork, and input the mapping features to the plurality of position encoding blocks respectively.
[0209] In some embodiments, the content generation network comprises a second encoding subnetwork, a third encoding subnetwork, and a content encoding subnetwork, and the content encoding subnetwork comprises a plurality of content encoding blocks.
[0210] The image generation module 803 is specifically configured to extract first encoding features of the target mask extended image by using the second encoding subnetwork; extract second encoding features of the target original image by using the third encoding subnetwork; the dimension of the second encoding features is aligned with the text input dimension of the plurality of content encoding blocks; input the first encoding features and noise information to a first content encoding block of the content encoding subnetwork, and input the second encoding features as text inputs of the plurality of content encoding blocks, to obtain content guidance information output by the plurality of content encoding blocks.
[0211] In some embodiments, the first acquisition module 801 has a mapping relationship construction module configured to determine a mapping relationship between a pixel point of the first sample extended image in a first dimensional space and a position point in a second dimensional space; perform spatial edge alignment processing on the mapping relationship, and obtain the first mapping information based on the mapping relationship after the spatial edge alignment processing.
[0212] In some embodiments, when the target extended image is a target panorama image, the information obtaining module 801 is configured to obtain a target original image and first mapping information; the first mapping information is obtained based on a first sample panorama image; the first mapping information comprises a mapping relationship between a pixel point of the first sample panorama image in a two-dimensional space and a position point in a spherical coordinate system;
[0213] The mask processing module 802 is configured to generate a target mask panorama image corresponding to the target original image.
[0214] The image generation module 803 is configured to generate a target panorama image based on the first mapping information, the target original image, and the target mask panorama image by using an image generation model.
[0215] The image generation model is trained based on a plurality of training sample data; each training sample data comprises a sample original image and second mapping information, and a second sample panorama image corresponding to the sample original image.
[0216] Figure 1 The image generation device can perform the image generation method of the embodiments shown in Figure 6 and Figure 9 The implementation principle and technical effects of the image generation method of the embodiments shown in the above are not described again. The specific manner in which each module and unit in the above embodiments performs operations has been described in detail in the embodiments related to the method, and will not be described in detail here.
[0217] Figure 9 A structural schematic diagram of a model training device provided by an exemplary embodiment of the present application is shown in the figure. The device comprises:
[0218] The second obtaining module 901 is configured to obtain second mapping information; the second mapping information is obtained based on a third sample extended image; the second mapping information comprises a mapping relationship between a pixel point of the third sample extended image in a first dimensional space and a position point in a second dimensional space.
[0219] The second obtaining module 901 is further configured to obtain a second sample extended image, and a sample original image and a sample mask extended image corresponding to the second sample extended image.
[0220] The first model training module 902 is configured to generate a predicted extended image by using an image generation model, based on the second mapping information, in combination with the second sample extended image, the sample original image, and the sample mask extended image; and train the image generation model according to difference information between the predicted extended image and the second sample extended image.
[0221] The image generation model is configured to generate a target extended image based on first mapping information, in combination with a target original image and a target mask extended image corresponding to the target original image.
[0222] In an optional embodiment, the first model training module 902 is specifically configured to extract, by using a feature extraction model, a predicted image feature of the predicted extended image; determine a fifth similarity between the predicted image feature and a sample text feature corresponding to a distorted extended image type, a sixth similarity between the predicted image feature and a sample text feature corresponding to a non-distorted extended image type, and a seventh similarity between the predicted image feature and a sample text feature corresponding to an original image type; determine a distortion loss based on the fifth similarity, the sixth similarity, and the seventh similarity; determine a reconstruction loss based on difference information between the predicted extended image and the second sample extended image; and train the image generation model based on the distortion loss and the reconstruction loss. The sample text features corresponding to the distorted extended image type, the non-distorted extended image type, and the original image type are respectively obtained by using the feature extraction model to extract description texts corresponding to the distorted extended image type, the non-distorted extended image type, and the original image type.
[0223] Figure 3 The model training device can perform Figure 10 The model training method of the embodiments described above has the same implementation principles and technical effects. The specific operation modes of each module and unit in the device described in the above embodiments have been described in detail in the embodiments related to the method, and will not be described in detail here.
[0224] Figure 10 Another model training device provided by the exemplary embodiments of the present application is provided. The device includes:
[0225] The third acquisition module 1001 is configured to acquire training samples corresponding to a distorted extended image type, a non-distorted extended image type, and an original image type. The training samples include sample training images, or sample training images and description texts.
[0226] The second model training module 1002 is configured to extract, by using a feature extraction model, sample features corresponding to different image types of the training samples. The sample features include sample image features, or sample image features and sample text features. The feature extraction model is trained based on the sample features, similarity between sample features corresponding to the same image type, and similarity between sample features corresponding to different image types, which satisfy a training requirement as a training target.
[0227] The feature extraction model is used to extract a target image feature of a target extended image, and extract a first image feature and / or a first text feature corresponding to a non-distorted extended image type; the target image feature and the first image feature and / or the first text feature are used to judge whether the target extended image meets a distortion requirement; the target extended image is generated by using an image generation model, based on first mapping information, in combination with a target original image and a target mask extended image corresponding to the target original image; the first mapping information is obtained based on a first sample extended image; and the first mapping information includes a mapping relationship between a pixel point of the first sample extended image in a first dimensional space and a position point in a second dimensional space.
[0228] In an optional embodiment, the second model training module 1002 is specifically configured to, if the sample feature includes a sample image feature and a sample text feature, train the feature extraction model based on the sample image feature and the sample text feature, similarity between sample image features corresponding to a same image type, and similarity between sample image features corresponding to different image types satisfying a first training requirement as a first training target, similarity between sample text features and sample image features corresponding to the same image type, and similarity between sample text features and sample image features corresponding to different image types satisfying a second training requirement as a second training target.
[0229] Figure 4 The model training apparatus can perform Figure 11 The model training method of the embodiments shown above has the implementation principles and technical effects which will not be described again. The specific operation modes of each module and unit in the 10 apparatus in the above embodiments have been described in detail in the embodiments related to the method, and will not be described in detail here.
[0230] Figure 11 An embodiment of a structure schematic diagram of a computing device provided in the present application is shown. As shown in the figure, Figure 1 In practice, the computing device can include a storage component 1101 and a processing component 1102.
[0231] The storage component 1101 is configured to store computer programs and can be configured to store other various data to support operations on the computing device. Examples of these data include instructions of any application program or method for operating on the computing device, data structures, contact data, phonebook data, messages, pictures, videos, etc.
[0232] The processing component 1102 is coupled to the storage component 1101 and is configured to execute the computer programs in the storage component 1101 for implementing the Figure 7 or Figure 3The image generation method shown in FIG. Or it can be used to implement Figure 4 or Figure 11 The model training method shown.
[0233] Further, if Figure 11 As shown, the computing device may further include: a communication component 1103, a display component 1104, a power component 1105, an audio component 1106 and other components. Figure 11 Only some components are shown schematically, which does not mean that the equipment only includes Figure 11 In addition, Figure 11 The components in the dotted box are optional components, not mandatory components, and depend on the specific product form of the computing device. The computing device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT (Internet of Things) device, or a server device such as a conventional server, a cloud server or a server array. If the computing device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it can include Figure 11 If the computing device of this embodiment is implemented as a server device such as a conventional server, a cloud server or a server array, it may not include the components in the dotted box; Components within the dotted box.
[0234] The processing component includes one or more processors to execute computer instructions to perform all or part of the steps in the above method. Of course, the processing component can also be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above method.
[0235] The above-mentioned storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0236] The communication component is configured to facilitate wired or wireless communication between the device on which the communication component is installed and other devices. The device on which the communication component is installed can access a wireless network based on a communication standard, such as a radio communication technology, a wireless local area network (WLAN) technology, a Bluetooth (BT) technology, a near field communication (NFC) technology, and / or a global positioning system (GPS) technology. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast managing system via a broadcast channel.
[0237] The display component can include a screen, which can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action.
[0238] The power supply component is configured to supply power to various components of the device on which the power supply component is installed. The power supply component can include a power management system, one or more power sources, and other components associated with generating, managing and distributing power for the device on which the power supply component is installed.
[0239] The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive an external audio signal when the device on which the audio component is installed is in a mode, such as a call mode, a recording mode and a voice recognition mode. The received audio signal can be further stored in a memory or transmitted via the communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0240] Accordingly, the embodiments of the present application also provide a computer readable storage medium storing a computer program, when the computer program is executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. The computer readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of the computer readable storage medium include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital video disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium
[0241] Accordingly, the embodiments of the present application also provide a computer program product, the computer program product includes a computer program or instructions, when the computer program or instructions are executed by a processor, the processor is enabled to implement each step in the above-mentioned method embodiments. It should be understood that each process or a combination of multiple processes in the above-mentioned method flow can be implemented by the computer program or instructions. In addition, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable data processing devices can be implemented as a device for implementing the corresponding functions in the above-mentioned method embodiments.
[0242] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-mentioned system, device and unit can refer to the corresponding process in the above-mentioned method embodiments, which will not be described here.
[0243] It should also be noted that the terms "comprising," "including," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0244] Finally, it should be noted that only the preferred embodiments of the application have been described, and that all modifications and changes, which come within the spirit of the application, are therefore intended to be protected.
Claims
1. An image generation method, characterized in that: include: Acquire a target original image and first mapping information; The first mapping information is constructed based on the first sample extended image; the first mapping information includes a mapping relationship between pixel points of the first sample extended image in the first dimensional space and position points in the second dimensional space; Generating a target mask extended image corresponding to the target original image; Generate a target extended image by using an image generation model based on the first mapping information and combining the target original image and the target mask extended image; Wherein, the image generation model is trained based on a plurality of training sample data; Each training sample data includes a sample original image, second mapping information, and a second sample extended image corresponding to the sample original image.
2. The method according to claim 1, characterized in that Also includes: Extracting target image features of the target extended image using a feature extraction model; Determining a first image feature and / or a first text feature corresponding to the non-distorted extended image type; determining whether the target extended image meets a distortion requirement based on a first similarity between the target image feature and the first image feature and / or a second similarity between the target image feature and the first text feature; The first image feature is obtained by extracting any image corresponding to the non-distorted extended image type using the feature extraction model; the first text feature is obtained by extracting the description text corresponding to the non-distorted extended image type using the feature extraction model; The feature extraction model extracts sample features corresponding to the distorted extended image type, the non-distorted extended image type and the original image type respectively, and trains the feature extraction model based on the sample features, with the similarity between the sample features corresponding to the same image type and the similarity between the sample features corresponding to different image types meeting the training requirements as the training goal; wherein the sample features include sample image features, or sample image features and sample text features.
3. The method according to claim 2, characterized in that Also includes: determining a second image feature corresponding to the reference image type; The reference image type includes: a distorted extended image type and / or an original image type; the second image feature is obtained by extracting any image corresponding to the reference image type using the feature extraction model; The determining, based on the first similarity between the target image feature and the first image feature, whether the target extended image meets the distortion requirement includes: Whether the target extended image meets a distortion requirement is determined based on a first similarity between the target image feature and the first image feature, and a third similarity between the target image feature and the second image feature.
4. The method according to claim 2 or 3, characterized in that Also includes: determining a second text feature corresponding to the reference image type; The reference image type includes: a distorted extended image type and / or an original image type; the second text feature is obtained by extracting the description text corresponding to the reference image type using the feature extraction model; The determining, based on the second similarity between the target image feature and the first text feature, whether the target extended image meets the distortion requirement includes: Whether the target extended image meets a distortion requirement is determined based on a second similarity between the target image feature and the first text feature and a fourth similarity between the target image feature and the second text feature.
5. The method according to claim 1, wherein The image generation model includes a position generation network, a content generation network and an image generation network; The generating a target extended image by using an image generation model based on the first mapping information and combining the target original image and the target mask extended image includes: Determining location guidance information based on the first mapping information using the location generation network; Determining content guidance information based on the target original image and the target masked expanded image using the content generation network; The target extended image is generated by utilizing the image generation network and combining the position guidance information and the content guidance information.
6. The method according to claim 5, characterized in that The position generation network includes a position encoding subnetwork, and the position encoding subnetwork includes a plurality of position encoding blocks; The determining, using the location generation network and based on the first mapping information, location guidance information includes: The noise information is input into the first position encoding block of the position encoding subnetwork, and the first mapping information is input into the multiple position encoding blocks respectively to obtain the position guidance information output by the position encoding subnetwork.
7. The method according to claim 6, characterized in that The position generation network further includes: a first encoding subnetwork; Inputting the first mapping information into the plurality of position encoding blocks respectively includes: The first encoding sub-network is used to extract mapping features of the first mapping information, and the mapping features are respectively input into the plurality of position encoding blocks.
8. The method according to claim 5, characterized in that The content generation network includes: a second encoding sub-network, a third encoding sub-network and a content encoding sub-network; the content encoding sub-network includes a plurality of content encoding blocks; The determining of content guidance information based on the target original image and the target masked extended image using the content generation network includes: Extracting a first encoding feature of the target mask extended image using the second encoding subnetwork; Utilizing the third encoding subnetwork, extracting a second encoding feature of the target original image; wherein a dimension of the second encoding feature is aligned with a text input dimension of the plurality of content encoding blocks; The first encoding feature and noise information are input into a first content encoding block of the content encoding subnetwork, and the second encoding feature is used as text input of the multiple content encoding blocks to obtain content guidance information output by the multiple content encoding blocks.
9. The method according to claim 1, characterized in that Obtaining the first mapping information includes: Determine a mapping relationship between pixel points of the first sample extended image in a first dimensional space and position points in a second dimensional space; A spatial edge alignment process is performed on the mapping relationship, and the first mapping information is obtained based on the mapping relationship after the spatial edge alignment process.
10. An image generation method, characterized in that: include: Acquire a target original image and first mapping information; The first mapping information is constructed based on the first sample panoramic image; the first mapping information includes a mapping relationship between pixel points of the first sample panoramic image in a two-dimensional space and position points in a spherical coordinate system; Generating a target mask panoramic image corresponding to the target original image; Generate a target panoramic image by using an image generation model based on the first mapping information and combining the target original image and the target mask panoramic image; Wherein, the image generation model is trained based on a plurality of training sample data; Each training sample data includes a sample original image and second mapping information, as well as a second sample panoramic image corresponding to the sample original image.
11. A model training method, characterized in that: include: Obtaining second mapping information; The second mapping information is constructed based on the third sample extended image; the second mapping information includes a mapping relationship between pixel points of the third sample extended image in the first dimensional space and position points in the second dimensional space; Acquire a second sample extended image, and a sample original image and a sample mask extended image corresponding to the second sample extended image; Generate a predicted extended image by using an image generation model based on the second mapping information and combining the second sample extended image, the sample original image, and the sample mask extended image; training the image generation model according to difference information between the predicted extended image and the second sample extended image; The image generation model is used to generate a target extended image based on the first mapping information, in combination with the target original image and the target mask extended image corresponding to the target original image.
12. The method according to claim 11, characterized in that Training the image generation model according to difference information between the predicted extended image and the second sample extended image includes: Extracting predicted image features of the predicted extended image using a feature extraction model; determining a fifth similarity between the predicted image feature and a sample text feature corresponding to the distorted extended image type, determining a sixth similarity between the predicted image feature and a sample text feature corresponding to the undistorted extended image type, and determining a seventh similarity between the predicted image feature and a sample text feature corresponding to the original image type; determining a distortion loss according to the fifth similarity, the sixth similarity, and the seventh similarity; determining a reconstruction loss according to difference information between the predicted extended image and the second sample extended image; Training the image generation model according to the distortion loss and the reconstruction loss; The sample text features corresponding to the distorted extended image type, the non-distorted extended image type and the original image type are respectively obtained by extracting the description texts corresponding to the distorted extended image type, the non-distorted extended image type and the original image type using the feature extraction model.
13. A model training method, characterized in that: include: Obtain training samples corresponding to the distorted extended image type, the non-distorted extended image type, and the original image type respectively; The training samples include sample training images, or sample training images and description texts; Extracting sample features corresponding to training samples of different image types using a feature extraction model; the sample features include sample image features, or sample image features and sample text features; Based on the sample features, the feature extraction model is trained with the similarity between sample features corresponding to the same image type and the similarity between sample features corresponding to different image types meeting the training requirements as a training goal; The feature extraction model is used to extract target image features of a target extended image, and to extract first image features and / or first text features corresponding to a non-distorted extended image type; the target image features, and the first image features and / or first text features are used to determine whether the target extended image meets distortion requirements; the target extended image is generated using an image generation model based on first mapping information, in combination with a target original image and a target mask extended image corresponding to the target original image; the first mapping information is constructed based on a first sample extended image; the first mapping information includes a mapping relationship between pixel points of the first sample extended image in a first dimensional space and position points in a second dimensional space.
14. The method according to claim 13, characterized in that Based on the sample features, the feature extraction model is trained with the similarity between sample features corresponding to the same image type and the similarity between sample features corresponding to different image types meeting the training requirements as a training goal, including: If the sample features include sample image features and sample text images, the feature extraction model is trained based on the sample image features and the sample text features, with the similarity between the sample image features corresponding to the same image type and the similarity between the sample image features corresponding to different image types satisfying the first training requirement as the first training goal, and with the similarity between the sample text features and sample image features corresponding to the same image type and the similarity between the sample text features and sample image features corresponding to different image types satisfying the second training requirement as the second training goal.
15. A computing device, characterized in that including processing components and storage components; The storage component stores a computer program; the computer program is used to be called and executed by the processing component to implement the image generation method as described in any one of claims 1-10, or the model training method as described in any one of claims 11-14.
16. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by the processing component, it implements the image generation method as described in any one of claims 1-10, or the model training method as described in any one of claims 11-14.
17. A computer program product, characterized in that It includes a computer program or instructions, which, when executed by a processing component, implements the image generation method as described in any one of claims 1 to 10, or the model training method as described in any one of claims 11 to 14.