Image generation method and device, model generation method and device, equipment and medium
By performing image processing and texture updates on the original image, combining descriptive text information and model images, high-quality target images are generated and rendered to the stereo model, the problem of poor three-dimensional model reconstruction results caused by low two-dimensional image quality is solved, and more efficient three-dimensional model reconstruction is achieved.
Patent Information
- Application Number
- CN202510330220.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is susceptible to the shooting angle, equipment and environment when reconstructing two-dimensional images into three-dimensional models, resulting in low image quality and poor reconstruction effect.
By image processing on the original image, visual information is extracted, and combined with descriptive text information and model images, the model image is textured update, the target image is generated, and the target image is then rendered into the stereo model to obtain the target model.
The three-dimensional model reconstruction effect in the case of quality problems in two-dimensional images is improved, the authenticity and texture details of the target image are enhanced, and the difficulty of reconstruction of the three-dimensional model is simplified.
Smart Images

Figure CN120198568A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly to the technical fields of computer vision, deep learning, large models, etc., and can be applied to scenarios such as AIGC artificial intelligence-based content generation. More specifically, it relates to an image generation method, a model generation method, a device, an intelligent agent, an electronic device, a storage medium, and a program product. Background Art
[0002] With the development of computer vision and artificial intelligence technologies, it is possible to reconstruct a three-dimensional model of an object from a two-dimensional image. This technology shows broad application prospects in the field of digital twins. However, in many cases, the reconstruction effect of the three-dimensional model cannot be guaranteed. Summary of the Invention
[0003] The present disclosure provides an image generation method, a model generation method, a device, an intelligent agent, an electronic device, a storage medium, and a program product.
[0004] According to one aspect of the present disclosure, an image generation method is provided, including: performing image processing on a target object in an original image to obtain visual information; obtaining descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object; and using the visual information and the descriptive text information to update the texture of the model image to obtain a target image.
[0005] According to another aspect of the present disclosure, a model generation method is provided, including: rendering a target image of a target object into a three-dimensional model of the target object to obtain a target model; wherein the target image is generated by the above-mentioned image method, and the three-dimensional model represents the three-dimensional contour of the target object.
[0006] According to another aspect of the present disclosure, an image processing device is provided, including: a visual processing module for performing image processing on a target object in an original image to obtain visual information; an obtaining module for obtaining descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object; and an updating module for using the visual information and the descriptive text information to update the texture of the model image to obtain a target image.
[0007] According to another aspect of the present disclosure, a model generation device is provided, including: a rendering module for rendering a target image of a target object into a three-dimensional model of the target object to obtain a target model; wherein the target image is generated by the above-mentioned image processing device, and the three-dimensional model represents the three-dimensional contour of the target object.
[0008] According to another aspect of the present disclosure, an intelligent agent is provided, wherein the intelligent agent is configured to execute the above-mentioned image processing method and model generation method.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method as described above.
[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method as described above.
[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0014] Figure 1 Schematically shows an exemplary system architecture to which the image generation method and apparatus according to the embodiments of the present disclosure can be applied;
[0015] Figure 2 Schematically shows a flowchart of the image generation method according to the embodiments of the present disclosure;
[0016] Figure 3 Schematically shows a flowchart of obtaining visual information from an initial image according to the embodiments of the present disclosure;
[0017] Figure 4 Schematically shows a flowchart of the image generation method according to another embodiment of the present disclosure;
[0018] Figure 5 Schematically shows a model generation method according to the embodiments of the present disclosure;
[0019] Figure 6 Schematically shows a generation diagram of a target model according to the embodiments of the present disclosure;
[0020] Figure 7 Schematically shows a block diagram of an image processing apparatus according to the embodiments of the present disclosure;
[0021] Figure 8A block diagram of a model generation device according to an embodiment of the present disclosure is schematically shown;
[0022] Figure 9 A structural block diagram of an agent according to an embodiment of the present disclosure is schematically shown; and
[0023] Figure 10 A schematic block diagram of an example electronic device that can be used to implement the embodiments of the present disclosure is shown. Detailed implementation manners
[0024] The following describes exemplary embodiments of the present disclosure in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] Thanks to the rapid progress of image-to-3D generation technology in recent years, a 3D image model can be quickly generated by inputting a 2D original image. However, when collecting 2D images of a target object, due to the influence of shooting angles, shooting devices, and shooting environments, the captured 2D images may have problems such as distortion, information occlusion, and strong exposure, which may lead to low-quality 2D images and ultimately affect the reconstruction effect of the 3D model.
[0026] In view of this, the present disclosure provides an image generation method, a model generation method, a device, an electronic device, an agent, a storage medium, and a program product. The aim is to generate a 3D model with a good reconstruction effect in the case where there are quality problems in the 2D image, such as distortion and information occlusion.
[0027] Figure 1 A schematic system architecture to which the image generation method and device according to an embodiment of the present disclosure can be applied is shown.
[0028] It should be noted that Figure 1 The example shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios. For example, in another embodiment, the exemplary system architecture to which the image generation method and device can be applied may include a terminal device, but the terminal device can implement the image generation method and the image processing device provided by the embodiments of the present disclosure without interacting with the server.
[0029] As Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0031] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0032] The server 105 may be a server providing various services, such as a background management server that supports the content browsed by users using the terminal devices 101, 102, and 103 (for example only). The background management server may analyze and process data such as received user requests, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0033] It should be noted that the image generation method provided by the embodiments of the present disclosure can generally be executed by the first terminal device 101, the second terminal device 102, and the third terminal device 103. Correspondingly, the content processing device provided by the embodiments of the present disclosure may also be provided in the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0034] Alternatively, the image generation method provided by the embodiments of the present disclosure can generally also be executed by the server 105. Correspondingly, the content processing device provided by the embodiments of the present disclosure can generally be disposed in the server 105. The image generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the content processing device provided by the embodiments of the present disclosure can also be disposed in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0035] For example, the first terminal device 101, the second terminal device 102, and the third terminal device 103 can be used to collect original images, and then the obtained original images are sent to the server 105. The server 105 processes the original images to obtain visual information, obtains descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object, and uses the visual information and the descriptive text information to update the texture of the model image to obtain the target image. Or a server or a server cluster capable of communicating with the terminal devices 101, 102, 103, and / or the server 105 processes the original images to obtain visual information, obtains descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object, and uses the visual information and the descriptive text information to update the texture of the model image to obtain the target image.
[0036] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in
[0037] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure, and application, etc. of the user's personal information all comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good customs.
[0038] In the technical solution of the present disclosure, the authorization or consent of the user is obtained before obtaining or collecting the user's personal information.
[0039] It should be noted that the sequence numbers of the various operations in the following methods are only used for the representation of the operations for description, and should not be regarded as indicating the execution sequence of the various operations. Unless explicitly stated, the method does not need to be executed exactly in the order shown.
[0040] Figure 2 The flowchart of the image generation method according to the embodiments of the present disclosure is schematically shown.
[0041] As Figure 2 shown, the method includes operations S210 to S230.
[0042] In operation S210, image processing is performed on the target object in the original image to obtain visual information.
[0043] In operation S220, descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object are obtained.
[0044] In operation S230, using the visual information and the descriptive text information, texture update is performed on the model image to obtain the target image.
[0045] In an embodiment of the present disclosure, the original image may be, for example, a street view image, and the target object may be a building, a road, a traffic facility, a tree, etc.
[0046] The visual information may include an image that removes background information and only retains the target object. However, it is not limited thereto. It may also include information on one or more of the shape, texture, color, and spatial position of the target object.
[0047] In some embodiments, a VAE (Variational Auto-Encoder) encoder may be used to perform image processing on the target object in the original image, map the high-dimensional original image to a low-dimensional latent space, and obtain visual information. The present disclosure is not limited thereto. A sparse auto-encoder, a denoising auto-encoder, etc. may also be used to extract the visual information of the target object from the original image.
[0048] The descriptive text information may be text information used to describe or illustrate the characteristics of the target object, and may include information such as the type, attribute, color, size, material, function, etc. of the target object. The descriptive text information may be used as feature constraint information for texture update of the model image.
[0049] The model image may be a projection of the target object in two-dimensional space, may represent the shape and boundary size of the target object, and the model image may accurately reflect the two-dimensional contour of the target object.
[0050] In some embodiments, a diffusion model may be used to process the visual information, the descriptive text information, and the model image to obtain the target image. Specifically, the visual information and the descriptive text information may be used as constraint conditions to guide the model image to perform multiple fusion and noise reduction processes to obtain the target image. The present disclosure is not limited thereto. In other embodiments, a generative adversarial network may also be used to process the visual information, the descriptive text information, and the model image to obtain the target image.
[0051] According to an embodiment of the present disclosure, by performing multimodal information fusion on the visual information and descriptive text information of the original image, the texture of the model image is updated, combining both the detailed features of the descriptive text information and the real texture information of the original image, thereby improving the processing efficiency and the authenticity of the target image. In addition, since the target image is generated using the model image, and the model image represents the two-dimensional contour of the target object, the two-dimensional contour of the target object in the texture-updated target image matches the two-dimensional contour in the model image, which is a front-facing view with horizontal and vertical lines, thereby improving the distortion correction effect of the target image.
[0052] In the related art, two-dimensional image reconstruction can be performed based on the original image, descriptive text information for describing the characteristics of the target object, and a noise image to obtain a texture-updated image. Then, the texture-updated image is pasted back into the three-dimensional model of the object to obtain a three-dimensional reconstruction model.
[0053] Optionally, the image generation method provided by the embodiment of the present disclosure can also be used to generate a front-facing view with updated texture. The front-facing view is pasted back into the three-dimensional model of the object to obtain a three-dimensional reconstruction model.
[0054] Compared with the method of updating the texture based on the original image, the noise image, and the descriptive text information, using the image generation method provided by the embodiment of the present disclosure can perform structural information constraint through the model image representing the two-dimensional contour of the target object, improving the matching of the two-dimensional contour of the target object in the target image with the two-dimensional contour in the model image, and avoiding the problem that the two-dimensional contour of the target object in the generated target image is distorted due to the distortion of the target object in the original image. This reduces the difficulty of pasting back into the three-dimensional model and improves the reconstruction effect of the three-dimensional model.
[0055] According to an embodiment of the present disclosure, for the operation S210 as Figure 2 shown, performing image processing on the target object in the original image to obtain visual information may include: when it is determined that the image quality of the original image meets the predetermined image processing conditions, performing image enhancement on the original image to obtain an information-enhanced image. Extracting visual information from the target object in the information-enhanced image to obtain visual information.
[0056] An image quality judgment model can be used to perform quality recognition on the original image to obtain the quality recognition result of the original image. The quality recognition result can be statistically counted as a percentage value. When the quality recognition result is lower than the predetermined quality threshold, it is determined that the image quality of the original image meets the predetermined image processing conditions. The type of the image quality judgment model is not limited. For example, it may include a pre-trained image classifier.
[0057] For example, due to the influence of the shooting angle, shooting device, and shooting environment, there are problems such as missing information, distortion, warping, occlusion, and strong exposure in the acquisition of the original image, resulting in the image quality meeting the predetermined image processing conditions.
[0058] In some embodiments, image enhancement of the original image may include image enhancement methods such as repairing the original image, denoising, and adjusting the grayscale value of the image. After obtaining the information-enhanced image, visual information of the target object can be extracted from the information-enhanced image using one of a VAE encoder, a sparse autoencoder, and a denoising autoencoder.
[0059] According to an embodiment of the present disclosure, by performing image enhancement on the original image with low quality and then extracting visual information from the information-enhanced image, the extracted visual information can be made closer to the true information of the original image, thereby improving the extraction effect of the visual information.
[0060] In some embodiments, the type of problem affecting the quality of the original image may be first identified, and then the image quality recognition result of the problem type may be determined, and further, whether the image quality of the original image meets the predetermined image processing conditions may be determined.
[0061] Specifically, the type of problem may include a distortion type, an information occlusion type, etc. Taking the original image as a street view image as an example, for the original image collected by a fish-eye camera, since the fish-eye camera is a panoramic camera, the building collected as the target object has a distortion type problem. In addition, since the fish-eye camera is generally installed on an information collection vehicle, and the information collection vehicle travels on the road, when collecting information on a building as the target object, it will be affected by the trees planted on both sides of the road, causing some information of the building to be occluded, and there is an information occlusion type problem.
[0062] Optionally, in the case where the type of problem affecting the image quality of the original image includes a distortion problem type, the predetermined image processing conditions may include a distortion correction condition. A distortion quality threshold may be set for the distortion correction condition, and in the case where the image quality recognition result of the distortion problem type is lower than the distortion quality threshold, it is determined that the image quality of the original image meets the distortion correction condition.
[0063] According to an embodiment of the present disclosure, performing image enhancement on the original image to obtain an information-enhanced image may include: in the case where the predetermined image processing conditions include a distortion correction condition, using a three-dimensional model of the target object to correct the distortion of the target object in the original image. The three-dimensional model represents the three-dimensional contour of the target object.
[0064] The three-dimensional model of the target object can be a three-dimensional geometric model of the target object. There is a corresponding spatial mapping relationship between the three-dimensional model and the target object. For example, when the target object is a building, the three-dimensional model can be the building's white membrane. The building's white membrane can refer to a building geometric model without material textures and decorative details obtained through three-dimensional scanning or BIM (Building Information Modeling) modeling, including key structural data such as the building's outer contour, floor height, and column grid.
[0065] Through the mapping relationship between the three-dimensional model of the target object and the target object, the feature points of the three-dimensional model and the original image can be matched. According to the matching results, the distortion parameters can be deduced inversely, and then the inverse compensation can be performed according to the distortion parameters to correct the distortion of the target object in the original image and obtain the corrected image.
[0066] According to an embodiment of the present disclosure, by using the mapping relationship between the target object and the three-dimensional model to correct the distortion of the original image, the number of pixels between the information-enhanced image and the three-dimensional model can be made to correspond, thus significantly improving the accuracy and efficiency of image processing.
[0067] According to another embodiment of the present disclosure, using the three-dimensional model of the target object to correct the distortion of the target object in the original image may further include: inputting the three-dimensional model and the original image into a three-dimensional texture generation model to obtain a three-dimensional model with updated texture. Based on the three-dimensional model with updated texture, an information-enhanced image with the same display perspective as the original image is obtained.
[0068] In an embodiment of the present disclosure, in the three-dimensional texture generation model, the original image and the three-dimensional model can be fused to generate a three-dimensional model with surface detail textures.
[0069] Specifically, first, the distortion parameters can be deduced inversely by using the mapping relationship between the three-dimensional model and the distorted original image, and the original image can be corrected for distortion using the distortion parameters to obtain the corrected image. Then, the corrected image is pasted back onto the side of the three-dimensional model with the same perspective as the original image, and the corrected image and the three-dimensional model are fused to make the space and texture of the three-dimensional image consistent, thus obtaining a three-dimensional model with updated texture. Then, an information-enhanced image is obtained from the three-dimensional model with updated texture. However, it is not limited to this. The texture information of the distorted original image can also be used to update the texture of the three-dimensional model to obtain an information-enhanced three-dimensional model. An image with the same perspective as the original image is obtained to obtain the information-enhanced image.
[0070] According to an embodiment of the present disclosure, by combining the original image and the three-dimensional model, a three-dimensional model with real textures is generated using the three-dimensional texture generation model, which not only retains the precise structure of the three-dimensional model but also realizes the efficient reconstruction of textures, thereby enabling the extraction of high-quality visual information.
[0071] Optionally, when the determined problem type affecting the image quality of the original image includes the information occlusion problem type, the predetermined image processing conditions may include information restoration conditions. An occlusion quality threshold may be set for the information restoration conditions. When the image quality recognition result of the information occlusion problem type is lower than the occlusion quality threshold, it is determined that the image quality of the original image meets the information restoration conditions.
[0072] According to another embodiment of the present disclosure, image enhancement of the original image to obtain an information-enhanced image may include: when the predetermined image processing conditions include information restoration conditions, information restoration is performed on the occluded area in the target object.
[0073] In some embodiments, the occluded area in the target object can be restored by automatically filling the occluded area using the information in the unoccluded area of the image.
[0074] In other embodiments, the pixel values of the occluded area can also be estimated by interpolation using the information of the pixels around the occluded area. Alternatively, the diffusion model is used to process the original image, treating the occluded area as noise, and the original image is gradually denoised using the diffusion model to restore the occluded area.
[0075] According to an embodiment of the present disclosure, by restoring the occluded area in the target object, the occluded part in the image can be filled, making the image content more complete and improving the visual quality of the original image.
[0076] Figure 3 A flowchart showing the process of obtaining visual information using the original image according to an embodiment of the present disclosure is schematically illustrated.
[0077] As Figure 3As shown, the image quality of the original image 310 is judged. When it is determined that the image quality of the original image 310 does not meet the predetermined image processing condition 320, the VAE encoder 330 is used to process the original image 310 to obtain visual information 390. When it is determined that the original image 310 meets the information restoration condition 350, information restoration processing is performed on the original image 310 to obtain an information-enhanced image 380, and feature extraction is performed on the information-enhanced image 380 to obtain visual information 390. When it is determined that the original image 310 meets the distortion correction condition 340, the original image 310 and the three-dimensional model 360 are input into the three-dimensional texture generation model 370 to perform distortion correction on the original image 310, generate a three-dimensional model with updated texture, and based on the three-dimensional model with updated texture, obtain an information-enhanced image 380 with the same display perspective as the original image 310, and perform feature extraction on the information-enhanced image 380 to obtain visual information 390. Thus, the quality of visual information extraction is improved.
[0078] According to an embodiment of the present disclosure, for the operation S220 as Figure 2 shown, obtaining descriptive text information for describing the characteristics of the target object may include: performing object feature recognition on the original image to obtain feature descriptive text information. Based on the feature descriptive text information and personalized information, descriptive text information is obtained.
[0079] In an embodiment of the present disclosure, performing object feature recognition on the original image may include recognizing features such as the type, attributes, color, size, material, and function of the original image. For example, if the target object is a building, specific feature information such as the height, width, number of floors, distribution of facade windows, facade color, main structure, and material of the building can be recognized, and combining these feature information constitutes the feature descriptive text information.
[0080] In an embodiment of the present disclosure, personalized information may include the user's preference information or demand information, etc. For example, the style, color, or identifying elements of the generated target image.
[0081] The feature descriptive text information can be directly used as the descriptive text information. However, it is not limited thereto. The feature descriptive text information can also be combined with personalized information to obtain the descriptive text information.
[0082] According to an embodiment of the present disclosure, by identifying the object features of the target object to obtain the feature descriptive information of the target object, the detailed features of the target object can be refined. In addition, by using personalized information, the reconstruction style expected by the user can be generated. Thus, the combination of two information types with the same information type, namely feature descriptive information and personalized information, can improve the generation effect of the target image from different aspects, such as improving the texture detail update effect and the personalized generation effect.
[0083] According to an embodiment of the present disclosure, for operation S220 as shown in Figure 2 Figure, obtaining a model image for characterizing the two-dimensional contour of the target object may include: obtaining an image with the same display perspective as the original image from the three-dimensional model of the target object to obtain the model image. The three-dimensional model characterizes the three-dimensional contour of the target object.
[0084] Specifically, in the three-dimensional model of the target object, the perspective is set to be exactly the same as the display perspective of the original image. Then, an image is rendered from the set perspective, thereby generating a model image with the same display perspective as the original image.
[0085] For example, if the target object is a building and the original image is the front view of the building, the front view of the building's white film can be obtained from the building's white film to obtain the model image.
[0086] According to an embodiment of the present disclosure, by obtaining a model image with the same display perspective as the original image from the three-dimensional model of the target object, the pixels and dimensions of the original image and the model image can be aligned. Furthermore, the generation efficiency of the three-dimensional model of the final target object can be improved.
[0087] According to an embodiment of the present disclosure, for operation S230 as shown in Figure 2 Figure, using visual information and descriptive text information to update the texture of the model image to obtain the target image may include: calling the control module of the large model to extract contour control information for texture update based on the visual information and the model image. Calling the texture update module of the large model to perform texture update on the contour control information based on the visual information and the descriptive text information to obtain the target image.
[0088] In an embodiment of the present disclosure, when using the large model to update the texture of the model image, the control module may be an independent control module added on the basis of the large model. The control module can receive external control signals, such as visual information, model image, etc., and use the control signals to guide the texture update process of the texture update module of the large model, so that the generated target image better meets the specified conditions.
[0089] By calling the control module, the texture information of the visual information can be fused with the contour information of the model image, thereby obtaining the contour control information for texture update.
[0090] According to an embodiment of the present disclosure, the large model may be a multimodal large model combined with a diffusion model, and the texture update module may be the original architecture of the diffusion model or a multimodal large model combined with a diffusion model. Using the visual information and the descriptive text information as constraint conditions, guiding the contour control information to perform multi-level fusion denoising processing to obtain the target image.
[0091] According to an embodiment of the present disclosure, by calling a control module to extract contour control information for texture update, more refined fusion control of visual information and model images can be achieved, and contour control information meeting requirements can be generated. Then, by calling a large model to process visual information and descriptive text information, and for the contour control information, a target image is obtained, which helps to improve the generation efficiency and generation effect of the target image.
[0092] Figure 4 A flowchart of an image generation method according to another embodiment of the present disclosure is schematically shown.
[0093] As Figure 4 shown, the image quality of the original image 401 is judged. When the image quality of the original image 401 meets a predetermined image processing condition 402, information enhancement processing is performed on the original image 401 to obtain an information-enhanced image 403, and feature extraction is performed on the information-enhanced image 403 to obtain visual information 405. When the original image 401 does not meet the predetermined image processing condition 402, the original image 401 is processed by a VAE encoder 404 to obtain visual information 405. Object feature recognition is performed on the original image 401 to obtain feature descriptive text information 406, and descriptive text information 408 is obtained according to the feature descriptive text information 406 and personalized information 407. An image with the same display perspective as the original image 401 is obtained from the three-dimensional model 409 of the target object to obtain a model image 410. The visual information 405, descriptive text information 408, and model image 410 are input into a large model 411. By calling the control module of the large model 411 based on the visual information 405 and model image 410, contour control information for texture update is extracted, and the texture update module of the large model 411 is called to perform texture update on the contour control information based on the visual information 405 and descriptive text information 408 to obtain a target image 412.
[0094] According to an embodiment of the present disclosure, after performing the operation S230 shown in Figure 2 to generate a target image, the image generation method may further include: using the descriptive text information and visual information to perform texture update on the target image to obtain an updated target image.
[0095] The descriptive text information, visual information, and target image may be used as input information and input into a large model to obtain an updated target image.
[0096] Specifically, the visual information and descriptive text information may be used as constraint conditions to guide the target image to perform multiple fusion noise reduction processes to obtain an updated target image.
[0097] In some other embodiments, a generative adversarial network or a diffusion model can also be used to process visual information, descriptive text information, and a target image to obtain an updated target image.
[0098] According to an embodiment of the present disclosure, by using descriptive text information and visual information to update the texture of a target image, the target image can be iteratively texture-updated to generate an image that meets the requirements, thereby improving the generation accuracy of the final target image.
[0099] According to an embodiment of the present disclosure, after generating the target image or the updated target image, the image generation method may further include: generating other perspective images with different display perspectives from the target image based on the target image.
[0100] Specifically, the target image can be input into a multi-perspective image transformation model to generate images with different perspectives corresponding to the input image.
[0101] In some embodiments, the multi-perspective image transformation model can be a multi-perspective image transformation model trained based on a generative adversarial network. It can also be a multi-perspective image transformation model trained based on a diffusion model. It can also be obtained by processing with other large models. As long as it can predict and generate other perspective images with different display perspectives based on a single-perspective target image.
[0102] For example, if the target image is in the due east perspective, other perspective images in the due south, due north, and due west perspectives can be generated based on the target image.
[0103] According to an embodiment of the present disclosure, by using the target image to generate images with other perspectives, the geometric constraints of a single-perspective image can be broken through, and multi-angle coherent visual information can be generated, thereby facilitating the improvement of the 3D reconstruction effect.
[0104] Figure 5 Schematically shows a model generation method according to an embodiment of the present disclosure.
[0105] As Figure 5 shown, the method includes step S510.
[0106] In operation S510, the target image of the target object is rendered into the three-dimensional model of the target object to obtain a target model.
[0107] According to an embodiment of the present disclosure, the three-dimensional model represents the three-dimensional contour of the target object.
[0108] According to an embodiment of the present disclosure, image processing can be performed on a target object in an original image to obtain visual information. Descriptive text information for describing the characteristics of the target object and a model image for characterizing the two-dimensional contour of the target object are acquired. Using the visual information and the descriptive text information, texture update is performed on the model image to obtain a target image.
[0109] After obtaining the target image, the target image can be attached to the surface of a three-dimensional model of the target object, and a suitable renderer and rendering parameters are selected to perform rendering on the three-dimensional model, and the texture of the target image is rendered into the three-dimensional model to obtain a target model of the target object.
[0110] In some embodiments, other perspective images with different display perspectives from the target image can be generated first according to the target image, and then the images of all perspectives are attached to the corresponding surfaces of the three-dimensional model of the target object, and the three-dimensional model is rendered to obtain a three-dimensional target model of the target object.
[0111] According to an embodiment of the present disclosure, since the target image for rendering the three-dimensional model is a high-quality target image obtained by performing texture update on the model image using visual information and descriptive text information, the contour of the target image matches the contour of the three-dimensional model. Thus, while improving the final rendering effect of the three-dimensional model, the difficulty of three-dimensional model reconstruction can be simplified and the reconstruction efficiency can be improved.
[0112] Figure 6 Schematically shows a schematic diagram of the generation of a target model according to an embodiment of the present disclosure.
[0113] As Figure 6 shown, the original image 610 may include a building as the target object 611. Since the original image 610 is acquired using a fish-eye camera, there is a distortion problem with the target object 611 in the original image 610. In addition, due to the occlusion of objects such as wires 612 and roadside obstacles 613, there is an information occlusion problem with the target object 611 in the original image 610. Processing the original image 610 using the above image generation method can obtain a target image 620 in a front view perspective, where the occluded information of the target object in the target image has been restored and the distortion has been corrected. Other perspective images 630 different from the display perspective of the target image 620 can be predicted and generated using the target image 620 and then reattached to the three-dimensional model 640 of the target object to obtain a three-dimensional target model 650 of the target object.
[0114] Figure 7 Schematically shows a block diagram of an image processing apparatus according to an embodiment of the present disclosure.
[0115] As Figure 7 shown, the apparatus includes a visual processing module 710, an acquisition module 720, and an update module 730.
[0116] A visual processing module 710, configured to perform image processing on a target object in an original image to obtain visual information.
[0117] An acquisition module 720, configured to acquire descriptive text information for describing the features of the target object and a model image for characterizing the two-dimensional contour of the target object.
[0118] An update module 730, configured to update the texture of the model image by using the visual information and the descriptive text information to obtain a target image.
[0119] According to an embodiment of the present disclosure, the visual processing module 710 includes an enhancement processing sub-module and an information extraction sub-module.
[0120] The enhancement processing sub-module is configured to perform image enhancement on the original image to obtain an information-enhanced image when it is determined that the image quality of the original image meets a predetermined image processing condition.
[0121] The information extraction sub-module is configured to extract visual information from the target object in the information-enhanced image to obtain visual information.
[0122] According to an embodiment of the present disclosure, the visual processing module 710 further includes a distortion correction sub-module.
[0123] The distortion correction sub-module is configured to perform distortion correction on the target object in the original image by using a three-dimensional model of the target object when the predetermined image processing condition includes a distortion correction condition, wherein the three-dimensional model characterizes the three-dimensional contour of the target object.
[0124] According to an embodiment of the present disclosure, the distortion correction sub-module includes a texture update unit and an information-enhanced image determination unit.
[0125] The texture update unit is configured to input the three-dimensional model and the original image into a three-dimensional texture generation model to obtain a three-dimensional model with updated texture.
[0126] The information-enhanced image determination unit is configured to obtain an information-enhanced image with the same display perspective as the original image based on the three-dimensional model with updated texture.
[0127] According to an embodiment of the present disclosure, the enhancement processing sub-module includes an information restoration unit.
[0128] The information restoration unit is configured to restore the information of the occluded area in the target object when the predetermined image processing condition includes an information restoration condition.
[0129] According to an embodiment of the present disclosure, the image processing apparatus further includes a target image update module.
[0130] The target image update module uses descriptive text information and visual information to update the texture of the target image and obtain the updated target image.
[0131] According to an embodiment of the present disclosure, the acquisition module 720 includes an identification unit and an information determination unit.
[0132] The identification unit is configured to perform object feature recognition on the original image to obtain feature descriptive text information.
[0133] The information determination unit is configured to obtain descriptive text information based on the feature descriptive text information and personalized information.
[0134] According to an embodiment of the present disclosure, the acquisition module 720 further includes a model image acquisition unit.
[0135] The model image acquisition unit is configured to obtain an image with the same display perspective as the original image from the three-dimensional model of the target object to obtain a model image, where the three-dimensional model represents the three-dimensional contour of the target object.
[0136] According to an embodiment of the present disclosure, the image processing device further includes a generation module.
[0137] The generation module is configured to generate other perspective images with different display perspectives from the target image based on the target image.
[0138] According to an embodiment of the present disclosure, the update module 720 includes a first call unit and a second call unit.
[0139] The first call unit is configured to call the control module to extract contour control information for texture update based on visual information and the model image.
[0140] The second call unit calls the large model to perform texture update on the contour control information based on visual information and descriptive text information to obtain the target image.
[0141] According to embodiments of the present disclosure, any plurality of the visual processing module 710, the acquisition module 720, and the update module 730 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to embodiments of the present disclosure, at least one of the visual processing module 710, the acquisition module 720, and the update module 730 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the visual processing module 710, the acquisition module 720, and the update module 730 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0142] Figure 8 Schematically shows a block diagram of a model generation device according to an embodiment of the present disclosure.
[0143] As Figure 8 shown, the image processing device includes a rendering module 810
[0144] The rendering module 810 is configured to render a target image of a target object into a three-dimensional model of the target object to obtain a target model.
[0145] According to embodiments of the present disclosure, the target image is generated by the image processing device 700, and the three-dimensional model represents the three-dimensional contour of the target object.
[0146] Figure 9 Schematically shows a block diagram of the structure of an agent of artificial intelligence according to an embodiment of the present disclosure.
[0147] In embodiments of the present disclosure, as Figure 9 shown, the agent 900 may include an input module 910, a processing module 920, and an output module 930.
[0148] The input module 910 is configured to receive input information.
[0149] The processing module 920 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the image processing method and the model generation method provided according to embodiments of the present disclosure by calling the large model.
[0150] An output module 930 for outputting the output information obtained by the processing module.
[0151] According to an embodiment of the present disclosure, the input module 910 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (such as a user or an external environment), and converting it into a format that the agent 900 can understand and process. The input module 910 is the primary link for the agent 900 to interact with the outside world, enabling the agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0152] In an example, the input module 910 can input the original image described above.
[0153] In an example, the processing module 920 is the core support for the agent 900's ability to handle complex tasks. The processing module 920 can execute the image processing method and the model generation method described above.
[0154] In an example, the performance of the processing module 920 can be closely related to the large model on which the agent 900 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 920 can be designed to be highly configurable and extensible to handle various different types of tasks and requirements in real-world scenarios.
[0155] In an example, after the agent 900 obtains the original image, the processing module 920 can use the large model to process the original image, obtain the target image, and transfer it to the output module 930.
[0156] In an example, the output module 930 can output the target object or the target model described above.
[0157] The agent 900 according to an embodiment of the present disclosure can simply and effectively improve the degree of intelligence, and improve flexibility and versatility.
[0158] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0159] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0160] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0161] According to an embodiment of the present disclosure, a computer program product includes a computer program which, when executed by a processor, implements the method as described above.
[0162] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0163] As Figure 10 shown, the device 1000 includes a computing unit 1001 which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0164] A plurality of components in the device 1000 are connected to the input / output (I / O) interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, an optical disk, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0165] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the image processing method and the model generation method. For example, in some embodiments, the image processing method and the model generation method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the image processing method and the model generation method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the image processing method and the model generation method by any other suitable means (e.g., by means of firmware).
[0166] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0167] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0168] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0169] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0170] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0171] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs that run on the respective computers and have a client-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0172] It should be understood that various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved. This is not limited herein.
[0173] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating an image, comprising: Perform image processing on the target object in the original image to obtain visual information; Acquire descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object; as well as The model image is texture updated using the visual information and the descriptive text information to obtain a target image.
2. The method according to claim 1, wherein: The step of performing image processing on the target object in the original image to obtain visual information includes: When it is determined that the image quality of the original image meets a predetermined image processing condition, performing image enhancement on the original image to obtain an information enhanced image; and Visual information is extracted from the target object in the information enhanced image to obtain the visual information.
3. The method according to claim 2, wherein: The performing image enhancement on the original image to obtain an information enhanced image includes: In a case where the predetermined image processing conditions include distortion correction conditions, distortion correction is performed on the target object in the original image using a stereo model for the target object, wherein the stereo model represents a stereo contour of the target object.
4. The method according to claim 3, wherein: The method of using the stereoscopic model of the target object to perform distortion correction on the target object in the original image includes: Inputting the three-dimensional model and the original image into a three-dimensional texture generation model to obtain a three-dimensional model with updated texture; and Based on the texture-updated stereo model, the information-enhanced image having the same display viewing angle as the original image is obtained.
5. The method according to any one of claims 2 to 4, wherein: The performing image enhancement on the original image to obtain an information enhanced image includes: When the predetermined image processing conditions include information restoration conditions, information restoration is performed on the blocked area of the target object.
6. The method according to any one of claims 1 to 5, further comprising: The texture of the target image is updated by using the descriptive text information and the visual information to obtain an updated target image.
7. The method according to any one of claims 1 to 6, wherein: Obtain descriptive text information for describing the characteristics of the target object, including: Performing object feature recognition on the original image to obtain feature descriptive text information; and The descriptive text information is obtained based on the feature descriptive text information and the personalized information.
8. The method according to any one of claims 1 to 7, wherein: Acquiring a model image for representing a two-dimensional contour of the target object, including: An image having the same display viewing angle as the original image is acquired from the three-dimensional model of the target object to obtain the model image, wherein the three-dimensional model represents the three-dimensional contour of the target object.
9. The method according to any one of claims 1 to 8, further comprising: Based on the target image, another viewing angle image having a different display viewing angle from that of the target image is generated.
10. The method according to any one of claims 1 to 9, wherein: The step of updating the texture of the model image by using the visual information and the descriptive text information to obtain a target image includes: Calling the control module of the large model to extract contour control information for texture updating based on the visual information and the model image; and The texture update module of the large model is called to perform texture update on the contour control information based on the visual information and the descriptive text information to obtain the target image.
11. A model generation method, comprising: Rendering a target image of a target object into a stereoscopic model of the target object to obtain a target model; The target image is generated by the method according to any one of claims 1 to 10, and the three-dimensional model represents the three-dimensional contour of the target object.
12. An image processing device, comprising: A visual processing module is used to perform image processing on the target object in the original image to obtain visual information; An acquisition module, used to acquire descriptive text information for describing the characteristics of the target object and a model image for representing the two-dimensional contour of the target object; as well as An updating module is used to update the texture of the model image using the visual information and the descriptive text information to obtain a target image.
13. A model generation device, comprising: A rendering module, used for rendering a target image of a target object into a three-dimensional model of the target object to obtain a target model; The target image is generated by the device as claimed in claim 12, and the three-dimensional model represents the three-dimensional outline of the target object.
14. An intelligent agent, wherein: The agent is configured to perform the method according to any one of claims 1 to 11.
15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.