Method, device and medium for generating extraterrestrial building image
By working in tandem with a large language model and an image diffusion model, a structured cue word framework is used to convert natural language into cue words that match the extraterrestrial environment, generating efficient and professional images of extraterrestrial architecture. This solves the problems of professionalism and efficiency in the visualization of extraterrestrial architecture design in existing technologies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ARCHITECTURAL DESIGN & RES INST OF TSINGHUA UNIV
- Filing Date
- 2026-02-03
- Publication Date
- 2026-06-02
AI Technical Summary
Existing architectural design visualization tools are unable to accurately reflect the architectural form, material adaptability, and spatial layout logic in extraterrestrial environments. They lack data support for space environments, resulting in unprofessional and inefficient generated images.
The system employs a large language model and an image diffusion model working together. Natural language input is converted into prompts that match the extraterrestrial environment through a structured prompt word framework to generate an initial building image. The image diffusion model is then driven by refined prompt words to iteratively generate the target building image.
It achieves automated and controllable image generation from macro planning to micro detail, improving the professional adaptability and generation efficiency of extraterrestrial architectural visualization. The generated images are highly controllable and professional in terms of details, materials and functions.
Smart Images

Figure CN122133252A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a method, apparatus, device, and medium for generating images of extraterrestrial architecture. Background Technology
[0002] Currently, with the deepening of human space exploration, establishing permanent structures on the surfaces of exoplanets such as the Moon has become an important development direction. In this process, architectural design needs to adapt to the extreme environments of exoplanets, such as low gravity, high radiation, long periods of alternating polar day and night, extreme temperature differences, and the lack of an atmosphere. Traditional architectural design methods, mostly based on Earth's environment, are insufficient to accurately address the unique needs of exoplanetary environments.
[0003] In existing technologies, visualization tools for architectural design are mainly designed for Earth-based architectural scenarios. They lack data support and physical constraints for space environments, making it difficult for the generated images to accurately reflect architectural forms, material adaptability, and spatial layout logic in extraterrestrial environments.
[0004] Therefore, there is an urgent need for a visualization solution that can directly serve the construction of extraterrestrial structures. Summary of the Invention
[0005] To address the aforementioned technical problems, this disclosure provides a method, apparatus, device, and medium for generating images of extraterrestrial architecture.
[0006] According to one aspect of this disclosure, a method for generating images of extraterrestrial architecture is provided, the method comprising: The system acquires an input statement describing the architecture of an alien planet and converts the input statement into a first prompt word that matches the environment of the alien planet using a preset structured prompt word framework. The geospatial image of the outer planet's surface and the first prompt word are input into a pre-trained image diffusion model, which generates multiple sets of initial building images. The structured cue word framework is used to generate second cue words to describe building attributes; The image diffusion model generates multiple sets of target building images based on the second prompt word and multiple sets of initial building images.
[0007] According to another aspect of this disclosure, an apparatus for generating images of extraterrestrial architecture is also provided, the apparatus comprising: The first prompt word generation module is used to obtain the input statement used to describe the alien planet's architecture, and convert the input statement into a first prompt word that matches the alien planet's environment through a preset structured prompt word framework; The first image generation module is used to input the geospatial image of the surface of the extraterrestrial planet and the first prompt word into a pre-trained image diffusion model, and generate multiple sets of initial building images through the image diffusion model; The second prompt word generation module is used to generate second prompt words for describing building attributes through the structured prompt word framework; The second image generation module is used to generate multiple sets of target building images based on the second prompt word and multiple sets of the initial building images using the image diffusion model.
[0008] This disclosure also provides an electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-described method for generating images of alien buildings.
[0009] This disclosure also provides a computer-readable storage medium storing a computer program for executing the above-described method for generating images of extraterrestrial structures.
[0010] The technical solution provided in this disclosure has the following advantages compared with the prior art: The technical solution provided in this disclosure includes: acquiring an input statement for describing extraterrestrial architecture; converting the input statement into a first prompt word that matches the extraterrestrial environment through a preset structured prompt word framework; inputting a geospatial image of the extraterrestrial surface and the first prompt word into a pre-trained image diffusion model; generating multiple sets of initial building images through the image diffusion model; generating a second prompt word for describing building attributes through the structured prompt word framework; and generating multiple sets of target building images based on the second prompt word and the multiple sets of initial building images through the image diffusion model.
[0011] In this technical solution, firstly, a structured cue word framework is used to convert natural language input statements into first cue words that conform to the image diffusion model. Secondly, by fusing geospatial images with the first cue words through the image diffusion model, multiple sets of initial building images with reasonable macro-layout and terrain integration can be generated quickly and in batches, significantly improving the efficiency and diversity of the conceptual design stage. Furthermore, based on the structured cue word framework, detailed second cue words describing building attributes are generated, driving the image diffusion model to deepen and iterate through multiple attributes of the initial building images, ultimately outputting highly controllable and professional target building images in terms of details, materials, and functions. Therefore, this solution can effectively achieve automated and controllable image generation from macro-planning to micro-detail, effectively improving the technical difficulties of poor professional adaptability, low generation efficiency, and limited solutions in traditional methods for visualizing extraterrestrial architecture. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0013] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of the method for generating images of extraterrestrial structures according to an embodiment of this disclosure; Figure 2 This is a schematic diagram of the initial building image described in an embodiment of this disclosure; Figure 3 This is a schematic diagram of the target building image described in the embodiments of this disclosure; Figure 4 This is a schematic diagram of the structure of the apparatus for generating images of extraterrestrial structures according to an embodiment of this disclosure; Figure 5 This is a schematic diagram of the structure of the electronic device described in an embodiment of this disclosure. Detailed Implementation
[0015] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0016] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0017] For visualization solutions of extraterrestrial architecture, one approach is to utilize AI image generation platforms to generate images based on natural language descriptions. However, due to a lack of specialized training data for extraterrestrial architecture or aerospace industry styles, the generated images suffer from inconsistent styles, insufficient professionalism, and poor adaptability to the real environment, making them unsuitable for directly serving the visualization design of lunar habitation sites.
[0018] In this context, embodiments of this disclosure provide a method, apparatus, device, and medium for generating images of extraterrestrial architecture. This disclosure primarily utilizes a large language model and an image diffusion model to collaboratively generate architectural images suitable for extraterrestrial environments. It is mainly applied to the design of human settlements in extraterrestrial spaces such as the Moon, effectively enabling the intelligent assisted generation of images such as sketches, bird's-eye views, and material composition diagrams of human settlements adapted to extraterrestrial environments. This provides a visually appealing, highly interactive, and efficient image expression platform for space architecture concepts.
[0019] Figure 1 This flowchart illustrates a method for generating images of buildings on an extraterrestrial planet, as provided in this embodiment. This method is applicable to scenarios involving the design of architectural images on extraterrestrial planets such as the Moon. The method for generating these images can be executed by an apparatus for generating such images, which can be implemented using software and / or hardware, specifically, for example, electronic devices or servers. The electronic devices may include in-vehicle hosts, tablets, desktop computers, laptops, and smartphones, etc., devices with communication capabilities. The server may be a cloud server or a server cluster, etc., devices with storage and computing capabilities.
[0020] like Figure 1 As shown, the method for generating images of extraterrestrial structures provided in this embodiment may include the following steps.
[0021] S102, Obtain the input statement used to describe the alien planet's architecture, and convert the input statement into a first prompt word that matches the alien planet's environment through a preset structured prompt word framework.
[0022] In this embodiment, a user can input natural language statements into a structured prompting framework to describe the design concept of an alien building, such as: a research station on the lunar surface.
[0023] This embodiment is not limited to a single large language model, but rather constructs a structured prompting framework applicable to general large language models (such as state-of-the-art models like GPT-4o, DeepSeek V3, and Qwen 2.5). This structured prompting framework guides the large language model through a chained reasoning process involving statement requirement analysis, constraint extraction, and instruction translation, using the mapping relationship between extraterrestrial environmental parameters and Earth's architectural logic, as well as a semantic alignment mechanism.
[0024] Specifically, when invoking the large language model, a pre-defined mapping relationship between structured extraterrestrial environmental parameters and Earth's architectural logic is established within the context of the large language model. This mapping relationship is presented in the form of tables, lists, or structured text, including key environmental parameters of the extraterrestrial planet; taking the Moon as an example, these environmental parameters include: gravitational acceleration approximately 1 / 6 that of Earth, high vacuum, strong cosmic ray radiation, and surface temperature differences reaching ±180℃, etc.
[0025] It is understandable that the environmental parameters of the aforementioned exoplanets differ significantly from those of Earth, which would impact Earth-based architectural logic. Therefore, this embodiment can correlate the aforementioned environmental parameters of exoplanets with the architectural design logic; these impacts include, for example, structural load calculation methods, airtightness requirements, radiation shielding strategies, and thermal control system design. Furthermore, a mapping relationship is established between exoplanetary environmental parameters and Earth-based architectural logic. This mapping relationship essentially endows the large language model with the prior knowledge required for domain-specific reasoning.
[0026] Regarding the semantic alignment mechanism, the structured cue word framework predefines conversion rules from natural language expressing architectural design needs to cue words for the image diffusion model. These conversion rules ensure that the output format of the large language model is stable and the content is accurate, enabling it to directly serve the downstream image diffusion model.
[0027] As a concrete example, when a user inputs the phrase "scientific research station on the lunar surface," the large language model, suitable for a structured cue word framework, automatically transforms the input into a sequence of English labels that conforms to best practices for image diffusion models, such as: "lunar habitat, hexagonal modules, regolith 3d printed texture, cold lighting, unreal engine 5 render, 8k." This sequence of English labels can then be used as the first cue word input into the image diffusion model.
[0028] A large language model suitable for a structured prompt word framework, based on the mapping relationship between the aforementioned extraterrestrial environment parameters and Earth's architectural logic, as well as a semantic alignment mechanism, converts the input statement into a first prompt word that matches the extraterrestrial environment. Specifically, this includes: First, based on a pre-defined structured prompt framework and a pre-defined mapping relationship between extraterrestrial environmental parameters and Earth's architectural logic, the system generates a list of space usage functions and a task list that conform to the extraterrestrial environment from the input statement. Then, according to pre-defined semantic alignment rules, the list of space usage functions and the task list are converted into the first prompt.
[0029] Specifically, taking the Moon as an example, large language models (such as ChatGPT) compare and analyze the differences between environmental parameters of buildings on the Moon and on Earth, automatically extracting constraints such as approximately 1 / 6 gravity, large temperature differences, and the absence of an atmosphere. Based on these constraints, a necessary list of spatial usage functions and a task specification are generated. The list of spatial usage functions and the task specification clearly define the requirements for adaptability to the extraterrestrial environment and the composition of basic functions that the building needs to meet.
[0030] After completing the above analysis, the large language model converts the spatial usage function list and task list into specialized first prompts suitable for image generation, based on preset semantic alignment rules. The first prompts are label sequences containing rich visual descriptors that conform to the best practices of image diffusion models, such as: semi-underground honeycomb units resistant to high radiation, modular communities under low gravity.
[0031] S104: Input the geospatial image of the alien planet's surface and the first prompt word into the pre-trained image diffusion model, and generate multiple sets of initial building images through the image diffusion model.
[0032] The image diffusion model in this embodiment is obtained by fine-tuning the pre-trained stable diffusion (SD) model based on the LoRA model. The fine-tuning process will be described in subsequent embodiments.
[0033] In this embodiment, the geospatial image of the exoplanet's surface and a first prompt word are first input into a pre-trained image diffusion model. The geospatial image of the exoplanet's surface can be a topographical graphic or geomorphic model of the requested surface, specifically an image showing geospatial features such as craters and plains. The first prompt word is, for example, "radial module group arranged based on the edge of the lunar maria."
[0034] Then, the geospatial image is converted into a terrain semantic map using an image diffusion model, and the first cue word is converted into a semantic vector; then, multiple sets of initial building images are generated based on the terrain semantic map and the semantic vector.
[0035] Specifically, the image diffusion model converts geospatial images into standardized topographic semantic maps. For example, it distinguishes different landform features, such as steep slopes, flat land, and the edges of crater mountains, through color coding or channel annotation. This data serves as a constraint on the spatial layout of the image diffusion model when generating building images, ensuring that the location relationships of the generated building clusters conform to topographic logic, such as avoiding placement on steep slopes and utilizing crater walls as natural barriers.
[0036] For the example of the first cue word "radial module clusters arranged along the edge of the lunar maria," it contains rich design semantics, including the layout strategy (radial), environmental relationship (maria edge), and unit type (module cluster). A text encoder using an image diffusion model (such as CLIP Text Encoder) converts the first cue word into a high-dimensional semantic vector, serving as a conditional signal controlling the overall style and content theme of the image.
[0037] Building upon this foundation, the image diffusion model, guided by global style and content representations from semantic vectors (e.g., generating lunar buildings and modular clusters), and spatial structure representations from terrain semantic maps (e.g., buildings must conform to the plain edge and be distributed radially), and by introducing random noise seeds, it iterates through multiple steps to rapidly generate multiple sets of bird's-eye view images that conform to the terrain conditions and design strategies, and vary in individual building forms, cluster density, lighting effects, and color schemes, all while maintaining core design constraints such as terrain and layout strategies. This results in multiple sets of initial building images. An example of an initial building image can be found here. Figure 2 As shown.
[0038] This embodiment combines the first prompt word with geospatial images to accurately and collaboratively generate multiple sets of high-quality initial architectural images in multiple dimensions such as space, style, and semantics.
[0039] S106, Generate a second cue word to describe the building attributes using a structured cue word framework.
[0040] This embodiment may include: inputting contextual information into a structured prompting framework; wherein the contextual information includes at least one of the following: a list of space usage functions and a task book, descriptive text of the initial building image, and user instructions; generating a second prompting word for describing the building attributes of the initial building image based on the contextual information through the structured prompting framework; wherein the building attributes include at least: building materials, structural form, skin form, and background elements that match the environment of the alien planet.
[0041] In a specific embodiment, based on the generated initial architectural image, the structured cue word framework drives the large language model to perform design-deepening reasoning and automatically generate second cue words to guide the generation of architectural details and higher professionalism.
[0042] To generate a second cue word, a large language model suitable for a structured cue word framework, in addition to the pre-defined mapping relationship between its injected extraterrestrial environmental parameters and Earth's architectural logic, can also input the following contextual information: Space usage function list and task list: These serve as background knowledge to ensure that the detailed design does not deviate from the initial environmental adaptability and functional objectives. Initial architectural image descriptive text: It's understandable that large language models typically cannot directly process image data; therefore, the initial architectural image can be converted into descriptive text. For example, a visual language model (such as GPT-4V, Qwen-VL, etc.) can be used to generate a structured, detailed descriptive text for the initial architectural image. The descriptive text includes: overall layout features (e.g., a cluster composed of six hexagonal modules arranged in a ring), topographical relationships (e.g., situated on a gentle slope at the edge of a crater), and the general shape and style of the architectural complex.
[0043] User instructions: These are used to express the user's detailed design requirements in natural language form. For example, the user instruction could be: Design three different skin texture schemes for the entire community.
[0044] Next, the large language model generates specific and visual descriptions of multiple architectural attributes, such as building materials, structural forms, surface forms, and background elements, based on contextual information. These descriptions are then integrated into second cue words suitable for generating professional architectural renderings. These second cue words refine the more detailed architectural attributes in the initial architectural image. The specific second cue words used to describe architectural attributes are as follows: The second cue word for building materials is used to describe building materials adapted to extraterrestrial environments and their surface visual characteristics; among them, building materials include 3D printed lunar soil, composite materials, and metals, and surface visual characteristics include texture, reflection, and wear.
[0045] The second cue word corresponding to the structural form is used to infer the special structural forms that may be used in low gravity and thermal expansion and contraction environments, as well as the visualized construction node details. Such structural forms include folding and unfolding, inflatable and nested cabins, etc.
[0046] The second prompt word corresponding to the skin form is used to describe how the building skin handles radiation shielding, thermal insulation, and the protection of openings such as doors, windows, and observation windows, such as multi-layer hatch covers and porthole sunshades.
[0047] The second cue word corresponding to the background element is used to describe the background element that can reflect the scene and architectural function of the alien planet, such as solar arrays, lunar rovers, astronauts, equipment boxes, exposed pipeline systems, etc.
[0048] S108: Based on the second prompt word and multiple initial building images, the image diffusion model generates multiple sets of target building images.
[0049] This embodiment, based on the generated initial architectural image, further utilizes a second set of design prompts generated by a large language model to refine the design, and then outputs multiple sets of target architectural images with different design directions in batches through an image diffusion model. An example of the target architectural images can be found here. Figure 3 As shown.
[0050] Regarding the image diffusion model in the above embodiments, this embodiment provides a training method for the image diffusion model.
[0051] In this embodiment, sample images are first acquired, including aerospace industrial style images and architectural effect images; and the sample images are standardized and text tags used to describe image information are added.
[0052] Specifically, real-life photos and high-resolution architectural renderings of completed projects can be obtained from collaborating architectural design firms, and these images can be used as architectural renderings. Architectural renderings can represent the compositional logic, visual expression techniques, and aesthetic tone required for the output images, ensuring that the images generated by the image diffusion model have a unified aesthetic style.
[0053] High-resolution images of the interior and exterior of real spacecraft, space station modules, and planetary probes were collected from publicly available resources, along with materials such as simulated bases on the moon and other exoplanets, and science fiction industrial design concept images. These images were then used as aerospace industrial style images. Aerospace industrial style images provide key visual elements such as equipment layout, interface forms, and material textures (e.g., metals, composite materials, fabrics) in zero-gravity or low-gravity environments, ensuring that the images generated by the image diffusion model have a unified aerospace industrial style.
[0054] All sample images, including those with aerospace industrial style and architectural renderings, are standardized, including but not limited to: uniform resolution (e.g., 1024×1024) and color space correction, to ensure the consistency of the training data.
[0055] Add text labels to the sample images to describe refined image information. Specifically, these text labels not only label the names of objects in the sample images (such as "habitat"), but also describe in more detail the iconic architectural aesthetic features of the sample images, such as visual style, material texture, structural features, lighting environment, and other image information. The lighting environment mentioned above is as follows: a metal cabin with heat dissipation fins, a matte white coating on the surface, blue auxiliary lighting inside, and a composition with a strong sense of depth.
[0056] Based on the above sample images, this embodiment proceeds to the training process of the image diffusion model, as shown below.
[0057] (1) With the weight parameters of the pre-trained stable diffusion model fixed, the LoRA model is injected into the cross-attention layer of the U-Net network; (2) Input the pre-acquired sample images into a stable diffusion model injected with the LoRA model, and output the predicted images; (3) Use a preset loss function to determine the loss value between the sample image and the predicted image; (4) Update the parameters in the LoRA model by backpropagation based on the loss value to obtain the LoRA fine-tuning update parameters; (5) Based on LoRA, fine-tune the updated parameters, return to the step of inputting the sample image with the above text labels into the stable diffusion model injected with the LoRA model and outputting the predicted image, until the loss value is less than the preset loss threshold, and output the LoRA model parameters. (6) Subtract the LoRA model parameters from the weight parameters of the stable diffusion model to obtain the updated stable diffusion model, and determine the updated stable diffusion model as the image diffusion model.
[0058] This embodiment uses the LoRA training method. While keeping the pre-trained weight parameters of the stable diffusion model frozen, a trainable low-rank decomposition matrix is injected into the cross-attention layer of U-Net. Based on this, the stable diffusion model is fine-tuned, and the fine-tuned stable diffusion model is determined as the final image diffusion model that can be used for image generation.
[0059] In the training examples above, the sample images used ensure that the images generated by the image diffusion model have a consistent and controllable professional visual style, effectively solving the problem of random and unprofessional output styles from general models. The LoRA mechanism makes the customization and updating of professional styles extremely cost-effective, eliminating the need to train a large image diffusion model from scratch each time.
[0060] In summary, the method for generating images of extraterrestrial buildings provided in this disclosure includes: acquiring an input statement for describing extraterrestrial buildings; converting the input statement into a first prompt word that matches the environment of the extraterrestrial planet through a preset structured prompt word framework; inputting a geospatial image of the extraterrestrial planet's surface and the first prompt word into a pre-trained image diffusion model; generating multiple sets of initial building images through the image diffusion model; generating a second prompt word for describing building attributes through the structured prompt word framework; and generating multiple sets of target building images based on the second prompt word and the multiple sets of initial building images through the image diffusion model.
[0061] In this technical solution, firstly, a structured cue word framework is used to convert natural language input statements into first cue words that conform to the image diffusion model. Secondly, by fusing geospatial images with the first cue words through the image diffusion model, multiple sets of initial building images with reasonable macro-layout and terrain integration can be generated quickly and in batches, significantly improving the efficiency and diversity of the conceptual design stage. Furthermore, based on the structured cue word framework, detailed second cue words describing building attributes are generated, driving the image diffusion model to deepen and iterate through multiple attributes of the initial building images, ultimately outputting highly controllable and professional target building images in terms of details, materials, and functions. Therefore, this solution can effectively achieve automated and controllable image generation from macro-planning to micro-detail, effectively improving the technical difficulties of poor professional adaptability, low generation efficiency, and limited solutions in traditional methods for visualizing extraterrestrial architecture.
[0062] Corresponding to the method for generating images of extraterrestrial structures provided in the foregoing embodiments, such as Figure 4 As shown, this embodiment provides an apparatus for generating images of extraterrestrial architecture, which includes the following modules: The first prompt word generation module 410 is used to obtain an input statement describing the alien planet's architecture and convert the input statement into a first prompt word that matches the alien planet's environment through a preset structured prompt word framework. The first image generation module 420 is used to input the geospatial image of the surface of an extraterrestrial planet and the first prompt word into a pre-trained image diffusion model, and generate multiple sets of initial building images through the image diffusion model; The second prompt word generation module 430 is used to generate second prompt words for describing building attributes through the structured prompt word framework; The second image generation module 440 is used to generate multiple sets of target building images based on the second prompt word and multiple sets of the initial building images using the image diffusion model.
[0063] The device provided in this embodiment has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0064] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Figure 5 As shown, the electronic device 500 includes one or more processors 501 and memory 502.
[0065] The processor 501 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 500 to perform desired functions.
[0066] The memory 502 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 501 may execute the program instructions to implement the method for generating images of extraterrestrial buildings according to the embodiments of this disclosure described above, and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.
[0067] In one example, the electronic device 500 may also include an input device 503 and an output device 504, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0068] In addition, the input device 503 may also include, for example, a keyboard, a mouse, etc.
[0069] The output device 504 can output various information to the outside, including determined distance information, direction information, etc. The output device 504 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0070] Of course, for the sake of simplicity, Figure 5 Only some of the components of the electronic device 500 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 500 may include any other suitable components depending on the specific application.
[0071] Furthermore, this embodiment also provides a computer-readable storage medium storing a computer program for executing the above-described method for generating images of extraterrestrial structures.
[0072] The present disclosure provides a computer program product for generating images of extraterrestrial structures, including a method, apparatus, electronic device, and medium. The program product includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0073] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0074] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating images of extraterrestrial architecture, characterized in that, The method includes: The system acquires an input statement describing the architecture of an alien planet and converts the input statement into a first prompt word that matches the environment of the alien planet using a preset structured prompt word framework. The geospatial image of the outer planet's surface and the first prompt word are input into a pre-trained image diffusion model, which generates multiple sets of initial building images. The structured cue word framework is used to generate second cue words to describe building attributes; The image diffusion model generates multiple sets of target building images based on the second prompt word and multiple sets of initial building images.
2. The method according to claim 1, characterized in that, The process of converting the input statement into a first prompt word that matches the environment of the alien planet using a preset structured prompt word framework includes: Based on a preset structured prompt framework and a preset mapping relationship between extraterrestrial environmental parameters and Earth building logic, the input statement generates a list of space usage functions and a task book that conforms to the extraterrestrial environment. According to the preset semantic alignment rules, the list of space usage functions and the task list are converted into the first prompt words.
3. The method according to claim 2, characterized in that, The generation of second cue words for describing building attributes through the structured cue word framework includes: Input contextual information into the structured prompt word framework; wherein, the contextual information includes at least one of the following: the list of space usage functions and the task book, the descriptive text of the initial building image, and user instructions; Based on the context information, the structured cue word framework generates a second cue word for describing the building attributes of the initial building image; wherein the building attributes include at least: building materials, structural form, skin form and background elements that match the environment of the alien planet.
4. The method according to claim 1, characterized in that, The process involves inputting a geospatial image of the extraterrestrial surface and the first prompt word into a pre-trained image diffusion model, which then generates multiple sets of initial building images, including: The geospatial image of the exoplanet's surface and the first prompt word are input into a pre-trained image diffusion model; The image diffusion model is used to convert the geospatial image into a terrain semantic map, and the first prompt word is converted into a semantic vector. Multiple sets of initial building images are generated based on the terrain semantic map and the semantic vector.
5. The method according to claim 1, characterized in that, The image diffusion model is obtained by fine-tuning a pre-trained stable diffusion model based on the LoRA model.
6. The method according to claim 1 or 5, characterized in that, The method further includes: With the weight parameters of a pre-trained stable diffusion model fixed, the LoRA model is injected into the cross-attention layer of the U-Net network; The pre-acquired sample images are input into a stable diffusion model injected with the LoRA model, and the predicted images are output. The loss value between the sample image and the predicted image is determined using a preset loss function; The parameters in the LoRA model are updated by backpropagation based on the loss value to obtain the LoRA fine-tuning update parameters; Based on the LoRA fine-tuning update parameters, return to the step of inputting the sample image with the above text labels into the stable diffusion model injected with the LoRA model and outputting the predicted image, until the loss value is less than the preset loss threshold, and output the LoRA model parameters. The LoRA model parameters are subtracted from the weight parameters of the stable diffusion model to obtain the updated stable diffusion model, and the updated stable diffusion model is determined as the image diffusion model.
7. The method according to claim 6, characterized in that, The method further includes: Acquire sample images, including: aerospace industrial style images and architectural effect images; The sample images are standardized and text labels describing the image information are added.
8. A device for generating images of extraterrestrial architecture, characterized in that, The device includes: The first prompt word generation module is used to obtain the input statement used to describe the alien planet's architecture, and convert the input statement into a first prompt word that matches the alien planet's environment through a preset structured prompt word framework; The first image generation module is used to input the geospatial image of the surface of the extraterrestrial planet and the first prompt word into a pre-trained image diffusion model, and generate multiple sets of initial building images through the image diffusion model; The second prompt word generation module is used to generate second prompt words for describing building attributes through the structured prompt word framework; The second image generation module is used to generate multiple sets of target building images based on the second prompt word and multiple sets of the initial building images using the image diffusion model.
9. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a terminal device, cause the terminal device to perform the method as described in any one of claims 1-7.