Multi-dimensional space-time city scene construction method and system based on generative model

Through the multi-dimensional space-time urban market landscape construction method based on the generative model, traditional block landscape with artistic style is generated and displayed, which solves the problem of poor user experience in the existing technology, and achieves a more efficient and user-friendly urban market landscape experience.

CN120107526APending Publication Date: 2025-06-06FUTURE CITY (SHANGHAI) ARCHITECTURAL PLANNING & DESIGN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510164453.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the existing technology, the protection and renewal methods of old urban areas and characteristic ancient towns have a long cycle and high cost, making it difficult to adapt to the diverse needs of modern tourists, resulting in a poor user experience.

Method used

A multi-dimensional space-time city market scene construction method based on a generative model is adopted, and the target model information and target style image information of the target building are obtained, and the optimization model information is generated and sent to the user terminal to realize the display on the AR device.

Benefits of technology

This method can quickly and accurately generate traditional block landscapes with artistic style, significantly improve the user experience, and solve the problem of poor user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005272034430000011
    Figure HDA0005272034430000011
  • Figure HDA0005272034430000012
    Figure HDA0005272034430000012
  • Figure HDA0005272034430000021
    Figure HDA0005272034430000021
Patent Text Reader

Abstract

The invention is suitable for the technical field of city design, and provides a multi-dimensional space-time city scene construction method and system based on a generative model, and the method comprises the steps: obtaining target model information and target style image information of a target building; the method comprises the following steps: firstly, obtaining target style image information, then migrating the ink style of the target style image information to target model information, quickly and accurately generating optimization model information, and finally sending the optimization model information to a specified user terminal. According to the method, the traditional block with a unique artistic style can be presented on the AR equipment of the user, various design options are provided for the building facade, meanwhile, the characteristic style of the historical block is continued in a digitization mode, new cultural vitality is injected into the ancient town block, and the user experience is improved. And the understanding of residents and tourists on urban cultural heritage protection is improved unconsciously, so that the tourists have a deep sense of belonging to the block, and the construction of the smart city is effectively promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of urban design, and in particular to a method and system for constructing a multi-dimensional spatiotemporal urban scene based on a generative model. Background Art

[0002] Old town areas and characteristic ancient town blocks are often the cultural business cards of a city, containing rich cultural heritage and unique regional features. However, with the advancement of urbanization, these traditional blocks are facing the dilemma of declining vitality. In addition, the aesthetic needs of modern tourists are also constantly increasing. They are eager to immerse themselves in urban spaces rich in history and cultural connotations.

[0003] At present, traditional protection and renewal methods often have long cycles and high costs, making it difficult to adapt to the diverse needs of tourists in a timely manner. They also have problems with poor user experience and need further improvement. Summary of the invention

[0004] Based on this, an embodiment of the present application provides a method and system for constructing a multi-dimensional spatiotemporal urban scene based on a generative model to solve the problem of poor user experience in the prior art.

[0005] In a first aspect, an embodiment of the present application provides a method for constructing a multidimensional spatiotemporal urban scene based on a generative model, the method comprising:

[0006] Obtain target model information and target style image information of the target building;

[0007] Generate optimization model information according to the target model information and the target style image information;

[0008] The optimization model information is sent to a designated user terminal.

[0009] Compared with the prior art, the beneficial effects are as follows: in the multi-dimensional spatiotemporal urban scene construction method based on the generative model provided in the embodiment of the present application, the terminal device can first obtain the target model information and the target style image information of the target building, and then migrate the ink style of the target style image information to the target model information, quickly and accurately generate the optimized model information, and finally send the optimized model information to the designated user terminal, thereby displaying the traditional block with artistic style on the user's AR device, greatly improving the user experience, and to a certain extent solving the current problem of poor user experience.

[0010] In a second aspect, an embodiment of the present application provides a system for constructing a multi-dimensional spatiotemporal urban scene based on a generative model, the system comprising:

[0011] Target model information acquisition module: used to obtain target model information and target style image information of the target building;

[0012] Optimization model information generation module: used to generate optimization model information according to the target model information and the target style image information;

[0013] Optimization model information sending module: used to send the optimization model information to a designated user terminal.

[0014] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of the first aspect described above when executing the computer program.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method of the first aspect described above are implemented.

[0016] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0018] Figure 1 It is a flowchart of a method for constructing a multi-dimensional spatiotemporal city scene provided by an embodiment of the present application;

[0019] Figure 2 It is a schematic diagram of the process before step S100 in the multi-dimensional spatiotemporal city scene construction method provided in one embodiment of the present application;

[0020] Figure 3 is a schematic diagram of using a drone to perform urban street scene modeling according to an embodiment of the present application;

[0021] Figure 4 It is a flowchart after step S103 in the multi-dimensional spatiotemporal city scene construction method provided in one embodiment of the present application;

[0022] Figure 5 It is a schematic diagram of the LoRA model training process provided by an embodiment of the present application;

[0023] Figure 6 It is a schematic diagram of a loss curve during LoRA model training provided in an embodiment of the present application;

[0024] Figure 7 It is a flowchart after step S105 in the multi-dimensional spatiotemporal city scene construction method provided in one embodiment of the present application;

[0025] Figure 8 is a schematic diagram of a target style image information generation process based on AI GC provided in an embodiment of the present application;

[0026] Fig. 9 is a schematic diagram of a Unity3D scene setting provided in an embodiment of the present application;

[0027] Fig.10 is a schematic diagram of a real AR application of a user provided by an embodiment of the present application;

[0028] Fig.11 It is a flowchart after step S300 in the multi-dimensional spatiotemporal city scene construction method provided in one embodiment of the present application;

[0029] Fig.12 It is a module block diagram of a multi-dimensional spatiotemporal city scene construction system provided by an embodiment of the present application;

[0030] Fig.13 It is a schematic diagram of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0031] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0032] In the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0033] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0034] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.

[0035] See also Figure 1 , Figure 1 : is a flow chart of a method for constructing a multi-dimensional spatio-temporal city scene based on a generative model provided in an embodiment of the present application. In this embodiment, the execution subject of the method for constructing a multi-dimensional spatio-temporal city scene is a terminal device. It is understood that the types of terminal devices include but are not limited to mobile phones, tablet computers, laptop computers, ultra-mobile personal computers (Ultra-Mobile Personal Computer, UMPC), netbooks, personal digital assistants (Personal Digital Assistant, PDA), etc., and the embodiments of the present application do not impose any restrictions on the specific types of terminal devices.

[0036] See also Figure 1 The multi-dimensional spatiotemporal city scene construction method provided in the embodiment of the present application includes but is not limited to the following steps:

[0037] In S100, target model information and target style image information of a target building are obtained.

[0038] Specifically, the terminal device can first obtain the target model information and target style image information of the target building, where the target building is used to describe a building in a traditional block, such as a temple, ancestral hall or wall; the target model information is used to describe a high-precision three-dimensional model constructed for the target building; the target style image information is used to describe a specified ink style image.

[0039] In some possible implementations, in order to obtain effective target model information and target style image information, please refer to Figure 2 Before step S100, the method further includes but is not limited to the following steps:

[0040] In S101, based on a preset scanning radar, laser scanning data information of the target building is obtained, and based on a preset camera, two-dimensional detailed image information of the target building is obtained.

[0041] Without loss of generality, the scanning radar can be pre-installed at a designated location of the UAV; the laser scanning data information is used to describe the data obtained by the scanning radar scanning the target building; and the two-dimensional detail image information is used to describe the data obtained by the camera photographing the target building.

[0042] For example, see Figure 3 The terminal device can first use the scanning radar in the drone to perform multi-angle high-resolution scanning of the target building, and effectively obtain the laser scanning data information of the target building; at the same time, the terminal device can obtain the two-dimensional detailed image information of the target building based on the preset camera, thereby combining ground photography to ensure that the structural details of the target building are fully reflected in the data set.

[0043] In S102, initial three-dimensional model information is generated according to the laser scanning data information.

[0044] Specifically, after the terminal device obtains the laser scanning data information, the terminal device can input the laser scanning data information into a preset reconstruction software to generate initial three-dimensional model information, thereby obtaining a highly detailed building model, wherein the reconstruction software can be RealityCapture or Pix4D; the initial three-dimensional model information is used to describe the three-dimensional model reconstructed based on the laser scanning data information.

[0045] In S103, the two-dimensional detail image information is texture mapped to the initial three-dimensional model information to generate target model information.

[0046] Specifically, after the terminal device generates the initial three-dimensional model information, the terminal device can first perform projection extraction on each front elevation of the initial three-dimensional model information, and then texture map the two-dimensional detail image information to each front elevation of the initial three-dimensional model information to generate target model information, so as to realize style transfer generation processing, ensure that the subsequent generated content accurately matches the real scene, provide a basis for the input image of the generative AI, and ensure that the virtual content in the block is consistent with the building structure.

[0047] In some possible implementations, in order to facilitate the subsequent migration of the ink painting style to the target building, please refer to Figure 4 After step S103, the method further includes but is not limited to the following steps:

[0048] In S104 , a plurality of reference style image information is obtained.

[0049] Specifically, the terminal device may first obtain a plurality of reference style image information, wherein the reference style image information is used to describe a reference image with an ink painting style, and the reference style image information may be collected from the art works of writers such as Wu Guanzhong.

[0050] In S105 , an untrained fine-tuning model is trained according to a plurality of reference style image information to generate a pre-trained fine-tuning model.

[0051] For example, see Figure 5 , the fine-tuning model can be a low-rank adaptation (LoRA) model; after the terminal device obtains multiple baseline style image information, the terminal device can use the multiple baseline style image information to train the untrained fine-tuning model to generate a pre-trained fine-tuning model, which is a pre-trained ink style LoRA model, so that the subsequently generated images have the visual characteristics of Wu Guanzhong's ink style, optimize the details of the subsequently generated images, and ensure the consistency of the artistic style and the beauty of the visual effect. Without loss of generality, please refer to Figure 6 , Figure 6 Used to show the loss curve of the LoRA model during training, Figure 6 It can be seen that the LoRA model has high stability and high convergence.

[0052] In some possible implementations, in order to make the generated stylized images better fit the cultural style of the target model information and provide content materials for augmented reality display, please refer to Figure 7 After step S105, the method further includes but is not limited to the following steps:

[0053] In S106, the target model information is projected to generate building facade information.

[0054] Specifically, the terminal device may first perform projection processing on the target model information to generate building facade information, wherein the building facade information is used to describe the front elevation image of the target model information.

[0055] In S107, the building facade information is input into the image generation model, and the image generation model is controlled based on the control network and preset structural constraints, and the internal parameters of the image generation model are adjusted based on the fine-tuning model to construct the target style image information.

[0056] For example, see Figure 8After the terminal device generates the building facade information, the terminal device can input the building facade information into the image generation model, and the image generation model can be a generative AI model, such as Stab le Diffusion; at the same time, the terminal device can control the image generation model based on the control network and preset structural constraints to ensure that the generated ink style image is accurately located at the building elements, such as roofs, windows and doors, and the control network is Control Net; and the terminal device can adjust the internal parameters of the image generation model based on the fine-tuning model to effectively construct the target style image information.

[0057] In S200, optimized model information is generated according to the target model information and the target style image information.

[0058] Specifically, after the terminal device obtains the target model information and the target style image information, the terminal device can migrate the target style image information with ink style to each front facade image of the target model information according to the target model information and the target style image information, and generate optimized model information, thereby combining the actual block scene with the virtual stylized content, constructing the target model information with ink style, and greatly enhancing the attractiveness and cultural expression of the traditional block.

[0059] In S300, the optimization model information is sent to a designated user terminal.

[0060] Specifically, the user terminal can be a mobile device in the user's hand or a lightweight AR glasses being worn; after the terminal device generates the optimization model information, the terminal device can send the optimization model information to the designated user terminal, thereby realizing the application of the generated ink style image to the augmented reality system, which is conducive to displaying virtual content in the actual block through AR glasses, so that users can experience stylized dynamic landscapes in the real block.

[0061] For example, see Fig. 9 and Fig.10 ,In order to ensure the stability and clarity of the generated content in the user’s field of vision, the terminal device can control the ,lightweight AR glasses to load the generated ink-style images, and then use geo-positioning technology to ,further accurately overlay the virtual images on the real scene.

[0062] In one possible implementation, the terminal device can also collect user feedback through questionnaires and interaction records to evaluate the attractiveness and experience quality of virtual content to tourists, facilitate quantitative analysis of feedback data, and optimize the display effect of the system in combination with user evaluations, thereby providing data support for future optimization iterations.

[0063] In some possible implementations, in order to optimize and iterate the multi-dimensional spatiotemporal urban scene construction method, please refer to Fig.11 After step S300, the method further includes but is not limited to the following steps:

[0064] In S400, based on a preset monitoring tool, delay rate information, frame rate information, and fluency information are obtained.

[0065] Specifically, the terminal device can first use the preset monitoring tool to obtain delay rate information, frame rate information and fluency information, among which the delay rate information is used to describe the time delay rate of the system's response to the request; the frame rate information is used to describe the number of frames processed by the system per second; the fluency information is used to describe the smoothness score of the user experience, and the fluency information can be obtained from user feedback.

[0066] In S410, real-time performance analysis data packet information is generated according to the delay rate information, the frame rate information and the fluency information.

[0067] Specifically, after the terminal device obtains the delay rate information, frame rate information and smoothness information, the terminal device can generate real-time performance analysis data packet information based on the delay rate information, frame rate information and smoothness information, which is conducive to evaluating the real-time performance of the generative AI and AR integrated system.

[0068] The implementation principle of the multi-dimensional spatiotemporal urban scene construction method based on the generative model in the embodiment of the present application is: the terminal device can first obtain the target model information and the target style image information of the target building, and then migrate the ink style of the target style image information to the target model information, quickly and accurately generate the optimized model information, and finally send the optimized model information to the designated user terminal, so as to display the old town or characteristic ancient town block with artistic style on the user's AR device, greatly improving the user experience.

[0069] It should be noted that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0070] The embodiment of the present application also provides a multi-dimensional spatiotemporal city scene construction system. The multi-dimensional spatiotemporal city scene construction system based on the generative model is shown for the sake of convenience. Fig.12 As shown, the system 120 includes:

[0071] Target model information acquisition module 121: used to acquire target model information and target style image information of a target building;

[0072] Optimization model information generation module 122: used to generate optimization model information according to target model information and target style image information;

[0073] The optimization model information sending module 123 is used to send the optimization model information to a designated user terminal.

[0074] Optionally, the system 120 further includes:

[0075] Laser scanning data information acquisition module: used to acquire laser scanning data information of the target building based on a preset scanning radar, and to acquire two-dimensional detailed image information of the target building based on a preset camera;

[0076] Initial 3D model information generation module: used to generate initial 3D model information according to laser scanning data information;

[0077] Target model information generation module: used to map the texture of the two-dimensional detail image information to the initial three-dimensional model information to generate the target model information.

[0078] Optionally, the system 120 further includes:

[0079] Baseline style image information acquisition module: used to obtain multiple baseline style image information;

[0080] Fine-tuning model generation module: used to train an untrained fine-tuning model based on multiple benchmark style image information to generate a pre-trained fine-tuning model.

[0081] Optionally, the system 120 further includes:

[0082] Building facade information generation module: used to project the target model information and generate building facade information;

[0083] Target style image information construction module: used to input building facade information into the image generation model, and control the image generation model based on the control network and preset structural constraints, and adjust the internal parameters of the image generation model based on the fine-tuning model to construct the target style image information.

[0084] Optionally, the system 120 further includes:

[0085] Delay rate information acquisition module: used to obtain delay rate information, frame rate information and fluency information based on preset monitoring tools;

[0086] Real-time performance analysis data packet information generation module: used to generate real-time performance analysis data packet information according to delay rate information, frame rate information and fluency information.

[0087] It should be noted that the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0088] The present application also provides a terminal device, such as Fig.13 As shown, the terminal device 130 of this embodiment includes: a processor 131, a memory 132, and a computer program 133 stored in the memory 132 and executable on the processor 131. When the processor 131 executes the computer program 133, the steps in the above-mentioned multi-dimensional spatiotemporal city scene construction method embodiment are implemented, such as Figure 1 Steps S100 to S300 shown; or, when the processor 131 executes the computer program 133, the functions of each module in the above device are realized, for example Fig.12 The functions of modules 121 to 123 are shown.

[0089] The terminal device 130 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server, and the terminal device 130 includes but is not limited to a processor 131 and a memory 132. Those skilled in the art will appreciate that Fig.13 It is only an example of the terminal device 130 and does not constitute a limitation of the terminal device 130. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device 130 may also include input and output devices, network access devices, buses, etc.

[0090] Among them, the processor 131 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.; the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0091] The memory 132 may be an internal storage unit of the terminal device 130, such as a hard disk or memory of the terminal device 130, or an external storage device of the terminal device 130, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. equipped on the terminal device 130; further, the memory 132 may also include both an internal storage unit and an external storage device of the terminal device 130, and the memory 132 may also store a computer program 133 and other programs and data required by the terminal device 130, and the memory 132 may also be used to temporarily store data that has been output or is to be output.

[0092] One embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of each of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc.; the computer-readable medium may include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0093] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, all equivalent changes made according to the methods, principles, and structures of the present application should be included in the protection scope of the present application.

Claims

1. A method for constructing a multidimensional spatiotemporal urban scene based on a generative model, characterized in that: The method comprises: Obtain target model information and target style image information of the target building; Generate optimization model information according to the target model information and the target style image information; The optimization model information is sent to a designated user terminal.

2. The method according to claim 1, characterized in that Before acquiring the target model information and the target style image information of the target building, the method further includes: Based on the preset scanning radar, the laser scanning data information of the target building is obtained, and based on the preset camera, the two-dimensional detailed image information of the target building is obtained; Generating initial three-dimensional model information according to the laser scanning data information; The two-dimensional detail image information is texture mapped to the initial three-dimensional model information to generate the target model information.

3. The method according to claim 2, characterized in that After mapping the texture of the two-dimensional detail image information to the initial three-dimensional model information to generate the target model information, the method further includes: Obtain multiple baseline style image information; According to the plurality of reference style image information, an untrained fine-tuning model is trained to generate a pre-trained fine-tuning model.

4. The method according to claim 3, characterized in that After training an untrained fine-tuning model according to the plurality of reference style image information to generate a pre-trained fine-tuning model, the method further includes: Performing projection processing on the target model information to generate building facade information; The building facade information is input into a picture generation model, and based on a control network and preset structural constraints, the picture generation model is controlled, and based on the fine-tuning model, internal parameters of the picture generation model are adjusted to construct target style image information.

5. The method according to claim 1, characterized in that After sending the optimization model information to a designated user terminal, the method further includes: Based on the preset monitoring tools, obtain the delay rate information, frame rate information and fluency information; Real-time performance analysis data packet information is generated according to the delay rate information, the frame rate information and the fluency information.

6. A multi-dimensional spatiotemporal urban scene construction system based on a generative model, characterized in that: The system comprises: Target model information acquisition module: used to obtain target model information and target style image information of the target building; Optimization model information generation module: used to generate optimization model information according to the target model information and the target style image information; Optimization model information sending module: used to send the optimization model information to a designated user terminal.

7. The system according to claim 6, characterized in that The system further comprises: Laser scanning data information acquisition module: used to acquire laser scanning data information of the target building based on a preset scanning radar, and to acquire two-dimensional detailed image information of the target building based on a preset camera; An initial three-dimensional model information generation module is used to generate initial three-dimensional model information according to the laser scanning data information; Target model information generation module: used for mapping the texture of the two-dimensional detail image information to the initial three-dimensional model information to generate the target model information.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.