Urban development plan proposal support method, information processing device and computer program
The information processing device generates urban development images and videos based on user input, addressing the lack of support for town planning concepts and enabling user-driven visualization and sharing of urban development plans.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2025-12-23
- Publication Date
- 2026-07-07
Smart Images

Figure 2026113439000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a method for supporting the proposal of town planning concepts, an information processing apparatus, and a computer program.
Background Art
[0002] There is a technique for generating a vacant house age prediction heat map using a machine learning algorithm (for example, Patent Document 1). The information processing apparatus according to Patent Document 1 inputs, as input data, at least one of the construction year of a house, household information of the house, the elapsed years of the infrastructure of the house, and the usage amount of infrastructure resources essential for life in the house, and uses, as teacher data, a learning model that performs vacant house prediction using the vacant house information in the vacant house database, thereby predicting the vacant house age of the house. A prediction unit that predicts the vacant house age of the house, and an image generation unit that overlays a vacant house age prediction heat map generated based on the prediction result by the prediction unit on a map.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, the information processing apparatus described in Patent Document 1 is not an apparatus for supporting the proposal of town planning concepts.
[0005] An object of the present disclosure is to provide a method for supporting the proposal of town planning concepts, an information processing apparatus, and a computer program that support the proposal of town planning concepts that visualize the town planning envisioned by users.
Means for Solving the Problems
[0006] A method for supporting urban development concept proposals according to one aspect of this disclosure involves acquiring request information that expresses the desired urban development, acquiring location information related to the desired urban development, acquiring material images corresponding to the acquired location information, and inputting the acquired request information and material images into a generation model to generate an urban development image or urban development video that reflects the content of the request information.
[0007] An information processing device according to one aspect of the present disclosure is an information processing device comprising a communication unit and a control unit, wherein the control unit acquires request information expressing a desired urban development, acquires location information related to the desired urban development, acquires material images corresponding to the acquired location information, and inputs the acquired request information and the material images into a generation model to generate an urban development image or urban development video that reflects the content of the request information.
[0008] A computer program according to one aspect of this disclosure causes a computer equipped with a communication unit and a control unit to execute a process that generates an image or video of a town development that reflects the content of the request information by acquiring request information that expresses a desired town development, acquiring location information related to the desired town development, acquiring material images corresponding to the acquired location information, and inputting the acquired request information and material images into a generation model. [Effects of the Invention]
[0009] According to this disclosure, it is possible to support the proposal of urban development concepts that visualize the urban development that users dream of. [Brief explanation of the drawing]
[0010] [Figure 1] This is a schematic diagram illustrating an example of the configuration of the urban development concept proposal support system according to this embodiment. [Figure 2] This is a block diagram showing an example configuration of the information processing device according to this embodiment. [Figure 3] This is a conceptual diagram showing an example of a database configuration. [Figure 4]This is a conceptual diagram illustrating the method for supporting urban development concept proposals according to this embodiment. [Figure 5] This flowchart shows the processing procedure for supporting urban development concept proposals according to this embodiment. [Figure 6] This is an example of a dream plan image. [Figure 7] This is a schematic diagram showing an example of the output screen of the urban development concept sharing platform. [Figure 8] This is a conceptual diagram illustrating the method for selecting dream plan images and dream plan videos based on layer type. [Figure 9] This is the first conceptual diagram showing the functional configuration of the urban development concept proposal support system. [Figure 10] This is the second conceptual diagram showing the functional configuration of the urban development concept proposal support system. [Figure 11] This is the third conceptual diagram showing the functional configuration of the urban development concept proposal support system. [Modes for carrying out the invention]
[0011] A town planning concept proposal support system, a town planning concept proposal support method, an information processing device, and a computer program according to embodiments of this disclosure will be described below with reference to the drawings. This disclosure is not limited to these examples, but is intended to include all modifications within the meaning and scope of the claims, as indicated by the claims. Furthermore, at least some of the embodiments described below may be combined in any way.
[0012] The purpose of this disclosure is to provide a concept proposal support method (hereinafter referred to as the "community development concept proposal support method"), an information processing device, and a computer program that support users in planning and developing community-based concept proposals by visualizing images of townscapes, products, services, etc., that they have envisioned and planned, and in order to share those images with others. This method of disclosure will also contribute to "promoting a self-governing society led by residents" and "democratizing community development."
[0013] The types of users include residents, the government, businesses, educational institutions, etc. Through the method for supporting the proposal of town planning concepts, each user can enjoy the following advantages. Residents can propose town planning concepts such as visualizing the image of the townscape they have envisioned and planned. When the government aims to formulate a resident participation-type town planning concept or publicize regional issue-solving themes such as disaster prevention and disaster response, and countermeasures for vacant houses, it can provide information and explanations to share the image with residents by visualizing it. Businesses can propose visualizations of products and services related to town planning that are rooted in the region and carried out as economic activities. Educational institutions can conduct active learning through visualization, such as town planning education, SDGs education, and exploratory learning through school education and local groups.
[0014] FIG. 1 is a schematic diagram for explaining a configuration example of a town planning concept proposal support system according to the present embodiment. The town planning concept proposal support system includes an information processing device 1, a generation AI server 2, a map server 3, a virtual globe server 4, and a user communication terminal 5. The various devices are communicably connected via a communication network N such as the Internet.
[0015] FIG. 2 is a block diagram showing a configuration example of the information processing device 1 according to the present embodiment. The information processing device 1 is a computer including a processing unit 11, a storage unit 12, and a communication unit 13. Note that the information processing device 1 may be configured by a plurality of computers for distributed processing, may be realized by a plurality of virtual machines provided in one server, may be realized using a cloud server, or a part thereof may be configured by a quantum computer.
[0016] The processing unit 11 has one or more arithmetic processing devices such as a CPU (Central Processing Unit) and an MPU (Micro-Processing Unit), and executes the town planning concept provision support processing according to the present embodiment by reading and executing the computer program 12a stored in the storage unit 12.
[0017] The storage unit 12 is a storage device such as a hard disk, an EEPROM (Electrically Erasable Programmable Read-Only Memory), or a flash memory. The storage unit 12 stores a computer program 12a necessary for the processing unit 11 to execute the town planning concept providing support process. Further, the storage unit 12 stores a base map 12b, a user DB 12c, a concept DB 12d, and an evaluation DB 12e. Note that the computer program 12a may be recorded in a computer-readable manner on the recording medium 10. The storage unit 12 stores the computer program 12a read from the recording medium 10 by a reading device (not shown). The recording medium 10 is a semiconductor memory such as a flash memory, an optical disk, a magnetic disk, a magneto-optical disk, or the like. Also, it may be an aspect in which various programs according to the present embodiment are downloaded from an external server (not shown) connected to the communication network N and stored in the storage unit 12. Furthermore, it may be an aspect in which the computer program 12a is stored in a server on the cloud.
[0018] The communication unit 13 includes a communication circuit that communicates with the generation AI server 2, the map server 3, the virtual globe server 4, the user communication terminal 5, etc. via the communication network N. The processing unit 11 transmits data related to the town planning concept via the communication unit 13.
[0019] The map server 3 is a server that provides map information. The map server 3 stores map information and an image associated with position information indicating a position on the earth. The image includes an image posted by the user to the map server 3. The image is, for example, an image obtained by imaging the location using a camera. The map server 3 provides an image associated with the map based on the position information. When the information processing device 1 executes processing related to the urban development plan, it requests images associated with location information from the map server 3 based on the location information. The map server 3 has the function of sending images associated with location information to the information processing device 1 in response to the request from the information processing device 1. Hereinafter, the images that the information processing device 1 requests from the map server 3 and obtains will be referred to as source images.
[0020] Virtual globe server 4 stores aerial images of various locations and provides aerial images of the Earth viewed from any point and altitude. Virtual globe server 4 can also provide animated videos (hereinafter referred to as "aerial videos") of any point, such as circling from above, zooming in, or zooming out. When the information processing device 1 executes processing related to the urban development plan, it requests an overhead image or overhead video of the location indicated by the location information from the virtual globe server 4. The virtual globe server 4 has the function of transmitting an overhead image or overhead video of the location indicated by the location information to the information processing device 1 in response to the request from the information processing device 1.
[0021] The user communication terminal 5 is a terminal such as a smartphone, tablet PC, or desktop PC with GPS functionality, which is connected to the communication network N and capable of sending and receiving data via the communication network N. The user communication terminal 5 comprises, for example, a control unit including one or more processors such as CPUs and MPUs, an input device such as a keyboard or mouse, a display device, a communication circuit for communicating with the information processing device 1, and a GPS receiver. The user communication terminal 5 in this embodiment is a terminal mainly used by ordinary citizens and sends and receives data related to urban development plans with the information processing device 1. Specifically, the user communication terminal 5 transmits request information expressing the desired urban development to the information processing device 1. Furthermore, the GPS function can be replaced with other satellite positioning systems, or it can be an integration of other positioning systems with other technologies.
[0022] Request information includes text data that expresses in language the atmosphere, name, ideals, etc., related to urban development desired by the user. For example, if there is an urban development plan by the government or private business, or products / services related to urban development, the request information expresses what kind of urban landscape the user wants, what kind of products / services would be desirable, the user's dreams for urban development, desired development plan image, or products / services. Hereinafter, such request information will be referred to as dream plan information. The request information may also be in audio format. The information processing device 1 simply needs to convert the audio request information into text request information. Request information can be image data, such as a user's hand-drawn sketch or an existing photograph / image. Request information can be in the form of text data generated using brain technology, which converts visual images conjured by a brain sensor worn by the user into text, or text data generated using a speech recognition app for people with speech impairments. In particular, it is an effective method for conveying requests for information for people with hearing impairments or speech difficulties, and combined with the effectiveness of voice data input support for the elderly and others, obtaining requests for information is a good way to provide reasonable accommodations in terms of usability and accessibility for people with various disabilities and those who have difficulty communicating.
[0023] The generation AI server 2 has a document generation model 21, an image generation model 22, and a video generation model 23. The hardware configuration of the generation AI server 2 is the same as that of the information processing device 1. Each model constitutes a so-called text generation AI, image generation AI, and video generation AI, and the generation AI server 2 provides the information processing device 1 with an API (Application Programming Interface) for the generation AI. The generation AI server 2 may be configured with each generation AI on a separate server, as a virtual machine, or as a virtual server packaged as a container. Furthermore, the information processing device 1 may be configured to include some or all of the document generation model 21, image generation model 22, and video generation model 23. The generation AI server 2 may also be a "multimodal" generation AI capable of processing various formats such as text, images, audio, and video, like "Gemini," or an "omnimodal" generation AI. It may also be an autonomously operating generation AI, such as an AI agent. Furthermore, each generating AI is trained not to output documents, images, and videos containing expressions with political connotations, antisocial connotations, expressions that promote discrimination or prejudice, expressions with violent connotations, expressions that may infringe on the rights, information, or privacy of specific individuals, expressions that may infringe on the rights or information of specific corporations, photographs, illustrations, comics, etc. that may infringe on copyright, video expressions that use part or all of photographs taken or videos created by others posted on social media, and other expressions that slander, insult, libel, discriminatory language or insults, offensive, sexual or exclusive expressions, baseless lies or rumors, or controversial content against specific individuals, companies, groups, religions, beliefs, ethnic groups, etc. All or part of these may be included as exclusion elements in the system instructions.
[0024] The document generation model 21 is a generative AI model that, when given dream plan information expressing a desired urban development, generates and outputs a prompt for image generation that appropriately reflects the dream plan information. The document generation model 21 is, for example, an API (Application Programming Interface) related to a Large Language Model (LLM). A Large Language Model is, for example, a machine learning model that uses a mechanism called a Transformer. The prompts generated by the document generation model 21 are instructions for the image generation model 22, described later, to accurately generate urban development images (hereinafter referred to as "dream plan images") that reflect the content of the dream plan information. Furthermore, the document generation model 21 can obtain request information in a dialogue format with the user in order to obtain information for generating prompts. For example, the system may be configured to collect dream plan information by instructing the document generation model 21 to generate prompts for generating images related to urban development using system instructions, to prepare predetermined anticipated questions in advance and have the model respond to them, to ask a predetermined number of questions, for example, 3 to 5 times, and to include predetermined questions. Request information is obtained primarily through conversational interactions via user communication terminal devices. For general users with limited specialized knowledge, it is best to utilize chatbots to obtain request information through conversation. The conversational format can be text input, voice, or images. A good procedure for a chatbot conversation would be: User information → Skill level mode selection → Location information → Target category → Target information → Required equipment → Outcome information → Persona setting → Environment information → Style expression information → Composition angle → Thumbnail display information. By allowing users to select from five generation modes tailored to their individuality, preferences, and skill levels, it would be beneficial to address the indigestion experienced by beginners and those with lower skill levels, as well as the lack of challenge for intermediate and advanced users, thereby increasing the likelihood of sustained enthusiasm and motivation. • Beginners: Minimize the number of choices and the content of the conversation, and aim for a single generated image as the answer. • Beginner level: Several items, open-ended answers → It is best to have users simply answer / select from several information options to generate the data. Intermediate level: All items, open-ended answers → It is best to have users answer, select, and modify all information options. • Intermediate to advanced users: Adjust all items → It is best to generate and correct all information options by answering, selecting, and adjusting them. • Advanced: It is best to fine-tune all items → generate and modify all information options by answering, selecting, and fine-tuning them. It would be beneficial to use a chatbot that combines scenario-based and AI-based learning through Retrieval-Augmented Generation (RAG) or fine-tuning to repeatedly search and answer questions regarding urban development proposals, not only with existing large-scale language models (LLMs), but also by storing contexts and scenarios specific to the urban development field in a local database beforehand (for example, as items in an "urban development proposal library," such as categories like buildings, spaces, and sections / urban development-related requirements / place names, characteristics, and facility names of the usage area / urban planning of the user administration, etc.). The chatbot would then repeatedly search and answer through a knowledge-based chatbot that combines scenario-based and AI-based learning, such as Retrieval-Augmented Generation (RAG), which selects and presents the most suitable answer to the user's request, and then generates and answers the most appropriate prompt sentences after contextual correction suitable for local urban development. The contexts and scenarios specific to the urban development field, which are stored in the local database, can be pre-listed by category, or they can be obtained by generating AI to estimate category items and specific content, or they can be obtained by crawling related web pages estimated by the generating AI. In particular, in local government administrations where regional identification is necessary, regionally specific names such as addresses, facility names, and road names, as well as administrative information such as urban planning and policy guidelines, are obtained by crawling relevant web pages and stored as region-specific contexts and scenarios, enabling contextual adjustments suitable for local urban development.
[0025] It is preferable to configure the document generation model 21 to interact with the user and acquire dream plan information by providing system instructions that include the objective of obtaining the following elements necessary for generating prompts suitable for generating images related to urban development. It is also preferable to generate prompts by providing system instructions to the document generation model 21 that instruct it to output prompts that include elements (1) to (3) and exclude element (4). (1) Basic elements (essential elements) Type and size of the town (city, rural town, village, metropolis, small town, etc.), architectural style (modern architecture, historical buildings, skyscrapers, wooden houses, brick buildings, etc.), era or period (medieval, modern, future, 19th century, Showa era, etc.), viewpoint and composition (aerial view, ground view, overhead view, view from an alley, etc.), weather and time of day (sunny, rainy, snowy, night, twilight, sunrise, etc.) (2) Detail elements (elements that enrich the expression) Natural elements (trees, rivers, mountains, sea, parks, etc.), transportation elements (cars, trains, buses, bicycles, pedestrians, etc.), sense of everyday life (people's activities, signs, laundry, trash, graffiti, etc.), colors and atmosphere (colorful, monochrome, nostalgic, futuristic, quiet, lively, etc.), place or landmark (Tokyo, New York, Paris, Eiffel Tower, Statue of Liberty, etc.) (3) Illustration elements (elements that specify the art style): Art style (realistic or photographic, illustrative, anime, watercolor, oil painting, game style, etc.) (4) Exclusion elements Words that deny any connection to urban development, negative expressions (e.g., "there is no...", "except..."), and contradictory information. Furthermore, user request information can include not only text, but also original sketches, images, photographs, and images of public spaces similar to the request, or a combination of these.
[0026] Furthermore, instead of using the document generation model 21, the system may be configured to obtain request information by preparing predetermined anticipated questions in advance and having the system respond to them.
[0027] The image generation model 22 is a generative AI model that, when a prompt generated by the document generation model 21 is input, generates and outputs a dream plan image that reflects the content of the dream plan information. The image generation model 22 includes, for example, a text encoder that converts the input prompt into low-dimensional vector information, and an image generator that generates an image based on that vector information. Furthermore, the image generation model 22 may be configured to generate and output a dream plan image that reflects the contents of the dream plan information when it receives a prompt generated by the document generation model 21 and a source image obtained from the map server 3. The image generation model 22 may incorporate adjustments and settings via an API when acquiring request information, such as using hyperparameters to adjust user request information and source images, setting parameters using Temperature and Top-p, or incorporating adjustments and settings via an API to approximate the image requested by the user. When modifying the generated image again, it is advisable to allow the user to adjust the degree of reproduction using parameters such as seed values to improve the reproducibility of the generated image. Image generation AI via APIs may use VAE (Variational Autoencoder: a variational autoencoder consisting of an encoder that takes in text and images, and a decoder that generates images), GAN (Generative Adversarial Networks: generative adversarial networks that pit generators and discriminators against each other), or CNN (Convolutional Neural Network: a convolutional neural network that extracts local features). If you want to modify an image generated by the user, for example, to ensure the reproducibility of the generated image, you can either use an already generated image as a new source image to generate a new image, or repeat this process. In urban development proposals, source images are often based on images of existing public spaces. Users may mask or crop specific parts of these images and then insert the resulting composite image for output. Alternatively, instead of specifying an image, users may describe the specific part using text and then output a generated image based on that description. Image generation models 22 include, for example, Stable Diffusion, Midjourney, Dall-E 3, Canva, ImageFX, and other generative AIs that combine these. As described above, the image generation model 22 generates images and videos from source information based on user requests. However, it is necessary to minimize instances where the generated images contradict the user's intentions or create a sense of incongruity between the requested object and the public space before compositing. As a method of correction, it is advisable to provide the generation agreement in advance as an arbitrary prompt in processing steps S19 and S21 of the information processing device 1 in Figure 5. For example, "When compositing the user's requested object onto a public space photograph of the provided location, do not alter the overall composition or characteristics outside the area to be replaced by the object. Instead, improve the overall sharpness and detail, maintain the original, natural color tone, avoid unnatural smoothing or processing, adjust the white balance to make the colors natural and vibrant (but not excessively vibrant), moderately increase the contrast and dynamic range, reduce noise, correct blur and softness, sharpen edges to create a clear and realistic impression, and the final image should be high resolution, faithful to the original public space, while the user's requested object is clearer and the intended meaning is conveyed to the viewer."
[0028] The video generation model 23 is a generation AI model that, when given a dream plan image generated by the image generation model 22, material images acquired from the map server 3, and an overhead video acquired from the virtual globe server 4 as input, generates and outputs a dream plan video that reflects the content of the dream plan information. The video generation model 23 may also be an AI model configured to generate and output a video with sound effects or music automatically added as background music (BGM) from sound source data, or a video with arbitrarily selected sound effects or music added. Examples of video generation models 23 include Runway Gen-2, Dream Machine, Pika, and Sora. The video generation model 23 can be, for example, a video generation AI such as Sora or Google Veo, a short clip AI for social media such as Canva or Filmora, Motionleap or MoviePic which move or animate specific areas, or a combination of various media elements such as still images, videos, audio, text, and animations, like a multimedia slideshow. To effectively express and communicate a user's concept proposal, it is effective to combine images and videos with user comments (text and audio) to concisely convey not only the current issues that led to the proposal, the purpose and selling points of the proposal, the expected effects, and the user's worldview and values.
[0029] Figure 3 is a conceptual diagram showing an example of a database (DB) configuration. User DB12c stores user information for the urban development concept proposal support system, associating each user's user ID, username, handle name, other basic user information, and user type information. The handle name is an arbitrary name assigned by the user. Basic user information includes the user's email address, etc. User type information indicates the user's attributes, such as individual, private business, and government agency. User type information is information that will be displayed and shared on the urban development concept sharing platform 6, which will be described later, and is also information for switching the display mode.
[0030] Concept DB12d stores the following information in association with the concept ID, user ID, title (proposal name), thumbnail document, posting date and time, location information, layer type, city classification, and dream plan image and dream plan video generated by the urban development concept proposal support system. The concept ID is an ID used to identify the dream plan image and dream plan video generated based on the dream plan information. The user ID is the ID of the user who provided the dream plan information. The title (proposal name) is the identifier name of the dream plan. The title is a name automatically generated from the administrative district and city classification to which the location indicated by the location information belongs. The thumbnail document is the title of the dream plan arbitrarily assigned by the user. The posting date and time is the date and time when the generated dream plan image and video were posted to the urban development concept sharing platform 6, which will be described later. The location information is information provided by the user along with the dream plan information, indicating the location targeted by the dream plan. Furthermore, in the resident layer, location information is limited to the town and block number (e.g., XX town, XX block) from the perspective of protecting personal information and privacy. On the map, it is represented as an approximate area enclosed by a circle with a diameter of approximately 100m, ensuring that the location of the dream plan does not indicate a specific point. The layer type indicates the map layer on which the dream plan images and videos are mapped. The town division indicates the division of buildings, facilities, plazas, vacant lots, roads, etc., on the map.
[0031] Furthermore, the concept ID may be associated with the SDGs development goals and the content related to supporting vulnerable groups. Users can upload their dream plan information or, at any appropriate time, select a goal for their project from the 17 SDGs development goals. The content of the selected SDGs development goal will be displayed as an SDGs pictogram in the dream plan proposal display unit 62, which will be described later. Users can upload their dream plan information or, at appropriate times, select options for accessibility features for vulnerable individuals (wheelchairs, strollers, accessible toilets, AEDs, etc.) to be included in their plan. The selected vulnerability features will be displayed as a pictogram in the dream plan proposal display unit 62, which will be described later. In the user's dream plan, it is possible to encourage users to become aware of the SDGs development goals and to raise awareness of addressing vulnerable groups as part of social infrastructure.
[0032] The evaluation DB12e stores the concept ID, the number of views, the poster ID, and the evaluations and comments on the dream plan images and videos in association with each other. It is a database that stores information regarding evaluations and comments on dream plan images and videos shared among multiple users through the urban development concept sharing platform 6. The number of views is the number of times the dream plan image and video have been viewed, and the poster ID is the ID of the user who posted the evaluation and comments on the dream plan image and video.
[0033] Figure 4 is a conceptual diagram showing the urban development concept proposal support method according to this embodiment, Figure 5 is a flowchart showing the processing procedure for urban development concept proposal support according to this embodiment, and Figure 6 is an example of a dream plan image. Assuming that the user communication terminal 5 is in a state of accessing and logging into the information processing device 1, the processing procedure related to urban development concept proposal support will be described below. Furthermore, in order to simplify the explanation, the processing without considering user type information and layer type will be described first.
[0034] The user communication terminal 5 transmits location information of the target to which the desired urban development should be represented to the information processing device 1 (step S11). For example, the user communication terminal 5 transmits location information detected by GPS. Alternatively, the user communication terminal 5 may accept input of an address or postal code and transmit the location information indicated by the received address or postal code to the information processing device 1. Furthermore, location information may be transmitted by taking a photograph using the user communication terminal 5 and transmitting the photograph data, which includes location information detected by GPS, to the information processing device 1. You may also transmit location information for points identified using a base map or similar.
[0035] The information processing device 1 acquires location information transmitted from the user communication terminal 5 (step S12). The information processing device 1 then transmits the location information to the map server 3 and requests source images associated with the location information, thereby acquiring source images from the map server 3 (step S13). The information processing device 1 also transmits the location information to the virtual globe server 4 and requests an overhead video associated with the location information, thereby acquiring an overhead video from the virtual globe server 4 (step S14).
[0036] Next, the user communication terminal 5 transmits information related to the user's dream of urban development, i.e., dream plan information, to the information processing device 1 (step S15), and the information processing device 1 acquires the dream plan information transmitted from the user communication terminal 5 (step S16). The information processing device 1 may also acquire document data obtained from the user as dream plan information through a dialogue-style exchange between the user and the document generation model 21 of the generation AI server 2, where the user expresses the desired urban development. Dream plan information is documented information expressed in narrative language, such as, "Wouldn't it be great to have a stylish cafe and barbecue garden in front of XXX Shrine, along the YYY River gorge!"
[0037] Next, the information processing device 1 sends the dream plan information to the generation AI server 2 or the document generation AI and requests the generation of a prompt (step S17). The generation AI server 2 inputs the dream plan information into the document generation model 21 to generate a prompt for generating a dream plan image and provides the generated prompt to the information processing device 1 (step S18). The information processing device 1 receives the prompt provided by the generation AI server 2. The generated prompts may be sentences or groups of words such as "along a stream, countryside, mountains, blue sky, cafe, restaurant, wooden, modern, sunny, realistic illustration."
[0038] The information processing device 1 may be configured to generate multiple prompts and send them to the user communication terminal 5, and to accept the selection and editing of the prompt to be used. The information processing device 1 then performs the following processing using the selected prompt and the selected and edited prompt.
[0039] Next, the information processing device 1 sends the prompt generated by the document generation model 21 to the generation AI server 2 or the image generation AI, requesting the generation of a dream plan image (step S19). The generation AI server 2 generates a dream plan image by inputting the prompt to the image generation model 22 and provides the generated dream plan image to the information processing device 1 (step S20). The information processing device 1 retrieves the dream plan image provided by the generation AI server 2. Figures 6A, 6B, and 6C are dream plan images actually generated by the image generation model 22.
[0040] The information processing device 1 may be configured to generate multiple dream plan images and send them to the user communication terminal 5, and to accept the selection of the dream plan image to be adopted. The information processing device 1 then performs the following processing using the selected dream plan image.
[0041] Next, the information processing device 1 sends the reference image and overhead video acquired in steps S13 and S14, along with the dream plan image generated by the image generation model 22, to the generation AI server 2 or the video generation AI, requesting the generation of a dream plan video (step S21). The generation AI server 2 generates a dream plan video by inputting the reference image, overhead video, and dream plan image into the video generation model 23, and provides the generated dream plan video to the information processing device 1 (step S22). The information processing device 1 acquires the dream plan video provided by the generation AI server 2.
[0042] The information processing device 1 may be configured to generate multiple dream plan videos and send them to the user communication terminal 5, and to accept the selection of the dream plan video to be adopted. The information processing device 1 then executes the following processes using the selected dream plan operation.
[0043] Next, the information processing device 1 associates the dream plan image and video on the base map 12b based on the location information acquired in step S12, and shares the dream plan image and video on the urban development concept sharing platform 6 so that they can be viewed (step S23). Furthermore, the information processing device 1 accepts requests to upload the dream plan images and dream plan videos to the urban development concept sharing platform 6, and uploads and shares the dream plan images and dream plan videos if it accepts the upload operation from the user. On a sharing platform, user-submitted images and videos can be displayed individually or on a map, and it would also be good to display them with various annotations. For example, it's a good idea to create and display content that's easy to share on other social media platforms, such as a curated collection of favorite images and videos suggested by other users, a collection of themes tagged with specific elements using AI, or a collection of trends tagged with popular elements using AI.
[0044] Figure 7 is a schematic diagram showing an example of the output screen of the urban development concept sharing platform 6. The urban development concept sharing platform 6 includes a base map display unit 61. The base map display unit 61 displays an image of the base map 12b and marker images 61a that indicate the locations where dream plan images and dream plan videos have been uploaded. Mark images 61a include, for example, a circle image drawn with a dashed line. Mark images 61a indicate an area with a diameter of approximately 100m and indicate the existence of dream plan images and videos associated with locations within that area. The reason why the area indicated by marker images 61a is so broad is to avoid problems with information related to specific locations from the perspective of privacy and personal information. In addition, the map image enclosed by marker images 61a is blurred.
[0045] When the cursor hovers over the mark image 61a, the Dream Plan Proposal Display Unit 62 pops up. The Dream Plan Proposal Display Unit 62 displays the Dream Plan image, Dream Plan video, its thumbnail document, handle name, proposal name, location, and posting date included in the area indicated by the mark image 61a. The Dream Plan Proposal Display Unit 62 also includes ratings and comments for the Dream Plan image and video.
[0046] Users of the urban development concept proposal support system can view dream plan images and dream plan videos shared through the urban development concept sharing platform 6.
[0047] Furthermore, as shown in Figure 5, users of the urban development concept proposal support system can post evaluations and comments on dream plan images and dream plan videos to the urban development concept sharing platform 6 using the user communication terminal 5 (step S24). The information processing device 1 receives evaluations and comments on specific dream plan images and dream plan videos posted and transmitted by other users via the user communication terminal 5, and shares the received evaluations and comments so that they can be viewed (step S25). As shown in Figure 7, the received evaluations and comments are displayed on the dream plan proposal display unit 62 of the urban development concept sharing platform 6 and shared among multiple users.
[0048] Up to this point, we have explained the processing without considering the layer type based on user type information. However, it is also possible to configure the system to distinguish layers based on user type and to select and share / view dream plan images and dream plan videos based on layer type.
[0049] When the information processing device 1 generates a dream plan image and a dream plan video, it associates the layer type with them and stores them in the storage unit 12. The layer type is determined by the user type information of the user who provided the dream plan information. If the user type is an individual, the layer type indicating the individual layer is associated with the dream plan image and video. If the user type is a government agency, the layer type indicating the government layer is associated, and if the user type is a private business, the layer type indicating the private business layer is associated.
[0050] Furthermore, the information processing device 1 stores a base map 12b corresponding to the layer type in the storage unit 12. Based on the user type information of the logged-in user, the information processing device 1 identifies the layer type to be shared. Then, the information processing device 1 displays a dream plan image and video representing the identified layer type on the base map 12b of that layer type.
[0051] Figure 8 is a conceptual diagram showing the selection method for dream plan images and dream plan videos based on layer type. When the logged-in user is an individual, the dream plan images and videos of the individual layer are displayed on the general map and shared among individual users through the urban development concept sharing platform 6. If the logged-in user is a government agency, the dream plan images and videos of the government layer will be displayed on the urban planning map and shared among government agencies through the urban development concept sharing platform 6. The dream plan images and videos of the government layer may also be shared with specific private businesses. Furthermore, government agency users may be configured to view dream plan images and videos of their personal layer by switching layers. Additionally, government agency users may be configured to view dream plan images and videos of specific private businesses by switching layers. Furthermore, administrative agencies may set themes for solving local issues and construct solution layers aimed at resolving them. For example, as part of disaster prevention measures and response, photos of disaster sites and evacuation routes could be incorporated, and then the appropriate on-site response and evacuation routes could be visualized using AI-generated images and provided to affected residents. Alternatively, photos of the interior of evacuation shelters, storage warehouses, and the congestion levels within evacuation shelters could be incorporated, and then the placement of temporary beds for evacuees could be planned using AI-generated images. The consumption / shortage status of storage warehouses and the congestion levels within evacuation shelters could also be visualized and controlled using AI-generated images, and this information could be provided to residents. In addition, as part of measures against vacant houses, photos of the current state of vacant houses could be incorporated, and the deterioration over time, improvements after countermeasures, and the owner's post-disaster response plans such as clearing the land, remodeling, renovation, or rebuilding could be visualized using AI-generated images and used as support measures for countermeasures. If the logged-in user is a business operator, the residential construction map will display dream plan images and videos from the private business operator layer, and these will be shared among individual users, specific private businesses, and specific government agencies through the urban development concept sharing platform 6. Businesses may establish a layer that sets themes related to sales and promotional activities for customers and provides sales promotion and customer service that elicits customer interest, requests, and needs. For example, a home builder could propose building plans for potential clients' plots of land to drive business negotiations, or they could solicit dream building plans from the general public to acquire new customers. For renovation companies, it would be beneficial to use AI-generated visualizations to show the process of renovations, remodeling, and rebuilding, and to use this information for customer proposals, or to use it to receive renovation offers from general customers. For game software manufacturers, it would be good to use it in role-playing city-building games that allow players to virtually experience urban planning and development on actual or fictional maps. Alternatively, they could offer games or services / products that feature characters and actors from anime, movies, manga, dramas, etc. (limited to those for which usage permission has been granted), by setting up cityscapes or specific locations and incorporating them into images and videos. Travel agencies and regional tourism promotion organizations could create promotional videos featuring local mascots or tourism ambassadors (only those for which usage permission has been granted) in images and videos set in tourist destinations, resorts, or promotional areas, or in streetscapes or specific locations. They could also offer services and products such as pilgrimage tours.
[0052] The selective sharing method of dream plan images and dream plan videos based on the layer type described above is just one example. Dream plan images and dream plan videos shared through the urban development concept sharing platform 6 may be configured to be arbitrarily selected using user type information and layer type. Furthermore, the accuracy of the location information indicated by the shared dream plan images and dream plan videos may be varied depending on the layer type. For administrative layers, where accurate location accuracy is required, it is advisable to associate the location information with the dream plan images and dream plan videos and display them on the base map 12b with an accuracy of a few meters. For personal layers, to avoid issues of personal information and privacy, it is advisable to associate the location information with the dream plan images and dream plan videos and display them on the base map 12b with an accuracy of approximately 100 meters. At a minimum, the location accuracy of the administrative layer's location information should be set higher than that of the personal layer's location information.
[0053] As described above, the urban development concept proposal support system, urban development concept proposal support method, information processing device 1, and computer program 12a according to this embodiment can support the proposal of urban development concepts that visualize the urban development dreams of ordinary citizens.
[0054] For ordinary citizens and other users, it is difficult to put their vision of urban development into words, and it is not easy to visualize it through illustrations, sketches, etc. While using generative AI is an option, generating prompts to visualize one's vision of urban development requires skill and practice.
[0055] According to this embodiment, since the document generation model 21 is used to generate prompts for image generation, even if the dream plan described by the user is vague, it is possible to accurately materialize that dream plan as a dream plan image. Furthermore, by acquiring dream plan information through dialogue with the user, the information processing device 1 can actively collect necessary information and materialize the urban development envisioned by the user as a dream plan image.
[0056] According to the information processing device 1 of this embodiment, based on the dream plan information provided by the user, a prompt for image generation can be generated, and a dream plan image that embodies the dream can be generated. In other words, the dream city development and city development concept envisioned by the user can be materialized as an image. Specifically, this disclosure will allow users to obtain the following benefits: Residents can propose urban development plans that visualize their envisioned and planned streetscapes and other images. When the government aims to formulate a community development plan with resident participation, or when it is raising awareness of local issues such as disaster response, it can provide information and explanations that visualize and share images with residents. Businesses can propose visual representations of their locally-rooted products and services as part of their economic activities. Educational institutions can implement proactive, active learning through video, including community development education, SDGs education, and inquiry-based learning through school education and local organizations.
[0057] Furthermore, a dream plan video can be generated based on a reference image (e.g., a captured image) representing the current urban development at a location specified by the user, along with a dream plan image and an overhead view video. The dream plan video is a video representation of the reference image, overhead view image, and dream plan image representing the current urban development at the location, allowing users to easily visualize their dream urban development or urban development concept in video form.
[0058] Furthermore, since the system generates dream plan videos using reference images related to location information obtained from map server 3 and overhead videos obtained from virtual globe server 4, it is possible to express more realistic urban development concepts as dream plan videos.
[0059] Furthermore, by enabling multiple users to share dream plan images and videos through the concept sharing platform 6, and to evaluate and comment on each other's plans, it is possible to stimulate the generation of dream plan images and videos and urban development concepts. Furthermore, users can also post to commonly used online information sharing platforms (SNS: Social Networking Service). Examples include social networking services such as LINE, YouTube (registered trademark), Instagram, X (formerly Twitter), Facebook, TikTok, Pinterest, and note. Dream plan images and videos are accompanied by thumbnails, and when users post them on social media, official hashtags (#yumecan, #yumemachicanvas, #generatingAIcommunitydevelopment, #communitydevelopmentDX, etc.) are automatically added. This official hashtag allows residents across the country and within local governments to connect with posters of dream plan images and videos from other local governments, fostering the formation of user communities that transcend regions and local governments, and further revitalizing community-based urban development initiatives. Furthermore, government agencies, businesses, and educational institutions can set themes periodically, solicit dream plan images and videos from individual users, and hold contests to award prizes. This allows the system to be used as a proposal support system with area marketing and research functions, enabling community-participatory urban development concepts and the creation of products and services that utilize the dream plans and ideas desired by individual users.
[0060] In this embodiment, we have described a configuration in which the source image is input to the video generation model 23, but it is also possible to configure it so that the source image is input to the image generation model 22 along with a prompt. This allows for obtaining a more realistic dream plan image.
[0061] Furthermore, although this embodiment describes an example of acquiring an overhead video from the virtual globe server 4, it is also possible to configure the system to acquire an overhead image from the virtual globe server 4 and input that overhead image as a source image to the video generation model 23.
[0062] Furthermore, although this embodiment describes an example in which the document generation model 21, image generation model 22, and video generation model 23 are configured separately, it is also possible to configure the system to generate dream plan images and dream plan videos from dream plan information using a multimodal generation model capable of processing text, images, and videos. In other words, when dream plan information acquired from the user communication terminal 5, material images acquired from the map server 3, and overhead images and / or overhead videos acquired from the virtual globe server 4 are input, the information processing device 1 may be configured using a generation model that generates dream plan images and / or dream plan videos that reflect various types of information.
[0063] The following describes the problem to be solved, the means for solving the problem, the effects of this disclosure, and a conceptual diagram of the embodiment relating to this disclosure. The problem and effects described below are not intended to reveal the technical scope of the disclosure in the embodiment. Furthermore, at least some of the configurations described below may be combined as desired.
[0064] [Problems to be solved] Traditional methods for citizens to participate in urban planning and local community development (facilities, streetscapes, etc.) involve submitting proposals to the government through public comment systems, primarily through written text, supplemented by hand-drawn sketches, to convey their ideas and concepts. However, written text alone is insufficient to fully convey citizens' visions, and for ordinary citizens without specialized drawing or design skills, it is difficult to effectively communicate their ideas and visions visually. Furthermore, it is difficult for those receiving the proposals to accurately grasp the proposer's true intentions and purpose. In citizen-participatory and community-led urban planning and local community development, traditional methods have failed to effectively and concretely communicate citizens' ideas and visions, resulting in a loss of opportunities for proactive participation, proposals, and improvements toward better urban planning and community development. Furthermore, methods for citizens to mutually share and evaluate ideas and visions have been limited. There is a fundamental problem with how citizens can effectively communicate their ideas and visions, and this needs to be resolved.
[0065] [Means for solving the problem] (Note 1) We obtain request information that expresses the desired urban development, We obtain location information related to the desired urban development project. Obtain source images corresponding to the acquired location information, By inputting the acquired request information and the source images into the generation model, a town development image or town development video that reflects the content of the request information is generated. Methods for supporting proposals for urban development plans.
[0066] (Note 2) The aforementioned generation model is When a prompt based on the aforementioned request information is entered, an image generation model generates a city planning image that reflects the content of the aforementioned request information, When the aforementioned urban development image and the aforementioned material image are input, a video generation model generates an urban development video that reflects the content of the urban development image and the aforementioned material image. Includes, By inputting the prompt based on the acquired request information into the image generation model, the urban development image is generated. The urban development video is generated by inputting the urban development image generated by the image generation model and the source image into the video generation model. The method for supporting urban development plan proposals as described in Appendix 1.
[0067] (Note 3) The aforementioned generation model is When a prompt based on the aforementioned request information and the aforementioned source image are input, an image generation model generates a city planning image that reflects the content of the aforementioned request information. When the aforementioned urban development image is input, a video generation model generates an urban development video that reflects the content of the aforementioned urban development image. Includes, By inputting the prompt based on the acquired request information and the source image into the image generation model, the urban development image is generated. The urban development video is generated by inputting the urban development images generated by the image generation model into the video generation model. The method for supporting urban development plan proposals as described in Appendix 1.
[0068] The aforementioned image generation model is Furthermore, if generated urban planning images and source images are input, a new urban planning image will be generated. The method for supporting urban development plan proposals as described in Appendix 2 or Appendix 3.
[0069] (Note 4) The aforementioned generation model is When the aforementioned request information is input, the document generation model includes a document generation model that generates a prompt related to the generation of urban development images using the image generation model, By inputting the acquired request information into the document generation model, a prompt based on the request information is generated. The method for supporting urban development plan proposals as described in Appendix 2 or Appendix 3.
[0070] (Note 5) The aforementioned request information is, This includes text data obtained through a dialogue-based exchange between a user who expresses their desired urban planning and the document generation model. The method for supporting urban development plan proposals is described in Appendix 4.
[0071] (Note 6) From a map server that provides map information including images associated with location information, images associated with the acquired location information are obtained as the source images. A method for supporting urban development plan proposals described in any one of the appendices 1 through 3.
[0072] (Note 7) From a virtual globe server that provides aerial images of various locations, we obtain an aerial image or aerial video of the location indicated by the acquired location information. By inputting the acquired overhead image or overhead video and the urban development image generated by the image generation model into the video generation model, an urban development video is generated. The method for supporting urban development plan proposals as described in Appendix 2 or Appendix 3.
[0073] (Note 8) The aforementioned urban development images and videos are output to a platform for sharing urban development concepts related to proposals for urban development images, linked to a base map based on acquired location information. A method for supporting urban development plan proposals described in any one of the appendices 1 through 3.
[0074] (Note 9) Multiple images and videos of urban development projects, generated from user requests expressing urban development, are output in a way that allows them to be viewed by each other. The method for supporting urban development plan proposals is described in Appendix 8.
[0075] (Note 10) We accept comments or evaluations regarding the aforementioned urban development images and videos. In connection with the aforementioned urban development images and videos, comments or evaluations of the aforementioned urban development images and videos are output to the urban development concept sharing platform. The method for supporting urban development plan proposals, as described in Appendix 9.
[0076] (Note 11) The system includes a user database that stores user identifiers indicating multiple users in association with user type information indicating the type of user, The generated urban development images and urban development videos are associated with the user type information of the user who provided the request information for generating the urban development images and urban development videos. Output the urban development images and urban development videos selected according to the user type information. The method for supporting urban development plan proposals, as described in Appendix 9.
[0077] (Note 12) An information processing device comprising a communication unit and a control unit, The control unit, We obtain request information that expresses the desired urban development, We obtain location information related to the desired urban development project. Obtain source images corresponding to the acquired location information, By inputting the acquired request information and the source images into the generation model, a town development image or town development video that reflects the content of the request information is generated. Information processing device.
[0078] (Note 13) A computer equipped with a communications unit and a control unit, We obtain request information that expresses the desired urban development, We obtain location information related to the desired urban development project. Obtain source images corresponding to the acquired location information, By inputting the acquired request information and the source images into the generation model, a town development image or town development video that reflects the content of the request information is generated. A computer program designed to execute a process.
[0079] The urban development plan proposal support method described in any one of the appendices 1 to 11 is preferably executed by an information processing device.
[0080] In any of the above appendices, it is preferable that by using public space photographs as source images, urban development proposals that reflect the user's dream plan can be synthesized and visualized within existing public spaces, streetscapes, and cities.
[0081] In any of the above appendices, it is preferable that the urban development image or urban development video be activated and displayed as AR content on the user's device using AR technology, either from a map, or triggered by a specific photograph or image, or actual public space or location information.
[0082] In any of the above appendices, it is preferable to display community development images or videos on a website's community space or shared gallery, or on social media, with annotations such as hashtags.
[0083] In any of the above appendices, it is preferable that the information processing device analyzes the user's request history, stores the generated urban development images and generation history, and has editing functions that allow for repeated editing, correction, partial modification, expansion, deletion, and synthesis, replacement, and addition with source images, and that the generation model combines reproducibility and variability so that the user's urban development request information is optimally selected, modified, and generated as a dream plan, either as it is or by expanding on the image.
[0084] In any of the above appendices, the information processing device possesses remarkable accessibility, allowing anyone, from children to the elderly, to visualize and share their dream plans using a cutting-edge generative AI model via their smartphone or other device, without requiring any special skills such as AI skills from the user. Furthermore, it is preferable that the device possesses usability that allows anyone, even first-time users, to easily, quickly, efficiently, comfortably visualize and share their ideas as they envision them, thereby increasing satisfaction.
[0085] In any of the above appendices, it is preferable that the information processing device uses a generative model with search extension capabilities to perform searches of a large-scale language model while referencing context specific to the urban development field, so that the prompts generated through interactive exchanges converge on proposals in the urban development field that meet the user's needs. In the course of interacting with the user, prompts that are suitable for the request information and highly reliable are generated.
[0086] In any of the above appendices, it is preferable to utilize RAG (Search Enhancement Generation), which pre-constructs and stores context specific to the urban development field (categories such as buildings, spaces, and sections / elements of requests related to the city / place names, characteristics, and facility names of the usage area / urban planning of the user administration, etc.) as a local database and references it, and to have the LLM (Large-Scale Language Model) perform an accurate vector search in response to user questions, thereby generating answers that are highly reliable and suitable for the user's urban development concept. In this way, by adapting the system to the specific requirements and regional characteristics of the urban development field, the accuracy of the image generation prompt can be improved, making it possible to bring the generated images and videos closer to the user's urban development vision.
[0087] It is preferable that the information processing device allows users to select from approximately five generation modes (with settings for the number of conversations, search depth, range and number of selection items, parameters, etc., set for each stage) according to their initiative, preferences, and level of proficiency. For example, the five stages could be (1) for first-time users, (2) for beginners, (3) for intermediate users, (4) for upper-intermediate users, and (5) for advanced users. This would alleviate the common problems that arise when using generation AI, such as insufficient information for beginners and those who find it unsatisfying for upper-intermediate users, and would increase the likelihood of continued use of generation AI.
[0088] In any of the above appendices, this proposed support method is a variable and interactive method that allows users to enjoy results and benefits tailored to various purposes, not limited to urban development, depending on the type of user utilizing it, and preferably the types of users are residents, government agencies, businesses, educational institutions, etc.
[0089] In any of the above appendices, it is preferable that resident users can propose to the government, local communities, etc., urban development plans that visualize their envisioned streetscapes and other images.
[0090] In any one of the above appendices, it is preferable that administrative users can solicit proposals for community-participatory urban development planning, explain urban planning and regional issue resolution themes to residents, and utilize them for support services for administrative users, etc.
[0091] In any one of the above appendices, it is preferable that users such as real estate businesses can provide customers with building-related products such as houses, shops, accommodation, medical facilities, cultural facilities, commercial facilities, tourist facilities, leisure facilities, and parking facilities, or renovation-related services such as remodeling and refurbishment, or utilize them for similar marketing and sales promotion, and support services for users such as real estate businesses.
[0092] In any one of the above appendices, it is preferable that users such as game businesses may provide customers with game software, such as role-playing city-building games that allow them to virtually experience urban planning and other city development on actual or fictional maps, or games and services such as fan activities that allow characters and performers from anime, movies, manga, dramas, etc. (limited to those for which usage licenses have been granted) to appear in images and videos by setting up cityscapes or specific locations, or utilize similar marketing, sales promotion, and support services for users of game businesses.
[0093] In any of the above appendices, it is preferable that users such as travel agencies and regional tourism promotion organizations may use travel products to their customers, such as promotional videos that feature local mascots or tourism ambassadors (limited to those for which usage permission has been granted), in images and videos set in the streetscapes or specific locations within tourist destinations, resorts, or promotional areas, or to provide planned services and products such as pilgrimages to sacred sites, as well as for similar marketing and sales promotion, and support services for users such as travel agencies and regional tourism promotion organizations.
[0094] In any of the above supplementary provisions, it is preferable that educational institution users can utilize the system in practical educational settings such as community-based urban development education, SDGs education, and inquiry-based learning, as well as in support services for educational institution users.
[0095] In any of the above appendices, it is preferable that the platform for obtaining generated images and videos, as well as comments and evaluations, produced from the proposed support method, etc., can form a specific economic sphere with the users and customers who utilize it (for example, a matching service between users and businesses, a contest sponsored by businesses, advertising on the platform, credit placement on images and videos, etc.).
[0096] In any of the above appendices, it is preferable that the images and videos generated by each of the above models are based on existing public spaces and cityscapes, and represent urban development proposals that reflect the user's dream plan as 2D spatial images, and are provided so that they can be viewed, shared, and evaluated on an urban development concept sharing platform.
[0097] In any one of the above appendices, Based on the aforementioned request information, data related to the request information is read from the RAG database. The aforementioned generation model is By inputting the acquired request information, the source images, and the data read from the RAG database into the generation model, a town development image or town development video that reflects the content of the request information is generated. It is preferable.
[0098] In the above note, it is preferable that the RAG database includes information about the type of user (individual, government, business, educational institution, etc.), information about multiple locations, and information about buildings.
[0099] In any of the above appendices, it is preferable to store multiple generated urban development images and provide the selected urban development image in response to user selection and requests.
[0100] In the above appendix, it is preferable to accept editing requests for the selected urban development image, edit the urban development image according to the accepted editing content, and store the edited urban development image.
[0101] In the above note, it is preferable that the editing includes the synthesis of multiple images of urban development.
[0102] In the above appendix, it is preferable to accept the selection of one or more urban development images from a plurality of stored urban development images (which may include edited urban development images), and to generate a new urban development image by inputting the request information and the selected urban development images into the generation model.
[0103] [Effects of this disclosure] This disclosure allows any citizen user to visualize their own ideas and visions in public spaces. Even ordinary citizens who possess no specialized knowledge (such as urban planning), specialized skills (such as drawing and design), or skills in utilizing advanced technologies (such as the ability to highly manipulate generative AI) can easily visualize their dream city development in just a few seconds, regardless of age, from children to the elderly, and project it into public spaces. These city development concepts can be proposed on shared platforms and evaluated by others, and a synergistic effect with the ideas and visions of others can be expected. This disclosure will not only enable citizen users to propose ideas on their own initiative, but will also allow the government to solicit proposals by outlining urban planning and specific issues. As such, this disclosure will contribute to the formation of a citizen-participatory, self-governing society, promoting the "democratization of urban development" and is expected to have the effect of increasing the potential for application in solving local issues, improving overall urban development literacy, and human resource development. Furthermore, it is expected to have various uses and effects depending on the user, such as stimulating economic activity when businesses utilize it for sales and promotional activities, or revitalizing active learning when educational institutions use it for experiential and inquiry-based learning in community development.
[0104] [Conceptual diagram of an embodiment relating to this disclosure]
[0105] Figures 9, 10, and 11 are the first to third conceptual diagrams showing the functional configuration of the urban development concept proposal support system. The first to third conceptual diagrams show the functions realized by this embodiment and the embodiments described in the appendix above.
[0106] The system's overview is as follows: We will provide a system and platform that allows any resident to easily propose their dreams and ideas for community development. By utilizing cutting-edge generative AI technology (text, images, and videos), we can quickly transform residents' wishful thinking into visuals and share them on a platform, providing a collaborative space for evaluation and exchange of opinions between the government and residents. By providing this system, the government can share and consider residents' voices as more concrete visual images to address local issues and develop urban planning plans, thus serving as a digital partner to promote resident-participatory urban development. As a highly public project that contributes to the urban development policies of local governments, it will play a role as a "digital public good" that contributes to regional revitalization throughout Japan.
[0107] The current challenges are as follows: (1) Residents' opinions are not easily reflected: The process of gathering public opinion is rigid and complex, and the barriers of outdated existing systems lead to the entrenchment of participants and the hollowing out of consensus-building, resulting in feelings of distance, alienation, and indifference. (2) Specialized knowledge is required: Current proposal methods, such as public comments and proposals, require skills and expertise to clearly communicate proposals, making it difficult for ordinary citizens to participate. (3) Gathering opinions takes time and costs money: Current methods such as surveys, information sessions, and town hall meetings place a heavy burden on the government, yet they standardize the way opinions are collected. As a result, the number of silent majority and indifferent people is increasing, creating a barrier between individuals and local communities, where "individuals cannot influence the government's urban development efforts."
[0108] The direction in which the issues to be resolved should proceed is as follows: (1) Providing a system that allows residents to easily make suggestions. (2) Promoting dialogue with the government through the visualization and video production of proposals (3) Revitalizing local communities and expanding the base of participants through exchange of opinions and evaluations among residents.
[0109] The design concept for the support method for proposing urban development plans is as follows: (1) Easy to use: User interface (UI) that is easy for anyone to use, from children to the elderly, and provides a comfortable and user-friendly experience (UX). (2) Speed: We utilize cutting-edge AI to quickly visualize proposals (images: approximately 15 seconds, videos: approximately 60 seconds). (3) Sharing: Provide a collaborative space on the platform where residents can share their proposals and where the administration and residents can evaluate and exchange opinions.
[0110] The information processing device 1 acquires request information, location information, and source images through interaction with the user communication terminal 5 by utilizing the generation AI server 2. The information processing device 1 may be configured to accept basic information settings in order to efficiently acquire request information. The basic information includes the following setting information. • Creative autonomy: Leave it all to me ← Create roughly ← Balance well → Create with attention to detail → Create with attention to every detail / Eliminate negative elements • Availability of provided material information: Location photos (site photos), hand-drawn sketches, image renderings (example photos, etc.) / annotations • Reproducibility parameters (seed value, etc.): Generated image is OK • Complete → Partial correction (selection and resetting of correction information) → Start all over (regeneration)
[0111] The information processing device 1 reads request information and data related to urban development from the RAG database based on at least the acquired request information. The RAG database stores context specific to the urban development field. The RAG database may store the context in a text-searchable format, a vector-searchable format, or it may store searchable data obtained by crawling relevant web pages estimated by a generative AI.
[0112] The following information is stored within the context specific to the urban development field and is provided as options in a dialogue format when obtaining user requests. • Who are you? (Subject information) = Individual / Government / Business / Educational institution What (1)? (Target category information) = Building facilities (houses, etc.) / Public spaces (parks, etc.) / Road sections (sidewalks, etc.). • What (2)? (Required facility information) = Required functions, elements, and facilities (e.g., a terrace with a good view, benches for relaxation, easy-to-walk paths). What to do? (Outcome information) = Create something new (new construction) / Rebuild (rebuilding / renovation) / Make adjustments (improvement / remodeling), etc. • Who are the users? (Assumed persona) = Age group, gender, family structure, purpose of use, mode of transportation (pedestrian, stroller), etc. • What do you want to do? (Request information) = The intentions of the person in charge / Atmosphere and scenery of the object, etc. / Comfort and user experience of the user, etc. How to describe the style? = Type of art style (realistic photography, illustration, animation, etc.) / Atmosphere (warm, etc.) / Artistic style (Japanese, Western, etc.). What is the composition and angle? = Point of view (overhead, exterior, interior, panoramic view, etc.) / Whether or not people are included, character settings, number of people, and appearance.
[0113] Each generation model of the generation AI server 2 generates urban development images and urban development videos based on the above-mentioned request information, location information, source images, and data read from the RAG database.
[0114] The generated prompts, urban development images, and urban development plans will be used as deliverable utilization data related to the urban development concept and dream plan. The deliverable utilization data will include the following information: Location information, base map, layer type data, mapping data Generation prompt, image / video / audio data, thumbnail data Image synthesis data / annotation with public spaces, map data search Comment rating information, voting / survey results, viewing / posting history
[0115] The data on the utilization of deliverables generated by users will be utilized by local communities, individuals, government agencies, businesses, corporations, and educational institutions through a shared platform for urban development concepts. Furthermore, it is desirable that the system operator make the data publicly available to society as a whole, including significant aggregated data, analytical reports, and examples of its use.
[0116] The information processing device 1 has administrator functions that control the following information processing. • Frontend: User access, My Page, Chatbot Backend: User identification and management, usage control, security, analytics. • Layers: Integrated management / Individual (local area) / Government (urban planning / local issues) Businesses (specific projects, urban development games) and educational institutions (inquiry-based learning) Base map, specific project map, generated image, video mapping • Platform: Templates, content creation and management, evaluation and sharing • API integration: Generation AI (parameter adjustments to improve the accuracy of image and video reproduction, etc.) Chatbot platform and SNS integration specializing in urban development [Explanation of symbols]
[0117] 1: Information Processing Device 2: Generation AI Server 3: Map Server 4: Virtual Globe Server 5: User communication terminal 6: Concept Sharing Platform 10: Recording media 11: Processing Section 12: Storage section 12a: Computer program 12b: Base map 12c: User DB 12d: Concept DB 12e: Evaluation DB 13: Communications Department
Claims
1. We obtain request information that expresses the desired urban development, We obtain location information related to the desired urban development project. Obtain source images corresponding to the acquired location information, By inputting the acquired request information and the source images into the generation model, a town development image or town development video that reflects the content of the request information is generated. Methods for supporting proposals for urban development plans.
2. The aforementioned generation model is When a prompt based on the aforementioned request information is entered, an image generation model generates a city planning image that reflects the content of the aforementioned request information, When the aforementioned urban development image and the aforementioned material image are input, a video generation model generates an urban development video that reflects the content of the urban development image and the aforementioned material image. Includes, By inputting the prompt based on the acquired request information into the image generation model, the urban development image is generated. The urban development video is generated by inputting the urban development image generated by the image generation model and the source image into the video generation model. The method for supporting the proposal of urban development plans as described in claim 1.
3. The aforementioned generation model is When a prompt based on the aforementioned request information and the aforementioned source image are input, an image generation model generates a city planning image that reflects the content of the aforementioned request information. When the aforementioned urban development image is input, a video generation model generates an urban development video that reflects the content of the aforementioned urban development image. Includes, By inputting the prompt based on the acquired request information and the source image into the image generation model, the urban development image is generated. The urban development video is generated by inputting the urban development images generated by the image generation model into the video generation model. The method for supporting the proposal of urban development plans as described in claim 1.
4. The aforementioned generation model is When the aforementioned request information is input, the document generation model includes a document generation model that generates a prompt related to the generation of urban development images using the image generation model, By inputting the acquired request information into the document generation model, a prompt based on the request information is generated. A method for supporting the proposal of a town development plan according to claim 2 or claim 3.
5. The aforementioned request information is, This includes text data obtained through a dialogue-based exchange between a user who expresses their desired urban planning and the document generation model. The method for supporting the proposal of urban development plans as described in claim 4.
6. From a map server that provides map information including images associated with location information, images associated with the acquired location information are obtained as the source images. A method for supporting the proposal of a town development plan according to any one of claims 1 to 3.
7. From a virtual globe server that provides aerial images of various locations, we obtain an aerial image or aerial video of the location indicated by the acquired location information. By inputting the acquired overhead image or overhead video and the urban development image generated by the image generation model into the video generation model, an urban development video is generated. A method for supporting the proposal of a town development plan according to claim 2 or claim 3.
8. The aforementioned urban development images and videos are output to a platform for sharing urban development concepts related to proposals for urban development images, linked to a base map based on acquired location information. A method for supporting the proposal of a town development plan according to any one of claims 1 to 3.
9. Multiple images and videos of urban development projects, generated from user requests expressing urban development, are output in a way that allows them to be viewed by each other. The method for supporting the proposal of a town development plan as described in claim 8.
10. We accept comments or evaluations regarding the aforementioned urban development images and videos. In connection with the aforementioned urban development images and videos, comments or evaluations of the aforementioned urban development images and videos are output to the urban development concept sharing platform. The method for supporting the proposal of urban development plans as described in claim 9.
11. The system includes a user database that stores user identifiers indicating multiple users in association with user type information indicating the type of user, The generated urban development images and urban development videos are associated with the user type information of the user who provided the request information for generating the urban development images and urban development videos. Output the urban development images and urban development videos selected according to the user type information. The method for supporting the proposal of urban development plans as described in claim 9.
12. An information processing device comprising a communication unit and a control unit, The control unit, We obtain request information that expresses the desired urban development, We obtain location information related to the desired urban development project. Obtain source images corresponding to the acquired location information, By inputting the acquired request information and the source images into the generation model, a town development image or town development video that reflects the content of the request information is generated. Information processing device.
13. A computer equipped with a communications unit and a control unit, We obtain request information that expresses the desired urban development, We obtain location information related to the desired urban development project. Obtain source images corresponding to the acquired location information, By inputting the acquired request information and the source images into the generation model, a town development image or town development video that reflects the content of the request information is generated. A computer program designed to execute a process.
Citation Information
Patent Citations
JP2024018823A